Historical data migration method, device, equipment and medium for oil and gas field database

By digitizing and extracting features from paper logging documents in oil and gas field databases, combined with adaptive graphics technology and machine learning models, the problems of insufficient digitization depth of paper logging documents and lack of multi-source data collaboration were solved, efficient data migration and cross-well area data association were achieved, and the searchability and credibility of the data were improved.

CN120216461BActive Publication Date: 2025-09-30DESHI ENERGY TECH GRP CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510694840.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-09-30
Estimated Expiration
2045-05-28

AI Technical Summary

Technical Problem

Existing technologies in oil and gas field databases have insufficient digitization depth for paper logging documents, lack of multi-source data collaboration, low accuracy in graphic feature recognition, and weak logic for data classification and migration, resulting in the inability to efficiently retrieve and utilize data and to associate data across well areas.

Method used

By digitizing paper logging documents to generate RGB images, extracting logging text and graphic features, and building a hierarchical storage structure based on oil and gas exploration business needs, adaptive graphic feature extraction technology is used to achieve collaborative integration and automated migration of multi-source data. The migration volume is dynamically adjusted to adapt to database performance, and machine learning models are used to optimize migration strategies.

Benefits of technology

It improves the retrievability and utilization of paper data, solves the island problem of multi-source data, improves the accuracy of logging curve symbol recognition, optimizes the efficiency of data migration process, ensures data integrity and credibility, and adapts to the multi-scenario compatibility of different oil and gas field projects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216461B_ABST
    Figure CN120216461B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of oil and gas field databases, and discloses a method, device, equipment and medium for migrating historical data of an oil and gas field database, comprising: digitizing paper logging documents in the oil and gas field database to generate RGB images; synchronously accessing semi-structured electronic documents in the oil and gas field database; extracting logging text features in the RGB images and electronic documents; extracting logging graphic features in the RGB images and electronic documents for logging curve symbols; determining a first data volume of logging text features and logging graphic features; determining a single migration data volume; judging whether the first data volume exceeds the single migration data volume; if so, dividing the first data volume into multiple migration data volumes; determining a second data volume of a new oil and gas field database; comparing the first data volume and the second data volume to determine whether they are consistent; if they are consistent, migrating the logging text features and the logging graphic features to a pre-constructed new oil and gas field database according to a preset classification of the logging data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of oil and gas field databases, and in particular to a method, device, equipment, and medium for migrating historical data of an oil and gas field database. Background Art

[0002] In the oil and gas field exploration and development sector, well logging data, the core basis for reservoir evaluation and production decision-making, is typically stored long-term in traditional databases in the form of paper documents and semi-structured electronic documents (such as PDFs and Excel spreadsheets). With the demand for digital upgrades, the migration of historical data to modern databases requires structured storage and the integration of multimodal features.

[0003] Current technical means mainly include: 1) scanning paper logging documents to generate electronic images; 2) directly importing electronic documents through the database interface; 3) extracting partial text data based on keyword matching; 4) manually marking the symbol coordinates of logging curves to realize graphic data migration.

[0004] However, the existing technology has significant defects:

[0005] Insufficient digitization of paper documents: Traditional scanning only generates low-structure RGB images and lacks intelligent recognition and semantic analysis of logging text features (such as well numbers and formation parameters). This prevents data from being efficiently retrieved and utilized in new databases.

[0006] Lack of multi-source data collaboration: The processing processes for paper and electronic documents are independent of each other. Non-standard logging parameters (such as annotation text and segmented data) in semi-structured electronic documents lack unified feature extraction rules, forming information islands.

[0007] Low accuracy in graphic feature recognition: The extraction of logging curve symbols relies on manual interpretation or fixed template matching, which cannot adapt to layout differences and symbol deformation (such as curve breakpoints and coordinate axis offsets) between different documents, resulting in distortion of the migrated data.

[0008] Weak data classification and migration logic: A well logging data classification system that matches oil and gas exploration business has not been established, and the migrated data lacks a hierarchical storage structure, which restricts in-depth analysis and cross-well area data correlation applications. Summary of the Invention

[0009] One or more embodiments of this specification provide a method, apparatus, device, and medium for migrating historical data of an oil and gas field database, which are used to solve the technical problems raised in the background art.

[0010] One or more embodiments of this specification adopt the following technical solutions:

[0011] One or more embodiments of this specification provide a method for migrating historical data from an oil and gas field database, the method comprising:

[0012] Digitize paper logging documents in oil and gas field databases to generate RGB images;

[0013] Synchronously accessing semi-structured electronic documents in the oil and gas field database;

[0014] Extracting well logging text features from the RGB image and the electronic document;

[0015] For the well logging curve symbols, extract the well logging graphic features in the RGB image and the electronic document;

[0016] determining a first data volume of the well logging text feature and the well logging graphic feature;

[0017] Determining a single migration data volume based on a single migration reception volume of the new oil and gas field database;

[0018] Determining whether the first data volume exceeds the single migration data volume;

[0019] If so, dividing the first data volume into multiple migration data volumes according to the single migration data volume;

[0020] determining a second data volume of the new oil and gas field database;

[0021] comparing whether the first data volume is consistent with the second data volume;

[0022] If they are consistent, the well logging text features and the well logging graphic features are migrated to a pre-built new oil and gas field database according to the preset classification of the well logging data.

[0023] It should be noted that the embodiments of this specification have the following beneficial effects through the above content:

[0024] Enhanced in-depth digitization capabilities for paper documents: By digitizing paper logging documents and generating RGB images, combined with intelligent extraction of logging text features, this technology breaks through the limitation of traditional scanning that only generates low-structured images, achieves semantic-level analysis (such as well numbers and formation parameters), and significantly improves the searchability and utilization of paper data in the new database.

[0025] Achieve collaborative integration of multi-source heterogeneous data: Synchronously access paper documents and semi-structured electronic documents (such as PDF, Excel), and extract text features (including non-standard annotations and segmentation parameters) through unified rules, breaking the isolation of paper and electronic document processing processes and solving the "information island" problem caused by the dispersion of multi-source data features.

[0026] Improve the intelligent recognition accuracy of logging curve symbols: Adaptive graphic feature extraction technology is used for logging curve symbols to accommodate symbol deformation (such as curve breakpoints and coordinate axis offsets) in documents of different layouts, reducing reliance on manual annotation and the risk of distortion during graphic data migration.

[0027] Build a business-oriented classification and migration system: Preset logging data classification based on oil and gas exploration business needs, migrate text and graphic features to the new database according to the hierarchical storage structure, strengthen the business relevance between data, and support cross-well area data comparison and exploration decision analysis.

[0028] Optimize the efficiency of the data migration process: Through the automated extraction of text and graphic features and classification migration logic, it replaces traditional manual labeling and template matching operations, reduces manual intervention, and shortens the historical data migration cycle.

[0029] Strengthening the business adaptability of the new database: The migrated data combines structured text and high-precision graphic features, meeting the needs of multimodal data fusion analysis (such as parameter text and curve symbols) in oil and gas field exploration and development, and providing a high-quality data foundation for reservoir modeling and production optimization.

[0030] Ensure the integrity of data migration: By comparing the data volumes before and after migration (the first data volume and the second data volume), verify whether the logging text features and graphic features have been completely migrated to the new database. This can avoid data loss due to omissions during feature extraction or transmission, and improve the reliability of data migration.

[0031] Enhance the fault tolerance of the data migration process: Automatically trigger the consistency check mechanism after data migration is completed. If the data volume is inconsistent, the migration process will be re-triggered. This reduces the risk of data corruption caused by external factors such as network interruptions and storage anomalies, and ensures the robustness of the migration process.

[0032] Reduce manual verification costs: An automated data volume comparison mechanism replaces traditional manual sampling checks, significantly reducing the workload of manual review after data migration and improving overall migration efficiency.

[0033] Enhance the credibility of migrated data: Through data consistency verification, the logging text and graphic features stored in the new database are ensured to strictly correspond to the original data, avoiding subsequent analysis errors caused by data misalignment or redundancy, and providing a highly reliable data foundation for oil and gas field exploration.

[0034] Optimize the closed-loop management of the migration process: Embed data consistency verification into the final stage of the migration process, forming a closed-loop management logic of "extract-migrate-verify", and achieving full-link controllability and traceability of the migration process.

[0035] Improved stability of large-scale data migration: By dynamically matching the amount of data to be received in a single migration to the new database, system overload or transmission interruption caused by exceeding the limit of a single migration volume is avoided, ensuring the continuity and stability of migration tasks in high-load scenarios.

[0036] Optimize resource allocation for the target database: Divide migration batches based on the new database's real-time ingestion capacity to prevent instantaneous saturation of memory, storage, or bandwidth resources, ensuring that normal database operations are not disrupted during the migration process.

[0037] Enhanced controllability of the migration process: By predicting whether the data volume exceeds the limit and automatically triggering batch migration, fine-grained control of the migration task granularity is achieved, making it easier for operations personnel to monitor and intervene in abnormal batches in real time.

[0038] Reduce the overall risk of migration failure: Split large-scale data into multiple subtasks that adapt to the target system's carrying capacity. Even if a batch migration fails, only that batch needs to be retried instead of the entire data, reducing data rollback costs and time losses.

[0039] Adapting to performance differences in heterogeneous databases: Dynamically adjust the amount of data migrated per session based on the hardware configuration and processing capabilities of the new oil and gas field database to avoid migration efficiency bottlenecks caused by performance mismatch between the source and target databases.

[0040] Furthermore, the determining of the single migration data volume based on the single migration reception volume of the new oil and gas field database includes:

[0041] The single migration data volume is determined based on the single migration reception volume of the new oil and gas field database and the ratio of the well logging text features to the well logging graphic features.

[0042] It should be noted that the embodiments of this specification have the following beneficial effects through the above content:

[0043] Optimize dynamic resource allocation for heterogeneous data: By combining the ratio of logging text features to graphic features (such as the proportion of text parameters and curve symbols), the amount of data migrated at a time is dynamically adjusted to avoid imbalanced allocation of database computing resources (such as memory and bandwidth) caused by excessive migration of a certain type of data, thereby improving resource utilization during the migration process.

[0044] Improve collaborative processing efficiency for data migration: Split migration batches based on the proportional relationship between text and graphic features, allowing the new database to balance text parsing and graphics rendering tasks during a single reception, reducing delays caused by data type processing conflicts and improving overall migration efficiency.

[0045] Reduce the risk of exceptions caused by data type mismatches: Based on the differences in processing text and graphic features (such as semantic analysis required for text parsing and coordinate mapping required for graphics), the single migration amount is proportionally adapted to prevent parsing errors or storage anomalies caused by excessive amounts of a single data type, thereby enhancing the stability of the migration process.

[0046] Adapting to the business processing needs of multimodal data: Through scale-aware migration volume division, ensure that the new database retains the original correlation between text and graphic features when receiving (such as the synchronous migration of parameter text and corresponding curve symbols), supporting subsequent fusion analysis and application of multimodal data.

[0047] Enhanced migration flexibility for different data structures: Dynamically adjust migration strategies to address the differences in the ratio of text and graphics in well logging data across different oil and gas field projects (e.g., old wells primarily use paper maps, new wells primarily use electronic data), improving the method's compatibility across multiple scenarios.

[0048] Furthermore, before determining the single migration data volume based on the single migration reception volume of the new oil and gas field database and the ratio of the well logging text features to the well logging graphic features, the method further includes:

[0049] The single migration reception volume is determined by testing the historical storage engine performance of the new oil and gas field database and establishing a migration volume prediction model in combination with historical migration logs.

[0050] It should be noted that the embodiments of this specification have the following beneficial effects through the above content:

[0051] Improve the scientific nature of single migration volume decision-making: Through historical storage engine performance testing and migration log modeling, we break through the traditional empirical threshold setting method. This allows the determination of the single migration volume to be more closely aligned with the actual hardware performance of the new database (such as storage I / O and concurrent processing capabilities), avoiding migration volume overload or idle resources caused by manual experience.

[0052] Enhanced dynamic adaptability of migration strategies: Predictive models built based on historical data can perceive database performance fluctuations (such as increased load during peak hours) and dynamically adjust the amount of data received in a single migration to adapt to differences in system status at different times, ensuring real-time matching of migration tasks with the database operating environment.

[0053] Reduce the trial-and-error cost of migration configuration: The model automatically recommends a single migration amount, replacing traditional manual parameter adjustment or multiple trial migrations. This reduces the risk of migration failure or performance degradation due to improper configuration and shortens the migration debugging cycle.

[0054] Accumulate the value of historical data assets: Convert historical migration logs and performance test data into the training basis of predictive models, enable data assets to continuously optimize migration strategies, and improve the method's ability to generalize to similar database migration scenarios.

[0055] Ensure system stability in high-load scenarios: By combining performance test results with model predictions, we precisely control the amount of migration per session to ensure it does not exceed the database's capacity limit, preventing service response delays or downtime caused by sudden data write pressure.

[0056] Furthermore, the historical migration log includes the amount of data for each migration, the migration time, and the performance indicators of the storage engine;

[0057] The historical storage engine performance test of the new oil and gas field database is conducted, and a migration volume prediction model is established in combination with historical migration logs to determine the single migration reception volume, including:

[0058] Performing a historical storage engine performance test on the new oil and gas field database to obtain historical performance test results and the historical migration log;

[0059] Analyze the historical migration logs to identify historical key factors that affect migration performance;

[0060] Establishing a performance prediction model based on the historical performance test results and the historical key factors, wherein the performance prediction model is a machine learning model;

[0061] The single migration reception amount is determined based on the performance prediction model.

[0062] It should be noted that the embodiments of this specification have the following beneficial effects through the above content:

[0063] Improving the accuracy and adaptability of migration volume predictions: By analyzing multi-dimensional data (data volume, time, and performance indicators) in historical migration logs and combining them with machine learning models to capture historical key factors (such as storage engine I / O bottlenecks and concurrent load thresholds), we dynamically generate migration volumes that match current database performance, overcoming prediction biases based on manual experience or fixed rules.

[0064] Enhance the model's ability to analyze complex performance correlations: A performance prediction model based on machine learning automatically learns the nonlinear relationship between historical performance test results and key migration factors (such as storage latency caused by data surges). This accurately quantifies the impact of storage engine performance on migration volume, enabling adaptive optimization of migration strategies.

[0065] Reduce the potential risk of migration performance fluctuations: By identifying historical key factors (such as performance drops triggered by specific data volume thresholds), predict and avoid migration volume configurations that may cause database instability, fundamentally reducing the probability of storage engine overload or response timeouts during the migration process.

[0066] Continuous self-optimization of migration strategies: Machine learning models can continuously update parameters and iteratively optimize as historical migration logs and performance test data accumulate. This allows the decision logic for a single migration to evolve autonomously with database hardware upgrades or changes in business scenarios, maintaining long-term effectiveness.

[0067] Enhanced explainability and controllability of the migration process: By explicitly identifying key factors affecting migration performance (such as the correlation between migration time and storage engine CPU usage), this provides decision-making support for operations and maintenance personnel, facilitating targeted adjustments to database configurations or migration plans, and improving the transparency and controllability of the migration process.

[0068] Furthermore, determining the single migration reception amount based on the performance prediction model includes:

[0069] Performing a performance test on the current storage engine of the new oil and gas field database to obtain current performance test results and current migration logs;

[0070] Analyze the current migration log to identify current key factors affecting migration performance;

[0071] The current performance test result and the current key factors are input into the performance prediction model to determine the single migration reception quantity.

[0072] It should be noted that the embodiments of this specification have the following beneficial effects through the above content:

[0073] Real-time dynamic optimization of migration strategies: By combining current storage engine performance test results with real-time migration log analysis, key influencing factors (such as sudden load fluctuations and hardware status changes) are dynamically identified. This allows the decision on the amount of data to be transferred to adapt to the instantaneous performance of the database in real time, avoiding degradation in migration efficiency caused by historical models being out of sync with the current environment.

[0074] Enhanced fault tolerance for sudden performance anomalies: Identifies migration risk factors (such as transient I / O bottlenecks in the storage engine) based on current performance data, and proactively avoids unsafe migration volume settings through model prediction, reducing the probability of migration interruptions caused by hardware failures or external interference.

[0075] Improve the environmental awareness accuracy of migration volume decisions: Synchronize the current performance test results and the operating status in the real-time log (such as the number of concurrent tasks and cache occupancy) into the model to ensure that the migration volume accurately matches the real-time resource reserve of the database, optimizing the balance between resource utilization and migration efficiency.

[0076] Supports closed-loop iterative optimization of migration strategies: By continuously feeding current migration logs and performance data into the prediction model, a closed-loop optimization chain of "decision-execution-verification-update" is formed. This enables the model to autonomously evolve as the database operating environment changes (such as hardware upgrades and business expansion), maintaining long-term effectiveness.

[0077] Enhanced full-cycle controllability of the migration process: Before each migration task is launched, a real-time data-driven model re-evaluation mechanism ensures that the single migration volume is always within the optimal load range of the database, achieving full-cycle stability assurance from task initiation to completion.

[0078] One or more embodiments of this specification provide a historical data migration device for an oil and gas field database, including:

[0079] An image generation unit digitizes paper logging documents in the oil and gas field database to generate RGB images;

[0080] An access unit for synchronously accessing a semi-structured electronic document in the oil and gas field database;

[0081] a text feature extraction unit for extracting features of well logging text in the RGB image and the electronic document;

[0082] A graphic feature extraction unit extracts the well logging graphic features in the RGB image and the electronic document based on the well logging curve symbols;

[0083] a first data volume determining unit, configured to determine a first data volume of the well logging text feature and the well logging graphic feature;

[0084] a migration data amount determining unit, configured to determine a single migration data amount based on a single migration reception amount of the new oil and gas field database;

[0085] a migration data volume determination unit, configured to determine whether the first data volume exceeds the single migration data volume;

[0086] a migration data volume division unit, if yes, dividing the first data volume into multiple migration data volumes according to the single migration data volume;

[0087] a second data volume determining unit, configured to determine a second data volume of the new oil and gas field database;

[0088] a data volume comparison unit, for comparing whether the first data volume is consistent with the second data volume;

[0089] Migration unit, if consistent, then according to the preset classification of the well logging data, the well logging text features and the well logging graphic features are migrated to a pre-built new oil and gas field database.

[0090] One or more embodiments of this specification provide a historical data migration device for an oil and gas field database, including:

[0091] at least one processor; and,

[0092] a memory communicatively connected to the at least one processor; wherein,

[0093] The memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to:

[0094] Digitize paper logging documents in oil and gas field databases to generate RGB images;

[0095] Synchronously accessing semi-structured electronic documents in the oil and gas field database;

[0096] Extracting well logging text features from the RGB image and the electronic document;

[0097] For the well logging curve symbols, extract the well logging graphic features in the RGB image and the electronic document;

[0098] determining a first data volume of the well logging text feature and the well logging graphic feature;

[0099] Determining a single migration data volume based on a single migration reception volume of the new oil and gas field database;

[0100] Determining whether the first data volume exceeds the single migration data volume;

[0101] If so, dividing the first data volume into multiple migration data volumes according to the single migration data volume;

[0102] determining a second data volume of the new oil and gas field database;

[0103] comparing whether the first data volume is consistent with the second data volume;

[0104] If they are consistent, the well logging text features and the well logging graphic features are migrated to a pre-built new oil and gas field database according to the preset classification of the well logging data.

[0105] One or more embodiments of this specification provide a non-volatile computer storage medium storing computer-executable instructions. When executed by a computer, the computer-executable instructions can achieve:

[0106] Digitize paper logging documents in oil and gas field databases to generate RGB images;

[0107] Synchronously accessing semi-structured electronic documents in the oil and gas field database;

[0108] Extracting well logging text features from the RGB image and the electronic document;

[0109] For the well logging curve symbols, extract the well logging graphic features in the RGB image and the electronic document;

[0110] determining a first data volume of the well logging text feature and the well logging graphic feature;

[0111] Determining a single migration data volume based on a single migration reception volume of the new oil and gas field database;

[0112] Determining whether the first data volume exceeds the single migration data volume;

[0113] If so, dividing the first data volume into multiple migration data volumes according to the single migration data volume;

[0114] determining a second data volume of the new oil and gas field database;

[0115] comparing whether the first data volume is consistent with the second data volume;

[0116] If they are consistent, the well logging text features and the well logging graphic features are migrated to a pre-built new oil and gas field database according to the preset classification of the well logging data.

[0117] At least one of the above technical solutions adopted in the embodiments of this specification can achieve the following beneficial effects:

[0118] Enhanced in-depth digitization capabilities for paper documents: By digitizing paper logging documents and generating RGB images, combined with intelligent extraction of logging text features, this technology breaks through the limitation of traditional scanning that only generates low-structured images, achieves semantic-level analysis (such as well numbers and formation parameters), and significantly improves the searchability and utilization of paper data in the new database.

[0119] Achieve collaborative integration of multi-source heterogeneous data: Synchronously access paper documents and semi-structured electronic documents (such as PDF, Excel), and extract text features (including non-standard annotations and segmentation parameters) through unified rules, breaking the isolation of paper and electronic document processing processes and solving the "information island" problem caused by the dispersion of multi-source data features.

[0120] Improve the intelligent recognition accuracy of logging curve symbols: Adaptive graphic feature extraction technology is used for logging curve symbols to accommodate symbol deformation (such as curve breakpoints and coordinate axis offsets) in documents of different layouts, reducing reliance on manual annotation and the risk of distortion during graphic data migration.

[0121] Build a business-oriented classification and migration system: Preset logging data classification based on oil and gas exploration business needs, migrate text and graphic features to the new database according to the hierarchical storage structure, strengthen the business relevance between data, and support cross-well area data comparison and exploration decision analysis.

[0122] Optimize the efficiency of the data migration process: Through the automated extraction of text and graphic features and classification migration logic, it replaces traditional manual labeling and template matching operations, reduces manual intervention, and shortens the historical data migration cycle.

[0123] Strengthening the business adaptability of the new database: The migrated data combines structured text and high-precision graphic features, meeting the needs of multimodal data fusion analysis (such as parameter text and curve symbols) in oil and gas field exploration and development, and providing a high-quality data foundation for reservoir modeling and production optimization. BRIEF DESCRIPTION OF THE DRAWINGS

[0124] In order to more clearly illustrate the embodiments of this specification or the technical solutions in the prior art, the following briefly introduces the drawings required for the embodiments or the description of the prior art. Obviously, the drawings described below are only some of the embodiments described in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without inventive work. In the drawings:

[0125] Figure 1 A schematic flow chart of a method for migrating historical data of an oil and gas field database provided in one or more embodiments of this specification;

[0126] Figure 2 A schematic diagram of the structure of a historical data migration device for an oil and gas field database provided in one or more embodiments of this specification;

[0127] Figure 3 A schematic diagram of the structure of a historical data migration device for an oil and gas field database provided in one or more embodiments of this specification. DETAILED DESCRIPTION

[0128] The embodiments of this specification provide a method, apparatus, device, and medium for migrating historical data of an oil and gas field database.

[0129] To help those skilled in the art better understand the technical solutions in this specification, the following will provide a clear and complete description of the technical solutions in the embodiments of this specification, in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of this specification, not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this specification without creative work should fall within the scope of protection of this specification.

[0130] Figure 1 This is a flow chart of a method for migrating historical data from an oil and gas field database, provided in one or more embodiments of this specification. This process can be executed by a historical data migration system for an oil and gas field database. Certain input parameters or intermediate results in the process can be manually adjusted to help improve accuracy.

[0131] The method steps of the embodiment of this specification are as follows:

[0132] S101, digitize paper logging documents in the oil and gas field database to generate RGB images.

[0133] In the examples of this specification, the following specific implementation schemes can be used:

[0134] 1. Document scanning and preprocessing

[0135] The paper well logging documents were scanned using a high-resolution scanner to generate original RGB images.

[0136] De-noise, de-skew, and enhance contrast on scanned images to ensure they are clear and readable.

[0137] 2. Standardized image storage

[0138] The processed RGB image is stored in a temporary directory according to a unified naming rule (such as "pound sign_date_page number") to provide input for subsequent feature extraction.

[0139] S102, synchronously accessing the semi-structured electronic documents in the oil and gas field database.

[0140] In the examples of this specification, the following specific implementation schemes can be used:

[0141] 1. Multi-format document access

[0142] Batch read semi-structured electronic documents such as PDF and Excel through database or file system interfaces. Access encrypted or permission-restricted documents after decryption or permission verification.

[0143] 2. Format unification

[0144] Convert documents of different formats into intermediate data formats (such as PDF to text, Excel to CSV) to eliminate the interference of format differences on feature extraction.

[0145] S103: Extracting well logging text features from the RGB image and the electronic document.

[0146] In the examples of this specification, the following specific implementation schemes can be used:

[0147] 1. Image Character Recognition (OCR)

[0148] The text area in the RGB image is located and the OCR technology is used to recognize the text content (such as well number, depth value, and lithology description).

[0149] 2. Electronic document text analysis

[0150] Perform structured analysis on the text content in semi-structured electronic documents and extract key fields (such as "porosity" and "permeability" parameters in the table).

[0151] 3. Semantic annotation and association

[0152] Through natural language processing (NLP) technology, the extracted text is annotated with semantic tags (such as "stratum name" and "test conclusion") according to business logic.

[0153] S104 , extracting well logging graphic features from the RGB image and the electronic document for the well logging curve symbols.

[0154] In the examples of this specification, the following specific implementation schemes can be used:

[0155] 1. Curve symbol recognition

[0156] Perform image segmentation on the logging curve area in the RGB image, identify the coordinate points, line type and scale information of the curve symbol. Extract the original coordinate data of vector graphics in electronic documents (such as embedded curve graphs in PDF).

[0157] 2. Graphic feature structuring

[0158] Convert curve symbols into standardized data formats (e.g., time-depth curves into JSON arrays), preserving the topological relationships and properties of the symbols.

[0159] S105: Determine a first data volume of the well logging text feature and the well logging graphic feature.

[0160] S106: Determine a single migration data volume based on a single migration reception volume of the new oil and gas field database.

[0161] S107: Determine whether the first data volume exceeds the single migration data volume.

[0162] S108: If yes, divide the first data volume into multiple migration data volumes according to the single migration data volume.

[0163] S109: Determine a second data volume of the new oil and gas field database.

[0164] S110: Compare the first data volume and the second data volume to see whether they are consistent.

[0165] S111: If they are consistent, the well logging text features and the well logging graphic features are migrated to a pre-built new oil and gas field database according to a preset classification of the well logging data.

[0166] In the examples of this specification, the following specific implementation schemes can be used:

[0167] 1. Classification rule matching

[0168] According to the preset logging data classification (such as stratification by "well area", "data type" and "acquisition time"), text features and graphic features are assigned to corresponding classification labels.

[0169] 2. Multimodal Data Fusion and Migration

[0170] Associate text features (structured text) with graphic features (standardized coordinate data) according to classification labels, and write them in batches to the designated storage module of the new database using ETL tools.

[0171] 3. Verify migration results

[0172] Perform sampling verification on the data written into the new database to ensure the integrity and logical consistency of text and graphic features.

[0173] It should be noted that the embodiments of this specification have the following beneficial effects through the above content:

[0174] Enhanced in-depth digitization capabilities for paper documents: By digitizing paper logging documents and generating RGB images, combined with intelligent extraction of logging text features, this technology breaks through the limitation of traditional scanning that only generates low-structured images, achieves semantic-level analysis (such as well numbers and formation parameters), and significantly improves the searchability and utilization of paper data in the new database.

[0175] Achieve collaborative integration of multi-source heterogeneous data: Synchronously access paper documents and semi-structured electronic documents (such as PDF, Excel), and extract text features (including non-standard annotations and segmentation parameters) through unified rules, breaking the isolation of paper and electronic document processing processes and solving the "information island" problem caused by the dispersion of multi-source data features.

[0176] Improve the intelligent recognition accuracy of logging curve symbols: Adaptive graphic feature extraction technology is used for logging curve symbols to accommodate symbol deformation (such as curve breakpoints and coordinate axis offsets) in documents of different layouts, reducing reliance on manual annotation and the risk of distortion during graphic data migration.

[0177] Build a business-oriented classification and migration system: Preset logging data classification based on oil and gas exploration business needs, migrate text and graphic features to the new database according to the hierarchical storage structure, strengthen the business relevance between data, and support cross-well area data comparison and exploration decision analysis.

[0178] Optimize the efficiency of the data migration process: Through the automated extraction of text and graphic features and classification migration logic, it replaces traditional manual labeling and template matching operations, reduces manual intervention, and shortens the historical data migration cycle.

[0179] Strengthening the business adaptability of the new database: The migrated data combines structured text and high-precision graphic features, meeting the needs of multimodal data fusion analysis (such as parameter text and curve symbols) in oil and gas field exploration and development, and providing a high-quality data foundation for reservoir modeling and production optimization.

[0180] Specifically, for the embodiment of this specification, determining a first data volume of the well logging text feature and the well logging graphic feature; determining a second data volume of the new oil and gas field database; comparing the first data volume and the second data volume to see whether they are consistent; if they are consistent, completing the migration of the well logging text feature and the well logging graphic feature to the new oil and gas field database can be implemented through the following specific implementation plan:

[0181] Step 1: Determine the first data volume of logging text features and graphic features

[0182] 1. Definition of data volume statistics rules

[0183] Text feature statistics: Traverse all extracted logging text features (such as well numbers, formation parameters), and count the total number of records and fields by field type (text, value).

[0184] Graphic feature statistics: Count the total number of logging curve symbols (such as coordinate points, line types, and scales) by graphic unit (single curve, symbol group).

[0185] 2. Automatic data volume calculation

[0186] Develop scripts or call the database statistics interface to traverse the temporary storage directory before migration and summarize the total data volume of text features and graphic features (such as the number of records and storage units).

[0187] Step 2: Determine the second data volume of the new oil and gas field database

[0188] 1. Target database data statistics

[0189] In the corresponding storage modules of the new database (such as the structured text table and the graphic feature table), a predefined data volume query operation is performed to count the number of records of migrated text and graphic features.

[0190] 2. Real-time feedback on data volume

[0191] After the migration task is completed, the data volume statistics program is automatically triggered to generate a data volume report after migration (such as "X text features migrated, Y groups of graphic features migrated").

[0192] Step 3: Compare the first data volume with the second data volume to see if they are consistent

[0193] 1. Consistency check logic design

[0194] Define the comparison rules: if the number of text feature records and the number of graphic feature units are equal, they are considered consistent; otherwise, they are inconsistent.

[0195] 2. Automated comparison process

[0196] The first data volume counted before the migration and the second data volume counted after the migration are input into a comparison program, a verification result (such as "consistent" or "inconsistent") is output, and difference details (such as missing well signs or curve numbers) are recorded.

[0197] Step 4: If they are consistent, the migration is completed.

[0198] 1. Migration completion confirmation mechanism

[0199] When the consistency check result is "consistent", the commit operation of the new database is automatically triggered, the temporarily stored migration data is officially written to the business table, and the temporary storage resources are released.

[0200] 2. Result notification and logging

[0201] A migration success notification is sent to the operation and maintenance system. The first and second data volumes of the migration and the verification results are recorded in the audit log for subsequent tracing.

[0202] It should be noted that the embodiments of this specification have the following beneficial effects through the above content:

[0203] Ensure the integrity of data migration: By comparing the data volumes before and after migration (the first data volume and the second data volume), verify whether the logging text features and graphic features have been completely migrated to the new database. This can avoid data loss due to omissions during feature extraction or transmission, and improve the reliability of data migration.

[0204] Enhance the fault tolerance of the data migration process: Automatically trigger the consistency check mechanism after data migration is completed. If the data volume is inconsistent, the migration process will be re-triggered. This reduces the risk of data corruption caused by external factors such as network interruptions and storage anomalies, and ensures the robustness of the migration process.

[0205] Reduce manual verification costs: An automated data volume comparison mechanism replaces traditional manual sampling checks, significantly reducing the workload of manual review after data migration and improving overall migration efficiency.

[0206] Enhance the credibility of migrated data: Through data consistency verification, the logging text and graphic features stored in the new database are ensured to strictly correspond to the original data, avoiding subsequent analysis errors caused by data misalignment or redundancy, and providing a highly reliable data foundation for oil and gas field exploration.

[0207] Optimize the closed-loop management of the migration process: Embed data consistency verification into the final stage of the migration process, forming a closed-loop management logic of "extract-migrate-verify", and achieving full-link controllability and traceability of the migration process.

[0208] Specifically, before migrating the well logging text features and the well logging graphic features to the pre-built new oil and gas field database, a single migration data volume can be determined based on a single migration reception volume of the new oil and gas field database; and a determination can be made as to whether the first data volume exceeds the single migration data volume. If so, the first data volume can be divided into multiple migration data volumes based on the single migration data volume. This can be achieved through the following specific implementation scheme:

[0209] Step 1: Determine the single migration volume for the new database

[0210] 1. Dynamically evaluate database load capacity

[0211] Conduct real-time performance tests on the new oil and gas field database (such as concurrent write stress tests and storage I / O throughput tests). Combined with the load threshold records in historical migration logs, comprehensively evaluate the upper limit of the current amount of data that can be safely received at a single time.

[0212] 2. Dynamic adjustment rules for receiving volume

[0213] Dynamically adjust the single migration volume based on the database's real-time resource usage (such as CPU, memory, and free disk space) to ensure that core database services are not affected during the migration process.

[0214] Step 2: Determine whether the first data volume exceeds the single migration receiving volume

[0215] 1. Comparison of Automation Data Volume

[0216] Use a script or migration tool to automatically compare the total amount of data to be migrated (the first amount of data) with the amount received in a single migration:

[0217] If the first data volume is less than or equal to the single-time received volume, a full one-time migration is triggered.

[0218] If the first data volume is greater than the single data volume received, the batch migration logic is triggered.

[0219] 2. Threshold Alarm Mechanism

[0220] When the data volume exceeds the limit, an alarm notification is sent to the operation and maintenance system, and the migration task is automatically frozen until manual confirmation or system resources are released.

[0221] Step 3: Divide the data volume into multiple migrations

[0222] 1. Data batching strategy

[0223] Split by data type ratio: Split the total data volume into multiple batches proportionally based on the data volume ratio of text features and graphic features (for example, text accounts for 70% and graphics accounts for 30%), ensuring that the ratio of the two types of data in each batch is consistent with the original data.

[0224] Group by business relevance: Based on the business attributes of logging data (such as well number and acquisition time), highly correlated text and graphic features are divided into the same batch to avoid data association breaks across batches.

[0225] 2. Batch migration sequence scheduling

[0226] A priority queue mechanism is used to prioritize the migration of high-value data (such as recently collected data and key well data), and the batch interval is controlled through the task scheduler to prevent concentrated writing from causing resource contention.

[0227] It should be noted that the embodiments of this specification have the following beneficial effects through the above content:

[0228] Improved stability of large-scale data migration: By dynamically matching the amount of data to be received in a single migration to the new database, system overload or transmission interruption caused by exceeding the limit of a single migration volume is avoided, ensuring the continuity and stability of migration tasks in high-load scenarios.

[0229] Optimize resource allocation for the target database: Divide migration batches based on the new database's real-time ingestion capacity to prevent instantaneous saturation of memory, storage, or bandwidth resources, ensuring that normal database operations are not disrupted during the migration process.

[0230] Enhanced controllability of the migration process: By predicting whether the data volume exceeds the limit and automatically triggering batch migration, fine-grained control of the migration task granularity is achieved, making it easier for operations personnel to monitor and intervene in abnormal batches in real time.

[0231] Reduce the overall risk of migration failure: Split large-scale data into multiple subtasks that adapt to the target system's carrying capacity. Even if a batch migration fails, only that batch needs to be retried instead of the entire data, reducing data rollback costs and time losses.

[0232] Adapting to performance differences in heterogeneous databases: Dynamically adjust the amount of data migrated per session based on the hardware configuration and processing capabilities of the new oil and gas field database to avoid migration efficiency bottlenecks caused by performance mismatch between the source and target databases.

[0233] Furthermore, when determining the single migration data volume based on the single migration reception volume of the new oil and gas field database, the single migration data volume can be determined based on the single migration reception volume of the new oil and gas field database and the ratio of the logging text features to the logging graphic features.

[0234] In the examples of this specification, the following specific implementation schemes can be used:

[0235] Step 1: Dynamically evaluate the single migration volume of the new database

[0236] 1. Real-time performance testing and historical data analysis

[0237] Perform short-term stress tests on the new database (such as simulating write peaks) to obtain the maximum stable reception capacity (such as the number of records processed per second or storage throughput) under the current hardware configuration.

[0238] Combined with the successful batch data volume records in historical migration logs, analyze the fluctuation range of the received volume under different load scenarios (such as business peaks and valleys).

[0239] 2. Dynamic adjustment mechanism of receiving volume

[0240] Based on the real-time database resource usage (such as CPU utilization and remaining memory), the upper limit of the amount of data received during a single migration is dynamically adjusted according to preset rules (for example, if resource usage exceeds 70%, the amount of data received is automatically reduced).

[0241] Step 2: Calculate the ratio of text and graphic features in well logging

[0242] 1. Calculation of data feature ratio

[0243] The data volume (such as the number of records and the number of storage units) of the extracted logging text features (such as well numbers and parameter tables) and graphic features (such as curve symbols and coordinate points) is counted respectively, and the ratio of the two is calculated (for example, text accounts for 60% and graphics accounts for 40%).

[0244] 2. Proportional Standardized Storage

[0245] Store ratio information (such as text:graphics = 3:2) as metadata in the migration task configuration file for subsequent batch migration calls.

[0246] Step 3: Dynamically allocate the amount of data to be migrated in a single batch

[0247] 1. Proportional Adaptive Batch Strategy

[0248] Data splitting logic: Based on the maximum amount of data that can be received in a single migration, data is split according to the ratio of text to graphics. For example, if the amount of data received in a single migration is 1000 units and the ratio is 3:2, then a single migration will include 600 units of text and 400 units of graphics.

[0249] Maintaining business relevance: Ensure that the text features and corresponding graphic features of the same logging data are assigned to the same batch (for example, the text description and curve symbols of Well A are migrated in the same batch) to avoid data relevance breaks.

[0250] 2. Automated batch generation

[0251] Develop a migration scheduling tool to extract text and graphic features from the data pool to be migrated in proportion, generate several complete batches, and mark the batch numbers and content lists.

[0252] Step 4: Migration execution and dynamic adjustment

[0253] 1. Batch migration priority control

[0254] High-priority data (such as recently active wells and key parameters) are migrated first, and the data volume allocation of subsequent batches is dynamically adjusted based on the ratio.

[0255] 2. Real-time feedback and elastic scaling

[0256] Monitor the actual execution effect of each migration (such as time consumption and resource usage). If high graphics processing latency is found, automatically reduce the proportion of graphics features in subsequent batches or the single reception quantity.

[0257] It should be noted that the embodiments of this specification have the following beneficial effects through the above content:

[0258] Optimize dynamic resource allocation for heterogeneous data: By combining the ratio of logging text features to graphic features (such as the proportion of text parameters and curve symbols), the amount of data migrated at a time is dynamically adjusted to avoid imbalanced allocation of database computing resources (such as memory and bandwidth) caused by excessive migration of a certain type of data, thereby improving resource utilization during the migration process.

[0259] Improve collaborative processing efficiency for data migration: Split migration batches based on the proportional relationship between text and graphic features, allowing the new database to balance text parsing and graphics rendering tasks during a single reception, reducing delays caused by data type processing conflicts and improving overall migration efficiency.

[0260] Reduce the risk of exceptions caused by data type mismatches: Based on the differences in processing text and graphic features (such as semantic analysis required for text parsing and coordinate mapping required for graphics), the single migration amount is proportionally adapted to prevent parsing errors or storage anomalies caused by excessive amounts of a single data type, thereby enhancing the stability of the migration process.

[0261] Adapting to the business processing needs of multimodal data: Through scale-aware migration volume division, ensure that the new database retains the original correlation between text and graphic features when receiving (such as the synchronous migration of parameter text and corresponding curve symbols), supporting subsequent fusion analysis and application of multimodal data.

[0262] Enhanced migration flexibility for different data structures: Dynamically adjust migration strategies to address the differences in the ratio of text and graphics in well logging data across different oil and gas field projects (e.g., old wells primarily use paper maps, new wells primarily use electronic data), improving the method's compatibility across multiple scenarios.

[0263] Furthermore, before determining the single migration data volume based on the single migration reception volume of the new oil and gas field database and the ratio of the logging text features to the logging graphic features, the historical storage engine performance test of the new oil and gas field database can be conducted, and a migration volume prediction model can be established in combination with historical migration logs to determine the single migration reception volume.

[0264] In the examples of this specification, the following specific implementation schemes can be used:

[0265] Step 1: Historical storage engine performance testing and log collection

[0266] 1. Performance test scenario design

[0267] Simulate different load scenarios (such as low, medium, and high concurrent writes and mixed read and write operations) in the new database's storage engine, and record key performance indicators (such as transactions per second, disk I / O latency, and memory usage) during the test.

[0268] The test results are categorized and stored, with marking of test time, load type, and corresponding performance thresholds (such as the maximum stable data reception amount).

[0269] 2. Historical migration log integration

[0270] Extract historical migration logs from the old migration system, including fields such as the data volume, duration, success / failure status, and error type of each migration, to build a structured log database.

[0271] Step 2: Identify key performance influencing factors

[0272] 1. Data correlation analysis

[0273] Align historical performance test results with migration logs by timestamp to analyze the correlation between performance metrics (such as I / O latency) and migration results (such as failure rate).

[0274] Identify key factors through statistical methods (for example, when disk I / O latency exceeds 50ms, the probability of migration failure increases by 80%).

[0275] 2. Feature Engineering

[0276] Derivate features from the raw data in the log (such as data volume and timestamp) to generate migration task features (such as the ratio of single data volume to total data volume and the number of concurrent tasks in the same time period).

[0277] Step 3: Construction of migration prediction model

[0278] 1. Model selection and training

[0279] Use supervised learning algorithms (such as random forests and gradient boosting trees) to train a regression model with historical performance test data as input features and the maximum amount of successfully migrated data as labels.

[0280] Cross-validate the model to ensure its generalization ability under various load scenarios.

[0281] 2. Model deployment and interface encapsulation

[0282] The trained model is encapsulated as an API service and integrated into the migration scheduling system, which supports real-time reception of performance parameters and returns recommended migration amounts.

[0283] Step 4: Dynamically predict migration volume based on current status

[0284] 1. Real-time performance testing and log collection

[0285] Before starting each migration task, perform a lightweight performance test (such as a 5-second stress test) on the new database to obtain the current storage engine's I / O, CPU, and memory status.

[0286] Synchronously collect real-time migration logs (such as the current number of concurrent tasks and the amount of migrated data).

[0287] 2. Model reasoning and migration output

[0288] The prediction model inputs current performance data (such as disk I / O latency and memory usage) and real-time log features, and outputs a recommended value for the single migration reception volume that is suitable for the current status.

[0289] Step 5: Continuous model optimization and feedback mechanism

[0290] 1. Online learning update

[0291] After each migration task is completed, the actual migration results (success / failure, time consumption) are compared with the model prediction value, and model parameter fine-tuning (such as incremental training) is automatically triggered.

[0292] 2. Abnormal scenario handling

[0293] When the deviation between the model prediction value and the actual performance exceeds the threshold, an alarm is triggered and the system switches to the backup rule engine (such as a conservative migration strategy based on historical averages). At the same time, the cause of the anomaly is recorded for iterative repair of the model.

[0294] It should be noted that the embodiments of this specification have the following beneficial effects through the above content:

[0295] Improve the scientific nature of single migration volume decision-making: Through historical storage engine performance testing and migration log modeling, we break through the traditional empirical threshold setting method. This allows the determination of the single migration volume to be more closely aligned with the actual hardware performance of the new database (such as storage I / O and concurrent processing capabilities), avoiding migration volume overload or idle resources caused by manual experience.

[0296] Enhanced dynamic adaptability of migration strategies: Predictive models built based on historical data can perceive database performance fluctuations (such as increased load during peak hours) and dynamically adjust the amount of data received in a single migration to adapt to differences in system status at different times, ensuring real-time matching of migration tasks with the database operating environment.

[0297] Reduce the trial-and-error cost of migration configuration: The model automatically recommends a single migration amount, replacing traditional manual parameter adjustment or multiple trial migrations. This reduces the risk of migration failure or performance degradation due to improper configuration and shortens the migration debugging cycle.

[0298] Accumulate the value of historical data assets: Convert historical migration logs and performance test data into the training basis of predictive models, enable data assets to continuously optimize migration strategies, and improve the method's ability to generalize to similar database migration scenarios.

[0299] Ensure system stability in high-load scenarios: By combining performance test results with model predictions, we precisely control the amount of migration per session to ensure it does not exceed the database's capacity limit, preventing service response delays or downtime caused by sudden data write pressure.

[0300] Furthermore, the historical migration log includes the data volume, migration time and performance indicators of the storage engine for each migration; the historical storage engine performance test of the new oil and gas field database is combined with the historical migration log to establish a migration volume prediction model to determine the single migration reception volume. The historical storage engine performance test of the new oil and gas field database can be used to obtain historical performance test results and the historical migration log; the historical migration log is analyzed to identify historical key factors affecting migration performance; based on the historical performance test results and the historical key factors, a performance prediction model is established, and the performance prediction model is a machine learning model; based on the performance prediction model, the single migration reception volume is determined.

[0301] In the examples of this specification, the following specific implementation schemes can be used:

[0302] Step 1: Integrate historical performance testing and migration logs

[0303] 1. Data collection and storage

[0304] Performance test data collection: Perform historical storage engine performance tests (such as I / O throughput, CPU peak, and memory usage under different loads) in the new oil and gas field database and store the test results in a structured data table by timestamp.

[0305] Migration log normalization: Clean historical migration logs, extract key fields (migration time, data volume, time consumption, error code), and associate them with performance test results by timestamp to build a unified historical data set.

[0306] Step 2: Analyze historical migration logs and identify key factors

[0307] 1. Data Exploration and Pattern Mining

[0308] Anomaly Detection: Use log analysis tools (such as Python Pandas or ELK Stack) to collect records of migration failures or abnormal migration times and analyze their corresponding performance indicators (for example, a significant increase in failure rate during periods of high I / O latency).

[0309] Correlation analysis: Calculate the correlation between the migration data volume and duration and performance indicators (such as CPU utilization and disk write rate) to identify strong correlation factors (for example, the correlation coefficient between disk I / O latency and migration duration is greater than 0.8).

[0310] 2. Extraction and annotation of key factors

[0311] Factors that significantly affect migration performance (such as "I / O latency > 100ms" and "number of concurrent tasks > 5") are marked as historical key factors and stored as feature labels.

[0312] Step 3: Performance prediction model construction and training

[0313] 1. Feature Engineering and Dataset Construction

[0314] Input feature design: Generate model input features (such as "I / O latency," "number of concurrent tasks," and "data volume / maximum throughput ratio") based on historical key factors and performance test results.

[0315] Label definition: Use the maximum single data reception amount in historical successful migration tasks as the label to construct a supervised learning dataset.

[0316] 2. Model selection and training

[0317] Algorithm selection: Use regression machine learning models (such as random forest and XGBoost) to adapt to numerical label prediction tasks.

[0318] Training and validation: Divide the training set and test set by time window (e.g., the first 80% of the time data is used for training, and the last 20% is used for validation), and evaluate the prediction accuracy of the model on historical data (e.g., MAE, R²).

[0319] Step 4: Model-based decision on the amount of money to be received in a single migration

[0320] 1. Model deployment and interface encapsulation

[0321] Encapsulate the trained model as a REST API or database built-in function and integrate it into the migration scheduling system. This system supports real-time reception of current performance parameters and returns migration amount recommendations.

[0322] 2. Dynamic Migration Amount Prediction Process

[0323] Input preprocessing: Converts current storage engine performance metrics (such as real-time I / O rate) and migration task attributes (such as the amount of data to be migrated) into the model input format.

[0324] Model inference: Call the model API to output the predicted value of the number of records received in a single migration (for example, "a maximum of 5,000 records received in a single migration").

[0325] Safety threshold limit: If the model output value exceeds the hardware load limit (such as insufficient disk space), a conservative value (such as 80% of the historical safe migration amount) is automatically triggered as the final decision.

[0326] Step 5: Continuous model optimization and feedback loop

[0327] 1. Online learning and iterative updates

[0328] After each migration task is completed, the actual migration results (such as the amount of successful data and time consumption) are compared with the model prediction value to generate incremental training data and trigger model retraining regularly (such as automatic weekly updates).

[0329] 2. Abnormal scenario handling mechanism

[0330] When the model prediction deviation exceeds the threshold (such as prediction error > 30%), an alarm is triggered and the rule engine is switched to (such as a static migration strategy based on historical averages). At the same time, the cause of the anomaly is recorded for model repair.

[0331] It should be noted that the embodiments of this specification have the following beneficial effects through the above content:

[0332] Improving the accuracy and adaptability of migration volume predictions: By analyzing multi-dimensional data (data volume, time, and performance indicators) in historical migration logs and combining them with machine learning models to capture historical key factors (such as storage engine I / O bottlenecks and concurrent load thresholds), we dynamically generate migration volumes that match current database performance, overcoming prediction biases based on manual experience or fixed rules.

[0333] Enhance the model's ability to analyze complex performance correlations: A performance prediction model based on machine learning automatically learns the nonlinear relationship between historical performance test results and key migration factors (such as storage latency caused by data surges). This accurately quantifies the impact of storage engine performance on migration volume, enabling adaptive optimization of migration strategies.

[0334] Reduce the potential risk of migration performance fluctuations: By identifying historical key factors (such as performance drops triggered by specific data volume thresholds), predict and avoid migration volume configurations that may cause database instability, fundamentally reducing the probability of storage engine overload or response timeouts during the migration process.

[0335] Continuous self-optimization of migration strategies: Machine learning models can continuously update parameters and iteratively optimize as historical migration logs and performance test data accumulate. This allows the decision logic for a single migration to evolve autonomously with database hardware upgrades or changes in business scenarios, maintaining long-term effectiveness.

[0336] Enhanced explainability and controllability of the migration process: By explicitly identifying key factors affecting migration performance (such as the correlation between migration time and storage engine CPU usage), this provides decision-making support for operations and maintenance personnel, facilitating targeted adjustments to database configurations or migration plans, and improving the transparency and controllability of the migration process.

[0337] Furthermore, when determining the single migration reception volume based on the performance prediction model, the current storage engine performance test of the new oil and gas field database can be used to obtain the current performance test results and the current migration log; the current migration log can be analyzed to identify the current key factors affecting the migration performance; the current performance test results and the current key factors can be input into the performance prediction model to determine the single migration reception volume.

[0338] In the examples of this specification, the following specific implementation schemes can be used:

[0339] Step 1: Perform a performance test on the current storage engine

[0340] 1. Lightweight real-time performance test design

[0341] Test scenario simulation: Perform a short (e.g., 10-second) write stress test on the new database, simulate migration tasks with different data volumes (e.g., small batch, medium batch), and record the storage engine's real-time response metrics (e.g., disk I / O rate, CPU usage, and memory peak).

[0342] Standardize test results: Store performance indicators in a unified format (such as JSON or database table), including fields such as timestamp, test type, and performance value.

[0343] 2. Real-time collection of migration logs

[0344] Log capture and filtering: Use the database's built-in logging tools or a third-party monitoring system (such as Prometheus) to capture log data for the current migration task (such as migration rate, error type, and task duration), and filter out irrelevant noise data (such as debug logs).

[0345] Step 2: Analyze the current migration log to identify key factors

[0346] 1. Key factor extraction process

[0347] Abnormal pattern identification: Use log analysis tools (such as ELK Stack) to detect frequent errors (such as "write timeout" and "connection interruption") in migration logs and correlate them with performance indicators when they occur (for example, increased timeout probability during periods of high CPU usage).

[0348] Performance bottleneck location: Analyze the correlation between migration rate and storage engine metrics (for example, for every 10ms increase in disk I / O latency, the migration rate decreases by 20%) to identify the current core bottleneck factor (such as I / O latency and insufficient memory).

[0349] 2. Prioritize key factors

[0350] Rank key factors based on their impact (such as the probability of migration failure and their contribution to migration time) (for example, I / O latency > insufficient memory > network bandwidth).

[0351] Step 3: Input the performance prediction model to determine the migration acceptance

[0352] 1. Model input data preprocessing

[0353] Feature alignment: Convert the current performance test results (e.g., disk I / O = 120MB / s) and key factors (e.g., “I / O latency is the bottleneck”) into the input format required by the model (e.g., normalized numerical vectors or classification labels).

[0354] Contextual enhancement: Combined with the current migration task attributes (such as the data type is "graphic feature"), business tags are added to the input data to improve the model's adaptability to scenarios.

[0355] 2. Model reasoning and result output

[0356] Model call: Call the pre-trained performance prediction model through the API or local library, input pre-processed data, and obtain the predicted value of the single migration reception volume (such as "maximum single reception volume = 5000 records").

[0357] Result verification and correction: If the model output value exceeds the preset safety range (such as exceeding the hardware load limit), a conservative strategy will be triggered (such as using historical safety values) and the anomaly will be recorded for iterative model optimization.

[0358] Step 4: Dynamic execution and feedback of migration reception

[0359] 1. Dynamic allocation of migration tasks

[0360] Based on the amount of model output received, the migration scheduler splits the data to be migrated into batches that adapt to the current performance and assigns them to the execution queue.

[0361] 2. Real-time feedback on execution results

[0362] Record the deviation between the actual migration results (such as time consumption and success rate) and the predicted values ​​as training data for subsequent model optimization.

[0363] It should be noted that the embodiments of this specification have the following beneficial effects through the above content:

[0364] Real-time dynamic optimization of migration strategies: By combining current storage engine performance test results with real-time migration log analysis, key influencing factors (such as sudden load fluctuations and hardware status changes) are dynamically identified. This allows the decision on the amount of data to be transferred to adapt to the instantaneous performance of the database in real time, avoiding degradation in migration efficiency caused by historical models being out of sync with the current environment.

[0365] Enhanced fault tolerance for sudden performance anomalies: Identifies migration risk factors (such as transient I / O bottlenecks in the storage engine) based on current performance data, and proactively avoids unsafe migration volume settings through model prediction, reducing the probability of migration interruptions caused by hardware failures or external interference.

[0366] Improve the environmental awareness accuracy of migration volume decisions: Synchronize the current performance test results and the operating status in the real-time log (such as the number of concurrent tasks and cache occupancy) into the model to ensure that the migration volume accurately matches the real-time resource reserve of the database, optimizing the balance between resource utilization and migration efficiency.

[0367] Supports closed-loop iterative optimization of migration strategies: By continuously feeding current migration logs and performance data into the prediction model, a closed-loop optimization chain of "decision-execution-verification-update" is formed. This enables the model to autonomously evolve as the database operating environment changes (such as hardware upgrades and business expansion), maintaining long-term effectiveness.

[0368] Enhanced full-cycle controllability of the migration process: Before each migration task is launched, a real-time data-driven model re-evaluation mechanism ensures that the single migration volume is always within the optimal load range of the database, achieving full-cycle stability assurance from task initiation to completion.

[0369] Figure 2 A schematic structural diagram of a historical data migration device for an oil and gas field database provided for one or more embodiments of this specification includes: an image generation unit 201, an access unit 202, a text feature extraction unit 203, a graphic feature extraction unit 204, a first data volume determination unit 205, a migration data volume determination unit 206, a migration data volume determination unit 207, a migration data volume division unit 208, a second data volume determination unit 209, a data volume comparison unit 210 and a migration unit 211.

[0370] The image generation unit 201 digitizes paper logging documents in the oil and gas field database to generate RGB images;

[0371] Access unit 202, synchronously accessing the semi-structured electronic document in the oil and gas field database;

[0372] A text feature extraction unit 203 extracts well logging text features from the RGB image and the electronic document;

[0373] The graphic feature extraction unit 204 extracts the well logging graphic features in the RGB image and the electronic document based on the well logging curve symbols;

[0374] A first data volume determining unit 205 is configured to determine a first data volume of the well logging text feature and the well logging graphic feature;

[0375] A migration data amount determining unit 206 determines a single migration data amount based on a single migration reception amount of the new oil and gas field database;

[0376] A migration data amount determination unit 207 is configured to determine whether the first data amount exceeds the single migration data amount;

[0377] The migration data volume division unit 208 , if yes, divides the first data volume into multiple migration data volumes according to the single migration data volume;

[0378] A second data volume determining unit 209 is configured to determine a second data volume of the new oil and gas field database;

[0379] A data volume comparison unit 210 compares the first data volume with the second data volume to see whether they are consistent;

[0380] If they are consistent, the migration unit 211 migrates the well logging text features and the well logging graphic features to a pre-built new oil and gas field database according to a preset classification of the well logging data.

[0381] Figure 3 A schematic diagram of a historical data migration device for an oil and gas field database provided in one or more embodiments of this specification includes:

[0382] at least one processor; and,

[0383] a memory communicatively connected to the at least one processor; wherein,

[0384] The memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to:

[0385] Digitize paper logging documents in oil and gas field databases to generate RGB images;

[0386] Synchronously accessing semi-structured electronic documents in the oil and gas field database;

[0387] Extracting well logging text features from the RGB image and the electronic document;

[0388] For the well logging curve symbols, extract the well logging graphic features in the RGB image and the electronic document;

[0389] determining a first data volume of the well logging text feature and the well logging graphic feature;

[0390] Determining a single migration data volume based on a single migration reception volume of the new oil and gas field database;

[0391] Determining whether the first data volume exceeds the single migration data volume;

[0392] If so, dividing the first data volume into multiple migration data volumes according to the single migration data volume;

[0393] determining a second data volume of the new oil and gas field database;

[0394] comparing whether the first data volume is consistent with the second data volume;

[0395] If they are consistent, the well logging text features and the well logging graphic features are migrated to a pre-built new oil and gas field database according to the preset classification of the well logging data.

[0396] One or more embodiments of this specification provide a non-volatile computer storage medium storing computer-executable instructions. When executed by a computer, the computer-executable instructions can achieve:

[0397] Digitize paper logging documents in oil and gas field databases to generate RGB images;

[0398] Synchronously accessing semi-structured electronic documents in the oil and gas field database;

[0399] Extracting well logging text features from the RGB image and the electronic document;

[0400] For the well logging curve symbols, extract the well logging graphic features in the RGB image and the electronic document;

[0401] determining a first data volume of the well logging text feature and the well logging graphic feature;

[0402] Determining a single migration data volume based on a single migration reception volume of the new oil and gas field database;

[0403] Determining whether the first data volume exceeds the single migration data volume;

[0404] If so, dividing the first data volume into multiple migration data volumes according to the single migration data volume;

[0405] determining a second data volume of the new oil and gas field database;

[0406] comparing whether the first data volume is consistent with the second data volume;

[0407] If they are consistent, the well logging text features and the well logging graphic features are migrated to a pre-built new oil and gas field database according to the preset classification of the well logging data.

[0408] The various embodiments in this specification are described in a progressive manner. Similar portions between the various embodiments can be referenced to each other, and each embodiment focuses on the differences from the other embodiments. In particular, the device, apparatus, and non-volatile computer storage medium embodiments are generally similar to the method embodiments, so their descriptions are relatively simplified. For relevant details, refer to the descriptions of the method embodiments.

[0409] The various embodiments in this specification are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences from other embodiments. In particular, the device embodiments are generally similar to the method embodiments, so the description is relatively simple. For relevant parts, refer to the description of the method embodiments.

[0410] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0411] In the embodiments provided in this application, it should be understood that the disclosed devices / network equipment and methods can be implemented in other ways. For example, the device / network equipment embodiments described above are merely illustrative. For example, the division of the modules or units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0412] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0413] In addition, the functional units in the various embodiments of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above units may be implemented in the form of hardware or software.

[0414] If the integrated module / unit is implemented as a software functional unit and sold or used as a standalone product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application can implement all or part of the process steps in the above-mentioned method embodiments by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When executed by a processor, the computer program can implement the steps of each of the above-mentioned method embodiments. The computer program includes computer program code, which can be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal, and software distribution medium. It should be noted that the content of the computer-readable medium can be appropriately increased or decreased based on the requirements of legislation and patent practice in a jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media does not include electric carrier signals and telecommunication signals.

[0415] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.

Claims

1. A method for migrating historical data of an oil and gas field database, characterized in that: include: Digitize paper logging documents in oil and gas field databases to generate RGB images; Synchronously accessing semi-structured electronic documents in the oil and gas field database; Extracting well logging text features from the RGB image and the electronic document; For the well logging curve symbols, extract the well logging graphic features in the RGB image and the electronic document; determining a first data volume of the well logging text feature and the well logging graphic feature; Determine the amount of data to be migrated in a single session based on the amount of data received in a single migration of the new oil and gas field database; Determining whether the first data volume exceeds the single migration data volume; If so, dividing the first data volume into multiple migration data volumes according to the single migration data volume; determining a second data volume of the new oil and gas field database; comparing whether the first data volume is consistent with the second data volume; If they are consistent, the well logging text features and the well logging graphic features are migrated to a pre-built new oil and gas field database according to the preset classification of the well logging data; The determining of the single migration data volume based on the single migration reception volume of the new oil and gas field database includes: Determining a single migration data volume based on a single migration reception volume of the new oil and gas field database and a ratio of the well logging text features to the well logging graphic features; Before determining the single migration data volume based on the single migration reception volume of the new oil and gas field database and the ratio of the well logging text features to the well logging graphic features, the method further includes: Through the historical storage engine performance test of the new oil and gas field database, a migration volume prediction model is established in combination with historical migration logs to determine the single migration reception volume; The historical migration log includes the amount of data migrated each time, the migration time, and the performance indicators of the storage engine; The historical storage engine performance test of the new oil and gas field database is conducted, and a migration volume prediction model is established in combination with historical migration logs to determine the single migration reception volume, including: Performing a historical storage engine performance test on the new oil and gas field database to obtain historical performance test results and the historical migration log; Analyze the historical migration logs to identify historical key factors that affect migration performance; Establishing a performance prediction model based on the historical performance test results and the historical key factors, wherein the performance prediction model is a machine learning model; Determining the single migration reception amount based on the performance prediction model; The determining the single migration reception amount based on the performance prediction model includes: Performing a performance test on the current storage engine of the new oil and gas field database to obtain current performance test results and current migration logs; Analyze the current migration log to identify current key factors affecting migration performance; The current performance test result and the current key factors are input into the performance prediction model to determine the single migration reception quantity.

2. A historical data migration device for an oil and gas field database, characterized in that: include: An image generation unit digitizes paper logging documents in the oil and gas field database to generate RGB images; An access unit for synchronously accessing a semi-structured electronic document in the oil and gas field database; a text feature extraction unit for extracting features of well logging text in the RGB image and the electronic document; A graphic feature extraction unit extracts the well logging graphic features in the RGB image and the electronic document based on the well logging curve symbols; a first data volume determining unit, configured to determine a first data volume of the well logging text feature and the well logging graphic feature; A migration data volume determination unit determines a single migration data volume based on a single migration reception volume of a new oil and gas field database; a migration data volume determination unit, configured to determine whether the first data volume exceeds the single migration data volume; a migration data volume division unit, if yes, dividing the first data volume into multiple migration data volumes according to the single migration data volume; a second data volume determining unit, configured to determine a second data volume of the new oil and gas field database; a data volume comparison unit, for comparing whether the first data volume is consistent with the second data volume; Migration unit, if consistent, then according to the preset classification of the well logging data, the well logging text features and the well logging graphic features are migrated to a pre-built new oil and gas field database; The determining of the single migration data volume based on the single migration reception volume of the new oil and gas field database includes: Determining a single migration data volume based on a single migration reception volume of the new oil and gas field database and a ratio of the well logging text features to the well logging graphic features; Before determining the single migration data volume based on the single migration reception volume of the new oil and gas field database and the ratio of the well logging text features to the well logging graphic features, the method further includes: Through the historical storage engine performance test of the new oil and gas field database, a migration volume prediction model is established in combination with historical migration logs to determine the single migration reception volume; The historical migration log includes the amount of data migrated each time, the migration time, and the performance indicators of the storage engine; The historical storage engine performance test of the new oil and gas field database is conducted, and a migration volume prediction model is established in combination with historical migration logs to determine the single migration reception volume, including: Performing a historical storage engine performance test on the new oil and gas field database to obtain historical performance test results and the historical migration log; Analyze the historical migration logs to identify historical key factors that affect migration performance; Establishing a performance prediction model based on the historical performance test results and the historical key factors, wherein the performance prediction model is a machine learning model; Determining the single migration reception amount based on the performance prediction model; The determining the single migration reception amount based on the performance prediction model includes: Performing a performance test on the current storage engine of the new oil and gas field database to obtain current performance test results and current migration logs; Analyze the current migration log to identify current key factors affecting migration performance; The current performance test result and the current key factors are input into the performance prediction model to determine the single migration reception quantity.

3. A historical data migration device for an oil and gas field database, characterized in that: include: at least one processor; as well as, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to: Digitize paper logging documents in oil and gas field databases to generate RGB images; Synchronously accessing semi-structured electronic documents in the oil and gas field database; Extracting well logging text features from the RGB image and the electronic document; For the well logging curve symbols, extract the well logging graphic features in the RGB image and the electronic document; determining a first data volume of the well logging text feature and the well logging graphic feature; Determine the amount of data to be migrated in a single session based on the amount of data received in a single migration of the new oil and gas field database; Determining whether the first data volume exceeds the single migration data volume; If so, dividing the first data volume into multiple migration data volumes according to the single migration data volume; determining a second data volume of the new oil and gas field database; comparing whether the first data volume is consistent with the second data volume; If they are consistent, the well logging text features and the well logging graphic features are migrated to a pre-built new oil and gas field database according to the preset classification of the well logging data; The determining of the single migration data volume based on the single migration reception volume of the new oil and gas field database includes: Determining a single migration data volume based on a single migration reception volume of the new oil and gas field database and a ratio of the well logging text features to the well logging graphic features; Before determining the single migration data volume based on the single migration reception volume of the new oil and gas field database and the ratio of the well logging text features to the well logging graphic features, the method further includes: Through the historical storage engine performance test of the new oil and gas field database, a migration volume prediction model is established in combination with historical migration logs to determine the single migration reception volume; The historical migration log includes the amount of data migrated each time, the migration time, and the performance indicators of the storage engine; The historical storage engine performance test of the new oil and gas field database is conducted, and a migration volume prediction model is established in combination with historical migration logs to determine the single migration reception volume, including: Performing a historical storage engine performance test on the new oil and gas field database to obtain historical performance test results and the historical migration log; Analyze the historical migration logs to identify historical key factors that affect migration performance; Establishing a performance prediction model based on the historical performance test results and the historical key factors, wherein the performance prediction model is a machine learning model; Determining the single migration reception amount based on the performance prediction model; The determining the single migration reception amount based on the performance prediction model includes: Performing a performance test on the current storage engine of the new oil and gas field database to obtain current performance test results and current migration logs; Analyze the current migration log to identify current key factors affecting migration performance; The current performance test result and the current key factors are input into the performance prediction model to determine the single migration reception quantity.

4. A non-volatile computer storage medium, characterized in that The computer-executable instructions are stored, and when the computer-executable instructions are executed by a computer, they can achieve: Digitize paper logging documents in oil and gas field databases to generate RGB images; Synchronously accessing semi-structured electronic documents in the oil and gas field database; Extracting well logging text features from the RGB image and the electronic document; For the well logging curve symbols, extract the well logging graphic features in the RGB image and the electronic document; Migrating the well logging text features and the well logging graphic features to a pre-built new oil and gas field database according to preset classifications of the well logging data; determining a first data volume of the well logging text feature and the well logging graphic feature; Determine the amount of data to be migrated in a single session based on the amount of data received in a single migration of the new oil and gas field database; Determining whether the first data volume exceeds the single migration data volume; If so, dividing the first data volume into multiple migration data volumes according to the single migration data volume; determining a second data volume of the new oil and gas field database; comparing whether the first data volume is consistent with the second data volume; If they are consistent, the well logging text features and the well logging graphic features are migrated to a pre-built new oil and gas field database according to the preset classification of the well logging data; The determining of the single migration data volume based on the single migration reception volume of the new oil and gas field database includes: Determining a single migration data volume based on a single migration reception volume of the new oil and gas field database and a ratio of the well logging text features to the well logging graphic features; Before determining the single migration data volume based on the single migration reception volume of the new oil and gas field database and the ratio of the well logging text features to the well logging graphic features, the method further includes: Through the historical storage engine performance test of the new oil and gas field database, a migration volume prediction model is established in combination with historical migration logs to determine the single migration reception volume; The historical migration log includes the amount of data migrated each time, the migration time, and the performance indicators of the storage engine; The historical storage engine performance test of the new oil and gas field database is conducted, and a migration volume prediction model is established in combination with historical migration logs to determine the single migration reception volume, including: Performing a historical storage engine performance test on the new oil and gas field database to obtain historical performance test results and the historical migration log; Analyze the historical migration logs to identify historical key factors that affect migration performance; Establishing a performance prediction model based on the historical performance test results and the historical key factors, wherein the performance prediction model is a machine learning model; Determining the single migration reception amount based on the performance prediction model; The determining the single migration reception amount based on the performance prediction model includes: Performing a performance test on the current storage engine of the new oil and gas field database to obtain current performance test results and current migration logs; Analyze the current migration log to identify current key factors affecting migration performance; The current performance test result and the current key factors are input into the performance prediction model to determine the single migration reception quantity.

Citation Information

Patent Citations

  • Data migration method and device, electronic equipment and computer readable storage medium

    CN116737693A

  • Historical data migration method and system

    CN116910022A

  • OCR (Optical Character Recognition) method for production logging-oriented archival data

    CN118587719A

  • Data migration method and device, equipment and medium

    CN119690939A