Historical data migration method and device of oil and gas field database, equipment and medium

By deeply digitizing the paper logging documents in the oil and gas field database, combined with the collaborative integration of multi-source data and adaptive graph feature extraction, the problems of insufficient data retrieval, lack of multi-source data synergy and low graph feature recognition accuracy in the existing technology are solved, efficient data migration and business correlation are achieved, and in-depth analysis and cross-well data correlation applications are supported.

CN120216461AActive Publication Date: 2025-06-27DESHI ENERGY TECH GRP CO LTD

Patent Information

Application Number
CN202510694840.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-06-27
Estimated Expiration
2045-05-28

AI Technical Summary

Technical Problem

The existing technology has insufficient digital processing depth for paper logging documents in oil and gas field databases, and cannot effectively identify the logging text features, resulting in data being unable to be efficiently retrieved and utilized; at the same time, the coordination of multi-source data is missing, the graphical feature recognition accuracy is low, and the data classification and migration logic is weak, forming information islands, affecting the in-depth analysis of data and the application of data cross-well area data correlation.

Method used

By digitizing the paper logging documents in the oil and gas field database, RGB images are generated, and combined with intelligent extraction of logging text features, semantic analysis is achieved; at the same time, semi-structured electronic documents are synchronized to extract text features uniformly, breaking the isolation of paper and electronic document processing processes; adopting adaptive graphic feature extraction technology to improve the intelligent identification accuracy of logging curve symbols; based on the needs of oil and gas exploration business, logging data classification is preset, and a hierarchical storage structure is built to realize the coordinated integration and migration of text and graphic features.

Benefits of technology

Enhance the deep digitalization capabilities of paper documents, realize semantic analysis of logging data, and improve the retrievalability and utilization of data; realize the coordinated integration of multi-source heterogeneous data to solve the problem of information islands; improve the intelligent identification accuracy of logging curve symbols, reduce manual labeling dependence; build a business-oriented classification and migration system, strengthen business correlation between data, and support cross-well area data comparison and exploration decision analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216461A_ABST
    Figure CN120216461A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of oil and gas field databases, and discloses a historical data migration method, device and equipment of an oil and gas field database, and a medium, and the method comprises the steps: carrying out the digital processing of a paper logging document in the oil and gas field database, and generating an RGB image; synchronously accessing a semi-structured electronic document in an oil and gas field database; extracting logging character features in the RGB image and the electronic document; aiming at the logging curve symbol, extracting logging graphic features in the RGB image and the electronic document; determining a first data volume of the logging character features and the logging graphic features; determining a single migration data volume; judging whether the first data volume exceeds the single migration data volume or not; if yes, dividing the first data volume into multiple migration data volumes; determining a second data volume of the new oil and gas field database; comparing whether the first data volume is consistent with the second data volume; and if the logging character features and the logging graphic features are consistent, migrating the logging character features and the logging graphic features to a pre-constructed new oil and gas field database according to the preset classification of the logging data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of oil and gas field databases, and in particular to a method, device, equipment and medium for migrating historical data of an oil and gas field database. Background Art

[0002] In the field of oil and gas field exploration and development, logging data, as the core basis for reservoir evaluation and production decision-making, is usually stored in a traditional database in the form of paper documents and semi-structured electronic documents (such as PDF and Excel spreadsheets) for a long time. With the demand for digital upgrading, when migrating historical data to a new database, structured storage and multi-modal feature fusion need to be achieved.

[0003] The current technical means mainly include: 1) Scanning paper logging documents to generate electronic images; 2) Directly importing electronic documents through a database interface; 3) Extracting some text data based on keyword matching; 4) Manually annotating the coordinates of logging curve symbols to achieve graphic data migration.

[0004] However, the existing technologies have significant defects: Insufficient digital depth of paper documents: Traditional scanning only generates low-structured RGB images, without intelligent recognition and semantic parsing of logging text features (such as well numbers and formation parameters), resulting in data that cannot be efficiently retrieved and utilized by the new database; Lack of multi-source data collaboration: The processing flows of paper documents and electronic documents are independent of each other, and there is a lack of unified feature extraction rules for non-standard logging parameters (such as annotation text and segmented data) in semi-structured electronic documents, forming information islands; Low accuracy of graphic feature recognition: The extraction of logging curve symbols relies on manual interpretation or fixed template matching, and cannot adapt to the layout differences and symbol deformations (such as curve breakpoints and axis offsets) of different documents, resulting in distorted migrated data; Weak logic for data classification and migration: A logging data classification system matching the oil and gas exploration business has not been established, and the migrated data lacks a hierarchical storage structure, restricting in-depth analysis and cross-well area data correlation applications. Summary of the Invention

[0005] One or more embodiments of this specification provide a method, device, equipment and medium for migrating historical data of an oil and gas field database to solve the technical problems raised in the background art.

[0006] One or more embodiments of this specification adopt the following technical solutions: A method for migrating historical data of an oil and gas field database provided by one or more embodiments of this specification, the method includes: Digitally processing paper logging documents in an oil and gas field database to generate RGB images; Synchronously access the semi-structured electronic documents in the oil and gas field database; Extract the well logging text features in the RGB image and the electronic document; For well logging curve symbols, extract the well logging graphic features in the RGB image and the electronic document; Determine the first data volume of the well logging text features and the well logging graphic features; Based on the single migration reception volume of the new oil and gas field database, determine the single migration data volume; Judge whether the first data volume exceeds the single migration data volume; If so, divide the first data volume into multiple migration data volumes according to the single migration data volume; Determine the second data volume of the new oil and gas field database; Compare whether the first data volume is consistent with the second data volume; If they are consistent, migrate the well logging text features and the well logging graphic features to the pre-constructed new oil and gas field database according to the preset classification of well logging data.

[0007] It should be noted that the embodiments of this specification have the following beneficial effects through the above content: Enhance the deep digitization ability of paper documents: Through the digital processing of paper well logging documents and the generation of RGB images, combined with the intelligent extraction of well logging text features, break through the limitation of traditional scanning that only generates low-structured images, realize semantic-level parsing (such as well numbers, formation parameters), and significantly improve the retrievability and utilization rate of paper data in the new database.

[0008] Realize the collaborative integration of multi-source heterogeneous data: Synchronously access paper documents and semi-structured electronic documents (such as PDF, Excel), and extract text features (including non-annotated explanations, segmented parameters) through unified rules, break the isolation of the processing processes of paper and electronic documents, and solve the "information island" problem caused by the dispersion of multi-source data features.

[0009] Improve the intelligent recognition accuracy of well logging curve symbols: For well logging curve symbols, adopt adaptive graphic feature extraction technology, be compatible with symbol deformations of different format documents (such as curve breakpoints, axis offsets), reduce the dependence on manual annotation, and reduce the risk of graphic data migration distortion.

[0010] Construct a business-oriented classification and migration system: Based on the business requirements of oil and gas exploration, preset the classification of well logging data, migrate the text and graphic features to the new database according to the hierarchical storage structure, strengthen the business relevance between data, and support cross-well area data comparison and exploration decision-making analysis.

[0011] Optimize the efficiency of the data migration process: By automating the extraction and classification migration logic of text and graphic features, replace traditional manual annotation and template matching operations, reduce manual intervention links, and shorten the historical data migration cycle.

[0012] Strengthen the business adaptability of the new database: The migrated data has both structured text and high-precision graphic features, meeting the requirements of the integrated analysis of multi-modal data (such as parameter text, curve symbols) in oil and gas field exploration and development, and providing a high-quality data foundation for reservoir modeling and production optimization.

[0013] Ensure the integrity of data migration: By comparing the data volumes before and after migration (the first data volume and the second data volume), verify whether the logging text features and graphic features are completely migrated to the new database, avoid data loss caused by omissions during feature extraction or transmission, and improve the reliability of data migration.

[0014] Enhance the fault tolerance of the data migration process: Automatically trigger the consistency verification mechanism after the data migration is completed. If the data volumes are inconsistent, re-trigger the migration process to reduce the risk of data corruption caused by external factors such as network interruption and storage anomalies, and ensure the robustness of the migration process.

[0015] Reduce the cost of manual verification: Use an automated data volume comparison mechanism to replace traditional manual sampling inspections, greatly reduce the manual review workload after data migration, and improve the overall migration efficiency.

[0016] Enhance the credibility of the migrated data: Through data volume consistency verification, ensure that the logging text and graphic features stored in the new database strictly correspond to the original data, avoid subsequent analysis errors caused by data misalignment or redundancy, and provide a high-credibility data foundation for oil and gas field exploration operations.

[0017] Optimize the closed-loop management of the migration process: Embed the data volume consistency verification into the final link of the migration process to form a closed-loop management logic of "extraction - migration - verification", realizing full-link controllability and traceability of the migration process.

[0018] Improve the stability of large-scale data migration: By dynamically matching the single-migration reception volume of the new database, avoid system overload or transmission interruption caused by exceeding the single-migration data volume limit, and ensure the continuity and stability of migration tasks in high-load scenarios.

[0019] Optimize the resource allocation of the target database: Divide the migration batches according to the real-time reception capacity of the new database to prevent the instantaneous occupation of memory, storage, or bandwidth resources, and ensure the normal business operation of the database during migration without interference.

[0020] Enhance the controllability of the migration process: By predicting whether the data volume exceeds the limit and automatically triggering batch-by-batch migration, fine-grained control of the migration task granularity is achieved, facilitating real-time monitoring and intervention in abnormal batches by operation and maintenance personnel.

[0021] Reduce the overall risk of migration failure: Split large-scale data into multiple subtasks that adapt to the bearing capacity of the target system. Even if a batch of migration fails, only that batch needs to be retried instead of the entire volume of data, reducing the data rollback cost and time loss.

[0022] Adapt to the performance differences of heterogeneous databases: Dynamically adjust the amount of data migrated in a single time according to the hardware configuration and processing capacity of the new oil and gas field database, avoiding bottlenecks in migration efficiency caused by performance mismatches between the source database and the target database.

[0023] Furthermore, determining the amount of data migrated in a single time based on the single-time reception volume of the new oil and gas field database includes: Determine the amount of data migrated in a single time based on the single-time reception volume of the new oil and gas field database and the ratio of the logging text features to the logging graphic features.

[0024] It should be noted that through the above content, the embodiments of this specification have the following beneficial effects: Optimize the dynamic allocation of resources for heterogeneous data: By combining the ratio of logging text features to graphic features (such as the data volume ratio of text parameters to curve symbols), dynamically adjust the amount of data migrated in a single time, avoiding the imbalance in the allocation of database computing resources (such as memory, bandwidth) caused by excessive migration of a certain type of data, and improving the resource utilization rate during the migration process.

[0025] Improve the collaborative processing efficiency of data migration: Split migration batches based on the proportional relationship between text and graphic features, enabling the new database to evenly process text parsing and graphic rendering tasks during a single reception, reducing delays caused by data type processing conflicts, and improving the overall migration efficiency.

[0026] Reduce the risk of exceptions caused by data type mismatches: For the processing differences between text and graphic features (such as semantic analysis required for text parsing and coordinate mapping required for graphics), adapt the single-time migration volume proportionally to prevent parsing errors or storage exceptions caused by over-quantity of a single data type, and enhance the stability of the migration process.

[0027] Adapt to the business processing requirements of multi-modal data: Through proportion-aware migration volume division, ensure that the new database retains the original relevance of text and graphic features during reception (such as synchronous migration of parameter text and corresponding curve symbols), supporting subsequent fusion analysis and application of multi-modal data.

[0028] Enhance the flexibility of migration for different data structures: Dynamically adjust the migration strategy according to the differences in the ratio of text to graphics in logging data in different oil and gas field projects (for example, old wells mainly use paper drawings, and new wells mainly use electronic data), and improve the compatibility of the method with multiple scenarios.

[0029] Further, before determining the amount of data migrated in a single time based on the amount of data received in a single migration of the new oil and gas field database and the ratio of the logging text features to the logging graphic features, the method further includes: Through the performance test of the historical storage engine of the new oil and gas field database, establish a migration volume prediction model in combination with the historical migration log to determine the amount of data received in a single migration.

[0030] It should be noted that the embodiments of this specification have the following beneficial effects through the above content: Improve the scientificity of the decision-making on the amount of data migrated in a single time: Through the performance test of the historical storage engine and the modeling of the migration log, break through the traditional empirical threshold setting method, and make the determination of the amount of data received in a single migration more in line with the actual hardware performance of the new database (such as storage I / O, concurrent processing ability), and avoid the risk of overloading the migration amount or resource idleness caused by manual experience.

[0031] Enhance the dynamic adaptability of the migration strategy: The prediction model built based on historical data can perceive the performance fluctuations of the database (such as the increase in load during peak hours), dynamically adjust the amount of data received in a single migration, adapt to the differences in the system state at different times, and ensure the real-time matching of the migration task and the database operating environment.

[0032] Reduce the trial-and-error cost of migration configuration: Automatically recommend the amount of data migrated in a single time through the model, replace the traditional manual parameter adjustment or multiple exploratory migrations, reduce the risk of migration failure or performance degradation caused by unreasonable configuration, and shorten the migration debugging cycle.

[0033] Precipitate the value of historical data assets: Convert the historical migration log and performance test data into the training basis of the prediction model, realize the continuous optimization feedback of data assets on the migration strategy, and improve the generalization ability of the method for migration scenarios of similar databases.

[0034] Ensure the system stability under high-load scenarios: Combine the performance test results and model prediction to accurately control the amount of data migrated in a single time not to exceed the bearing limit of the database, and prevent service response delays or outages caused by sudden data write pressure.

[0035] Further, the historical migration log includes the amount of data migrated each time, the migration time, and the performance metrics of the storage engine; The step of establishing a migration volume prediction model through the performance test of the historical storage engine of the new oil and gas field database and combining the historical migration log to determine the amount of data received in a single migration includes: Through the performance test of the historical storage engine of the new oil and gas field database, the historical performance test results and the historical migration log are obtained; Analyze the historical migration log to identify the historical key factors affecting migration performance; Based on the historical performance test results and the historical key factors, establish a performance prediction model, and the performance prediction model is a machine learning model; Based on the performance prediction model, determine the single migration reception volume.

[0036] It should be noted that the embodiments of this specification have the following beneficial effects through the above content: Improve the accuracy and adaptability of migration volume prediction: By analyzing multi-dimensional data (data volume, time, performance indicators) in the historical migration log and combining with a machine learning model to capture historical key factors (such as storage engine I / O bottlenecks, concurrent load thresholds), dynamically generate a migration reception volume that matches the current database performance, overcoming the prediction deviation of manual experience or fixed rules.

[0037] Enhance the model's ability to analyze complex performance correlations: Based on machine learning, build a performance prediction model that can automatically learn the non-linear relationship between historical performance test results and migration key factors (such as storage latency caused by a sharp increase in data volume), accurately quantify the influence weight of storage engine performance on the migration volume, and achieve adaptive optimization of migration strategies.

[0038] Reduce the potential risk of migration performance fluctuations: By identifying historical key factors (such as a sharp drop in performance triggered by a specific data volume threshold), anticipate and avoid migration volume configurations that may cause database instability, and fundamentally reduce the occurrence probability of storage engine overload or response timeout during the migration process.

[0039] Achieve continuous self-optimization of migration strategies: The machine learning model can continuously update parameters and iterate and optimize with the accumulation of historical migration logs and performance test data, enabling the decision-making logic of the single migration reception volume to evolve autonomously with database hardware upgrades or business scenario changes, maintaining long-term effectiveness.

[0040] Strengthen the interpretability and controllability of the migration process: By explicitly identifying the key factors affecting migration performance (such as the correlation between migration time and storage engine CPU occupancy), provide a decision-making basis for operation and maintenance personnel, facilitate targeted adjustment of database configurations or migration plans, and enhance the transparency and intervenability of the migration process.

[0041] Furthermore, the determining of the single migration reception volume based on the performance prediction model includes: Through the performance test of the current storage engine of the new oil and gas field database, obtain the current performance test results and the current migration log; Analyze the current migration log to identify the current key factors affecting migration performance; Input the current performance test results and the current key factors into the performance prediction model to determine the single migration reception volume.

[0042] It should be noted that the embodiments of this specification have the following beneficial effects through the above content: Realize real-time dynamic optimization of the migration strategy: By combining the current storage engine performance test results with real-time migration log analysis, dynamically identify the current key influencing factors (such as sudden load fluctuations, hardware state changes), so that the decision of the single migration reception volume can be adapted to the instantaneous performance state of the database in real time, and avoid the decline in migration efficiency caused by the disconnection between the historical model and the current environment.

[0043] Enhance the fault tolerance ability for sudden performance anomalies: Identify migration risk factors (such as instantaneous I / O bottlenecks in the storage engine) based on the current performance data, and actively avoid unsafe migration volume settings through model prediction, reducing the probability of migration interruption caused by hardware failures or external interferences.

[0044] Improve the environmental perception accuracy of the migration volume decision: Synchronously input the current performance test results and the running status in the real-time log (such as the number of concurrent tasks, cache occupancy rate) into the model to ensure that the migration reception volume accurately matches the real-time resource margin of the database, and optimize the balance between resource utilization and migration efficiency.

[0045] Support the closed-loop iterative optimization of the migration strategy: By continuously feeding back the current migration log and performance data to the prediction model, form a closed-loop optimization link of "decision - execution - verification - update", so that the model can evolve autonomously following the changes in the database running environment (such as hardware upgrades, business expansion), and maintain long-term effectiveness.

[0046] Strengthen the full-cycle controllability of the migration process: Before each migration task is started, through a real-time data-driven model re-evaluation mechanism, ensure that the single migration volume is always within the optimal load range of the database, and achieve full-cycle stability guarantee from task start to completion.

[0047] A historical data migration device for an oil and gas field database provided by one or more embodiments of this specification includes: An image generation unit that digitally processes paper logging documents in the oil and gas field database to generate RGB images; An access unit that synchronously accesses semi-structured electronic documents in the oil and gas field database; A text feature extraction unit that extracts logging text features from the RGB images and the electronic documents; A graphic feature extraction unit extracts logging graphic features from the RGB image and the electronic document for logging curve symbols; A first data volume determination unit determines a first data volume of the logging text features and the logging graphic features; A migration data volume determination unit determines a single migration data volume based on the single migration reception volume of the new oil and gas field database; A migration data volume determination unit determines whether the first data volume exceeds the single migration data volume; A migration data volume division unit, if so, divides the first data volume into multiple migration data volumes according to the single migration data volume; A second data volume determination unit determines a second data volume of the new oil and gas field database; A data volume comparison unit compares whether the first data volume is consistent with the second data volume; A migration unit, if consistent, migrates the logging text features and the logging graphic features to a pre-constructed new oil and gas field database according to a preset classification of logging data.

[0048] A historical data migration device for an oil and gas field database provided by one or more embodiments of this specification includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to: Digitally process a paper logging document in an oil and gas field database to generate an RGB image; Synchronously access a semi-structured electronic document in the oil and gas field database; Extract logging text features from the RGB image and the electronic document; For logging curve symbols, extract logging graphic features from the RGB image and the electronic document; Determine a first data volume of the logging text features and the logging graphic features; Based on the single migration reception volume of the new oil and gas field database, determine a single migration data volume; Determine whether the first data volume exceeds the single migration data volume; If so, divide the first data volume into multiple migration data volumes according to the single migration data volume; Determine a second data volume of the new oil and gas field database; Compare whether the first data volume is consistent with the second data volume; If they are consistent, transfer the well logging text features and the well logging graphic features to a pre-constructed new oil and gas field database according to the preset classification of well logging data.

[0049] A non-volatile computer storage medium provided by one or more embodiments of this specification stores computer-executable instructions, and when the computer-executable instructions are executed by a computer, they can implement: Digitally process the paper well logging documents in the oil and gas field database to generate RGB images; Synchronously access the semi-structured electronic documents in the oil and gas field database; Extract the well logging text features from the RGB images and the electronic documents; For well logging curve symbols, extract the well logging graphic features from the RGB images and the electronic documents; Determine the first data volume of the well logging text features and the well logging graphic features; Based on the single transfer reception volume of the new oil and gas field database, determine the single transfer data volume; Judge whether the first data volume exceeds the single transfer data volume; If so, divide the first data volume into multiple transfer data volumes according to the single transfer data volume; Determine the second data volume of the new oil and gas field database; Compare whether the first data volume is consistent with the second data volume; If they are consistent, transfer the well logging text features and the well logging graphic features to a pre-constructed new oil and gas field database according to the preset classification of well logging data.

[0050] The above at least one technical solution adopted by the embodiments of this specification can achieve the following beneficial effects: Enhance the deep digitalization ability of paper documents: Through the digital processing of paper well logging documents and the generation of RGB images, combined with the intelligent extraction of well logging text features, break through the limitation of traditional scanning that only generates low-structured images, realize semantic-level parsing (such as well numbers, formation parameters), and significantly improve the retrievability and utilization rate of paper data in the new database.

[0051] Realize the collaborative integration of multi-source heterogeneous data: Synchronously access paper documents and semi-structured electronic documents (such as PDF, Excel), and extract text features (including non-annotated explanations, segmented parameters) through unified rules, break the isolation of the processing processes of paper and electronic documents, and solve the "information island" problem caused by the dispersion of multi-source data features.

[0052] Improve the intelligent recognition accuracy of logging curve symbols: For logging curve symbols, adopt adaptive graphic feature extraction technology, which is compatible with symbol deformations (such as curve breakpoints and axis offsets) in different format documents, reduce the dependence on manual annotation, and reduce the risk of graphic data migration distortion.

[0053] Build a business-oriented classification and migration system: Preset the classification of logging data based on the requirements of oil and gas exploration business, migrate text and graphic features to a new database according to a hierarchical storage structure, strengthen the business relevance between data, and support cross-well area data comparison and exploration decision-making analysis.

[0054] Optimize the efficiency of the data migration process: Through the automated extraction of text and graphic features and the classification and migration logic, replace the traditional manual annotation and template matching operations, reduce the manual intervention links, and shorten the historical data migration cycle.

[0055] Strengthen the business adaptability of the new database: The migrated data has both structured text and high-precision graphic features, meets the requirements of integrated analysis of multi-modal data (such as parameter text and curve symbols) in oil and gas field exploration and development, and provides a high-quality data basis for reservoir modeling and production optimization. Description of the Drawings

[0056] In order to more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in this specification. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. In the drawings: Figure 1 It is a schematic flowchart of a method for migrating historical data of an oil and gas field database provided by one or more embodiments of this specification; Figure 2 It is a schematic structural diagram of a device for migrating historical data of an oil and gas field database provided by one or more embodiments of this specification; Figure 3 It is a schematic structural diagram of a device for migrating historical data of an oil and gas field database provided by one or more embodiments of this specification. Detailed Embodiments

[0057] Embodiments of this specification provide a method, device, equipment and medium for migrating historical data of an oil and gas field database.

[0058] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on the embodiments of this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this specification.

[0059] Figure 1 It is a schematic flowchart of a method for migrating historical data of an oil and gas field database provided by one or more embodiments of this specification. This process can be executed by a historical data migration system of the oil and gas field database. Some input parameters or intermediate results in the process allow manual intervention and adjustment to help improve accuracy.

[0060] The method process steps of the embodiments of this specification are as follows: S101, digitize the paper logging documents in the oil and gas field database to generate RGB images.

[0061] In the embodiments of this specification, the following specific implementation solutions can be adopted: 1. Document scanning and preprocessing Use a high-resolution scanner to scan the paper logging documents to generate original RGB images.

[0062] Denoise, correct the tilt, and enhance the contrast of the scanned images to ensure that the images are clear and readable.

[0063] 2. Image standardized storage Store the processed RGB images in a temporary directory according to a unified naming rule (such as "well number_date_page number") to provide input for subsequent feature extraction.

[0064] S102, synchronously access the semi-structured electronic documents in the oil and gas field database.

[0065] In the embodiments of this specification, the following specific implementation solutions can be adopted: 1. Multi-format document access Batch read semi-structured electronic documents such as PDF and Excel through a database interface or a file system interface. Decrypt or verify the permissions for encrypted or permission-restricted documents before access.

[0066] 2. Format unification processing Convert documents in different formats to an intermediate data format (such as converting PDF to text and Excel to CSV) to eliminate the interference of format differences on feature extraction.

[0067] S103. Extract the logging text features in the RGB image and the electronic document.

[0068] In the embodiments of this specification, the following specific implementation solutions can be adopted: 1. Optical Character Recognition (OCR) for images Locate the text areas in the RGB image and use OCR technology to recognize the text content (such as well numbers, depth values, lithology descriptions).

[0069] 2. Parsing of electronic document text Perform structured parsing on the text content in the semi-structured electronic document to extract key fields (such as "porosity" and "permeability" parameters in a table).

[0070] 3. Semantic annotation and association Through Natural Language Processing (NLP) technology, annotate semantic labels (such as "formation name", "test conclusion") for the extracted text according to business logic.

[0071] S104. For logging curve symbols, extract the logging graphic features in the RGB image and the electronic document.

[0072] In the embodiments of this specification, the following specific implementation solutions can be adopted: 1. Recognition of curve symbols Perform image segmentation on the logging curve areas in the RGB image to recognize the coordinate points, line types, and scale information of the curve symbols. Extract the original coordinate data for the vector graphics (such as PDF embedded curve graphs) in the electronic document.

[0073] 2. Structuring of graphic features Convert the curve symbols into a standardized data format (such as converting a time-depth curve into a JSON array), and retain the topological relationships and attributes of the symbols.

[0074] S105. Determine the first data volume of the logging text features and the logging graphic features.

[0075] S106. Based on the single migration reception volume of the new oil and gas field database, determine the single migration data volume.

[0076] S107. Determine whether the first data volume exceeds the single migration data volume.

[0077] S108. If so, divide the first data volume into multiple migration data volumes according to the single migration data volume.

[0078] S109. Determine the second data volume of the new oil and gas field database.

[0079] S110, Compare whether the first data volume is consistent with the second data volume.

[0080] S111, If they are consistent, then transfer the well logging text features and the well logging graphic features to a pre-constructed new oil and gas field database according to the preset classification of well logging data.

[0081] In the embodiments of this specification, the following specific implementation solutions can be adopted: 1. Classification rule matching According to the preset classification of well logging data (such as stratifying by "well area", "data type", "acquisition time"), assign the text features and graphic features to the corresponding classification labels.

[0082] 2. Multi-modal data fusion and transfer Associate the text features (structured text) and the graphic features (standardized coordinate data) according to the classification labels, and batch write them into the specified storage module of the new database through an ETL tool.

[0083] 3. Verification of transfer results Sample and verify the written data in the new database to ensure the integrity and logical consistency of the text and graphic features.

[0084] It should be noted that through the above content, the embodiments of this specification have the following beneficial effects: Enhance the deep digitalization ability of paper documents: Through the digital processing of paper well logging documents and the generation of RGB images, combined with the intelligent extraction of well logging text features, break through the limitation of traditional scanning that only generates low-structured images, achieve semantic-level parsing (such as well numbers, formation parameters), and significantly improve the retrievability and utilization rate of paper data in the new database.

[0085] Realize the collaborative integration of multi-source heterogeneous data: Synchronously access paper documents and semi-structured electronic documents (such as PDF, Excel), and extract text features (including non-annotated explanations, segmented parameters) through unified rules, break the isolation of the processing processes of paper and electronic documents, and solve the "information island" problem caused by the dispersion of multi-source data features.

[0086] Improve the intelligent recognition accuracy of well logging curve symbols: For well logging curve symbols, adopt an adaptive graphic feature extraction technology, be compatible with the symbol deformation of different format documents (such as curve breakpoints, axis offsets), reduce the dependence on manual annotation, and reduce the risk of distortion in the transfer of graphic data.

[0087] Construct a business-oriented classification and transfer system: Based on the business requirements of oil and gas exploration, preset the classification of well logging data, transfer the text and graphic features to the new database according to a hierarchical storage structure, strengthen the business relevance between data, and support cross-well area data comparison and exploration decision-making analysis.

[0088] Optimize the efficiency of the data migration process: By automating the extraction and classification migration logic of text and graphic features, replace the traditional manual annotation and template matching operations, reduce the manual intervention links, and shorten the historical data migration cycle.

[0089] Strengthen the business adaptability of the new database: The migrated data has both structured text and high-precision graphic features, meeting the requirements of the integrated analysis of multi-modal data (such as parameter text, curve symbols) in oil and gas field exploration and development, and providing a high-quality data foundation for reservoir modeling and production optimization.

[0090] Specifically, for the embodiments of this specification, determine the first data volume of the well logging text features and the well logging graphic features; determine the second data volume of the new oil and gas field database; compare whether the first data volume is consistent with the second data volume; if they are consistent, complete the migration of the well logging text features and the well logging graphic features to the new oil and gas field database, which can be implemented through the following specific implementation plans: Step 1: Determine the first data volume of well logging text features and graphic features 1. Define the data volume statistics rules Text feature statistics: Traverse all the extracted well logging text features (such as well numbers, formation parameters), and count the total number of records and the number of fields according to the field types (text, numerical value).

[0091] Graphic feature statistics: Count the total number of well logging curve symbols (such as coordinate points, line types, scales) according to the graphic units (single curve, symbol group).

[0092] 2. Automatically calculate the data volume Develop a script or call the database statistics interface, traverse the temporary storage directory before migration, and summarize the total data volume (such as the number of records, the number of storage units) of text features and graphic features.

[0093] Step 2: Determine the second data volume of the new oil and gas field database 1. Statistic the data of the target database In the corresponding storage modules (such as structured text tables, graphic feature tables) of the new database, execute the predefined data volume query operation to count the number of records of the migrated text and graphic features.

[0094] 2. Real-time feedback of the data volume After the migration task is completed, automatically trigger the data volume statistics program to generate a data volume report after migration (such as "X text features are migrated, Y groups of graphic features are migrated").

[0095] Step 3: Compare whether the first data volume is consistent with the second data volume 1. Design the consistency verification logic Define the comparison rule: If the number of text features records and the number of graphic feature units are equal, it is determined to be consistent; otherwise, it is inconsistent.

[0096] 2. Automated comparison process Input the first data volume counted before migration and the second data volume after migration into the comparison program, output the verification result (such as "consistent" or "inconsistent"), and record the details of the differences (such as missing well numbers or curve numbers).

[0097] Step 4: If it is consistent, the migration is completed. 1. Migration completion confirmation mechanism When the consistency verification result is "consistent", automatically trigger the submission operation of the new database, officially write the temporarily stored migration data into the business table, and release the temporarily stored resources.

[0098] 2. Result notification and log recording Send a migration success notification to the operation and maintenance system, and at the same time record the first data volume, the second data volume and the verification result of this migration in the audit log for subsequent traceability.

[0099] It should be noted that the embodiments of this specification have the following beneficial effects through the above content: Ensure the integrity of data migration: By comparing the data volumes before and after migration (the first data volume and the second data volume), verify whether the logging text features and graphic features are completely migrated to the new database, avoid data loss caused by omissions during feature extraction or transmission, and improve the reliability of data migration.

[0100] Strengthen the fault tolerance ability of the data migration process: Automatically trigger the consistency verification mechanism after the data migration is completed. If the data volumes are inconsistent, re-trigger the migration process to reduce the risk of data damage caused by external factors such as network interruption and storage anomalies, and ensure the robustness of the migration process.

[0101] Reduce the cost of manual verification: Use an automated data volume comparison mechanism to replace the traditional manual sampling inspection, greatly reduce the manual review workload after data migration, and improve the overall migration efficiency.

[0102] Enhance the credibility of the migrated data: Through the data volume consistency verification, ensure that the logging text and graphic features stored in the new database strictly correspond to the original data, avoid subsequent analysis errors caused by data misalignment or redundancy, and provide a high-credibility data basis for oil and gas field exploration operations.

[0103] Optimize the closed-loop management of the migration process: Embed the data volume consistency verification into the final link of the migration process to form a closed-loop management logic of "extraction - migration - verification", and achieve full-link controllability and traceability of the migration process.

[0104] Specifically, before migrating the well logging text features and the well logging graphic features to the pre-constructed new oil and gas field database, the amount of data migrated in a single time can be determined based on the single-time migration reception capacity of the new oil and gas field database; it can be judged whether the first data volume exceeds the single-time migration data volume; if so, according to the single-time migration data volume, the first data volume can be divided into multiple migration data volumes, and the following specific implementation plans can be adopted: Step 1: Determine the single-time migration reception capacity of the new database 1. Dynamically evaluate the database load capacity Conduct real-time performance tests on the new oil and gas field database (such as concurrent write pressure tests, storage I / O throughput tests), and comprehensively evaluate the upper limit of the amount of data that can be safely received currently in combination with the load threshold records in the historical migration logs.

[0105] 2. Dynamic adjustment rules for the reception capacity Dynamically lower or raise the single-time migration reception capacity according to the real-time resource occupancy rate of the database (such as CPU, memory, remaining disk space) to ensure that the core business of the database is not affected during the migration process.

[0106] Step 2: Judge whether the first data volume exceeds the single-time migration reception capacity 1. Automated data volume comparison Automatically compare the total amount of data to be migrated (the first data volume) with the single-time migration reception capacity through scripts or migration tools: If the first data volume ≤ the single-time reception volume, trigger a one-time full-volume migration; If the first data volume > the single-time reception volume, trigger the batch migration logic.

[0107] 2. Threshold warning mechanism When the data volume exceeds the limit, send an alarm notification to the operation and maintenance system, and automatically freeze the migration task until manual confirmation or system resources are released.

[0108] Step 3: Divide the multiple migration data volumes 1. Data batching strategy Split according to the data type ratio: According to the data volume ratio of text features and graphic features (such as 70% for text and 30% for graphics), split the total data volume into multiple batches in proportion to ensure that the ratio of the two types of data in each batch is the same as the original data.

[0109] Group according to business relevance: Based on the business attributes of well logging data (such as by well number, acquisition time), divide the text and graphic features with strong relevance into the same batch to avoid the breakage of cross-batch data associations.

[0110] 2. Batch migration order scheduling Adopt a priority queue mechanism to preferentially migrate high-value data (such as recently collected data, key well data), and control the batch interval time through a task scheduler to prevent resource contention caused by centralized writing.

[0111] It should be noted that the embodiments of this specification have the following beneficial effects through the above content: Improve the stability of large-scale data migration: By dynamically matching the single migration reception volume of the new database, avoid system overload or transmission interruption caused by exceeding the single migration data volume limit, and ensure the continuity and stability of migration tasks in high-load scenarios.

[0112] Optimize the resource allocation of the target database: Divide the migration batches according to the real-time reception capacity of the new database to prevent the memory, storage, or bandwidth resources from being instantaneously filled, and ensure that the normal business operation of the database during the migration process is not interfered.

[0113] Enhance the controllability of the migration process: By predicting whether the data volume exceeds the limit and automatically triggering batch-by-batch migration, achieve fine-grained control of the migration task granularity, which is convenient for operation and maintenance personnel to monitor and intervene in abnormal batches in real time.

[0114] Reduce the overall risk of migration failure: Split the large-scale data into multiple subtasks that adapt to the bearing capacity of the target system. Even if a certain batch of migration fails, only this batch needs to be retried instead of the full amount of data, reducing the data rollback cost and time loss.

[0115] Adapt to the performance differences of heterogeneous databases: Dynamically adjust the single migration data volume according to the hardware configuration and processing capacity of the new oil and gas field database to avoid the migration efficiency bottleneck caused by the performance mismatch between the source database and the target database.

[0116] Furthermore, when determining the single migration data volume based on the single migration reception volume of the new oil and gas field database, the single migration data volume can be determined based on the single migration reception volume of the new oil and gas field database, as well as the ratio of the logging text features to the logging graphic features.

[0117] In the embodiments of this specification, the following specific implementation schemes can be adopted: Step 1: Dynamically evaluate the single migration reception volume of the new database 1. Real-time performance testing and historical data analysis Conduct a short-term stress test on the new database (such as simulating the write peak) to obtain the maximum stable reception volume (such as the number of records processed per second or the storage throughput) under the current hardware configuration.

[0118] Combine the successful batch data volume records in the historical migration logs to analyze the reception volume fluctuation range under different load scenarios (such as business peak / low valley).

[0119] 2. Dynamic Adjustment Mechanism for Reception Volume According to the real-time resource occupancy rate of the database (such as CPU utilization rate, remaining memory), dynamically adjust the upper limit of the reception volume for a single migration according to preset rules (for example, automatically lower the reception volume when the resource occupancy exceeds 70%).

[0120] Step 2: Statistic the ratio of logging text and graphic features 1. Calculation of data feature ratio Statistically count the data volume (such as the number of records, the number of storage units) of the extracted logging text features (such as well numbers, parameter tables) and graphic features (such as curve symbols, coordinate points) respectively, and calculate the ratio between the two (for example, text accounts for 60% and graphics account for 40%).

[0121] 2. Standardized storage of ratio Store the ratio information (such as text:graphics = 3:2) as metadata in the migration task configuration file for subsequent batch migration calls.

[0122] Step 3: Dynamically allocate the data volume for a single migration according to the ratio 1. Batch strategy adapted to the ratio Data volume splitting logic: Split the data according to the upper limit of the reception volume for a single migration and the ratio of text to graphics. For example, if the single reception volume is 1000 units and the ratio is 3:2, then a single migration includes 600 units of text and 400 units of graphics.

[0123] Maintaining business relevance: Ensure that the text features and corresponding graphic features of the same logging data are assigned to the same batch (such as the text description and curve symbols of Well A are migrated in the same batch) to avoid data association breakage.

[0124] 2. Automatic batch generation Develop a migration scheduling tool to extract text and graphic features from the data pool to be migrated according to the ratio, generate several complete batches, and mark the batch numbers and content lists.

[0125] Step 4: Migration execution and dynamic adjustment 1. Control of batch migration priority Data with high priority (such as recently active wells, key parameters) is migrated first, and the data volume allocation of subsequent batches is dynamically adjusted according to the ratio.

[0126] 2. Real-time feedback and elastic scaling Monitor the actual execution effect of each migration (such as time consumption, resource occupancy). If it is found that the graphic processing delay is relatively high, automatically reduce the ratio of graphic features or the single reception volume in subsequent batches.

[0127] It should be noted that through the above content, the embodiments of this specification have the following beneficial effects: Optimize the dynamic allocation of resources for heterogeneous data: By combining the ratio of well logging text features and graphic features (such as the data volume ratio of text parameters to curve symbols), dynamically adjust the amount of data migrated at one time, avoid the imbalance of database computing resources (such as memory, bandwidth) allocation caused by excessive migration of a certain type of data, and improve the resource utilization rate during the migration process.

[0128] Improve the collaborative processing efficiency of data migration: Based on the proportional relationship between text and graphic features, split the migration batches, so that the new database can evenly process text parsing and graphic rendering tasks when receiving at one time, reduce the latency caused by data type processing conflicts, and improve the overall migration efficiency.

[0129] Reduce the risk of exceptions caused by data type mismatches: For the processing differences between text and graphic features (such as semantic analysis required for text parsing and coordinate mapping required for graphics), adapt the amount of data migrated at one time according to the ratio, prevent parsing errors or storage exceptions caused by excessive amounts of a single data type, and enhance the stability of the migration process.

[0130] Adapt to the business processing requirements of multi-modal data: Through the proportional perception of the amount of data migrated, ensure that the original relevance of text and graphic features is retained when the new database receives (such as the synchronous migration of parameter text and corresponding curve symbols), and support the subsequent fusion analysis and application of multi-modal data.

[0131] Enhance the flexibility of migration for different data structures: For the differences in the ratio of well logging text and graphics in different oil and gas field projects (such as old wells mainly using paper maps and new wells mainly using electronic data), dynamically adjust the migration strategy to improve the compatibility of the method with multiple scenarios.

[0132] Furthermore, before determining the amount of data migrated at one time based on the one-time migration reception volume of the new oil and gas field database and the ratio of the well logging text features to the well logging graphic features, a migration volume prediction model can be established by combining the historical storage engine performance test of the new oil and gas field database and the historical migration log to determine the one-time migration reception volume.

[0133] In the embodiments of this specification, the following specific implementation methods can be adopted: Step 1: Historical storage engine performance test and log collection 1. Performance test scenario design Simulate different load scenarios (such as low / medium / high concurrent writes, mixed read and write operations) in the storage engine of the new database, and record the key performance indicators during the test (such as the number of transactions processed per second, disk I / O latency, memory occupancy rate).

[0134] Classify and store the test results, and mark the test time, load type, and corresponding performance thresholds (such as the maximum stable data reception volume).

[0135] 2. Historical Migration Log Integration Extract historical migration logs from the old migration system, including fields such as the amount of data migrated each time, the time taken, the success / failure status, and the error type, to build a structured log database.

[0136] Step 2: Identification of Key Performance Impact Factors 1. Data Relevance Analysis Align the historical performance test results with the migration logs by timestamp, and analyze the relevance between performance metrics (such as I / O latency) and migration results (such as failure rate).

[0137] Identify key factors through statistical methods (e.g., when the disk I / O latency exceeds 50ms, the migration failure probability increases by 80%).

[0138] 2. Feature Engineering Processing Derive features from the raw data in the logs (such as the amount of data and timestamp) to generate migration task features (such as the ratio of the amount of data per single migration to the total amount of data, and the number of concurrent tasks in the same time period).

[0139] Step 3: Construction of Migration Volume Prediction Model 1. Model Selection and Training Adopt supervised learning algorithms (such as random forest, gradient boosting tree), use historical performance test data as input features, and the maximum amount of data successfully migrated as labels to train a regression model.

[0140] Perform cross-validation on the model to ensure its generalization ability in various load scenarios.

[0141] 2. Model Deployment and Interface Encapsulation Package the trained model as an API service and integrate it into the migration scheduling system to support real-time reception of performance parameters and return recommended migration volumes.

[0142] Step 4: Dynamic Prediction of Migration Volume Based on the Current State 1. Real-time Performance Testing and Log Collection Before each migration task starts, conduct a lightweight performance test (such as a 5-second stress test) on the new database to obtain the I / O, CPU, and memory status of the current storage engine.

[0143] Synchronously collect real-time migration logs (such as the current number of concurrent tasks and the amount of data already migrated).

[0144] 2. Model Inference and Migration Volume Output Input the current performance data (such as disk I / O latency, memory occupancy rate) and real-time log features into the prediction model, and output the recommended value of the single migration reception volume adapted to the current state.

[0145] Step 5: Model Continuous Optimization and Feedback Mechanism 1. Online Learning and Update After each migration task is completed, compare the actual migration result (success / failure, time consumption) with the model prediction value, and automatically trigger fine-tuning of model parameters (such as incremental training).

[0146] 2. Abnormal Scenario Handling When the deviation between the model prediction value and the actual performance exceeds the threshold, trigger an alarm and switch to a backup rule engine (such as a conservative migration volume strategy based on historical average values), and record the reason for the anomaly for model iteration and repair at the same time.

[0147] It should be noted that through the above content, the embodiments of this specification have the following beneficial effects: Improve the scientificity of the decision-making on the single migration volume: Through the performance test of the historical storage engine and the modeling of migration logs, break through the traditional empirical threshold setting method, and make the determination of the single migration reception volume more in line with the actual hardware performance of the new database (such as storage I / O, concurrent processing ability), avoiding migration volume overload or resource idleness caused by manual experience.

[0148] Enhance the dynamic adaptability of migration strategies: The prediction model built based on historical data can perceive the performance fluctuations of the database (such as the increase in load during peak hours), dynamically adjust the single migration reception volume, adapt to the system state differences at different times, and ensure the real-time matching of migration tasks and the database operating environment.

[0149] Reduce the trial-and-error cost of migration configuration: Through the model's automatic recommendation of the single migration volume, replace the traditional manual parameter tuning or multiple exploratory migrations, reduce the risk of migration failure or performance degradation caused by unreasonable configuration, and shorten the migration debugging cycle.

[0150] Precipitate the value of historical data assets: Convert historical migration logs and performance test data into the training basis of the prediction model, realize the continuous optimization and feedback of data assets to migration strategies, and improve the generalization ability of the method for migration scenarios of similar databases.

[0151] Ensure system stability under high-load scenarios: Combine performance test results and model predictions to accurately control the single migration volume not to exceed the bearing limit of the database, and prevent service response delays or outages caused by sudden data write pressure.

[0152] Further, the historical migration log includes the amount of data migrated each time, the migration time, and the performance metrics of the storage engine; when establishing a migration volume prediction model through the historical storage engine performance test of the new oil and gas field database and determining the single - migration reception volume, the historical performance test results and the historical migration log can be obtained through the historical storage engine performance test of the new oil and gas field database; analyze the historical migration log to identify the historical key factors affecting the migration performance; based on the historical performance test results and the historical key factors, establish a performance prediction model, where the performance prediction model is a machine - learning model; based on the performance prediction model, determine the single - migration reception volume.

[0153] In the embodiments of this specification, the following specific implementation solutions can be adopted: Step 1: Integration of historical performance test and migration log 1. Data collection and storage Performance test data collection: Perform historical storage engine performance tests (such as I / O throughput, CPU peak value, memory occupancy rate under different loads) in the new oil and gas field database, and store the test results as a structured data table according to the time stamp.

[0154] Migration log normalization: Clean the historical migration log, extract key fields (migration time, data volume, time consumption, error code), and associate them with the performance test results according to the time stamp to construct a unified historical data set.

[0155] Step 2: Analysis of historical migration log and identification of key factors 1. Data exploration and pattern mining Anomaly detection: Through log analysis tools (such as Python Pandas or ELK Stack), count the records of migration failures or abnormal time consumption, and analyze the corresponding performance metrics (such as the failure rate significantly increases during high I / O latency).

[0156] Correlation analysis: Calculate the correlation between the amount of migrated data, time consumption, and performance metrics (such as CPU utilization rate, disk write rate), and filter out strongly correlated factors (for example: the correlation coefficient between disk I / O latency and migration time consumption > 0.8).

[0157] 2. Extraction and annotation of key factors Mark the factors that significantly affect the migration performance (such as "I / O latency > 100ms", "number of concurrent tasks > 5") as historical key factors and store them as feature labels.

[0158] Step 3: Construction and training of performance prediction model 1. Feature engineering and data set construction Input Feature Design: Based on historical key factors and performance test results, generate model input features (such as "I / O latency", "number of concurrent tasks", "data volume / maximum throughput ratio").

[0159] Label Definition: Use the maximum single data reception volume in historical successful migration tasks as the label to construct a supervised learning dataset.

[0160] 2. Model Selection and Training Algorithm Selection: Adopt regression-based machine learning models (such as Random Forest, XGBoost) to adapt to numerical label prediction tasks.

[0161] Training and Validation: Divide the training set and test set according to time windows (such as using the first 80% of the time data for training and the last 20% for validation), and evaluate the prediction accuracy of the model on historical data (such as MAE, R²).

[0162] Step 4: Decision on Single Migration Reception Volume Based on the Model 1. Model Deployment and Interface Encapsulation Package the trained model as a REST API or a built-in database function, and integrate it into the migration scheduling system to support real-time reception of current performance parameters and return migration volume suggestions.

[0163] 2. Dynamic Migration Volume Prediction Process Input Preprocessing: Convert the current storage engine performance metrics (such as real-time I / O rate) and migration task attributes (such as data volume to be migrated) into the model input format.

[0164] Model Inference: Call the model API to output the predicted value of the single migration reception volume (for example, "maximum 5000 records received in a single time").

[0165] Safety Threshold Limit: If the model output value exceeds the hardware capacity limit (such as insufficient remaining disk space), automatically trigger a conservative value (such as 80% of the historical safe migration volume) as the final decision.

[0166] Step 5: Model Continuous Optimization and Feedback Loop 1. Online Learning and Iterative Update After each migration task is completed, compare the actual migration results (such as the amount of successful data, elapsed time) with the model prediction value to generate incremental training data, and regularly trigger model retraining (such as automatic update every week).

[0167] 2. Abnormal Scenario Handling Mechanism When the model prediction deviation exceeds the threshold (such as prediction error > 30%), trigger an alarm and switch to the rule engine (such as a static migration volume strategy based on historical averages), and record the reason for the anomaly for model repair.

[0168] It should be noted that the embodiments of this specification have the following beneficial effects through the above content: Improve the accuracy and adaptability of migration volume prediction: By analyzing multi-dimensional data (data volume, time, performance metrics) in historical migration logs and combining machine learning models to capture historical key factors (such as storage engine I / O bottlenecks, concurrent load thresholds), dynamically generate a migration reception volume that matches the current database performance, overcoming the prediction biases of manual experience or fixed rules.

[0169] Enhance the model's ability to analyze complex performance correlations: Based on machine learning, construct a performance prediction model that can automatically learn the non-linear relationship between historical performance test results and migration key factors (such as storage latency caused by a sharp increase in data volume), accurately quantify the impact weight of storage engine performance on the migration volume, and achieve adaptive optimization of migration strategies.

[0170] Reduce the potential risk of migration performance fluctuations: By identifying historical key factors (such as a sharp drop in performance triggered by a specific data volume threshold), anticipate and avoid migration volume configurations that may cause database instability, and fundamentally reduce the probability of storage engine overload or response timeouts during the migration process.

[0171] Achieve continuous self-optimization of migration strategies: The machine learning model can continuously update parameters and iterate and optimize with the accumulation of historical migration logs and performance test data, enabling the decision-making logic of the single migration reception volume to evolve autonomously with database hardware upgrades or changes in business scenarios, maintaining long-term effectiveness.

[0172] Strengthen the interpretability and controllability of the migration process: By explicitly identifying key factors that affect migration performance (such as the correlation between migration time and storage engine CPU occupancy), provide decision-making basis for operation and maintenance personnel, facilitate targeted adjustment of database configurations or migration plans, and enhance the transparency and intervenability of the migration process.

[0173] Furthermore, when determining the single migration reception volume based on the performance prediction model, the current performance test results and the current migration log can be obtained through the current storage engine performance test of the new oil and gas field database; analyze the current migration log to identify the current key factors that affect migration performance; input the current performance test results and the current key factors into the performance prediction model to determine the single migration reception volume.

[0174] In the embodiments of this specification, the following specific implementation schemes can be adopted: Step 1: Execute the current storage engine performance test 1. Lightweight real-time performance test design Test scenario simulation: Perform a short-term (e.g., 10 seconds) write stress test on a new database, simulate migration tasks with different data volumes (e.g., small batches, medium batches), and record the real-time response metrics of the storage engine (e.g., disk I / O rate, CPU occupancy rate, memory peak).

[0175] Standardization of test results: Store the performance metrics in a unified format (e.g., JSON or database table), including fields such as timestamp, test type, and performance values.

[0176] 2. Real-time collection of migration logs Log capture and filtering: Through the built-in log tool of the database or a third-party monitoring system (e.g., Prometheus), capture the log data of the current migration task (e.g., migration rate, error type, task duration), and filter out irrelevant noise data (e.g., debug logs).

[0177] Step 2: Analyze the current migration logs to identify key factors 1. Key factor extraction process Abnormal pattern recognition: Through a log analysis tool (e.g., ELK Stack), detect high-frequency errors in the migration logs (e.g., "write timeout", "connection interruption"), and correlate the performance metrics at the time of their occurrence (e.g., the probability of timeout increases during high CPU occupancy).

[0178] Performance bottleneck location: Statistically analyze the correlation between the migration rate and the storage engine metrics (e.g., for every 10ms increase in disk I / O latency, the migration rate decreases by 20%), and determine the current core bottleneck factors (e.g., I / O latency, memory shortage).

[0179] 2. Key factor priority ranking Rank the key factors according to the degree of influence (e.g., the probability of causing migration failure, the contribution to the duration) (e.g., I / O latency > memory shortage > network bandwidth).

[0180] Step 3: Determine the migration reception volume by inputting into the performance prediction model 1. Preprocessing of model input data Feature alignment: Convert the current performance test results (e.g., disk I / O = 120MB / s) and key factors (e.g., "I / O latency is the bottleneck") into the input format required by the model (e.g., standardized numerical vectors or classification labels).

[0181] Enhanced context association: Combine the current migration task attributes (e.g., data type is "graphical features"), add business labels to the input data, and improve the model's scenario adaptability.

[0182] 2. Model inference and result output Model call: Call the pre-trained performance prediction model through the API or local library, input the pre-processed data, and obtain the predicted value of the single migration reception volume (such as "the maximum single reception volume = 5000 records").

[0183] Result verification and correction: If the model output value exceeds the preset safety range (such as exceeding the hardware load limit), then trigger a conservative strategy (such as using historical safety values), and record the anomalies for model iteration and optimization.

[0184] Step 4: Dynamic execution and feedback of migration reception volume 1. Dynamic dispatching of migration tasks According to the reception volume output by the model, the data to be migrated is split into batches adapted to the current performance by the migration scheduler and assigned to the execution queue.

[0185] 2. Real-time feedback of execution results Record the deviation between the actual migration results (such as time consumption, success rate) and the predicted value as the training data for subsequent model optimization.

[0186] It should be noted that the embodiments of this specification have the following beneficial effects through the above content: Realize real-time dynamic optimization of migration strategies: By combining the current storage engine performance test results and real-time migration log analysis, dynamically identify the current key influencing factors (such as sudden load fluctuations, hardware state changes), so that the decision of the single migration reception volume can be adapted to the instantaneous performance state of the database in real time, and avoid the decline in migration efficiency caused by the disconnection between the historical model and the current environment.

[0187] Enhance the fault tolerance ability for sudden performance anomalies: Identify migration risk factors (such as instantaneous I / O bottleneck of the storage engine) based on the current performance data, and actively avoid unsafe migration volume settings through model prediction, reducing the probability of migration interruption caused by hardware failures or external interferences.

[0188] Improve the environmental perception accuracy of migration volume decision-making: Synchronously input the current performance test results and the running status in the real-time log (such as the number of concurrent tasks, cache occupancy rate) into the model to ensure that the migration reception volume accurately matches the real-time resource margin of the database, and optimize the balance between resource utilization and migration efficiency.

[0189] Support the closed-loop iterative optimization of migration strategies: By continuously feeding back the current migration log and performance data to the prediction model, form a closed-loop optimization link of "decision - execution - verification - update", so that the model can evolve autonomously following the changes in the database running environment (such as hardware upgrades, business expansion), and maintain long-term effectiveness.

[0190] Enhance the full-cycle controllability of the migration process: Before each migration task is launched, through a real-time data-driven model re-evaluation mechanism, ensure that the single migration volume is always within the optimal load range of the database, and achieve full-cycle stability guarantee from task startup to completion.

[0191] Figure 2 The structural schematic diagram of a historical data migration device for an oil and gas field database provided by one or more embodiments of this specification includes: an image generation unit 201, an access unit 202, a text feature extraction unit 203, a graphic feature extraction unit 204, a first data volume determination unit 205, a migration data volume determination unit 206, a migration data volume determination unit 207, a migration data volume division unit 208, a second data volume determination unit 209, a data volume comparison unit 210, and a migration unit 211.

[0192] The image generation unit 201 digitally processes the paper logging documents in the oil and gas field database to generate RGB images; The access unit 202 synchronously accesses the semi-structured electronic documents in the oil and gas field database; The text feature extraction unit 203 extracts the logging text features in the RGB images and the electronic documents; The graphic feature extraction unit 204 extracts the logging graphic features in the RGB images and the electronic documents for logging curve symbols; The first data volume determination unit 205 determines the first data volume of the logging text features and the logging graphic features; The migration data volume determination unit 206 determines the single migration data volume based on the single migration reception volume of the new oil and gas field database; The migration data volume determination unit 207 determines whether the first data volume exceeds the single migration data volume; The migration data volume division unit 208, if so, divides the first data volume into multiple migration data volumes according to the single migration data volume; The second data volume determination unit 209 determines the second data volume of the new oil and gas field database; The data volume comparison unit 210 compares whether the first data volume is consistent with the second data volume; The migration unit 211, if consistent, migrates the logging text features and the logging graphic features to a pre-constructed new oil and gas field database according to the preset classification of logging data.

[0193] Figure 3 The structural schematic diagram of a historical data migration device for an oil and gas field database provided by one or more embodiments of this specification includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to: Digitally process paper logging documents in the oil and gas field database to generate RGB images; Synchronously access semi-structured electronic documents in the oil and gas field database; Extract logging text features from the RGB images and the electronic documents; For logging curve symbols, extract logging graphic features from the RGB images and the electronic documents; Determine a first data volume of the logging text features and the logging graphic features; Based on the single migration reception volume of the new oil and gas field database, determine the single migration data volume; Judge whether the first data volume exceeds the single migration data volume; If so, divide the first data volume into multiple migration data volumes according to the single migration data volume; Determine a second data volume of the new oil and gas field database; Compare whether the first data volume is consistent with the second data volume; If they are consistent, migrate the logging text features and the logging graphic features to a pre-constructed new oil and gas field database according to the preset classification of logging data.

[0194] A non-volatile computer storage medium provided by one or more embodiments of this specification, storing computer-executable instructions, and when the computer-executable instructions are executed by a computer, they can implement: Digitally process paper logging documents in the oil and gas field database to generate RGB images; Synchronously access semi-structured electronic documents in the oil and gas field database; Extract logging text features from the RGB images and the electronic documents; For logging curve symbols, extract logging graphic features from the RGB images and the electronic documents; Determine a first data volume of the logging text features and the logging graphic features; Based on the single migration reception volume of the new oil and gas field database, determine the single migration data volume; Judge whether the first data volume exceeds the single migration data volume; If so, divide the first data volume into multiple migration data volumes according to the single migration data volume; Determine the second data volume of the new oil and gas field database; Compare whether the first data volume is consistent with the second data volume; If they are consistent, transfer the logging text features and the logging graphic features to the pre-constructed new oil and gas field database according to the preset classification of the logging data.

[0195] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiments of the device, equipment, and non-volatile computer storage medium, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiments.

[0196] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiments.

[0197] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0198] In the embodiments provided in this application, it should be understood that the disclosed device / network device and method can be implemented in other ways. For example, the device / network device embodiments described above are only illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in an electrical, mechanical, or other forms.

[0199] The unit described as a separation component may or may not be physically separated. The component displayed as a unit may or may not be a physical unit, that is, it may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0200] In addition, each functional unit in various embodiments of the present application may be integrated in a processing unit, may exist separately as individual physical units, or two or more units may be integrated in one unit. The above units can be implemented in the form of hardware or in the form of software.

[0201] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present application, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0202] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A method for migrating historical data of an oil and gas field database, characterized in that Including: Digitally process the paper logging documents in the oil and gas field database to generate RGB images; Synchronously access the semi-structured electronic documents in the oil and gas field database; Extract the logging text features from the RGB images and the electronic documents; For logging curve symbols, extract the logging graphic features from the RGB images and the electronic documents; Determine the first data volume of the logging text features and the logging graphic features; Based on the single migration reception volume of the new oil and gas field database, determine the single migration data volume; Judge whether the first data volume exceeds the single migration data volume; If so, divide the first data volume into multiple migration data volumes according to the single migration data volume; Determine the second data volume of the new oil and gas field database; Compare whether the first data volume is consistent with the second data volume; If they are consistent, migrate the logging text features and the logging graphic features to a pre-constructed new oil and gas field database according to the preset classification of logging data.

2. The method according to claim 1, wherein The determining the single migration data volume based on the single migration reception volume of the new oil and gas field database includes: Based on the single migration reception volume of the new oil and gas field database and the ratio of the logging text features and the logging graphic features, determine the single migration data volume.

3. The method according to claim 2, wherein Before the determining the single migration data volume based on the single migration reception volume of the new oil and gas field database and the ratio of the logging text features and the logging graphic features, the method further includes: Through the historical storage engine performance test of the new oil and gas field database, establish a migration volume prediction model in combination with the historical migration log to determine the single migration reception volume.

4. The method according to claim 3, characterized in that The historical migration log includes the data volume, migration time and performance indicators of the storage engine for each migration; The through the historical storage engine performance test of the new oil and gas field database, establishing a migration volume prediction model in combination with the historical migration log to determine the single migration reception volume includes: Through the historical storage engine performance test of the new oil and gas field database, obtain the historical performance test results and the historical migration log; Analyze the historical migration log to identify the historical key factors affecting the migration performance; Based on the historical performance test results and the historical key factors, establish a performance prediction model, and the performance prediction model is a machine learning model; Based on the performance prediction model, determine the single migration reception volume.

5. The method according to claim 4, characterized in that, The determining the single migration reception volume based on the performance prediction model includes: Through the current storage engine performance test of the new oil and gas field database, obtain the current performance test results and the current migration log; Analyze the current migration log to identify the current key factors affecting the migration performance; Input the current performance test results and the current key factors into the performance prediction model to determine the single migration reception volume.

6. An apparatus for migrating historical data of an oil and gas field database, characterized in that Including: An image generation unit that digitally processes the paper logging documents in the oil and gas field database to generate RGB images; An access unit that synchronously accesses the semi-structured electronic documents in the oil and gas field database; A text feature extraction unit that extracts the logging text features from the RGB images and the electronic documents; A graphic feature extraction unit extracts the logging graphic features in the RGB image and the electronic document for logging curve symbols; A first data volume determination unit determines the first data volume of the logging text features and the logging graphic features; A migration data volume determination unit determines the single - migration data volume based on the single - migration reception volume of the new oil and gas field database; A migration data volume judgment unit determines whether the first data volume exceeds the single - migration data volume; A migration data volume division unit, if so, divides the first data volume into multiple migration data volumes according to the single - migration data volume; A second data volume determination unit determines the second data volume of the new oil and gas field database; A data volume comparison unit compares whether the first data volume is consistent with the second data volume; A migration unit, if they are consistent, migrates the logging text features and the logging graphic features to a pre - constructed new oil and gas field database according to the preset classification of logging data.

7. An historical data migration device for an oil and gas field database, characterized in that Comprising: At least one processor; And, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to: Digitally process the paper logging documents in the oil and gas field database to generate RGB images; Synchronously access the semi - structured electronic documents in the oil and gas field database; Extract the logging text features in the RGB image and the electronic document; For logging curve symbols, extract the logging graphic features in the RGB image and the electronic document; Determine the first data volume of the logging text features and the logging graphic features; Determine the single - migration data volume based on the single - migration reception volume of the new oil and gas field database; Determine whether the first data volume exceeds the single - migration data volume; If so, divide the first data volume into multiple migration data volumes according to the single - migration data volume; Determine the second data volume of the new oil and gas field database; Compare whether the first data volume is consistent with the second data volume; If they are consistent, migrate the logging text features and the logging graphic features to a pre - constructed new oil and gas field database according to the preset classification of logging data.

8. A non-volatile computer storage medium, characterized in that, Stores computer - executable instructions, and when the computer - executable instructions are executed by a computer, they can implement: Digitally process the paper logging documents in the oil and gas field database to generate RGB images; Synchronously access the semi - structured electronic documents in the oil and gas field database; Extract the logging text features in the RGB image and the electronic document; For logging curve symbols, extract the logging graphic features in the RGB image and the electronic document; Migrate the logging text features and the logging graphic features to a pre - constructed new oil and gas field database according to the preset classification of logging data; Determine the first data volume of the logging text features and the logging graphic features; Determine the single - migration data volume based on the single - migration reception volume of the new oil and gas field database; Determine whether the first data volume exceeds the single - migration data volume; If so, divide the first data volume into multiple migration data volumes according to the single migration data volume; Determine the second data volume of the new oil and gas field database; Compare whether the first data volume is consistent with the second data volume; If they are consistent, migrate the well logging text features and the well logging graphic features to the pre-constructed new oil and gas field database according to the preset classification of the well logging data.

Citation Information

Patent Citations

  • Information extraction method and device for well logging interpretation curve

    CN108648245A

  • Database migration method and device, electronic equipment and storage medium

    CN111104392A

  • Data migration method and device, electronic equipment and computer readable storage medium

    CN116737693A

  • Historical data migration method and system

    CN116910022A

  • Paper version logging curve data acquisition method

    CN117313184A

Cited By

  • Data migration method, system and equipment for historical data of oiling laboratory and medium

    CN121387859A