Shale reservoir dynamic monitoring method and system based on multi-source data fusion
By integrating multi-source data and conducting collaborative verification with multiple algorithms, the problem of inaccurate shale reservoir boundary identification under complex geological conditions was solved, enabling accurate simulation and range determination of reservoir boundaries, and improving the scientific nature and feasibility of resource development.
Patent Information
- Application Number
- CN202511899706.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-16
- Publication Date
- 2026-03-06
AI Technical Summary
Existing technologies struggle to effectively integrate different types of detection information in complex geological environments, leading to inaccurate identification of shale reservoir boundaries and impacting the efficiency and economy of resource development.
A multi-source data fusion method was adopted, which collected data through multi-source detection equipment, used the support vector machine algorithm to classify the data characteristics, used the weighted average method to fuse inconsistent data, combined with the random forest algorithm to verify the actual contact interface location, and iteratively optimized and adjusted the coordinates of candidate points to build an optimized interface model, and finally determined the reservoir boundary range.
It significantly improves the accuracy and reliability of shale reservoir boundary identification, providing efficient and reliable support for underground resource exploration.
Smart Images

Figure CN121613530A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of information processing technology, and in particular to a method and system for dynamic monitoring of shale reservoirs based on multi-source data fusion. Background Technology
[0002] Shale reservoir research is a crucial area in energy development, holding irreplaceable value for ensuring energy security and promoting a green and low-carbon transformation. Especially in complex geological environments, accurately understanding the location and variation patterns of reservoirs directly impacts the efficiency and economics of resource extraction. However, current research and practice still face numerous challenges, urgently requiring innovative methods to overcome existing bottlenecks.
[0003] While various detection methods are used for reservoir monitoring in existing technologies, there is often a lack of comprehensive utilization of information from different sources, leading to inaccurate detection results in complex environments. Especially when faced with diverse geological structures and blurred boundaries between reservoirs and surrounding rock strata, relying solely on a single detection method is insufficient to fully capture the true characteristics of the reservoir. This limitation frequently results in misjudging the reservoir's extent during actual development, thus affecting subsequent planning and operations.
[0004] A deeper technical challenge lies in effectively integrating different types of detection information to improve the ability to identify reservoir boundaries. This is particularly true when processing diverse information, as the physical characteristics of various data differ significantly. For example, some data excel at reflecting deep features of subsurface structures, while others are better suited for capturing subtle shallow changes. This inconsistency makes it difficult to find a unified analytical standard during the fusion process. Furthermore, this inconsistency can lead to misidentification of the interface between the reservoir and surrounding rock strata. For instance, in some complex geological areas, the lack of effective cross-validation between data may result in the reservoir boundary being misjudged as other geological boundaries, causing significant errors in resource assessment.
[0005] Therefore, overcoming the challenges of data fusion due to differences in data characteristics from multiple information sources, and accurately identifying the contact interface between the reservoir and surrounding rock strata, has become a critical issue that urgently needs to be addressed. Solving this problem not only concerns the accurate determination of the reservoir's spatial extent but also directly affects the scientific validity and feasibility of resource development plans. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to provide a method and system for dynamic monitoring of shale reservoirs based on multi-source data fusion.
[0007] To achieve the above objectives, the present invention adopts the following technical solution: A dynamic monitoring method for shale reservoirs based on multi-source data fusion includes: Acquire various types of detection information; Based on the various types of detection information obtained, a classified dataset is obtained; If there are inconsistencies in physical characteristics among the classified datasets, the inconsistent data are merged to determine the unified dataset after merging. Reservoir boundary features and surrounding rock features are extracted from the fused unified dataset to obtain candidate points of boundary contact interface; For the obtained candidate points of the boundary contact interface, determine the location of the actual contact interface; If the determined actual contact interface position deviates from the initial boundary judgment, the candidate point coordinates are adjusted through iterative optimization to obtain an optimized interface model. The final reservoir boundary range is determined by simulating the changes in reservoir space range using an optimized interface model.
[0008] As a preferred method, multi-source detection equipment is used to collect underground structure data and shallow change data to obtain various types of detection information.
[0009] As a preferred approach, based on the various types of detection information obtained, the support vector machine algorithm is used to classify the differences in data characteristics, resulting in a classified dataset.
[0010] As a preferred approach, if there are inconsistencies in physical characteristics among the classified datasets, the inconsistent data are merged using a weighted average method to determine the merged unified dataset.
[0011] As a preferred approach, the random forest algorithm is used to verify and analyze the obtained candidate boundary contact interfaces to determine the actual contact interface location.
[0012] Preferably, if the determined actual contact interface position deviates from the initial boundary determination, the candidate point coordinates are adjusted through iterative optimization to obtain an optimized interface model, including: By collecting the location data of the initial interface, the original coordinate information is obtained, and the preliminary interface location distribution is determined. Based on the preliminary interface location distribution, a comparison method is used to match it with the actual contact data. If the matching result shows a deviation, the deviation range is recorded to obtain the deviation distribution data. For the deviation distribution data, obtain a set of candidate coordinates, use an iterative adjustment method to update the coordinate positions, and determine whether the updated coordinates meet the preset convergence conditions; If the updated coordinates meet the convergence criteria, a temporary interface model is constructed based on the adjusted coordinate data to obtain a preliminary optimized interface framework. Based on the preliminary optimized interface framework, the support vector machine algorithm is used to further train the model, determine the model parameters, and construct the optimized interface structure. The optimized interface structure is used to obtain a validation dataset for testing. If the test results do not reach the preset threshold, the process returns to the iterative adjustment stage to update the candidate coordinates and obtain the final interface model.
[0013] This invention also provides a dynamic monitoring system for shale reservoirs based on multi-source data fusion, comprising: The first processing module is used to acquire various types of detection information; The second processing module is used to obtain a classified dataset based on the various types of detection information acquired. The third processing module is used to merge inconsistent data and determine the unified dataset after merging if there are inconsistencies in physical characteristics in the classified dataset. The fourth processing module is used to extract reservoir boundary features and surrounding rock features from the fused unified dataset to obtain candidate points of the boundary contact interface; The fifth processing module is used to determine the actual contact interface location based on the obtained candidate boundary contact interface points. The sixth processing module is used to adjust the coordinates of candidate points through iterative optimization if there is a deviation between the determined actual contact interface position and the initial boundary judgment, so as to obtain an optimized interface model. The seventh processing module is used to simulate the changes in reservoir space range based on the optimized interface model and determine the final reservoir boundary range.
[0014] As a preferred embodiment, the second processing module uses a support vector machine algorithm to classify the differences in data characteristics based on the acquired various types of detection information, thereby obtaining a classified dataset.
[0015] Preferably, the third processing module is used to merge inconsistent data by weighted averaging if there are inconsistencies in physical characteristics in the classified dataset, and to determine the merged unified dataset.
[0016] As a preferred option, the fifth processing module is used to perform verification analysis on the obtained candidate points of the boundary contact interface using the random forest algorithm to determine the location of the actual contact interface.
[0017] This invention addresses the challenges of diverse data sources, inconsistent physical characteristics, and boundary judgment errors in underground structure exploration. Through a systematic process of multi-source data acquisition, classification and fusion, and optimization verification, it achieves accurate simulation and range determination of reservoir boundaries. First, the invention collects underground structure and shallow layer variation data from multiple sources. A support vector machine algorithm is used to classify the data characteristics. If inconsistencies in physical characteristics exist, a weighted average method is used to fuse the data into a unified dataset. Then, reservoir and surrounding rock strata features are extracted, candidate boundary contact interfaces are selected, and the actual interface location is verified using a random forest algorithm. If deviations exist, the coordinates of the candidate points are iteratively optimized to construct an optimized interface model, ultimately simulating and determining the reservoir boundary range. This invention significantly improves the accuracy of boundary determination through data fusion and multi-algorithm collaborative verification, providing efficient and reliable technical support for underground resource exploration. Attached Figure Description
[0018] Figure 1 A flowchart illustrating a dynamic monitoring method for shale reservoirs based on multi-source data fusion according to the present invention; Figure 2 This is a schematic diagram of the dynamic monitoring method for shale reservoirs based on multi-source data fusion according to the present invention; Figure 3 This is another schematic diagram of the dynamic monitoring method for shale reservoirs based on multi-source data fusion according to the present invention. Detailed Implementation
[0019] To further understand the content of this invention, a detailed description of the invention is provided in conjunction with the accompanying drawings and embodiments. The specific embodiments described herein are for illustrative purposes only and are not intended to limit the invention. It should also be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings.
[0020] Example 1: like Figures 1 to 3 As shown, this embodiment of the invention provides a method for dynamic monitoring of shale reservoirs based on multi-source data fusion, including: S101. Collect underground structure data and shallow change data through multi-source detection equipment to obtain various types of detection information.
[0021] Data on underground structures and shallow changes were collected using multi-source detection equipment, yielding various types of data. The collected data underwent preprocessing techniques for cleaning and format standardization, resulting in standardized data. Based on this standardized data, a support vector machine (SVM) algorithm was used to classify and extract features from the underground structure information, determining the main categories of structure distribution. Specifically, shale gas reservoir parameters were selected as training samples using various types of detection information; the Q-factor was calculated using shallow change characteristics; the SVM model parameters were optimized; a regional SVM model was established; and the established model was used to identify shale reservoirs as prediction samples. If the classified structure distribution categories did not match the preset threshold range, the standardized data underwent secondary calibration to obtain calibrated data. Using the calibrated data and shallow change information, time series analysis was employed to monitor the change trends and determine the periodicity of the changes. Based on the periodicity of the changes and the structure distribution categories, a data integration model was constructed to obtain comprehensive analysis results. Based on the comprehensive analysis results, information processing technology was used to verify the correlation of multi-source detection data and determine the final mapping relationship between underground structure and shallow changes.
[0022] For example, when collecting data on underground structures and shallow changes using multi-source detection equipment, a scenario of monitoring an urban underground pipe network can be envisioned. The detection equipment includes ground-based radar, acoustic detectors, and electromagnetic induction devices, which respectively collect information on the depth, density, and material composition of the underground structure, as well as data on shallow soil moisture and displacement changes. After data acquisition, noise interference or inconsistent formats may occur, thus requiring data preprocessing. Assuming that outliers due to ground debris exist in the radar data, median filtering can be used to remove noise. Simultaneously, data from different devices can be standardized using standard timestamps and spatial coordinate formats to obtain standardized data.
[0023] In one possible implementation, when using the Support Vector Machine (SVM) algorithm to classify and extract features from underground structural information, standardized data can be input into the model to extract the material characteristics of the underground pipe network, such as metal pipes and non-metal pipes, and classify them into main distribution categories. Assuming the classification results show that metal pipes account for 80%, while a preset threshold requires the proportion of metal pipes to be between 60% and 70%, a secondary calibration of the data is necessary. The calibration process may include recalibrating the sensitivity of the detection equipment or manually verifying abnormal data points, ultimately obtaining calibrated data.
[0024] For example, by combining calibrated data and shallow soil change information, time series analysis can be used to monitor trends. This analysis can examine the fluctuations in shallow soil moisture over the past 12 months, revealing a 3-month cycle in moisture changes. This cyclical pattern may be related to seasonal rainfall, and the monitoring results can help predict the stability risks of underground structures. Furthermore, by combining structural distribution categories to construct a data integration model, the distribution of metal pipes can be correlated with moisture change trends, revealing a comprehensive result: increased corrosion risk of metal pipes with rising moisture levels.
[0025] In one possible implementation, when verifying the correlation of multi-source detection data using information processing technology based on the comprehensive analysis results, correlation analysis can be used to confirm that the consistency between the depth detection results of radar data and acoustic data reaches over 90%, thereby determining the mapping relationship between underground structures and shallow changes, such as the indirect impact of humidity changes on pipeline depth. This verification ensures the reliability of the analysis results and provides a scientific basis for subsequent maintenance decisions. Through this method, not only is the accuracy of data processing improved, but underground structural risks can also be effectively predicted, significantly enhancing monitoring efficiency and safety.
[0026] S102. Based on the various types of detection information obtained, the support vector machine algorithm is used to classify the differences in data characteristics to obtain the classified data set.
[0027] Multiple types of probe information are acquired. Through preliminary organization of the information sources, the original datasets of different information types are identified. For the original datasets, the Support Vector Machine (SVM) algorithm is used to analyze data characteristics, obtaining preliminary classification results based on characteristic differences. Based on the preliminary classification results, and considering characteristic differences and classification criteria, the dataset is grouped to determine if classification bias exists. If the classification results do not match a preset threshold, the classification parameters are readjusted to obtain an optimized classification set. By combining the optimized classification set with information type and data processing methods, the core characteristics of each data group are determined, resulting in structured data groups. For the structured data groups, a comparative analysis of information sources and classification results is conducted to determine if the groups conform to the initial distribution of the probe information. If the distribution is uneven, the dataset is further corrected to obtain a corrected dataset. Based on the corrected dataset, and combining the algorithm model and SVM method, the final classification result is determined, resulting in a stable dataset usable for subsequent analysis. By combining the stable dataset with classification criteria and characteristic differences, a data mapping relationship is generated after classification, determining the final data classification system.
[0028] For example, in the field of underground structure and shallow subsurface change detection, the acquisition and processing of various types of detection information can begin with the initial organization of the information sources. Suppose that three types of raw data—sound waves, electromagnetic waves, and gravity fields—are collected using multi-source devices. These data differ significantly in format and magnitude. During the initial organization, the source and acquisition environment of each type of data can be labeled. For instance, sound wave data primarily reflects subsurface density, electromagnetic wave data relates to conductivity, and gravity field data indicates mass distribution, thus forming a clear raw dataset.
[0029] Specifically, when using the Support Vector Machine (SVM) algorithm for characteristic analysis of the original dataset, data features such as signal strength and frequency range can be used as input to initially classify the differences in characteristics between hard and loose layers of underground structures. Assuming that the signal strength of hard layers is generally higher than 80 units, while that of loose layers is lower than 50 units, the algorithm can initially delineate the boundary. This classification method helps to quickly identify data characteristics, but the boundaries may be ambiguous.
[0030] For example, when judging classification bias, if 10% of the data points in the classification results of the hard and loose layers do not match the preset thresholds of 80 and 50, respectively, the classification parameters need to be readjusted. This can be achieved by adding feature dimensions, such as signal duration, which can improve the classification accuracy from 85% to 92%, resulting in an optimized classification set. This adjustment effectively reduces misclassifications.
[0031] Specifically, when determining the core characteristics of the optimized classification set, the density characteristics of the acoustic wave data group and the conductivity characteristics of the electromagnetic wave data group can be extracted separately based on the information type, forming structured data groups. For example, data groups with high density correspond to underground rock strata, while data groups with high conductivity may point to aquifers. This grouping method provides a clear basis for subsequent analysis.
[0032] For example, when comparing and analyzing grouped data with the initial distribution, if a data distribution deviates from the initial detection range—for instance, only 30% of the acoustic data group matches the expected rock layer distribution—a secondary correction is required. This can be done by correcting the location deviations of the sampling points, adjusting the distribution ratio to 70%, thus obtaining a corrected dataset. This correction ensures the representativeness of the data.
[0033] Specifically, when determining the final classification result for the corrected dataset, the support vector method can be used to map the data to a high-dimensional space, further refining the classification. For example, underground structures can be divided into three categories: rock strata, aquifers, and cavities, forming a stable dataset. This refined classification lays the foundation for subsequent trend analysis.
[0034] For example, when generating the data mapping relationships after classification, rock strata can be associated with high density, and aquifers with high conductivity, forming a clear classification system. This mapping relationship helps transform complex data into intuitive results, facilitating the correlation analysis between underground structures and shallow changes. Through the above multi-level processing, the accuracy and usability of the data are significantly improved, providing reliable support for underground exploration.
[0035] S103. If there are inconsistencies in physical characteristics in the classified dataset, the inconsistent data are merged using a weighted average method to determine the merged unified dataset.
[0036] By initially screening the categorized data, records with physical attribute inconsistencies are identified from the dataset, determining the initial range of data to be processed. If the number of records with physical attribute inconsistencies exceeds a preset threshold, a weighted average method is used to initially fuse the inconsistent data, yielding a preliminary fusion result. Based on the preliminary fusion result, a subset of data with significant deviations is identified to determine if local inconsistencies exist. If local inconsistencies exist, a second weighted average method is used to deeply fuse the subset of data with significant deviations, determining the corrected subset. Based on the corrected subset, an overall consistency index with the unified dataset is obtained to determine if it meets the preset consistency standard. Through analysis of the consistency index, data integration techniques are used to make final adjustments to data with minor deviations, resulting in the final unified dataset. If the consistency of the final unified dataset still does not meet the preset standard, data traces of the adjustment process are saved through logging to determine the reference basis for subsequent optimization.
[0037] For example, when processing categorized data, initial screening is a crucial first step. For records with inconsistencies in physical properties, the range to be processed can be determined by comparing the data acquisition environment and measurement standards. Suppose that in an industrial sensor dataset, temperature readings show a significant difference of 5 degrees Celsius and 25 degrees Celsius at the same time point. This inconsistency may stem from sensor calibration issues or environmental interference. By recording these outliers, the initial range of data to be processed can be determined as records with temperature reading deviations exceeding 10 degrees Celsius.
[0038] For example, when inconsistencies exceed a preset threshold, a weighted average method is used for initial fusion. Assuming the threshold is set at 10% of the total data, and the actual proportion of inconsistent records reaches 15%, different weights are assigned to these data. For instance, weights are allocated based on sensor reliability; high-precision sensor data has a weight of 0.7, while low-precision data has a weight of 0.3, thus obtaining the initial fusion result. This method effectively smooths out abnormal fluctuations and improves the overall reliability of the data.
[0039] For example, when analyzing consistency deviations, if a subset shows a large deviation—such as an average deviation of 8 degrees Celsius for a certain group of temperature data, far exceeding the overall average of 2 degrees Celsius—it is necessary to determine whether there are local inconsistencies. This could be due to sensors in specific areas being exposed to high temperatures for extended periods, leading to inflated readings. Identifying these local issues can provide targeted guidance for subsequent deep fusion.
[0040] For example, the double-weighted average method can be used for deep fusion to address local inconsistencies. Suppose that weights are reassigned to subsets with large biases, taking into account time and environmental factors. For instance, the weight of recent readings is increased to 0.6, while the weight of older data is reduced to 0.4. The corrected subset bias may then be reduced to within 3 degrees Celsius. This approach better reflects the true trend of the data.
[0041] For example, when evaluating overall consistency metrics, assuming the preset standard is a deviation of less than 2 degrees Celsius, if the overall deviation between the corrected subset and the unified dataset is 2.5 degrees Celsius, further adjustments are needed. Through data integration techniques, such as removing extreme outliers or introducing smoothing algorithms, the deviation can ultimately be reduced to 1.8 degrees Celsius, meeting the standard. This adjustment ensures high data consistency.
[0042] For example, if the final unified dataset still does not meet the standard, such as the deviation remaining at 2.2 degrees Celsius, the adjustment process is recorded in logs, including the basis for each weight allocation and outlier removal. These logs provide important references for subsequent optimizations, avoid repeated and ineffective adjustments, and lay the foundation for continuous improvement of data quality.
[0043] For example, throughout the entire process, whether in initial or deep fusion, the consistency of the data's physical attributes remains a core concern. Through multi-level fusion and adjustments, not only can the stability of the dataset be improved, but a reliable foundation can also be provided for subsequent classification analysis. This method is particularly important in industrial data processing, effectively addressing data fluctuations in complex environments.
[0044] S104. Extract reservoir boundary features and surrounding rock features from the fused unified dataset to obtain candidate points of boundary contact interface.
[0045] By integrating the dataset, preliminary classification results of reservoir and strata data are obtained. Based on these preliminary classification results, a hierarchical analysis method is used to determine the distribution of interface features. For the distribution of interface features, a support vector machine algorithm is used to determine possible boundary locations. If the distribution of boundary locations meets a preset threshold, candidate coordinates of the contact point set are obtained. Using the candidate coordinates of the contact point set, combined with geological structure information, point distributions consistent with interlayer relationships are selected. Based on the point distribution results, the spatial characteristics of interlayer relationships are analyzed to obtain the final set of boundary contact points. Using information classification techniques, interface descriptions consistent with reservoir and strata data are determined for the final set of boundary contact points.
[0046] For example, in the initial classification of reservoir information and strata data, sandstone and mudstone layers can be preliminarily distinguished by grouping key parameters such as porosity and permeability within the dataset. Suppose that in a reservoir dataset, the porosity of sandstone layers ranges from 15% to 25%, while that of mudstone layers ranges from 5% to 10%. This parameter range division allows for a rapid preliminary classification. This approach helps subsequent analyses to more accurately focus on the characteristics of different strata.
[0047] For example, to determine the distribution of interface features using stratification analysis methods, stratification can be based on differences in rock layer thickness and density.
[0048] In one possible implementation, assuming that the rock strata data of a certain block indicate that the upper sandstone layer is 20 meters thick and has a density of 2.2 grams per cubic centimeter, while the lower mudstone layer is 15 meters thick and has a density of 2.5 grams per cubic centimeter, by comparing these parameters, it is possible to infer the locations where interface features may occur at abrupt density changes. This analysis helps to clarify the preliminary distribution range of interlayer interfaces.
[0049] For example, when using the Support Vector Machine (SVM) algorithm to determine boundary locations, the acoustic velocity and resistivity of rock strata data can be used as input features to construct a classification model. Assuming the acoustic velocity is 4000 m / s in sandstone and 3000 m / s in mudstone, and the resistivity data also shows significant differences, the algorithm can distinguish possible boundary locations using these features. This method can effectively improve the accuracy of boundary identification.
[0050] For example, when obtaining candidate coordinates for the contact point set, points that meet a preset threshold can be selected by combining the boundary position distribution. Assuming the threshold is set to a boundary position deviation of less than 1 meter, points with deviations within 0.5 meters can be used as candidate coordinates. This selection method ensures high reliability of the points analyzed subsequently.
[0051] For example, when selecting point distributions that conform to inter-stratal relationships based on geological structure information, the distribution of faults and fold characteristics within the region can be referenced. Suppose a block contains a set of near-vertical faults; points near these faults in the candidate coordinates may not meet the requirements for inter-stratal continuity and therefore need to be eliminated. This approach ensures that the point distribution more closely reflects the actual geological conditions.
[0052] For example, when analyzing the spatial characteristics of interlayer relationships, the continuity and changing trends of interlayer contacts can be observed through a three-dimensional spatial projection of the point distribution. If a boundary segment exhibits a significant tilting characteristic in the projection map, it can be inferred that the area may have been affected by tectonic movement. This analysis helps to more comprehensively understand the spatial distribution patterns of boundary contact points.
[0053] For example, when using information classification techniques to determine interface descriptions, the final set of boundary contact points can be correlated and matched with information such as reservoir porosity and permeability. For instance, if a sandstone layer near an interface has high porosity, reaching 20%, it can be described as a high-porosity contact interface. This descriptive method provides more intuitive data support for subsequent reservoir evaluation.
[0054] S105. For the obtained candidate points of the boundary contact interface, the random forest algorithm is used for verification and analysis to determine the location of the actual contact interface.
[0055] The proposed random forest algorithm is an improved version of the CART-AMV algorithm. This improved version effectively reduces algorithm complexity and improves detection accuracy, demonstrating significant improvements. First, the CART algorithm is used to construct binary decision trees by calculating the Gini coefficient of the feature labels in the training dataset, instead of the multi-branch decision trees used in the random forest algorithm. This effectively reduces the logarithmic operation used in information gain-based decision tree construction (such as in C4.5 and ID3 algorithms) to a quadratic operation, thus reducing complexity. Next, the generated binary decision trees are pruned. This pruning process removes less influential feature labels, further reducing complexity and computational intensity and time for random forests with a large number of decision trees. Finally, the AMV algorithm is used to combine the pruned binary decision trees into a random forest. The AMV algorithm stands for Absolute Value Management. MajorityVote is an absolute majority voting algorithm that votes according to the principle of majority rule, requiring that the number of votes for the majority voter be no less than half the number of decision trees in the random forest. If the number of votes for the majority voter does not meet the above conditions, the binary decision trees in the random forest are regenerated. This can avoid overfitting when classifying datasets with high noise levels.
[0056] Initial data is obtained from the candidate point set to construct point set information for subsequent processing. A feature extraction method is used to separate key descriptive parameters from the acquired point set information, resulting in a feature set for model training. A random forest algorithm is applied to train the model on the separated feature set to determine the classification prediction model. If the output of the classification prediction model does not match a preset threshold, the feature set is readjusted to obtain an updated feature combination. Based on the updated feature combination, the classification prediction model is reapplied to determine the initial location range of the contact interface. Interface filtering is performed on the initial location range to obtain more accurate point data and determine the true location of the contact interface. For the determined contact interface location, the results are finally verified using a location confirmation process to obtain the final judgment result.
[0057] For example, in the process of obtaining initial data from the candidate point set, preliminary screening can be performed to extract point information that conforms to geological distribution patterns. Suppose that in a study of the interface between a reservoir and surrounding rock strata, the initial candidate point set contains 1000 point data points derived from the analysis results of seismic wave reflection signals. By conducting preliminary analysis of the height, depth, and spatial distribution characteristics of these points, outliers that significantly deviate from geological patterns, such as points with depth values exceeding reasonable ranges, can be eliminated, ultimately resulting in 800 relatively reliable initial data points.
[0058] For example, when constructing point set information and employing feature extraction methods, we can focus on parameters such as the spatial coordinates of the points, the intensity of reflected signals, and the correlation with neighboring points. Suppose we extract the three-dimensional coordinates and corresponding signal intensity values of each of these 800 points, forming a data table containing multi-dimensional features. This feature extraction method helps subsequent models better understand the spatial relationships between points, laying the foundation for further processing.
[0059] For example, when training a model using the random forest algorithm on a separated feature set, the feature set can be divided into a training group and a validation group in a 7:3 ratio. By learning from the training group data, the model can initially identify which feature combinations are more likely to correspond to the actual contact interface locations. If the model's classification accuracy for the validation group reaches 85% after training, but does not reach the preset 90% threshold, the feature set needs to be adjusted, such as by adding distance features between points or removing some noisy signal strength data.
[0060] For example, when reapplying the classification prediction model after feature set adjustment, an improvement in model performance can be observed. Assuming the accuracy improves to 92% after adjustment, the model output can be used to preliminarily determine the location range of the contact interface, such as determining that the interface may be located within a depth range of 500 to 550 meters. This approach helps narrow down the scope of subsequent analysis.
[0061] For example, when screening the initial location range, geological sedimentary patterns can be used to further refine the location. Assuming a range of 500 to 550 meters, by analyzing point density and signal continuity, a depth of 520 to 530 meters can be ultimately selected as a more precise contact interface location. This screening method effectively improves the reliability of the results.
[0062] For example, when performing final verification of the contact interface location using the location confirmation process, additional geological data, such as borehole sampling results, can be introduced and compared with the model's predicted location. Assuming the borehole data indicates a significant lithological change at a depth of 525 meters, highly consistent with the model prediction, this location can be confirmed as the final contact interface. This multi-verification approach significantly improves the reliability of the results, providing a solid basis for subsequent geological analysis.
[0063] For example, the implementation of this overall process can effectively improve the accuracy of contact interface identification, while multi-level data processing and model optimization ensure that the results are highly consistent with actual geological conditions. This method has significant application value in reservoir boundary research and can provide reliable technical support for resource exploration and development.
[0064] S106. If there is a deviation between the determined actual contact interface position and the initial boundary judgment, the candidate point coordinates are adjusted through iterative optimization to obtain an optimized interface model.
[0065] By collecting initial interface location data, raw coordinate information is obtained to determine the preliminary interface location distribution. Based on this distribution, a comparison method is used to match it with actual contact data. If the matching results show a deviation, the deviation range is recorded, resulting in deviation distribution data. For this deviation distribution data, a candidate coordinate set is obtained, and an iterative adjustment method is used to update the coordinate positions. The updated coordinates are then checked to see if they meet preset convergence conditions. If the updated coordinates meet the convergence conditions, a temporary interface model is constructed based on the adjusted coordinate data, resulting in a preliminary optimized interface framework. Based on this preliminary optimized framework, a support vector machine algorithm is used to further train the model, determine the model parameters, and construct the optimized interface structure. Using the optimized interface structure, a validation dataset is obtained for testing. If the test results do not reach a preset threshold, the iterative adjustment process is returned to update the candidate coordinates, resulting in the final interface model. Finally, a data comparison tool is used to verify the final interface model's matching degree with the actual contact data, determining the model's applicability.
[0066] For example, when processing the initial interface location data acquisition, raw coordinate information can be obtained through high-precision sensors. Assume the acquired coordinate points are distributed within the X-axis range of 10.5 to 15.5 and the Y-axis range of 20.0 to 25.0. This method ensures data comprehensiveness and lays the foundation for subsequent analysis.
[0067] It should be noted that during the data collection process, attention should be paid to the impact of environmental interference on data accuracy. For example, changes in lighting may cause coordinate shifts, so it is recommended to operate under stable conditions.
[0068] For example, in the matching process between the initial interface location distribution and actual contact data, a distance-based comparison method can be used. A deviation threshold of 0.3 units can be set. If a deviation of 0.5 units is found at a point, the deviation range of that point and its surrounding area is recorded. This method facilitates rapid location of problem areas and provides data support for subsequent adjustments.
[0069] For example, in processing biased distribution data, when obtaining a set of candidate coordinates, clustering methods can be used to categorize points with large biases. Assuming there are 10 candidate points in a certain area, the coordinate positions are updated iteratively, and it is determined whether convergence conditions are met, such as a coordinate change of less than 0.1 units. This iterative approach can gradually approximate a more accurate position.
[0070] For example, when constructing a temporary interface model, a framework structure containing key points can be generated based on the adjusted coordinate data, assuming the framework covers 90% of the target area. This framework provides a preliminary basis for subsequent optimization and helps improve the model's stability.
[0071] For example, when further training the model using the Support Vector Machine (SVM) algorithm, the data points can be divided into contact and non-contact classes by adjusting the classification boundary parameters. Assuming the training dataset contains 1000 sample points, with 800 used for training and 200 for validation, this method can effectively distinguish the characteristics of different regions.
[0072] For example, when testing the validation dataset, if the test results show that the accuracy does not reach the preset threshold of 85%, the process returns to the iterative adjustment stage to update the candidate coordinates again, assuming that the accuracy improves to 88% after the adjustment. This iterative optimization mechanism can continuously improve the model performance.
[0073] For example, to validate the final interface model, data comparison tools can be used to compare the model's predictions with actual contact data point by point. If the matching degree reaches 95% or higher, the model is considered applicable. This validation method ensures the reliability of the results and provides a guarantee for practical applications.
[0074] S107. Simulate the changes in reservoir space range based on the optimized interface model to determine the final reservoir boundary range.
[0075] Initial range information is obtained from reservoir space data using a pre-established interface model. Data mapping techniques are then used to initially organize the spatial distribution, resulting in a basic spatial range dataset. Based on this dataset, dynamic simulation methods are employed to calculate range changes in real time. Optimization methods are used to adjust for deviations during the simulation process, determining the dynamically changed range data set. If abnormal fluctuations exist in the dynamically changed range data set, model analysis techniques are used to identify and correct outliers, obtaining a corrected range data set. For the corrected range data set, spatial derivation techniques are used to calculate the potential boundaries of the reservoir space. These calculations are then verified using simulation results to determine the accuracy of the potential boundaries. Based on the accuracy assessment of the potential boundaries, range adjustment techniques are used to refine the boundary positions, resulting in an adjusted boundary dataset. Finally, using the adjusted boundary dataset and boundary determination techniques, the final boundary range is locked, obtaining the final reservoir boundary range data.
[0076] For example, when obtaining the initial range information of the reservoir space through a pre-established interface model.
[0077] Understandably, initial extent information is often based on historical exploration data and geological feature analysis. Suppose that in a reservoir space, the initial extent data extracted from the model is a rectangular area 10 kilometers long and 8 kilometers wide. This preliminary extent may include basic parameters such as formation thickness and porosity. Such data extraction provides a foundation for subsequent processing, ensuring that spatial distribution analysis is based on solid evidence.
[0078] For example, in the initial organization of spatial distribution using data mapping techniques, reservoir spatial data can be projected onto a two-dimensional planar map to form a basic spatial dataset. Assuming the original data points are unevenly distributed, mapping techniques can group the data points according to geological units, forming a uniformly distributed dataset. This approach helps to identify potential discontinuities within the reservoir space, providing a clear starting point for subsequent dynamic simulations.
[0079] For example, when using dynamic simulation methods to calculate range changes in real time, the dynamic adjustment of the range can be predicted by simulating fluid flow or pressure changes within the reservoir. If, during the simulation, significant pressure fluctuations are observed in a certain area, the optimization method will adjust the simulation parameters according to the magnitude of the fluctuations, ensuring that the range data set more closely reflects the actual formation conditions. This real-time adjustment effectively improves the reliability of range prediction.
[0080] For example, if abnormal fluctuations exist in the dynamically changed range dataset, model analysis techniques can identify outliers through data filtering. Suppose a boundary point deviates from the average by 2 kilometers; the correction method will interpolate data from surrounding points to generate a corrected range dataset. This correction reduces the interference of outliers on the overall range assessment.
[0081] For example, in the spatial extrapolation technique for estimating potential reservoir boundaries, possible boundary lines can be derived based on a corrected dataset and combined with geological trend analysis. Suppose the calculation shows a boundary shifting 1.5 kilometers northward; simulation verification would then compare this to historical drilling data to confirm the reasonableness of the calculation. This verification method helps improve the reliability of boundary determination.
[0082] For example, when using extent adjustment techniques to refine boundary locations, high-resolution data can be used to fine-tune the boundary lines. Suppose a boundary line has a 0.5-kilometer ambiguous area; the adjustment technique will then incorporate stratigraphic features to re-divide the boundary, creating a more accurate boundary dataset. This refinement process allows the boundary extent to better reflect actual geological conditions.
[0083] For example, to determine the final reservoir boundary extent using boundary delineation techniques, multi-source data fusion can be used. This involves combining the adjusted boundary dataset with seismic exploration data to generate the final extent data. Assuming the final extent data covers an area of 75 square kilometers, this method ensures the comprehensiveness and consistency of the boundary extent, providing a reliable basis for subsequent development planning.
[0084] Example 2: This invention also provides a dynamic monitoring system for shale reservoirs based on multi-source data fusion, comprising: The first processing module is used to acquire various types of detection information; The second processing module is used to obtain a classified dataset based on the various types of detection information acquired. The third processing module is used to merge inconsistent data and determine the unified dataset after merging if there are inconsistencies in physical characteristics in the classified dataset. The fourth processing module is used to extract reservoir boundary features and surrounding rock features from the fused unified dataset to obtain candidate points of the boundary contact interface; The fifth processing module is used to determine the actual contact interface location based on the obtained candidate boundary contact interface points. The sixth processing module is used to adjust the coordinates of candidate points through iterative optimization if there is a deviation between the determined actual contact interface position and the initial boundary judgment, so as to obtain an optimized interface model. The seventh processing module is used to simulate the changes in reservoir space range based on the optimized interface model and determine the final reservoir boundary range.
[0085] As one embodiment of the present invention, the second processing module uses a support vector machine algorithm to classify the differences in data characteristics based on the acquired various types of detection information, and obtains a classified data set.
[0086] As one embodiment of the present invention, the third processing module is used to merge inconsistent data by weighted average method if there are inconsistencies in physical characteristics in the classified data set, and to determine the merged unified dataset.
[0087] As one embodiment of the present invention, the fifth processing module is used to perform verification analysis on the obtained candidate points of the boundary contact interface using the random forest algorithm to determine the location of the actual contact interface.
[0088] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Various modifications and improvements made to the technical solutions of the present invention by those skilled in the art without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.
Claims
1. A shale reservoir dynamic monitoring method based on multi-source data fusion, characterized in that, The method comprises the following steps: acquiring multiple types of detection information; obtaining a classified data set according to the acquired multiple types of detection information; if there is inconsistency in physical characteristics in the classified data set, fusing the inconsistent data to determine a fused uniform data set; extracting reservoir boundary features and surrounding rock features from the fused uniform data set to obtain boundary contact interface candidate points; judging a real contact interface position according to the obtained boundary contact interface candidate points; if the judged real contact interface position deviates from an initial boundary judgment, adjusting the coordinates of the candidate points through iterative optimization to obtain an optimized interface model; simulating reservoir space range changes according to the optimized interface model to determine a final reservoir boundary range.
2. The method for shale reservoir dynamic monitoring based on multi-source data fusion of claim 1, wherein, Underground structure data and shallow change data are collected by multiple source detection equipment to acquire multiple types of detection information.
3. The method for shale reservoir dynamic monitoring based on multi-source data fusion of claim 2, wherein, According to the acquired multiple types of detection information, the support vector machine algorithm is used to classify the data characteristic differences to obtain a classified data set.
4. The method for shale reservoir dynamic monitoring based on multi-source data fusion of claim 3, wherein, If there is inconsistency in physical characteristics in the classified data set, the inconsistent data is fused by the weighted average method to determine a fused uniform data set.
5. The method for shale reservoir dynamic monitoring based on multi-source data fusion of claim 4, wherein, For the obtained boundary contact interface candidate points, a random forest algorithm is used for verification analysis to judge the real contact interface position.
6. The method for shale reservoir dynamic monitoring based on multi-source data fusion of claim 5, wherein, If the judged real contact interface position deviates from an initial boundary judgment, the coordinates of the candidate points are adjusted through iterative optimization to obtain an optimized interface model, which comprises the following steps: By collecting the position data of the initial interface, the original coordinate information is acquired to determine the preliminary interface position distribution; According to the preliminary interface position distribution, the comparison method is used to match the actual contact data, and if the matching result shows deviation, the deviation range is recorded to obtain deviation distribution data; For the deviation distribution data, a candidate coordinate set is acquired, and the iterative adjustment method is used to update the coordinate position to judge whether the updated coordinate meets the preset convergence condition; If the updated coordinate meets the convergence condition, a temporary interface model is constructed based on the adjusted coordinate data to obtain a preliminary optimized interface framework; According to the preliminary optimized interface framework, the support vector machine algorithm is used to further train the model to determine the model parameters and construct an optimized interface structure; Through the optimized interface structure, the verification data set is acquired for testing, and if the test result does not reach the preset threshold, the iterative adjustment link is returned to update the candidate coordinates to obtain a final interface model.
7. A shale reservoir dynamic monitoring system based on multi-source data fusion, characterized in that, The method comprises the following steps: a first processing module is used to acquire multiple types of detection information; a second processing module is used to obtain a classified data set according to the acquired multiple types of detection information; a third processing module is used to fuse inconsistent data to determine a fused uniform data set if there is inconsistency in physical characteristics in the classified data set; a fourth processing module is used to extract reservoir boundary features and surrounding rock features from the fused uniform data set to obtain boundary contact interface candidate points; a fifth processing module is used to judge a real contact interface position according to the obtained boundary contact interface candidate points; The sixth processing module is configured to, if the real contact interface position and the initial boundary judgment exist deviation, adjust the candidate point coordinates through iterative optimization to obtain an optimized interface model. The seventh processing module is configured to simulate reservoir space range variation according to the optimized interface model, and determine a final reservoir boundary range.
8. The multi-source data fusion based shale reservoir dynamic monitoring system of claim 7, wherein, The second processing module is configured to, according to the obtained multiple types of detection information, classify data characteristic differences by using a support vector machine algorithm to obtain a classified data set.
9. The multi-source data fusion based shale reservoir dynamic monitoring system of claim 8, wherein, The third processing module is configured to, if there is physical characteristic inconsistency in the classified data set, fuse the inconsistent data by using a weighted average method to determine a fused unified data set.
10. The multi-source data fusion based shale reservoir dynamic monitoring system of claim 9, wherein, The fifth processing module is configured to, for the obtained boundary contact interface candidate points, perform verification analysis by using a random forest algorithm to determine a real contact interface position.
Citation Information
Cited By
Method for generating regional physical space boundary based on multi-dimensional index analysis
CN122152957A