Stratum state recognition method and system based on multi-parameter comprehensive analysis
By integrating stratigraphic data through multi-parameter comprehensive analysis and employing deviation correction and breakpoint identification techniques, the problem of fusion and correction of multi-source stratigraphic parameter data was solved, enabling high-precision real-time assessment and monitoring of stratigraphic conditions.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-26
- Publication Date
- 2026-04-07
AI Technical Summary
Existing technologies struggle to effectively integrate multi-source stratigraphic parameter data, especially in dynamically changing environments where data fusion and real-time correction are difficult to achieve. This leads to deviations in stratigraphic state identification results, affecting the scientific validity and reliability of engineering decisions.
By employing a multi-parameter integrated analysis method, data is acquired using multiple source sensors. An initial feature vector is formed using a parameter aggregation algorithm. The deviation is quantified by combining a benchmark model, and a deviation compensation coefficient is generated. Potential breakpoints are identified using a gradient field model, and a breakpoint attribute database is constructed. The window length is adjusted using a sliding window algorithm to generate the final formation state assessment report.
It significantly improves the data analysis accuracy and storage efficiency of formation condition identification, providing reliable technical support for real-time monitoring and assessment of formation conditions.
Smart Images

Figure CN121807852A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of geological exploration and resource development, and in particular relates to a method and system for identifying stratigraphic states based on multi-parameter comprehensive analysis. Background Technology
[0002] Stratigraphic condition identification, as a crucial research area in geological exploration and resource development, plays an irreplaceable role in ensuring engineering safety and improving resource extraction efficiency. Accurately determining stratigraphic characteristics not only affects the stability of underground structures but also directly impacts the scientific validity of related decisions. However, current research and applications often face challenges in assessing stratigraphic conditions due to complex environments and multi-source information, necessitating breakthroughs in existing technologies. Existing methods often struggle to effectively integrate diverse data from different sources when processing stratigraphic information, especially in dynamically changing environments, lacking a comprehensive consideration of the correlations between information and spatiotemporal variations. This leads to biased identification results under complex stratigraphic conditions, failing to meet the accuracy and real-time requirements of practical engineering. Particularly in multi-parameter environments, balancing the reliability and importance of different data becomes a pressing issue.
[0003] In addition, the core technical challenges lie in the fusion and dynamic analysis of multi-parameter data. First, formation parameters such as pressure and temperature come from different equipment, resulting in significant differences in data characteristics and complex interrelationships, making it difficult to form a unified characteristic representation. This can lead to the overlooking of key trends when comprehensively assessing formation conditions. Second, this complexity further exacerbates the problem of error accumulation in dynamic monitoring, especially during long-term or large-scale monitoring. Parameter deviations amplify over time and space, ultimately affecting the reliability of the overall assessment. For example, in underground drilling monitoring, if temperature data in a certain area is too high due to equipment interference and not corrected in time, it may lead to an incorrect judgment of formation stability, thus affecting the formulation of subsequent construction plans.
[0004] Therefore, how to effectively fuse data in a multi-parameter environment and correct deviations in dynamic changes in real time has become a key issue in the field of formation state identification. Summary of the Invention
[0005] To address the aforementioned technical problems, this invention provides a method and system for formation state identification based on multi-parameter comprehensive analysis. Specifically, a method for formation state identification based on multi-parameter comprehensive analysis includes:
[0006] Based on multi-source sensor data, time series and spatial location correlations are integrated through parameter aggregation algorithms to form an initial feature vector set.
[0007] For the measurement deviation in the initial feature vector set, a benchmark model is used to compare the measured values with the theoretical values, quantify the degree of deviation, and obtain the deviation compensation coefficient;
[0008] Based on the deviation compensation coefficient, the adjusted feature vector is dynamically monitored. If the deviation compensation coefficient exceeds a preset threshold, a corrected feature vector is generated by correcting it using historical data.
[0009] Obtain parameter mutation information from the corrected feature vector, calculate the gradient anomaly of the parameters using the gradient field model, determine the potential breakpoint location, and generate a breakpoint attribute database.
[0010] Based on the breakpoint attribute database, location data is extracted, and the sliding window algorithm is used to adaptively adjust the window length along the time dimension. Statistical features within the window are calculated to obtain the trend curve.
[0011] Based on the trend curve, a hierarchical storage architecture is constructed, data is migrated to different storage layers according to access frequency, and the storage index structure is determined.
[0012] The trend curve is retrieved from the storage index structure. If the retrieval result shows an abnormal trend, the multidimensional parameter data is re-integrated and the feature vector set is updated.
[0013] The updated feature vector set is obtained, and secondary correction is performed in combination with the deviation compensation coefficient. It is then determined whether the window needs to be expanded to capture subtle change trends, and the final formation state assessment report is generated.
[0014] Preferably, the process of integrating time series and spatial location correlations through a parameter aggregation algorithm to form an initial feature vector set includes:
[0015] Multidimensional parameter data is collected through multi-source sensors to construct raw data records;
[0016] Based on the original data records, the correlation between time series and spatial location is analyzed to generate a location-time mapping table;
[0017] For the aforementioned location-time mapping table, the multidimensional parameters are grouped using a parameter aggregation method to obtain aggregated parameter groups;
[0018] If there is a parameter value in the aggregated parameter group that exceeds the preset threshold range, then the parameter is marked as abnormal, and the abnormal parameter set is determined.
[0019] The abnormal parameter set is compared and analyzed with the initial feature vector, and the support vector machine algorithm is used to classify the degree of deviation and determine the distribution of deviation categories.
[0020] Based on the deviation category distribution, a deviation correction rule table is generated, and a set of corrected feature vectors is obtained.
[0021] Preferably, the process of obtaining the deviation compensation coefficient includes:
[0022] The measured data in the initial feature vector set are obtained and compared with the pre-established theoretical data to determine the preliminary measurement deviation range;
[0023] Based on the preliminary measurement deviation range, the measured data and theoretical data are compared item by item using a benchmark model to calculate the degree of deviation.
[0024] Based on the deviation value, the corresponding deviation compensation coefficient is obtained through a preset mapping rule;
[0025] If the deviation compensation coefficient exceeds the preset threshold range, then the deviation compensation coefficient is subjected to boundary constraint processing to determine the final applicable deviation compensation coefficient.
[0026] The initial feature vector set is adjusted element by element using the final applicable deviation compensation coefficient to obtain the adjusted feature vector set.
[0027] Preferably, the process of generating a corrected feature vector by correcting historical data includes:
[0028] Obtain the adjusted feature vector data, extract the deviation compensation coefficient, and determine the initial state of the deviation compensation coefficient;
[0029] The initial state of the deviation compensation coefficient is dynamically monitored. If the deviation compensation coefficient exceeds a preset threshold, a correction process is triggered to obtain monitoring result data.
[0030] Based on the monitoring results, historical data is retrieved, and a matching analysis is performed on the deviation compensation coefficient that exceeds the preset threshold to determine the correlation between the historical data and the current data.
[0031] Based on the correlation judgment results, the feature vectors are processed to generate corrected feature vector data.
[0032] Preferably, the process of generating the breakpoint attribute database includes:
[0033] By inputting the corrected feature vector, the gradient field model is used to process multiple parameters, calculate the gradient change value of each parameter, and obtain preliminary gradient anomaly distribution results.
[0034] Based on the gradient anomaly distribution results, obtain the parameter range with prominent outliers, perform a depth scan on the parameter range, and determine the candidate set of potential breakpoint locations.
[0035] Based on the potential breakpoint locations in the candidate set, a preset threshold is used for filtering. If the gradient outlier at a certain location exceeds the preset threshold, it is marked as a critical breakpoint, and a list of filtered breakpoint locations is obtained.
[0036] Based on the filtered list of breakpoint locations, obtain the time-series data segment corresponding to each breakpoint, construct a breakpoint attribute database, record the specific parameter information of each breakpoint, and obtain a complete attribute data record.
[0037] Preferably, the process of obtaining the trend curve includes:
[0038] Location data is extracted from the breakpoint attribute database, and preliminary sorting is performed on the time dimension to obtain a set of location data arranged in chronological order, resulting in structured time series data.
[0039] Based on the time series data, the data is segmented using a sliding window method, and the window size is dynamically adjusted to adapt to data fluctuations in different time periods, thereby determining the data range within each window.
[0040] Based on the data range within each window, calculate the statistical characteristic values, obtain the mean and variance data for each window, and obtain the window feature set;
[0041] If the variance data of a certain window in the window feature set exceeds a preset threshold, the window is subdivided, the window size is reduced, and the statistical feature values are recalculated to determine whether the feature values are stable.
[0042] By summarizing all window feature sets, trend data over time is generated, and a continuous trend curve is obtained.
[0043] Preferably, the process of determining the storage index structure includes:
[0044] Analyze the trend curve, extract the access frequency distribution characteristics of the data from the historical access records, classify the access frequency, and obtain the preliminary division results of high-frequency data and low-frequency data;
[0045] Based on the preliminary division results, a hierarchical storage architecture is constructed, with high-frequency data allocated to the hot data layer and low-frequency data allocated to the cold data layer, and the allocation ratio of storage resources for each layer is determined.
[0046] Based on the data distribution of the hot data layer and the cold data layer, a dynamic adjustment mechanism is implemented. If the access frequency of a certain data exceeds a preset threshold within a certain period of time, data migration is triggered to transfer it from the cold data layer to the hot data layer, and the storage status after migration is obtained.
[0047] Data location information is extracted from the migrated storage state, an index structure is designed, and a B-tree index is used to quickly locate data in the hot data layer, thus obtaining the access path after the index is built.
[0048] Preferably, the process of updating the feature vector set includes:
[0049] Based on the storage index structure, change trend curve data are obtained from the preset data warehouse, and historical records are extracted using a batch reading method to obtain a preliminary trend dataset.
[0050] Based on the preliminary trend dataset, anomaly detection is performed on the trend curve. If an abnormal trend is detected, the location and range of the anomaly point are determined.
[0051] Obtain multidimensional data records corresponding to abnormal trends, and use parameter aggregation methods to recombine the multidimensional data to generate an updated dataset;
[0052] Based on the updated dataset, feature vectors are calculated, the feature vectors are classified, and it is determined whether there are persistent anomalies.
[0053] If the feature vector classification results show persistent anomalies, the feature vector set is updated to generate a new feature vector set.
[0054] Through a cyclic monitoring process, the new set of feature vectors is compared with the historical trend curve to obtain the difference data.
[0055] Preferably, the process of generating the final formation condition assessment report includes:
[0056] The latest feature vector set is obtained from the data update stage, and the latest feature vector set is initially cleaned and standardized to obtain a normalized vector data set.
[0057] The normalized vector data set is corrected by applying a deviation compensation coefficient, and the corrected vector data set is determined by comparing it with a preset deviation threshold.
[0058] Based on the corrected vector data set, a secondary correction operation is performed, and a weighted average method is used to smooth out subtle changes, resulting in a smoothed feature data set.
[0059] The smoothed feature data group is analyzed for change trends. If the change trend exceeds the preset fluctuation range, the window expansion mechanism is triggered to determine the adjusted window range.
[0060] Based on the adjusted window range, data segments with subtle changes are re-extracted, and the data segments are classified to determine potential anomalies in the formation state.
[0061] Based on the classification results, integrate various indicators of the formation state, perform state analysis operations, and output comprehensive evaluation data results;
[0062] Based on the comprehensive assessment data, a dynamic monitoring record of the formation state is generated, and detailed information on the changing trends is presented through data visualization tools.
[0063] This invention also provides a formation state identification system based on multi-parameter integrated analysis, comprising:
[0064] The parameter aggregation module is used to acquire multi-dimensional parameter data based on multi-source sensors, and integrate time series and spatial location correlations through parameter aggregation algorithms to form an initial feature vector set.
[0065] The deviation compensation module is used to compare the measured values with the theoretical values using a benchmark model to quantify the degree of deviation and obtain the deviation compensation coefficient for the measurement deviation in the initial feature vector set.
[0066] The dynamic monitoring module is used to dynamically monitor the adjusted feature vector according to the deviation compensation coefficient. If the deviation compensation coefficient exceeds a preset threshold, a corrected feature vector is generated by correcting it using historical data.
[0067] The breakpoint detection module is used to obtain parameter mutation information in the correction feature vector, calculate the gradient anomaly of the parameters through the gradient field model, determine the potential breakpoint location, and generate a breakpoint attribute database.
[0068] The time series analysis module is used to extract location data based on the breakpoint attribute database, adaptively adjust the window length along the time dimension using a sliding window algorithm, calculate the statistical characteristics within the window, and obtain the trend curve.
[0069] The storage optimization module is used to construct a hierarchical storage architecture based on the trend curve, migrate data to different storage layers according to access frequency, and determine the storage index structure.
[0070] The loop monitoring module is used to retrieve the trend curve from the storage index structure. If the retrieval result shows an abnormal trend, the multidimensional parameter data is re-integrated and the feature vector set is updated.
[0071] The state assessment module is used to obtain the updated feature vector set, perform secondary correction in combination with the deviation compensation coefficient, determine whether it is necessary to expand the window to capture subtle change trends, and generate the final formation state assessment report.
[0072] Compared with the prior art, the present invention has the following advantages and technical effects:
[0073] This invention discloses a formation state assessment method based on multi-source sensor data fusion and deviation correction. It addresses the correlation between multi-dimensional parameter data in time series and spatial location, as well as the logical correlation challenges in business scenarios such as measurement deviation, parameter mutation, and storage optimization. The method integrates data to form an initial feature vector through a parameter aggregation algorithm, and uses a benchmark model to quantify the degree of deviation, generating compensation coefficients to adjust the feature vector.
[0074] This invention generates a corrected feature vector by dynamically monitoring deviations and combining them with historical data. It identifies potential breakpoint locations using a gradient field model, constructs a breakpoint attribute database, and then adaptively adjusts the window length using a sliding window algorithm to extract trend curves. A hierarchical storage architecture optimizes data retrieval speed. If an abnormal trend is detected, the data is re-fused and the feature vector is corrected a second time, ultimately generating a formation condition assessment report.
[0075] This invention significantly improves the accuracy of data analysis and storage efficiency, providing reliable technical support for real-time monitoring and assessment of formation conditions. Attached Figure Description
[0076] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings:
[0077] Figure 1 This is a schematic diagram of the method flow according to an embodiment of the present invention;
[0078] Figure 2 This is a schematic diagram of the system structure according to an embodiment of the present invention. Detailed Implementation
[0079] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.
[0080] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.
[0081] Example 1
[0082] like Figure 1 As shown, this embodiment provides a formation state identification method based on multi-parameter comprehensive analysis, including:
[0083] Based on multi-source sensor data, time series and spatial location correlations are integrated through parameter aggregation algorithms to form an initial feature vector set.
[0084] For the measurement deviation in the initial feature vector set, a benchmark model is used to compare the measured values with the theoretical values, quantify the degree of deviation, and obtain the deviation compensation coefficient;
[0085] Based on the deviation compensation coefficient, the adjusted feature vector is dynamically monitored. If the deviation compensation coefficient exceeds the preset threshold, the corrected feature vector is generated by correcting it using historical data.
[0086] Obtain parameter mutation information in the corrected feature vector, calculate the gradient anomaly of the parameters through the gradient field model, determine the potential breakpoint location, and generate a breakpoint attribute database;
[0087] Location data is extracted from the breakpoint attribute database. The sliding window algorithm is used to adaptively adjust the window length along the time dimension, calculate the statistical characteristics within the window, and obtain the trend curve.
[0088] Based on the trend curve, construct a hierarchical storage architecture, migrate data to different storage layers according to access frequency, and determine the storage index structure.
[0089] Retrieve the trend curve from the storage index structure. If the retrieval results show an abnormal trend, re-integrate the multidimensional parameter data and update the feature vector set.
[0090] The updated feature vector set is obtained, and secondary correction is performed in combination with the deviation compensation coefficient. It is then determined whether the window needs to be expanded to capture subtle change trends, and the final formation state assessment report is generated.
[0091] Furthermore, the process of integrating time series and spatial location correlations through parameter aggregation algorithms to form an initial feature vector set includes:
[0092] Multidimensional parameter data is collected through multi-source sensors to construct raw data records;
[0093] Based on the original data records, analyze the correlation between time series and spatial location, and generate a location-time mapping table;
[0094] For the location-time mapping table, the multidimensional parameters are grouped using the parameter aggregation method to obtain the aggregated parameter group;
[0095] If there is a parameter value in the aggregated parameter group that exceeds the preset threshold range, then the parameter is marked as abnormal, and the abnormal parameter set is determined.
[0096] The abnormal parameter set is compared and analyzed with the initial feature vector, and the support vector machine algorithm is used to classify the degree of deviation and determine the distribution of deviation categories.
[0097] Based on the distribution of deviation categories, a deviation correction rule table is generated, and a set of corrected feature vectors is obtained.
[0098] Furthermore, this embodiment utilizes temperature and humidity sensors at different depths and in different regions to collect real-time data on multidimensional parameters such as formation pressure, temperature, density, porosity, and permeability. These sensors collect data once per minute, forming raw data records. For example, a record might include a timestamp of 2023-10-01 12:00:00, location coordinates (latitude 39.9°, longitude 116.3°), temperature of 25.5°C, and humidity of 60%, thus constructing a raw dataset containing thousands of records, providing a foundation for subsequent analysis. Based on this, a time-series correlation analysis method is used to generate a location-time mapping table.
[0099] Specifically, this method first sorts the original data records by time series, and then uses spatial clustering algorithms, such as K-, which means associating nearby locations, for example, mapping sensor data of a certain area to the same grid cell to generate a table, where each row corresponds to a location grid and the columns include continuous data points of time series such as 12:00 to 13:00, thereby revealing the distribution pattern of parameters in the spatiotemporal dimension.
[0100] In one embodiment, for the location-time mapping table, this embodiment utilizes a parameter aggregation method to group multidimensional parameters, obtaining aggregated parameter groups. This aggregation method can be based on statistical averaging or weighted summation. For example, averaging the temperature data of a location grid over one hour calculates an average temperature of 26.0 degrees Celsius, while simultaneously grouping humidity into a high-humidity group (above 50%), thus forming aggregated parameter groups such as {Average Temperature: 26.0, Average Humidity: 62%, PM2.5 Peak: 50}. This helps simplify data complexity and highlight key trends. If some parameter values in the aggregated parameter group exceed a preset threshold range, the parameters in that group are marked as anomalies, identifying an abnormal parameter set.
[0101] Specifically, by comparing the abnormal parameter set with the initial feature vector, the Support Vector Machine (SVM) algorithm is used to classify the degree of deviation and determine the distribution of deviation categories. As a supervised learning model, SVM classifies data by finding the maximum margin hyperplane. Here, the initial feature vector may be a vector based on historical normal data, such as [temperature 25, humidity 55]. The comparison process involves calculating the Euclidean distance deviation. For example, if the deviation vector [1, 7] is input into the SVM model, after training, the model classifies it into mild, moderate, or severe deviation categories, with distributions such as 70% mild, 20% moderate, and 10% severe, thus quantifying the severity of the anomaly. Based on the deviation category distribution, a deviation correction rule table is generated, and the corrected feature vector set is obtained.
[0102] For example, for the moderate deviation category, the rule table may specify the correction of the data, and the result after application is a set of correction vectors, which can improve the accuracy of the data and enhance the reliability of environmental early warning.
[0103] In one embodiment, data storage and index construction are performed on the corrected feature vector set to obtain a queryable feature database. Specifically, the process includes storing the set in a NoSQL database such as MongoDB and building a B-tree index to support fast queries. For example, when a user queries the corrected temperature for a specific location, the system can efficiently retrieve and return the result, thereby enabling real-time data access and providing decision support for technical effectiveness.
[0104] Furthermore, the process of obtaining the deviation compensation coefficient includes:
[0105] The measured data in the initial feature vector set are obtained and compared with the pre-established theoretical data to determine the preliminary measurement deviation range;
[0106] Based on the preliminary measurement deviation range, the measured data and theoretical data are compared item by item using a benchmark model to calculate the degree of deviation.
[0107] Based on the degree of deviation, the corresponding deviation compensation coefficient is obtained through a preset mapping rule;
[0108] If the deviation compensation coefficient exceeds the preset threshold range, the deviation compensation coefficient is subjected to boundary constraint processing to determine the final applicable deviation compensation coefficient.
[0109] The initial feature vector set is adjusted element by element using the final applicable deviation compensation coefficient to obtain the adjusted feature vector set.
[0110] Furthermore, in this embodiment, when processing multi-dimensional parameter data acquired by multiple source sensors, comparing the measured data with the theoretical data of the initial feature vector set is a crucial step. This comparison method intuitively reflects the deviation of the parameter states, laying the foundation for subsequent analysis.
[0111] For example, when comparing measured and theoretical data item by item using a benchmark model to determine the initial measurement deviation range, temperature can be used as an example to further refine the analysis.
[0112] In one possible implementation, the benchmark model comprehensively considers factors such as the equipment's operating environment and historical data trends to calculate the degree of temperature deviation. Assuming that comparisons reveal that the temperature deviation is not only related to the equipment itself but may also be affected by fluctuations in the external ambient temperature, the deviation degree may be quantified into a specific range, such as moderate deviation. This analysis helps to more accurately pinpoint the root cause of the problem.
[0113] For example, when mapping the deviation compensation coefficient based on the degree of deviation, we can imagine a scenario where, for the degree of deviation of vibration parameters, a preset mapping rule maps a moderate deviation to a compensation coefficient of 0.8. If this coefficient exceeds a preset threshold range, such as 0.5 to 1.0, boundary constraint processing is required, and it may eventually be adjusted to 0.75. This process ensures the rationality of the compensation coefficient and avoids excessively large or small adjustments from interfering with the data.
[0114] For example, in the element-by-element adjustment of the initial feature vector set, assuming the original value of the pressure parameter is too high, after adjustment with a final compensation coefficient of 0.75, the data is closer to the theoretical expectation. This adjustment method can effectively smooth out abnormal fluctuations and improve data consistency.
[0115] For example, when using the Support Vector Machine (SVM) algorithm to classify the adjusted feature vector set, we can take equipment status classification as an example to determine whether the data meets the normal operating standards. Assuming the classification result shows that 90% of the equipment statuses are normal, then the adjusted data can be considered to have met the expected standards. This classification process provides a reliable basis for subsequent decision-making.
[0116] For example, in the final output data determination stage, if the classification results meet the standards, the adjusted feature vector set can be directly used to update the equipment condition monitoring database. This approach ensures data availability and supports long-term monitoring and analysis. Through these multi-stage refinement processes, the accuracy and reliability of data processing can be significantly improved, providing strong support for deviation correction in the field of formation condition monitoring.
[0117] Furthermore, the process of generating corrected feature vectors by correcting historical data includes:
[0118] Obtain the adjusted feature vector data, extract the deviation compensation coefficient, and determine the initial state of the deviation compensation coefficient;
[0119] The initial state of the deviation compensation coefficient is dynamically monitored. If the deviation compensation coefficient exceeds the preset threshold, the correction process is triggered to obtain the monitoring result data.
[0120] Based on the monitoring results, historical data is retrieved, and a matching analysis is performed on the deviation compensation coefficient that exceeds the preset threshold to determine the correlation between historical data and current data.
[0121] Based on the correlation judgment results, the feature vectors are processed to generate corrected feature vector data.
[0122] Furthermore, in the deviation correction process for feature vector data in this embodiment, the process of extracting the deviation compensation coefficient for the adjusted feature vector data starts with an analysis of the initial state of the data. In this embodiment, considering the sensor data processing scenario, sensor measurements may deviate due to environmental interference. The extracted deviation compensation coefficient might be 1.2, indicating that the original data needs to be amplified and adjusted to a certain extent. By comparing and analyzing the historical standard value of 1.0, it can be preliminarily determined that the current coefficient is too high and requires further monitoring.
[0123] For example, in the dynamic monitoring process of the deviation compensation coefficient, if the preset threshold range is 0.8 to 1.1, and the monitored coefficient is 1.2, a correction process will be triggered. The monitoring process can be based on real-time data streams, collecting coefficient fluctuation data every 5 minutes to generate monitoring result data. This method can detect anomalies in a timely manner, ensuring the timeliness of subsequent corrections.
[0124] For example, when retrieving historical data for matching analysis, suppose the repository contains deviation coefficient records for the past 30 days, and the current coefficient of 1.2 is highly correlated with a historical record of 1.25. By analyzing the environmental conditions between the two, such as differences in temperature or humidity, it can be determined that the current deviation may be caused by similar external factors. This correlation analysis helps to accurately pinpoint the root cause of the problem.
[0125] For example, regarding the method for correcting feature vectors, this embodiment smooths the data based on the correlation results. Assuming an element in the original feature vector has a value of 10, it is adjusted to 9.5 after correction to reduce the impact of bias. This processing method improves the stability of the data.
[0126] For example, in implementing the breakpoint detection process, this embodiment locates outliers by scanning and correcting feature vector data. Suppose that in 100 data points, the value of the 25th point suddenly jumps from 5 to 15, then it can be marked as a breakpoint. This detection helps to discover discontinuities in the data.
[0127] For example, in this embodiment, when classifying breakpoints using the support vector machine algorithm, breakpoints are divided into two types: those caused by equipment failure and those caused by environmental interference. Assuming the 25th breakpoint is related to equipment voltage fluctuations, it is classified as a failure-type breakpoint. This classification method facilitates subsequent targeted processing.
[0128] For example, in this embodiment, when generating the monitoring log file and storing it in the database, breakpoint information, classification results, and timestamps are recorded together. Assume the log file contains a detailed description of the 25th breakpoint and is stored in the cloud database for easy subsequent tracing and analysis. This storage method improves data manageability.
[0129] Furthermore, the process of generating the breakpoint attribute database includes:
[0130] By correcting the eigenvector data input, a gradient field model is used to process multiple parameters, and the gradient change value of each parameter is calculated to obtain preliminary gradient anomaly distribution results.
[0131] Based on the gradient anomaly distribution results, parameter intervals with prominent outliers are obtained, and a deep scan of the parameter intervals is performed to determine a candidate set of potential breakpoint locations.
[0132] Based on the potential breakpoint locations in the candidate set, a preset threshold is used for filtering. If the gradient outlier at a certain location exceeds the preset threshold, it is marked as a critical breakpoint, and a list of filtered breakpoint locations is obtained.
[0133] Based on the filtered list of breakpoint locations, obtain the time series data segment corresponding to each breakpoint, construct a breakpoint attribute database, record the specific parameter information of each breakpoint, and obtain a complete attribute data record.
[0134] Furthermore, for example, in processing the corrected feature vector data, this embodiment processes these parameters through a gradient field model to capture the changing trends of the parameters over time.
[0135] For example, when calculating gradient changes and obtaining preliminary results of gradient anomalies, consider this scenario: the gradient change of the amplitude parameter is found to be significantly higher than other parameters over a certain time period, reaching an increase of 1.2 units per second, while the increase of other parameters is only 0.3 units per second. This anomaly suggests focusing on a specific interval of the amplitude parameter. By deeply scanning this interval, assuming the scan results show that the gradient value is consistently high between 10:00 and 10:05, this interval can be listed as a candidate set for potential breakpoint locations.
[0136] For example, in this embodiment, when performing threshold filtering on potential breakpoint locations in the candidate set, assuming a preset threshold of 1.0 unit increase per second, the location where the gradient value reaches 1.5 units at time point 10:02 will be marked as a critical breakpoint. Only points meeting the criteria are retained in the filtered breakpoint location list, reducing redundant data in subsequent analysis. Next, the time-series data segment corresponding to each breakpoint is obtained. For example, the segment at 10:02 shows abnormal fluctuations in both amplitude and frequency. This information is stored in the breakpoint attribute database, recording detailed parameters such as the abnormal duration of 30 seconds and the frequency offset of 0.8 units.
[0137] For example, in this embodiment, when extracting and classifying breakpoint features related to time series data analysis, the features are divided into two categories: one category is features that directly reflect the equipment status, such as amplitude anomalies; the other category is indirect features, such as frequency shifts. Through classification, it is determined that amplitude anomalies are more indicative in time series analysis, while frequency shifts serve as an auxiliary reference. Finally, based on the application requirements of time series analysis, the input to the analysis model is generated, for example, using amplitude anomalies as the primary input data and frequency shifts as a weighting adjustment factor to construct the final analysis data structure. This approach ensures that the analysis model better reflects actual needs.
[0138] For example, in constructing the breakpoint attribute database and classifying breakpoint features, this embodiment emphasizes structured data storage and classification, which significantly improves the efficiency of subsequent analysis. Based on the classified feature set, this embodiment can help quickly locate potential time points of equipment failure, providing a basis for maintenance decisions. This method reduces misjudgments and provides a more reliable data foundation for time series analysis.
[0139] Furthermore, the process of obtaining the trend curve includes:
[0140] Location data is extracted from the breakpoint attribute database, and preliminary processing is performed on the time dimension to obtain a set of location data arranged in chronological order, resulting in structured time series data.
[0141] Based on time series data, a sliding window method is used to segment the data, dynamically adjust the window size to adapt to data fluctuations in different time periods, and determine the data range within each window;
[0142] Based on the data range within each window, calculate the statistical characteristic values, obtain the mean and variance data for each window, and obtain the window feature set;
[0143] If the variance data of a certain window in the window feature set exceeds the preset threshold, the window will be subdivided, the window size will be reduced, and the statistical feature values will be recalculated to determine whether the feature values are stable.
[0144] By summarizing all window feature sets, trend data over time is generated, and a continuous trend curve is obtained.
[0145] Furthermore, in this embodiment, when extracting location data from the breakpoint attribute database and organizing it by the time dimension, the database stores breakpoint location information corresponding to multiple time points. During the initial organization, these location data are arranged in ascending order of timestamps to form a time series dataset, such as one data point per hour from 8:00 AM to 8:00 PM, for a total of 12 location data points, forming a structured sequence for subsequent analysis.
[0146] For example, the sliding window method for time series data involved in this embodiment can dynamically adjust the window size to adapt to data fluctuations. In this embodiment, the initial window size is set to 2 hours, covering the data range from 8:00 to 10:00. If significant data fluctuations are found within this time period, such as a sudden increase in storage load from 50% to 80%, the window is reduced to 1 hour, and the data from 8:00 to 9:00 and from 9:00 to 10:00 are re-analyzed to ensure more accurate capture of fluctuations.
[0147] For example, in this embodiment, when calculating the statistical feature values within each window, regarding the acquisition of the mean and variance, taking one window as an example, assuming the data load mean from 9:00 to 10:00 is 60% and the variance is 5%. If the preset variance threshold is 3%, then the variance of this window exceeds the threshold, and the window needs to be subdivided, for example, the time period can be divided into 9:00 to 9:30 and 9:30 to 10:00, and the feature values can be recalculated until the variance stabilizes within the threshold. This allows for more detailed identification of abnormal fluctuations.
[0148] For example, when generating trend data by summarizing the feature set of the summarized windows, this embodiment forms a continuous trend curve by connecting the mean data points of each window. Assuming that the mean data from point 8 to point 12 are 50%, 60%, 75%, and 80% respectively, plotting the trend curve using these points can visually show the gradual increase in storage load, providing a basis for subsequent analysis.
[0149] For example, when analyzing storage space usage patterns, high-load periods can be identified based on trend curves. If the average load from 12:00 to 14:00 reaches 85%, which is much higher than the 60% of other periods, then this period is marked as a high-load interval, and more storage resources are allocated to it to ensure stable system operation.
[0150] For example, when adjusting the storage space allocation strategy in this embodiment, by comparing the trend curve with the actual load data, assuming that the trend curve predicts a load of 90% at 14 o'clock, while the actual load is 95%, it indicates that the existing allocation is insufficient and the storage capacity needs to be increased by 10% to form an optimized configuration scheme and improve resource utilization efficiency.
[0151] For example, through the above series of operations, we can more accurately grasp the operating status of the storage system, adjust resource allocation in a timely manner, avoid system bottlenecks caused by sudden load increases, optimize the utilization efficiency of storage space, and provide reliable support for continuous monitoring and analysis of time-series data.
[0152] Furthermore, the process of determining the storage index structure includes:
[0153] By analyzing the trend curves, extracting the access frequency distribution characteristics from historical access records, classifying the access frequencies, and obtaining preliminary results of the division between high-frequency and low-frequency data;
[0154] Based on the preliminary division results, a hierarchical storage architecture was constructed, allocating high-frequency data to the hot data layer and low-frequency data to the cold data layer, and determining the allocation ratio of storage resources for each layer;
[0155] Based on the data distribution of the hot data layer and the cold data layer, a dynamic adjustment mechanism is implemented. If the access frequency of a certain data exceeds a preset threshold within a certain period of time, data migration is triggered to transfer it from the cold data layer to the hot data layer, and the storage status after migration is obtained.
[0156] Extract data location information from the migrated storage state, design an index structure, and use a B-tree index to quickly locate data in the hot data layer to obtain the access path after index construction.
[0157] Furthermore, this embodiment extracts the access frequency distribution characteristics of data from historical access records when analyzing the trend curve, to gain a preliminary understanding of the storage system's load at different time periods. Assuming the storage system records show that the access frequency during weekdays from 8:00 AM to 6:00 PM is significantly higher than at night, after statistical classification, daytime data can be divided into high-frequency data, and nighttime data into low-frequency data. This classification provides a basis for the subsequent construction of a tiered storage architecture.
[0158] For example, in constructing a tiered storage architecture, this embodiment allocates high-frequency data to the hot data layer, using high-performance storage devices, while low-frequency data is allocated to the cold data layer, using lower-cost storage media. Assuming the hot data layer is allocated 70% of its fast storage resources and the cold data layer is allocated 30% of its slow storage resources, this proportion balances performance and cost.
[0159] For example, in the implementation of the dynamic adjustment mechanism, if the access frequency of a cold data layer file suddenly increases from 10 times per day to 100 times per day in a certain week, exceeding the preset threshold of 50 times, the migration mechanism is triggered to transfer it to the hot data layer. After migration, the system will record the new storage state to ensure that the data location information is updated, providing an accurate basis for subsequent retrieval.
[0160] For example, in designing the index structure, this embodiment uses a B-tree index for fast location of data in the hot data layer. Assuming the hot data layer stores 100,000 records, the B-tree index can locate them in milliseconds, significantly improving retrieval efficiency. The access path after index construction records the storage location of each data record, facilitating rapid system response to requests.
[0161] For example, in optimizing retrieval efficiency, if a retrieval request points to the hot data layer, the system will prioritize locating the target data through the index path. Assuming the preset retrieval time standard is 0.5 seconds, if the actual retrieval time is 0.3 seconds, the requirement is met; if it exceeds 0.5 seconds, further optimization is needed, such as adjusting the index structure or adding caching.
[0162] For example, in cases where retrieval time exceeds the standard, this embodiment preloads data from the cold data layer. If the system predicts that certain low-frequency data may be frequently accessed within the next 24 hours, it preloads this data and stores it near the hot data layer, reducing temporary migration time and improving response speed.
[0163] For example, when continuously monitoring changes in access frequency, if a sudden increase in the frequency of accessing a certain data is detected, from 20 times per day to 200 times per day, exceeding the preset range, the tiering and migration process is re-triggered. The updated storage architecture configuration will adjust the resource allocation ratio according to the latest frequency distribution to ensure that the system always maintains a high-efficiency operating state. This continuous monitoring and dynamic adjustment mechanism can effectively cope with changes in data access patterns and ensure the stability and efficiency of the storage system.
[0164] Furthermore, the process of updating the feature vector set includes:
[0165] Based on the storage index structure, trend curve data is obtained from the preset data warehouse, and historical records are extracted using a batch reading method to obtain a preliminary trend dataset.
[0166] Based on the preliminary trend dataset, anomaly detection is performed on the trend curve. If an abnormal trend is detected, the location and range of the anomaly point are determined.
[0167] Obtain multidimensional data records corresponding to abnormal trends, and use parameter aggregation methods to recombine the multidimensional data to generate an updated dataset;
[0168] Based on the updated dataset, calculate feature vectors, classify the feature vectors, and determine whether there are persistent anomalies.
[0169] If the feature vector classification results show persistent anomalies, the feature vector set is updated to generate a new feature vector set.
[0170] By using a cyclical monitoring process, the new set of feature vectors is compared with the historical trend curve to obtain the difference data.
[0171] Furthermore, this embodiment designs specific implementation methods for processing trend curves and detecting anomalies from multiple perspectives. It combines the business background of hierarchical storage of historical access frequency and explains the data warehouse, trend dataset, anomaly detection and subsequent processing.
[0172] For example, regarding the process of retrieving trend curve data from a data warehouse and batch reading historical records, this embodiment designs a scheduled task system to batch read access records from the past 24 hours every morning. Assuming a system generates approximately 1 million access records daily, the batch reading can be processed in batches of 100,000 records each time, ensuring balanced system load. This approach effectively organizes a preliminary trend dataset, laying the foundation for subsequent analysis.
[0173] For example, when detecting anomalies in trend curves, this embodiment employs a time-series-based threshold judgment method. Assuming the normal fluctuation range of a certain data indicator is between 5,000 and 10,000 daily visits, if the number of visits suddenly reaches 20,000 on a certain day, the system will mark this point as an anomaly and record its time range as a specific hour within that day. This method facilitates the rapid location and range of anomalies, providing a basis for subsequent processing.
[0174] For example, to acquire multidimensional data records corresponding to abnormal trends and perform parameter aggregation, this embodiment extracts multidimensional information such as user ID, access time, and access type related to the anomaly points, and recombines this information according to the time dimension to generate an aggregated dataset in hours. This aggregation method helps to clarify the specific context of the anomaly occurrence, facilitating further analysis.
[0175] For example, when using the support vector machine algorithm to classify feature vectors, this embodiment extracts features such as the access frequency and duration of outliers into vectors. Assuming an outlier's access frequency is 100 times per minute and its duration is 30 minutes, the system will determine whether it is a persistent anomaly based on preset rules. This classification method helps distinguish between temporary fluctuations and long-term problems.
[0176] For example, when updating the feature vector set, this embodiment collects the latest access data in real time, assuming the feature vectors are updated hourly, and incorporates the latest outlier data into the set. This dynamic update mechanism ensures the timeliness of the analysis results.
[0177] For example, in the cyclical monitoring process, when comparing a new set of feature vectors with historical trend curves, this embodiment sets a difference threshold. If the difference exceeds 10%, the current trend is considered unstable, and the monitoring strategy needs further adjustment. This comparison method helps to identify potential problems in a timely manner.
[0178] For example, regarding adjusting the frequency and range of cyclic monitoring, this embodiment dynamically adjusts the frequency based on the magnitude of the difference in data. For instance, if the difference is high, the monitoring frequency is adjusted from once per hour to once every 30 minutes, while simultaneously expanding the monitored data range from a single indicator to multiple related indicators. This adjustment method improves the accuracy of monitoring and ensures the reliability of the final output results.
[0179] Through the above multi-faceted implementation methods, it can be seen that each step is closely centered around the business needs of data storage and analysis, progressing step by step to jointly support the complete process of formation state anomaly detection and handling. This design not only improves data processing efficiency but also provides a reliable basis for subsequent optimization.
[0180] Furthermore, the process of generating the final stratigraphic state assessment report includes:
[0181] The latest feature vector set is obtained from the data update stage, and the latest feature vector set is initially cleaned and standardized to obtain a normalized vector data set.
[0182] The normalized vector data set is corrected by applying a deviation compensation coefficient, and the corrected vector data set is determined by comparing it with a preset deviation threshold.
[0183] Based on the corrected vector data set, a second correction operation is performed, and a weighted average method is used to smooth out subtle changes, resulting in a smoothed feature data set.
[0184] The smoothed feature data set is analyzed for trend changes. If the trend changes exceed the preset fluctuation range, the window expansion mechanism is triggered to determine the adjusted window range.
[0185] Based on the adjusted window range, data fragments with subtle changes are re-extracted, and the data fragments are classified to identify potential anomalies in the formation state.
[0186] Based on the classification results, integrate various indicators of the formation state, perform state analysis operations, and output comprehensive evaluation data results;
[0187] Based on the comprehensive assessment data, a dynamic monitoring record of the formation state is generated, and detailed information on the changing trends is presented through data visualization tools.
[0188] For example, in the operational fields of data processing and formation condition monitoring, this embodiment addresses the processing and analysis of feature vector sets from multiple stages to ensure data accuracy and monitoring reliability. For the feature vector set obtained during the data update stage, preliminary cleaning and standardization are crucial steps. Assuming 1000 sets of feature vector data are extracted from the data warehouse, containing noise values and non-standard units, invalid data points are removed through cleaning, and then the data units are standardized, such as converting all values to percentages, resulting in a standardized vector data set. This step contributes to the accuracy of subsequent analysis.
[0189] For example, in the deviation correction stage, this embodiment introduces a deviation compensation coefficient to adjust the data. Assuming the deviation value of a certain set of vectors is above a preset threshold of 0.05, after correction using the compensation coefficient, the deviation value is reduced to within 0.02, ensuring data stability. This correction method can effectively reduce the impact of data acquisition errors.
[0190] For example, this embodiment employs a weighted average method for secondary correction and data smoothing. Assuming a segment of characteristic data fluctuates significantly within a short period, such as a value jumping from 10 to 20 within 5 minutes, the weighted average method smooths the fluctuation into a gradual trend, resulting in a smoothed characteristic data set. This helps avoid misjudgments caused by short-term fluctuations. In trend analysis, if the fluctuation of the smoothed data set exceeds a preset range, such as a change rate exceeding 15% for three consecutive hours, a sliding window expansion mechanism is triggered, extending the window range from 1 hour to 3 hours to capture changes over a longer period. This adjustment method more comprehensively reflects the dynamic changes in formation conditions.
[0191] For the re-extraction and classification of data with subtle changes, assuming 500 data fragments are extracted within an expanded window, a support vector machine algorithm is used for classification to determine that 10% of the data fragments may indicate potential anomalies in the formation state. This classification method helps to accurately locate problem areas.
[0192] For a comprehensive assessment of formation conditions, after integrating classification results and various indicators, if a region's stability index is found to be below 80 points, a comprehensive assessment result is output by combining other parameters such as pressure values and displacement data. This step provides a comprehensive basis for subsequent decision-making.
[0193] Finally, in the dynamic monitoring record generation and visualization stage, a line graph is used to display the trend of geological state changes over the past 24 hours, marking the time and magnitude of anomalies, such as the anomaly value reaching a peak of 25% at a certain moment, which facilitates a clear understanding of the situation for relevant personnel. This approach significantly improves the readability and usability of the monitoring information.
[0194] Example 2
[0195] like Figure 2 As shown, based on the same inventive concept, this embodiment also provides a formation state identification system based on multi-parameter comprehensive analysis, including:
[0196] The parameter aggregation module is used to acquire multi-dimensional parameter data based on multi-source sensors, and integrate time series and spatial location correlations through parameter aggregation algorithms to form an initial feature vector set.
[0197] The deviation compensation module is used to compare the measured values with the theoretical values using a benchmark model to quantify the degree of deviation and obtain the deviation compensation coefficient for the measurement deviation in the initial feature vector set.
[0198] The dynamic monitoring module is used to dynamically monitor the adjusted feature vector based on the deviation compensation coefficient. If the deviation compensation coefficient exceeds the preset threshold, the corrected feature vector is generated by correcting it using historical data.
[0199] The breakpoint detection module is used to obtain parameter mutation information in the correction feature vector, calculate the gradient anomaly of the parameters through the gradient field model, determine the potential breakpoint location, and generate a breakpoint attribute database.
[0200] The time series analysis module is used to extract location data based on the breakpoint attribute database, and adopts a sliding window algorithm to adaptively adjust the window length along the time dimension, calculate the statistical characteristics within the window, and obtain the trend curve.
[0201] The storage optimization module is used to build a tiered storage architecture based on the trend curve, migrate data to different storage layers according to access frequency, and determine the storage index structure.
[0202] The loop monitoring module is used to retrieve trend curves from the storage index structure. If the retrieval results show abnormal trends, the multidimensional parameter data is re-integrated and the feature vector set is updated.
[0203] The state assessment module is used to obtain the updated feature vector set, perform secondary correction by combining the deviation compensation coefficient, determine whether it is necessary to expand the window to capture subtle change trends, and generate the final formation state assessment report.
[0204] The formation state identification system based on multi-parameter comprehensive analysis provided in this embodiment has all the advantages of the formation state identification method based on multi-parameter comprehensive analysis provided in Embodiment 1.
[0205] Example 3
[0206] This embodiment also discloses a computer device, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the method described in Embodiment 1.
[0207] Example 4
[0208] This embodiment also discloses a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the method described in Embodiment 1.
[0209] Example 5
[0210] This embodiment also discloses a computer program product, including a computer program that, when executed by a processor, implements the steps of the method described in Embodiment 1.
[0211] The above are merely preferred embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for identifying formation states based on multi-parameter comprehensive analysis, characterized in that, include: Based on multi-source sensor data, time series and spatial location correlations are integrated through parameter aggregation algorithms to form an initial feature vector set. For the measurement deviation in the initial feature vector set, a benchmark model is used to compare the measured values with the theoretical values, quantify the degree of deviation, and obtain the deviation compensation coefficient; Based on the deviation compensation coefficient, the adjusted feature vector is dynamically monitored. If the deviation compensation coefficient exceeds a preset threshold, a corrected feature vector is generated by correcting it using historical data. Obtain parameter mutation information from the corrected feature vector, calculate the gradient anomaly of the parameters using the gradient field model, determine the potential breakpoint location, and generate a breakpoint attribute database. Based on the breakpoint attribute database, location data is extracted, and the sliding window algorithm is used to adaptively adjust the window length along the time dimension. Statistical features within the window are calculated to obtain the trend curve. Based on the trend curve, a hierarchical storage architecture is constructed, data is migrated to different storage layers according to access frequency, and the storage index structure is determined. The trend curve is retrieved from the storage index structure. If the retrieval result shows an abnormal trend, the multidimensional parameter data is re-integrated and the feature vector set is updated. The updated feature vector set is obtained, and secondary correction is performed in combination with the deviation compensation coefficient. It is then determined whether the window needs to be expanded to capture subtle change trends, and the final formation state assessment report is generated.
2. The method according to claim 1, characterized in that, The process of integrating time series data and spatial location correlations using parameter aggregation algorithms to form an initial feature vector set includes: Multidimensional parameter data is collected through multi-source sensors to construct raw data records; Based on the original data records, the correlation between time series and spatial location is analyzed to generate a location-time mapping table; For the aforementioned location-time mapping table, the multidimensional parameters are grouped using a parameter aggregation method to obtain aggregated parameter groups; If there is a parameter value in the aggregated parameter group that exceeds the preset threshold range, then the parameter is marked as abnormal, and the abnormal parameter set is determined. The abnormal parameter set is compared and analyzed with the initial feature vector, and the support vector machine algorithm is used to classify the degree of deviation and determine the distribution of deviation categories. Based on the deviation category distribution, a deviation correction rule table is generated, and a set of corrected feature vectors is obtained.
3. The method according to claim 1, characterized in that, The process of obtaining the deviation compensation coefficient includes: The measured data in the initial feature vector set are obtained and compared with the pre-established theoretical data to determine the preliminary measurement deviation range; Based on the preliminary measurement deviation range, the measured data and theoretical data are compared item by item using a benchmark model to calculate the degree of deviation. Based on the deviation value, the corresponding deviation compensation coefficient is obtained through a preset mapping rule; If the deviation compensation coefficient exceeds the preset threshold range, then the deviation compensation coefficient is subjected to boundary constraint processing to determine the final applicable deviation compensation coefficient. The initial feature vector set is adjusted element by element using the final applicable deviation compensation coefficient to obtain the adjusted feature vector set.
4. The method according to claim 1, characterized in that, The process of generating corrected feature vectors by correcting historical data includes: Obtain the adjusted feature vector data, extract the deviation compensation coefficient, and determine the initial state of the deviation compensation coefficient; The initial state of the deviation compensation coefficient is dynamically monitored. If the deviation compensation coefficient exceeds a preset threshold, a correction process is triggered to obtain monitoring result data. Based on the monitoring results, historical data is retrieved, and a matching analysis is performed on the deviation compensation coefficient that exceeds the preset threshold to determine the correlation between the historical data and the current data. Based on the correlation judgment results, the feature vectors are processed to generate corrected feature vector data.
5. The method according to claim 1, characterized in that, The process of generating a breakpoint attribute database includes: By inputting the corrected feature vector, the gradient field model is used to process multiple parameters, calculate the gradient change value of each parameter, and obtain preliminary gradient anomaly distribution results. Based on the gradient anomaly distribution results, obtain the parameter range with prominent outliers, perform a depth scan on the parameter range, and determine the candidate set of potential breakpoint locations. Based on the potential breakpoint locations in the candidate set, a preset threshold is used for filtering. If the gradient outlier at a certain location exceeds the preset threshold, it is marked as a critical breakpoint, and a list of filtered breakpoint locations is obtained. Based on the filtered list of breakpoint locations, obtain the time-series data segment corresponding to each breakpoint, construct a breakpoint attribute database, record the specific parameter information of each breakpoint, and obtain a complete attribute data record.
6. The method according to claim 1, characterized in that, The process of obtaining the trend curve includes: Location data is extracted from the breakpoint attribute database, and preliminary sorting is performed on the time dimension to obtain a set of location data arranged in chronological order, resulting in structured time series data. Based on the time series data, the data is segmented using a sliding window method, and the window size is dynamically adjusted to adapt to data fluctuations in different time periods, thereby determining the data range within each window. Based on the data range within each window, calculate the statistical characteristic values, obtain the mean and variance data for each window, and obtain the window feature set; If the variance data of a certain window in the window feature set exceeds a preset threshold, the window is subdivided, the window size is reduced, and the statistical feature value is recalculated to determine whether the feature value is stable. By summarizing all window feature sets, trend data over time is generated, and a continuous trend curve is obtained.
7. The method according to claim 1, characterized in that, The process of determining the storage index structure includes: Analyze the trend curve, extract the access frequency distribution characteristics of the data from the historical access records, classify the access frequency, and obtain the preliminary division results of high-frequency data and low-frequency data; Based on the preliminary division results, a hierarchical storage architecture is constructed, with high-frequency data allocated to the hot data layer and low-frequency data allocated to the cold data layer, and the allocation ratio of storage resources for each layer is determined. Based on the data distribution of the hot data layer and the cold data layer, a dynamic adjustment mechanism is implemented. If the access frequency of a certain data exceeds a preset threshold within a certain period of time, data migration is triggered to transfer it from the cold data layer to the hot data layer, and the storage status after migration is obtained. Data location information is extracted from the migrated storage state, an index structure is designed, and a B-tree index is used to quickly locate data in the hot data layer, thus obtaining the access path after the index is built.
8. The method according to claim 1, characterized in that, The process of updating the feature vector set includes: Based on the storage index structure, change trend curve data are obtained from the preset data warehouse, and historical records are extracted using a batch reading method to obtain a preliminary trend dataset. Based on the preliminary trend dataset, anomaly detection is performed on the trend curve. If an abnormal trend is detected, the location and range of the anomaly point are determined. Obtain multidimensional data records corresponding to abnormal trends, and use parameter aggregation methods to recombine the multidimensional data to generate an updated dataset; Based on the updated dataset, feature vectors are calculated, the feature vectors are classified, and it is determined whether there are persistent anomalies. If the feature vector classification results show persistent anomalies, the feature vector set is updated to generate a new feature vector set. Through a cyclic monitoring process, the new set of feature vectors is compared with the historical trend curve to obtain the difference data.
9. The method according to claim 1, characterized in that, The process of generating the final formation condition assessment report includes: The latest feature vector set is obtained from the data update stage, and the latest feature vector set is initially cleaned and standardized to obtain a normalized vector data set. The normalized vector data set is corrected by applying a deviation compensation coefficient, and the corrected vector data set is determined by comparing it with a preset deviation threshold. Based on the corrected vector data set, a secondary correction operation is performed, and a weighted average method is used to smooth out subtle changes, resulting in a smoothed feature data set. The smoothed feature data group is analyzed for change trends. If the change trend exceeds the preset fluctuation range, the window expansion mechanism is triggered to determine the adjusted window range. Based on the adjusted window range, data segments with subtle changes are re-extracted, and the data segments are classified to determine potential anomalies in the formation state. Based on the classification results, integrate various indicators of the formation state, perform state analysis operations, and output comprehensive evaluation data results; Based on the comprehensive assessment data, a dynamic monitoring record of the formation state is generated, and detailed information on the changing trends is presented through data visualization tools.
10. A formation state identification system based on multi-parameter integrated analysis, characterized in that, include: The parameter aggregation module is used to acquire multi-dimensional parameter data based on multi-source sensors, and integrate time series and spatial location correlations through parameter aggregation algorithms to form an initial feature vector set. The deviation compensation module is used to compare the measured values with the theoretical values using a benchmark model to quantify the degree of deviation and obtain the deviation compensation coefficient for the measurement deviation in the initial feature vector set. The dynamic monitoring module is used to dynamically monitor the adjusted feature vector according to the deviation compensation coefficient. If the deviation compensation coefficient exceeds a preset threshold, a corrected feature vector is generated by correcting it using historical data. The breakpoint detection module is used to obtain parameter mutation information in the correction feature vector, calculate the gradient anomaly of the parameters through the gradient field model, determine the potential breakpoint location, and generate a breakpoint attribute database. The time series analysis module is used to extract location data based on the breakpoint attribute database, adaptively adjust the window length along the time dimension using a sliding window algorithm, calculate the statistical characteristics within the window, and obtain the trend curve. The storage optimization module is used to construct a hierarchical storage architecture based on the trend curve, migrate data to different storage layers according to access frequency, and determine the storage index structure. The loop monitoring module is used to retrieve the trend curve from the storage index structure. If the retrieval result shows an abnormal trend, the multidimensional parameter data is re-integrated and the feature vector set is updated. The state assessment module is used to obtain the updated feature vector set, perform secondary correction in combination with the deviation compensation coefficient, determine whether it is necessary to expand the window to capture subtle change trends, and generate the final formation state assessment report.