Geotechnical engineering survey data archiving and backup method, system and media

By identifying the target period of geotechnical engineering survey data and configuring threshold values, the problem of high cost of survey data storage and transmission is solved, and efficient data compression and precise data retention are achieved.

CN120179613BActive Publication Date: 2025-08-22BEIJING GO TO TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510637282.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-08-22
Estimated Expiration
2045-05-19

AI Technical Summary

Technical Problem

The prior art is difficult to set accurate data compression thresholds in geotechnical engineering survey data, resulting in high storage and transmission costs and low compression efficiency, which cannot meet the accuracy requirements of engineering analysis.

Method used

By analyzing the fluctuation characteristics of the survey data, identifying multiple target periods, and configuring threshold values ​​according to the abnormal importance, data compression and archive backup are implemented for the target period.

Benefits of technology

It improves the compression accuracy of survey data, retains more original data points, reduces storage space and bandwidth requirements, and meets the accuracy requirements of engineering analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179613B_ABST
    Figure CN120179613B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of data processing, and in particular to a method, system, and medium for archiving and backing up survey data for geotechnical engineering. This technical solution obtains multiple types of survey data for geotechnical engineering, each type of survey data corresponding to a single survey category of geotechnical engineering; obtains multiple target time periods of multiple types of survey data based on the fluctuation characteristics of the data values ​​of each type of survey data as the acquisition time changes; since the target time period is the time period with the greatest degree of data anomaly represented by each type of survey data, the survey category corresponding to each target time period is analyzed for the cause of the anomaly to obtain the anomaly importance of each target time period; according to the anomaly importance of each target time period, a threshold value for data compression is configured for the corresponding target time period; according to the threshold value of each target time period, the survey data of the corresponding target time period is compressed and archived and backed up, thereby achieving accurate compression of the survey data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and in particular to a method, system and medium for archiving and backing up survey data for geotechnical engineering. Background Art

[0002] Geotechnical engineering faces a variety of risks, such as geological disasters and foundation settlement. Survey data can be used for risk assessment. For example, extensive geological pre-survey work is often required in the construction planning of large-scale public transportation facilities and urban highways. This data can be recorded and stored by surveying and recording each borehole in the geotechnical project according to a pre-defined borehole distribution map.

[0003] The data involved in geotechnical engineering surveys includes geological profiles, soil samples, groundwater levels, and seismic wave velocities. Storing and transmitting this data in its raw form would consume enormous amounts of storage space and bandwidth resources. In practical geotechnical engineering applications, there is a certain tolerance for data accuracy. As long as the compressed data still meets the requirements of engineering analysis, a certain degree of information loss can be tolerated. Using lossy compression can reduce data storage and transmission costs, which is particularly important for large-scale engineering projects and can help reduce overall budgets. Existing revolving door lossy compression algorithms, when the data compression threshold is not appropriately selected, can result in reduced compression efficiency or compressed information loss exceeding the set tolerance. Summary of the Invention

[0004] In order to solve the technical problem of how to achieve accurate compression of survey data, the present invention aims to provide a survey data archiving and backup method, system and medium for geotechnical engineering. The technical solution adopted is as follows:

[0005] In a first aspect, an embodiment of the present invention provides a method for archiving and backing up survey data for geotechnical engineering, the method comprising:

[0006] Obtain multiple types of geotechnical engineering survey data, where each type of survey data corresponds to a single survey category of geotechnical engineering;

[0007] According to the fluctuation characteristics of the data values ​​of each type of survey data as the acquisition time changes, multiple target time periods of multiple types of survey data are obtained, wherein the target time period is the time period with the greatest degree of data anomaly represented by each type of survey data;

[0008] Perform set partitioning and intersection processing on the survey categories corresponding to each target period in turn, and determine whether each target period is a single-type anomaly source or a diverse anomaly source based on the processing results; obtain the anomaly factor of each target period based on the overlapping anomaly duration represented by the single-type anomaly source or the diverse anomaly source; and obtain the anomaly importance of each target period based on the anomaly factor of each target period;

[0009] According to the abnormal importance of each target period, configure the threshold value for implementing data compression in the corresponding target period;

[0010] According to the threshold value of each target period, the survey data of the target period is compressed and archived for backup.

[0011] In an optional embodiment, multiple target time periods of multiple types of survey data are obtained based on the fluctuation characteristics of the data values ​​of each type of survey data as the acquisition time changes, including:

[0012] If the abnormal fluctuation degree of the data value of each type of survey data in the preset collection period is greater than a first threshold, the collection time period of the corresponding survey data is determined to be an abnormal period;

[0013] If there are survey data of multiple survey categories in the same abnormal period, the multiple survey categories in the same abnormal period are determined as an abnormal category set;

[0014] According to the Pearson correlation coefficient, time overlap length and abnormal period difference of the two survey categories in the abnormal category set, the correlation degree of the corresponding two survey categories is obtained;

[0015] Comprehensive anomaly analysis and clustering processing are performed based on all correlation degrees, and the time period corresponding to the clustering results is determined as the target time period.

[0016] In an optional embodiment, the correlation degree between the two survey categories in the abnormal category set is obtained based on the Pearson correlation coefficient, the time overlap length, and the abnormal period difference of the two survey categories, including:

[0017] According to the formula , obtain the correlation degree between any two survey categories ,in, is the Pearson correlation coefficient between survey category w and survey category x at time t, is the abnormal period difference between survey category w and survey category x at time t, is the time overlap length between survey category w and survey category x at time t, is the normalization function.

[0018] In an optional embodiment, comprehensive anomaly analysis and clustering processing are performed based on all correlation degrees, and the time period corresponding to the clustering result is determined as the target time period, including:

[0019] When the correlation degree is greater than a second threshold, determining the corresponding abnormal period as a reference period;

[0020] According to the formula , obtain the comprehensive abnormality degree at time t ,in, is the number of first abnormal categories in the abnormal period at time t, W1 is the number of second abnormal categories in the abnormal category set at time t, is the abnormality degree of survey category w at time t, is the average value of the correlation between the survey category w and other survey categories in the abnormal category set at time t, is the number of reference periods at time t;

[0021] Clustering is performed based on the absolute value of the difference between any two comprehensive abnormality levels among all comprehensive abnormality levels as the clustering distance, and the time period corresponding to the clustering result is determined as the target time period.

[0022] In an optional embodiment, the abnormality cause analysis is performed on the survey category corresponding to each target period to obtain the abnormality importance of each target period, including:

[0023] The survey categories corresponding to each target period are divided into sets and intersections are processed in turn, and each target period is determined to be a single type of abnormal source or a diverse type of abnormal source based on the processing results;

[0024] Obtain the anomaly factor for each target period based on the duration of overlapping anomalies represented by a single type of anomaly source or multiple anomaly sources;

[0025] According to the abnormal factor of each target period, the abnormal importance of each target period is obtained.

[0026] In an optional embodiment, the survey categories corresponding to each target period are sequentially divided into sets and intersection-finding processes are performed, and each target period is determined to be a single type of abnormal source or a multiple type of abnormal source based on the processing results, including:

[0027] Divide the survey categories of each abnormal moment in each target period into a single set to obtain an abnormal category set for each abnormal moment in each target period;

[0028] Calculate the intersection of all abnormal category sets of the preset duration in each target period to obtain the processing results of each target period;

[0029] If the duration of the overlapping anomalies represented by the processing result is greater than the third threshold, the target period is determined to be a single type of anomaly source;

[0030] If the overlapping anomaly duration represented by the processing result is less than or equal to the third threshold, the target period is determined to be the source of the diversity anomaly.

[0031] In an optional embodiment, the abnormality factor for each target period is obtained based on the overlapping abnormality duration represented by a single abnormality source or multiple abnormality sources, including:

[0032] According to the formula , get the abnormal factor of the e-th target period ,in, is the maximum abnormal degree of all abnormal moments in the e-th target period, K is the number of overlapping abnormal durations in the e-th target period, The duration of overlapping anomalies represented by a single type of anomaly source or multiple anomaly sources, is the number of survey categories in the abnormal category set with the kth overlapping abnormal duration in the eth target period, It represents the number of all survey categories with the kth overlapping abnormal duration in the eth target period, where both e and k are natural numbers greater than 0.

[0033] In an optional embodiment, obtaining the abnormal importance of each target period according to the abnormal factor of each target period includes:

[0034] According to the formula , obtain the abnormal importance of the e-th target period ,in, is the abnormal factor of the e-th target period, Q is the number of target periods after the e-th target period, is the time interval between the e-th target period and the q-th target period, is the number of difference categories between the e-th target period and the q-th target period, where e is a natural number greater than 0.

[0035] In a second aspect, an embodiment of the present invention further provides a geotechnical engineering survey data archiving and backup system, the system comprising:

[0036] The collection terminal is used to collect multiple types of geotechnical engineering survey data and output them to the outside world, where each type of survey data corresponds to a single survey category of geotechnical engineering;

[0037] The processing terminal is connected to the acquisition terminal; the processing terminal is used to obtain multiple target time periods of multiple types of survey data based on the fluctuation characteristics of the data value of each type of survey data as the acquisition time changes, wherein the target time period is the time period in which the data anomaly represented by each type of survey data is the largest;

[0038] The processing terminal is also used to perform set division and intersection processing on the survey categories corresponding to each target time period in turn, and determine whether each target time period is a single-type abnormal source or a diverse abnormal source based on the processing results; obtain the abnormal factor of each target time period based on the overlapping abnormality duration represented by the single-type abnormal source or the diverse abnormal source; obtain the abnormal importance of each target time period based on the abnormal factor of each target time period; configure the threshold value for implementing data compression for the corresponding target time period based on the abnormal importance of each target time period; and compress the survey data of the target time period based on the threshold value of each target time period and archive and back up it.

[0039] In a third aspect, an embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of any one of the methods in the first aspect when executed by a processor.

[0040] The present invention has the following beneficial effects:

[0041] The technical solution of the present invention obtains multiple types of survey data for geotechnical engineering, each type of survey data corresponds to a single survey category of geotechnical engineering; based on the fluctuation characteristics of the data values ​​of each type of survey data as the acquisition time changes, multiple target time periods of the multiple types of survey data are obtained; since the target time period is the time period with the greatest degree of data anomaly represented by each type of survey data, the cause of the anomaly is analyzed for the survey category corresponding to each target time period to obtain the anomaly importance of each target time period; based on the anomaly importance of each target time period, a threshold value for implementing data compression for the corresponding target time period is configured; based on the threshold value of each target time period, the survey data for the target time period is compressed and archived and backed up. This technical solution parses the target time period corresponding to the abnormal data in the complex survey data of geotechnical engineering, configures the threshold value for implementing data compression based on the anomaly importance of the target time period, can improve the compression accuracy of such data, can retain more original data points, and thus achieve accurate compression of the survey data. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the prior art descriptions. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0043] Figure 1 A flowchart of a geotechnical engineering survey data archiving and backup method provided by one embodiment of the present invention;

[0044] Figure 2A schematic diagram of the structure of a geotechnical engineering survey data archiving and backup system provided by one embodiment of the present invention. DETAILED DESCRIPTION

[0045] To further illustrate the technical means and effectiveness of the present invention in achieving its intended purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, describes in detail a method, system, and medium for archiving and backing up survey data for geotechnical engineering projects, including its specific implementation, structure, features, and effectiveness. In the following description, references to "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics of one or more embodiments may be combined in any suitable manner.

[0046] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.

[0047] The following describes in detail a method, system and medium for archiving and backing up survey data for geotechnical engineering provided by the present invention in conjunction with the accompanying drawings.

[0048] During its formation, rock and soil undergo various geological processes and are influenced by multiple factors. As a result, it is a heterogeneous medium characterized by spatial variability. Due to the diversity, complexity, and uncertainty of this rock and soil medium, it can have a significant negative impact on engineering construction. Therefore, it is necessary to analyze the actual distribution patterns of geotechnical engineering projects based on survey data. Currently, it is difficult to set an accurate data compression threshold for survey data compression, resulting in geotechnical engineering survey data requiring significant storage space and bandwidth resources during archiving and backup. The following embodiments of the present invention will specifically address these issues.

[0049] See also Figure 1 , Figure 1 This is a flow chart of a method for archiving and backing up survey data for geotechnical engineering, provided in accordance with one embodiment of the present invention. The method can be applied to a survey data processing terminal to implement data archiving and backing up. The processing terminal can be a computer or server that can perform data processing. The method includes:

[0050] S11. Acquire multiple types of geotechnical engineering survey data, wherein each type of survey data corresponds to a single survey category of the geotechnical engineering.

[0051] Specifically, survey data can be obtained based on on-site survey tasks of geotechnical engineering, such as drilling holes at predetermined survey sites using a drilling rig, obtaining geotechnical samples from different depths, and collecting multi-dimensional data from the geotechnical samples. The survey data of each dimension is the survey data of a survey category. The multiple categories of survey data specifically include groundwater level monitoring data, soil moisture data, slope displacement, soil stress data, and surface settlement data. Among them, the multiple categories of survey data can be processed by a decimal calibration standardization method so that they have similar scales and distributions for better comparison. Decimal calibration standardization is a well-known technology, and the specific method will not be introduced here.

[0052] S12. Obtain multiple target time periods for multiple types of survey data based on fluctuation characteristics of data values ​​of each type of survey data as the data is collected over time. The target time period is a time period in which the degree of data anomaly represented by each type of survey data is the greatest.

[0053] Specifically, many phenomena in geotechnical engineering, such as soil settlement and groundwater changes, usually change slowly. If the survey data suddenly changes drastically, that is, the absolute value of the difference is large, it may mean that the soil or other engineering parameters have undergone a mutation, which is a warning signal for the safety of geotechnical engineering. Therefore, this situation requires special attention.

[0054] The fluctuation characteristics analysis of data can be performed based on each type of survey data, for example, the fluctuation characteristics of the w-th category survey data in the time series can be calculated. data minus the The absolute value of the difference between the data is normalized, and the moments with a difference absolute value greater than 0.6 are recorded as abnormal moments, and the consecutive adjacent abnormal moments are recorded as a period, namely the target period.

[0055] Exemplarily, step S12 includes sub-steps S12-1 to S12-4, which are described in detail as follows:

[0056] S12-1. If the degree of abnormal fluctuation in the data values ​​of each type of survey data during a preset collection period is greater than a first threshold, the collection period of the corresponding survey data is determined to be an abnormal period. The preset collection period and the first threshold can be set based on actual needs, as long as they can accurately filter out abnormal data. Fluctuations in the data values ​​of the survey data during the preset collection period indicate a possible risk of geological changes, requiring special attention to the survey data collected under such changes. The collection period of the corresponding survey data is then determined to be an abnormal period.

[0057] When calculating the abnormal fluctuation degree of each type of survey data, the formula , obtain the abnormal fluctuation degree of the survey category w in the preset collection period s ;in, represents the duration of the preset collection period s for survey category w; It represents the maximum abnormal value of all abnormal moments of survey category w in the preset collection period s. The larger the maximum value, the greater the possibility of abnormality in this period. represents the range of survey category w within the preset collection period s. A larger range indicates greater fluctuation in the survey data for that category, and a higher likelihood of anomalies. Based on the above formula, we can accurately calculate the degree of abnormal fluctuation for each type of survey data and, in turn, determine the abnormal period. The abnormal period represents the collection period corresponding to the data with the largest fluctuation characteristics within the acquired survey data.

[0058] Understandably, in geotechnical engineering survey data spanning multiple categories (or dimensions), each category has its own characteristics and normal range of variation, and abnormalities may only occur in one or a few specific categories. First, analyzing the degree of abnormality in a single category can quickly pinpoint the time period when the abnormality occurred. Then, combining the abnormal time periods of time series data from multiple categories can determine the target time period with the greatest analytical value.

[0059] S12-2. If survey data from multiple survey categories exists during the same abnormal time period, it indicates that geotechnical engineering may be affected by multiple dimensions, and the survey data during this period needs to be focused on. For example, heavy rainfall can cause the groundwater level to rise, which in turn increases the water content of the rock and soil, thereby increasing soil moisture data. At the same time, the increased water content of the rock and soil may change its mechanical properties, causing changes in the strain data that characterizes the stress of the rock and soil. If the rock and soil is located on a slope, the increased water content may also reduce the shear strength of the slope soil, thereby causing changes in slope displacement. It is important to pay attention to survey data from multiple survey categories during the same abnormal time period. Therefore, multiple survey categories within the same abnormal time period are identified as an abnormal category set.

[0060] S12-3. Based on the Pearson correlation coefficient, time overlap length, and anomaly period difference between the two survey categories in the anomaly category set, the degree of correlation between the two survey categories is obtained. This degree of correlation can be used to determine whether there is a strong correlation between the two survey categories in the anomaly category set. It can be understood that the stronger the correlation between the survey data of multiple categories within the anomaly period, the more significant the impact of a certain issue causing anomalies in multiple categories of geotechnical engineering data, which requires more focused analysis. If the correlation between the survey data of multiple categories within the anomaly period is weak, it means that the issue only caused anomalies in a single category of survey data, resulting in a relatively small impact.

[0061] It should be noted that for any two abnormal survey categories at time t, the interpolation algorithm can be used to make the time periods of the two categories at the t-th moment equal in length, that is, to interpolate the shorter time period and calculate the Pearson correlation coefficient of the two equal-length time periods. The time overlap length is the overlapping time length between two abnormal categories in the same abnormal time period, and the abnormal time period difference is the absolute value of the difference between the time lengths of the two abnormal categories.

[0062] For example, the correlation between two survey categories is calculated according to the formula Calculated, is the correlation degree between any two survey categories, and the correlation degree between survey category w and survey category x at time t. is the Pearson correlation coefficient between survey category w and survey category x at time t, is the abnormal period difference between survey category w and survey category x at time t, is the time overlap length between survey category w and survey category x at time t, is a normalization function. t is a natural number greater than 0, and w and x are natural numbers greater than 0 but not equal. It should be noted that, in order to ensure that the calculation results are meaningful, when performing fractional operations in the embodiment of the present invention, when encountering a situation where the denominator is 0, a parameter adjustment factor greater than 0 needs to be added to the denominator to prevent the denominator from being 0. The value of the parameter adjustment factor is set by the implementer according to the actual situation, and this application does not impose any special restrictions.

[0063] From the above formula, we can see that the Pearson correlation coefficient The larger the value is, the stronger the correlation between the data of the two periods is, and the higher the abnormality at time t is. In the subsequent compression process, the survey data of multiple categories corresponding to this moment should be retained; The larger the difference in the length of the representation, the more data the interpolation algorithm supplements, and the less credible the data; the length of the time overlap Represents the intersection of two time periods. When there is no intersection, let is 1.

[0064] S12-4. Perform comprehensive anomaly analysis and clustering based on all correlation levels, and determine the time period corresponding to the clustering results as the target time period. Performing comprehensive anomaly analysis across all correlation levels can more accurately determine the correlation between categories. Clustering can then be performed to achieve category division and determine the target time period.

[0065] Exemplarily, step S12-4 includes:

[0066] In the first step, when the degree of correlation is greater than the second threshold, the corresponding abnormal period is determined as the reference period. The second threshold can be set based on actual needs, for example, it is set to 0.7. According to the above method, the correlation between the survey category w and the survey category x at time t is obtained, that is, the degree of correlation. The correlation degree greater than 0.7 is recorded as the reference period of the survey category w at time t. The reference period represents the time period with more important reference value for the abnormal survey category w. For example, in rock engineering, landslides will be correlated with the anomalies caused by multiple dimensions. In order to obtain the time period in which the survey data of multiple categories show abnormalities, it is necessary to analyze which time periods have a greater correlation with each category, that is, the reference period of each time period. The more reference time periods corresponding to each time period under each category at the same time, the more likely a landslide will occur at that moment.

[0067] The second step is based on the formula , obtain the comprehensive abnormality degree at time t .in, is the number of first abnormal categories in the abnormal period at time t. The greater the number, the greater the weight of the moment. W1 is the number of second abnormal categories in the abnormal category set at time t. is the abnormality degree of survey category w at time t. The greater the abnormality degree, the more abnormal the survey data of this category is during the period. is the average value of the correlation between the survey category w and other survey categories in the abnormal category set at time t. The larger the mean value, the more similar the survey data of multiple categories are; is the number of reference periods at time t. The larger the number, the more similar other categories in the abnormal category set are to the survey category w.

[0068] In the third step, clustering is performed using the absolute difference between any two of the comprehensive anomaly levels as the clustering distance. The time period corresponding to the clustering result is determined as the target time period. Clustering uses the Kmeans clustering algorithm to cluster the absolute difference values, resulting in several clusters. Each cluster represents a relatively close absolute difference value. After clustering, the absolute difference values ​​of multiple clusters are clustered together. The time period corresponding to the abnormal moment contained in each cluster is the target time period.

[0069] S13. Analyze the causes of abnormalities in the survey categories corresponding to each target period to obtain the importance of abnormalities in each target period.

[0070] Specifically, the above-mentioned target time periods are divided according to the comprehensive anomaly level at each moment, and the comprehensive anomaly level at all moments within each target time period is similar. Anomalies in geotechnical engineering survey data can be caused by a variety of factors, such as human error and changes in natural conditions. In geotechnical engineering survey data, if the categories of abnormal data are stable, this may mean that the anomaly originates from a fixed source, and the characteristics and patterns of the abnormal data are relatively easy to identify and track. If the categories of abnormal data are constantly changing, it means that the causes of the anomaly are diverse, and in this case, more attention is needed.

[0071] By analyzing the causes of anomalies within the target survey period, we can determine whether the survey data within that period warrants focused analysis and quantify the importance of the anomaly. To perform anomaly analysis, we can build and train a convolutional neural network model. After training, we input the survey categories and corresponding survey data for the target period into the convolutional neural network model. Based on the model's output, we determine the importance of the anomaly for that period.

[0072] Exemplarily, step S13 includes sub-steps S13-1 to S13-3, which are described in detail as follows:

[0073] S13-1. Perform set partitioning and intersection processing on the survey categories corresponding to each target period. Based on the processing results, determine whether each target period is a single-type anomaly source or a diverse anomaly source. Each moment in each target period can be partitioned into sets, and the resulting sets can be intersected. This intersection processing can determine whether the anomalous moment persists in the partitioned sets, thereby analyzing whether the target period is a single-type anomaly source or a diverse anomaly source.

[0074] Sub-step S13-1 specifically includes:

[0075] The first step is to divide the survey categories of each abnormal moment in each target period into a single set to obtain the abnormal category set of each abnormal moment in each target period. For example, is the category set at the first moment in the e-th target period, is the category set at the second moment in the e-th target period, is the category set at the third moment in the e-th target period, is the category set at the fourth moment in the e-th target period, is the category set at the 5th moment in the eth target period.

[0076] The second step is to calculate the intersection of all abnormal category sets of the preset duration in each target period to obtain the processing results of each target period. For example, the intersection of all category sets is {B, C}, Denote it as the target set for intersection processing, that is, the processing result of the target period e. The duration when both survey categories B and C exist is the overlapping anomaly duration.

[0077] In the third step, if the duration of the overlapping anomalies represented by the processing results is greater than the third threshold, the survey category of the abnormal data is stable, and the target period is determined to be a single-type anomaly source. If the duration of the overlapping anomalies represented by the processing results is less than or equal to the third threshold, the survey category of the abnormal data is constantly changing, and the target period is determined to be a diversified anomaly source.

[0078] S13-2. Obtain an anomaly factor for each target time period based on the duration of overlapping anomalies characterized by a single anomaly source or multiple anomaly sources. The cause of the survey data anomaly can be determined based on the duration of overlapping anomalies. The anomaly factor can then be determined based on the cause of the anomaly. A correspondence table can be constructed based on the correspondence between the duration of overlapping anomalies and the anomaly factors. After determining the current duration of overlapping anomalies, the anomaly factor can be determined from the correspondence table.

[0079] The anomaly factor can also be calculated according to the formula , get the abnormal factor of the e-th target period .in, is the maximum abnormal degree of all abnormal moments in the e-th target period, K is the number of overlapping abnormal durations in the e-th target period, The duration of overlapping anomalies represented by a single type of anomaly source or multiple anomaly sources, is the number of survey categories in the abnormal category set with the kth overlapping abnormal duration in the eth target period, It represents the number of all survey categories with the kth overlapping abnormal duration in the eth target period. It can be understood that The larger the value is, the more survey categories with the kth overlapping abnormal duration in the eth target period are, and the more factors are involved in the abnormal data. Taking landslide in geotechnical engineering as an example, The larger it is, the more dimensions are affected by the landslide and the greater the anomaly during that period.

[0080] S13-3. Obtain the abnormal importance of each target period based on the abnormal factor of each target period. The abnormal importance of each target period can be derived based on the corresponding relationship between the abnormal factor and the abnormal importance. The abnormal importance can also be further normalized to meet the requirements of the subsequent configuration threshold.

[0081] In geotechnical engineering survey data, anomalies in one dimension may accumulate over time due to cumulative effects, gradually affecting other dimensions. For example, a long-term rise in groundwater levels, such as natural disasters like earthquakes and floods, may first affect data in one dimension and then, through complex geological processes, affect other dimensions. Therefore, further analysis is needed to determine how abnormal trends in geotechnical engineering survey data lead to increased soil saturation, which in turn affects soil mechanical properties. Therefore, determining the corresponding anomaly importance based on the anomaly factor during the target period, as described above, may be inaccurate.

[0082] Based on this, in a specific embodiment, according to the formula , obtain the abnormal importance of the e-th target period .in, is the abnormal factor of the e-th target period; Q is the number of target periods after the e-th target period; is the time interval between the e-th target period and the q-th target period; is the number of difference categories between the e-th target period and the q-th target period. The greater the number of difference categories, the more the survey category in the e-th target period gradually increases the influence of subsequent other category data. The smaller it is, the faster the impact is and the more important the e-th period is. In the time series, the target period e is before the target period q, and there is at least one target period between them.

[0083] S14. According to the abnormal importance of each target period, a threshold value for implementing data compression in the corresponding target period is configured.

[0084] Specifically, set the minimum threshold , the threshold adaptive size is .Will As the threshold value for implementing data compression in the e-th target period, the threshold values ​​of all target periods are obtained in the same way.

[0085] S15. According to the threshold value of each target period, the survey data of the target period is compressed and archived for backup.

[0086] Specifically, a revolving door algorithm is used to compress collected anomaly survey data to reduce storage capacity. The more important the target period, the smaller the corresponding threshold value, which reduces the compression rate but improves compression accuracy, preserving more original data points. The less important the target period, the larger the corresponding threshold value, and more data points are considered to be within an acceptable error range. Consequently, more data points are compressed, increasing the compression rate.

[0087] It should be noted that the above method is for compressing abnormal data in survey data. This type of data has greater analytical value and requires retaining more original data points. For compressing normal data in survey data, you can set a maximum threshold value. The non-target time period (normal time period, i.e. the possibility of landslide occurrence is very small) is compressed.

[0088] The following is an overall description of the archiving and backup method for geotechnical engineering survey data, specifically as follows:

[0089] S101. Obtain various types of survey data for geotechnical engineering.

[0090] S102. Filter out several abnormal time periods from the time series of survey data of each category, derive a target time period based on the abnormal time period with the maximum abnormality, and calculate the abnormality degree of each category of survey data in the abnormal time period.

[0091] S103: Calculate the correlation between abnormal time periods of any two categories in which the target time period is located.

[0092] S104. Calculate the comprehensive abnormality level at each moment in the target period.

[0093] S105. Filter out several target time periods based on the comprehensive abnormality degree.

[0094] S106. Calculate the abnormal factor for each target period.

[0095] S107: Calculate the abnormal importance of each target period.

[0096] S108: Configure threshold values ​​to compress abnormal data and archive and back up the data.

[0097] Based on the same technical concept as the backup method, the embodiment of the present invention also provides a geotechnical engineering survey data archiving and backup system, please refer to Figure 2 , Figure 2 This is a structural diagram of a backup system, which includes a collection terminal 1 and a processing terminal 2.

[0098] The acquisition terminal 1 is used to collect multiple types of survey data of geotechnical engineering and output them to the outside, wherein each type of survey data corresponds to a single survey category of geotechnical engineering. The acquisition terminal can be various sensors, etc., which obtain survey data based on the survey tasks of geotechnical engineering.

[0099] The processing terminal 2 is connected to the collection terminal 1; the processing terminal 2 is used to obtain multiple target time periods of multiple types of survey data based on the fluctuation characteristics of the data values ​​of each type of survey data as the collection time changes, wherein the target time period is the time period with the greatest degree of data anomaly represented by each type of survey data; the processing terminal is also used to analyze the causes of anomalies in the survey categories corresponding to each target time period to obtain the importance of anomalies in each target time period; according to the importance of anomalies in each target time period, a threshold value for implementing data compression in the corresponding target time period is configured; and according to the threshold value of each target time period, the survey data of the corresponding target time period is compressed and archived for backup.

[0100] Based on the same technical concept as the backup method, an embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of any one of the backup methods when executed by a processor.

[0101] It should be noted that the order in which the embodiments of the present invention are described above is for illustrative purposes only and does not necessarily represent the superiority or inferiority of the embodiments. The processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0102] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments.

Claims

1. A method for archiving and backing up survey data for geotechnical engineering, characterized in that: The method comprises: Acquire multiple types of geotechnical engineering survey data, wherein each type of survey data corresponds to a single survey category of the geotechnical engineering; Obtaining multiple target time periods of the multiple types of survey data based on fluctuation characteristics of the data values ​​of each type of survey data as the data is collected over time, wherein the target time period is a time period in which the degree of data anomaly represented by each type of survey data is the greatest; Perform set partitioning and intersection processing on the survey categories corresponding to each target period in turn, and determine whether each target period is a single-type anomaly source or a diverse anomaly source based on the processing results; obtain the anomaly factor of each target period based on the overlapping anomaly duration represented by the single-type anomaly source or the diverse anomaly source; and obtain the anomaly importance of each target period based on the anomaly factor of each target period; According to the abnormal importance of each target period, configure the threshold value for implementing data compression in the corresponding target period; According to the threshold value of each target period, the survey data of the target period is compressed and archived; The step of obtaining a plurality of target time periods of the plurality of types of survey data according to the fluctuation characteristics of the data values ​​of each type of survey data as the data values ​​change with the acquisition time includes: If the abnormal fluctuation degree of the data value of each type of survey data in the preset collection period is greater than a first threshold, determining the collection time period of the corresponding survey data as an abnormal period; If there are survey data of multiple survey categories in the same abnormal period, the multiple survey categories in the same abnormal period are determined as an abnormal category set; Obtaining the degree of correlation between the two survey categories according to the Pearson correlation coefficient, time overlap length, and abnormal period difference of the two survey categories in the abnormal category set; Comprehensive anomaly analysis and clustering processing are performed according to all correlation degrees, and the time period corresponding to the clustering result is determined as the target time period.

2. The geotechnical engineering survey data archiving and backup method according to claim 1, characterized in that: The obtaining of the correlation degree between the two survey categories according to the Pearson correlation coefficient, the time overlap length and the abnormal period difference of the two survey categories in the abnormal category set includes: According to the formula , obtain the correlation degree between any two survey categories ,in, is the Pearson correlation coefficient between survey category w and survey category x at time t, is the abnormal period difference between survey category w and survey category x at time t, is the time overlap length between survey category w and survey category x at time t, is the normalization function.

3. The geotechnical engineering survey data archiving and backup method according to claim 1, characterized in that: The comprehensive anomaly analysis and clustering process is performed according to all correlation degrees, and the time period corresponding to the clustering result is determined as the target time period, including: When the correlation degree is greater than a second threshold, determining the corresponding abnormal time period as a reference time period; According to the formula , obtain the comprehensive abnormality degree at time t ,in, is the number of first abnormal categories in the abnormal period at time t, W1 is the number of second abnormal categories in the abnormal category set at time t, is the abnormality degree of survey category w at time t, is the average value of the correlation between the survey category w and other survey categories in the abnormal category set at time t, is the number of reference periods at time t; Clustering processing is performed according to the absolute value of the difference between any two comprehensive abnormality levels among all comprehensive abnormality levels as the clustering distance, and the time period corresponding to the clustering result is determined as the target time period.

4. The geotechnical engineering survey data archiving and backup method according to claim 1, characterized in that: The survey categories corresponding to each target period are sequentially divided into sets and intersection-finding processes are performed, and each target period is determined to be a single type of abnormal source or a plurality of abnormal sources according to the processing results, including: Divide the survey categories of each abnormal moment in each target period into a single set to obtain an abnormal category set for each abnormal moment in each target period; Calculate the intersection of all abnormal category sets of the preset duration in each target period to obtain the processing results of each target period; If the overlapping anomaly duration represented by the processing result is greater than a third threshold, determining that the target period is a single type of anomaly source; If the overlapping abnormality duration represented by the processing result is less than or equal to a third threshold, the target time period is determined to be a diversity abnormality source.

5. The geotechnical engineering survey data archiving and backup method according to claim 1, characterized in that: Obtaining the abnormality factor for each target period according to the overlapping abnormality duration represented by the single abnormality source or the diverse abnormality sources includes: According to the formula , get the abnormal factor of the e-th target period ,in, is the maximum abnormal degree of all abnormal moments in the e-th target period, K is the number of overlapping abnormal durations in the e-th target period, The duration of overlapping anomalies represented by the single type of anomaly source or the multiple anomaly sources, is the number of survey categories in the abnormal category set with the kth overlapping abnormal duration in the eth target period, Represents the number of all survey categories with the kth overlapping abnormal duration in the eth target period.

6. The geotechnical engineering survey data archiving and backup method according to claim 1, characterized in that: The step of obtaining the abnormal importance of each target period according to the abnormal factor of each target period includes: According to the formula , obtain the abnormal importance of the e-th target period ,in, is the abnormal factor of the e-th target period, Q is the number of target periods after the e-th target period, is the time interval between the e-th target period and the q-th target period, is the number of difference categories between the e-th target period and the q-th target period, where e is a natural number greater than 0.

7. A geotechnical engineering survey data archiving and backup system, characterized in that: The system comprises: A collection terminal is used to collect multiple types of geotechnical engineering survey data and output them to the outside, wherein each type of survey data corresponds to a single survey category of the geotechnical engineering; a processing terminal connected to the acquisition terminal; the processing terminal is used to obtain multiple target time periods of the multiple types of survey data based on the fluctuation characteristics of the data values ​​of each type of survey data changing with the acquisition time, wherein the target time period is the time period in which the data anomaly represented by each type of survey data is the greatest; The step of obtaining a plurality of target time periods of the plurality of types of survey data according to the fluctuation characteristics of the data values ​​of each type of survey data as the data values ​​change with the acquisition time includes: If the abnormal fluctuation degree of the data value of each type of survey data in the preset collection period is greater than a first threshold, determining the collection time period of the corresponding survey data as an abnormal period; If there are survey data of multiple survey categories in the same abnormal period, the multiple survey categories in the same abnormal period are determined as an abnormal category set; Obtaining the degree of correlation between the two survey categories according to the Pearson correlation coefficient, time overlap length, and abnormal period difference of the two survey categories in the abnormal category set; Performing comprehensive anomaly analysis and clustering processing based on all correlation levels, and determining the time period corresponding to the clustering results as the target time period; The processing terminal is also used to perform set division and intersection processing on the survey categories corresponding to each target time period in turn, and determine whether each target time period is a single-type abnormal source or a diverse abnormal source based on the processing results; obtain the abnormal factor of each target time period based on the overlapping abnormality duration represented by the single-type abnormal source or the diverse abnormal source; obtain the abnormal importance of each target time period based on the abnormal factor of each target time period; configure the threshold value for implementing data compression for the corresponding target time period based on the abnormal importance of each target time period; and compress the survey data of the target time period based on the threshold value of each target time period and archive and back up it.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Internal arteriovenous fistula remote monitoring and warning system

    CN116719983A

  • Abnormal data cleaning and eliminating method and system for big data analysis

    CN118797243A