Urban geological survey method based on big data cloud computing technology
By adopting big data cloud computing technology in urban geological surveys, data hierarchical storage and compression and dynamically dispatching computing resources, the problem of low data storage and processing efficiency in traditional urban geological surveys is solved, and data processing efficiency and accuracy are improved.
Patent Information
- Application Number
- CN202510118141.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-06-20
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In traditional urban geological surveys, the data storage method is fixed, and the importance of data is not distinguished, resulting in the risk of data quality loss in key areas, and data in non-critical areas occupy too much storage space; during data processing, the proportion of computing resource allocation is fixed, resulting in excessive loading of some processing units and insufficient utilization of the processing capacity of the other part.
The urban geological survey method based on big data cloud computing technology is adopted, and a multi-level storage partition mechanism is established by grouping urban geological data by geographic grid units and time points to realize data hierarchical storage and compression; key monitoring areas are screened according to the threshold value of the geological parameter change rate, and differentiated compression storage strategies are adopted for different regions; the geological parameter data flow is classified according to the calculation load, similar computing tasks are merged, and data processing pipelines are set up, computing resources are dynamically allocated, and a cache prefetch mechanism is established.
The data storage space is optimized and data access time is reduced. Through differentiated compression storage strategies, the overall storage efficiency is optimized. Through dynamic task scheduling and resource allocation, the computing node load is balanced, the processing capacity is fully utilized, and data processing efficiency and accuracy are improved.
Smart Images

Figure CN120179644A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of urban geological surveys, and particularly to an urban geological survey method based on big data cloud computing technology. Background Art
[0002] Urban geological survey is a professional technical field that systematically investigates and evaluates the stratigraphic structure, geotechnical properties, groundwater, geological disasters, mineral resources, etc. in urban areas through means such as geological drilling, geophysical exploration, and surveying and mapping. This field involves multiple branches such as geological structure, hydrogeology, engineering geology, and environmental geology. By collecting and analyzing geological data, it reveals the characteristics of the urban underground space structure, providing basic geological data and decision-making basis for urban planning, engineering construction, underground space development and utilization, geological disaster prevention and control, etc. In traditional urban geological surveys, the data storage method is unified and fixed, without distinguishing the importance of data, resulting in a risk of quality loss of data in key areas, and excessive storage space occupied by data in non-key areas; during the data processing process, the characteristics of different types of task loads are not considered, and the calculation resource allocation ratio is fixed, causing some processing units to be overloaded and the processing capacity of some other parts to be not fully utilized. Therefore, improvements are needed. Summary of the Invention
[0003] The purpose of the present invention is to solve the deficiencies existing in the prior art, and to propose an urban geological survey method based on big data cloud computing technology.
[0004] To achieve the above purpose, the present invention adopts the following technical solutions. An urban geological survey method based on big data cloud computing technology includes the following steps:
[0005] Based on urban geological borehole detection data, ground settlement monitoring data, and underground pipeline network distribution data, group the data according to geographical grid units and time points, and perform statistical analysis to obtain a geological parameter statistical set; based on the geological parameter statistical set, set multiple storage partitions according to the data update frequency, and establish a geological data partition storage table;
[0006] Based on the geological data partition storage table, screen the areas where the change rate exceeds the geological parameter change threshold to obtain a set of key monitoring areas; perform lossless compression storage on the set of key monitoring areas, and perform lossy compression storage on the data in other areas to generate a geological risk area distribution matrix;
[0007] Based on the geological data partition storage table and the geological risk area distribution matrix, classify the geological parameter data streams according to the calculation load to obtain a geological data calculation classification table; merge similar calculations for the geological data calculation classification table, set up a data processing pipeline, establish a cache prefetch queue, and generate a geological parameter processing flow chart;
[0008] Based on the geological risk area distribution matrix and the geological parameter processing flow chart, allocate computing resources to the geological parameters according to the processing flow to obtain a geological computing node distribution table.
[0009] Preferably, the steps for obtaining the statistical set of geological parameters are as follows:
[0010] Obtain geological borehole exploration data, ground settlement monitoring data, and underground pipeline network distribution data, screen the information and classify it according to geographical grid units and time points to form a preliminary classification set of geological data;
[0011] Perform statistical analysis on the geological parameters of each geographical grid unit in the preliminary classification set of geological data, calculate the mean and standard deviation of the geological parameters of each geographical grid unit to obtain a statistical result set of geological parameters;
[0012] Based on the statistical result set of geological parameters, integrate the statistical data of all geographical grid units, and form a statistical set of geological parameters through data aggregation.
[0013] Preferably, the steps for obtaining the geological data partition storage table are as follows:
[0014] Based on the statistical set of geological parameters, count the number of data updates and distribution characteristics in each geographical grid unit to generate an analysis result of data update frequency;
[0015] According to the analysis result of data update frequency, design a multi-level storage partition scheme, specify the matching storage level by determining the update frequency of each grid unit, and generate a multi-level storage partition setting;
[0016] Based on the multi-level storage partition setting, summarize the storage locations and levels of data of different geographical grid units stored according to the update frequency, and construct a geological data partition storage table.
[0017] Preferably, the steps for obtaining the key monitoring area set are as follows:
[0018] Based on the geological data partition storage table, extract the historical data of geological parameters in each geographical grid unit to generate a set of geological parameter change data;
[0019] According to the set of geological parameter change data, calculate the geological parameter change rate in the geographical grid unit. The calculation formula is:
[0020]
[0021] Among them, R i represents the geological parameter change rate of the i-th grid unit, ΔP i,t represents the difference in geological parameters of the i-th grid unit at two consecutive time points, represents the standard deviation of the difference, Pi,k is the geological parameter value of the i-th unit at the k-th time point, is the average value of the values, and N is the number of time points;
[0022] Based on the geological parameter change rate, set a geological parameter change threshold, screen the geographical grid units whose change rate exceeds the geological parameter change threshold, and form a set of key monitoring areas.
[0023] Preferably, the steps for obtaining the geological risk area distribution matrix are as follows:
[0024] Based on the set of key monitoring areas, analyze data redundancy items and remove irrelevant content, perform lossless compression on each geographical grid unit, and generate a set of compressed data for key monitoring areas;
[0025] Based on the set of compressed data for key monitoring areas, perform lossy compression processing on the data of other geographical grid units, merge the processing results into the set of compressed data for key monitoring areas, and generate a geological risk area distribution matrix according to geographical location and risk level.
[0026] Preferably, the steps for obtaining the geological data calculation classification table are as follows:
[0027] Based on the geological data partition storage table and the geological risk area distribution matrix, extract the geological parameter data streams of each geographical grid unit, analyze the calculation requirements of each parameter in the geological parameter data streams, identify the processing characteristics of calculation-intensive and IO-intensive data, and generate a set of geological parameter data streams including calculation requirement classifications;
[0028] According to the set of geological parameter data streams, classify the data streams according to calculation requirements. For calculation-intensive data, perform item-by-item vectorization operations to gradually optimize the calculation efficiency. For IO-intensive data, integrate read and write operations and execute them in batches to generate a set of geological parameter data streams after classification processing according to the load;
[0029] Based on the set of geological parameter data streams, organize and standardize all classified data, merge the processing results of calculation-intensive and IO-intensive data, and form a geological data calculation classification table.
[0030] Preferably, the steps for obtaining the geological parameter processing flow chart are as follows:
[0031] Based on the geological data calculation classification table, analyze the task characteristics within each calculation category, group the tasks according to calculation types, and integrate the processing methods for the same type of calculations to generate a set of merged calculation classification tasks;
[0032] According to the set of merged calculation classification tasks, calculate the optimization value of the data processing pipeline. The calculation formula is:
[0033]
[0034] Among them, P opt is the optimized value of the data processing pipeline, CT j is the set of processing times for the j-th type of task, Q j is the computing density of the j-th type of task, CD j,k is the k-th data processing delay of the j-th type of task, m is the number of task classifications, and l is the number of operations included in each type of task;
[0035] Based on the optimized value of the data processing pipeline, a cache prefetch queue is constructed to optimize the scheduling and distribution of the data stream, and the execution order and flow direction of all tasks are organized into a geological parameter processing flow chart.
[0036] Preferably, the steps for obtaining the geological computing node distribution table are as follows:
[0037] Based on the geological risk area distribution matrix and the geological parameter processing flow chart, the computing task flow information is extracted, and the computing requirements of each area, including the task volume, processing complexity, and time distribution, are analyzed to generate a regional computing requirement analysis table;
[0038] According to the regional computing requirement analysis table, the number of computing node requirements is calculated, and the calculation formula is:
[0039]
[0040] Among them, N i is the number of computing nodes required for the i-th monitoring area, T i,t is the total task volume of the i-th area at time point t, R i,t is the task complexity parameter of the i-th area at time point t, H i,t is the data flow volume of the i-th area at time point t, A t and B t are the system load and data redundancy parameter at time point t respectively, D i,t is the data transmission delay of the i-th area at time point t, and K is the total number of time points;
[0041] Based on the number of node requirements, computing nodes are allocated to form a geological computing node distribution table.
[0042] Compared with the prior art, the advantages and positive effects of the present invention are as follows:
[0043] By grouping and processing urban geological data according to geographical grid units and time points, the present invention establishes a multi-level storage partition mechanism to achieve hierarchical data storage and compression, reduce storage space occupancy, and accelerate data access speed; based on the threshold of geological parameter change rate, key monitoring areas are screened, and different differential compression storage strategies are adopted for different areas to optimize the overall storage efficiency on the basis of ensuring the integrity of important data; the geological parameter data stream is classified according to the computing load, similar computing tasks are merged, and a data processing pipeline is set up to reduce duplicate calculations and improve data processing efficiency; computing resources are dynamically allocated according to the processing flow direction, and a cache prefetch mechanism is established to accelerate data reading and computing processes. The multi-level data compression storage and distributed parallel processing mechanism make the calculation of geological parameters more efficient, the identification of risk areas more timely, the data processing accuracy higher, and the space utilization rate better. Through dynamic task scheduling and resource allocation strategies, the load of computing nodes is more balanced, the processing capacity is fully exerted, and the overall operation state is more stable. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 It is a schematic diagram of the steps of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0045] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0046] Please refer to Figure 1 , the present invention provides a technical solution, a method for urban geological survey based on big data cloud computing technology, including the following steps:
[0047] Based on urban geological borehole exploration data, ground settlement monitoring data and underground pipeline network distribution data, the data is grouped according to geographical grid units and time points, and statistical analysis is carried out to obtain a statistical set of geological parameters; based on the statistical set of geological parameters, multi-level storage partitions are set according to the data update frequency, and a geological data partition storage table is established;
[0048] Based on the geological data partition storage table, areas where the change rate exceeds the geological parameter change threshold are screened to obtain a set of key monitoring areas; the set of key monitoring areas is stored in a lossless compression manner, and the data in other areas is stored in a lossy compression manner to generate a geological risk area distribution matrix;
[0049] Based on the geological data partition storage table and the geological risk area distribution matrix, the geological parameter data stream is classified according to the computing load to obtain a geological data calculation classification table; similar calculations in the geological data calculation classification table are merged, a data processing pipeline is set up, and a cache prefetch queue is established to generate a geological parameter processing flow diagram;
[0050] Based on the geological risk area distribution matrix and the geological parameter processing flow chart, allocate computing resources for geological parameters according to the processing flow to obtain the geological computing node distribution table.
[0051] The steps for obtaining the statistical set of geological parameters are as follows:
[0052] Obtain geological borehole detection data, ground settlement monitoring data, and underground pipeline network distribution data, screen the information and classify it according to geographical grid units and time points to form a preliminary classification set of geological data;
[0053] Perform statistical analysis on the geological parameters of each geographical grid unit in the preliminary classification set of geological data, calculate the mean and standard deviation of the geological parameters of each geographical grid unit to obtain the statistical result set of geological parameters;
[0054] Based on the statistical result set of geological parameters, integrate the statistical data of all geographical grid units and form a statistical set of geological parameters through data aggregation.
[0055] Specifically, based on the obtained geological borehole detection data, ground settlement monitoring data, and underground pipeline network distribution data, first compare the numerical values of each data record with a pre-set interval. For example, the hole depth interval can be set from 0 meters to 3000 meters, and the settlement amount interval can be set from 0 meters to 10 centimeters. The upper limits of these intervals are formed by historical geological survey data and on-site evaluation, and the lower limits are determined by the minimum range of the instrument. When the hole depth or settlement amount of any record exceeds the corresponding interval, mark the record as abnormal and register the relevant time point. Subsequently, split the information of all the marked normal data and group it according to geographical grid units. During this process, a 500-meter square area can correspond to one grid unit, and multiple records at the same time point are merged into the same classification group according to a pre-determined time interval. Then, combine the pipeline route and laying depth information in the underground pipeline network distribution data to append pipeline position descriptions to each grid unit. If there are multiple data with similar coordinates but a large difference in recording time within the same grid unit, keep them as independent entries and supplement the time difference information. If there are entries with overlapping coordinates but incomplete parameter fields, store them separately for subsequent comparison. Then, establish an index at intervals of days or hours. When indexing, it is necessary to check the coordinates, time, and pipeline association fields one by one to confirm the matching relationship between the borehole information, settlement monitoring values, and pipeline routes within the grid unit, and there is no need to perform repeated cleaning operations. In this way, the three types of data are integrated and all archived at the same classification level, and finally, a preliminary classification set of geological data is obtained.
[0056] Based on the geological parameters of each geographical grid unit in the preliminary classification set of the obtained geological data, first summarize the geological parameters of the same type within the same grid unit and eliminate the marked abnormal data records. These eliminated records may include entries with hole depths exceeding the range of 0 meters to 3000 meters or settlement amounts greater than 10 centimeters. If more refined thresholds are required, they are adjusted according to the results of historical statistical analysis. Subsequently, when calculating the mean of all normal records within the grid unit, use the method of ∑x i / n, where x i represents the value of the geological parameter of the i-th record within the grid unit, and n represents the total number of records that meet the range conditions. Then calculate the standard deviation according to the method of , where represents the mean obtained in the previous step. When some records are repeated in multiple time periods, they are processed as independent records and classified and marked before calculating the mean and standard deviation. If missing or incomplete values are encountered, interpolation is performed based on the means of other records in the same period or the data of this record is directly excluded. Then integrate values such as the mean and standard deviation into the statistical fields of the corresponding grid unit, and maintain the index using the coordinate information of the same grid unit. No additional fine-grained interval divisions are generated. Through the above statistical analysis process, summarize the centralized distribution intervals of the geological parameters of each grid unit and record their fluctuation characteristics to obtain the geological parameter statistical result set.
[0057] When integrating the statistical data of all geographical grid units based on the geological parameter statistical result set, first conduct a vertical comparison of the mean and standard deviation information of each grid unit obtained in the previous step. If the means and standard deviations of multiple adjacent grid units show relatively close ranges in value, these grid units are classified into the same aggregation group, and record the corresponding coordinate index and time distribution of each aggregation group. Then, if there are some records in the same aggregation group that are eliminated due to abnormal marking, indicate the missing amount within the aggregation group, and make further processing according to whether the missing information can be compensated by the historical data of nearby grid units. If a unified threshold needs to be set to judge the aggregation effect, a reference value can be selected from the historical distribution and comprehensively inferred based on the ranges of the mean and standard deviation. Then assign corresponding identification codes to each aggregation group and combine them with the time points of the records. Subsequently, conduct cross-checks on all aggregation groups to confirm whether there are duplicate statistics or coordinate overlaps between different grid units, and record them in a new index set after verifying the information of each group. Through this aggregation method, the statistical data of all grid units can be further classified, and no further in-depth screening of the existing abnormal data is performed. Finally, generate the geological parameter statistical set through the data aggregation method.
[0058] The steps to obtain the geological data partition storage table are as follows:
[0059] Based on the statistical set of geological parameters, the number of data updates and distribution characteristics in each geographic grid unit are counted to generate data update frequency analysis results;
[0060] According to the data update frequency analysis results, a multi-level storage partitioning scheme is designed. The matching storage level is specified by determining the update frequency of each grid unit, and a multi-level storage partition setting is generated;
[0061] Based on the multi-level storage partition setting, the locations and levels of data stored in different geographic grid units according to the update frequency are summarized to construct a geological data partition storage table.
[0062] Specifically, based on the statistical set of geological parameters, the update records of each geographic grid unit in different time periods are first selected and the number of occurrences is summarized. The daily or hourly geological data change frequency is registered one by one under the corresponding grid. When the number of changes in a grid in a specific period exceeds a certain empirical threshold set according to historical statistical results, it can be judged as frequently updated. For example, if the empirical threshold is based on the comprehensive mean of the observation data in the past year plus 1 to 2 standard deviations to obtain an average daily update number of 10 times as the high-frequency limit, then when the number of grid updates on a certain day is higher than this value, it is marked as a high-frequency band, and the low-frequency band can correspond to the low daily update number. In the case of 2 times, the benchmark of 2 times is obtained by referring to the mean of the historical statistical analysis results minus 1 standard deviation. The middle range is used as the medium frequency mark. When dividing the update frequency categories, it is also necessary to take into account the time span of each grid and record the differences in the distribution of update times caused by factors such as weekends and weekdays or day and night. If some grids have extremely high frequencies or are not updated for a long period of time, they are marked separately for subsequent comparison and the possible reasons are annotated. Finally, the sorted update times and their distribution profiles are classified and summarized to form a set of corresponding tables with multiple labels such as high frequency, medium frequency, and low frequency, and indexes are established to generate data update frequency analysis results.
[0063] According to the results of data update frequency analysis, we first determine the priority divisions corresponding to the three update frequencies and list the number of grids contained in high frequency, medium frequency and low frequency. High frequency usually refers to grids with an average daily update frequency of more than 10 times, medium frequency falls between 2 and 10 times, and low frequency is less than 2 times. Each numerical threshold is set based on the statistical results of the previous stage and the average level of observations over the years. Then, hardware locations with faster read and write responses are configured for high-frequency grids, and appropriate cache space is reserved for medium-frequency grids. For low-frequency grids, basic hardware locations are allocated to meet the usual access requirements. When designing multi-level storage partitions, it is also necessary to record whether each type of grid will have significant fluctuations within a period of 24 hours or a week. For example, if some high-frequency grids experience a sudden drop in updates at night, their storage priority can be dynamically switched. If there are medium-frequency grids that frequently rise to high frequency during seasonal changes, corresponding storage space must also be reserved and the corresponding switching timing must be set for them. In addition to hardware selection, the storage level definition of each grid should also indicate the specific triggering conditions of the read and write strategy and cache strategy. In this way, all grids are traversed and relevant information is uniformly summarized to generate multi-level storage partition settings.
[0064] Based on the multi-level storage partition setting, each geographic grid is first allocated a storage area corresponding to high frequency or low frequency, and the medium frequency grid is marked as a flexibly adjustable state. Then, the storage level determined in the previous step, whether it carries a cache strategy, whether there is a timed switching scheme, and other information are read one by one according to the grid number or geographic coordinate order. These items are arranged in a fixed field order and their storage location and required hardware description are recorded. If some grids are found to be in a state of frequent fluctuation, they are separately included in the adjustable partition and their fluctuation cycle and the factors that may cause fluctuations are noted. If the average daily update number quickly crosses from low frequency to high frequency, a record is added to illustrate the historical evolution curve of the grid and mark the corresponding time range. If some grids have completely different update modes in different seasons, additional instructions are added to the medium frequency configuration. Finally, all grid information is deduplicated and consistency tested, and then archived and corresponding access identifiers or shared tags are established for each storage level. In this way, all partition information is centrally summarized and maintained in correspondence with the geographic grid number to construct a geological data partition storage table.
[0065] The steps to obtain the key monitoring area set are:
[0066] Based on the geological data partition storage table, the historical data of geological parameters in each geographic grid unit is extracted to generate a set of geological parameter change data;
[0067] According to the geological parameter change data set, the geological parameter change rate within the geographic grid unit is calculated using the following formula:
[0068]
[0069] where R i represents the change rate of the geological parameters of the i-th grid cell, and ΔP i,t represents the difference in geological parameters of the i-th grid cell at two consecutive time points. represents the standard deviation of the difference, and P i,k is the value of the geological parameter of the i-th cell at the k-th time point. is the average value of the values, and N is the number of time points;
[0070] Based on the change rate of geological parameters, a threshold for geological parameter change is set, and geographical grid cells with a change rate exceeding the threshold for geological parameter change are screened to form a set of key monitoring areas.
[0071] Specifically, based on the historical geological parameter data in the geological data partition storage table, first, the records associated with each geographical grid cell are selected one by one from the existing partition storage table and compared. When comparing, it is necessary to pay attention to whether the time stamps and coordinate information corresponding to each record are complete. If there is a record whose coordinates do not match the grid, it is excluded. If there is a situation where the time stamp differs from the previous record by less than 1 hour or more than 24 hours, corresponding marks are added and the continuity of this period is recorded. Subsequently, the geological parameter entries of the same grid cell within different time segments are centralized and compared on a daily or weekly basis as a time segment. During the comparison process, a reference interval can be set for parameters such as hole depth and settlement. For example, the hole depth interval can be selected from 0 meters to 3000 meters, and the settlement interval can be selected from 0 millimeters to 100 millimeters. These interval values are statistically obtained based on local earlier construction monitoring and multiple tests and are integrated. If there is a record exceeding the specified interval, the source of the abnormality is noted and the entry is separated for separate observation. After that, the mean or median of the parameters of the normal records within each time period is simply sorted and incorporated into an indexable time series set, and then the evolution process of the parameter values of the grid cell is listed in sequence in combination with the time stamp and arranged in chronological order. In this way, a continuous change sequence of geological parameters is formed. When the historical data of all grids have undergone the same processing process, the time series records corresponding to each grid are uniformly numbered and summarized, and finally, the continuous time series data of all grids are combined into a set of geological parameter change data.
[0072] The advantage of the formula is that it combines the absolute difference, standard deviation, and relative fluctuations in multiple time periods for comprehensive judgment. It can take into account the average level differences of different grid cells while analyzing the cumulative changes of geological parameters, and reflects the fluctuation degree of geological parameters between consecutive time periods by multiplying the square root of the mean square of the differences over the entire time period.
[0073] For the acquisition steps of the ΔP i,t parameter, in the previous step, the geological parameter values of each time point of the grid cell have been sorted out and recorded as etc. First, calculate the difference based on the records adjacent in timestamp and record the change amount between each adjacent time point. For example, when collecting data continuously for 30 days and recording the porosity or settlement once a day, 29 adjacent differences can be obtained. If the data quality on a certain day is deviated, mark the data source on the difference of that day and note the abnormal impact in the record. In order to obtain ΔP i,t , it is necessary to strictly confirm that the coordinates and collection time of the current period are exactly matched with those of the previous period, and then subtract the values collected twice. The values can be formed by merging the periodic monitoring results of multiple local collection devices. For example, for the settlement data of 30 consecutive days, the measurement results before and after are in millimeters, and after multiple calibrations, the measurement error is guaranteed to be within 0.2 millimeters. If 8.6 millimeters are measured in a certain period and 7.1 millimeters are measured in the previous period, then ΔP i,t = 1.5 millimeters. The ΔP i,t obtained in this process usually distributes in the range of 0 millimeters to 5 millimeters.
[0074] The steps for obtaining the parameter are as follows: First, perform a variance operation on all ΔP i,t differences within the same grid cell, and then take the square root of the variance to obtain the standard deviation. If 29 ΔP i,t differences are generated within 30 days, and let their mean be then we can let can be obtained by summing all ΔP i,t and then dividing by 29. The value is usually around 0.5 millimeters to 2 millimeters, and may be slightly larger for areas with obvious land subsidence. Once the statistics are completed, record it in the index table corresponding to the grid number for subsequent substitution into the formula.
[0075] P i,k The steps for obtaining the parameter are as follows: During the process of extracting geological parameters for consecutive periods, record the geological parameter values collected at each time point. This value can represent indicators such as porosity, settlement, or soil layer water content. For example, the settlement is measured in millimeters, and the porosity is recorded in decimal form. If the collection days are 30 days, then 30 P i,1 , P i,2 , …, P i,30 values will be generated. These values are usually directly read by on-site sensors and saved after one-time denoising. If it is found that the sensor has an error in a certain period, it is necessary to compare the records of adjacent periods or adjacent positions to exclude abnormal values. Each P i,k usually locates in the range of 0 millimeters to 10 millimeters in areas with relatively stable geological conditions.
[0076] The steps for obtaining the parameter are as follows: When obtaining all P i,kAfter that, add up their values and divide by the number of records actually involved in the calculation. If all records are valid for 30 days, then n = 30. If records in certain time periods are marked as abnormal, the corresponding data need to be removed. The obtained value will vary according to different ranges of geographical features. For example, in a certain place, the settlement range is from 0.2 mm to 5 mm. If the average of all records is 2.4 mm, then mm, and this parameter will ultimately be used for normalizing the differences between adjacent time periods.
[0077] The steps to obtain the N parameter are as follows: count the time points or the number of records required for calculation and ensure that the time intervals between records are consistent. If measured in days and the observation period is 30 days, then N = 29 available differences; if the frequency is once per hour and the observation period is 7 days, then N = 168. The value size is directly related to the observation duration, observation frequency, and whether there are abnormal records. This parameter needs to have a one-to-one correspondence with ΔP i,t and other contents to facilitate the summation calculation in.
[0078] Calculation process:
[0079] The following gives a practical example. When the observation period of a certain grid cell is 30 days and the acquisition frequency is once a day, then N = 29. Now select the ΔP of one day i,t = 1.5 mm. Assume that after variance calculation, we get mm, and mm. To calculate First, divide the difference between each pair of adjacent measurement values by 2.4, then square the results and sum them up. For example, assume that after performing the above operations on all 29 differences, the sum obtained is approximately 25.52, then:
[0080]
[0081] Then we can let:
[0082] This result indicates that the difference between the current grid cell on a certain day and the previous day has a relative change of 1.76 times compared to the average level and fluctuation of the full-cycle data of this grid. When R i is between 1.0 and 2.0, it may mean a normal moderate change. When R i exceeds 2.0, it indicates a relatively large change amplitude and further attention is needed.
[0083] Based on the geological parameter change rate, for all grid cells corresponding to R iValues are read sequentially and compared one by one with the pre-set thresholds for geological parameter changes, which are often determined based on statistical results formed from historical monitoring data. For example, in the observed data of the past year, the most obvious subsidence area can be selected, and the maximum R i value and the average R i value are subtracted and combined with the range of 1 to 2 standard deviations to delimit the numerical interval. If the lower limit within this interval falls around 1.0 and the upper limit falls near 2.0 or 2.5, grid cells with values greater than 2.0 or 2.5 can be marked as significantly abnormal and included in the key screening range, and grid cells within the range of 1.0 to 2.0 are classified as the medium change range. If in some areas, the R i value still exceeds 2.5, these areas are recorded separately and noted as possibly having excessive foundation settlement or rapid soil layer disturbance. Subsequently, the grid numbers of all grid cells exceeding the set threshold are summarized and cross-indexed with their corresponding coordinate positions, and the start and end times of the observation and the duration in days or hours are recorded to form preliminary abnormal information. Finally, these marked abnormal grids are grouped into a set of key monitoring areas.
[0084] The steps to obtain the geological risk area distribution matrix are as follows:
[0085] Based on the set of key monitoring areas, analyze the data redundancy items and remove the irrelevant content, perform lossless compression on each geographical grid cell, and generate a set of compressed data for key monitoring areas;
[0086] Based on the set of compressed data for key monitoring areas, perform lossy compression on the data of other geographical grid cells, merge the processing results into the set of compressed data for key monitoring areas, and generate the geological risk area distribution matrix according to the geographical location and risk level.
[0087] Specifically, based on the key monitoring area set, first select the data entries corresponding to the key monitoring area set obtained previously and confirm the redundant field content therefrom, including records that appear repeatedly but have the same timestamp or geographic location and additional information marked as low credibility, etc., and then set a standard threshold in the process of counting the number of occurrences of these redundant fields. For example, refer to the repetition distribution range of similar data in the same grid unit in the past three months and select a repetition rate of 80% as the basis for cleaning. When the repetition rate of a record reaches or exceeds 80%, it is removed from subsequent calculations and marked. After that, when the remaining data is losslessly compressed, a one-by-one comparison method can be used to reduce repeated paragraphs. For example, consecutive identical settlement or porosity records can be merged into one paragraph and carry the number of occurrences. If it is found that there are only slight differences in the records of adjacent time intervals, such as a numerical fluctuation of 0.1 mm to 0.2 mm, it is judged whether it is sufficient to distinguish according to the previously determined threshold. If it is not enough to distinguish, the paragraph is regarded as the same data and merged, and an additional mark is added to the list to indicate that this part of the data belongs to the same observation sequence. Through this compression method, all grid cells in the key monitoring area can be losslessly processed without deleting valid data. If information such as coordinates, timestamps, key stratigraphic features, etc. that need to maintain complete accuracy is encountered, its numerical intervals will not be merged. Then all the losslessly compressed contents will be uniformly sorted and attached with grid numbers and time series indexes. Finally, these processed data will be archived to generate a compressed data set for the key monitoring area.
[0088] Based on the compressed data set of key monitoring areas, the grid information corresponding to the non-key monitoring areas is first extracted from the previously obtained classified records and grouped according to the characteristics of pipeline distribution, settlement parameters or porosity parameters, and then a lossy compression strategy is selected to reduce the precision of the records in each group, such as reducing the numerical precision from three decimal places to two decimal places or one lower, and setting an applicable range in advance based on the field monitoring historical records, such as settlement between 0 mm and 100 mm, porosity between 0 and 1. If the collected value exceeds this range, it will be marked separately and the lossy compression process will not be performed. If it is within the range, Adjacent records are considered to be of the same type when the difference is less than 0.5 mm. When merging, only one representative value is retained and the timestamp corresponding to the representative value is recorded. If the difference in the porosity record is less than 0.05, the merging process is also performed. The processed lossy compression results are summarized and cross-checked with the compressed data set of the key monitoring area. If abnormal conditions such as timestamp overlap or coordinate errors are found, notes are added. Finally, all merged data entries are arranged into a matrix according to geographic location and risk level, with grid numbers in rows and risk levels in columns, and coordinate indexes and time span information are added to generate a geological risk area distribution matrix.
[0089] The steps for obtaining the geological data calculation classification table are as follows:
[0090] Based on the geological data partition storage table and the geological risk area distribution matrix, extract the geological parameter data streams of each geographical grid unit, analyze the calculation requirements of each parameter in the geological parameter data streams, identify the processing characteristics of calculation-intensive and IO-intensive data, and generate a set of geological parameter data streams containing calculation requirement classifications;
[0091] According to the set of geological parameter data streams, classify and process the data streams according to the calculation requirements. For calculation-intensive data, perform item-by-item vectorization operations to gradually optimize the calculation efficiency. For IO-intensive data, integrate read and write operations and execute them in batches to generate a set of geological parameter data streams after classification processing according to the load;
[0092] Based on the set of geological parameter data streams, sort and standardize all classified data, and merge the processing results of calculation-intensive and IO-intensive data to form a geological data calculation classification table.
[0093] Specifically, based on the geological data partition storage table and the geological risk area distribution matrix, first lock the geological parameter data streams of each geographical grid unit from them and sequentially read the original entries of multiple parameters such as settlement, porosity, and soil layer thickness in these data streams. Then, check these data streams in combination with the coordinate identification and timestamp of the grid unit. If it is found that the recorded time span or spatial coordinates are incomplete, mark the source of the anomaly and conduct subsequent investigations. Then, compare each parameter item by item with the predefined range. For example, compare the settlement with the range of 0 mm to 100 mm, compare the porosity with the range of 0 to 1, and compare the soil layer thickness with the range of 0 m to 2000 m. All entries that exceed these ranges obtained by adding and subtracting 2 standard deviations from the mean of historical measurement data are listed separately and carefully processed in subsequent analyses. Subsequently, check the processing requirements and IO requirements of these data entries during high-concurrency operations. The settlement data that requires frequent numerical calculations and matrix operations can be regarded as calculation-intensive, while the parameters that require frequent reading or writing of a large number of time series from the database tend to be IO-intensive. To avoid confusion, it is also necessary to continuously track whether multiple parameters under the same grid unit are updated at the same moment during this process and judge the final calculation load type based on the sampling frequency of the sensor and historical fluctuations. In this way, screen all grid units and their geological parameter entries one by one and mark the calculation-intensive or IO-intensive labels in the records. Finally, merge the entries containing these processing labels and arrange the table according to the grid numbers to generate a set of geological parameter data streams containing calculation requirement classifications.
[0094] According to the computationally intensive and IO-intensive labels in the geological parameter data stream collection, all data streams are first grouped and uniformly recorded in different operation pools. The computationally intensive grouping places parameters such as settlement and porosity that require a large number of multiplication, addition or matrix solving operations. Before performing vectorized operations, each record can be arranged in a one-dimensional array in chronological order and the time dimensions such as date or hour can be separately marked. If a grid cell contains dozens to hundreds of continuous settlement records, batch normalization is first performed on these records and mapped to a vector structure that can be executed quickly. After completing the accumulation and difference calculations item by item, the results are summarized. IO-intensive grouping classifies data that mainly rely on high-frequency read and write operations and need to be repeatedly accessed or appended from the storage medium, and merges entries with adjacent time periods or similar coordinates into batches during execution to reduce the overhead caused by multiple single accesses. If there are repeated coordinates in the same period of time, they can be directly integrated into a batch read and write action. If the coordinates are adjacent and the time is similar, they can also be merged. Finally, the computationally intensive results of the completed item-by-item vectorized operations and the IO-intensive results of the integrated read and write operations are output separately, and the grid unit number, timestamp and coordinate information are kept from being confused, generating a set of geological parameter data streams processed by load classification.
[0095] Based on the set of geological parameter data streams, when reorganizing and standardizing all records marked as computationally intensive or IO-intensive, fixed rows and columns can be set up according to each grid unit or each time period, and each record can be sorted according to the grid number and time sequence to ensure that the values of settlement, porosity, etc. are in the correct column. If multiple sets of settlement or porosity data appear in the same grid within the same hour, they can be marked in the order of records and merged into the corresponding rows. If there are missing records, they can be processed according to a certain interpolation strategy or null value mark. Then, when merging the computationally intensive results and the IO-intensive results, it is necessary to The connection order of the two in the same grid unit is confirmed according to the time distribution or coordinates. If the former contains multiple calculation results in one grid unit and the latter has multiple continuous reading actions, it is necessary to ensure that the order of the two is consistent during the integration stage. After all geological parameter items are uniformly arranged, necessary standardized conversions such as normalization or dimensional unification are performed on each field. Then, the final data table generation process is carried out according to the standardized information, and the corresponding grid number and completion time record are attached to each column or each record, so that the settlement or porosity results of the grid in a certain period of time can be quickly located in subsequent queries to form a geological data calculation classification table.
[0096] The steps to obtain the flow diagram of geological parameter processing are as follows:
[0097] Calculate a classification table based on geological data, analyze the task characteristics within each calculated category, group the tasks by calculation type, and integrate the processing methods for the same type of calculations to generate a merged set of calculation classification tasks;
[0098] Based on the merged set of calculation classification tasks, calculate the optimization value of the data processing pipeline. The calculation formula is:
[0099]
[0100] where P opt is the optimization value of the data processing pipeline, CT j is the set of processing times for the j-th type of task, Q j is the calculation density of the j-th type of task, CD j,k is the k-th data processing delay of the j-th type of task, m is the number of task classifications, and l is the number of operations included in each type of task;
[0101] Based on the optimization value of the data processing pipeline, construct a cache prefetch queue, optimize the scheduling and distribution of the data stream, and organize the execution order and flow direction of all tasks into a geological parameter processing flow diagram.
[0102] Specifically, based on the geological data calculation classification table, first split out the records related to settlement monitoring, pipeline distribution comparison, porosity statistics, etc. item by item according to the geological task requirements of different categories. During the execution process, it is necessary to first read the grid number, timestamp, and main parameter fields under each category, and compare these parameter entries with the pre-set reference intervals item by item. For example, compare the soil layer thickness with the range between 0 meters and 3000 meters, compare the settlement amount with the range between 0 millimeters and 100 millimeters, and compare the porosity with the range between 0 and 1. These range values are determined by adding and subtracting several standard deviations from the historical mean obtained from multiple local surveys and long-term observations. When the value of a record entry exceeds these ranges, it is marked as abnormal and added to the subsequent processing list. Then, arrange all data entries in order according to the time given by the monitoring equipment or recorders or the batch processing frequency. If there are multiple items with similar processing methods in the same category, they are merged within the category and the processing type is uniformly identified. If it is observed that both the settlement amount statistics requirement and the porosity calculation requirement need to perform batch numerical calculations and their sampling times or grid ranges are similar, they are regarded as an integrable task unit. Through this merging method, multiple scattered records under the same category are concentrated and arranged, and their key fields are confirmed. At the same time, entries with missing coordinates or timestamps are excluded, and content with less serious missing is tried to be interpolated or segmented. After all items have completed this splitting and checking, the processing methods for the same type of calculations can be sorted out and finally merged into an overall executable list to generate a merged set of calculation classification tasks.
[0103] The advantage of the formula is that it incorporates the processing time, computational density, and data processing latency of each operation of each type of task into the operation. By combining the time scale and resource load in the form of a numerator and a denominator, it can more clearly measure the overall efficiency of different task combinations in the pipeline.
[0104] CT j A parameter that represents the set of processing times for the j-th type of task. Specifically, it can be obtained by monitoring and summarizing the execution durations of multiple operations included in this task. For example, when performing batch operations on settlement data, the start time and end time of each operation within the past 30 days can be recorded, and then these time differences form a set of processing duration sequences. If the j-th type of task contains 20 operation processes, 20 duration values can be obtained, each value measured by on-site hardware devices, and the device accuracy can be achieved through millisecond-level timing in the log. Then, all duration values are put into the set as CT. j If some operations occur dozens of times a day, all duration values need to be integrated and then statistically analyzed to obtain the set of processing times for the j-th type of task.
[0105] Q j A parameter that represents the computational density of the j-th type of task and can be measured based on the amount of data operations within the unit processing time. For example, when performing vectorized operations on settlement amounts, if 1000 records need to be differenced and summed within 5 seconds, the amount of data operations can be set to 1000 key numerical calculations, and the time can be set to 5 seconds, thus obtaining a computational density of 200 per second for this period. Then, based on the overall performance of this task in multiple runs, an equilibrium calculation is performed. For example, the number of entries and the elapsed time recorded in all operation logs within the past week are weighted and averaged to obtain a relatively stable Q. j If it is found that the ratio of the number of entries to time fluctuates by 3 times during multiple runs, the highest and lowest values are excluded and then the remaining values are averaged to ensure that Q j can reflect the most common computational intensity of the j-th type of task.
[0106] CD j,k A parameter that represents the k-th data processing latency of the j-th type of task. When there are multiple operations under each type of task, each operation has a latency from the start request to the completion of reading and writing. For example, when performing batch reading of settlement records, a latency of 20 milliseconds to 50 milliseconds may occur, while when calculating the porosity matrix, a latency of 100 milliseconds to 200 milliseconds may occur. By frequently collecting these processing latencies and merging the most common latency values of the same operations during multiple runs, a stable value for calculation can be obtained, and these stable values are listed item by item as CD. j,k, if such a task contains l operations, l delay values need to be obtained, and attention should be paid to excluding the abnormal durations caused by extreme outliers or special faults. Finally, all delay values should have the same dimension of milliseconds or seconds so as to enter the formula for calculation.
[0107] m represents the number of task categories, which specifically depends on how many categories of tasks are divided in total when analyzing the merged set of computational classification tasks. For example, the batch operation of settlement amount is 1 category, the item-by-item comparison of porosity is the 2nd category, the data retrieval of soil layer thickness is the 3rd category, etc. The processing method, input type, and output scale can be divided and the total number can be counted through manual or system automatic determination methods. If three major categories are finally identified, then m = 3; if more branches are refined, then m may increase to 4 or 5, and the classification needs to be completed based on the actual characteristics of the task itself.
[0108] l represents the number of operations included in each category of tasks, and it needs to be disassembled in detail according to the execution process of each category of tasks. For example, the batch operation of settlement amount includes 3 operations: data difference, sum of squares, and average value statistics, so l = 3. If another task includes data reading, data interpolation, and data visualization, then l = 3. Similarly, if there are more processing steps, they are accumulated. Finally, after making the same analysis for each category of tasks, the corresponding l value is obtained.
[0109] Calculation process:
[0110] Here is an example scenario: There are currently m = 2 categories of geological data tasks, namely the batch operation of settlement amount and the matrix calculation of porosity. Let the average processing duration of the processing time set CT1 of the first category of tasks be 10 seconds, and the average processing duration of the processing time set CT2 of the second category of tasks be 15 seconds. Let Q1 = 200 and Q2 = 150, corresponding to the number of records or the amount of operations processed per second for the two categories of tasks. If the first category of tasks contains l = 2 operations and the data delay is CD 1,1 = 0.03 seconds, CD 1,2 = 0.05 seconds; the second category of tasks contains l = 2 operations and the data delay is CD 2,1 = 0.1 seconds, CD 2,2 = 0.08 seconds. Now, substitute the above values into the formula:
[0111] First step, first calculate seconds;
[0112] Then calculate
[0113] Second step, |CT1| = 10, so for the first category of tasks, we get
[0114] Third step, seconds;
[0115]
[0116] In the fourth step, if |CT2| = 15, then the second type of task gets
[0117] Finally, P opt = 0.05 + 0.1 = 0.15;
[0118] This result indicates that when the P opt value reaches 0.15, it means that both types of tasks are processed according to their corresponding computing densities and latencies within the given time. If the P opt value is larger, it means that more high-density computing or higher-latency operations can be accommodated simultaneously in the same pipeline. If the value is lower, it means that the efficiency of the overall task pipeline is low, and the operation strategy or latency control measures need to be readjusted.
[0119] Based on the optimized value of the data processing pipeline, compare the P opt obtained in the previous step with the corresponding task classification information, and allocate a dedicated preloading space to prepare data reading or computing in advance before each task call. During execution, scan the processing duration, computing density, latency, etc. recorded in the previous step item by item according to the queuing order of each type of task. If it is observed that some tasks frequently have a high computing density and an operation latency between 0.1 second and 0.2 seconds in a certain time period, a certain priority will be assigned to them and a prefetch action will be arranged for this priority, and the data that originally needed to be read dispersedly multiple times will be centrally pre-loaded into the cache area in advance. If another type of task only needs sporadic reading and writing but has a high concurrency in a short time, it will be merged with the existing prefetch queue to avoid congestion caused by repeated write operations or single-item reads. Finally, generate a complete execution order list during the data scheduling process and indicate the flow direction of each task, and use a combination of grid numbers and timestamps to manage this data transmission link, so as to connect all tasks in the same pipeline and determine the final execution order, and output it as a geological parameter processing flow chart.
[0120] The steps to obtain the geological computing node distribution table are as follows:
[0121] Based on the geological risk area distribution matrix and the geological parameter processing flow chart, extract the computing task flow information, analyze the computing requirements of each area, including the task volume, processing complexity, and time distribution, and generate a regional computing requirement analysis table;
[0122] According to the regional computing requirement analysis table, calculate the required number of computing nodes. The calculation formula is:
[0123]
[0124] where N i is the number of computing nodes required for the i-th monitoring area, and T i,tis the total task volume of the i-th region at time point t, R i,t is the task complexity parameter of the i-th region at time point t, H i,t is the data flow volume of the i-th region at time point t, A t and B t are the system load and data redundancy parameter at time point t respectively, D i,t is the data transmission delay of the i-th region at time point t, K is the total number of time points;
[0125] Based on the number of node requirements, computing nodes are allocated to form a geological computing node distribution table.
[0126] Specifically, based on the geological risk area distribution matrix and the geological parameter processing flow chart, first read the coordinate index and time distribution records of each monitoring area, and match these records with the previously determined geological parameter entries. When matching, it is necessary to distinguish items such as the settlement amount, porosity, and pipeline distribution of each area, and combine the daily or hourly observation frequency to identify the actual task volume of each area at different time periods. For example, the high-frequency updated settlement data volume in a certain area during the day can reach 30 to 50 pieces per hour, while at night it may only be updated once every 2 to 3 hours. In order to determine the processing complexity, the demand for data comparison, data accumulation, or differential operation in this area can be further subdivided. If an area needs to count multiple items such as settlement amount changes, pipeline loads, and soil layer thickness every day and each item involves batch operations, it can be considered that this area has a high processing complexity. If another area mostly involves simple pipeline reading and writing operations or lightweight numerical summation, its complexity is relatively low. Then, the above information is distinguished according to the 24-hour or weekday and weekend cycles, and it is marked which high-load areas hardly update parameters or update occasionally from early morning to early morning, and which medium-load areas still maintain an average level during holidays. Through this comparison and registration, the task volume of each area and the change range at different time points can be roughly grasped. If a large-scale data processing occurs in a certain area during the peak period, the processing duration and hardware resource occupancy are noted in the record and it is listed as the first priority. If there are some areas where the load suddenly increases occasionally, they are included in the elastic scheduling range. Finally, a unified index table is formed according to the total task entries, processing complexity, and time distribution of the monitoring areas, and the demand of all areas at different time points is summarized and the specific values are marked, and it is sorted into a regional computing demand analysis table.
[0127] The advantage of the formula is that by integrating multiple factors such as the task volume, task complexity, data flow volume, system load, data redundancy, and transmission delay of the region, it can three-dimensionally present the true demand situation of each region for computing nodes, and make the node allocation not only depend on the task volume, but also combine real-world impacts such as the load environment and network delay, so as to more accurately determine the required number of nodes.
[0128] T i,t Steps to obtain the parameter, which represents the total task volume of the i-th region at time point t. First, record whether there are activities such as batch operations of multiple settlement amounts, pipeline displacement comparison, or porosity matrix calculation in the region at the specified time point, and then count the number of records or data scale corresponding to each activity together, using the number of entries or the total amount of data blocks as the measurement unit. For example, if a region has batch operation activities of settlement amounts in 6 time periods within a day, and the number of entries generated by each activity is between 300 and 500, after recording, accumulate them and combine with the computational volume of soil layer thickness analysis. Finally, merge these values to obtain T. i,t If the number of entries is selected as the measurement standard, take the sum of the entries; if the data block size is selected, add them up in units of MB or GB.
[0129] R i,t Steps to obtain the parameter, which represents the task complexity of the i-th region at time point t. It is necessary to classify the arithmetic operations performed in this region during this period and assign a complexity score to each arithmetic operation. For example, ordinary calculation operations such as differential summation of settlement amounts are set to 1 to 3 points, and operations such as depth iteration or large matrix multiplication are set to 3 to 7 points. Then, add the scores according to the number of occurrences, and perform a comprehensive calculation on all operations at this time point to generate a total complexity value. If it is found in this region that pixel-by-pixel analysis of multi-dimensional data or high-dimensional interpolation calculation is required, the complexity score can also be adjusted upward to 8 to 10 points. Finally, count the scores of multiple types of operations and combine them with the number of occurrences to obtain the total task complexity at this time period.
[0130] H i,t Steps to obtain the parameter, which represents the data flow volume of the i-th region at time point t. Usually, MB or GB is used as the measurement unit. It is necessary to count the total amount of data uploaded and downloaded on the network and the data scale read and written on the local hard disk during the actual execution of the arithmetic task, and either weight or accumulate the two. For example, when batch processing settlement amounts, if 100MB of data needs to be read from the storage medium and 20MB of results need to be written after calculation, the data flow volume during this period can be temporarily recorded as 120MB. If there are other programs performing a large number of write operations on the porosity of the same region during the same period, then merge and accumulate them. Note that when adding different types of operations, ensure that the units are consistent and exclude some duplicate read and write records to obtain a more accurate H.i,t 。
[0131] A t and B t Steps for obtaining parameters of A t represents the system load at time point t. It is necessary to collect indicators such as CPU utilization, memory occupancy, or I / O busy ratio at the system level, linearly transform or normalize these indicators to obtain a value between 0 and 1. For example, if the CPU utilization is 60% and the memory occupancy is 70%, a weighted processing can be performed on the two to obtain a system load value of about 0.65; B t represents the data redundancy at this time point. If backups or multiple copies of the same data are maintained in the system, different copies can be accessed in parallel at the same time point to reduce the occupancy congestion of a single storage source. Here, the number of multiple copies can be converted into a redundancy score. For example, when there are two copies, it can be recorded as 1.5 points, and when there are three copies, it can be recorded as 2.0 points to represent different degrees of data availability. Then, these data are converted into values such as 0.5 or 1.0 and added to A t When adding, it is necessary to ensure the same dimension.
[0132] D i,t Steps for obtaining parameters of D. This parameter represents the data transmission delay of the i-th area at time point t, which is obtained based on the measured information of the network and hardware devices. If a complete data reading is performed on this area within the same time period and a response delay of 30 milliseconds to 60 milliseconds is generated, the average value of multiple measurements can be taken and extreme abnormal items can be excluded to obtain a relatively stable delay value such as 45 milliseconds. Then, it is converted into 0.045 seconds to unify the dimension. When the network quality or hardware bandwidth of this area or this time point is poor, this value may rise to 0.1 second or 0.15 second. It is also necessary to collect multiple times within a long period and summarize the results to obtain a final value.
[0133] Steps for obtaining parameter K. K is the total number of time points and needs to be clearly defined according to the actual monitoring period and observation frequency. If observations are made at the hourly level every day, 24 time points can be divided within a day. If it covers a week, K = 168. If it is extended to a month, K can reach 720 or even more, but it is necessary to keep the data collection rules and patterns relatively consistent for each time period. Each time point has corresponding T i,t 、R i,t 、H i,t 、A t 、B t and D i,t values, ensuring that both the numerator and denominator in the formula calculation have complete numerical sources.
[0134] Calculation process:
[0135] Taking the i-th region at K = 4 time points as an example, let T i,1 = 800 items, T i,2 = 600 items, T i,3 = 900 items, T i,4 = 750 items. These items can come from the total amount of settlement statistics and pipeline load comparison operations performed in the region at 4 different time periods. Let R i,1 = 5, R i,2 = 3, R i,3 = 7, R i,4 = 6, indicating that the calculation complexity score ranges for each time period are obtained by on-site engineering personnel; let H i,1 = 300MB, H i,2 = 220MB, H i,3 = 350MB, H i,4 = 280MB, indicating the data flow volume in the region during the corresponding time period; let A1 = 0.7, A2 = 0.65, A3 = 0.8, A4 = 0.75, which is the combined comprehensive normalization result of CPU and memory occupancy rates, and B1 = 1.0, B2 = 1.5, B3 = 1.0, B4 = 2.0, indicating the data redundancy for the corresponding time period; let D i,1 = 0.03 seconds, D i,2 = 0.04 seconds, D i,3 = 0.02 seconds, D i,4 = 0.05 seconds, which is obtained here by summarizing multiple measurements of network response and device bandwidth.
[0136] For ease of calculation, the H in i,t can be converted to a quantity dimension at the same numerical level. Here, it can be assumed that the MB data is scaled by a coefficient to a unit value. For example, 300MB corresponds to 300 units. Then:
[0137]
[0138] Calculate the results for other time points in sequence. In the denominator it is also necessary to first take the square root of A t +B t and then add D i,t . For example, at time point 1: A1 + B1 = 0.7 + 1.0 = 1.7, plus 0.03 = 1.33. Then calculate the numerator and denominator corresponding to time point 1:
[0139]
[0140] After performing the same calculation for the remaining 3 time points, sum them up and divide by K = 4, and finally take the square root of the result to obtain N i .
[0141] The result shows that the calculated N i reflects the number of computing nodes required for the i-th region considering factors such as task volume, task complexity, data flow volume, system load, data redundancy, and network latency at all time points. When this value is high, it means that more nodes need to be configured to maintain the normal operation of this region during high-load periods. If this value is low, it indicates that only a small number of nodes are required to meet the demand.
[0142] Based on the number of node requirements, the node demand values of each previous region at different time periods are extracted and registered sequentially in the background. For the registration process, it is necessary to clarify the coordinates and grid numbers corresponding to each region and match the previously recorded peak or off-peak task periods. If the node demand of a certain region reaches 8 in the morning and only 3 in the afternoon, these time intervals can be marked in the registration form and the reasons can be annotated. For example, the batch operation of settlement amount is concentrated in the morning or the comparison of porosity data is often carried out in the afternoon. Then, after the node demand information of all regions is completed, they are unified and summarized. If it is found that the demand of most regions is only 1 to 2 nodes at night, these nodes can be allocated to other regions with higher load. Through this allocation method, the number of nodes is listed in the wide-area or local area network according to the region number and time sequence, and it is indicated which time periods need to enable more parallel computing or prepare cache reading and writing. All allocation situations are combined into a geological computing node distribution table.
[0143] The above is only a preferred embodiment of the present invention, and it does not limit the present invention in other forms. Any person skilled in the art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes and apply them to other fields. However, as long as it does not depart from the technical solution content of the present invention, any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention still fall within the protection scope of the technical solution of the present invention.
Claims
1. An urban geological survey method based on big data cloud computing technology, characterized in that: The following steps are involved: Based on urban geological drilling detection data, ground subsidence monitoring data and underground pipe network distribution data, the data are grouped according to geographic grid units and time points, and statistical analysis is performed to obtain a statistical set of geological parameters; Based on the statistical set of geological parameters, multi-level storage partitions are set according to the data update frequency to establish a geological data partition storage table; Based on the geological data partition storage table, the areas with change rates exceeding the geological parameter change threshold are screened to obtain a set of key monitoring areas; The key monitoring area set is losslessly compressed and stored, and other area data is losslessly compressed and stored to generate a geological risk area distribution matrix; Based on the geological data partition storage table and the geological risk area distribution matrix, the geological parameter data stream is classified according to the calculation load to obtain a geological data calculation classification table; Merge similar calculations on the geological data calculation classification table, set up a data processing pipeline, establish a cache pre-fetch queue, and generate a geological parameter processing flow diagram; Based on the geological risk area distribution matrix and the geological parameter processing flow diagram, the geological parameters are allocated computing resources according to the processing flow to obtain a geological computing node distribution table.
2. The urban geological survey method based on big data cloud computing technology according to claim 1 is characterized in that: The steps for obtaining the statistical set of geological parameters are as follows: Obtain geological drilling detection data, ground subsidence monitoring data and underground pipe network distribution data, filter the information and classify it according to geographic grid units and time points to form a preliminary classification set of geological data; Performing statistical analysis on the geological parameters of each geographic grid unit in the preliminary classification set of geological data, calculating the mean and standard deviation of the geological parameters of each geographic grid unit, and obtaining a statistical result set of geological parameters; Based on the geological parameter statistical result set, the statistical data of all geographic grid cells are integrated to form a geological parameter statistical set through data aggregation.
3. The urban geological survey method based on big data cloud computing technology according to claim 1 is characterized in that: The steps for obtaining the geological data partition storage table are: Based on the statistical set of geological parameters, the number of data updates and distribution characteristics in each geographic grid unit are counted to generate data update frequency analysis results; According to the data update frequency analysis result, a multi-level storage partitioning scheme is designed, and a matching storage level is specified by determining the update frequency of each grid unit to generate a multi-level storage partitioning setting; Based on the multi-level storage partition setting, the locations and levels of different geographic grid unit data stored according to the update frequency are summarized to construct a geological data partition storage table.
4. The urban geological survey method based on big data cloud computing technology according to claim 1 is characterized in that: The steps for obtaining the key monitoring area set are: Based on the geological data partition storage table, extract the geological parameter historical data in each geographic grid unit to generate a geological parameter change data set; According to the geological parameter change data set, the geological parameter change rate in the geographic grid unit is calculated, and the calculation formula is: Among them, R i represents the rate of change of geological parameters of the i-th grid cell, ΔP i,t represents the difference in geological parameters of the ith grid cell at two consecutive time points, represents the standard deviation of the difference, P i,k is the geological parameter value of the i-th unit at the k-th time point, is the mean of the values, and N is the number of time points; Based on the geological parameter change rate, a geological parameter change threshold is set, and geographic grid units whose change rate exceeds the geological parameter change threshold are screened to form a set of key monitoring areas.
5. The urban geological survey method based on big data cloud computing technology according to claim 1 is characterized in that: The steps for obtaining the geological risk area distribution matrix are as follows: Based on the key monitoring area set, analyzing data redundancy items and removing irrelevant content, performing lossless compression on each geographic grid unit, and generating a key monitoring area compressed data set; Based on the compressed data set of the key monitoring area, lossy compression processing is performed on other geographic grid unit data, the processing results are merged into the compressed data set of the key monitoring area, and a geological risk area distribution matrix is generated according to geographical location and risk level.
6. The urban geological survey method based on big data cloud computing technology according to claim 1 is characterized in that: The steps for obtaining the geological data calculation classification table are as follows: Based on the geological data partition storage table and the geological risk area distribution matrix, the geological parameter data stream of each geographic grid unit is extracted, the calculation requirements of each parameter in the geological parameter data stream are analyzed, the processing characteristics of calculation-intensive and IO-intensive data are identified, and a geological parameter data stream set containing a calculation requirement classification is generated; According to the geological parameter data stream set, the data stream is classified and processed according to the computing requirements, and for computing-intensive data, vectorized operations are performed item by item to gradually optimize the computing efficiency. For IO-intensive data, read and write operations are integrated and executed in batches to generate a geological parameter data stream set after being processed by load classification; Based on the geological parameter data stream set, all classified data are sorted and standardized, and the calculation-intensive and IO-intensive data processing results are merged to form a geological data calculation classification table.
7. The urban geological survey method based on big data cloud computing technology according to claim 1 is characterized in that: The steps for obtaining the geological parameter processing flow diagram are as follows: Based on the geological data calculation classification table, analyze the characteristics of tasks in each calculation category, group the tasks by calculation type, and integrate the processing methods of similar calculations to generate a merged calculation classification task set; According to the combined computing classification task set, the optimization value of the data processing pipeline is calculated, and the calculation formula is: Among them, P opt is the optimized value of the data processing pipeline, CT j is the processing time set of the j-th task, Q j is the computational density of the j-th task, CD j,k is the k-th data processing delay of the j-th task, m is the number of task categories, and l is the number of operations included in each category of task; Based on the optimized value of the data processing pipeline, a cache prefetch queue is built to optimize the scheduling and distribution of data flows, and the execution order and flow direction of all tasks are organized into a geological parameter processing flow diagram.
8. The urban geological survey method based on big data cloud computing technology according to claim 1 is characterized in that: The steps for obtaining the geological calculation node distribution table are as follows: Based on the geological risk regional distribution matrix and the geological parameter processing flow diagram, the computing task flow information is extracted, the computing requirements of each region are analyzed, including the task volume, processing complexity and time distribution, and a regional computing requirement analysis table is generated; According to the regional calculation demand analysis table, the required number of nodes is calculated, and the calculation formula is: Among them, N i is the number of computing nodes required for the i-th monitoring area, T i,t is the total task volume of the ith region at time point t, R i,t is the task complexity parameter of the ith region at time point t, H i,t is the data flow of the ith region at time point t, A t and B t are the system load and data redundancy parameters at time point t, D i,t is the data transmission delay of the ith region at time point t, and K is the total number of time points; Based on the required number of nodes, computing nodes are allocated to form a geological computing node distribution table.
Citation Information
Cited By
Fine-grained per-vector scaling for neural network quantization
CN114118347A
Fine-grained per-vector scaling for neural network quantization
CN114118347B
Electroencephalogram data transmission method combining lossless and lossy compression
CN120751020A
Internet of Things platform data processing system
CN121330848A