Municipal road construction section data integrated management method and system based on distributed database
Through the data integration management method of municipal road construction sections based on distributed databases, traditional data warehouses are solved to solve the problem of scalability and real-time construction data, and efficient storage and processing of data are realized to meet the real-time and high-performance data requirements of construction services.
Patent Information
- Application Number
- CN202510556192.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-29
AI Technical Summary
Traditional data warehouse architectures are difficult to meet the scalability requirements of municipal road construction data, and there are shortcomings in data reading speed, processing time and adaptability to different network conditions, which cannot meet the construction business's requirements for real-time and high performance of data.
The data integration management method of municipal road construction sections based on distributed database is adopted, and the data cleaning, classification and priority storage strategies are used to calculate the frequency of data usage and increase frequency, evaluate the storage performance of the data warehouse, and prioritize storage and matching of data and data warehouses based on these indicators.
It improves the data reading speed and processing efficiency, meets the construction business's requirements for real-time and high performance data, optimizes the data storage layout, and reduces the cost of data integration.
Smart Images

Figure CN120067112A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data integration management, and specifically to a method and system for data integration management of municipal road construction sections based on a distributed database. Background Art
[0002] With the continuous expansion of the scale of municipal road construction, the amount of data generated during the construction process has increased explosively. Traditional data management methods are unable to cope when faced with massive, complex, and diverse-source construction data. Construction data includes various types such as construction equipment operation parameters, construction logs, quality inspection results, and on-site temporary change information, and comes from different channels such as sensor devices, mobile applications, and manual records.
[0003] According to the patent application with the publication number CN118885476B, a method for intelligent management of data integration services is disclosed. Before data integration, the present invention detects the consistency and accuracy of data from each data source and performs corresponding processing, thereby ensuring the data quality after data integration and being conducive to providing reliable data support for enterprise business decision-making; before data integration, it identifies sensitive data from each data source and performs data desensitization, thereby ensuring the security of data during the integration process; according to the data volume and data update frequency, it evaluates the priority of the data middle platform to allocate computing resources to transmit data from the data source, which is conducive to optimizing the configuration of the data integration process and resources and reducing the cost of data integration.
[0004] However, the amount of municipal road construction data is huge and the types are complex, and the traditional data warehouse architecture is difficult to meet the scalability requirements of data storage. At the same time, the classified management of different types of data and how to quickly and accurately obtain and process the required data are also one of the challenges. However, the existing technologies have deficiencies in data reading speed, processing duration, and adaptability to different network conditions, and cannot well meet the requirements of construction operations for data real-time performance and high performance. Summary of the Invention
[0005] In view of the deficiencies of the prior art, the present invention provides a method and system for data integration management of municipal road construction sections based on a distributed database, which solves the problems of deficiencies in data reading speed, processing duration, and adaptability to different network conditions, and being unable to well meet the requirements of construction operations for data real-time performance and high performance.
[0006] To achieve the above objectives, the present invention is realized through the following technical solutions: A method for data integration management of municipal road construction sections based on a distributed database, the method specifically includes the following steps: Collect construction data, and through data cleaning and data conversion, obtain preprocessed data; Classify the preprocessed data, calculate the sum of the usage frequency and the increase frequency of the same type of data within the time period to obtain the real-time value, calculate the sum of the maximum delay standard and the minimum processing speed to obtain the demand value; add the real-time value and the demand value to obtain the storage priority value, and generate data sorting information by sorting from largest to smallest. Construct a data warehouse, select the analysis object, respectively obtain its reading speeds when the network is stable and unstable, and calculate the average value of the two to obtain the reading value. Obtain the concurrent processing duration and the response duration, respectively match them with the judgment criteria to obtain the processing duration assignment and the response duration assignment, add the two to obtain the processing assignment, add the obtained value and the processing assignment to obtain the storage performance value, and sort from largest to smallest according to the storage performance value to generate data warehouse sorting information. Find the preprocessed data with the largest storage priority value from the data sorting information, that is, the priority storage data; determine the data warehouse with the largest storage performance value from the data warehouse sorting information, that is, the priority storage warehouse, and match and store the two to generate priority storage information. Obtain the associated data of the priority storage data, calculate the processing time based on the amount of associated data and the speed of processing the same type of data by each remaining data warehouse, and screen the remaining data warehouses with the minimum processing speed corresponding to the associated data to obtain the preselected storage warehouses. If the number of preselected storage warehouses is zero, store the associated data and the priority storage data together. If not, select the preselected storage warehouse with the shortest processing time to store the associated data, generate associated storage information, and combine the associated and priority storage information to obtain storage management information.
[0007] As a further solution of the present invention, the specific method for calculating the sum of the usage frequency and the increase frequency of the same type of data within the time period to obtain the real-time value is as follows: Classify the obtained preprocessed data to obtain the same type of data. Then, obtain the preprocessed data in the same type of data and label it as i, and i = 1, 2,..., j, where j represents the number of preprocessed data. Taking time T as the period, obtain the number of views of the preprocessed data i within the time period T, and calculate the usage frequency corresponding to the preprocessed data i. At the same time, obtain the increase amount corresponding to the preprocessed data i within the time period T, and calculate the increase frequency corresponding to the preprocessed data i according to the formula increase frequency = increase amount ÷ time T, and calculate the numerical sum of the usage frequency and the increase frequency, which is recorded as the real-time value of the preprocessed data i.
[0008] As a further solution of the present invention, the specific method for calculating the sum of the maximum delay standard and the minimum processing speed to obtain the demand value is as follows: Obtain the same type of data corresponding to the preprocessed data i, and obtain the corresponding maximum delay standard. At the same time, obtain the minimum processing speed corresponding to the preprocessed data i, and then calculate the numerical sum of the two to obtain the demand value of the preprocessed data i.
[0009] As a further solution of the present invention, the specific manner of generating data sorting information by sorting from large to small is as follows: Sum the real-time value and the required value of the preprocessed data i to obtain the storage priority value of the preprocessed data i, and sort them from large to small according to the storage priority value to generate data sorting information.
[0010] As a further solution of the present invention, the specific manner of calculating the read value of the average of the two is as follows: Obtain all data warehouses and label them as a, and a = 1, 2,..., b, where b represents the number of data warehouses. Obtain any type of data corresponding to the data warehouses, and record it as the analysis object. Obtain the read speed of the analysis object in different network conditions, and at the same time calculate the average value of the read speeds corresponding to different network conditions of the analysis object. Specifically, calculate the sum of the read speeds corresponding to stability and instability, and calculate the average value, which is recorded as the read value of the analysis object.
[0011] As a further solution of the present invention, the specific manner of generating data warehouse sorting information is as follows: Analyze the concurrent processing duration, obtain the processing durations corresponding to n real-time data processing tasks, and here n is specifically set by the operator. At the same time, match the obtained processing durations with the judgment criteria, and obtain the corresponding assignment, which is recorded as the processing duration assignment. Similarly, analyze the response duration to obtain the corresponding assignment, which is recorded as the response duration assignment, and calculate the sum of the numerical values of the processing duration assignment and the response duration assignment, which is recorded as the processing assignment. Then calculate the sum of the numerical values of the read value and the processing assignment, which is recorded as the storage performance value of the data warehouse; And sort them from large to small according to the obtained storage performance value to generate data warehouse sorting information.
[0012] As a further solution of the present invention, the specific manner of generating priority storage information is as follows: Based on the obtained data sorting information and data warehouse sorting information, conduct comprehensive analysis. Determine the associated data by analyzing the relevance of the data. Obtain the preprocessed data with the largest storage priority value in the data sorting information, which is recorded as the priority storage data. At the same time, obtain the data warehouse with the largest storage performance value in the data warehouse sorting information, which is recorded as the priority storage warehouse, and match and store the two to generate priority storage information.
[0013] As a further solution of the present invention, the specific manner of obtaining the preselected storage warehouse is as follows: Then obtain the associated data of the priority storage data, and at the same time conduct matching analysis on the remaining data warehouses based on the associated data; Obtain the processing speed corresponding to the remaining data warehouse, where the processing speed here is expressed as the processing speed of the same type as the associated data. At the same time, obtain the data volume corresponding to the associated data, calculate the processing time of the remaining data warehouse for the associated data, then obtain the minimum processing speed corresponding to the associated data, and screen the remaining data warehouse based on the minimum processing speed to obtain the preselected storage warehouse.
[0014] As a further solution of the present invention, the specific method for obtaining the storage management information is as follows: If the number of preselected storage warehouses is zero, store the associated data together with the priority storage data and generate associated storage information. If the number of preselected storage warehouses is not zero, take the preselected storage warehouse with the shortest processing time as the standard to generate associated storage information, and at the same time, obtain the storage management information by integrating the associated storage information and the priority storage information.
[0015] A data integration management system for municipal road construction sections based on a distributed database, including: A data acquisition module, which is used to acquire construction data, perform data cleaning on it to obtain preprocessed data, and transmit it to the storage analysis module at the same time; A data storage and analysis module, which is used to store and analyze the acquired preprocessed data. By classifying the preprocessed data, the same type of data is obtained. At the same time, calculate the usage frequency and increase frequency of the same type of data within a time period to obtain real-time values, and obtain demand values by calculating the maximum delay standard and the minimum processing speed. Then calculate the sum of the real-time value and the demand value to obtain the storage priority value of the preprocessed data, sort it from largest to smallest to generate data sorting information, and transmit it to the data management and analysis module at the same time; A data warehouse storage analysis module, which is used to establish a corresponding data warehouse, and at the same time analyze the reading speed and processing duration of different data corresponding to the data warehouse. Select one of the different data in the data warehouse as the analysis object, respectively obtain its reading speed when the network is stable and unstable, calculate the sum of the two reading speeds, and then calculate the average value to obtain the reading value of the analysis object. The processing duration analysis covers the concurrent processing duration and the response duration. Obtain the concurrent processing duration corresponding to the task, match it with the judgment standard numerical range obtained from a large amount of data analysis to obtain the processing duration assignment. Similarly, match the response duration to obtain the response duration assignment. Add the processing duration assignment and the response duration assignment to obtain the processing assignment. Add the reading value and the processing assignment to obtain the storage performance value of the data warehouse. Sort according to the storage performance value from largest to smallest to generate data warehouse sorting information, and transmit it to the data management and analysis module; A data management and analysis module, which is used to find out the pre - processed data with the largest storage priority value from the data sorting information, denoted as the priority storage data. At the same time, determine the data warehouse with the largest storage performance value in the data warehouse sorting information, called the priority storage warehouse, match and store the two, record the priority storage information, obtain the associated data of the priority storage data, based on the associated data, obtain the processing speed of each remaining data warehouse for processing the same type of data, combine with the amount of associated data, calculate the processing time, and screen the remaining data warehouses according to the minimum processing speed corresponding to the associated data to obtain the pre - selected storage warehouses; If the number of pre - selected storage warehouses is zero, store the associated data and the priority storage data together to generate associated storage information; if not zero, select the pre - selected storage warehouse with the shortest processing time to store the associated data, generate associated storage information, synthesize the associated and priority storage information to obtain storage management information, analyze all the pre - processed data according to this process, finally generate complete storage management information, and at the same time transmit the storage management information to the management information output module; A management information output module, which is used to manage the obtained storage management information.
[0016] The present invention provides a method and system for integrated management of municipal road construction section data based on a distributed database. Compared with the prior art, it has the following beneficial effects: Through comprehensive data cleaning operations, including removing duplicate data, handling missing values, correcting incorrect data, and performing data conversion operations such as format conversion, data type unification, and encoding conversion, the present invention ensures the accuracy, integrity, and availability of construction data, providing a reliable data basis for subsequent analysis and decision - making.
[0017] Comprehensively analyze the storage performance of data warehouses, including calculating the object reading speed and evaluating the processing duration under different network conditions. By reasonably screening data warehouses and storing data with high storage priority values in data warehouses with excellent storage performance values, the data reading speed and processing efficiency are effectively improved, meeting the strict requirements of construction operations for data real - time performance and high performance.
[0018] Accurately identify the associated data of the priority storage data, and according to the characteristics and processing speed requirements of the associated data, screen and match - analyze the remaining data warehouses. By reasonably arranging the storage location of the associated data, whether it is stored together with the priority storage data or stored in the pre - selected storage warehouse with the shortest processing time, the data storage layout is optimized, and the overall efficiency of data access and processing is improved, providing a more scientific and efficient solution for the integrated management of municipal road construction section data. Brief Description of the Drawings
[0019] Figure 1 It is a flowchart of the method steps of the present invention; Figure 2This is the system principle block diagram of the present invention. Specific Embodiments
[0020] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0021] Example 1. Please refer to Figure 1 , this application provides a method for integrated management of municipal road construction section data based on a distributed database. The method specifically includes the following steps: Step S1: Use a variety of data collection means, including but not limited to sensor devices, mobile applications, and manual records, etc., to comprehensively collect various types of construction data. For example, deploy sensors at the construction site to monitor the operating parameters of construction equipment in real time, such as the working hours and fuel consumption of excavators, as well as the mixing speed and temperature of concrete mixers, etc.; at the same time, construction personnel upload daily construction logs, quality inspection results and other information through mobile applications; for some data that cannot be obtained through automatic collection, such as temporary change information at the construction site, it is manually recorded and then entered into the system.
[0022] Remove duplicate data: During the data collection process, due to reasons such as network fluctuations and equipment failures, duplicate data records may occur. For example, the sensor sends the same construction equipment operating status data twice continuously. By using the deduplication function of the database and based on the unique identifier of the data (such as equipment number, timestamp, etc.), these duplicate data can be effectively identified and deleted to ensure the accuracy and uniqueness of the data.
[0023] Handle missing values: Missing values are inevitable in construction data. For example, due to sensor failures, the road surface temperature data at a certain moment is not recorded. For the missing values of numerical data, the mean filling method can be used, that is, calculate the average value of this type of data at other normal moments and fill the missing values with this average value; for non-numerical data, such as the missing name of the construction personnel in a certain construction log record, it can be supplemented by querying relevant attendance records or communicating with the on-site person in charge.
[0024] Correct incorrect data: Incorrect data may result from data entry errors or sensor failures, etc. For example, when entering construction progress data, the actual completed project quantity is wrongly written as an obviously unreasonable value. By setting reasonable data ranges and logical rules for verification, such as judging according to the construction plan and the completed construction stages that the project quantity in this stage should be within a certain range, if it exceeds this range, it is regarded as incorrect data and verified and corrected in time.
[0025] Format conversion: The data formats collected from different data sources may vary. For example, data obtained from construction equipment sensors provided by different suppliers may be stored in CSV format or JSON format. These data need to be uniformly converted into a format suitable for subsequent processing, such as converting CSV format data into a format that can be directly imported into a database, facilitating data storage and query.
[0026] Unifying data types: Ensure the consistency of data types to improve data processing efficiency. For example, for the quantity data of construction materials, some data sources may store it as a string type, while in subsequent data analysis, it needs to be converted into a numerical type for mathematical operations and statistical analysis.
[0027] Encoding conversion: If the data contains Chinese characters or other special characters, there may be encoding inconsistency problems. For example, the operation manuals of some construction equipment imported from abroad contain Chinese instructions, and their encoding may be different from the local system encoding. Encoding conversion is required to unify it into UTF-8 encoding to avoid garbled characters and ensure the readability and usability of the data.
[0028] Step S2: Store and manage the obtained preprocessed data. By establishing a corresponding data warehouse, classify the obtained preprocessed data to obtain data of the same type. Then, obtain the preprocessed data in the data of the same type and label it as i, where i = 1, 2,..., j, and j represents the quantity of preprocessed data. Then, analyze the real-time nature and requirements of the preprocessed data i in the data of the same type separately; Adopt a mature data warehouse architecture, such as the Hive data warehouse based on the Hadoop ecosystem, combined with the distributed file system HDFS to store a large amount of preprocessed data. Its advantage lies in strong scalability and can easily handle the continuous growth of construction data volume. For example, in a large road construction project, various construction data generated daily can reach several GB, and the Hive data warehouse can effectively store and manage these data.
[0029] Classify according to construction business logic and data characteristics. Taking road construction as an example, the preprocessed data can be divided into construction progress categories, quality inspection categories, equipment status categories, etc. For example, the construction progress category covers data such as the daily completed engineering volume and cumulative progress of each section; the quality inspection category includes test result data such as pavement compaction degree and concrete strength; the equipment status category includes data such as the operating duration and failure times of construction equipment.
[0030] Taking the time period T (e.g., T = 1 day) as a cycle, with the help of the log records of the database or a dedicated user behavior monitoring tool, accurately obtain the number of views n of the pre - processed data i (such as the real - time progress data of a certain construction section on the same day) within this time period T. The usage frequency f use The calculation formula of use is f , where T needs to be converted into a time unit consistent with the view count statistical period (e.g., when the unit is day, T = 1 day; if the view count is within hours, T is converted into hours). For example, within a certain day, the construction management staff viewed the real - time progress data of a certain section 50 times, then the usage frequency is 50 times / day; For the same - type data to which the pre - processed data i belongs (such as the daily progress data of all construction sections), within the time period T, through the incremental records of the database or the statistical function of the data acquisition system, obtain its increment , and the increment frequency f add The calculation formula of add is , for example, 3 new sections started construction on the same day, and the progress data generated by these 3 sections belongs to the increment of this type of data. If T = 1 day, then the increment frequency f add = 3 times / day. Add the usage frequency f use and the increment frequency f add to get the real - time value V of the pre - processed data i real-time , that is, V real-time = f use + f add . Taking the above example, the real - time value V real-time = 50 + 3 = 53.
[0031] According to the road construction industry standards, construction management specifications and actual business requirements, determine the maximum delay standard D max allowed during the reading process of the same - type data to which the pre - processed data i belongs (for example, for real - time construction equipment status data, it is stipulated that the maximum delay cannot exceed 5 seconds). Combining the performance indicators and business process requirements of the construction data processing system, determine the minimum processing speed S min of the pre - processed data i (such as 10 construction quality inspection data need to be processed per second).
[0032] Add the maximum delay standard D max and the minimum processing speed S min to get the required value V requirement of the pre - processed data i, that is, V requirement = D max + S min . Suppose the maximum delay standard of a certain type of construction data is 3 seconds and the minimum processing speed is 8 pieces / second, then the required value Vrequirement = 3 + 8 = 11. At the same time, sort them in descending order according to the storage requirement values to generate corresponding data sorting information.
[0033] Step S3: Label all the data warehouses involved in the analysis as a, where a = 1, 2,..., b, and b represents the total number of data warehouses. For example, in a data management system for a large road construction project, 5 data warehouses are deployed to store construction data of different stages and types. At this time, b = 5, and the data warehouses are labeled 1, 2, 3, 4, and 5 in sequence.
[0034] For each data warehouse, clarify the various types of data stored in it, including construction progress data, quality inspection data, equipment operation data, etc. For example, in the third data warehouse, select the construction progress data as the analysis object.
[0035] Set that within a unit time interval (such as 1 minute), if the network fluctuation difference is within the allowable range (such as ±5 Mbps), the network is stable. For example, within a certain 1 minute, the network download speed fluctuates from 100 Mbps to 103 Mbps, and the difference is 3 Mbps, which is within the allowable range, so it is determined that the network is stable.
[0036] When the network fluctuation difference within the unit time interval exceeds the allowable range, it means the network is unstable. For example, within another 1 minute, the network download speed drops suddenly from 90 Mbps to 70 Mbps, and the difference is 20 Mbps, which exceeds the allowable range of ±5 Mbps, so it is determined that the network is unstable.
[0037] Under different network conditions, obtain the reading speeds of the analysis object (construction progress data) multiple times. For example, under stable network conditions, conduct 10 reading speed tests, and the results are 102 Mbps, 105 Mbps, 101 Mbps, etc. Calculate the sum of these speeds as 1030 Mbps, and the average reading speed is 103 Mbps; under unstable network conditions, conduct 10 tests in the same way, and the results are 80 Mbps, 85 Mbps, 78 Mbps, etc., and the sum is 820 Mbps, and the average reading speed is 82 Mbps.
[0038] Calculate the average reading speed of the analysis object corresponding to different network conditions, that is, (103 Mbps + 82 Mbps) ÷ 2 = 92.5 Mbps. This value is recorded as the reading value of the analysis object. The operator sets n real-time data processing tasks. In this scenario, n = 3. For example, for the construction progress data, initiate 3 real-time processing tasks simultaneously, namely, count the progress of each section on the same day, calculate the cumulative progress of this week, and compare the actual progress with the planned progress.
[0039] Obtain the processing durations corresponding to these 3 tasks, assuming they are 2 seconds, 3 seconds, and 2.5 seconds respectively.
[0040] Based on a large amount of historical data analysis by the operator, determine the numerical range of the judgment criteria. For example, through analysis, it is obtained that for such tasks, a processing duration of 1 - 3 seconds is excellent and is assigned a value of 3; 3 - 5 seconds is good and is assigned a value of 2; 5 - 7 seconds is medium and is assigned a value of 1; more than 7 seconds is poor and is assigned a value of 0. The average processing duration of the above 3 tasks is (2 + 3 + 2.5) ÷ 3 = 2.5 seconds, which is in the 1 - 3 second range, and the processing duration is assigned a value of 3.
[0041] Similarly for the above 3 tasks, record their response durations, assuming they are 0.5 seconds, 0.6 seconds, and 0.4 seconds respectively, with an average of 0.5 seconds. Determine the response duration judgment criteria through historical data analysis. For example, 0 - 1 second is excellent and is assigned a value of 3. This average is within this range, and the response duration is assigned a value of 3.
[0042] Combining the assignment of the concurrent processing duration and the assignment of the response duration, the total processing duration assignment is obtained as 3 (assuming the weights of the two are the same and a simple average is taken); Calculate the sum of the read value and the processing assignment value to obtain the storage performance value of the data warehouse. In the above example, the read value is 92.5 Mbps, the processing duration assignment is 3, and the storage performance value = 92.5 + 3 = 95.5.
[0043] Step S4: From the data sorting information, find the pre - processed data with the largest storage priority value and mark it as the priority storage data. For example, in the road construction data, the pre - processed data regarding the progress of key construction nodes has the highest storage priority value among all pre - processed data due to its importance for construction decisions. At the same time, in the data warehouse sorting information, determine the data warehouse with the largest storage performance value and record it as the priority storage warehouse. Assume that among 5 data warehouses, the storage performance value of the second data warehouse is evaluated as the highest. Subsequently, store this priority storage data in the priority storage warehouse to form the priority storage information, ensuring that important data can be stored in the data warehouse with the best performance to guarantee efficient access and processing of the data.
[0044] For the priority storage data, obtain the associated data that has a subordinate relationship with it. Taking road construction as an example, if the priority storage data is the actual construction progress of a certain section on the same day, its associated data may include the planned construction progress of this section, the real - time inventory data of the construction materials used, etc., because when viewing the actual construction progress on the same day, the planned construction progress and the material inventory data often need to be referred to together, and there is a subordinate relationship between them. For the remaining data warehouses except the priority storage warehouse, obtain their processing speeds for processing data of the same type as the associated data. For example, if the associated data is construction material inventory data, the processing speeds of the remaining data warehouses for such data vary. At the same time, clarify the data volume of the associated data. Suppose the volume of associated construction material inventory data is 1000 records. According to the formula "processing time = data volume ÷ processing speed", calculate the time required for each remaining data warehouse to process the associated data. For example, if the processing speed of the first data warehouse for construction material inventory data is 200 records per second, then the time for it to process these 1000 associated data is 1000÷200 = 5 seconds.
[0045] Determine the minimum processing speed requirement corresponding to the associated data, which is usually determined by the construction business process and real-time requirements. For example, according to the construction management specification, the construction material inventory data must be processed within 3 seconds to ensure that the construction progress is not affected, that is, the minimum processing speed is approximately 1000÷3≈333 records per second. Taking this minimum processing speed as the standard, screen the remaining data warehouses, and regard the data warehouses with processing speeds reaching or exceeding this standard as preselected storage warehouses. If the processing speed of the third data warehouse for construction material inventory data is 400 records per second, which meets the minimum processing speed requirement, it becomes one of the preselected storage warehouses.
[0046] If after screening, the number of preselected storage warehouses is zero, this means that none of the remaining data warehouses can meet the minimum processing speed requirement of the associated data. At this time, store the associated data together with the priority storage data in the priority storage warehouse, and generate associated storage information to record this special storage arrangement for subsequent query and management.
[0047] When the number of preselected storage warehouses is not zero, select the data warehouse with the shortest processing time from these preselected storage warehouses. For example, there are 3 preselected storage warehouses, and the processing times for their associated data are 4 seconds, 3 seconds, and 5 seconds respectively. Select the data warehouse with a processing time of 3 seconds as the storage location for the associated data, and generate associated storage information.
[0048] Integrate the priority storage information and the associated storage information to form storage management information. This information comprehensively records the storage locations and arrangements of the priority storage data and its associated data, providing clear guidance for the storage management of road construction data. For example, the priority storage data (the actual construction progress of a certain road section on the current day) is stored in the second data warehouse, and its associated data (the construction plan progress, material inventory data, etc. of this road section) is stored in the preselected storage warehouse with the shortest processing time (such as data warehouse 4). These information are integrated to form complete storage management information.
[0049] Embodiment 2, please refer to Figure 2 This application provides a municipal road construction section data integration management system based on a distributed database, including: A data acquisition module, which is used to acquire construction data, clean the data to obtain preprocessed data, and transmit it to the storage and analysis module. The specific processing method is the same as the processing process of step S1 in Embodiment 1; A data storage and analysis module, which is used to store and analyze the acquired preprocessed data. By classifying the preprocessed data, the same type of data is obtained. At the same time, the usage frequency and increase frequency of the same type of data within a time period are calculated to obtain real-time values. The demand value is obtained by calculating the maximum delay standard and the minimum processing speed. Then, the sum of the real-time value and the demand value is calculated to obtain the storage priority value of the preprocessed data, and the data sorting information is generated by sorting from large to small. At the same time, it is transmitted to the data management and analysis module. The specific processing method is the same as the processing process of step S2 in Embodiment 1; A data warehouse storage and analysis module, which is used to establish a corresponding data warehouse. At the same time, analyze the reading speed and processing duration of different data corresponding to the data warehouse. Select one of the different data in the data warehouse as the analysis object, and respectively obtain its reading speed when the network is stable and unstable, calculate the sum of the two reading speeds, and then calculate the average value to obtain the reading value of the analysis object. The processing duration analysis covers the concurrent processing duration and the response duration. Obtain the concurrent processing duration corresponding to the task, match it with the judgment standard value interval obtained from a large amount of data analysis to obtain the processing duration assignment. Similarly, match the response duration to obtain the response duration assignment. Add the processing duration assignment and the response duration assignment to obtain the processing assignment. Add the reading value and the processing assignment to obtain the storage performance value of the data warehouse. Sort according to the storage performance value from large to small to generate data warehouse sorting information, and transmit it to the data management and analysis module. The specific processing method is the same as the processing process of step S3 in Embodiment 1; A data management and analysis module, which is used to find the preprocessed data with the largest storage priority value from the data sorting information, denoted as the priority storage data. At the same time, determine the data warehouse with the largest storage performance value in the data warehouse sorting information, called the priority storage warehouse, match and store the two, record the priority storage information, obtain the associated data of the priority storage data, and based on the associated data, obtain the processing speed of each remaining data warehouse for the same type of data. Combine the associated data volume to calculate the processing time. Screen the remaining data warehouses according to the minimum processing speed corresponding to the associated data to obtain the preselected storage warehouses. If the number of preselected storage warehouses is zero, store the associated data and the priority storage data together to generate associated storage information; if not, select the preselected storage warehouse with the shortest processing time to store the associated data to generate associated storage information. Combine the associated and priority storage information to obtain the storage management information, and analyze all the preprocessed data according to this process to finally generate the complete storage management information. The specific processing method is the same as the processing process of step S4 in Embodiment 1, and at the same time, transmit the storage management information to the management information output module; The management information output module is used to manage the obtained storage management information.
[0050] For some data in the above formula, only their numerical values are taken for calculation, and the parameter units are not substituted for calculation. At the same time, the content not described in detail in this specification belongs to the prior art well-known to those skilled in the art.
[0051] The above embodiments are only used to illustrate the technical method of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical method of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical method of the present invention.
Claims
1. A municipal road construction section data integration management method based on a distributed database, characterized in that: The method specifically comprises the following steps: Collect construction data, and obtain pre-processed data through data cleaning and data conversion; Classify the pre-processed data, calculate the real-time value of the sum of the usage frequency and the increase frequency of the same type of data within the time period, calculate the sum of the maximum delay standard and the minimum processing speed, and get the demand value; add the real-time value and the demand value to get the storage priority value, and generate data sorting information by sorting from large to small; Build a data warehouse, select the analysis object, obtain its reading speed when the network is stable and unstable, and calculate the reading value of both; Acquire concurrent processing duration and response duration, match them with the judgment criteria to obtain processing duration value and response duration value, add them together to obtain processing value, add the obtained value and processing value together to obtain storage performance value, sort them from large to small according to storage performance value, and generate data warehouse sorting information; Find the pre-processed data with the largest storage priority value from the data sorting information, that is, the priority storage data; determine the data bin with the largest storage performance value from the data bin sorting information, that is, the priority storage bin, match the two and store them to generate priority storage information; Obtain the associated data of the priority storage data, calculate the processing time according to the amount of associated data and the speed of each remaining data bin processing the same type of data, and select the remaining data bins according to the minimum processing speed corresponding to the associated data to obtain the pre-selected storage bin; If the number of pre-selected storage bins is zero, the associated data is stored together with the priority storage data. If it is not zero, the pre-selected storage bin with the shortest processing time is selected to store the associated data, generate associated storage information, and combine the associated and priority storage information to obtain storage management information.
2. The municipal road construction section data integration management method based on distributed database according to claim 1 is characterized in that: The specific method of calculating the real-time value of the sum of the usage frequency and the increase frequency of the same type of data within a time period is: The obtained preprocessed data are classified to obtain data of the same type, and then the preprocessed data in the data of the same type are obtained and labeled as i, where i=1, 2, ..., j, where j represents the number of preprocessed data; Taking time T as a period, obtain the number of views of preprocessed data i within time period T, and calculate the usage frequency corresponding to preprocessed data i. At the same time, obtain the increase corresponding to preprocessed data i within time period T, and calculate the increase frequency corresponding to preprocessed data i according to the formula increase frequency = increase amount ÷ time T. Calculate the sum of the numerical values of usage frequency and increase frequency and record it as the real-time value of preprocessed data i.
3. The municipal road construction section data integration management method based on distributed database according to claim 1 is characterized in that: The specific method of calculating the sum of the maximum delay standard and the minimum processing speed to obtain the required value is: The same type of data corresponding to the preprocessed data i is obtained, and the corresponding maximum delay standard is obtained, and the minimum processing speed corresponding to the preprocessed data i is obtained, and then the sum of the two values is calculated to obtain the required value of the preprocessed data i.
4. The municipal road construction section data integration management method based on distributed database according to claim 1 is characterized in that: The specific method of generating data sorting information by sorting from largest to smallest is: The real-time value and the demand value of the preprocessed data i are summed up to obtain the storage priority value of the preprocessed data i, and the preprocessed data are sorted from large to small according to the storage priority value to generate data sorting information.
5. The municipal road construction section data integration management method based on distributed database according to claim 1 is characterized in that: The specific method of calculating the reading value of both is as follows: Get all the data bins and label them as a, where a=1, 2, ..., b, where b represents the number of data bins. Get any type of data from the different data bins and record them as analysis objects. Get the reading speed of the analysis object under different network conditions. Meanwhile, calculate the average reading speed under different network conditions. Specifically, calculate the sum of the reading speeds corresponding to stability and instability, and calculate the average value, which is recorded as the reading value of the analysis object.
6. The municipal road construction section data integration management method based on distributed database according to claim 1 is characterized in that: The specific method of generating the data bin sorting information is: Analyze the concurrent processing time to obtain the processing time corresponding to n real-time data processing tasks, where n is specifically set by the operator. At the same time, match the obtained processing time with the judgment standard, and obtain the corresponding assignment as the processing time assignment. Similarly, analyze the response time to obtain the corresponding assignment as the response time assignment, and calculate the sum of the processing time assignment and the response time assignment as the processing assignment. Then calculate the sum of the reading value and the processing assignment as the storage performance value of the data warehouse. And sort the obtained storage performance values from large to small to generate data warehouse sorting information.
7. The municipal road construction section data integration management method based on distributed database according to claim 1 is characterized in that: The specific method of generating the priority storage information is: A comprehensive analysis is performed based on the obtained data sorting information and data warehouse sorting information. The associated data is determined by analyzing the correlation of the data. The preprocessed data with the largest storage priority value in the data sorting information is obtained and recorded as the priority storage data. At the same time, the data warehouse with the largest storage performance value in the data warehouse sorting information is obtained and recorded as the priority storage warehouse. The two are matched and stored to generate priority storage information.
8. The municipal road construction section data integration management method based on distributed database according to claim 1 is characterized in that: The specific method of pre-selecting the storage bin is as follows: Then, the associated data of the data stored in priority is obtained, and the remaining data bins are matched and analyzed based on the associated data; Obtain the processing speed corresponding to the remaining data bins, and the processing speed here is expressed as the processing speed of the same type as the associated data. At the same time, obtain the data volume corresponding to the associated data, and calculate the processing time of the remaining data bins for the associated data. Then, obtain the minimum processing speed corresponding to the associated data, and use the minimum processing speed as the standard to filter the remaining data bins to obtain the pre-selected storage bins.
9. The municipal road construction section data integration management method based on distributed database according to claim 1 is characterized in that: The specific method of obtaining the storage management information is: If the number of pre-selected storage bins is zero, the associated data is stored together with the priority storage data, and associated storage information is generated. If the number of pre-selected storage bins is not zero, the pre-selected storage bin with the shortest processing time is used as the standard to generate associated storage information, and the storage management information is obtained by combining the associated storage information and the priority storage information.
10. A municipal road construction section data integrated management system based on a distributed database, used to execute the municipal road construction section data integrated management method based on a distributed database as claimed in any one of claims 1 to 9, characterized in that: include: A data acquisition module is used to acquire construction data, clean the data to obtain pre-processed data, and transmit the data to the storage and analysis module; The data storage and analysis module is used to store and analyze the acquired pre-processed data, classify the pre-processed data to obtain data of the same type, calculate the usage frequency and increase frequency of the same type of data within a time period to obtain the real-time value, and calculate the maximum delay standard and the minimum processing speed to obtain the demand value, then calculate the sum of the real-time value and the demand value to obtain the storage priority value of the pre-processed data, and generate data sorting information by sorting from large to small, and transmit it to the data management and analysis module; Data warehouse storage analysis module, which is used to establish the corresponding data warehouse, and analyze the reading speed and processing time of different data corresponding to the data warehouse at the same time, select one of the different data in the data warehouse as the analysis object, obtain its reading speed when the network is stable and unstable, calculate the sum of the two reading speeds, and then calculate the average to obtain the reading value of the analysis object. The processing time analysis covers concurrent processing time and response time. The concurrent processing time corresponding to the task is obtained, and it is matched with the judgment standard numerical range obtained based on a large amount of data analysis to obtain the processing time assignment. Similarly, the response time is matched to obtain the response time assignment, and the processing time assignment is added to the response time assignment to obtain the processing assignment. The reading value is added to the processing assignment to obtain the storage performance value of the data warehouse, and the storage performance value is sorted from large to small to generate the data warehouse sorting information, and it is transmitted to the data management analysis module; Data management and analysis module, which is used to find the pre-processed data with the largest storage priority value from the data sorting information, record it as the priority storage data, and at the same time determine the data bin with the largest storage performance value in the data bin sorting information, called the priority storage bin, match and store the two, record the priority storage information, obtain the associated data of the priority storage data, obtain the speed of each remaining data bin processing the same type of data based on the associated data, calculate the processing time in combination with the amount of associated data, and screen the remaining data bins according to the minimum processing speed corresponding to the associated data to obtain the pre-selected storage bin; If the number of pre-selected storage bins is zero, the associated data and the priority storage data are stored in one place to generate associated storage information; If it is not zero, the pre-selected storage bin with the shortest processing time is selected to store the associated data, generate associated storage information, and integrate the associated and priority storage information to obtain storage management information. All pre-processed data are analyzed according to this process, and finally complete storage management information is generated, and the storage management information is transmitted to the management information output module; A management information output module is used to manage the obtained storage management information.
Citation Information
Patent Citations
A data integration service intelligent management method
CN118885476B
Industrial operation system data lake construction method based on data warehouse
CN114490886A
Decision analysis-based data center system
CN115600849A
Method and System for Constructing Data Warehouse Based on Wireless Communication Network, and Device and Medium
US20240273116A1