Method and System for Data Integration Management of Municipal Road Construction Sections Based on Distributed Database
Through the data integration management method of distributed database, the high performance and real-time problems of municipal road construction data are solved. Through data cleaning and storage optimization, data reading speed and processing efficiency are improved, and construction business needs are met.
Patent Information
- Application Number
- CN202510556192.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-04-29
AI Technical Summary
Traditional data management methods are difficult to meet the high performance and real-time requirements of municipal road construction data, especially in terms of data reading speed, processing time, and adaptability to different network conditions.
Adopt a data integration management method based on distributed databases, through data cleaning, classification, storage priority value calculation and data warehouse performance analysis, reasonably filter the data warehouse storage location, optimize the data storage layout, and improve data access and processing efficiency.
It realizes the accuracy and completeness of construction data, improves data reading speed and processing efficiency, and meets the requirements of construction business for real-time and high performance.
Smart Images

Figure CN120067112B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data integration management, and specifically to a method and system for data integration management of municipal road construction sections based on a distributed database. Background Art
[0002] With the continuous expansion of the scale of municipal road construction, the amount of data generated during the construction process has increased explosively. Traditional data management methods are unable to cope when faced with massive, complex, and diverse sources of construction data. Construction data includes various types such as construction equipment operation parameters, construction logs, quality inspection results, and on-site temporary change information, and comes from different channels such as sensor devices, mobile applications, and manual records.
[0003] According to the patent application with the publication number CN118885476B, a method for intelligent management of data integration services is disclosed. Before data integration, the present invention detects the consistency and accuracy of data from each data source and performs corresponding processing, thereby ensuring the data quality after data integration, which is conducive to providing reliable data support for enterprise business decisions; before data integration, it identifies sensitive data from each data source and performs data desensitization, thereby ensuring the security of data during the integration process; according to the data volume and data update frequency, it evaluates the priority of the data middle platform to allocate computing resources to transmit data from the data source, which is conducive to optimizing the configuration of the data integration process and resources and reducing the cost of data integration.
[0004] However, the amount of municipal road construction data is huge and the types are complex, and the traditional data warehouse architecture is difficult to meet the scalability requirements of data storage. At the same time, the classified management of different types of data and how to quickly and accurately obtain and process the required data are also one of the challenges. However, the existing technologies have deficiencies in data reading speed, processing duration, and adaptability to different network conditions, and cannot well meet the requirements of construction operations for data real-time performance and high performance. Summary of the Invention
[0005] In view of the deficiencies of the prior art, the present invention provides a method and system for data integration management of municipal road construction sections based on a distributed database, which solves the problems of deficiencies in data reading speed, processing duration, and adaptability to different network conditions, and cannot well meet the requirements of construction operations for data real-time performance and high performance.
[0006] To achieve the above objectives, the present invention is realized through the following technical solutions: A method for data integration management of municipal road construction sections based on a distributed database, the method specifically includes the following steps:
[0007] Collect construction data, and through data cleaning and data conversion, obtain preprocessed data;
[0008] Classify the preprocessed data, calculate the sum of the usage frequency and the increase frequency of the same type of data within the time period to obtain the real-time value, calculate the sum of the maximum delay standard and the minimum processing speed to obtain the demand value; add the real-time value and the demand value to obtain the storage priority value, and sort them from largest to smallest to generate data sorting information;
[0009] Construct a data warehouse, select the analysis object, respectively obtain its reading speeds when the network is stable and unstable, and calculate the average value of the two to obtain the reading value;
[0010] Obtain the concurrent processing duration and the response duration, respectively match them with the judgment criteria to obtain the processing duration assignment and the response duration assignment, add the two to obtain the processing assignment, add the value and the processing assignment to obtain the storage performance value, and sort them from largest to smallest according to the storage performance value to generate data warehouse sorting information;
[0011] Find the preprocessed data with the largest storage priority value from the data sorting information, that is, the priority storage data; determine the data warehouse with the largest storage performance value from the data warehouse sorting information, that is, the priority storage warehouse, and match and store the two to generate priority storage information;
[0012] Obtain the associated data of the priority storage data, and calculate the processing time based on the amount of associated data and the speed of the remaining data warehouses in processing the same type of data. Screen the remaining data warehouses with the minimum processing speed corresponding to the associated data to obtain the preselected storage warehouses;
[0013] If the number of preselected storage warehouses is zero, store the associated data and the priority storage data together. If not, select the preselected storage warehouse with the shortest processing time to store the associated data, generate associated storage information, and combine the associated and priority storage information to obtain storage management information.
[0014] As a further solution of the present invention, the specific method for calculating the sum of the usage frequency and the increase frequency of the same type of data within the time period to obtain the real-time value is as follows:
[0015] Classify the obtained preprocessed data to obtain the same type of data. Then, obtain the preprocessed data in the same type of data and label it as i, and i = 1, 2,..., j, where j represents the number of preprocessed data;
[0016] Taking time T as the period, obtain the number of views of the preprocessed data i within the time period T, calculate the usage frequency corresponding to the preprocessed data i, and at the same time obtain the increase amount corresponding to the preprocessed data i within the time period T. According to the formula increase frequency = increase amount ÷ time T, calculate the increase frequency corresponding to the preprocessed data i, and calculate the sum of the values of the usage frequency and the increase frequency, which is denoted as the real-time value of the preprocessed data i.
[0017] As a further solution of the present invention, the specific method for calculating the demand value by adding the maximum delay standard and the minimum processing speed is as follows:
[0018] Obtain the same type of data corresponding to the preprocessed data i, obtain the corresponding maximum delay standard, and at the same time obtain the minimum processing speed corresponding to the preprocessed data i, and then calculate the sum of the two values to obtain the demand value of the preprocessed data i.
[0019] As a further solution of the present invention, the specific method for generating data sorting information in descending order is as follows:
[0020] Sum the real-time value and the demand value of the preprocessed data i to obtain the storage priority value of the preprocessed data i, and sort them in descending order according to the storage priority value to generate data sorting information.
[0021] As a further solution of the present invention, the specific method for calculating the average value of the two to obtain the reading value is as follows:
[0022] Obtain all data warehouses and label them as a, and a = 1, 2,..., b, where b represents the number of data warehouses. Obtain any type of data corresponding to the data warehouse, and record it as the analysis object. Obtain the reading speed of the analysis object in different network conditions, and at the same time calculate the average value of the reading speeds corresponding to different network conditions of the analysis object. Specifically, calculate the sum of the reading speeds corresponding to stability and instability, and calculate the average value, which is recorded as the reading value of the analysis object.
[0023] As a further solution of the present invention, the specific method for generating data warehouse sorting information is as follows:
[0024] Analyze the concurrent processing duration, obtain the processing durations corresponding to n real-time data processing tasks, and here n is specifically set by the operator. At the same time, match the obtained processing durations with the judgment standard, and obtain the corresponding assignment, which is recorded as the processing duration assignment. Similarly, analyze the response duration to obtain the corresponding assignment, which is recorded as the response duration assignment, and calculate the sum of the values of the processing duration assignment and the response duration assignment, which is recorded as the processing assignment. Then calculate the sum of the reading value and the processing assignment, which is recorded as the storage performance value of the data warehouse;
[0025] And sort them in descending order according to the obtained storage performance value to generate data warehouse sorting information.
[0026] As a further solution of the present invention, the specific method for generating priority storage information is as follows:
[0027] Based on the obtained data sorting information and the data warehouse sorting information, comprehensive analysis is carried out. By analyzing the relevance of the data, associated data is determined. The preprocessed data with the largest storage priority value in the data sorting information is recorded as the priority storage data. At the same time, the data warehouse with the largest storage performance value in the data warehouse sorting information is recorded as the priority storage warehouse, and the two are matched and stored to generate priority storage information.
[0028] As a further solution of the present invention, the specific method for obtaining the preselected storage warehouse is as follows:
[0029] Then, the associated data of the priority storage data is obtained, and at the same time, the remaining data warehouses are matched and analyzed based on the associated data;
[0030] The processing speed corresponding to the remaining data warehouses is obtained, and the processing speed here is expressed as the processing speed of the same type as the associated data. At the same time, the data volume corresponding to the associated data is obtained, and the processing time of the remaining data warehouses for the associated data is calculated. Then, the minimum processing speed corresponding to the associated data is obtained, and the remaining data warehouses are screened based on the minimum processing speed to obtain the preselected storage warehouse.
[0031] As a further solution of the present invention, the specific method for obtaining the storage management information is as follows:
[0032] If the number of preselected storage warehouses is zero, the associated data and the priority storage data are stored together to generate associated storage information. If the number of preselected storage warehouses is not zero, the preselected storage warehouse with the shortest processing time is used as the standard to generate associated storage information. At the same time, the storage management information is obtained by integrating the associated storage information and the priority storage information.
[0033] A data integration management system for municipal road construction sections based on a distributed database includes:
[0034] A data acquisition module, which is used to acquire construction data, perform data cleaning on it to obtain preprocessed data, and transmit it to the storage analysis module at the same time;
[0035] A data storage and analysis module, which is used to store and analyze the obtained preprocessed data. By classifying the preprocessed data, the same type of data is obtained. At the same time, the usage frequency and increase frequency of the same type of data within a time period are calculated to obtain real-time values, and the demand values are obtained by calculating the maximum delay standard and the minimum processing speed. Then, the sum of the real-time value and the demand value is calculated to obtain the storage priority value of the preprocessed data, and it is sorted from largest to smallest to generate data sorting information, and it is transmitted to the data management and analysis module at the same time;
[0036] Data warehouse storage analysis module, which is used to establish a corresponding data warehouse, and at the same time analyze the reading speed and processing duration of different data in the data warehouse. Select one of the different data in the data warehouse as the analysis object, and respectively obtain its reading speed when the network is stable and unstable, calculate the sum of the two reading speeds, and then calculate the average value to obtain the reading value of the analysis object. The processing duration analysis covers the concurrent processing duration and the response duration. Obtain the concurrent processing duration corresponding to the task, match it with the judgment standard value interval obtained from a large amount of data analysis to obtain the processing duration assignment. Similarly, match the response duration to obtain the response duration assignment. Add the processing duration assignment and the response duration assignment to obtain the processing assignment. Add the reading value and the processing assignment to obtain the storage performance value of the data warehouse. Sort according to the storage performance value from large to small to generate the data warehouse sorting information, and transmit it to the data management analysis module;
[0037] Data management analysis module, which is used to find the preprocessed data with the largest storage priority value from the data sorting information, denoted as the priority storage data. At the same time, determine the data warehouse with the largest storage performance value in the data warehouse sorting information, called the priority storage warehouse, match and store the two, record the priority storage information, obtain the associated data of the priority storage data, and based on the associated data, obtain the processing speed of each remaining data warehouse for the same type of data. Combine the associated data volume to calculate the processing time. Screen the remaining data warehouses according to the minimum processing speed corresponding to the associated data to obtain the preselected storage warehouses;
[0038] If the number of preselected storage warehouses is zero, store the associated data and the priority storage data together to generate associated storage information; if not, select the preselected storage warehouse with the shortest processing time to store the associated data to generate associated storage information. Combine the associated and priority storage information to obtain the storage management information. Analyze all the preprocessed data according to this process, finally generate the complete storage management information, and at the same time transmit the storage management information to the management information output module;
[0039] Management information output module, which is used to manage the obtained storage management information.
[0040] The present invention provides a method and system for integrated management of municipal road construction section data based on a distributed database. Compared with the prior art, it has the following beneficial effects:
[0041] Through comprehensive data cleaning operations, including removing duplicate data, handling missing values and correcting incorrect data, as well as data conversion operations such as format conversion, data type unification and encoding conversion, the present invention ensures the accuracy, integrity and availability of the construction data, providing a reliable data basis for subsequent analysis and decision-making.
[0042] Conduct a comprehensive analysis of the storage performance of the data warehouse, including calculating the object reading speed and evaluating the processing duration under different network conditions. By reasonably screening the data warehouse, storing data with high storage priority values in the data warehouse with excellent storage performance values, the data reading speed and processing efficiency are effectively improved, meeting the strict requirements of the construction business for data real-time and high performance.
[0043] Accurately identify the associated data of the preferentially stored data, and perform screening and matching analysis on the remaining data warehouses according to the characteristics and processing speed requirements of the associated data. By reasonably arranging the storage locations of the associated data, whether it is stored together with the preferentially stored data or stored in the preselected storage warehouse with the shortest processing time, the data storage layout is optimized, and the overall efficiency of data access and processing is improved, providing a more scientific and efficient solution for the integrated management of municipal road construction section data. Brief Description of the Drawings
[0044] Figure 1 It is a flowchart of the method steps of the present invention;
[0045] Figure 2 It is a block diagram of the system principle of the present invention. Detailed Embodiments
[0046] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0047] Embodiment 1, please refer to Figure 1 , the present application provides a method for integrated management of municipal road construction section data based on a distributed database. The method specifically includes the following steps:
[0048] Step S1: Use a variety of data collection means, including but not limited to sensor devices, mobile applications, and manual records, etc., to comprehensively collect various types of construction data. For example, deploy sensors at the construction site to monitor the operating parameters of construction equipment in real time, such as the working hours and fuel consumption of excavators, as well as the mixing speed and temperature of concrete mixers; at the same time, construction personnel upload information such as daily construction logs and quality inspection results through mobile applications; for some data that cannot be obtained through automatic collection, such as temporary change information at the construction site, it is manually recorded and then entered into the system.
[0049] Removing duplicate data: During the data collection process, due to reasons such as network fluctuations and equipment failures, duplicate data records may occur. For example, the sensor sends the same construction equipment operation status data twice in a row. By using the deduplication function of the database and based on the unique identifier of the data (such as equipment number, timestamp, etc.), these duplicate data can be effectively identified and deleted to ensure the accuracy and uniqueness of the data.
[0050] Handling missing values: Missing values are inevitable in construction data. For example, due to sensor failures, the road surface temperature data at a certain moment is not recorded. For the missing values of numerical data, the mean filling method can be used, that is, calculate the average value of this type of data at other normal times and fill the missing values with this average value; for non-numerical data, such as the missing construction worker name in a construction log, it can be supplemented by querying relevant attendance records or communicating with the on-site person in charge.
[0051] Correcting incorrect data: Incorrect data may result from data entry errors or sensor failures, etc. For example, when entering construction progress data, the actual completed project volume is miswritten as an obviously unreasonable value. By setting reasonable data ranges and logical rules for verification, such as judging that the project volume at this stage should be within a certain range according to the construction plan and the completed construction stage, if it exceeds this range, it is regarded as incorrect data and verified and corrected in a timely manner.
[0052] Format conversion: The data formats collected from different data sources may vary. For example, the data obtained from construction equipment sensors provided by different suppliers is stored in different formats, some in CSV format and some in JSON format. These data need to be uniformly converted into a format suitable for subsequent processing, such as converting CSV format data into a format that can be directly imported into the database for easy data storage and query.
[0053] Unifying data types: Ensure the consistency of data types to improve data processing efficiency. For example, for the quantity data of construction materials, some data sources may store it as a string type, while in subsequent data analysis, it needs to be converted into a numerical type for mathematical operations and statistical analysis.
[0054] Encoding conversion: If the data contains Chinese characters or other special characters, there may be encoding inconsistency problems. For example, the operation manuals of some construction equipment imported from abroad contain Chinese instructions, and their encoding may be different from the local system encoding. Encoding conversion is required to unify it into UTF-8 encoding to avoid garbled characters and ensure the readability and usability of the data.
[0055] Step S2: Store and manage the obtained preprocessed data. By establishing a corresponding data warehouse, classify the obtained preprocessed data to get data of the same type. Then, obtain the preprocessed data in the data of the same type and label it as i, where i = 1, 2, …, j, and j represents the quantity of the preprocessed data. Next, analyze the real-time nature and requirements of the preprocessed data i in the data of the same type separately;
[0056] Adopt a mature data warehouse architecture, such as the Hive data warehouse based on the Hadoop ecosystem, and combine it with the distributed file system HDFS to store a large amount of preprocessed data. Its advantage lies in strong scalability and can easily cope with the continuous growth of construction data volume. For example, in a large road construction project, various construction data generated daily can reach several GB, and the Hive data warehouse can effectively store and manage these data.
[0057] Classify according to the construction business logic and data characteristics. Taking road construction as an example, the preprocessed data can be divided into construction progress type, quality inspection type, equipment status type, etc. For example, the construction progress type covers data such as the daily completed engineering volume and cumulative progress of each road section; the quality inspection type includes test result data such as pavement compaction degree and concrete strength; the equipment status type includes data such as the operating duration and failure times of construction equipment.
[0058] Taking time T (such as T = 1 day) as a cycle, with the help of the database's log records or specialized user behavior monitoring tools, accurately obtain the number of views n of the preprocessed data i (such as the real-time progress data of a certain construction section on the same day) within this time cycle T. The usage frequency f use is calculated by the formula f use = , where T needs to be converted into a time unit consistent with the view count statistical cycle (such as when the unit is day, T = 1 day; if the view count is statistically within hours, T is converted into hours). For example, within a certain day, the construction management personnel viewed the real-time progress data of a certain road section 50 times, then the usage frequency is 50 times / day;
[0059] For the data of the same type to which the preprocessed data i belongs (such as the same-day progress data of all construction sections), within the time cycle T, obtain its increase amount , and the increase frequency f add is calculated by the formula f add = , for example, 3 new road sections started construction on the same day, and the progress data generated by these 3 road sections belongs to the increase amount of this type of data. If T = 1 day, then the increase frequency f add = 3 times / day. Compare the usage frequency f use and the increase frequency f addAdd them up to obtain the real-time value V of the preprocessed data i real-time , that is, V real-time = f use + f add . Taking the above example, the real-time value V real-time = 50 + 3 = 53
[0060] According to the road construction industry standards, construction management specifications, and actual business requirements, determine the maximum delay standard D allowed for the same type of data to which the preprocessed data i belongs during the reading process max (For example, for real-time construction equipment status data, it is stipulated that the maximum delay cannot exceed 5 seconds). Combining the performance indicators and business process requirements of the construction data processing system, determine the minimum processing speed S corresponding to the preprocessed data i min (such as 10 construction quality inspection data need to be processed per second).
[0061] Add the maximum delay standard D max and the minimum processing speed S min to obtain the required value V of the preprocessed data i requirement , that is, V requirement = D max + S min . Suppose the maximum delay standard for a certain type of construction data is 3 seconds and the minimum processing speed is 8 pieces / second, then the required value V requirement = 3 + 8 = 11. At the same time, sort according to the storage required values from large to small to generate the corresponding data sorting information
[0062] Step S3: Label all the data warehouses participating in the analysis, marked as a, where a = 1, 2,..., b, and b represents the total number of data warehouses. For example, in the data management system of a large road construction project, 5 data warehouses are deployed to store construction data of different stages and types. At this time, b = 5, and the data warehouses are numbered 1, 2, 3, 4, 5 in sequence
[0063] For each data warehouse, clarify the various types of data it stores, including construction progress data, quality inspection data, equipment operation data, etc. For example, in the third data warehouse, select the construction progress data as the analysis object
[0064] Set that within the unit time interval (such as 1 minute), the network fluctuation difference within the allowable range (such as ±5 Mbps) is considered network stability
[0065] For example, within a certain 1 minute, the network download speed fluctuates from 100 Mbps to 103 Mbps, and the difference is 3 Mbps, which is within the allowable range, and it is determined to be network stable
[0066] When the network fluctuation difference exceeds the allowable range within a unit time interval, the network is considered unstable. For example, within another minute, the network download speed drops suddenly from 90 Mbps to 70 Mbps, with a difference of 20 Mbps, exceeding the allowable range of ±5 Mbps, and it is determined that the network is unstable.
[0067] Under different network conditions, the reading speeds of the analysis object (construction progress data) are obtained multiple times. For example, under stable network conditions, 10 reading speed tests are conducted, and the results are 102 Mbps, 105 Mbps, 101 Mbps, etc. The sum of these speeds is 1030 Mbps, and the average reading speed is 103 Mbps; under unstable network conditions, 10 tests are also conducted, and the results are 80 Mbps, 85 Mbps, 78 Mbps, etc., with a sum of 820 Mbps and an average reading speed of 82 Mbps.
[0068] Calculate the average reading speed corresponding to different network conditions of the analysis object, that is, (103 Mbps + 82 Mbps) ÷ 2 = 92.5 Mbps, and this value is recorded as the reading value of the analysis object;
[0069] The operator sets n real-time data processing tasks. In this scenario, n = 3. For example, for construction progress data, 3 real-time processing tasks are initiated simultaneously, namely, counting the progress of each section on the same day, calculating the cumulative progress of this week, and comparing the actual progress with the planned progress.
[0070] Obtain the processing durations corresponding to these 3 tasks, assuming they are 2 seconds, 3 seconds, and 2.5 seconds respectively.
[0071] Based on a large amount of historical data analysis by the operator, the numerical interval of the judgment standard is determined. For example, through analysis, it is obtained that for such tasks, the processing duration is excellent in the range of 1 - 3 seconds, assigned a value of 3; good in the range of 3 - 5 seconds, assigned a value of 2; medium in the range of 5 - 7 seconds, assigned a value of 1; poor above 7 seconds, assigned a value of 0. The average processing duration of the above 3 tasks is (2 + 3 + 2.5) ÷ 3 = 2.5 seconds, which is in the 1 - 3 second interval, and the processing duration is assigned a value of 3.
[0072] Similarly for the above 3 tasks, record their response durations, assuming they are 0.5 seconds, 0.6 seconds, and 0.4 seconds respectively, with an average value of 0.5 seconds. Through historical data analysis, the response duration judgment standard is determined. For example, 0 - 1 second is excellent, assigned a value of 3, and this average value is in this interval, so the response duration is assigned a value of 3.
[0073] Combining the assignment of the concurrent processing duration and the assignment of the response duration, the total processing duration assignment is obtained as 3 (assuming the weights of the two are the same and a simple average is taken);
[0074] Calculate the sum of the read value and the processed assignment value to obtain the storage performance value of the data warehouse. In the above example, the read value is 92.5 Mbps, the processed duration assignment is 3, and the storage performance value = 92.5 + 3 = 95.5.
[0075] Step S4: From the data sorting information, find the preprocessed data with the largest storage priority value and mark it as the priority storage data. For example, in the road construction data, the preprocessed data regarding the progress of key construction nodes has the highest storage priority value among all preprocessed data due to its importance for construction decisions. At the same time, in the data warehouse sorting information, determine the data warehouse with the largest storage performance value and record it as the priority storage warehouse. Suppose among 5 data warehouses, the storage performance value of the second data warehouse is evaluated to be the highest. Subsequently, store this priority storage data in the priority storage warehouse to form the priority storage information, ensuring that important data can be stored in the data warehouse with the best performance to guarantee efficient access and processing of the data.
[0076] For the priority storage data, obtain the associated data that has a subordinate relationship with it. Taking road construction as an example, if the priority storage data is the actual construction progress of a certain section on the same day, its associated data may include the construction plan progress of this section, the real-time inventory data of the construction materials used, etc., because when viewing the actual construction progress on the same day, the construction plan progress and the material inventory data often need to be referred to together, and there is a subordinate relationship between them.
[0077] For the remaining data warehouses except the priority storage warehouse, obtain their processing speeds for processing data of the same type as the associated data. For example, the associated data is the construction material inventory data, and the processing speeds of the remaining data warehouses for such data are different. At the same time, clarify the data volume of the associated data. Suppose the associated construction material inventory data volume is 1000 record entries. According to the formula "processing time = data volume ÷ processing speed", calculate the time required for each remaining data warehouse to process the associated data. For example, if the processing speed of the first data warehouse for the construction material inventory data is 200 entries per second, then the time for it to process these 1000 associated data is 1000 ÷ 200 = 5 seconds.
[0078] Determine the minimum processing speed requirement corresponding to the associated data, which is usually determined by the construction business process and real-time requirements. For example, according to the construction management regulations, the construction material inventory data must be processed within 3 seconds to ensure that the construction progress is not affected, that is, the minimum processing speed is 1000 ÷ 3 ≈ 333 entries per second. Taking this minimum processing speed as the standard, screen the remaining data warehouses, and use the data warehouses whose processing speeds reach or exceed this standard as the preselected storage warehouses. If the processing speed of the third data warehouse for the construction material inventory data is 400 entries per second, which meets the minimum processing speed requirement, it becomes one of the preselected storage warehouses.
[0079] If, after screening, the number of preselected storage bins is zero, this means that none of the remaining data bins can meet the minimum processing speed requirements for the associated data. In this case, the associated data is stored together with the priority storage data in the priority storage bin, and associated storage information is generated to record this special storage arrangement for subsequent query and management.
[0080] When the number of preselected storage bins is not zero, select the data bin with the shortest processing time from these preselected storage bins. For example, if there are 3 preselected storage bins with processing times for associated data of 4 seconds, 3 seconds, and 5 seconds respectively, select the data bin with a processing time of 3 seconds as the storage location for the associated data, and generate associated storage information.
[0081] Integrate the priority storage information and the associated storage information to form storage management information. This information comprehensively records the storage locations and arrangements of the priority storage data and its associated data, providing clear guidance for the storage management of road construction data. For example, the priority storage data (the actual construction progress of a certain road section on the current day) is stored in the second data bin, and its associated data (the construction plan progress, material inventory data, etc. of this road section) is stored in the preselected data bin with the shortest processing time (such as data bin 4). These pieces of information are integrated to form complete storage management information.
[0082] Embodiment 2. Please refer to Figure 2 , this application provides a data integration management system for municipal road construction sections based on a distributed database, including:
[0083] A data acquisition module, which is used to acquire construction data, perform data cleaning on it to obtain preprocessed data, and at the same time transmit it to the storage and analysis module. The specific processing method is the same as the processing process in step S1 of Embodiment 1;
[0084] A data storage and analysis module, which is used to store and analyze the acquired preprocessed data. By classifying the preprocessed data, the same type of data is obtained. At the same time, the usage frequency and increase frequency of the same type of data within a time period are calculated to obtain real-time values, and the demand values are obtained by calculating the maximum delay standard and the minimum processing speed. Then, the sum of the real-time value and the demand value is calculated to obtain the storage priority value of the preprocessed data, and data sorting information is generated by sorting from largest to smallest. At the same time, it is transmitted to the data management and analysis module. The specific processing method is the same as the processing process in step S2 of Embodiment 1;
[0085] Data warehouse storage analysis module, which is used to establish the corresponding data warehouse, and at the same time analyze the reading speed and processing duration of different data in the data warehouse. Select one of the different data in the data warehouse as the analysis object, and respectively obtain its reading speed when the network is stable and unstable, calculate the sum of the two reading speeds, and then calculate the average value to obtain the reading value of the analysis object. The processing duration analysis covers the concurrent processing duration and the response duration. Obtain the concurrent processing duration corresponding to the task, match it with the judgment standard numerical interval obtained from a large amount of data analysis to obtain the processing duration assignment. Similarly, match the response duration to obtain the response duration assignment. Add the processing duration assignment and the response duration assignment to obtain the processing assignment. Add the reading value and the processing assignment to obtain the storage performance value of the data warehouse. Sort the data warehouses according to the storage performance value from large to small, generate the data warehouse sorting information, and transmit it to the data management analysis module. The specific processing method is the same as the processing process of step S3 in Embodiment 1;
[0086] Data management analysis module, which is used to find the preprocessed data with the largest storage priority value from the data sorting information, denoted as the priority storage data. At the same time, determine the data warehouse with the largest storage performance value in the data warehouse sorting information, called the priority storage warehouse, match and store the two, record the priority storage information, obtain the associated data of the priority storage data, and based on the associated data, obtain the processing speed of each remaining data warehouse for the same type of data. Combine the associated data volume to calculate the processing time. Screen the remaining data warehouses according to the minimum processing speed corresponding to the associated data to obtain the preselected storage warehouses. If the number of preselected storage warehouses is zero, store the associated data and the priority storage data together to generate the associated storage information; if not, select the preselected storage warehouse with the shortest processing time to store the associated data to generate the associated storage information. Combine the associated and priority storage information to obtain the storage management information. Analyze all the preprocessed data according to this process, and finally generate the complete storage management information. The specific processing method is the same as the processing process of step S4 in Embodiment 1, and at the same time transmit the storage management information to the management information output module;
[0087] Management information output module, which is used to manage the obtained storage management information.
[0088] For some data in the above formula, only their numerical values are taken for calculation, and the parameter units are not substituted for calculation. At the same time, the content not described in detail in this specification belongs to the prior art well-known to those skilled in the art.
[0089] The above embodiments are only used to illustrate the technical method of the present invention and not to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical method of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical method of the present invention.
Claims
1. A method for integrated management of data in municipal road construction sections based on a distributed database, characterized in that, The method specifically includes the following steps: Collect construction data, and through data cleaning and data transformation, obtain preprocessed data; Classify the preprocessed data, calculate the sum of the usage frequency and the increase frequency of the same type of data within the time period to obtain the real-time value, calculate the sum of the maximum delay standard and the minimum processing speed to obtain the demand value, and the maximum delay standard refers to the maximum tolerance threshold for data reading delay in the business process during the processing of municipal road construction data, which is determined by industry standards or enterprise self-defined specifications, and the minimum processing speed represents the lower limit of the data processing volume that must be completed per unit time, which is determined by the real-time requirements of the construction business; add the real-time value and the demand value to obtain the storage priority value, and generate data sorting information in descending order; Construct a data warehouse, select the analysis object, respectively obtain its reading speeds when the network is stable and unstable, and calculate the average value of the two to obtain the reading value; Obtain the concurrent processing duration and the response duration. The concurrent processing duration refers to the average time taken for a single task to be submitted to completion when the distributed database processes multiple data tasks simultaneously, which is used to evaluate the batch processing ability of the data warehouse. The response duration refers to the time interval from when the user requests data to when the system returns the result, which is used to measure the real-time response ability of the data warehouse. Respectively match them with the judgment criteria to obtain the processing duration assignment and the response duration assignment. The judgment criteria refer to the pre-set performance evaluation threshold, which is obtained through the statistics of a large amount of historical data and is used to quantitatively evaluate the processing efficiency of the data warehouse. Add the two to obtain the processing assignment, and add the obtained value to the processing assignment to obtain the storage performance value. Sort in descending order according to the storage performance value to generate the data warehouse sorting information; Find the preprocessed data with the largest storage priority value from the data sorting information, that is, the priority storage data; determine the data warehouse with the largest storage performance value from the data warehouse sorting information, that is, the priority storage warehouse, and match the two for storage to generate the priority storage information; Obtain the associated data of the priority storage data. The associated data refers to the data that has a business logic association with the priority storage data, and the two have a strong correlation in the construction business and need to be used collaboratively. Calculate the processing time based on the amount of associated data and the speed of each remaining data warehouse in processing the same type of data, and screen the remaining data warehouses with the minimum processing speed corresponding to the associated data to obtain the preselected storage warehouses; If the number of preselected storage warehouses is zero, store the associated data and the priority storage data together. If not, select the preselected storage warehouse with the shortest processing time to store the associated data, generate the associated storage information, and comprehensively combine the associated and priority storage information to obtain the storage management information.
2. The method for data integration management of municipal road construction sections based on a distributed database according to claim 1, characterized in that, The specific method for calculating the sum of the usage frequency and the increase frequency of the same type of data within the time period to obtain the real-time value is as follows: Classify the obtained preprocessed data to obtain the same type of data. Then, obtain the preprocessed data in the same type of data and label it as i, where i = 1, 2,..., j, and j represents the number of preprocessed data. Taking the time period T as a cycle, obtain the number of views of the preprocessed data i within the time period T, calculate the usage frequency corresponding to the preprocessed data i, and at the same time obtain the increase amount corresponding to the preprocessed data i within the time period T. According to the formula increase frequency = increase amount ÷ time T, calculate the increase frequency corresponding to the preprocessed data i, and calculate the sum of the values of the usage frequency and the increase frequency, denoted as the real-time value of the preprocessed data i.
3. The method for integrated management of municipal road construction section data based on a distributed database according to claim 1, characterized in that The specific method for calculating the sum of the maximum delay standard and the minimum processing speed to obtain the required value is as follows: Obtain the same type of data corresponding to the preprocessed data i, and obtain the corresponding maximum delay standard. At the same time, obtain the minimum processing speed corresponding to the preprocessed data i, and then calculate the sum of the two values to obtain the required value of the preprocessed data i.
4. The method for data integration management of municipal road construction sections based on a distributed database according to claim 1, characterized in that, The specific method for generating data sorting information in descending order is as follows: Sum the real-time value and the required value of the preprocessed data i to obtain the storage priority value of the preprocessed data i, and sort it in descending order according to the storage priority value to generate data sorting information.
5. The method for integrating and managing data of municipal road construction sections based on a distributed database according to claim 1, characterized in that, The specific method for calculating the average value of the two to obtain the read value is as follows: Obtain all data warehouses and label them as a, where a = 1, 2, …, b, and b represents the number of data warehouses. Obtain any type of data corresponding to the data warehouse, denoted as the analysis object. Obtain the read speed of the analysis object under different network conditions, and at the same time calculate the average value of the read speeds corresponding to different network conditions of the analysis object. Specifically, calculate the sum of the read speeds corresponding to stability and instability, and calculate the average value, denoted as the read value of the analysis object.
6. The method for integrated management of municipal road construction section data based on a distributed database according to claim 1, wherein The specific method for generating data warehouse sorting information is as follows: Analyze the concurrent processing duration, obtain the processing durations corresponding to n real-time data processing tasks, where n is specifically set by the operator. At the same time, match the obtained processing durations with the judgment standard, and obtain the corresponding assignment denoted as the processing duration assignment. Similarly, analyze the response duration to obtain the corresponding assignment denoted as the response duration assignment, and calculate the sum of the values of the processing duration assignment and the response duration assignment denoted as the processing assignment. Then calculate the sum of the read value and the processing assignment, denoted as the storage performance value of the data warehouse; And sort in descending order according to the obtained storage performance value to generate data warehouse sorting information.
7. The method for integrated management of municipal road construction section data based on a distributed database according to claim 1, characterized in that, The specific method for generating the priority storage information is as follows: Based on the obtained data sorting information and data warehouse sorting information, conduct a comprehensive analysis. Determine the associated data by analyzing the relevance of the data. Obtain the preprocessed data with the largest storage priority value in the data sorting information, denoted as the priority storage data. At the same time, obtain the data warehouse with the largest storage performance value in the data warehouse sorting information, denoted as the priority storage warehouse, and match and store the two to generate the priority storage information.
8. The method for integrated management of municipal road construction section data based on a distributed database according to claim 1, characterized in that, The specific method for obtaining the preselected storage warehouse is as follows: Then obtain the associated data of the priority storage data, and at the same time conduct a matching analysis on the remaining data warehouses based on the associated data; Obtain the processing speed corresponding to the remaining data warehouse, where the processing speed here is expressed as the processing speed of the same type as the associated data. At the same time, obtain the data volume corresponding to the associated data, calculate the processing time of the remaining data warehouse for the associated data, then obtain the minimum processing speed corresponding to the associated data, and screen the remaining data warehouses based on the minimum processing speed to obtain the preselected storage warehouses.
9. The method for integrated management of municipal road construction section data based on a distributed database according to claim 1, characterized in that, The specific method for obtaining the storage management information is as follows: If the number of preselected storage warehouses is zero, store the associated data together with the priority storage data and generate associated storage information. If the number of preselected storage warehouses is not zero, take the preselected storage warehouse with the shortest processing time as the standard to generate associated storage information, and at the same time, obtain the storage management information by integrating the associated storage information and the priority storage information.
10. A data integration management system for municipal road construction sections based on a distributed database, which is used to execute the data integration management method for municipal road construction sections based on a distributed database according to any one of claims 1-9, characterized in that, It includes: A data acquisition module, which is used to acquire construction data, perform data cleaning on it to obtain preprocessed data, and transmit it to the storage analysis module at the same time; A data storage analysis module, which is used to perform storage analysis on the acquired preprocessed data. By classifying the preprocessed data, the same type of data is obtained. At the same time, calculate the usage frequency and increase frequency of the same type of data within the time period to obtain real-time values, and calculate the demand values by calculating the maximum delay standard and the minimum processing speed. Then calculate the sum of the real-time value and the demand value to obtain the storage priority value of the preprocessed data, sort it from largest to smallest to generate data sorting information, and transmit it to the data management analysis module at the same time; A data warehouse storage analysis module, which is used to establish a corresponding data warehouse, and at the same time analyze the reading speed and processing duration of different data corresponding to the data warehouse. Select one of the different data in the data warehouse as the analysis object, respectively obtain its reading speeds when the network is stable and unstable, calculate the sum of the two reading speeds, and then calculate the average value to obtain the reading value of the analysis object. The processing duration analysis covers the concurrent processing duration and the response duration. Obtain the concurrent processing duration corresponding to the task, match it with the judgment standard numerical range obtained from a large amount of data analysis to obtain the processing duration assignment. Similarly, match the response duration to obtain the response duration assignment. Add the processing duration assignment and the response duration assignment to obtain the processing assignment. Add the reading value and the processing assignment to obtain the storage performance value of the data warehouse. Sort according to the storage performance value from largest to smallest to generate data warehouse sorting information, and transmit it to the data management analysis module; A data management analysis module, which is used to find the preprocessed data with the largest storage priority value from the data sorting information, denoted as the priority storage data. At the same time, determine the data warehouse with the largest storage performance value in the data warehouse sorting information, called the priority storage warehouse, match and store the two, record the priority storage information, obtain the associated data of the priority storage data, based on the associated data, obtain the speed of each remaining data warehouse for processing the same type of data, combine the associated data volume, calculate the processing time, and screen the remaining data warehouses according to the minimum processing speed corresponding to the associated data to obtain the preselected storage warehouses; If the number of preselected storage warehouses is zero, store the associated data and the priority storage data together to generate associated storage information; If it is not zero, select the preselected storage bin with the shortest processing time to store the associated data, generate associated storage information, synthesize the associated and priority storage information to obtain storage management information, analyze all the preprocessed data according to this process, finally generate complete storage management information, and at the same time transmit the storage management information to the management information output module; The management information output module is used to manage the obtained storage management information.
Citation Information
Patent Citations
A data integration service intelligent management method
CN118885476B
Industrial operation system data lake construction method based on data warehouse
CN114490886A
Decision analysis-based data center system
CN115600849A