Sandstone system multi-source data integration method and system based on task scheduling and distribution
By constructing a multi-source data configuration database and a distributed task scheduling module, combined with multi-protocol parsing and multi-objective planning models, the problem of multi-source data integration and distribution in the sand and gravel system was solved, realizing unified integration and efficient distribution of data, improving the real-time performance and stability of data, and supporting the refined management of the sand and gravel system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA GEZHOUBA GROUP CO LTD
- Filing Date
- 2026-03-30
- Publication Date
- 2026-07-31
AI Technical Summary
The difficulty in integrating multi-source heterogeneous data in sand and gravel systems, inefficient data acquisition and scheduling, disconnect between distribution and scheduling, and unstable transmission in high-altitude environments lead to insufficient real-time data, redundancy, or missing data, affecting the implementation of production monitoring and fault early warning services.
A multi-source data configuration database is constructed, a multi-protocol parsing module library is built, and the xxl-job distributed task scheduling module is used to capture data according to the configured frequency. The data is then cleaned, transformed, and standardized. Finally, the optimal scheduling and ordering of data distribution is achieved through a multi-objective linear programming model, and the data is distributed to the target routing application.
It has achieved unified integration and precise control of multi-source data, improved the comprehensiveness, real-time performance and stability of data collection, and promoted the refined and intelligent operation of the sand and gravel system.
Smart Images

Figure CN122489635A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of data processing technology, and more specifically, relates to a method and system for multi-source data integration of sand and gravel systems based on task scheduling and distribution. Background Technology
[0002] The sand and gravel system is a core production link in high-altitude hydropower projects, mining operations, and other projects, and its operational efficiency directly affects the project's progress and quality. Currently, sand and gravel systems involve various types of equipment, including PLC control systems, smart meters, smart weighbridges, video surveillance, and environmental monitoring equipment. The data sources are scattered, and the protocol formats are diverse (such as OPC, Modbus, and SDK proprietary protocols), forming a multi-source heterogeneous data structure.
[0003] In traditional sand and gravel systems, data acquisition often relies on manual recording or independent data collection by a single device, lacking a unified integrated management and control mechanism. On the one hand, the data transmission protocols of different devices are incompatible, making it difficult to seamlessly connect data between the device layer and the system layer, forming "data silos." On the other hand, data acquisition lacks flexible and efficient task scheduling strategies, making it impossible to dynamically adjust the acquisition frequency according to production needs. This results in problems such as insufficient real-time performance, data redundancy, or missing data, which in turn affects the implementation of core businesses such as production monitoring, fault early warning, and quality traceability.
[0004] Furthermore, high-altitude sand and gravel systems face harsh conditions such as extreme cold and large temperature differences, placing higher demands on the stability and reliability of data acquisition. Traditional data integration methods struggle to cope with special scenarios such as frequent equipment offline and data transmission delays. Therefore, there is an urgent need for a multi-source data integration method based on task scheduling to overcome challenges such as multi-protocol compatibility, efficient data aggregation, and stable transmission, thereby achieving unified integration and precise control of multi-source data from sand and gravel systems. Summary of the Invention
[0005] This invention aims to solve the problems of difficult integration of multi-source heterogeneous data in sand and gravel systems, inefficient data acquisition and scheduling, disconnect between distribution and scheduling, and unstable transmission in high-altitude environments. It achieves unified integration and precise control of multi-source data, improves the comprehensiveness, real-time performance, and stability of data acquisition, promotes collaborative optimization of scheduling and distribution, and helps sand and gravel systems operate in a refined and intelligent manner.
[0006] To address the aforementioned deficiencies or improvement needs of existing technologies, as a first aspect of this invention, the present invention provides a method for multi-source data integration in a sand and gravel system based on task scheduling and distribution, comprising: S1. Construct a configuration database for multi-source data in the sand and gravel system, configuring the acquisition frequency, acquisition priority, and data transmission protocol for each data source; based on the different data transmission protocols covered in the sand and gravel system, write and construct corresponding multi-protocol data parsing code modules to form a parsing module library that can be called on demand; S2. Start the distributed task scheduling module based on xxl-job. The scheduling module queries the collection frequency of each data source from the configuration database and captures multi-source raw data of the sand and gravel system at regular intervals according to the collection frequency. S3. The parsed data is transmitted to the data processing module, which cleans, transforms and standardizes the collected raw data, and adds associated labeling information such as equipment ownership, production process and collection task ID to the standardized data to form standardized multi-source data of the sand and gravel system and complete the unified integration of heterogeneous data. S4. Standardized multi-source data from the sand and gravel system is pushed to a distributed real-time data bus built on Kafka. The data enters the queue of the data bus and waits for distribution. Based on a multi-objective linear programming model of task scheduling priority and data distribution timeliness, combined with the priority coefficient corresponding to the collection task ID of each data to be distributed, the business response weight coefficient of the target routing application, and the maximum effective data transmission timeliness threshold, an objective function is constructed and the optimal scheduling order for data distribution is obtained by solving it. S5. Read the data to be distributed from the data queue according to the optimal scheduling sorting. Based on the association labeling information of the data to be distributed, query the corresponding business distribution route from the preset routing database. After checking the liveness status of the target routing application, distribute the data to the target routing application according to the optimal scheduling sorting, thereby realizing the coordinated optimization of task scheduling and data distribution.
[0007] Furthermore, the data source in S1 includes equipment data and software data of the sand and gravel system. The equipment data covers the operation data of PLC control system, smart meters, smart weighbridge, environmental monitoring equipment, and video surveillance equipment. The software data covers database change logs, application operation data, and production business data.
[0008] Furthermore, in S2, the scheduling module queries the collection priority of each data source from the configuration database, sorts the processing queue of the multi-source raw data according to the priority, queries the transmission protocol matching each data source, and calls the corresponding data parsing module from the parsing module library according to the transmission protocol to perform protocol parsing on the sorted multi-source raw data and output the parsed data.
[0009] Furthermore, the multi-objective linear programming model in S4 is specifically as follows: A multi-objective programming method with time-effect decay and resource elasticity constraints is constructed. It integrates the mathematical characteristics of task priority dynamic decay, distribution time-effect actual benefits, and resource occupation elasticity threshold in the actual production scenario of sand and gravel system. The multi-dimensional objective dimension is unified through normalization processing, and multi-objective collaborative optimization is completed by relying on the mathematical coupling between objectives and the constraint boundary limit. The overall objective function of the model is: This includes dynamic priority optimization terms. Distribution timeliness optimization items Resource usage optimization items The three core components correspond to the dynamic matching requirements of task scheduling priorities, the timeliness and effectiveness requirements of data distribution, and the rational utilization requirements of system bus resources, respectively. The model simultaneously satisfies the preset constraints, which are as follows: in This is the actual waiting time after the data to be distributed enters the Kafka data bus queue. This represents the maximum effective transmission time threshold for this data. For the first Estimated transmission time for the data to be distributed; This is the resource elasticity coefficient, which is a preset threshold for limiting the over-allocation of bus resources based on the actual resource redundancy of the sand and gravel system. For the first Distribution scheduling decision variables for data to be distributed; For the first Resource consumption for distributing data to be distributed; This refers to the basic resource usage of the Kafka data bus. This represents the maximum available distribution resources on the Kafka data bus. and After normalization After normalization of the difference ratio, the range of values for all three terms is [missing information]. , This is the sequence identifier for the data to be distributed, with a value of [value]. arrive Positive integers.
[0010] Furthermore, the dynamic priority optimization term for: in For the first The basic priority coefficient for each piece of data to be distributed corresponds to the collection task ID, and its value is assigned by the collection priority in the database configured by the claims. , Assign the highest priority value to the data collection task; For the first The business response weight coefficient of the target routing application for the data to be distributed is assigned differently based on the importance of the core production business and auxiliary management business of the sand and gravel system; This is a priority timeliness decay factor. This factor decays exponentially as the proportion of data waiting time to the maximum effective timeliness threshold increases, thereby realizing dynamic adjustment of task scheduling priority. For the first Distribution scheduling decision variables for data to be distributed. Indicates the execution of dispatch scheduling. This indicates that the action will not be taken for the time being. This represents the total number of data items to be distributed in the Kafka data bus queue.
[0011] Furthermore, the distribution timeliness optimization item for: in For the first The estimated transmission time of the data to be distributed is determined by the network link latency of the target routing application and the data size through a linear fitting model. Calculations show that For link transmission coefficient, For the first The amount of data in each data item For link base latency; For the first The timeliness of the data to be distributed is a valid decision variable. express The data remains within its valid timeframe after transmission. express The data will expire after transmission.
[0012] Furthermore, the distribution timeliness optimization item for: in For the first The resource consumption for distributing each piece of data to be distributed is quantified by the actual bandwidth consumption for data transmission. This represents the basic resource usage of the Kafka data bus, and the minimum bandwidth usage required to ensure the basic operation of the system. This represents the maximum available distribution resources of the Kafka data bus, quantified by the maximum available bandwidth of the bus.
[0013] Furthermore, the state detection process for the corresponding service distribution route in S5 is as follows: If the target business application is alive, establish a data transmission connection with the target business application and distribute the data to be distributed to the corresponding business modules such as production process management, equipment operation management, visual monitoring, and platform basic management. If the target business application is not alive, check whether the data to be distributed has expired. If it has not expired, put the data back to the end of the data queue to wait for redistribution. If it has expired, discard the data directly to complete the on-demand flow of multi-source data in the sand and gravel system.
[0014] As a second aspect of the present invention, a multi-source data integration system for sand and gravel systems based on task scheduling and distribution is also provided, comprising: The database and parsing module configuration unit is used to build a configuration database for multi-source data in the sand and gravel system, configure the collection frequency, collection priority and data transmission protocol of each data source; according to the different data transmission protocols covered in the sand and gravel system, write and build the corresponding multi-protocol data parsing code module to form a parsing module library that can be called on demand; The distributed scheduling data acquisition unit is used to start the distributed task scheduling module based on xxl-job. The scheduling module queries the acquisition frequency of each data source from the configuration database and captures multi-source raw data of the sand and gravel system at regular intervals according to the acquisition frequency. The heterogeneous data standardization and integration unit is used to transmit the parsed data to the data processing module, clean, transform and standardize the collected raw data, and add associated labeling information such as equipment ownership, production process and collection task ID to the standardized data to form standardized multi-source data of the sand and gravel system and complete the unified integration of heterogeneous data. The planning model distribution and sorting unit is used to push standardized multi-source data from the sand and gravel system to a distributed real-time data bus built on Kafka. The data enters the queue of the data bus and waits for distribution. Based on a multi-objective linear programming model that combines task scheduling priority and data distribution timeliness, the objective function is constructed and the optimal scheduling and sorting of data distribution is obtained by combining the priority coefficients corresponding to the collection task IDs of each data to be distributed, the business response weight coefficients of the target routing application, and the maximum effective transmission timeliness threshold of the data. The targeted data distribution unit is used to read the data to be distributed from the data queue according to the optimal scheduling sorting, query the corresponding business distribution route from the preset routing database according to the association label information of the data to be distributed, detect the liveness status of the target routing application, and then distribute the data to the target routing application according to the optimal scheduling sorting, thereby realizing the coordinated optimization of task scheduling and data distribution.
[0015] As a third aspect of the invention, a computer-readable storage medium is also provided, on which a computer program is stored, which is executed by a processor, according to any one of the claims, a method for multi-source data integration of a sand and gravel system based on task scheduling and distribution.
[0016] In summary, compared with the prior art, the above-described technical solutions conceived by this invention can achieve the following beneficial effects: 1. The multi-source data integration method for sand and gravel systems based on task scheduling and distribution of this invention constructs a multi-source data configuration database and builds a multi-protocol parsing module library. Simultaneously, relying on the xxl-job distributed task scheduling module, it completes the timed acquisition and protocol parsing of multi-source raw data according to the configured acquisition frequency and priority, achieving standardized management and control of multi-source heterogeneous data from the data acquisition source. This method specifically solves the data integration problem caused by the dispersed data sources and diverse transmission protocols in sand and gravel systems. It allows data sources from different devices and with different protocols to be collected in an orderly manner according to preset rules, effectively breaking down data barriers between the device layer and the system layer. It achieves comprehensive acquisition of data from various devices in the sand and gravel system, such as PLC control systems, smart meters, and environmental monitoring equipment, ensuring the standardization and comprehensiveness of the data acquisition process and laying a unified acquisition foundation for subsequent data processing and application.
[0017] 2. The multi-source data integration method for sand and gravel systems based on task scheduling and distribution of the present invention cleans, transforms, and standardizes the parsed raw data, adding associated annotation information such as equipment affiliation, production process, and collection task ID to the standardized data, thus achieving unified integration of heterogeneous data in the sand and gravel system. This processing method solves the problems of inconsistent data formats and incomplete data information in traditional data acquisition modes, transforming heterogeneous data of different types and formats into standardized data that the system can recognize. At the same time, through associated annotation, each data point has clear traceability information, effectively improving the regularity and traceability of the data, avoiding interference from abnormal and redundant data to subsequent business applications, and providing standardized data support for accurate data distribution and efficient utilization.
[0018] 3. The multi-source data integration method for sand and gravel systems based on task scheduling and distribution of this invention pushes standardized data to a Kafka distributed real-time data bus. It solves the optimal scheduling order for data distribution based on a multi-objective linear programming model that considers task scheduling priority and data distribution timeliness. Data is read according to the order and the corresponding route is queried. After detecting the liveness status of the target application, targeted distribution is completed. This distribution method achieves coordinated optimization of task scheduling and data distribution, solving the problems of insufficient real-time performance and disconnect between scheduling and distribution in traditional data distribution. It allows data to be accurately and orderly distributed to the corresponding business modules according to production business needs. Simultaneously, the liveness detection of the target application ensures the effectiveness of data transmission, realizing on-demand data flow and efficient sharing, promoting data interoperability between various business modules of the sand and gravel system, and improving the response efficiency of production scheduling. Attached Figure Description
[0019] Figure 1 This is a flowchart of a multi-source data integration method for a sand and gravel system based on task scheduling and distribution, according to an embodiment of the present invention. Figure 2 This is a schematic diagram of the overall architecture design of an embodiment of the present invention; Figure 3 This is a schematic diagram of the data acquisition task scheduling configuration based on xxl-job according to an embodiment of the present invention; Figure 4 This is a schematic diagram illustrating data parsing and conversion in an embodiment of the present invention; Figure 5 This is a schematic diagram of data distribution based on Kafka according to an embodiment of the present invention; Figure 6 This is a schematic diagram of the system units in an embodiment of the present invention. Detailed Implementation
[0020] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.
[0021] Example 1 Please refer to Figure 1 This embodiment 1 provides a method for integrating multi-source data of a sand and gravel system based on task scheduling and distribution, including: S1. Construct a configuration database for multi-source data in the sand and gravel system, configuring the acquisition frequency, acquisition priority, and data transmission protocol for each data source; based on the different data transmission protocols covered in the sand and gravel system, write and construct corresponding multi-protocol data parsing code modules to form a parsing module library that can be called on demand; S2. Start the distributed task scheduling module based on xxl-job. The scheduling module queries the collection frequency of each data source from the configuration database and captures multi-source raw data of the sand and gravel system at regular intervals according to the collection frequency. S3. The parsed data is transmitted to the data processing module, which cleans, transforms and standardizes the collected raw data, and adds associated labeling information such as equipment ownership, production process and collection task ID to the standardized data to form standardized multi-source data of the sand and gravel system and complete the unified integration of heterogeneous data. S4. Standardized multi-source data from the sand and gravel system is pushed to a distributed real-time data bus built on Kafka. The data enters the queue of the data bus and waits for distribution. Based on a multi-objective linear programming model of task scheduling priority and data distribution timeliness, combined with the priority coefficient corresponding to the collection task ID of each data to be distributed, the business response weight coefficient of the target routing application, and the maximum effective data transmission timeliness threshold, an objective function is constructed and the optimal scheduling order for data distribution is obtained by solving it. S5. Read the data to be distributed from the data queue according to the optimal scheduling sorting. Based on the association labeling information of the data to be distributed, query the corresponding business distribution route from the preset routing database. After checking the liveness status of the target routing application, distribute the data to the target routing application according to the optimal scheduling sorting, thereby realizing the coordinated optimization of task scheduling and data distribution.
[0022] Please refer to Figure 2 , Figure 3 , Figure 4 as well as Figure 5 This embodiment 1 further elaborates on the above steps.
[0023] (1) Database and parsing module configuration Given the dispersed data sources and diverse transmission protocols in sand and gravel systems, to achieve standardized management and control of multi-source data from the collection source, it is first necessary to build a multi-source data configuration database adapted to the operational requirements of the sand and gravel system. For each data source within the system, corresponding collection frequencies, collection priorities, and data transmission protocols should be configured for their respective application scenarios, providing a clear configuration basis for subsequent data collection scheduling and execution. Simultaneously, based on the various data transmission protocols actually covered in the sand and gravel system, corresponding multi-protocol data parsing code modules should be written and constructed. All parsing modules should be integrated into a parsing module library that can be called on demand, ensuring that data sources with different protocols can be effectively parsed.
[0024] The database configuration covers a complete range of data sources, including equipment and software data for the sand and gravel system. The equipment data consists of real-time monitoring and operational data from various production and operating equipment, specifically covering operational data from PLC control systems, smart meters, smart weighbridges, environmental monitoring equipment, and video surveillance equipment. The software data includes business-related data generated by various system software, such as database change logs, application operation data, and production business data, thus achieving full coverage of all data sources for the sand and gravel system.
[0025] After configuration and module setup are complete, start the task scheduling module to perform data collection and parsing. The scheduling module will automatically retrieve the collection frequency of each data source from the configuration database, for example: Configure high-frequency timed tasks (e.g., every 10 seconds) for real-time monitoring data of equipment (such as the operating status of conveyor belts and the vibration frequency of crushers) to ensure data timeliness; configure low-frequency timed tasks (e.g., every hour) for periodic statistical data (such as hourly output and energy consumption statistics) to reduce data redundancy; set high priority for data acquisition tasks of key equipment (such as crushers and mixing hosts) to ensure priority allocation of resources.
[0026] Subsequently, the scheduling module continues to retrieve the acquisition priority of each data source, sorts the external data processing queue according to priority, allocates priority processing resources to the data collected by key equipment such as crushers and mixers, queries the transmission protocol corresponding to each data, and accurately calls the corresponding parsing module from the parsing module library to process the data according to the protocol type. Finally, it outputs standardized data that has been parsed, realizing the orderly and efficient parsing of data from different sources and with different protocols, and meeting the needs of multi-source data acquisition and parsing in the sand and gravel system.
[0027] (2) Distributed scheduling data acquisition Given that the sand and gravel system has multiple data sources that are scattered and use various protocols, and that different data have different requirements for real-time performance and importance, a simple centralized scheduling approach is insufficient to meet the needs of efficient data collection and orderly processing. Therefore, the xxl-job distributed task scheduling framework is adopted as the core control module in the data integration process. It is specifically responsible for formulating and executing data collection task strategies to ensure the orderliness and efficiency of data collection and parsing.
[0028] After the distributed task scheduling module is started, it first automatically establishes a connection with the multi-source data configuration database built in the early stage, queries the collection frequency corresponding to each data source, and, in combination with the production scenario requirements of the sand and gravel system, strictly triggers data collection tasks on a timed basis according to the collection frequency, captures multi-source raw data in the sand and gravel system, and ensures that different types of data can be collected in a timely manner as required.
[0029] Simultaneously, the scheduling module queries the configuration database for the collection priorities of each data source. Based on these priorities, it sorts the captured multi-source raw data into processing queues, prioritizing data sources from critical equipment such as crushers and mixers to ensure the priority processing of core production data. After sorting, the scheduling module further queries the data transmission protocols matching each data source. Based on the determined transmission protocol type, it calls the corresponding data parsing module from the previously built multi-protocol parsing module library to perform targeted protocol parsing on the sorted multi-source raw data. This transforms the heterogeneous raw data into a format that the system can recognize and process, ultimately outputting the parsed data to provide reliable support for subsequent data processing stages.
[0030] This scheduling method allows for customized configuration of the execution frequency, triggering conditions, retry mechanisms, and failure handling strategies for data acquisition tasks based on data type, device priority, and production scenario requirements. This enables dynamic scheduling and flexible management of data acquisition tasks, effectively adapting to the complex operational needs of sand and gravel systems.
[0031] (3) Standardization and integration of heterogeneous data The raw data generated by the sand and gravel system during actual operation suffers from problems such as disordered formats, inconsistent types, and partial data distortion. Direct use of this data can affect the accuracy of subsequent business applications. Therefore, it is necessary to perform standardization processing on the parsed data. After scheduling and protocol parsing, the data is transmitted to a dedicated data processing module, which performs cleaning, transformation, and standardization operations on the data sequentially according to preset rules.
[0032] During the processing, firstly, based on reasonable thresholds for normal equipment operation, abnormal data exceeding the normal range is removed. Simultaneously, redundant information from repeated collections is filtered out, and missing key operational data is appropriately supplemented to ensure data integrity and validity. Subsequently, various data types undergo format standardization processing, converting heterogeneous data (numerical, character, Boolean, etc.) generated by different devices into a unified specified format and establishing unified data field specifications to ensure a consistent structural system for data from different sources.
[0033] After standardizing the format, each data entry is labeled with relevant information such as the equipment's production stage and task ID, giving the data a clear source and business attribute, facilitating subsequent data distribution and traceability. Through these processes, the originally heterogeneous and scattered sand and gravel system data is integrated into standardized data with a unified structure and complete information, providing stable and reliable data support for subsequent data distribution and business applications.
[0034] (4) Distribution and sorting of planning models Given that the standardized multi-source data of the sand and gravel system needs to be accurately and efficiently distributed to various business application modules, and that different data have different priorities, timeliness requirements and resource consumption, if there is no effective scheduling and sorting distribution mechanism, problems such as chaotic data distribution, insufficient timeliness or waste of resources are likely to occur. Therefore, it is necessary to achieve orderly data distribution through a reasonable scheduling model and distribution process.
[0035] After standardization, the multi-source data from the sand and gravel system is pushed to a distributed real-time data bus built on Kafka, where it enters a queue awaiting distribution. To achieve optimal data distribution, a multi-objective programming method with time-based decay and resource elasticity constraints is employed. This method integrates the mathematical characteristics of dynamic task priority decay, actual distribution time-based benefits, and resource occupancy elasticity thresholds in the actual production scenario of the sand and gravel system. Through normalization, the multi-dimensional objective dimensions are unified, and multi-objective collaborative optimization is achieved by relying on the mathematical coupling between objectives and the constraint boundary limits. The overall objective function of the model is: This includes dynamic priority optimization terms. Distribution timeliness optimization items Resource usage optimization items The three core components correspond to the dynamic matching requirements of task scheduling priorities, the timeliness and effectiveness requirements of data distribution, and the rational utilization requirements of system bus resources, respectively. Among them, the dynamic priority optimization term for: in For the first The basic priority coefficient for each piece of data to be distributed corresponds to the collection task ID, and its value is assigned by the collection priority in the database configured by the claims. , Assign the highest priority value to the data collection task; For the first The business response weight coefficient of the target routing application for the data to be distributed is assigned differently based on the importance of the core production business and auxiliary management business of the sand and gravel system; This is a priority timeliness decay factor. This factor decays exponentially as the proportion of data waiting time to the maximum effective timeliness threshold increases, thereby realizing dynamic adjustment of task scheduling priority. For the first Distribution scheduling decision variables for data to be distributed. Indicates the execution of dispatch scheduling. This indicates that the action will not be taken for the time being. This represents the total number of data items to be distributed in the Kafka data bus queue.
[0036] In addition, the distribution timeliness optimization item for: in For the first The estimated transmission time of the data to be distributed is determined by the network link latency of the target routing application and the data size through a linear fitting model. Calculations show that For link transmission coefficient, For the first The amount of data in each data item For link base latency; For the first The timeliness of the data to be distributed is a valid decision variable. express The data remains within its valid timeframe after transmission. express The data will expire after transmission.
[0037] Meanwhile, the distribution timeliness optimization item for: in For the first The resource consumption for distributing each piece of data to be distributed is quantified by the actual bandwidth consumption for data transmission. This represents the basic resource usage of the Kafka data bus, and the minimum bandwidth usage required to ensure the basic operation of the system. This represents the maximum available distribution resources of the Kafka data bus, quantified by the maximum available bandwidth of the bus.
[0038] Overall, the model needs to simultaneously satisfy the preset constraints, which are as follows: in This is the actual waiting time after the data to be distributed enters the Kafka data bus queue. This represents the maximum effective transmission time threshold for this data. For the first Estimated transmission time for the data to be distributed; This is the resource elasticity coefficient, which is a preset threshold for limiting the over-allocation of bus resources based on the actual resource redundancy of the sand and gravel system. For the first Distribution scheduling decision variables for data to be distributed; For the first Resource consumption for distributing data to be distributed; This refers to the basic resource usage of the Kafka data bus. This represents the maximum available distribution resources on the Kafka data bus. and After normalization After normalization of the difference ratio, the range of values for all three terms is [missing information]. , This is the sequence identifier for the data to be distributed, with a value of [value]. arrive Positive integers.
[0039] After the optimal scheduling order for data distribution is obtained by solving the model, the data distribution work is carried out according to this order.
[0040] (5) Targeted data distribution Considering that the standardized sand and gravel system's multi-source data needs to match the requirements of each business module to achieve efficient flow and sharing, avoid problems such as misaligned data distribution and invalid transmission, and ensure the coordination of task scheduling and data distribution, data distribution work needs to be carried out according to the preset optimal scheduling order to ensure that data can be delivered to the corresponding business applications as needed.
[0041] The real-time data bus, built on the Kafka distributed messaging system, has received and stored standardized data after pre-processing. This data carries associated labeling information such as device ownership, production process, and collection task ID, providing a clear basis for distribution. During data distribution, the data is first read from the data queue according to optimal scheduling, ensuring that high-priority and time-sensitive data are processed first.
[0042] After reading the data, based on the associated annotation information of each data entry, the corresponding business distribution route is accurately queried from the preset routing database to determine the target application for which the data needs to be distributed. These target applications cover the core business modules of the sand and gravel system, such as production process management, equipment operation management, visual monitoring, and platform basic management, thus achieving the matching of data with business needs.
[0043] After retrieving the corresponding distribution route, the first step is to check the liveness status of the target application, which is a crucial step in ensuring effective data transmission. If the target application is detected to be alive, a stable data transmission connection is immediately established with the application's consumers, and the data is distributed to the corresponding business modules according to the optimal scheduling order to meet the business data needs of each module. If the target business application is detected to be not alive, promptly check whether the data to be distributed has expired. If the data has not exceeded the maximum effective transmission time threshold and is still within the valid range, put it back to the end of the data queue to wait for the next distribution cycle, so as to avoid loss of valid data. If data has expired and can no longer meet the timeliness requirements of business applications, it is discarded directly to reduce the consumption of invalid data on system resources. Through this series of standardized processes, the on-demand flow and efficient sharing of multi-source data in the sand and gravel system are realized, achieving collaborative optimization of task scheduling and data distribution, and providing data support for the stable operation of various business modules.
[0044] Example 2 Please refer to Figure 6 This embodiment 2 provides a multi-source data integration system for sand and gravel systems based on task scheduling and distribution, including: The database and parsing module configuration unit is used to build a configuration database for multi-source data in the sand and gravel system, configure the collection frequency, collection priority and data transmission protocol of each data source; according to the different data transmission protocols covered in the sand and gravel system, write and build the corresponding multi-protocol data parsing code module to form a parsing module library that can be called on demand; The distributed scheduling data acquisition unit is used to start the distributed task scheduling module based on xxl-job. The scheduling module queries the acquisition frequency of each data source from the configuration database and captures multi-source raw data of the sand and gravel system at regular intervals according to the acquisition frequency. The heterogeneous data standardization and integration unit is used to transmit the parsed data to the data processing module, clean, transform and standardize the collected raw data, and add associated labeling information such as equipment ownership, production process and collection task ID to the standardized data to form standardized multi-source data of the sand and gravel system and complete the unified integration of heterogeneous data. The planning model distribution and sorting unit is used to push standardized multi-source data from the sand and gravel system to a distributed real-time data bus built on Kafka. The data enters the queue of the data bus and waits for distribution. Based on a multi-objective linear programming model that combines task scheduling priority and data distribution timeliness, the objective function is constructed and the optimal scheduling and sorting of data distribution is obtained by combining the priority coefficients corresponding to the collection task IDs of each data to be distributed, the business response weight coefficients of the target routing application, and the maximum effective transmission timeliness threshold of the data. The targeted data distribution unit is used to read the data to be distributed from the data queue according to the optimal scheduling sorting, query the corresponding business distribution route from the preset routing database according to the association label information of the data to be distributed, detect the liveness status of the target routing application, and then distribute the data to the target routing application according to the optimal scheduling sorting, thereby realizing the coordinated optimization of task scheduling and data distribution.
[0045] Example 3 This embodiment 3 also provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it can implement any step of a multi-source data integration method for a sand and gravel system based on task scheduling and distribution.
[0046] The computer-readable storage medium may include various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0047] For a description of the computer-readable storage medium provided in this application, please refer to the above method embodiments; further details will not be repeated here.
[0048] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for multi-source data integration of a sandstone system based on task scheduling and distribution, characterized in that, include: S1. Construct a configuration database for multi-source data of the sand and gravel system, and configure the acquisition frequency, acquisition priority and data transmission protocol for each data source; Based on the different data transmission protocols covered in the sand and gravel system, corresponding multi-protocol data parsing code modules were written and built to form a parsing module library that can be called on demand; S2. Start the distributed task scheduling module based on xxl-job. The scheduling module queries the collection frequency of each data source from the configuration database and captures multi-source raw data of the sand and gravel system at regular intervals according to the collection frequency. S3. The parsed data is transmitted to the data processing module, which cleans, transforms and standardizes the collected raw data, and adds associated labeling information such as equipment ownership, production process and collection task ID to the standardized data to form standardized multi-source data of the sand and gravel system and complete the unified integration of heterogeneous data. S4. Standardized multi-source data from the sand and gravel system is pushed to a distributed real-time data bus built on Kafka. The data enters the queue of the data bus and waits for distribution. Based on a multi-objective linear programming model of task scheduling priority and data distribution timeliness, combined with the priority coefficient corresponding to the collection task ID of each data to be distributed, the business response weight coefficient of the target routing application, and the maximum effective data transmission timeliness threshold, an objective function is constructed and the optimal scheduling order for data distribution is obtained by solving it. S5. Read the data to be distributed from the data queue according to the optimal scheduling sorting. Based on the association labeling information of the data to be distributed, query the corresponding business distribution route from the preset routing database. After checking the liveness status of the target routing application, distribute the data to the target routing application according to the optimal scheduling sorting, thereby realizing the coordinated optimization of task scheduling and data distribution.
2. The method for multi-source data integration of a sand and gravel system based on task scheduling and distribution according to claim 1, characterized in that, The data source in S1 includes equipment data and software data of the sand and gravel system. The equipment data covers the operation data of PLC control system, smart meter, smart weighbridge, environmental monitoring equipment, and video surveillance equipment. The software data covers database change logs, application operation data, and production business data.
3. The method for multi-source data integration of a sand and gravel system based on task scheduling and distribution according to claim 1, characterized in that, In S2, the scheduling module queries the collection priority of each data source from the configuration database, sorts the processing queue of multi-source raw data according to the priority, queries the transmission protocol matching each data source, and calls the corresponding data parsing module from the parsing module library according to the transmission protocol to perform protocol parsing on the sorted multi-source raw data and output the parsed data.
4. The method for multi-source data integration of a sand and gravel system based on task scheduling and distribution according to claim 1, characterized in that, The multi-objective linear programming model in S4 is specifically as follows: A multi-objective programming method with time-effect decay and resource elasticity constraints is constructed. It integrates the mathematical characteristics of task priority dynamic decay, distribution time-effect actual benefits, and resource occupation elasticity threshold in the actual production scenario of sand and gravel system. The multi-dimensional objective dimension is unified through normalization processing, and multi-objective collaborative optimization is completed by relying on the mathematical coupling between objectives and the constraint boundary limit. The overall objective function of the model is: This includes dynamic priority optimization terms. Distribution timeliness optimization items Resource usage optimization items The three core components correspond to the dynamic matching requirements of task scheduling priorities, the timeliness and effectiveness requirements of data distribution, and the rational utilization requirements of system bus resources, respectively. The model simultaneously satisfies the preset constraints, which are as follows: in This is the actual waiting time after the data to be distributed enters the Kafka data bus queue. This represents the maximum effective transmission time threshold for this data. For the first Estimated transmission time for the data to be distributed; This is the resource elasticity coefficient, which is a preset threshold for limiting the over-allocation of bus resources based on the actual resource redundancy of the sand and gravel system. For the first Distribution scheduling decision variables for data to be distributed; For the first Resource consumption for distributing data to be distributed; This refers to the basic resource usage of the Kafka data bus. This represents the maximum available distribution resources on the Kafka data bus. and After normalization After normalization of the difference ratio, the range of values for all three terms is [missing information]. , This is the sequence identifier for the data to be distributed, with a value of [value]. arrive Positive integers.
5. The method for multi-source data integration of a sand and gravel system based on task scheduling and distribution according to claim 4, characterized in that, The dynamic priority optimization item for: in For the first The basic priority coefficient for each piece of data to be distributed corresponds to the collection task ID, and its value is assigned by the collection priority in the database configured by the claims. , Assign the highest priority value to the data collection task; For the first The business response weight coefficient of the target routing application for the data to be distributed is assigned differently based on the importance of the core production business and auxiliary management business of the sand and gravel system; This is a priority timeliness decay factor. This factor decays exponentially as the proportion of data waiting time to the maximum effective timeliness threshold increases, thereby realizing dynamic adjustment of task scheduling priority. For the first Distribution scheduling decision variables for data to be distributed. Indicates the execution of dispatch scheduling. This indicates that the action will not be taken for the time being. This represents the total number of data items to be distributed in the Kafka data bus queue.
6. The method for multi-source data integration of a sand and gravel system based on task scheduling and distribution according to claim 4, characterized in that, The distribution timeliness optimization item for: in For the first The estimated transmission time of the data to be distributed is determined by the network link latency of the target routing application and the data size through a linear fitting model. Calculations show that For link transmission coefficient, For the first The amount of data in each data item For link base latency; For the first The timeliness of the data to be distributed is a valid decision variable. express The data remains within its valid timeframe after transmission. express The data will expire after transmission.
7. The method for multi-source data integration of a sand and gravel system based on task scheduling and distribution according to claim 4, characterized in that, The distribution timeliness optimization item for: in For the first The resource consumption for distributing each piece of data to be distributed is quantified by the actual bandwidth consumption for data transmission. This represents the basic resource usage of the Kafka data bus, and the minimum bandwidth usage required to ensure the basic operation of the system. This represents the maximum available distribution resources of the Kafka data bus, quantified by the maximum available bandwidth of the bus.
8. The method for multi-source data integration of a sand and gravel system based on task scheduling and distribution according to claim 1, characterized in that, The state detection process for the corresponding service distribution route in S5 is as follows: If the target business application is alive, establish a data transmission connection with the target business application and distribute the data to be distributed to the corresponding business modules such as production process management, equipment operation management, visual monitoring, and platform basic management. If the target business application is not alive, check whether the data to be distributed has expired. If it has not expired, put the data back to the end of the data queue to wait for redistribution. If it has expired, discard the data directly to complete the on-demand flow of multi-source data in the sand and gravel system.
9. A multi-source data integration system for sand and gravel systems based on task scheduling and distribution, characterized in that, include: The database and parsing module configuration unit is used to build a configuration database for multi-source data in the sand and gravel system, configure the collection frequency, collection priority and data transmission protocol of each data source; according to the different data transmission protocols covered in the sand and gravel system, write and build the corresponding multi-protocol data parsing code module to form a parsing module library that can be called on demand; The distributed scheduling data acquisition unit is used to start the distributed task scheduling module based on xxl-job. The scheduling module queries the acquisition frequency of each data source from the configuration database and captures multi-source raw data of the sand and gravel system at regular intervals according to the acquisition frequency. The heterogeneous data standardization and integration unit is used to transmit the parsed data to the data processing module, clean, transform and standardize the collected raw data, and add associated labeling information such as equipment ownership, production process and collection task ID to the standardized data to form standardized multi-source data of the sand and gravel system and complete the unified integration of heterogeneous data. The planning model distribution and sorting unit is used to push standardized multi-source data from the sand and gravel system to a distributed real-time data bus built on Kafka. The data enters the queue of the data bus and waits for distribution. Based on a multi-objective linear programming model that combines task scheduling priority and data distribution timeliness, the objective function is constructed and the optimal scheduling and sorting of data distribution is obtained by combining the priority coefficients corresponding to the collection task IDs of each data to be distributed, the business response weight coefficients of the target routing application, and the maximum effective transmission timeliness threshold of the data. The targeted data distribution unit is used to read the data to be distributed from the data queue according to the optimal scheduling sorting, query the corresponding business distribution route from the preset routing database according to the association label information of the data to be distributed, detect the liveness status of the target routing application, and then distribute the data to the target routing application according to the optimal scheduling sorting, thereby realizing the coordinated optimization of task scheduling and data distribution.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, The computer program is executed by a processor as described in any one of claims 1-8: a method for integrating multi-source data of a sand and gravel system based on task scheduling and distribution.