A method for data transmission analysis and processing in a construction enterprise information system
By monitoring and optimizing the data transmission tasks of construction companies' information systems, the problem of resource shortages caused by data anomalies was solved, real-time monitoring and efficient processing of data transmission were achieved, and the reliability and responsiveness of data transmission were improved.
Patent Information
- Application Number
- CN202511247519.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-03
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2045-09-03
AI Technical Summary
In existing construction enterprise information systems, data transmission efficiency is low, and frequent data anomalies lead to system resource shortages, affecting the timeliness and reliability of data transmission. There is a lack of real-time monitoring and processing of data anomalies and optimization of resource allocation related to delays.
By monitoring the total number of sent tasks, the number of failures, and the locking period, and combining this with system utilization, steady-state probability, average queue length, and average dwell time, multi-objective optimization is performed to generate a sending task queue. This enables real-time monitoring and location of data anomalies, optimizes the adjustment ratio of the task queue, and prevents duplicate data uploads and delays.
It enables real-time monitoring of data anomalies, quickly identifies data conflicts, improves data transmission efficiency, reduces the average queue dwell time, and ensures the real-time responsiveness of the task queue and the accuracy of data upload.
Smart Images

Figure CN120811995B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data transmission technology, specifically a data transmission analysis and processing method for an information system of a construction enterprise. Background Technology
[0002] When uploading building data, batch processing and scheduled tasks are often used. However, this approach can easily lead to low data transmission efficiency when dealing with large volumes of data and frequent changes. Since data transmission tasks are concentrated at fixed times, it can strain system resources and affect the timeliness of data transmission. Furthermore, the time difference between calculating and sending data changes may cause some changed data to fail to be sent in a timely manner, affecting data integrity and the reliability of data transmission.
[0003] For example, Chinese Patent Publication No. CN117834307A discloses a data transmission protection method and system based on communication status, which relates to the field of communication data security technology. The method includes: acquiring information about several communication network nodes corresponding to the communication network; determining all communication lines of each communication network node; acquiring line risk data of the communication network node; acquiring the historical data throughput of each communication network node; analyzing and obtaining the predicted data throughput of the communication network node; comprehensively evaluating the data transmission security status of each communication network node; and adjusting the security protection strategy for the communication network node.
[0004] For example, Chinese Patent Publication No. CN119603126A discloses a multi-terminal serial data transmission interaction system and method based on a communication network. The system includes a data initiator, a serial data processing node module, a data receiver, a communication monitoring and management module, an emergency decision-making module, and a background monitoring terminal. In this invention, the data initiator initially encapsulates the data to be transmitted and sends it to the serial data processing node module. Each data processing node in the serial data processing node module is responsible for a specific data processing task. After processing by all data processing nodes, the data receiver receives and parses the data, improving the efficiency of data processing and transmission. Furthermore, the communication monitoring and management module monitors the entire data transmission process in real time to identify anomalies in the data transmission process.
[0005] Existing technologies quantify the risks of communication lines by specifying the test intervals of communication network nodes and describing the data status of communication through anomaly detection values during data transmission. However, existing technologies lack real-time data monitoring, ignore resource allocation issues related to parameters such as service rate and the number of receiving systems when data transmission failures occur, and cannot directly obtain the cause of data anomalies. This results in the uncertainty of the status of multiple receiving systems during data anomalies, leading to reduced response efficiency and reduced efficiency in locating transmission failures during data transmission analysis. Summary of the Invention
[0006] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is: a data transmission analysis and processing method for information systems of construction enterprises, including: S1, collecting the sending task information configured by the communication node, monitoring the data anomalies before and after the execution of the sending task, and recording the total number of sending tasks, the number of sending failures, and the lock period of sending failures when the current sending task is executed.
[0007] S2, analyze the number of sends and the number of failures for each batch of sending tasks, and combine the lock period updated when the sending task changes to determine the task execution status when the data changes.
[0008] S3, based on the task execution status and the last timestamp of the data update, applies multi-objective constraints to the current sending task, calculates the data increment of the sending task under the multi-objective constraints, sorts the sending tasks according to the data increment, and generates the task queue corresponding to the sending task.
[0009] S4, based on the task queue of the sending task, performs correlation analysis between the sending task and the receiving system, and calculates the adjustment ratio of each sending task during queuing transmission.
[0010] S5 optimizes batch division based on the adjustment ratio of the tasks sent in each batch, sets hierarchical response actions in response to the distribution of each batch, and adjusts the queue data volume of the tasks sent in each batch.
[0011] The beneficial effects of this invention are as follows: First, this invention quantifies data changes under data anomalies by monitoring the total number of transmissions, failures, and locking periods of transmission tasks, combined with system utilization, steady-state probability, average queue length, and average dwell time. Furthermore, it divides the current transmission task into multiple groups of transmission task information based on the processing latency, packet arrival rate, and service rate during execution. Then, by combining the corresponding transmission tags of the transmission tasks, it traces and expands the task execution status, achieving real-time monitoring and data location of uploaded data during data anomalies, quickly identifying data conflicts during data anomalies, and avoiding blind expansion when data anomalies fail.
[0012] Second, this invention optimizes subtasks by taking average queue length and dwell time as core objectives, combined with the joint solution of the number of receiving systems and service rate, and generates task queues by sorting according to data increment. This reduces the blocking phenomenon of data upload caused by sorting by a single indicator, thereby improving the efficiency of data upload in task queues under multi-objective processing and reducing the average dwell time of the queue.
[0013] Third, this invention generates an adjustment ratio by calculating the similarity between historical failure identifiers and the initial identifiers of the current task queue. If the similarity is greater than or equal to a threshold, the initial identifiers are updated with the least common subset of failure identifiers. This further quantifies the correlation between the current task queue and historical failure patterns, prevents repeated data upload failures, and improves the accuracy of data upload pattern recognition. Then, based on the adjustment ratio, sub-batch is divided, and the amount of queue data output in the current corresponding batch is adjusted according to the adjustment ratio of the sub-batch. This ensures the real-time responsiveness of task queue uploads, forming a full-link optimization of data anomaly monitoring → multi-target generation → failure pattern association → dynamic batch response, and completing the queuing adjustment of multiple data queues under data anomalies. Attached Figure Description
[0014] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0015] Figure 1 This is a flowchart illustrating a data transmission analysis and processing method for an information system in a construction company.
[0016] Figure 2 This is a flowchart illustrating step S1 of a data transmission analysis and processing method for an information system of a construction enterprise.
[0017] Figure 3 This is a flowchart illustrating step S2 of a data transmission analysis and processing method for an information system in a construction enterprise.
[0018] Figure 4 This is a flowchart illustrating step S3 of a data transmission analysis and processing method for an information system in a construction enterprise.
[0019] Figure 5 This is a flowchart illustrating step S4 of a data transmission analysis and processing method for an information system in a construction enterprise.
[0020] Figure 6 This is a flowchart illustrating step S5 of a data transmission analysis and processing method for an information system in a construction enterprise. Detailed Implementation
[0021] The embodiments of the present invention are described in detail below. The embodiments described below are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention. Where specific techniques or conditions are not specified in the embodiments, they shall be performed in accordance with the techniques or conditions described in the literature in the art or in accordance with the product manual.
[0022] See Figure 1A data transmission analysis and processing method for an information system of a construction enterprise includes: S1, collecting the sending task information configured by the communication node, monitoring the data anomalies before and after the execution of the sending task, and recording the total number of sending tasks, the number of sending failures, and the lock period for sending failures when the current sending task is executed.
[0023] S2, analyze the number of sends and the number of failures for each batch of sending tasks, and combine the lock period updated when the sending task changes to determine the task execution status when the data changes.
[0024] S3, based on the task execution status and the last timestamp of the data update, applies multi-objective constraints to the current sending task, calculates the data increment of the sending task under the multi-objective constraints, sorts the sending tasks according to the data increment, and generates the task queue corresponding to the sending task.
[0025] S4, based on the task queue of the sending task, performs correlation analysis between the sending task and the receiving system, and calculates the adjustment ratio of each sending task during queuing transmission.
[0026] S5 optimizes batch division based on the adjustment ratio of the tasks sent in each batch, sets hierarchical response actions in response to the distribution of each batch, and adjusts the queue data volume of the tasks sent in each batch.
[0027] The aforementioned communication nodes represent devices that upload data, such as computers and cameras. The sending task distributes various types of basic business data in the system, such as equipment usage and material consumption rates during construction, to multiple interface addresses according to their respective business functions. The sending task follows a pre-set sending task configuration table in the system, extracting the data type and receiving system for each type of basic business data. The data distributed to each interface is recorded as sending task information based on its transmission address. This sending task information also includes three parameters: the maximum number of data items sent per batch, the number of failed sending requests, and the lockout period after exceeding the failure limit. The lockout period is described according to the data fluctuation cycle to determine the relative situation of the sending task during fluctuations. This emphasizes the conflict between multiple data transmissions in scenarios with multiple data transmission fluctuations, as well as the portion of data transmission actually completed after abnormal transmission.
[0028] When allocating tasks, the sending task distributed by the current communication node is mainly determined based on the data type dictionary and the receiving system dictionary. The data type dictionary includes various descriptions of the construction progress, such as cost summary, schedule deviation, material consumption, and distribution of safety hazards. The receiving system dictionary describes various data storage and output systems, such as ERP system, material management system, progress monitoring platform, and BI analysis system.
[0029] The task information set at this time is used to verify issues such as slow data updates and data update conflicts caused by frequent data changes in multi-data distribution and uploading scenarios, in order to ensure the efficiency of data upload and reduce disturbances during data upload.
[0030] like Figure 2 As shown, the implementation of step S1 also includes: S11, at any time when the data transmission changes, check the transmission task configured by the current communication node, use the processing delay, packet arrival rate and service rate of the transmission task as judgment indicators, and conduct horizontal and vertical attribution analyses respectively with business analysis as the guide.
[0031] S12 uses the rate of change of the judgment index to describe the horizontal attribution, and records the system utilization and steady-state probability corresponding to the current sending task as the calculation result of the horizontal attribution.
[0032] S13. Perform vertical attribution based on the business events corresponding to the current sending task, and record the average queue length and average dwell time of the sending task during data anomalies as the calculation result of the vertical attribution.
[0033] S14 synchronizes the calculation results of horizontal and vertical attribution to the sending task information, quantifying the processing of each sending task.
[0034] Preferably, the packet arrival rate is calculated as the proportion of successfully sent packets to indicate the portion of logs and on-site data that are normally transmitted when uploaded in data packet form. The failure count is calculated as the number of packets that failed to be sent. The lock period is calculated as the time period during which the corresponding system interface is locked when there are consecutive transmission failures. The data locked by the corresponding interface for consecutive transmission failures can be directly queried from the database. The lock period is also directly queried from the database. These data are pre-set. The statistics of the lock period are used to further quantify the portion of data that did not complete the transmission normally, so as to indicate the abnormal situation in the configured sending task under the current multi-task transmission scenario.
[0035] The aforementioned service rate represents the system's service capacity per unit time for receiving and transmitting data. This service rate can be directly obtained from the service rate configured in the M / M / C model, while processing latency emphasizes the time delay when data transmission is completed.
[0036] Preferably, when setting the judgment indicators, the processing method further includes: obtaining the number of current sending tasks, and after obtaining the receiving system corresponding to the sending task, verifying the processing delay, packet arrival rate and service rate of the sending task during execution based on the number of sending tasks corresponding to each receiving system.
[0037] The locking mechanism handles latency, packet arrival rate, and service rate during lockout cycles that fail. It treats the sending tasks within each lockout cycle as the primary data to be processed and divides the current sending task into multiple groups of sending task information based on the number of receiving systems.
[0038] At this point, statistics will be compiled on the systems receiving and sending tasks to illustrate the performance of each system in receiving data, and the task information sent under different numbers of receiving systems will be statistically analyzed to further determine the current data transmission situation.
[0039] During the above processing, the data analyzed will be added to the data set corresponding to the task information, or the corresponding data will be stored in the task execution status table to illustrate the execution process of the current task under the data change scenario.
[0040] The aforementioned data anomalies generally indicate that during the data transmission of the current sending task, there are abnormal situations such as data loss or large data fluctuations in the output and storage of the corresponding system. This emphasizes a scenario where there are certain problems in the data upload process. It is necessary to analyze the sending task part of the transmission to improve the real-time responsiveness of the corresponding building information system and prevent problems such as large data upload delays and long data queues in a single receiving system, which can cause serious queuing issues during data upload.
[0041] In one embodiment of the present invention, since the basic business data in the construction industry changes frequently and the amount of data uploaded is large, each set sending task needs to check the execution status under the data change and write the relevant execution status into the task execution status to record the current sending task allocation and execution status.
[0042] When determining the current task execution status, the data portion of the task execution status that is attributed horizontally and vertically should be used as the main basis for judgment. This portion will directly reflect the time and data queue length during the data transmission process.
[0043] The task execution status will be categorized into several states at this point: First, normal execution, indicating that the task completed normally outside the locking period, with both the number of transmissions and failures meeting expectations; this indicates that the number of transmissions reached the expected batch size, the number of failures was below the threshold, and there was no data anomaly or locking period interference with data transmission. Second, waiting for locking, indicating that the task triggered the locking mechanism due to data anomaly, entered the waiting queue, and paused execution; this indicates that data anomaly occurred, such as data updates or modifications, and the locking period took effect, the transmission task was paused, and the task changed from in progress to waiting for locking. Third, execution failure, indicating situations unrelated to locking, such as technical failures, insufficient resources, or configuration errors, unrelated to the locking period; this indicates that the number of failures exceeded the threshold, displayed error types such as network timeout or insufficient permissions, or the locking period was not triggered or had ended. Fourth, execution failure due to locking, indicating a failure related to the locking period, representing a task interruption or timeout failure due to the locking period; this indicates that the task could not acquire resources within the locking period, the locking time exceeded the maximum allowed waiting time for the task, or displayed a lock timeout or similar error. A partial success rate indicates that some batches of the task succeeded, while others failed due to locking or other reasons. This suggests that the transmission quantity was partially completed, but the failures were concentrated in specific batches, and these failed batches were related to data anomalies or locking periods. The relative task status represented by these data points is the main focus of the analysis at this stage. This information is then used to examine the task queue for subsequent data transmission to determine whether the current task is transmitting normally.
[0044] It should be noted that the threshold for the number of failures and the length of the lock period mentioned above are set according to the system receiving the current data, and these data are configured in the database in advance.
[0045] like Figure 3 As shown, the implementation of step S2 also includes: S21, performing source tracing for changes in the sending tasks under the current batch, and setting the execution flow of the sending tasks using the transmission tags of sending quantity, number of failures and locking period.
[0046] S22 extends the running status of the sending task by using the previous batch of data and the next batch of data in the current execution process as the pre-part and post-part respectively.
[0047] S23, update the running state with the extended periodic component, and use the updated running state as the task execution state when data changes.
[0048] The aforementioned transmission tag is used to generate a tag for each batch of sending tasks. This tag describes the data corresponding to the number of sends, the number of failures, and the lock period, as well as the current progress of the sending task, such as sending or resending. The execution flow is a process-oriented data composed of the contents of multiple transmission tags, which describes the current status of the sending task.
[0049] Preferably, the aforementioned pre- and post-parts are geared towards describing the data format sent by adjacent batches of transmission tasks after selecting the same receiving system. This determines the adjacent data portions of transmission tasks targeting the same goal, and considers the execution flow of the current batch of transmission tasks under the previous and subsequent batches as the running state of the current transmission task. The previous batch is considered as the portion that has been transmitted after the last data transmission was completed. As for the next batch, it indicates whether there is any remaining untransmitted data after the current data transmission. If there is, the remaining data is considered as the next batch; otherwise, the data with tags such as "completed" or "failed" for the current data transmission is considered as the data of the next batch, thus expanding the running state of the current transmission task.
[0050] The periodic component of the running status during expansion indicates the start time of the task, the end time of the task, the time when the lock period is released, and the time when the lock period begins, representing the current execution time point. It then records the current task execution status after each corresponding time and outputs this as the task execution status. The periodic component is set to ensure that the current task execution status directly corresponds to one of the five preset task execution statuses: normal execution, waiting for lock, execution failure due to non-locking reasons, execution failure due to locking reasons, and partial success, thus indicating the relative completion status of the task.
[0051] In one embodiment of the present invention, such as Figure 4 As shown, the implementation of step S3 includes: S31, for each subtask of the sending task accessing the network, the task execution status of each subtask when it is updated is counted, and the constraints of each subtask when it is queuing for uploading are extracted; the constraints described here represent that when the subtask is sent to the receiving system, the number of receiving systems, the service rate of the receiving system, the average queue length of the subtask and the average dwell time are regarded as the four parameters of the extracted constraints.
[0052] S32, multi-objective solution of the constraints of the subtask, takes the solution values of average queue length and average dwell time as the first objective, the solution value of service rate when the number of receiving systems is fixed as the second objective, the solution value of number of receiving systems when service rate is fixed as the third objective, and the joint solution value of number of receiving systems and service rate as the fourth objective, to achieve multi-objective solution of the subtask; at this time, the multi-objective solution is to minimize the weighted value of average queue length and average dwell time allocated to the sending task subtask, so as to prevent the system from abnormal due to frequent input under multiple data changes.
[0053] S33, with the first objective as the core, sequentially calculates the solutions for the second and third objectives based on the solution set of the first objective. The solutions for the third and second objectives are then used as inputs for the fourth objective, and the solution for the fourth objective is used as the output for the multi-objective solution. A Pareto front approach is employed here. First, the data contained in the solution set of the first objective is used as the initial Pareto front. The data from the initial Pareto front is then used to calculate the second and third objectives. At this point, both the third and second objectives represent a set of data. Then, the corresponding datasets are used as a method to calculate the fourth objective, thus completing the construction of the final Pareto front. Finally, the number of receiving systems and the service rate jointly optimized in the final Pareto front are selected according to system requirements.
[0054] S34. Based on the output value of the multi-objective solution, calculate the data increment of each subtask under the corresponding constraints when the multi-objective solution is completed, and sort the subtasks into a task queue according to the objectives when solving the subtasks, in descending order of data increment.
[0055] Data increment refers to the amount of data added or modified by each subtask relative to the baseline state, such as the state of the last transmission or processing, under the constraints of multi-objective optimization. It is used to explain the amount of data adjusted by each subtask after satisfying the constraints when adjusting the sending task; that is, it explains the current adjusted data amount by comparing the data block satisfying the multi-objective solution with the initial data block of the subtask. This data increment directly indicates the size of the data packet that needs to be uploaded and adjusted in the task execution states corresponding to waiting for locking, execution failure due to non-locking reasons, execution failure due to locking reasons, and partial success. It allows for the transmission of the incremental portion as needed, saving upload resources and preventing data upload anomalies. For normally executing tasks, it explains the difference between the previously uploaded data packet and the data packet being uploaded this time, preventing duplicate data uploads.
[0056] In one embodiment of the present invention, after normalizing the data transmitted by each communication node, the statistics of the uploaded data of each communication node are counted in a time series manner according to a fixed sliding window size; and the sending task configured for each communication node is identified by the receiving system of the sending task, and the corresponding system's processing delay, packet arrival rate and service rate are identified, wherein the service rate is expressed as the reciprocal of the processing delay.
[0057] The M / M / C model used is a queuing model for identifying multiple data transmissions. It transmits data in packets and identifies the number of packets sent to the receiving system. When the number of packets on each receiving system exceeds the size that can be processed in a single transmission, queuing occurs. Queuing is handled in this case to address the adjustment of the system's receiving data queue when construction site information related to the current building is sent to the corresponding system from multiple sources in a big data processing scenario. This reduces the time that data stays in the message channel and improves the response efficiency for different data.
[0058] The implementation method for obtaining the sending task using the receiving system can be as follows.
[0059] For example, system utilization can be expressed as: ;in, This represents the system utilization rate, which indicates the utilization rate of the current sending task configured under the corresponding receiving system. These data are used to measure the specific situation of the current data anomaly, and facilitate subsequent judgment of the stability of the current sending task execution by combining horizontal and vertical attribution. Indicates the arrival rate of the group; Indicates service rate; This indicates the number of receiving systems; at this time, the system utilization rate is less than 1, which indicates the status of the current sending task after each sending task is transmitted to the corresponding receiving system under data anomalies.
[0060] As for the steady-state probability This is represented as: ;in, Number of groups This represents the normalization constant, used to ensure that the sum of all steady-state probabilities is 1. This value can be obtained by solving for the current steady-state probability. The steady-state probability obtained here describes the stable state of the sending task after execution. It describes the probability that the data uploaded by the sending task will be in a state of waiting for processing in n data groups during long-term operation under the current data anomaly scenario. It reflects the buffering capacity when data anomalies occur, and prevents data congestion and increased data latency caused by data anomalies.
[0061] For average queue length This is represented as: When calculating the average queue length, we emphasize the number of packets waiting to be transmitted under each sending task during data transmission. This measures the waiting pressure of multiple data packets in the current sending task and prevents the current system's data processing efficiency and receiving capacity from decreasing due to sudden task changes or anomalies.
[0062] As for the average length of stay This is represented as: The average dwell time indicates the total time it takes for the data to be processed during the execution of the sending task. This time further illustrates the time consumed by the sending task after transmitting the corresponding data under frequent data movement, preventing problems such as warning delays caused by excessive time when the current sending task is processing data.
[0063] Then, when determining the execution of the current sending task, it is necessary to minimize the average queue length and average dwell time in order to analyze the relative situation of different sending tasks under the scenario of frequent data changes during the current data transmission process.
[0064] For example, objective function Represented as: ;in, , These are represented as weight values, which are described based on the average queue length and average dwell time. For example, the weight is set at this time as the ratio of the average queue length and average dwell time of the data uploaded by the current sending task to the sum of the average queue length and average dwell time of the total data. This represents the average dwell time of the i-th object. i is an index variable used to describe the total number of objects K that the current objective function needs to optimize. It refers to each sending task or the part of the sending task that needs to upload data. The average queue length of the i-th object represents the average length of the data queue uploaded in each sending task, illustrating the transmission portion after the current sending task is executed; The total number of objects represents the number of sending tasks that need to be processed. The value of i ranges from 1 to K. The goal is to minimize the data increment of the current sending task while minimizing the objective function, thus completing the transmission of all data with minimal data fluctuation. The data corresponding to the average queue length and average dwell time obtained under this minimized solution will be used as the solution set for the first objective.
[0065] Further optimization can be performed on the service rate and the number of receiving systems. While minimizing the objective function, the set of optimal solutions for multiple sending tasks with a fixed service rate and the number of receiving systems is obtained, given the total number of objects. For example, when fixing the number of receiving systems, to find the optimal solution for the service rate, the objective function is differentiated with respect to the service rate, and the derivative is set to zero to determine the optimal solution for the number of receiving systems under the corresponding sending task. When fixing the service rate, to find the optimal solution for the number of receiving systems, since the number of receiving systems is monotonic within a certain range, a binary search is used to solve for the range of the number of receiving systems. The midpoint of the range of receiving system values is selected to calculate the objective function. If the calculated objective function value is smaller than that obtained with the same number of receiving systems in historical data, the calculated value of the number of receiving systems is continuously updated until the objective function is calculated to its minimum. The number of receiving systems at this minimum is considered the optimal solution output. At this point, by fixing the service rate and the number of receiving systems respectively, the solution sets of the second and third objectives under the current calculation can be obtained in sequence. The solution sets will represent multiple sets of data when these two data are at the optimal solution, which will facilitate subsequent joint optimization and solution.
[0066] Next, joint optimization is needed for the service rate and the number of receiving systems. This involves multiplying the current service rate (considered the optimal solution) and the number of receiving systems. If the sum of the products is less than a preset value, the current service rate and the number of receiving systems are considered as output data. Based on the current number of receiving systems and the corresponding optimized average queue length, the current sending task is divided to determine the data to be sent to the corresponding system.
[0067] During joint optimization, the Lagrange multiplier method is used to construct the Lagrange function and calculate its partial derivatives. All partial derivatives are zero, which is used to obtain the current output service rate and the number of receiving systems.
[0068] For example, in joint optimization, it is represented as: ;in, This represents the number of receiving systems for the i-th object. This represents the service rate of the i-th object; This represents a global constraint value, indicating a preset value that must be less than the current joint optimization. The initial value of this constraint value represents the sum of the product of the number of receiving systems and the service rate before optimization. It is required that the sum of the product of the number of receiving systems and the service rate be less than this constraint value during each optimization. A Lagrange function is introduced for joint partial derivatives to optimize the current service rate and the number of receiving systems.
[0069] The Lagrangian function can be expressed as: ;in, This represents the new objective function formed after introducing Lagrange multipliers. This function takes the number of receiving systems and service rate of the i-th object as input during joint optimization, and adds the corresponding Lagrange multipliers to obtain the constructed Lagrange function. The objective function value is obtained by inputting the number of receiving systems and the service rate of the i-th object. The Lagrange multipliers are used to represent the objective function. Then, the partial derivative of the new objective function is set to zero. The partial derivatives refer to the number of receiving systems, service rate, and Lagrange multipliers specified in parentheses within the new objective function, thus yielding the current number of receiving systems and the associated service rate. Symbolic computation libraries like Sympy can then be used to extrapolate the current data, and the resulting data is considered the solution set for the fourth objective.
[0070] Based on data such as the number of receiving systems, service rate, and packet arrival rate under the current total number of sending tasks, the sending tasks are sent when the current system utilization, average queue length, and average dwell time are all at their optimal settings. This completes the basic processing of the sending tasks. These data can be used as the basis for the current configuration of the sending tasks, or as the data configuration of the sending tasks after analysis in the current task scenario. This data is used for subsequent auxiliary analysis of the sending task queue, as well as correlation analysis and other related content.
[0071] In the above implementation of multi-objective constraints, the processing can also be carried out according to the priority corresponding to the data increment. For example, step S34 also includes: in response to the output value after solving the multi-objective problem, the normalized value of the product of the service rate of the current subtask and the number of receiving systems is used as the priority to mark each subtask; and the priority of each subtask relative to the task queue is determined.
[0072] If the priority of each subtask is different under multiple data increment calculations, the average priority of each subtask is taken as the priority of the current subtask, and the task queue is sorted according to the priority of the subtask.
[0073] The normalized value described here is obtained by summing the products of the current subtask's service rate and the number of receiving systems in the task sequence, and then multiplying them by the products of all sorted subtasks in the task sequence to obtain the priority of each subtask.
[0074] When sorting the task queue by data increment from largest to smallest, the system tends to prioritize tasks with larger data increments to quickly reduce system load. However, this may overlook the actual system load capacity, leading to issues related to the number of receiving systems and the content corresponding to the service rate, causing delays in some resources. Subsequently, the sorted task queues are prioritized using a normalized value that is the product of the service rate and the number of receiving systems to balance the system load.
[0075] At this point, based on the multiple sets of values obtained from the multi-objective solution, the subtasks need to be sorted using data increments. If there are identical data increments, they should be sorted according to their corresponding priorities to complete the relative adjustment of the task queue.
[0076] In one embodiment of the present invention, such as Figure 5 As shown, the implementation of step S4 includes: S41, generating the initial identifier of the current task queue based on the execution status of multiple tasks corresponding to the current task queue.
[0077] S42, associate the number of sending failures and the locking period during the execution of the sending task in the historical data with the receiving system corresponding to each sending task, and obtain the failure flag on the receiving system.
[0078] After completing multi-objective optimization based on the number of receiving systems, service rate, and corresponding processing latency, it's necessary to verify the relative situation of data under the corresponding failure count and locking period. This allows for further analysis of upload latency and transmission errors during data upload queuing. Data experiencing consecutive failures leading to database locking, as well as other data failures, is represented using failure identifiers. These failure identifiers specify the specific data, time, and subtasks involved in the failed transmission after the transmission task was configured to the communication node. The initial identifier includes the specific data after solving the multi-objective constraints, distinguishing the portion sent to each receiving system and indicating the task execution state in which the initial identifier was set.
[0079] S43, read the number of failure flags on the receiving system, write the failure flags into the current task queue, calculate the similarity between the failure flags and the corresponding initial flags, and regard the calculated average similarity as the output adjustment ratio.
[0080] When writing a failure flag, the failure flag is written to the corresponding subtask in the task queue to indicate that there was a failure during the processing of the sending task. At this time, it is necessary to further analyze the similarity between the data corresponding to the failure flag and the data corresponding to the current optimized initial flag to evaluate the data situation under the previous transmission failure. At this time, the similarity can be calculated using the Pearson correlation coefficient method, with the number of receiving systems, the service rate of the receiving systems, the average queue length of the subtask, and the average dwell time as the dimensions for calculating the similarity. The average value of these four values is taken as the adjustment ratio of the output to obtain the similarity between the failure flag and the corresponding flag under the optimized processing.
[0081] The current output adjustment ratio is used to indicate whether the values of the number of receiving systems, the service rate of the receiving systems, the average queue length of the subtasks, and the average dwell time after the current optimization analysis match the historical failure patterns, so as to avoid repeated failures. At the same time, the adjustment ratio reflects the failure risk and indicates which subtasks in the task queue are at risk of failure.
[0082] Preferably, step S43 further includes: comparing each initial identifier in the task queue; when the similarity between the initial identifier and the failure identifier is greater than or equal to a preset threshold, it is considered that the current initial identifier has a corresponding failure identifier, and the least common subset of the failure identifiers is added to the corresponding initial identifier, thus updating the current task queue. It should be noted that the least common subset of failure identifiers represents the intersection of the initial identifier and the failure identifier in the data dimension, and this intersection is the smallest indivisible unit. Using only the least common subset is crucial for accurately locating the key parameter combinations that led to historical failures, avoiding a full update of the initial identifiers. This ensures that the output initial identifiers can identify which parts of the data in the currently transmitted task queue are prone to problems during upload.
[0083] If the similarity between the current initial identifier and the failure identifier is less than a preset threshold, it is assumed that the current initial identifier does not have a corresponding failure identifier, and the data corresponding to the current initial identifier is directly output to obtain the output adjustment ratio. When the initial identifier in the task queue does not correspond to a failure identifier, it indicates that the current output task queue has a low risk, and it can be directly output based on the receiving system configured for the current task queue. If new failure cases occur later, the similarity is recalculated to identify the similarity between the current task queue transmission and historical failure patterns.
[0084] Preferably, the preset threshold set above is set based on the average cosine similarity calculated from the failure identifiers in historical data to illustrate the data similarity in the failure mode. If the similarity between the initial identifier in the current task queue and the failure identifier in the failure mode reaches the corresponding preset threshold, it indicates that the current task queue has a similar pattern to historical data in terms of data transmission failure. At this time, the output adjustment ratio will be too large, emphasizing that the current task queue needs too much adjustment during data transmission.
[0085] It should be noted that when outputting the adjustment ratio under the corresponding task queue, the task queue will be formed according to the sorting method of the subtasks in the task. In other words, the task queue represents a transmission queue form of data formed by each sending task. Then, the adjustment ratio value calculated in the task queue will be synchronized to the corresponding sending task.
[0086] In one embodiment of the present invention, such as Figure 6As shown, the implementation of step S5 includes: S51, based on the adjustment ratio of the sending tasks under each batch, clustering the current batch according to the adjustment ratio, and dividing the current batch into multiple sub-batches.
[0087] S52, using the value range of the adjustment ratio under each sub-batch, sets the hierarchical response action for each sub-batch.
[0088] S53: When a hierarchical response action is triggered in any sub-batch, the queue length of each sub-batch is adjusted according to the hierarchical response action, and the queue data volume of the task to be sent in each batch is combined and output according to the data volume contained in the queue length.
[0089] The tiered response action, based on the adjustment ratio within the historical data range, divides data combinations exhibiting significant upload failures and successful uploads into two confidence intervals. The standard deviation and mean of each confidence interval are calculated, and a value range for the combination is set in the form of mean ± three standard deviations. Based on this value range, the tiered response action corresponding to the value range is extracted from the database. This includes increasing resource allocation, pausing queuing, and reducing data allocation, and the tiered response action is set for the corresponding sub-batch. For example, if the adjustment ratio of the sub-batch is less than the upper limit of the confidence interval for normal uploads, no adjustment is needed, and the original queue is maintained. If it is greater than the upper limit of the confidence interval for normal uploads but less than the lower limit of the confidence interval for upload failures, the resources allocated to the receiving system in the sub-batch are increased, such as increasing the number of CPU cores by 10-20% as the adjustment for the corresponding sub-batch resource allocation, without directly changing its queue length. When the value is between the lower and upper limits of the confidence interval for upload failures, the current sub-batch queue data volume is multiplied by (1 - adjustment ratio × reduction factor). The scaling factor is used to reduce the amount of data currently being uploaded, preventing the risk of high concurrency and upload failures. Its value range can be set from 0.2 to 1, and 0.5 can be selected to reduce the amount of data currently being uploaded. When there is data exceeding the upper limit of the confidence interval for upload failure, the data under the corresponding sub-batch is used as a queue for paused upload, the length of the corresponding sub-batch is set to 0, the upload of the corresponding sub-batch is paused, the queue data is retained, and manual intervention is triggered for investigation. After the investigation is completed, the current data is returned to the task queue, and the corresponding adjustment ratio calculation and output continue.
[0090] The implementation of step S53 also includes: verifying the queue data volume under the corresponding sub-batch, where the queue data volume is the total amount of data contained in the task queue under the corresponding sub-batch; within a preset waiting time, determining whether the number of sub-batches that trigger the same hierarchical response action is greater than 1, and if so, combining the sub-batches within the preset time period one by one to form an output batch.
[0091] If not, the adjustment ratio of each batch is used to combine the sub-batch into an output batch, and the queue data volume of the combined output batch is regarded as the queue data volume of the task to be sent in each batch.
[0092] When merging batches according to the preset waiting time, if there are more than one tiered response action within the waiting time, multiple batches that trigger the same tiered response action will be merged. For example, if two sub-batches both need to be paused, the corresponding batches will be merged. If the tiered response actions within the preset waiting time are different, they will be sorted and combined according to the adjustment ratio, and processed in ascending order of the adjustment ratio value.
[0093] When setting the preset waiting time, if the current task is a fast response, such as financial data updates and real-time video data updates, 30 seconds is set as the preset waiting time; when processing offline data, such as log analysis and offline reports, 10 minutes is set as the preset waiting time to determine the data to be processed in the corresponding time period.
[0094] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention, which are still covered within the protection scope of the present invention.
Claims
1. A method for data transmission analysis and processing in an information system of a construction enterprise, characterized in that, include: S1, collect the sending task information configured by the communication node, monitor the data changes before and after the sending task is executed, and record the total number of sending tasks, the number of sending failures, and the lock period for sending failures when the current sending task is executed. S2, analyze the number of sends and the number of failures for each batch of sending tasks, and combine the lock period updated when the sending task changes to determine the task execution status when the data changes; S3, based on the task execution status and the last timestamp of the data update, applies multi-objective constraints to the current sending task, calculates the data increment of the sending task under the multi-objective constraints, sorts the sending tasks according to the data increment, and generates the task queue corresponding to the sending task. S4, based on the task queue of the sending task, performs correlation analysis between the sending task and the receiving system, and calculates the adjustment ratio of each sending task during queuing transmission. S5 optimizes batch division based on the adjustment ratio of the sending tasks in each batch, sets hierarchical response actions in response to the distribution of each batch, and adjusts the queue data volume of the sending tasks in each batch. The implementation of step S1 further includes: S11, at any moment when the data transmission changes, checking the transmission tasks configured on the current communication node, using the processing latency, packet arrival rate, and service rate of the transmission tasks as judgment indicators, and conducting horizontal and vertical attribution analyses guided by business analysis; S12, describing horizontal attribution using the rate of change of the judgment indicators, recording the system utilization and steady-state probability corresponding to the current transmission task as the calculation result of horizontal attribution; S13, performing vertical attribution using the business events corresponding to the current transmission task, recording the average queue length and average dwell time of the transmission task during data changes as the calculation result of vertical attribution; S14, synchronizing the calculation results of horizontal and vertical attribution to the transmission task information to quantify the processing of each transmission task. When setting judgment indicators, the processing method also includes: obtaining the number of current sending tasks; after obtaining the receiving system corresponding to the sending task, verifying the processing latency, packet arrival rate and service rate of the sending task during execution based on the number of sending tasks corresponding to each receiving system; locking the processing latency, packet arrival rate and service rate under the locking period where the failure occurs; taking the sending tasks in each locking period as the data to be processed, and dividing the current sending task into multiple groups of sending task information based on the number of receiving systems; The implementation of step S3 includes: S31, for each subtask of the sending task accessing the network, the task execution status of each subtask when it is updated is counted, and the constraints of each subtask when it is queuing for uploading are extracted; S32, the constraints of the subtask are solved for multiple objectives, the solution values of average queue length and average dwell time are taken as the first objective, the solution value of service rate when the number of receiving systems is fixed is taken as the second objective, the solution value of number of receiving systems when service rate is fixed is taken as the third objective, and the joint solution value of number of receiving systems and service rate is taken as the fourth objective, so as to realize the multi-objective solution of the subtask; S33, taking the first objective as the core, according to the solution set of the first objective, the solution values of the second objective and the third objective are obtained in sequence, and the solution values of the third objective and the second objective are used as the input values of the fourth objective. The solution value of completing the fourth objective is used as the output value of the multi-objective solution; S34, according to the output value of the multi-objective solution, the data increment of each sub-task under the corresponding constraint condition when the multi-objective solution is completed is calculated, and the sub-tasks are sorted into a task queue according to the objective when the sub-task is solved, from the largest to the smallest data increment.
2. The data transmission analysis and processing method for a construction enterprise information system according to claim 1, characterized in that, The implementation of step S2 also includes: S21, perform source tracking for changes in the sending tasks under the current batch, and set the execution flow of the sending tasks using the transmission tags of sending quantity, number of failures and lock period; S22, using the previous batch of data and the next batch of data in the current execution process as the pre-part and post-part respectively, to extend the running status of the sending task; S23, update the running state with the extended periodic component, and use the updated running state as the task execution state when data changes.
3. The data transmission analysis and processing method for a construction enterprise information system according to claim 1, characterized in that, Step S34 also includes: In response to the output value after multi-objective solution, each subtask is labeled with priority based on the normalized value of the product of the service rate of the current subtask and the number of receiving systems; the priority of each subtask relative to the task queue is determined. If the priority of each subtask is different under multiple data increment calculations, the average priority of each subtask is taken as the priority of the current subtask, and the task queue is sorted according to the priority of the subtask.
4. The data transmission analysis and processing method for a construction enterprise information system according to claim 1, characterized in that, Step S4 can be implemented in the following ways: S41, Generate the initial identifier of the current task queue based on the execution status of multiple tasks corresponding to the current task queue; S42, associate the number of sending failures and the locking period during the execution of the sending task in the historical data with the receiving system corresponding to each sending task, and obtain the failure flag on the receiving system; S43, read the number of failure flags on the receiving system, write the failure flags into the current task queue, calculate the similarity between the failure flags and the corresponding initial flags, and regard the calculated average similarity as the output adjustment ratio.
5. The data transmission analysis and processing method for a construction enterprise information system according to claim 4, characterized in that, The implementation of step S43 also includes: The initial identifiers in the task queue are compared one by one. When the similarity between the initial identifier and the failure identifier is greater than or equal to a preset threshold, it is considered that the current initial identifier has a corresponding failure identifier. The least common subset of the failure identifiers is added to the corresponding initial identifier, and the current task queue is updated.
6. The data transmission analysis and processing method for a construction enterprise information system according to claim 1, characterized in that, Step S5 can be implemented in the following ways: S51, based on the adjustment ratio of the sending tasks in each batch, the current batch is clustered according to the adjustment ratio and divided into multiple sub-batches; S52, using the value range of the adjustment ratio under each sub-batch, set the hierarchical response action for each sub-batch; S53: When a hierarchical response action is triggered in any sub-batch, the queue length of each sub-batch is adjusted according to the hierarchical response action, and the queue data volume of the task to be sent in each batch is combined and output according to the data volume contained in the queue length.
7. The data transmission analysis and processing method for a construction enterprise information system according to claim 6, characterized in that, The implementation of step S53 also includes: Verify the queue data volume under the corresponding sub-batch. Within the preset waiting time, determine whether the number of sub-batches that trigger the same hierarchical response action is greater than 1. If so, combine the sub-batches within the preset time period one by one to form the output batch. If not, the adjustment ratio of each batch is used to combine the sub-batch into an output batch, and the queue data volume of the combined output batch is regarded as the queue data volume of the task to be sent in each batch.
Citation Information
Patent Citations
Data transmission protection method and system for data transmission network based on communication state
CN117834307A
Multi-terminal series data transmission interaction system and method based on communication network
CN119603126A
Resource planning method and device, electronic equipment and computer readable medium
CN119739508A
System and method for hybrid processing of mass offline data and mass real-time data
CN120256070A