Data processing method and system for centralized management platform

By collecting historical data transmission volume and real-time data in the centralized water management platform, dividing transmission levels and managing resource occupancy rates, the problems of data accumulation and delay are solved, and more efficient resource utilization and real-time response are achieved.

CN120067199AActive Publication Date: 2025-05-30TIANJIN DEV ZONE ESINT NETWORK SYST

Patent Information

Application Number
CN202510525513.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-05-30
Estimated Expiration
2045-04-25

AI Technical Summary

Technical Problem

The lack of feedback regulation capabilities of data processing in the existing centralized water management platform, resulting in the problem of data accumulation and delays for too long.

Method used

By collecting historical data transmission volumes, dividing idle time and normal time periods, determining the transmission level of real-time data based on important data models, monitoring resource occupancy and transmission speed, estimating transmission pressure, and performing data compression and waiting time management during data accumulation, increasing transmission priority and storing redundant data.

Benefits of technology

It realizes dynamic identification of system resource usage status, flexibly divides data transmission periods, improves the ability to identify the importance of different data, prevents data accumulation problems, and improves the resource utilization efficiency and real-time response capabilities of the data management platform.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067199A_ABST
    Figure CN120067199A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, and discloses a data processing method and system for a centralized management platform, and the method comprises the steps: collecting historical data transmission quantity, defining an idle time period and a normal time period, collecting real-time to-be-transmitted data, and determining the transmission level of the real-time to-be-transmitted data; obtaining a to-be-transmitted data volume of the first transmission level in a normal time period to determine a resource occupancy rate and a real-time transmission speed; obtaining the to-be-transmitted data volume in the idle time period, comparing the to-be-transmitted data volume with a data volume threshold value, judging whether data accumulation exists or not according to a comparison result, and if so, determining the waiting time limit of the to-be-transmitted data in each idle time period; and when the waiting time limit is reached, performing data compression on the untransmitted data, improving the transmission level of the backbone data, and storing the redundant data to the edge collaborative database. The problem of data accumulation is prevented, and the resource utilization efficiency and the real-time response capability of the data management platform are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and in particular, to a data processing method and system for a centralized management platform. Background Art

[0002] With the rapid development of information technology, various data acquisition and transmission systems are more widely used in all walks of life. In the water service system, there are usually multiple dispersed data acquisition terminals, business management systems, and user service platforms. For example, data between different secondary and tertiary units. These systems respectively record and process different types of data, such as equipment operation data, business data, customer data, etc. Due to the lack of a unified data processing platform, data is managed dispersedly, resulting in problems such as data redundancy, inconsistency, and easy loss. Therefore, it is necessary to establish a unified water service group data management platform.

[0003] However, there are some problems in the existing unified management platform: The current system has a scheduling mechanism for data transmission based on priorities, but usually only relies on preset priority strategies, lacking the ability to dynamically evaluate and adjust according to real-time situations, and cannot adjust the transmission strategy in real time according to actual situations. During peak hours or high-load situations, in order to ensure the transmission of core data, the transmission resources of non-core data will be compressed, which is likely to cause data accumulation, resulting in the system being unable to process data in time. Although the idle time period can be used to transmit low-priority data, high-priority data in the new data may further occupy resources, resulting in a large delay of low-priority data.

[0004] Therefore, it is necessary to design a data processing method and system for a centralized management platform to solve the problems existing in the current technology. Summary of the Invention

[0005] In view of this, the present invention proposes a data processing method and system for a centralized management platform, aiming to solve the problems that the data processing in the current water service centralized management platform lacks feedback adjustment ability, is prone to data accumulation, and has a too long delay time.

[0006] On the one hand, the present invention proposes a data processing method for a centralized management platform, including: Collecting historical data transmission volume to define idle time periods and normal time periods, collecting real-time data to be transmitted, and determining the transmission level of the real-time data to be transmitted based on an important data model; Obtaining the data volume to be transmitted at the first transmission level during the normal time period to determine the resource occupancy rate, and collecting the real-time transmission speed of the data volume to be transmitted at non-first transmission levels; Obtaining the data volume to be transmitted during the idle time period according to the real-time transmission speed, the growth rate of non-first transmission level data, and the resource occupancy rate; Compare the data volume to be transmitted during the idle period with the data volume threshold, and determine whether there is data accumulation according to the comparison result. When it is determined that there is data accumulation, determine the waiting time limit for the data to be transmitted during each idle period according to the data characteristics of the data to be transmitted during the idle period; When the waiting time limit is reached, perform data compression on the untransmitted data to obtain backbone data and redundant data, improve the transmission level of the backbone data, and store the redundant data in the edge collaborative database.

[0007] Further, when collecting the historical data transmission volume to define the idle period and the normal period, it includes: Analyze the historical data transmission volume in the recent 7 / 30 days based on the sliding window statistics and time series analysis model, and establish a distribution map according to the average transmission volume in each time period within the 24-hour cycle; Compare the average transmission volume in each period in the distribution map with the transmission volume threshold respectively. When the average transmission volume is lower than the transmission volume threshold, determine this period as the idle period; when the average transmission volume is higher than or equal to the transmission volume threshold, determine this period as the normal period.

[0008] Further, when obtaining the data volume to be transmitted at the first transmission level during the normal period to determine the resource occupancy rate, it includes: The resource occupancy rate is directly proportional to the data volume to be transmitted at the first transmission level during the normal period, and the value range of the resource occupancy rate is [0, 1].

[0009] Further, when obtaining the data volume to be transmitted during the idle period according to the real-time transmission speed, the growth rate of non-first transmission level data, and the resource occupancy rate, it includes: Collect the historical transmission records during the normal period, and divide the historical transmission records into short-time windows. Each short-time window includes the growth rate of historical non-first transmission level data, the historical transmission speed, and the historical resource occupancy rate; Use several consecutive short-time windows as the training set to train the long short-term memory (LSTM) prediction model; Input the real-time transmission speed, the growth rate of non-first transmission level data, and the resource occupancy rate into the trained prediction model to obtain the data volume to be transmitted during the idle period.

[0010] Further, when determining whether there is data accumulation according to the comparison result, it includes: When the data volume to be transmitted during the idle period is greater than the data volume threshold, it is determined that there is data accumulation; When the data volume to be transmitted during the idle period is less than or equal to the data volume threshold, it is determined that there is no data accumulation.

[0011] Further, when determining the waiting time limit for the data to be transmitted in each idle period according to the data characteristics of the data to be transmitted in the idle period, it includes: Taking the transmission level, the amount of data to be transmitted, and the data source of each data to be transmitted in the data to be transmitted in the idle period as data characteristics, comparing the data characteristics with the historical data characteristic set, and determining the waiting time limit according to the comparison result. The historical data characteristic set includes a number of historical data characteristics and a number of historical waiting time limits, and each historical data characteristic corresponds to a historical waiting time limit; When there is data in the historical data characteristic set whose similarity to the data characteristics is greater than the similarity threshold, taking the historical waiting time limit corresponding to the maximum similarity as the waiting time limit; When the similarity between the historical data characteristics in the historical data characteristic set and the data characteristics is less than or equal to the similarity threshold, determining a basic time limit according to the transmission level, determining a proportionality coefficient according to the amount of data to be transmitted and the data source, and determining the waiting time limit according to the basic time limit and the proportionality coefficient.

[0012] Further, when determining the basic time limit according to the transmission level, it includes: When the transmission level of the data to be transmitted in the idle period is the second transmission level, determining the basic time limit as the first duration; when the transmission level of the data to be transmitted in the idle period is the third transmission level, determining the basic time limit as the second duration; when the transmission level of the data to be transmitted in the idle period is the fourth transmission level, determining the basic time limit as the third duration; wherein, the first duration is less than the second duration, and the second duration is less than the third duration.

[0013] Further, when determining the proportionality coefficient according to the amount of data to be transmitted and the data source, and determining the waiting time limit according to the basic time limit and the proportionality coefficient, it includes: Inputting the amount of data to be transmitted and the data source into a preset fuzzy matching rule to obtain the proportionality coefficient. The value range of the proportionality coefficient is (0.8, 1.2), and determining the waiting time limit according to the basic time limit and the proportionality coefficient. The waiting time limit is the product of the basic time limit and the proportionality coefficient.

[0014] Further, when performing data compression on the untransmitted data to obtain backbone data and redundant data, it includes: Performing structured processing on the untransmitted data; the structured processing includes: segmenting the time-series data in the untransmitted data by device or time window, and determining high-probability data for each segment using Gaussian mixture distribution, and taking the high-probability data, inflection points, mutation points, and peaks as the backbone data; Cluster the event - type data in the untransmitted data by event type or recording time, obtain any data in each cluster according to the clustering result, and use the fault samples and alarm samples as the backbone data.

[0015] Compared with the prior art, the beneficial effects of the present invention are as follows: Through the distinction mechanism between idle time periods and normal time periods, combined with the historical data transmission volume, the dynamic recognition of the system resource usage status is realized. Furthermore, in actual operation, the data transmission time periods can be flexibly divided, and the real - time data to be transmitted is classified based on the important data model, improving the recognition ability of the importance of different data. By monitoring the resource occupancy rate of high - priority data and the real - time transmission speed and growth rate of low - priority data during normal time periods, the prediction of the transmission pressure during idle time periods is realized. By judging the data accumulation situation and introducing a waiting time limit mechanism, the long - term stagnant data is compressed, and the backbone data is extracted to improve its transmission priority, so that the redundant data is diverted to the edge collaborative database for storage. Thus, while ensuring the timely transmission of important information, the storage and transmission pressure are minimized, the problem of low - priority data accumulation caused by high - priority data continuously occupying resources during idle time periods is prevented, and the resource utilization efficiency and real - time response ability of the data management platform are improved.

[0016] On the other hand, the present application also provides a data processing system for a centralized management platform, which is used to apply the above - mentioned data processing method for a centralized management platform, including: An acquisition unit, configured to collect the historical data transmission volume to define idle time periods and normal time periods, collect real - time data to be transmitted, and determine the transmission level of the real - time data to be transmitted based on the important data model; obtain the resource occupancy rate by getting the amount of data to be transmitted at the first transmission level during the normal time period, and collect the real - time transmission speed of the amount of data to be transmitted at non - first transmission levels; A processing unit, configured to obtain the amount of data to be transmitted during the idle time period based on the real - time transmission speed, the growth rate of non - first - level data, and the resource occupancy rate; A judgment unit, configured to compare the amount of data to be transmitted during the idle time period with a data volume threshold, judge whether there is data accumulation according to the comparison result, and when it is determined that there is data accumulation, determine the waiting time limit for each piece of data to be transmitted during the idle time period according to the data characteristics of the data to be transmitted during the idle time period; when the waiting time limit is reached, perform data compression on the untransmitted data to obtain backbone data and redundant data, and improve the transmission level of the backbone data; An edge collaborative database, used to store the redundant data.

[0017] It can be understood that the above - mentioned data processing method and system for a centralized management platform have the same beneficial effects, which will not be elaborated here. Description of the Drawings

[0018] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings: Figure 1 It is a flowchart of a data processing method for a centralized management platform provided by an embodiment of the present invention; Figure 2 It is a functional block diagram of a data processing system for a centralized management platform provided by an embodiment of the present invention. Detailed embodiments

[0019] Hereinafter, exemplary embodiments of the present disclosure will be described in more detail with reference to the drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be completely conveyed to those skilled in the art. It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other. Hereinafter, the present invention will be described in detail with reference to the drawings and in combination with the embodiments.

[0020] In some embodiments of the present application, referring to Figure 1 as shown, a data processing method for a centralized management platform includes: S100: Collect historical data transmission volume to define idle time periods and normal time periods, collect real-time data to be transmitted, and determine the transmission level of the real-time data to be transmitted based on an important data model.

[0021] S200: Obtain the data volume to be transmitted at the first transmission level during the normal time period to determine the resource occupancy rate, and collect the real-time transmission speed of the data volume to be transmitted at non-first transmission levels.

[0022] S300: Obtain the data volume to be transmitted during the idle time period according to the real-time transmission speed, the growth rate of non-first transmission level data, and the resource occupancy rate.

[0023] S400: Compare the data volume to be transmitted during the idle time period with a data volume threshold, and determine whether there is data accumulation according to the comparison result. When it is determined that there is data accumulation, determine the waiting time limit for each piece of data to be transmitted during the idle time period according to the data characteristics of the data to be transmitted during the idle time period.

[0024] S500: When the waiting time limit is reached, perform data compression on the untransmitted data to obtain backbone data and redundant data, increase the transmission level of the backbone data, and store the redundant data in the edge collaboration database.

[0025] Specifically, in S100, historical data transmission volumes are first collected to divide the "idle time period" and the "normal time period", and the resource status characteristics of these two time periods are defined. At the same time, the currently real-time data to be transmitted is collected, and the constructed "important data model" is used to classify and score the data to determine its transmission level. The important data model can be constructed based on multi-dimensional information such as historical data transmission records, service types, data content characteristics, and user access frequencies. Specifically, a classification algorithm based on machine learning can be used to train the historical transmission data, extract key feature factors affecting data importance, such as the urgency of data generation time, the importance of associated devices, the priority of the business involved in the data, and the relevance to core indicators, assign corresponding weight scores to different data, and then establish an adaptively adjustable data importance evaluation model to achieve dynamic judgment and classification of the transmission levels of real-time data to be transmitted. In S200, the transmission volume of data with the first transmission level (i.e., the highest priority) is focused on, and the current resource occupancy rate is estimated from this, while the real-time transmission speed of low-priority data is collected. In S300, by integrating the real-time transmission speed, growth rate of non-first-level data, and system resource occupancy rate, the total amount of data to be transmitted that can be carried during the idle time period is calculated. In S400, the calculated amount of data to be transmitted during the idle time period is compared with a preset data volume threshold to determine whether there is a risk of data accumulation currently. If there is accumulation, combined with the characteristics of the type, priority, generation time, etc. of the data to be transmitted, a waiting time limit is intelligently set for each type of data to avoid indefinite detention. In S500, when the data has not been successfully transmitted when it reaches the set waiting time limit, data compression operations are performed on it, the backbone data representing the core business is extracted and its transmission level is raised for priority transmission, while the redundant data that can be discarded or post-processed is stored in the edge collaborative database and waits to be completed when there is a transmission gap.

[0026] It can be understood that by introducing time period identification driven by historical data, data level division supported by an importance evaluation model, and a dynamic scheduling mechanism that integrates transmission speed, data growth rate, and resource occupancy rate, a reasonable allocation, compression, transfer, and storage of data with different priorities are achieved, effectively preventing extreme skew usage of system resources and data transmission congestion problems. While improving data processing efficiency and response capabilities, the stability and robustness under peak loads are enhanced.

[0027] In some embodiments of the present application, when collecting the historical data transmission volume to define the idle time period and the normal time period, it includes: analyzing the historical data transmission volume in the recent 7 / 30 days based on the sliding window statistics and the time series analysis model, and establishing a distribution map according to the average transmission volume in each time period within a 24-hour cycle. Comparing the average transmission volume in each time period in the distribution map with the transmission volume threshold respectively. When the average transmission volume is lower than the transmission volume threshold, this time period is determined as the idle time period. When the average transmission volume is higher than or equal to the transmission volume threshold, this time period is determined as the normal time period.

[0028] In some embodiments of the present application, when obtaining the data volume to be transmitted at the first transmission level during the normal time period to determine the resource occupancy rate, it includes: the resource occupancy rate is directly proportional to the data volume to be transmitted at the first transmission level during the normal time period, and the value range of the resource occupancy rate is [0, 1].

[0029] It can be understood that by introducing the sliding window and the time series analysis, modeling the historical data transmission law, the scientificity and dynamics of the time period division are realized; at the same time, by directly quantifying the resource occupancy rate with the high-priority data volume, the resource evaluation is made more concise and efficient. The self-perception ability of the operating state is improved, providing accurate support for subsequent scheduling decisions, and enhancing the adaptability and intelligence of the system in complex operating scenarios.

[0030] In some embodiments of the present application, when obtaining the data volume to be transmitted during the idle time period according to the real-time transmission speed, the growth rate of non-first transmission level data, and the resource occupancy rate, it includes: collecting the historical transmission records during the normal time period, and dividing the historical transmission records into short time windows, each short time window including the historical growth rate of non-first transmission level data, the historical transmission speed, and the historical resource occupancy rate. Using a continuous number of short time windows as the training set to train the long short-term memory (LSTM) prediction model. Inputting the real-time transmission speed, the growth rate of non-first transmission level data, and the resource occupancy rate into the trained prediction model to obtain the data volume to be transmitted during the idle time period.

[0031] Specifically, the forgetting gate of the LSTM model is pre-constructed: ; Input gate: ; ; Cell state update: ; Output gate: ; ; Among them, represents the output of the forgetting gate, denotes the weight matrix of the forget gate, with a value range of [-1, 1], represents the hidden state at the previous moment, represents the input data, denotes the bias term of the forget gate, and σ is the sigmoid function, represents the output of the input gate, represents the candidate cell state, with a value range of (-1, 1), denotes the weight matrix of the input gate, with a value range of [-1, 1], denotes the weight matrix of the candidate cell state, with a value range of [-1, 1], 、 respectively denote the bias terms of the input gate and the candidate cell state, represents the cell state at the current moment, represents the cell state at the previous moment, represents the output of the output gate, represents the predicted data, denotes the weight matrix of the output gate, with a value range of [-1, 1], denotes the bias term of the output gate.

[0032] Specifically, the real-time transmission speed, the growth rate of non-first transmission level data, and the resource occupancy rate are used as input data to obtain the amount of data to be transmitted during idle periods.

[0033] It is understandable that the historical transmission records in normal periods are collected and analyzed, and divided into several short-time windows. Each window contains three key parameters: the growth rate of non-first transmission level data in history, the historical transmission speed, and the historical resource occupancy rate. This short-time window division method can effectively capture the dynamic change trend in system operation and enhance the model's perception ability of local patterns. Multiple consecutive short-time windows are used as training samples and input into the LSTM model for training. As a recurrent neural network with long-term and short-term memory capabilities, the structures such as the "forget gate", "input gate", "cell state update", and "output gate" of the LSTM can respectively achieve the shielding of invalid information, the introduction and reinforcement of valid information, the smooth update of the state, and the control of the prediction output, ensuring the model's ability to model the dynamic dependencies of time series. After the training is completed, during actual operation, the three parameters collected in real time - the real-time transmission speed, the growth rate of non-first transmission level data, and the current resource occupancy rate - are input into the LSTM model, and the predicted value output is the amount of data to be transmitted during the idle period. This prediction result not only reflects the comprehensive influence of the current state but also takes into account the evolution trend of historical data, thus realizing the early perception and response to the future data accumulation risk. By constructing a multivariate time series prediction model based on LSTM, the high-precision prediction of the amount of data to be transmitted during the idle period is realized, and the ability to predict the upcoming data congestion is improved.

[0034] In some embodiments of the present application, when judging whether there is data accumulation according to the comparison result, it includes: when the amount of data to be transmitted during the idle period is greater than the data volume threshold, it is determined that there is data accumulation. When the amount of data to be transmitted during the idle period is less than or equal to the data volume threshold, it is determined that there is no data accumulation.

[0035] In some embodiments of the present application, when determining the waiting time limit for each piece of data to be transmitted during the idle period according to the data characteristics of the data to be transmitted during the idle period, it includes: taking the transmission level, the amount of transmitted data, and the data source of each piece of data to be transmitted during the idle period as data characteristics, comparing the data characteristics with the historical data characteristic set, and determining the waiting time limit according to the comparison result. The historical data characteristic set includes several historical data characteristics and several historical waiting time limits, and each historical data characteristic corresponds to a historical waiting time limit.

[0036] Specifically, when there is data in the historical data characteristic set whose similarity to the data characteristics is greater than the similarity threshold, the historical waiting time limit corresponding to the maximum similarity value is taken as the waiting time limit. When the similarity between the historical data characteristics in the historical data characteristic set and the data characteristics is less than or equal to the similarity threshold, the basic time limit is determined according to the transmission level, the proportional coefficient is determined according to the amount of transmitted data and the data source, and the waiting time limit is determined according to the basic time limit and the proportional coefficient.

[0037] It is understandable that the characteristic information of each data to be transmitted is extracted, including its transmission level (priority), data volume, and data source (such as the device or area it belongs to), to form a complete feature vector. This feature vector is compared with a pre-constructed historical data feature set, which contains a large number of records of the transmission behaviors of similar types of data in history and their corresponding waiting time limits. When there is historical data with a similarity higher than the set threshold, the waiting time limit corresponding to the one with the highest similarity is preferentially used as a reference value to achieve "prediction based on experience"; if no sufficiently similar historical data is found, it switches to a rule-based method to set the waiting time limit: a basic time limit is set according to the transmission level, and then a proportionality coefficient is generated by combining the data volume and source information, and finally the two are combined to determine the final waiting time. By integrating the threshold judgment mechanism of the data accumulation state and the intelligent waiting time limit setting of multi-feature matching, the recognition of the load state is realized, ensuring the personalization and optimal configuration of data scheduling in resource-constrained situations. It improves the resource regulation ability and response flexibility in a high-concurrency environment, helps reduce the risk of data retention, and improves the utilization rate of the transmission link.

[0038] In some embodiments of the present application, when determining the basic time limit according to the transmission level, it includes: when the transmission level of the data to be transmitted during the idle period is the second transmission level, determining the basic time limit as the first duration. When the transmission level of the data to be transmitted during the idle period is the third transmission level, determining the basic time limit as the second duration. When the transmission level of the data to be transmitted during the idle period is the fourth transmission level, determining the basic time limit as the third duration. Among them, the first duration is less than the second duration, and the second duration is less than the third duration.

[0039] In some embodiments of the present application, when determining the proportionality coefficient according to the transmission data volume and the data source, and determining the waiting time limit according to the basic time limit and the proportionality coefficient, it includes: inputting the transmission data volume and the data source into a preset fuzzy matching rule to obtain the proportionality coefficient, and the value range of the proportionality coefficient is (0.8, 1.2). The waiting time limit is determined according to the basic time limit and the proportionality coefficient, and the waiting time limit is the product of the basic time limit and the proportionality coefficient.

[0040] Specifically, the transmission data volume and the data source are input into the preset fuzzy matching rule. The proportionality coefficient is determined based on the input values. The fuzzy matching rule converts the input transmission data volume and data source into a proportionality coefficient through setting a rule base and a membership function, which is used to adjust the basic time limit. For example: if the transmission data volume is greater than 70% of the transmission volume threshold or the data source is sensor data, the proportionality coefficient is 1.1. If the transmission data volume is greater than 50% of the transmission volume threshold or the data source is business data, the proportionality coefficient is 0.9.

[0041] It can be understood that by introducing a transmission level stratification mechanism to determine the basic time limit and supplementing it with a fuzzy logic matching mechanism to dynamically adjust the proportional coefficient, the setting of the waiting time limit is both regular, flexible and intelligent. This avoids the problems of important data delay or resource waste caused by fixed strategies, improves the adaptability of the scheduling strategy to complex data characteristics, and helps to achieve better data transmission efficiency and stability in a high-concurrency environment.

[0042] In some embodiments of the present application, when data compression is performed on the untransmitted data to obtain backbone data and redundant data, it includes: performing structured processing on the untransmitted data. The structured processing includes: segmenting the time series data in the untransmitted data by device or time window, and using Gaussian mixture distribution to determine the high-probability data for each segment. The high-probability data, inflection points, mutation points, and peak values are used as the backbone data. Clustering the event data in the untransmitted data by event type or recording time, and obtaining any data in each cluster according to the clustering result, and using the fault samples and alarm samples as the backbone data.

[0043] Specifically, the untransmitted data is structurally processed and divided into two major categories: time series data and event data according to the data type, and targeted compression strategies are adopted. For time series data (such as device operation logs, monitoring data, etc.), the original data is segmented based on the device number or time window. Gaussian mixture distribution is used to model the data probability distribution within each segment, and the data points in the high-probability region are identified from it. Such data usually represents the normal operation trend or key behavior patterns. Data points with significant features are also extracted, including inflection points (trend change turns), mutation points (sharp data changes), peak values (extreme value states), etc. These highly representative data points are jointly classified as backbone data. Ensure that the main structure and abnormal key points in the time series are retained during the compression process, facilitating subsequent analysis or quick restoration of the original appearance.

[0044] For event data (such as fault records, alarm logs, etc.), a clustering analysis method is adopted, and the data is clustered according to the event type (such as device alarm, communication interruption) or recording time (such as related events within the same time period). Representative data is selected and retained from each clustering cluster, and at the same time, the involved fault samples and alarm samples are completely retained as backbone data to ensure that the system has the support of a complete event chain in retrospective and diagnostic analysis. The remaining unselected data is defined as redundant data, which can be temporarily stored in the edge collaborative database according to the resource situation for subsequent processing.

[0045] It can be understood that through Gaussian mixture modeling and clustering analysis techniques, the structured compression processing of different data types is realized. While successfully extracting the backbone data, the transmission and storage requirements of redundant information are significantly reduced. This ensures the priority transmission of key data in resource-constrained environments and avoids the operation and maintenance risks caused by information loss. On the other hand, through intelligent compression, the overall data load is reduced, the transmission efficiency and response speed are improved, and it is applicable to the data scheduling optimization in complex scenarios of large-scale water systems.

[0046] In the above embodiments, through the differentiation mechanism between idle time periods and normal time periods, combined with the historical data transmission volume, the dynamic recognition of the system resource usage status is realized. Furthermore, in actual operation, the data transmission time periods can be flexibly divided, and the real-time data to be transmitted is classified based on the important data model, improving the recognition ability of the importance of different data. By monitoring the resource occupancy rate of high-priority data and the real-time transmission speed and growth rate of low-priority data during normal time periods, the prediction of the transmission pressure during idle time periods is realized. By judging the data accumulation situation and introducing a waiting time limit mechanism, the long-term stranded data is compressed, and the backbone data is extracted to improve its transmission priority, while the redundant data is diverted to the edge collaborative database for storage. Thus, while ensuring the timely transmission of important information, the storage and transmission pressure are minimized, and the problem of low-priority data accumulation caused by high-priority data continuously occupying resources during idle time periods is prevented, improving the resource utilization efficiency and real-time response ability of the data management platform.

[0047] In another preferred manner based on the above embodiments, refer to Figure 2 As shown, this embodiment provides a data processing system for a centralized management platform, which is used to apply the above data processing method for a centralized management platform, including: An acquisition unit, configured to acquire the historical data transmission volume to define idle time periods and normal time periods, acquire the real-time data to be transmitted, and determine the transmission level of the real-time data to be transmitted based on the important data model. Obtain the data volume to be transmitted at the first transmission level during normal time periods to determine the resource occupancy rate, and acquire the real-time transmission speed of the data volume to be transmitted at non-first transmission levels.

[0048] A processing unit, configured to obtain the data volume to be transmitted during idle time periods based on the real-time transmission speed, the growth rate of non-first transmission level data, and the resource occupancy rate.

[0049] A judgment unit, configured to compare the data volume to be transmitted during idle time periods with the data volume threshold, and judge whether there is data accumulation according to the comparison result. When it is determined that there is data accumulation, determine the waiting time limit for each piece of data to be transmitted during idle time periods according to the data characteristics of the data to be transmitted during idle time periods. When the waiting time limit is reached, compress the untransmitted data to obtain the backbone data and redundant data, and improve the transmission level of the backbone data.

[0050] An edge collaborative database for storing redundant data.

[0051] It can be understood that through the differentiation mechanism between idle time periods and normal time periods, combined with historical data transmission volumes, dynamic identification of the system resource usage status is achieved. Furthermore, during actual operation, data transmission time periods can be flexibly divided, and real-time data to be transmitted can be classified based on an important data model, enhancing the ability to identify the importance of different data. By monitoring the resource occupancy rate of high-priority data and the real-time transmission speed and growth rate of low-priority data during normal time periods, an estimation of the transmission pressure during idle time periods is realized. By judging the data accumulation situation and introducing a waiting time limit mechanism, long-term stagnant data is compressed, and the backbone data is extracted to enhance its transmission priority, while redundant data is diverted to the edge collaborative database for storage. Thus, while ensuring the timely transmission of important information, the storage and transmission pressures are minimized, preventing the problem of low-priority data accumulation caused by high-priority data continuously occupying resources during idle time periods, and enhancing the resource utilization efficiency and real-time response ability of the data management platform.

[0052] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0053] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowcharts and / or block diagrams can also be implemented. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0054] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions in the flow Figure 1One or more processes and / or blocks Figure 1 The functions specified in one or more blocks.

[0055] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one Figure 1 One or more processes and / or blocks Figure 1 The steps of the functions specified in one or more blocks.

[0056] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: modifications or equivalent replacements can still be made to the specific implementation manners of the present invention. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention shall be covered by the protection scope of the claims of the present invention.

Claims

1. A data processing method for a centralized management platform, characterized in that: include: Collect historical data transmission volume to define idle time periods and normal time periods, collect real-time data to be transmitted, and determine the transmission level of the real-time data to be transmitted based on the important data model; Acquire the amount of data to be transmitted at the first transmission level during the normal period to determine the resource occupancy rate, and collect the real-time transmission speed of the amount of data to be transmitted that is not at the first transmission level; Obtaining the amount of data to be transmitted during the idle time period according to the real-time transmission speed, the non-first transmission level data growth rate, and the resource occupancy rate; Comparing the amount of data to be transmitted during the idle time period with the data amount threshold, judging whether there is data accumulation according to the comparison result, and when it is judged that there is data accumulation, determining the waiting time limit of the data to be transmitted in each idle time period according to the data characteristics of the data to be transmitted in the idle time period; When the waiting time limit is reached, data compression is performed on the untransmitted data to obtain backbone data and redundant data, and the transmission level of the backbone data is improved, and the redundant data is stored in the edge collaborative database.

2. The data processing method for a centralized management platform according to claim 1, characterized in that: When collecting historical data transmission volume to define off-peak and normal periods, it includes: The historical data transmission volume of the past 7 / 30 days is analyzed based on the sliding window statistics and time series analysis model, and a distribution chart is established based on the average transmission volume in each time period within a 24-hour cycle; The average transmission volume of each time period in the distribution diagram is compared with the transmission volume threshold. When the average transmission volume is lower than the transmission volume threshold, the time period is determined as the idle time period; when the average transmission volume is higher than or equal to the transmission volume threshold, the time period is determined as the normal time period.

3. The data processing method for a centralized management platform according to claim 1, characterized in that: When obtaining the amount of data to be transmitted at the first transmission level in the normal period to determine the resource occupancy rate, it includes: The resource occupancy rate is proportional to the amount of data to be transmitted at the first transmission level in the normal period, and the resource occupancy rate has a value range of [0, 1].

4. The data processing method for a centralized management platform according to claim 1, characterized in that: When the amount of data to be transmitted during the idle time period is obtained according to the real-time transmission speed, the non-first transmission level data growth rate, and the resource occupancy rate, it includes: Collecting historical transmission records during the normal period, and dividing the historical transmission records into short time windows, each of which includes a historical non-first transmission level data growth rate, a historical transmission speed, and a historical resource occupancy rate; Using a plurality of consecutive short time windows as training sets to train the long-short time prediction model LSTM; The real-time transmission speed, non-first transmission level data growth rate and resource occupancy rate are input into the trained prediction model to obtain the amount of data to be transmitted during the idle time period.

5. The data processing method for a centralized management platform according to claim 1, characterized in that: When judging whether there is data accumulation based on the comparison results, it includes: When the amount of data to be transmitted during the idle time period is greater than the data amount threshold, it is determined that data accumulation exists; When the amount of data to be transmitted during the idle time period is less than or equal to the data amount threshold, it is determined that there is no data accumulation.

6. The data processing method for a centralized management platform according to claim 5, characterized in that: Determining the waiting time limit for data to be transmitted in each idle time period according to the data characteristics of the data to be transmitted in the idle time period includes: The transmission level, transmission data volume and data source of each data to be transmitted in the idle time period are used as data features, and the data features are compared with a historical data feature set, and the waiting time limit is determined according to the comparison result, wherein the historical data feature set includes a plurality of historical data features and a plurality of historical waiting time limits, and each of the historical data features corresponds to a historical waiting time limit; When there is data in the historical data feature set whose similarity with the data feature is greater than a similarity threshold, the historical waiting time limit corresponding to the maximum similarity value is used as the waiting time limit; When the similarity between the historical data features in the historical data feature set and the data features is less than or equal to the similarity threshold, the basic time limit is determined according to the transmission level, the proportional coefficient is determined according to the transmission data volume and the data source, and the waiting time limit is determined according to the basic time limit and the proportional coefficient.

7. The data processing method for a centralized management platform according to claim 6, characterized in that: When determining the basic time limit according to the transmission level, it includes: When the transmission level of the data to be transmitted in the idle time period is the second transmission level, the basic time limit is determined to be the first time length; when the transmission level of the data to be transmitted in the idle time period is the third transmission level, the basic time limit is determined to be the second time length; when the transmission level of the data to be transmitted in the idle time period is the fourth transmission level, the basic time limit is determined to be the third time length; wherein the first time length is smaller than the second time length, and the second time length is smaller than the third time length.

8. The data processing method for a centralized management platform according to claim 7, characterized in that: Determining the proportionality coefficient according to the amount of transmitted data and the data source, and determining the waiting time limit according to the basic time limit and the proportionality coefficient, includes: The transmission data volume and data source are input into a preset fuzzy matching rule to obtain the proportional coefficient, the value range of which is (0.8, 1.2). The waiting time limit is determined according to the basic time limit and the proportional coefficient, and the waiting time limit is the product of the basic time limit and the proportional coefficient.

9. The data processing method for a centralized management platform according to claim 8, characterized in that: When the untransmitted data is compressed to obtain the backbone data and redundant data, it includes: The untransmitted data is subjected to structured processing; the structured processing includes: segmenting the time series data in the untransmitted data according to the device or time window, determining high probability data for each segment using Gaussian mixture distribution, and using the high probability data, inflection points, mutation points and peak values ​​as the backbone data; The event data in the untransmitted data are clustered according to event type or recording time, and any data in each cluster is acquired according to the clustering result, and the fault samples and the alarm samples are used as the backbone data.

10. A data processing system for a centralized management platform, used for applying the data processing method for a centralized management platform as claimed in any one of claims 1 to 9, characterized in that: include: A collection unit is configured to collect historical data transmission volume to define idle time periods and normal time periods, collect real-time data to be transmitted, and determine a transmission level of the real-time data to be transmitted based on an important data model; Acquire the amount of data to be transmitted at the first transmission level during the normal period to determine the resource occupancy rate, and collect the real-time transmission speed of the amount of data to be transmitted that is not at the first transmission level; A processing unit is configured to obtain the amount of data to be transmitted during the idle time period by using the real-time transmission speed, the non-first transmission level data growth rate and the resource occupancy rate; The judging unit is configured to compare the amount of data to be transmitted in the idle time period with a data amount threshold, and judge whether there is data accumulation according to the comparison result; when it is determined that there is data accumulation, determine the waiting time limit of the data to be transmitted in each idle time period according to the data characteristics of the data to be transmitted in the idle time period; when the waiting time limit is reached, compress the data not transmitted to obtain the backbone data and redundant data, and improve the transmission level of the backbone data; The edge collaborative database is used to store the redundant data.

Citation Information

Patent Citations

  • Electric power communication data management system and method

    CN113438116A

  • Data transmission method and system and edge service equipment

    CN113783798A

  • Data transmission method and device, storage medium and electronic equipment

    CN116708142A

  • Method and system for monitoring ship for transporting hazardous chemicals

    CN118770485A

  • Intelligent traffic data transmission method and system based on industrial switch

    CN119094461A

Cited By

  • Underlying data reporting optimization method and device based on cloud platform

    CN120512469A