A power detection system based on multidimensional data
By introducing data acquisition, processing, allocation and redundant adjustment units into the power detection system, data allocation is optimized based on the data state and node state, the problem of inefficient data processing in the prior art is solved, and more efficient and secure data transmission and processing is achieved.
Patent Information
- Application Number
- CN202411233979.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-04
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2044-09-04
AI Technical Summary
The existing power detection system cannot make targeted adjustments based on the correlation status and data volume between the actual collected data, resulting in poor data processing efficiency.
By determining the data state and node state, a reasonable allocation method and compression method are selected to optimize data transmission processing by optimizing data transmission processing.
It improves data processing efficiency, avoids inefficiency and data loss caused by overloading of processing nodes, ensures the rationality and security of data allocation, and improves the accuracy of redundant data identification.
Smart Images

Figure CN119166347B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing, and in particular to a power detection system based on multi-dimensional data. Background Art
[0002] With the continuous increase in power equipment, the amount of data that the power detection system needs to process is also growing continuously. Due to the limited processing capacity of the processing nodes, it is impossible to choose a reasonable way to distribute the data when the data volume surges, resulting in poor data transmission efficiency. Therefore, how to adjust the data transmission processing according to the correlation status and data volume between the actual collected data to improve data processing efficiency is a technical problem that needs to be urgently solved by technical personnel in this field.
[0003] Chinese Patent Publication No. CN116680423B discloses a method, device, equipment and medium for managing multi-source heterogeneous data in the power supply chain, including: collecting multi-source heterogeneous data under the power supply chain, and classifying and storing the multi-source heterogeneous data according to data type; using an encoder of the CLIP algorithm to vectorize data of different data types to obtain multi-dimensional feature vectors; calculating the similarity of the multi-dimensional feature vectors obtained by converting structured data in the multi-source heterogeneous data, and merging the fields of the multi-dimensional feature vectors whose similarity exceeds a preset threshold; clustering the multi-dimensional feature vectors obtained by converting unstructured data in the multi-source heterogeneous data, and classifying and storing the clustered multi-dimensional feature vectors. It can be seen that the above technical solution has the following problems: it is impossible to make targeted adjustments to the data transmission and processing according to the correlation status and data volume between the actual collected data, resulting in poor data processing efficiency. Summary of the Invention
[0004] To this end, the present invention provides an electric power detection system based on multidimensional data to overcome the problem in the prior art that the data transmission and processing cannot be specifically adjusted according to the correlation status and data volume between the actual collected data, resulting in poor data processing efficiency.
[0005] To achieve the above objectives, the present invention provides a power detection system based on multidimensional data, comprising:
[0006] A data acquisition unit, comprising a plurality of data acquisition modules for performing data acquisition;
[0007] a data processing unit connected to the data acquisition unit, configured to determine a data state according to a data reference value of the data acquisition unit and a processing reference value of the data processing unit, and to determine a processing method according to the data state;
[0008] A first allocation unit, connected to the data processing unit, is used to determine, based on a reference quantity ratio, a processing method for allocating associated combined data and non-associated data based on an associated influence coefficient or a node coefficient;
[0009] a second allocating unit connected to the data processing unit, for executing a processing method for allocating the data to be allocated according to the node processing reference value;
[0010] a redundancy analysis unit connected to the first allocation unit and the second allocation unit, configured to, when performing redundant data analysis on a single processing node, divide the incoming data into a plurality of data blocks of equal size, generate a hash value corresponding to each data block using a secure hash algorithm, and determine the node status based on the maximum difference between the data blocks and the ratio of the number of similar data blocks;
[0011] The redundancy adjustment unit is connected to the redundancy analysis unit and is used to determine the adjustment method according to the node status, that is, to adjust the sliding window length or to adjust the reference value of the number of data blocks, and to directly compress or differentially compress the data blocks according to the difference of the data blocks.
[0012] Furthermore, the data processing unit determines the data state according to the data reference value of the data acquisition unit and the processing reference value of the data processing unit;
[0013] The data state includes a first data state in which the data reference value is greater than or equal to a preset data reference value or the processing reference value is less than a preset processing reference value, and a second data state in which the data reference value is less than the preset data reference value and the processing reference value is greater than or equal to the preset processing reference value.
[0014] Furthermore, the data processing unit determines a processing mode according to the data status;
[0015] If the data state is the first data state, the data processing unit determines that the processing method is to determine the associated allocation method according to the reference quantity ratio;
[0016] If the data state is the second data state, the data processing unit determines that the processing mode is to distribute the data to be distributed according to the node processing reference value.
[0017] Furthermore, the first allocation unit determines the associated allocation method according to the reference quantity ratio;
[0018] If the reference quantity ratio is greater than or equal to the preset reference quantity ratio, the associated allocation method is to allocate according to the associated influence coefficient;
[0019] If the reference quantity ratio is less than the preset reference quantity ratio, the association allocation method is to allocate according to the node coefficient;
[0020] The reference quantity ratio is equal to the ratio of the sum of the numbers of data acquisition modules corresponding to each associated combination to the number of all data acquisition modules.
[0021] Furthermore, the first allocation unit sequentially performs association analysis on each data acquisition module in a preset order, and records the data acquisition module that undergoes association analysis as a target data acquisition module. The association analysis includes detecting a correlation coefficient between the target data acquisition module and other data acquisition modules, and uniformly recording the data acquisition modules whose correlation coefficients with the target data acquisition module are greater than or equal to a preset correlation coefficient as a correlation combination with the target data acquisition module. After the association analysis for the target data acquisition module is completed, the association analysis is continued for the data acquisition modules that are not recorded as a correlation combination.
[0022] The calculation formula of the correlation coefficient r is:
[0023]
[0024] Among them, n is the number of data points contained in a single data column, x i and y i are the values of the i-th data point in the two data columns respectively, is x i The average value of the corresponding data column, y i The average value of the corresponding data column, i = 1, 2, 3, ..., n;
[0025] The default order is from large to small data reference amount.
[0026] Furthermore, the first allocation unit performs association allocation on each associated combination data according to an association preset order, wherein the first allocation unit allocates the associated combination data to a processing node with the smallest association influence coefficient in association allocation for a single associated combination data;
[0027] After the first allocation unit completes the allocation of the associated combined data, it sequentially allocates the non-associated data in an associated manner according to a preset order, and for a single non-associated data, allocates the non-associated data to a processing node with the smallest associated influence coefficient;
[0028] The associated combination data is the data to be allocated of the data acquisition module corresponding to a single associated combination, and the non-associated data is the data to be allocated of a single data acquisition module that has not been associated with the combination;
[0029] The preset association order is the order from large to small of the combined data reference amount.
[0030] Furthermore, the first allocating unit allocates the associated combined data and the non-associated data according to the node coefficient;
[0031] The first allocation unit allocates each associated combination data to a corresponding processing node with the smallest node coefficient, and allocates each non-associated data to a corresponding processing node with the smallest node coefficient.
[0032] Furthermore, the second allocating unit allocates the data to be allocated according to the node processing reference value, and for a single piece of data to be allocated, selects a processing node with the largest node processing reference value for allocation.
[0033] Furthermore, when the redundancy analysis unit performs redundant data analysis on a single processing node, it divides the incoming data into a number of data blocks of the same size, generates a hash value corresponding to each data block using a secure hash algorithm, and determines the node status based on the maximum difference between the data blocks and the proportion of similar data blocks;
[0034] The node status includes a first node status in which the maximum difference of data blocks is greater than or equal to the preset maximum difference of data blocks or the proportion of similar data blocks is greater than or equal to the preset proportion of similar data blocks, and a second node status in which the maximum difference of data blocks is less than the preset maximum difference of data blocks and the proportion of similar data blocks is less than the preset proportion of similar data blocks.
[0035] Furthermore, the redundancy adjustment unit determines an adjustment mode according to the node status;
[0036] In the first node state, the redundancy adjustment unit determines that the adjustment method is to reduce the sliding window length;
[0037] In the second node state, the redundancy adjustment unit determines that the adjustment method is to increase the reference value of the number of data blocks;
[0038] The relationship between the reduction value of the sliding window length and the comprehensive difference is negatively correlated;
[0039] The relationship between the increase in the number of data blocks and the comprehensive difference is a positive correlation.
[0040] Compared with the prior art, the beneficial effect of the present invention lies in that, in the technical solution of the present invention, the data processing unit determines the data status based on the data reference value of the data acquisition unit and the processing reference value of the data processing unit, effectively reflects the processing capability of the current processing node through the data reference value of the data acquisition unit and the processing reference value of the data processing unit, and adaptively selects different data allocation methods according to the data status, so that the selection of data allocation method is more in line with the actual working scenario, avoiding the problem of poor processing efficiency caused by excessive processing of data by the processing node, and thus improving the processing efficiency of the processing node by reasonably allocating data.
[0041] Furthermore, the first allocation unit in the present invention determines the processing method based on the reference quantity ratio, effectively reflects the correlation of the data to be allocated through the reference quantity ratio, and then adaptively selects the correlation allocation method, so that the selection of the correlation allocation method is more in line with the actual working scenario, and the data analysis efficiency is guaranteed by combining and allocating highly correlated data. At the same time, the weakly correlated data is allocated to other processing nodes, avoiding the problem of excessive data correlation in a single node leading to serious data loss when the processing node fails, thereby ensuring the security of the data.
[0042] Furthermore, the first allocation unit in the present invention effectively reflects the data throughput of the current processing node and the distance between the data acquisition module and the node through the node coefficient, and then allocates the associated combination data and non-associated data according to the node coefficient, thereby ensuring that each associated combination data and each non-associated data can be allocated to a node with a stronger data throughput and a closer distance, avoiding the problems of data transmission delay and long waiting time, and thus improving data processing efficiency.
[0043] Furthermore, the second allocation unit in the present invention determines the allocation method based on the node processing reference value of each processing node, and effectively reflects the processing capacity of each processing node through the node processing reference value, avoiding the problem of allocating the data to be allocated to nodes with poor processing capacity, thereby ensuring the rationality of data allocation.
[0044] Furthermore, the redundancy analysis unit in the present invention determines the node status based on the maximum difference of data blocks and the proportion of similar data blocks, effectively reflects the data redundancy situation of the current processing node through the maximum difference of data blocks and the proportion of similar data blocks, and adaptively selects different adjustment methods according to the node status, so that the selection of adjustment method is more in line with the actual working scenario, avoiding the problem of misjudging normal data as redundant data and inaccurate redundant data identification due to excessive length of data blocks, thereby improving the accuracy of redundant data identification. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 A unit connection diagram of the multi-dimensional data-based power detection system of the present invention;
[0046] Figure 2 A flow chart of determining a data state according to a data reference value of a data acquisition unit and a processing reference value of a data processing unit according to the present invention;
[0047] Figure 3 This is a flow chart of the present invention for determining a processing method according to data status;
[0048] Figure 4 This is a flow chart of the present invention for determining an adjustment method according to a node status. DETAILED DESCRIPTION
[0049] In order to make the objects and advantages of the present invention more clearly understood, the present invention is further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are merely used to explain the present invention and are not intended to limit the present invention.
[0050] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood by those skilled in the art that these embodiments are only used to explain the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.
[0051] It should be noted that, in the description of the present invention, terms such as "up", "down", "left", "right", "inside", and "outside" indicating directions or positional relationships are based on the directions or positional relationships shown in the accompanying drawings. This is only for the convenience of description and does not indicate or imply that the device or element must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it cannot be understood as a limitation on the present invention.
[0052] Furthermore, it should be noted that, in the description of the present invention, unless otherwise expressly specified or limited, the terms "mounted," "connected," and "connected" should be understood in a broad sense. For example, they may refer to fixed connections, detachable connections, or integral connections; mechanical connections or electrical connections; direct connections or indirect connections through an intermediate medium; and internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on specific circumstances.
[0053] See also Figures 1 to 4 As shown, the present invention provides a power detection system based on multi-dimensional data, comprising:
[0054] A data acquisition unit, comprising a plurality of data acquisition modules for performing data acquisition;
[0055] a data processing unit connected to the data acquisition unit, configured to determine a data state according to a data reference value of the data acquisition unit and a processing reference value of the data processing unit, and to determine a processing method according to the data state;
[0056] A first allocation unit, connected to the data processing unit, is used to determine, based on a reference quantity ratio, a processing method for allocating associated combined data and non-associated data based on an associated influence coefficient or a node coefficient;
[0057] a second allocating unit connected to the data processing unit, for executing a processing method for allocating the data to be allocated according to the node processing reference value;
[0058] a redundancy analysis unit connected to the first allocation unit and the second allocation unit, configured to, when performing redundant data analysis on a single processing node, divide the incoming data into a plurality of data blocks of equal size, generate a hash value corresponding to each data block using a secure hash algorithm, and determine the node status based on the maximum difference between the data blocks and the ratio of the number of similar data blocks;
[0059] The redundancy adjustment unit is connected to the redundancy analysis unit and is used to determine the adjustment method according to the node status, that is, to adjust the sliding window length or to adjust the reference value of the number of data blocks, and to directly compress or differentially compress the data blocks according to the difference of the data blocks.
[0060] The data acquisition unit in the present invention is provided with a continuous cycle monitoring period. At the end of each monitoring period, the data processing unit determines the data status once. The length of the monitoring period can be set according to the needs of the user. The greater the user's demand for data allocation and processing efficiency, the shorter the length of the monitoring period. A value of a monitoring period is provided, and the monitoring period is 1 minute. The data acquisition module of the present invention includes a voltage sensor, a wind speed sensor, a temperature sensor, and a vibration sensor, etc. The data collected by the data acquisition module is the data to be allocated. The data processing unit includes several processing nodes with processing capabilities. The specific setting position and setting number of the processing nodes are not specifically limited, and are adaptively set according to the actual application scenario. This is content that is easy for technical personnel in this field to understand and will not be elaborated here.
[0061] In the present invention, direct compression compresses the entire data content contained in the entire data block. For a single data block, differential compression uses a sliding window technique to find duplicate data between the data block and a reference data block, and compresses the duplicate data content to reduce storage space. The compression method used may include, but is not limited to, Huffman coding, arithmetic coding, or a traditional compression algorithm, which the user can select based on actual needs. The reference data block is the historical data block that is closest to the characteristic value of the data block and is stored in the redundancy analysis unit. If the data block difference is within a first preset data block difference range, direct compression is performed on the data block; if the data block difference is within a second preset data block difference range, differential compression is performed on the data block; if the data block difference is within a third preset data block difference range, no compression is required for the data block.
[0062] Among them, the values within the first preset data block difference range are all less than or equal to the first preset data block difference, the values within the second preset data block difference range are all greater than the first preset data block difference and less than or equal to the second preset data block difference, and the values within the third preset data block difference range are all greater than the second preset data block difference. The values of the first preset data block difference and the second preset data block difference can be set by the user according to the actual application scenario. The greater the user's demand for precise processing of redundant data, the greater the values of the first preset data block difference and the second preset data block difference. Provided are values of the first preset data block difference and the second preset data block difference: the first preset data block difference is 30% and the second preset data block difference is 80%.
[0063] Specifically, the data processing unit determines the data state according to the data reference value of the data acquisition unit and the processing reference value of the data processing unit;
[0064] The data state includes a first data state in which the data reference value is greater than or equal to a preset data reference value or the processing reference value is less than a preset processing reference value, and a second data state in which the data reference value is less than the preset data reference value and the processing reference value is greater than or equal to the preset processing reference value.
[0065] Among them, the data reference value of the data acquisition unit is the sum of the data volume to be allocated collected by each data processing module within a single monitoring cycle, and the processing reference value of the data processing unit is the difference between the maximum data volume that each processing node can process and the data volume being processed by each processing node within a single monitoring cycle. The unit of data volume is GB.
[0066] The values of the preset data reference value and the preset processing reference value can be set by the user according to the actual application scenario. The greater the user's demand for the stability of the processing node, the smaller the value of the preset data reference value and the larger the value of the preset processing reference value. One value of a preset data reference value and a preset processing reference value is provided. The preset data reference value is 50GB. The method for confirming the preset processing reference value is to detect the processing reference value corresponding to the record with the same data reference value in the historical monitoring period, and record the average value of the processing reference value whose processing speed meets the user's requirements as the preset processing reference value. The historical monitoring period is all monitoring periods before the current monitoring period.
[0067] Specifically, the data processing unit determines a processing mode according to the data status;
[0068] If the data state is the first data state, the data processing unit determines that the processing method is to determine the associated allocation method according to the reference quantity ratio;
[0069] If the data state is the second data state, the data processing unit determines that the processing mode is to distribute the data to be distributed according to the node processing reference value.
[0070] Among them, for a single processing node, the node processing reference value is confirmed by recording the difference between the maximum amount of data that a single processing node can process and the amount of data being processed by the processing node in a single monitoring cycle as the node processing reference value of the processing node.
[0071] Specifically, the first allocation unit determines the associated allocation method according to the reference quantity ratio;
[0072] If the reference quantity ratio is greater than or equal to the preset reference quantity ratio, the associated allocation method is to allocate according to the associated influence coefficient;
[0073] If the reference quantity ratio is less than the preset reference quantity ratio, the association allocation method is to allocate according to the node coefficient;
[0074] The reference quantity ratio is equal to the ratio of the sum of the numbers of data acquisition modules corresponding to each associated combination to the number of all data acquisition modules.
[0075] The ratio=the sum of the number of data acquisition modules corresponding to each associated combination / the number of all data acquisition modules.
[0076] Specifically, the first allocation unit performs association analysis on each data acquisition module in a preset order, and records the data acquisition module that undergoes association analysis as a target data acquisition module. The association analysis includes detecting a correlation coefficient between the target data acquisition module and other data acquisition modules, and recording the data acquisition modules whose correlation coefficients with the target data acquisition module are greater than or equal to a preset correlation coefficient as a correlation combination with the target data acquisition module. After the association analysis for the target data acquisition module is completed, the association analysis is continued for the data acquisition modules that are not recorded as a correlation combination.
[0077] The calculation formula of the correlation coefficient r is:
[0078]
[0079] Among them, n is the number of data points contained in a single data column, x i and y i are the values of the i-th data point in the two data columns respectively, is x i The average value of the corresponding data column, y i The average value of the corresponding data column, i = 1, 2, 3, ..., n;
[0080] The default order is from large to small data reference amount.
[0081] The data reference amount is the amount of data to be allocated collected by a single data acquisition module in a single monitoring cycle; the single data to be allocated is taken as the starting point of the starting data, and an interval point is set every 1s in chronological order, and the starting point and all interval points are recorded as time points. The data column is the data value corresponding to each time point, and the data point is the data value corresponding to a single time point; x i The average value of the corresponding data column The calculation formula is: y i The average value of the corresponding data column The calculation formula is:
[0082]
[0083] The value of the preset correlation coefficient can be set according to the actual application scenario. The greater the user's demand for data analysis efficiency, the larger the value of the preset correlation coefficient. It can be understood that the value range of r is [0,1]. The higher the correlation degree of the data, the closer the value of r is to 1. A value of the preset correlation coefficient is provided, r = 0.8.
[0084] Specifically, the first allocation unit performs association allocation on each associated combination data according to an association preset order, wherein the first allocation unit allocates the associated combination data to a processing node with the smallest association influence coefficient in association allocation for a single associated combination data;
[0085] After the first allocation unit completes the allocation of the associated combined data, it sequentially allocates the non-associated data in an associated manner according to a preset order, and for a single non-associated data, allocates the non-associated data to a processing node with the smallest associated influence coefficient;
[0086] The associated combination data is the data to be allocated of the data acquisition module corresponding to a single associated combination, and the non-associated data is the data to be allocated of a single data acquisition module that has not been associated with the combination;
[0087] The preset association order is the order from large to small of the combined data reference amount.
[0088] The combined data reference quantity is the sum of the data reference quantities of the data acquisition modules corresponding to a single associated combination in a single monitoring period.
[0089] The correlation influence coefficient is denoted as Φ. It can be understood that the correlation influence coefficient between a single correlation combination and a single processing node is determined differently from the correlation influence coefficient between a single data acquisition module without a correlation combination and a single processing node. The calculation formula for the correlation influence coefficient Φ between a single correlation combination and a single processing node is: The calculation formula for the correlation influence coefficient Φ between a single unassociated data acquisition module and a single processing node is: Among them, n is the number of data acquisition modules corresponding to a single association combination, m is the number of data acquisition modules corresponding to a single processing node, and the data acquisition modules in a single association combination and the data acquisition modules in a single processing node are sorted in descending order according to the data reference amount, r pd is the correlation coefficient between the pth data acquisition module in a single correlation combination and the dth data acquisition module in a single processing node, r d is the correlation coefficient between a single data acquisition module that has not been associated and combined and the dth data acquisition module in a single processing node.
[0090] Specifically, the first allocation unit allocates the associated combination data and the non-associated data according to the node coefficient;
[0091] The first allocation unit allocates each associated combination data to a corresponding processing node with the smallest node coefficient, and allocates each non-associated data to a corresponding processing node with the smallest node coefficient.
[0092] Among them, for a single processing node, the calculation formula of the node coefficient μ is: μ=lnT α1 +lnL α2 , where T is the data throughput of the processing node, data throughput of the processing node = amount of data processed by the processing node in a single monitoring cycle / time required for data transmission, L is the reference distance. It can be understood that the confirmation method of the reference distance for the association combination and the data acquisition module is different. For a single association combination, the reference distance is the average of the shortest distances from each data acquisition module in the association combination to the processing node. For a single data acquisition module, the reference distance is the shortest distance from the data acquisition module to the processing node. α1 is the weight coefficient corresponding to the data throughput of the processing node, and α2 is the weight coefficient corresponding to the reference distance.
[0093] It can be understood that the weight coefficient corresponding to the data throughput of the processing node and the weight coefficient corresponding to the reference distance are both set by the user according to the importance of the data throughput of the processing node and the reference distance, where α1+α2=1, and one value of α1 and α2 is provided, α1=0.5, α2=0.5.
[0094] Specifically, the second allocating unit allocates the data to be allocated according to the node processing reference value, and for a single piece of data to be allocated, selects a processing node with the largest node processing reference value for allocation.
[0095] The node processing reference value is the maximum amount of data that a single processing node can process in a single detection cycle.
[0096] Specifically, when the redundancy analysis unit performs redundant data analysis on a single processing node, it divides the incoming data into several data blocks of the same size, generates a hash value corresponding to each data block using a secure hash algorithm, and determines the node status based on the maximum difference between the data blocks and the proportion of similar data blocks;
[0097] The node status includes a first node status in which the maximum difference of data blocks is greater than or equal to the preset maximum difference of data blocks or the proportion of similar data blocks is greater than or equal to the preset proportion of similar data blocks, and a second node status in which the maximum difference of data blocks is less than the preset maximum difference of data blocks and the proportion of similar data blocks is less than the preset proportion of similar data blocks.
[0098] The incoming data is divided into several data blocks of equal size. For each data block, a sliding window is used to scan the contents of the data block. The sliding window is initially located at the beginning of the data block and then slides character by character in chronological order to the end of the data block. Each time the sliding window slides, the hash fingerprint of the byte sequence within the window is calculated. All hash fingerprints of the data block are linearly transformed N times, and the maximum hash value of the hash fingerprint after N linear transformations is used as the feature value of the data block. The value of N can be set by the user according to the actual application scenario, and a value of N = 15 is provided.
[0099] For a single data block, the data block difference = the absolute value of the difference between the characteristic value of the data block and the characteristic value of the reference data block / the maximum value between the characteristic value of the data block and the characteristic value of the reference data block; for the data corresponding to a single processing node, the data within a single monitoring cycle is divided into several data blocks of the same size, and the maximum data block difference of the processing node is the maximum value of the data block difference corresponding to each data block corresponding to the processing node.
[0100] The ratio of similar data blocks = the reference value of the number of similar data blocks / the reference value of the number of data blocks.
[0101] The reference value for the number of similar data blocks is the number of data blocks whose data block difference is less than or equal to the second preset data block difference, and the reference value for the number of data blocks is the number of data blocks divided for data of a single processing node in a single monitoring cycle;
[0102] The values of the preset maximum difference of data blocks and the ratio of the preset number of similar data blocks can be set according to the actual application scenario. The greater the user's demand for the accuracy of redundant data identification, the greater the values of the preset maximum difference of data blocks and the ratio of the preset number of similar data blocks. Provided are values of the preset maximum difference of data blocks and the ratio of the preset number of similar data blocks. The preset maximum difference of data blocks is 60%, and the preset ratio of similar data blocks is 50%.
[0103] Specifically, the redundancy adjustment unit determines an adjustment mode according to the node status;
[0104] In the first node state, the redundancy adjustment unit determines that the adjustment method is to reduce the sliding window length;
[0105] In the second node state, the redundancy adjustment unit determines that the adjustment method is to increase the reference value of the number of data blocks;
[0106] The relationship between the reduction value of the sliding window length and the comprehensive difference is negatively correlated;
[0107] The relationship between the increase in the number of data blocks and the comprehensive difference is a positive correlation.
[0108] The comprehensive difference is determined by summing the block differences corresponding to each data block segmented by a single processing node within a single monitoring cycle. The sliding window length is the number of characters that the sliding window can accommodate. Decreasing the sliding window length improves the accuracy of redundant data identification; increasing the reference value for the number of data blocks increases the size of the data blocks, avoiding inaccurate redundant data identification caused by insufficient data blocks.
[0109] Thus far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art may make equivalent changes or substitutions to the relevant technical features, and the technical solutions after such changes or substitutions will fall within the scope of protection of the present invention.
[0110] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that the present invention is susceptible to various modifications and variations. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.
Claims
1. A power detection system based on multidimensional data, characterized in that: include: A data acquisition unit, comprising a plurality of data acquisition modules for performing data acquisition; a data processing unit connected to the data acquisition unit, configured to determine a data state according to a data reference value of the data acquisition unit and a processing reference value of the data processing unit, and to determine a processing method according to the data state; A first allocation unit, connected to the data processing unit, is used to determine, based on a reference quantity ratio, a processing method for allocating associated combined data and non-associated data based on an associated influence coefficient or a node coefficient; a second allocating unit connected to the data processing unit, for executing a processing method for allocating the data to be allocated according to the node processing reference value; a redundancy analysis unit connected to the first allocation unit and the second allocation unit, configured to, when performing redundant data analysis on a single processing node, divide the incoming data into a plurality of data blocks of equal size, generate a hash value corresponding to each data block using a secure hash algorithm, and determine the node status based on the maximum difference between the data blocks and the ratio of the number of similar data blocks; a redundancy adjustment unit connected to the redundancy analysis unit, configured to determine, based on the node state, whether to adjust the sliding window length or the reference value of the number of data blocks, and to perform direct compression or differential compression on the data blocks based on the difference between the data blocks; The data reference value is the sum of the data volumes to be allocated collected by each data processing module within a single monitoring cycle; The processing reference value is the difference between the maximum amount of data that each processing node can process and the amount of data being processed by each processing node in a single monitoring cycle; The reference quantity ratio is equal to the ratio of the sum of the number of data acquisition modules corresponding to each associated combination to the number of all data acquisition modules; The calculation formula of the node coefficient μ is: μ=lnT α1 +lnL α2 , where T is the data throughput of the processing node, data throughput of the processing node = the amount of data processed by the processing node in a single monitoring cycle / the time required to transmit data, L is the reference distance, α1 is the weight coefficient corresponding to the data throughput of the processing node, and α2 is the weight coefficient corresponding to the reference distance; The node processing reference value is the maximum amount of data that a single processing node can process in a single monitoring cycle; The confirmation method of the correlation influence coefficient is: For a single association combination and a single processing node, the calculation formula of the association influence coefficient Ф is: For a single data acquisition module and a single processing node that are not associated, the calculation formula for the association influence coefficient Ф is: Among them, n is the number of data acquisition modules corresponding to a single association combination, m is the number of data acquisition modules corresponding to a single processing node, and r pd is the correlation coefficient between the pth data acquisition module in a single correlation combination and the dth data acquisition module in a single processing node, r d is the correlation coefficient between a single data acquisition module that has not been associated and combined and the dth data acquisition module in a single processing node.
2. The power detection system based on multidimensional data according to claim 1, characterized in that: The data processing unit determines the data state according to the data reference value of the data acquisition unit and the processing reference value of the data processing unit; The data state includes a first data state in which the data reference value is greater than or equal to a preset data reference value or the processing reference value is less than a preset processing reference value, and a second data state in which the data reference value is less than the preset data reference value and the processing reference value is greater than or equal to the preset processing reference value.
3. The power detection system based on multidimensional data according to claim 2, characterized in that: The data processing unit determines a processing mode according to the data status; If the data state is the first data state, the data processing unit determines that the processing method is to determine the associated allocation method according to the reference quantity ratio; If the data state is the second data state, the data processing unit determines that the processing mode is to distribute the data to be distributed according to the node processing reference value.
4. The power detection system based on multidimensional data according to claim 3, characterized in that: The first allocating unit determines an associated allocation method according to a reference quantity ratio; If the reference quantity ratio is greater than or equal to the preset reference quantity ratio, the associated allocation method is to allocate according to the associated influence coefficient; If the reference quantity ratio is less than the preset reference quantity ratio, the association allocation method is to allocate according to the node coefficient; The reference quantity ratio is equal to the ratio of the sum of the numbers of data acquisition modules corresponding to each associated combination to the number of all data acquisition modules.
5. The multi-dimensional data-based power detection system according to claim 4, characterized in that: The first allocation unit sequentially performs association analysis on each data acquisition module in a preset order, and records the data acquisition module that undergoes association analysis as a target data acquisition module. The association analysis includes detecting a correlation coefficient between the target data acquisition module and other data acquisition modules, and uniformly recording the data acquisition modules whose correlation coefficients with the target data acquisition module are greater than or equal to the preset correlation coefficient as a correlation combination with the target data acquisition module. After the association analysis for the target data acquisition module is completed, the association analysis is continued for the data acquisition modules that are not recorded as a correlation combination. The calculation formula of the correlation coefficient r is: Among them, n is the number of data points contained in a single data column, x i and y i are the values of the i-th data point in the two data columns respectively, is x i The average value of the corresponding data column, y i The average value of the corresponding data column, i = 1, 2, 3, ..., n; The default order is from large to small data reference amount.
6. The multi-dimensional data-based power detection system according to claim 5, characterized in that: The first allocation unit performs association allocation on each associated combination data according to an association preset order, wherein the first allocation unit allocates the associated combination data to a processing node with the smallest association influence coefficient in association allocation for a single associated combination data; After the first allocation unit completes the allocation of the associated combined data, it sequentially allocates the non-associated data in an associated manner according to a preset order, and for a single non-associated data, allocates the non-associated data to a processing node with the smallest associated influence coefficient; The associated combination data is the data to be allocated of the data acquisition module corresponding to a single associated combination, and the non-associated data is the data to be allocated of a single data acquisition module that has not been associated with the combination; The preset association order is the order from large to small of the combined data reference amount.
7. The multi-dimensional data-based power detection system according to claim 4, characterized in that: The first allocating unit allocates the associated combined data and the non-associated data according to the node coefficient; The first allocation unit allocates each associated combination data to a corresponding processing node with the smallest node coefficient, and allocates each non-associated data to a corresponding processing node with the smallest node coefficient.
8. The power detection system based on multi-dimensional data according to claim 3, characterized in that: The second allocating unit allocates the data to be allocated according to the node processing reference value, and selects a processing node with the largest node processing reference value for allocation for a single piece of data to be allocated.
9. The power detection system based on multi-dimensional data according to claim 8, characterized in that: When the redundancy analysis unit performs redundant data analysis on a single processing node, it divides the incoming data into several data blocks of the same size, generates a hash value corresponding to each data block using a secure hash algorithm, and determines the node status based on the maximum difference between the data blocks and the proportion of similar data blocks; The node status includes a first node status in which the maximum difference of data blocks is greater than or equal to the preset maximum difference of data blocks or the proportion of similar data blocks is greater than or equal to the preset proportion of similar data blocks, and a second node status in which the maximum difference of data blocks is less than the preset maximum difference of data blocks and the proportion of similar data blocks is less than the preset proportion of similar data blocks.
10. The multi-dimensional data-based power detection system according to claim 9, characterized in that: The redundancy adjustment unit determines an adjustment mode according to a node state; In the first node state, the redundancy adjustment unit determines that the adjustment method is to reduce the sliding window length; In the second node state, the redundancy adjustment unit determines that the adjustment method is to increase the reference value of the number of data blocks; The relationship between the reduction value of the sliding window length and the comprehensive difference is negatively correlated; The relationship between the increase in the number of data blocks and the comprehensive difference is a positive correlation.
Citation Information
Patent Citations
Management methods, devices, equipment and media for multi-source heterogeneous data in the power supply chain
CN116680423B
Data compression method and device
CN111061428A
System and method for screening redundant data in power system
CN118227613A