Distributed storage method for electric power CPS heterogeneous data
By constructing power vectors and clustering algorithms to determine high and low relationship clusters, dynamically allocate the number of virtual nodes, and improving the consistent hashing algorithm, the problem of uneven distribution of heterogeneous data of power CPS is solved, and the utilization rate of storage space and data distribution uniformity are improved.
Patent Information
- Application Number
- CN202510320866.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-03-18
AI Technical Summary
In the existing distributed storage method of power CPS heterogeneous data, the hash value is concentrated due to the similarity of power data, resulting in uneven data distribution, which reduces the storage space utilization and data distribution uniformity.
By constructing power vectors, the similarity and differences between devices are calculated, the clustering algorithm is used to determine high and low relationship clusters, combined with the difference threshold fusion clusters, dynamically allocate the number of virtual nodes, and the consistent hashing algorithm is improved for storage.
Improves the uniformity of data distribution, avoids node overload, reduces management complexity, and improves storage space utilization and scalability.
Smart Images

Figure CN120276669A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of computer data storage, and particularly to a distributed storage method for serving heterogeneous data of power CPS. Background Art
[0002] Power CPS, namely Cyber-Physical Systems for Power, is a complex system that deeply integrates the physical power grid and the information network. It realizes real-time perception, dynamic control, and information services of the power system by integrating computing, communication, and control technologies, so as to improve the efficiency, reliability, and security of the power system.
[0003] Heterogeneous data refers to data from different sources, such as various links of power generation, transmission, transformation, distribution, power consumption, and dispatching, with different structures and characteristics. These data may include structured data (such as tabular data in a database), semi-structured data (such as data in XML or JSON format), and unstructured data (such as text files, images, videos, etc.).
[0004] For traditional technologies of power CPS heterogeneous data, distributed storage is generally carried out through the consistent hashing algorithm. In order to ensure the uniformity of the distribution of heterogeneous data, the consistent hashing algorithm is solved by introducing virtual nodes. However, due to the large amount of similarity in power CPS heterogeneous data, such as the operation changes of power in a line, when calculating the hash value of heterogeneous data, there may be a possibility that the hash value is in one area, resulting in a relatively concentrated data distribution. And too many or too few virtual nodes in the algorithm reduce the utilization rate of storage space and the uniformity of data distribution. Summary of the Invention
[0005] In order to solve the above technical problems, this application provides a distributed storage method for serving heterogeneous data of power CPS to solve the existing problems.
[0006] The distributed storage method for serving heterogeneous data of power CPS in this application adopts the following technical solutions:
[0007] An embodiment of this application provides a distributed storage method for serving heterogeneous data of power CPS, and this method includes the following steps:
[0008] In power CPS, collect the power data of each front-end device at all acquisition moments within a preset time period, and form a power vector from the binary data of all power data;
[0009] Based on the difference in the length of the power vector between each front-end device and any other front-end device, and the similarity of the power vector, determine the power similarity between each front-end device and any other front-end device;
[0010] Based on the power similarity between each front-end device and all the other front-end devices, a clustering algorithm is used to determine the relative high and low relationship clusters of each front-end device; based on the distribution of all elements in the relative high relationship cluster of each front-end device, the high relationship value of each front-end device is determined; based on the dispersion degree of the high relationship values of all front-end devices, the heterogeneous data dispersion degree of the power CPS is determined.
[0011] Based on the difference between the maximum and minimum values among the high relationship values of all front-end devices, and the heterogeneous data dispersion degree, the difference threshold of the power CPS is determined; based on the average distribution of all elements in the relative low relationship cluster of each front-end device, and the high relationship value, the difference degree of each front-end device is determined, and the relative high and low relationship clusters are fused in combination with the difference threshold to obtain the power high relationship cluster of each front-end device.
[0012] Based on the total number of elements in the power high relationship cluster, the dynamic virtual node number of the power CPS is determined, and the power data of all front-end devices of the power CPS are stored.
[0013] Preferably, the method for determining the power similarity between each front-end device and any other front-end device is as follows:
[0014] In the power CPS, compare the lengths of the power vectors of each front-end device and the power vectors of any other front-end device, obtain each sub-power vector composed of elements that are equal in length and continuous with the power vector with the smallest length among the power vectors with the largest length, calculate the similarity between the power vector with the smallest length and all sub-power vectors, and count the maximum similarity.
[0015] The ratio of the maximum similarity to the length difference of the power vectors is used as the power similarity between each front-end device and any other front-end device.
[0016] Preferably, the method for determining the relative high and low relationship clusters of each front-end device is as follows:
[0017] Use a clustering algorithm to cluster the power similarities between each front-end device and all the other front-end devices to obtain two clustering clusters. Denote the clustering cluster where the clustering center with the maximum power similarity is located as the relative high relationship cluster, and the other clustering cluster as the relative low relationship cluster.
[0018] Preferably, the high relationship value of each front-end device is the average value of the power similarities corresponding to all front-end devices in the relative high relationship cluster of each front-end device.
[0019] Preferably, the expression for the heterogeneity data dispersion degree of the power CPS is: k = 1 - exp(-V); where k represents the heterogeneity data dispersion degree of the power CPS; V represents the dispersion degree of the high-relationship mean values of all front-end devices in the power CPS; exp() represents the exponential function with the natural constant as the base.
[0020] Preferably, the expression for the difference threshold of the power CPS is: T = E × k; T represents the difference threshold of the power CPS; E represents the range of the high-relationship values of all front-end devices in the power CPS.
[0021] Preferably, the difference degree of each front-end device is the absolute value of the difference between the mean value of all elements in the relatively low-relationship cluster of each front-end device and the high-relationship value.
[0022] Preferably, the method for determining the power high-relationship cluster of each front-end device is as follows:
[0023] If the difference degree is less than the difference threshold, the cluster obtained by merging the relatively low-relationship cluster and the relatively high-relationship cluster is used as the power high-relationship cluster of each front-end device; otherwise, if the difference degree is greater than or equal to the difference threshold, the relatively high-relationship cluster is used as the power high-relationship cluster of each front-end device.
[0024] Preferably, the expression for the dynamic virtual node number of the power CPS is: where Num represents the dynamic virtual node number of the power CPS; N represents the preset virtual node number; n represents the pre-designed number of computers; m represents the number of all front-end devices in the power CPS; LS i represents the total number of elements in the power high-relationship cluster of front-end device i; ceil() represents the ceiling function; norm() represents the normalization function.
[0025] Preferably, the storage of the power data of all front-end devices of the power CPS includes:
[0026] Taking the power vectors of all front-end devices in the power CPS as the input of the consistent hashing algorithm, where the number of virtual nodes of the consistent hashing algorithm is set as the dynamic virtual node number of the power CPS, and outputting all hash values and storing them in the computer.
[0027] This application has at least the following beneficial effects:
[0028] This application determines the power similarity between each front-end device and any other front-end device by constructing power vectors, the difference in the length of the power vectors between each front-end device and any other front-end device, and the similarity of the power vectors. The beneficial effect is that it can solve the problem of uneven load caused by many similar data clustering together in the hash space to form data hotspots, and improve the uniformity of data distribution. This application uses a clustering algorithm to determine the relative high and low relationship clusters of each front-end device; based on the distribution of all elements in the relative high relationship cluster of each front-end device, the high relationship value is determined; based on the dispersion degree of the high relationship values of all front-end devices, the heterogeneous data dispersion degree of the power CPS is determined; based on the difference between the maximum and minimum values among the high relationship values of all front-end devices and the heterogeneous data dispersion degree, combined with the average distribution of all elements in the relative low relationship cluster of each front-end device and the high relationship value, the difference degree of each front-end device is determined, and the relative high and low relationship clusters are fused in combination with the difference threshold to obtain the power high relationship cluster of each front-end device. The beneficial effect is that it can disperse the power data to more physical nodes, thus avoiding the overload of a single node. This application determines the dynamic virtual node number of the power CPS based on the total number of elements in the power high relationship cluster and the total number of elements in the relative high relationship cluster, and stores the power data of all front-end devices of the power CPS. The beneficial effect is that it can reduce the complexity of power CPS management, improve the flexibility during the expansion of the power CPS and the uniformity of data distribution. This application improves the consistency hashing algorithm by dynamically allocating the number of virtual nodes, improving the utilization rate of storage space and the uniformity of data distribution. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0030] Figure 1 It is a flowchart of the steps of a distributed storage method for serving heterogeneous data of a power CPS provided by an embodiment of the present application;
[0031] Figure 2 It is a schematic diagram of the power similarity extraction process provided by an embodiment of the present application;
[0032] Figure 3 It is a schematic diagram of the process of obtaining the power high relationship cluster provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0033] To further elaborate on the technical means and effects adopted by this application to achieve the intended invention purpose, the following combines the accompanying drawings and preferred embodiments to detail the specific implementation manner, structure, features, and effects of the distributed storage method for power CPS heterogeneous data proposed according to this application. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.
[0034] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs.
[0035] The following specifically describes the specific solution of the distributed storage method for power CPS heterogeneous data provided by this application in combination with the accompanying drawings.
[0036] A distributed storage method for power CPS heterogeneous data provided by an embodiment of this application specifically provides the following distributed storage method for power CPS heterogeneous data. Please refer to Figure 1 , and the method includes the following steps:
[0037] Step S1: In the power CPS, collect the power data of each front-end device at all acquisition times within a preset duration, and form a power vector from the binary data of all the power data.
[0038] Heterogeneous data refers to data with different data types, structures, and formats. The data may come from different sources and exist in different formats. In this embodiment, the power data from different front-end devices in the power CPS belongs to heterogeneous data. For the power CPS, (such as various sensors, video devices, etc.) collect the power data of each front-end device at all acquisition times within a preset duration. For example, the power data collected by a current sensor is current data, and the power data collected by a voltage sensor is voltage data. In addition, it also includes power data such as the patch log information generated when the computer control system maintains, updates, or repairs the front-end device, and the operation information (i.e., power fault information) when monitoring the device.
[0039] Furthermore, through wireless or wired network transmission, all the power data of each front-end device is transmitted to the central server of the power CPS, and the central server converts the power data into binary data through a built-in program. The binary data of each front-end device forms the power vector of each front-end device.
[0040] Step S2: Based on the difference in the length of the power vectors between each front-end device and any other front-end device, and the similarity of the power vectors, determine the power similarity between each front-end device and any other front-end device.
[0041] For the distributed storage of heterogeneous data in the power CPS, in this embodiment, the power data includes current, voltage, service information, repair log information, and power failure information. For each type of data in the power data, there is a large amount of repetition in the data generated during normal operation. For example, for power failure information, under normal circumstances, when the power failure information indicates normal operation and there are no abnormalities, the power failure information is roughly the same except for the operation time.
[0042] Before distributed storage, the power data is usually located in a centralized database or data warehouse, and this data may come from different front-end devices, such as sensors, smart meters, or other data acquisition systems. For the data collected by the same front-end or the same type of front-end devices, there may be a high degree of similarity. For example, current and voltage data usually fluctuate around a certain expected normal range. Therefore, for different front-end devices in the power system, due to the same operation mode, there may also be a large number of similar characteristics in the collected power data.
[0043] Due to the possible differences in the models and operating frequencies of different front-end devices, in the power CPS, the lengths of the power data collected at the same moment may also be different (note: the length of the power data refers to the length of its binary data), but there may be a large number of similar characteristics between the power data. Thus, based on the difference in the power vector lengths between each front-end device and any other front-end device, as well as the similarity of the power vectors, the power similarity between each front-end device and any other front-end device is determined. Specifically:
[0044] (1) In the power CPS, compare the lengths of the power vectors of each front-end device with the power vectors of any other front-end device. In the power vector with the maximum length, obtain each sub-power vector composed of elements that are equal in length and continuous with the power vector with the minimum length. Calculate the similarity between the power vector with the minimum length and all sub-power vectors, and count the maximum similarity.
[0045] (2) Further, analyze the difference in the lengths of the power vectors between each front-end device and any other front-end device, and use the ratio of the maximum similarity to the difference in the lengths of the power vectors as the power similarity between each front-end device and any other front-end device.
[0046] It should be noted that there are many methods to measure the similarity between vectors. In this embodiment, the cosine similarity between the power vector with the minimum length and each sub-power vector is calculated to measure the similarity between any two power vectors. Implementers can also use other methods to measure the similarity between vectors, such as the reciprocal of the Euclidean distance. There is no special limitation on the selection of the method for measuring the similarity between vectors in this embodiment.
[0047] Among them, the calculation process of cosine similarity is a well-known technology, and its specific calculation process will not be elaborated here.
[0048] Furthermore, it can be understood from the power similarity between each front-end device and any other front-end device that in a power CPS, the greater the similarity between the power data collected by the front-end devices, the higher the consistency of these power data in terms of characteristics and patterns, the greater the correlation between the power vector with the minimum length and each sub-power vector, and the more similar the power data are, which also indicates that the difference in the lengths of the power vectors is smaller, and thus the power similarity between the front-end devices is greater; conversely, in a power CPS, the smaller the similarity between the power data collected by the front-end devices, the greater the difference in the characteristics and patterns of these power data, the smaller the correlation between the power vector with the minimum length and each sub-power vector, and the greater the difference between the power data, which also indicates that the difference in the lengths of the power vectors is greater, and thus the power similarity between the front-end devices is smaller.
[0049] Preferably, the schematic diagram of the power similarity extraction process provided in this embodiment is as Figure 2 shown.
[0050] Step S3: Based on the power similarity between each front-end device and all other front-end devices, use a clustering algorithm to determine the relative high and low relationship clusters of each front-end device; based on the distribution of all elements in the relative high relationship cluster of each front-end device, determine the high relationship value of each front-end device; based on the dispersion degree of the high relationship values of all front-end devices, determine the heterogeneous data dispersion degree of the power CPS; based on the difference between the maximum and minimum values among the high relationship values of all front-end devices and the heterogeneous data dispersion degree, determine the difference threshold of the power CPS; based on the average distribution of all elements in the relative low relationship cluster of each front-end device and the high relationship value, determine the difference degree of each front-end device, and fuse the relative high and low relationship clusters in combination with the difference threshold to obtain the power high relationship cluster of each front-end device.
[0051] When the similarity degree between the power data of a front-end device and the power data of other front-end devices is relatively high, the power data similarity has a certain aggregation characteristic, that is, it is distributed within a numerical range, and the numerical range is relatively small. Due to the characteristics of power data similarity calculation, the power data similarity between power data with a relatively high similarity degree is distributed on the side with a larger value. When a fault occurs in the front-end device of a power CPS or abnormal power data appears in the collected object, compared with normal power data, the similarity degree between the abnormal power data and the normal power data is relatively low, so the similarity between the power data is relatively small.
[0052] Therefore, in order to exclude the interference caused by abnormal data generated by front-end device failures and classify front-end devices with a high degree of similarity in power data into one category by analyzing the similarity of power data between different front-end devices, specifically:
[0053] (1) Cluster the power similarity between the power vector of each front-end device and the power vectors of all other front-end devices. Set the number of clustering clusters to 2, and output two clustering clusters. Denote the clustering cluster where the clustering center with the maximum power similarity is located as the relatively high relationship cluster, and the other clustering cluster as the relatively low relationship cluster.
[0054] It should be noted that there are many common clustering algorithms. In this embodiment, the k-means clustering algorithm is used to cluster the power similarity. Implementers can also use other clustering methods such as fuzzy C-means clustering. There is no special limitation on the selection of clustering algorithms in this embodiment.
[0055] (2) For power data with a relatively high degree of similarity, they all exist in the relatively high relationship cluster. However, it is possible that the similarity degree between the power data and other types of different power data is relatively high. The elements in the relatively high relationship cluster and the elements in the relatively low relationship cluster belong to the same category, and it is unreasonable for the clustering algorithm to divide them into two categories. Therefore, it is necessary to judge the relatively low relationship cluster. The specific process is as follows:
[0056] Calculate the mean value of all elements in the relatively high relationship cluster of each front-end device, and denote it as the high relationship value of each front-end device;
[0057] The expression of the heterogeneity data dispersion k of the power CPS is: k = 1 - exp(-V); where V represents the dispersion degree of the high relationship values of all front-end devices in the power CPS; exp() represents the exponential function with the natural constant as the base.
[0058] It should be noted that there are many methods to measure the dispersion degree of a set of data. In this embodiment, the dispersion coefficient of the high relationship mean value of all front-end devices is calculated to measure the dispersion degree of the high relationship mean value. Implementers can also use other methods to measure the dispersion degree of data, such as variance and standard deviation. There is no special limitation on the selection of methods to measure the dispersion degree of data in this embodiment.
[0059] Among them, the calculation process of the dispersion coefficient is a well-known technology, and its specific calculation steps will not be elaborated here.
[0060] Furthermore, it can be understood from the heterogeneity data dispersion degree of the power CPS that when the dispersion degree of the high relationship values of all front-end devices is smaller, it indicates that the similarity degree between power data is greater, that is, the heterogeneity data dispersion degree of the power CPS is smaller; on the contrary, when the dispersion degree of the high relationship values of all front-end devices is greater, it indicates that the difference between power data is greater, that is, the heterogeneity data dispersion degree of the power CPS is greater.
[0061] (3) Furthermore, in order to classify front-end devices with relatively large similarity into one category, based on the difference between the maximum and minimum values among the high relationship values of all front-end devices and the heterogeneity data dispersion degree, a difference threshold of the power CPS is determined, specifically:
[0062] The expression of the difference threshold T of the power CPS is: T = E × k; E represents the range of the high relationship values of all front-end devices in the power CPS.
[0063] It can be understood from the difference threshold of the power CPS that when the range of the high relationship values of all front-end devices is greater, it indicates that the similarity degree between different front-end devices in the power CPS is lower, the correlation degree between power data is lower, and the heterogeneity difference dispersion degree of the power CPS is greater, then the difference threshold of the power CPS is greater, indicating that the difference between power data of different front-end devices is greater; on the contrary, when the range of the high relationship values of all front-end devices is smaller, it indicates that the similarity degree between different front-end devices in the power CPS is higher, the correlation degree between power data is higher, and the heterogeneity difference dispersion degree of the power CPS is smaller, then the difference threshold of the power CPS is smaller, indicating that the difference between power data of different front-end devices is smaller.
[0064] (4) Furthermore, based on the average distribution of all elements in the relatively low relationship cluster of each front-end device and the high relationship value, the difference degree of each front-end device is determined, specifically:
[0065] The difference degree of each front-end device is the absolute value of the difference between the mean value of all elements in the relatively low relationship cluster of each front-end device and the high relationship value.
[0066] It can be understood from the difference degree of each front-end device that when the similarity degree of the power data between the front-end devices is relatively high, the difference between the relatively low relationship cluster and the relatively high relationship cluster of the front-end devices is smaller, that is, the difference between the mean value of all elements in the relatively low relationship cluster and the high relationship value is smaller, indicating that the similarity between the elements in the relatively low relationship cluster and the elements in the relatively high relationship cluster is higher, and the relatively low relationship cluster and the relatively high relationship cluster should be merged into one cluster; on the contrary, when the similarity degree of the power data between the front-end devices is relatively low, the difference between the relatively low relationship cluster and the relatively high relationship cluster of the front-end devices is larger, that is, the difference between the mean value of all elements in the relatively low relationship cluster and the high relationship value is larger, indicating that the similarity between the elements in the relatively low relationship cluster and the elements in the relatively high relationship cluster is lower, and only the relatively high relationship cluster should be retained.
[0067] (5) Further, based on the difference degree and the difference threshold, a power high relationship cluster is constructed. Specifically, if the difference degree is less than the difference threshold, it indicates that the similarity between the elements in the relatively low relationship cluster and the elements in the relatively high relationship cluster is higher, and the cluster obtained by merging the relatively low relationship cluster and the relatively high relationship cluster is used as the power high relationship cluster of each front-end device; on the contrary, if the difference degree is greater than or equal to the difference threshold, the relatively high relationship cluster is used as the power high relationship cluster of the corresponding front-end device.
[0068] Preferably, the schematic diagram of the power high relationship cluster acquisition process provided in this embodiment is as Figure 3 shown.
[0069] Step S4: Based on the total number of elements in the power high relationship cluster of each front-end device, determine the dynamic virtual node number of the power CPS, and store the power data of all front-end devices of the power CPS.
[0070] For the power high relationship cluster, the number of elements in the cluster represents the number of front-end devices similar to the corresponding front-end device. When the number of similar front-end devices is larger, a larger number of virtual nodes should be used in the process of adopting the consistent hashing algorithm to improve the distribution uniformity of the heterogeneous data of the power CPS.
[0071] Therefore, based on the total number of elements in the power high relationship cluster and the total number of elements in the relatively high relationship cluster, the dynamic virtual node number of the power CPS is determined. Specifically:
[0072] The expression of the dynamic virtual node number Num of the power CPS is: In the formula, N represents the preset virtual node number; n represents the preset number of computers; m represents the number of all front-end devices in the power CPS; LS i represents the total number of elements in the power high relationship cluster of the front-end device i; ceil() represents the ceiling function; norm() represents the normalization function.
[0073] It should be noted that the values of the preset number of virtual nodes and the preset number of computers are both artificially set. In this embodiment, the value of the preset number of virtual nodes is 160, and the value of the preset number of computers is 5. Implementers can also set them by themselves according to specific situations, and this embodiment does not make special restrictions.
[0074] Furthermore, the power vectors of all front-end devices in the power CPS are used as the input of the consistent hashing algorithm. The number of virtual nodes of the consistent hashing algorithm is set as the dynamic virtual node number of the power CPS, and all the output hash values are stored in the computers according to the size of the storage capacity of the computers.
[0075] Among them, the consistent hashing algorithm is a well-known technology, and the specific process of obtaining the hash value will not be elaborated here. It should be noted that: the above sequence of the embodiments of the present application is only for description and does not represent the advantages or disadvantages of the embodiments. And the above describes specific embodiments of this specification. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0076] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. The key point of each embodiment is to illustrate the differences from other embodiments.
[0077] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; modifying the technical solutions recorded in the foregoing embodiments, or equivalently replacing some of the technical features, does not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of each embodiment of the present application, and should all be included in the protection scope of the present application.
Claims
1. A distributed storage method for heterogeneous data of power CPS, characterized in that The method includes the following steps: In the power CPS, collect the power data of each front-end device at all acquisition moments within a preset time period, and form a power vector by combining the binary data of all the power data. Based on the difference in the lengths of the power vectors between each front-end device and any other front-end device, and the similarity of the power vectors, determine the power similarity between each front-end device and any other front-end device. Based on the power similarity between each front-end device and all other front-end devices, use a clustering algorithm to determine the relative high and low relationship clusters of each front-end device; based on the distribution of all elements in the relative high relationship cluster of each front-end device, determine the high relationship value of each front-end device; based on the dispersion degree of the high relationship values of all front-end devices, determine the heterogeneous data dispersion degree of the power CPS. Based on the difference between the maximum and minimum values among the high relationship values of all front-end devices, and the heterogeneous data dispersion degree, determine the difference threshold of the power CPS; based on the average distribution of all elements in the relative low relationship cluster of each front-end device, and the high relationship value, determine the difference degree of each front-end device, and fuse the relative high and low relationship clusters in combination with the difference threshold to obtain the power high relationship cluster of each front-end device. Based on the total number of elements in the power high relationship cluster, determine the number of dynamic virtual nodes of the power CPS, and store the power data of all front-end devices of the power CPS.
2. The distributed storage method for heterogeneous data of power CPS according to claim 1, wherein The method for determining the power similarity between each front-end device and any other front-end device is as follows: In the power CPS, compare the lengths of the power vectors of each front-end device with the power vectors of any other front-end device, and obtain each sub-power vector composed of elements that are equal in length and continuous to the power vector with the smallest length from the power vector with the largest length. Calculate the similarity between the power vector with the smallest length and all sub-power vectors, and count the maximum similarity. Take the ratio of the maximum similarity to the length difference of the power vectors as the power similarity between each front-end device and any other front-end device.
3. The distributed storage method for heterogeneous data of power CPS according to claim 1, characterized in that, The method for determining the relative high and low relationship clusters of each front-end device is as follows: Use a clustering algorithm to cluster the power similarity between each front-end device and all other front-end devices to obtain two clustering clusters. Denote the clustering cluster where the clustering center with the maximum power similarity is located as the relative high relationship cluster, and the other clustering cluster as the relative low relationship cluster.
4. The distributed storage method for heterogeneous data of power CPS according to claim 1, characterized in that The high relationship value of each front-end device is the average value of the power similarities corresponding to all front-end devices in the relative high relationship cluster of each front-end device.
5. The distributed storage method for heterogeneous data of power CPS according to claim 1, characterized in that The expression for the heterogeneous data dispersion degree of the power CPS is: k = 1 - exp(-V); where k represents the heterogeneous data dispersion degree of the power CPS; V represents the dispersion degree of the high relationship means of all front-end devices in the power CPS; exp() represents the exponential function with the natural constant as the base.
6. The distributed storage method for heterogeneous data of power CPS according to claim 5, characterized in that, The expression for the difference threshold of the power CPS is: T = E × k; T represents the difference threshold of the power CPS; E represents the range of the high relationship values of all front-end devices in the power CPS.
7. The distributed storage method for heterogeneous data of power CPS according to claim 1, characterized in that, The difference degree of each front-end device is the absolute value of the difference between the mean value of all elements in the relatively low relationship cluster of each front-end device and the high relationship value.
8. The distributed storage method for heterogeneous data of power CPS according to claim 1, characterized in that The method for determining the high power relationship cluster of each front-end device is as follows: If the difference degree is less than the difference threshold, the cluster obtained by merging the relatively low relationship cluster and the relatively high relationship cluster is used as the high power relationship cluster of each front-end device; On the contrary, if the difference degree is greater than or equal to the difference threshold, the relatively high relationship cluster is used as the high power relationship cluster of each front-end device.
9. The distributed storage method for heterogeneous data of power CPS according to claim 1, characterized in that The expression for the number of dynamic virtual nodes of the power CPS is as follows: In the formula, Num represents the number of dynamic virtual nodes of the power CPS; N represents the preset number of virtual nodes; n represents the preset number of computers; m represents the number of all front-end devices in the power CPS; LS i represents the total number of elements in the power high-relationship cluster of the front-end device i; ceil() represents the ceiling function; norm() represents the normalization function.
10. The distributed storage method for heterogeneous data of power CPS according to claim 1, characterized in that, The storage of the power data of all front-end devices in the power CPS includes: Taking the power vectors of all front-end devices in the power CPS as the input of the consistent hashing algorithm, where the number of virtual nodes of the consistent hashing algorithm is set as the dynamic virtual node number of the power CPS, and outputting all hash values and storing them in the computer.
Citation Information
Patent Citations
Edge cloud dynamic scheduling method for multistage space-time analysis task
CN118484287A
Digital power service sensitive information desensitization processing method and system
CN119089506A
Method for processing computation in heterogeneous cluster system, and heterogeneous cluster system for performing the same
KR1020180084682A
Image retrieval method and apparatus, and computer-readable storage medium
WO2020199773A1
Cited By
Modeling simulation analysis method for interlocking fault of power CPS and related equipment
CN120914775A