Data management method and system for distributed data architecture
By encoding spatiotemporal feature of multimodal data and building multi-level correlation bitmaps, the problems of spatiotemporal correlation fragmentation and cross-modal logical correlation fracture in distributed data architecture are solved, and collaborative calculations between edge nodes and central nodes are realized, which improves the efficiency and accuracy of data correlation analysis.
Patent Information
- Application Number
- CN202510828415.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-20
AI Technical Summary
Under the distributed data architecture, it is difficult to achieve accurate correlation analysis due to the fragmentation of space-time associations and cross-modal logical associations caused by edge node shard storage.
By encoding the multimodal data generated by the device through spatiotemporal feature code, a multi-level correlation bitmap is constructed, and physical storage locations, cross-modal logical associations within the device and topological relationships between devices are integrated to realize collaborative calculations between edge physical nodes and central governance nodes.
The correlation analysis efficiency of multimodal data is improved, and the data correlation fracture and analysis distortion problems in traditional architectures are solved, and efficient support for equipment failure prediction, abnormal propagation and positioning, and system-level collaborative decision-making are achieved.
Smart Images

Figure CN120336332A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and more specifically, to a data governance method and system for a distributed data architecture. Background Art
[0002] In today's digital age, fields such as industrial Internet and Internet of Things are developing rapidly, and the amount of multimodal data generated by devices is increasing explosively. Under the distributed data architecture, how to efficiently store and analyze this data has become a key problem to be solved urgently.
[0003] In the prior art, a Chinese patent application with the publication number CN116257509A discloses a method and device for realizing data governance, which governs data by means of data access, integration, classification, and constructing a data model, etc., and improves the value and utilization rate of data assets. However, this prior art does not address the problem of spatio-temporal correlation fragmentation caused by edge node sharding storage of multimodal data under a distributed architecture. In actual application scenarios, when data is scattered and stored in multiple edge nodes, the correlation of data in the time and space dimensions is broken, and it is difficult to achieve accurate correlation analysis.
[0004] A Chinese patent application with the publication number CN119829551A discloses a method for constructing a data index for distributed data storage, which mainly focuses on the index construction of a distributed data storage system, and uses high-order graph theory, consistent hashing algorithm, etc. to optimize the load balancing of storage nodes and data storage strategies to ensure the availability, scalability, and reliability of the system. But this prior art does not consider the problem of cross-modal logical association breakage. In a multimodal data environment, such as in an intelligent factory where there are various modalities such as image data and sensor data of devices at the same time, the logical association between different modality data is crucial for in-depth analysis of the operating state of the devices.
[0005] However, when dealing with multimodal data in the current distributed data architecture, it faces severe challenges of spatio-temporal correlation fragmentation, cross-modal logical association breakage, and technical fragmentation between data storage and association calculation. Summary of the Invention
[0006] The present invention is applicable to distributed monitoring scenarios with dense devices in industrial Internet of Things, such as equipment monitoring on intelligent factory production lines, status management of wind power equipment clusters, etc. In these scenarios, multi-source heterogeneous devices generate a large amount of multimodal data in real time, and efficient data governance is required to achieve early warning of faults, positioning of abnormal propagation, and system-level collaborative decision-making.
[0007] To overcome the above-mentioned defects of the prior art, the present invention provides a data governance method and system for a distributed data architecture. By constructing a multi-level associated bitmap through spatio-temporal feature encoding (hash spatio-temporal anchoring and dynamic time window slicing), and integrating the physical storage location, cross-modal logical association within the device, and topological relationship between devices, a three-dimensional association network covering the physical layer, logical layer, and topological layer is formed. This solution realizes the efficient cooperation between edge physical node sharding storage and central governance node global computing, solves the problems of data association breakage and analysis distortion in the traditional architecture, significantly improves the association analysis efficiency of multi-modal data, provides underlying governance support for device fault prediction, abnormal propagation positioning, and system-level collaborative decision-making, and breaks through the technical gap between storage and analysis in the distributed architecture.
[0008] To achieve the above object, the present invention provides the following technical solutions:
[0009] A data governance method for a distributed data architecture, comprising:
[0010] Performing spatio-temporal feature encoding on multi-modal data generated by a device to obtain device data carrying spatio-temporal feature encoding; the spatio-temporal feature encoding includes hash spatio-temporal anchoring and dynamic time window slicing;
[0011] Based on the device data carrying spatio-temporal feature encoding, constructing a multi-level associated bitmap;
[0012] Based on the constructed multi-level associated bitmap, realizing collaborative computing between edge physical nodes and central governance nodes of the distributed data architecture.
[0013] Further, the device data carrying spatio-temporal feature encoding is written in the form of data blocks, and the data blocks at least include vibration data blocks and thermal imaging data blocks.
[0014] Further, the method for performing the hash spatio-temporal anchoring includes:
[0015] Obtaining a device ID, generating a 128-bit hash value according to the device ID, and recording it as the device hash value;
[0016] Intercepting the lower 8 bits of the device hash value as the physical node mapping identifier, and recording it as the node mapping code;
[0017] Marking the multi-modal data generated by the same device within a preset period T with the same node mapping code; the period T is the basic time unit for data governance.
[0018] Further, the method for performing the dynamic time window slicing includes:
[0019] Taking the device's first activation moment as the time origin, dividing the device's operating cycle into multiple sliding time windows according to the preset period T;
[0020] When writing the data blocks generated by the device, calculate the deviation Δτ between the timestamp of the data block and the end of the current time window. Based on the deviation Δτ, determine the time window to which the data block belongs and perform dynamic data sharding.
[0021] Further, the multi-level association bitmap includes a first-level bitmap, a second-level bitmap, and a third-level bitmap; the first-level bitmap is a physical association bitmap, the second-level bitmap is a logical association bitmap, and the third-level bitmap is a topological association bitmap.
[0022] Further, the method for constructing the first-level bitmap includes:
[0023] Create a two-dimensional bit matrix for the device data carrying the spatio-temporal feature encoding of each device, denoted as the device association matrix; according to the arrival time window of the device data carrying the spatio-temporal feature encoding, mark the association status at the corresponding row and column positions in the device association matrix to form the first-level bitmap.
[0024] Further, the method for constructing the second-level bitmap includes:
[0025] Perform a logical association determination on the vibration data blocks and thermal imaging data blocks within the same time window;
[0026] If the result of the logical association determination is that there is a logical association, create a logical association channel within the corresponding time window row in the device association matrix;
[0027] Calculate the spatio-temporal coupling degree of the logical association channel, and mark the corresponding bit positions of the logical association channel according to the spatio-temporal coupling degree to generate the second-level bitmap.
[0028] Further, the method for performing a logical association determination on the vibration data blocks and thermal imaging data blocks within the same time window includes:
[0029] Extract the vibration frequency feature of the vibration data block and the temperature difference feature of the thermal imaging data block;
[0030] When writing the data block, if the vibration frequency feature of the vibration data block within the same time window exceeds the preset frequency threshold and the temperature difference feature of the thermal imaging data block exceeds the preset temperature difference threshold, it is determined that there is a logical association between the vibration data block and the thermal imaging data block within the time window.
[0031] Further, the method for realizing the collaborative calculation between the edge physical node and the central governance node of the distributed data architecture includes:
[0032] Based on the second-level bitmap and the third-level bitmap, construct a global anomaly propagation graph;
[0033] Based on the global anomaly propagation graph, implement distributed root cause location through querying the graph database;
[0034] Based on the root cause localization result, determine the multimodal data and logical association channels corresponding to the root cause localization result in the first-level bitmap and the second-level bitmap, and converge the data distributed on the edge physical nodes to the central governance node;
[0035] The central governance node receives the converged data and performs unified association calculations.
[0036] A data governance system for a distributed data architecture, which is used to implement the data governance method for a distributed data architecture described above. The system includes:
[0037] Data encoding module: used to perform spatio-temporal feature encoding on the multimodal data generated by the device to obtain device data carrying spatio-temporal feature encoding; the spatio-temporal feature encoding includes hash spatio-temporal anchoring and dynamic time window slicing;
[0038] Bitmap construction module: based on the device data carrying spatio-temporal feature encoding, construct a multi-level association bitmap; the multi-level association bitmap includes a first-level bitmap, a second-level bitmap, and a third-level bitmap; the first-level bitmap is a physical association bitmap, the second-level bitmap is a logical association bitmap, and the third-level bitmap is a topological association bitmap;
[0039] Collaborative computing module: based on the constructed multi-level association bitmap, realize the collaborative computing between the edge physical nodes and the central governance node of the distributed data architecture.
[0040] Compared with the prior art, the beneficial effects of the present invention are:
[0041] The present invention effectively solves the problems of spatio-temporal association fragmentation and cross-modal logical association breakage caused by data sharding storage of edge nodes in a distributed architecture by performing spatio-temporal feature encoding on multimodal data and constructing a multi-level association bitmap. This method integrates multi-dimensional information such as physical storage location, cross-modal logical association within the device, and topological relationship between devices, constructs a three-dimensional association network covering the physical layer, logical layer, and topological layer, and breaks through the technical fragmentation of data storage and association analysis in traditional distributed architectures. Based on the constructed multi-level association bitmap, the present invention realizes the efficient collaboration between the data sharding storage of edge physical nodes and the global association calculation of the central governance node, ensuring both the load balance of edge nodes and the globality and accuracy of association analysis. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0043] Figure 1 It is the method flow chart of a data governance method for a distributed data architecture in the present invention;
[0044] Figure 2 It is the method flow chart for performing hash spatio-temporal anchoring on multi-modal data generated by devices in a data governance method for a distributed data architecture of the present invention;
[0045] Figure 3 It is the principle flow chart for judging the time window to which data belongs in the embodiment of the present invention;
[0046] Figure 4 It is the structural schematic diagram of a multi-level associated bitmap in the embodiment of the present invention;
[0047] Figure 5 It is the method flow chart for constructing a global anomaly propagation graph in a data governance method for a distributed data architecture of the present invention;
[0048] Figure 6 It is the functional module diagram of a data governance system for a distributed data architecture in the present invention. Detailed implementation manners
[0049] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0050] Embodiment 1:
[0051] Please refer to Figure 1 as shown, this embodiment provides a data governance method for a distributed data architecture, including:
[0052] Step S1000, performing spatio-temporal feature encoding on multi-modal data generated by devices to obtain device data carrying spatio-temporal feature encoding; the spatio-temporal feature encoding includes hash spatio-temporal anchoring and dynamic time window slicing;
[0053] Further, step S1000 includes:
[0054] Step S1100, performing hash spatio-temporal anchoring on multi-modal data generated by devices;
[0055] Further, as Figure 2 shown, step S1100 includes:
[0056] Step S1110, obtaining the device ID, generating a 128-bit hash value according to the device ID, and recording it as the device hash value;
[0057] Step S1120, intercepting the lower 8 bits of the device hash value as a physical node mapping identifier, recorded as a node mapping code;
[0058] Step S1130, marking the multimodal data generated by the same device within a preset period T as the same node mapping code; the period T is the basic time unit for data governance.
[0059] Specifically, the purpose of step S1100 is to solve the problems of data physical node mapping confusion and time dimension correlation break caused by identification differences and storage fragmentation in multi-source heterogeneous devices in the Industrial Internet through standardized hash processing of device identification and data clustering within a time period.
[0060] As the unique identifier of the device, the device ID has common forms such as IMEI, MAC address, etc., and its format and length vary greatly. In order to achieve unified data processing and subsequent efficient operation, a hash operation is used to convert device IDs of different formats and lengths into a fixed-length 128-bit hash value. The hash operation process is to use a specific hash algorithm to take the device ID as input, and after a series of mathematical transformations, output a fixed-length hash value. Taking a certain industrial equipment as an example, its IMEI number is "IMEI_00123456789". After being processed by the hash algorithm, a 128-bit hash value "a1b2c3d4e5f6..." is obtained. The purpose of this is to convert the uniqueness and diversity of the device ID into a unified and easy-to-process data format, so that in subsequent operations such as data storage, query and association analysis, there is no need to consider the original format differences of the device ID, thereby improving the versatility and efficiency of data processing. Step S1110 solves the problem of data processing difficulties caused by differences in device ID format and length, simplifies the data processing process, and improves the versatility of data processing. By unifying the device ID format, subsequent operations can process data from different devices in a consistent manner, avoiding errors and complexity caused by format differences. For example, when storing data, there is no need to design different storage strategies for device IDs in different formats, reducing the complexity and error probability of system design. The hash value in a unified format facilitates the subsequent interception of fixed-length fields (such as the lower 8 bits) as the basis for physical node mapping, ensuring that data from different devices can be distributed to edge physical nodes according to unified rules, providing a prerequisite for load balancing of distributed storage.
[0061] The described distributed data architecture consists of several edge physical nodes and a central governance node; the edge physical nodes are the underlying data storage units of the distributed data architecture, directly connected to industrial devices, and undertake the tasks of real-time access and sharded storage of multi-modal data. Each edge physical node corresponds to a physical storage entity (such as an industrial gateway, an edge server), and the central governance node is the high-level management unit of the distributed data architecture, responsible for integrating the fragmented data of the edge physical nodes, constructing a global association model and performing complex analysis. The reason for selecting the lower 8 bits of the device hash value as the physical node mapping identifier is that the lower 8 bits of the device hash value have good hashability for node mapping. Hashability means that the hash values can be more evenly distributed within a certain range. In this embodiment, it is to evenly distribute the data of different devices to different physical nodes. Taking a distributed system with 8 edge physical nodes as an example, when 100 devices generate data, by intercepting the lower 8 bits of the device hash value as the node mapping code, the amount of data received by each physical node is roughly equal, and there will be no situation where the data volume of a certain node is too large or too small. The purpose of doing this is to achieve load balancing at the physical node level and avoid affecting the overall performance of the system due to uneven data distribution, which may cause some nodes to be overloaded. Step S1120 solves the problem of possible load imbalance that may occur when data is stored in edge physical nodes, and improves the performance and stability of the distributed system. By balancing the load, each physical node can work under a relatively stable load, reducing the risk of system failures caused by the overload of a certain node, and improving the reliability and response speed of the entire distributed storage system. For example, when querying data, each node can respond to requests more quickly, reducing query latency. As the direct mapping identifier between the device and the physical node, the node mapping code can determine the data storage location without a complex routing algorithm, reducing the communication overhead and coordination cost between edge physical nodes. Traditional distributed storage often uses static allocation or simple hashing (such as modulo operation), which is prone to node load imbalance due to uneven distribution of device IDs. Step S1120 utilizes the strong hashability of cryptographic hashing to significantly improve the mapping uniformity and solve the problem of low storage efficiency caused by "data sharding skew" in the prior art.
[0062] Multimodal data includes vibration data, thermal imaging data, and text data; the operating state of industrial equipment needs to be comprehensively monitored from three core dimensions: mechanical vibration, thermal characteristics, and historical records; vibration data directly reflects the operating state of the mechanical components of the equipment (such as bearings, gears), and abnormal vibration frequencies often indicate faults such as component wear and imbalance, which are key indicators for early warning of equipment faults. Thermal imaging data is used to monitor the thermal state of the heat-generating components of the equipment (such as motors, terminal blocks), and abnormal temperature differences can indicate problems such as poor contact and heat dissipation failure. It forms a complement to vibration data. For example, abnormal vibration accompanied by local overheating can more accurately locate the fault location (such as bearing wear causing increased vibration and temperature rise). Text data (such as maintenance logs, sensor configuration parameters) provides historical maintenance information of the equipment (such as washing records, component replacement times) and static configurations (such as rated power), supplements the context semantics of the data, and is used as a background reference for correlation analysis. For example, combining maintenance logs can distinguish between temporary anomalies (such as temporary load fluctuations) and persistent faults (such as long-term fouling causing efficiency decline).
[0063] The preset period T is the basic time unit for data governance, and its length is determined according to specific scenarios, usually set to 5 to 15 minutes. Taking the production equipment of a certain factory as an example, if T is set to 10 minutes, within these 10 minutes, multimodal data such as vibration data, thermal imaging data, and text data generated by the equipment will be marked with the same node mapping code. The purpose of this is to achieve the clustering of data of the same equipment at the physical node level, facilitating subsequent correlation analysis and query. Because the data generated by the same equipment in a short period of time often has strong correlations, clustering and storing them can enable more rapid acquisition of relevant data during correlation analysis, improving the analysis efficiency. For example, when predicting equipment faults, it is necessary to analyze various data characteristics of the equipment over a period of time. The clustered data can reduce the time for data search and reading, improving the accuracy and timeliness of fault prediction. Step S1130 solves the problem of difficult correlation analysis caused by the lack of reasonable organization of data during storage, and improves the efficiency of data correlation analysis and query. By clustering the data of the same equipment within the preset period, the scope and time of data retrieval are reduced, the efficiency of data processing is improved, and it helps to more accurately perform operations such as equipment status monitoring and fault prediction. For example, when diagnosing equipment faults, relevant data can be obtained more quickly for analysis, shortening the fault diagnosis time and reducing equipment downtime.
[0064] Step S1100 unifies the device ID format through hashing operations to solve the data management chaos caused by the incompatibility of multi-source device identifiers. The hashing property of the lower 8 bits of the hash value is used to achieve uniform data sharding, avoiding overloading of a single node. Through the periodic T marking, it is ensured that multi-modal data within the same time period is associated at the physical storage layer, providing a basis for subsequent dynamic time window slicing. If S1110 is missing, heterogeneous device IDs cannot be unified, and the subsequent generation of node mapping codes loses standardized input, resulting in chaotic data sharding and unbalanced loads on edge nodes. If S1120 is missing, data cannot be evenly distributed to edge nodes, which may cause storage overload and increased computing latency on some nodes, violating the original intention of the distributed data architecture design. If S1130 is not implemented, multi-modal data within the same cycle is stored dispersedly, and when constructing the first-level bitmap in the subsequent step S2000, it will be impossible to accurately associate the time window with the data modality, resulting in distorted spatio-temporal correlation analysis.
[0065] Step S1200 performs dynamic time window slicing on the multi-modal data generated by the device;
[0066] Furthermore, step S1200 includes:
[0067] Step S1210 takes the moment of the device's first activation as the time origin, divides the device operation cycle into multiple sliding time windows according to the preset period T; at the same time, an absolute time interval is generated for each time window;
[0068] In step S1220, the device data carrying spatio-temporal feature encoding is written in the form of data blocks. When a data block is written, the deviation Δτ between the timestamp of the data block and the end of the current time window is calculated. Through the deviation Δτ, the time window to which the data block belongs is judged, and dynamic data sharding is performed. The data blocks include vibration data blocks, thermal imaging data blocks, and text data blocks.
[0069] Furthermore, step S1220 includes:
[0070] Step S1221, if Δτ is less than or equal to the allowable time deviation threshold Δt, then the data block is forcibly written to the physical node corresponding to the current time window;
[0071] Step S1222, if Δt < Δτ < T, then cross-window sharding compensation is started, and the data block is split into a main data block and a mirror data block carrying spatio-temporal coordinate metadata. The main data block is written to the current physical node, and the mirror data block is written to the physical node in the adjacent cycle;
[0072] Step S1223, if Δτ ≥ T, then the data block is determined as an abnormal sharding, triggering abnormal sharding fusing and performing hierarchical abnormal processing.
[0073] The method of the hierarchical abnormal processing includes:
[0074] When T ≤ Δτ < 2T, trigger the automatic calibration request for the device clock, and temporarily store the data block at the physical node on the edge of the time window. After calibration, remap the time window;
[0075] When Δτ ≥ 2T, transfer it to the central governance node and mark "severe clock anomaly", and at the same time trigger the device offline alarm.
[0076] Specifically, the dynamic time window slicing of the multi-modal data block generated by the device in step S1200 is a key link in spatio-temporal feature encoding, aiming to solve the problem of spatio-temporal correlation fragmentation caused by device clock drift and inconsistent data arrival time sequences. For example, when monitoring the running state of a device, if there is a slight deviation in the device clock, the data stored according to the fixed time stamp may be misclassified, resulting in deviations in the subsequent analysis of the device running trend and the failure to detect potential faults in a timely manner. By constructing a sliding time window system and implementing a dynamic sharding strategy, the continuity and relevance of multi-modal data in the time dimension are ensured, providing a time sequence basis for the subsequent construction of multi-level correlation bitmaps.
[0077] The sliding time window mechanism can break through the limitations of absolute time stamps and provide a unified time reference system; during the operation of the device, data is continuously generated. If only relying on absolute time stamps, when the device clock drifts, the time sequence and relevance of the data will become chaotic. The sliding time window mechanism can classify the data generated at different times according to a fixed period T. Regardless of how the device clock drifts, as long as the data generation time is within a certain time window range, it can be correctly classified and processed. The use of the sliding time window conforms to the characteristics of continuous data generation and time-sequence sensitivity of device data in the industrial Internet scenario, effectively solving the problem of data classification chaos caused by inconsistent time references, making the data clearly divided and organized in the time dimension, providing an orderly basis for subsequent data correlation analysis, and helping to more accurately grasp the change law of the device running state over time.
[0078] When the data block is written into the distributed system, the time stamp it carries reflects the data generation time. By calculating the difference Δτ between this time stamp and the end of the current time window, the time window to which the data block should belong is determined, which is the core operation of dynamic data sharding. For example, such as Figure 3As shown, the current time window is [10, 19], and the corresponding absolute time interval is [2024-01-01 08:10:00, 2024-01-01 08:19:00]. If the timestamp of a data block is 2024-01-01 08:15:00, by calculation, the deviation Δτ between the timestamp of this data block and the end of the time window is 4 minutes. Based on this, it is determined that this data block should belong to the current [10, 19] time window. This method of determining the time window to which data belongs according to the timestamp deviation can store data in slices in real-time and dynamically, adapting to the uncertainty of the data generation time of industrial equipment, avoiding storage and analysis errors caused by chaotic time order of data, solving the problem of difficult establishment of correlation due to time order issues during data storage, providing an accurate data slicing basis for subsequent time series-based data correlation analysis, and helping to improve the accuracy and reliability of data analysis.
[0079] The allowed time deviation threshold Δt is set to tolerate a certain degree of clock drift. It is very difficult for device clocks to achieve completely precise synchronization, and there will always be a certain error. For example, if Δt is set to 30 seconds, when the calculated deviation Δτ between the timestamp of the data block and the end of the current time window is less than or equal to 30 seconds, it means that the time deviation of this data block is within the acceptable range. At this time, the data block is directly written to the physical node corresponding to the current time window. This is because within this deviation range, the data block can still be considered to be generated within the current time window. Directly writing to the corresponding node can ensure the correct time attribution of the data, reduce unnecessary cross-window operations, and avoid the increase in complexity of data storage and correlation analysis caused by cross-window operations. This processing method solves the problem of possible data storage errors caused by small clock drifts of devices, ensures the continuity and consistency of data in the time dimension, enables time series-based data correlation analysis to more accurately reflect the operating state of the device, and improves the accuracy and stability of data processing. Avoid unnecessary slice splitting and improve storage efficiency.
[0080] Δt < Δτ < T indicates that the time deviation of the data block exceeds the range that can be directly written into the current time window but does not belong to a severe deviation. At this time, the cross-window sharding compensation mechanism is started, and the data block is split into a main data block and a mirror data block carrying spatio-temporal coordinate metadata. The spatio-temporal coordinate metadata includes the time information of the data block (such as the approximate time range when the data is generated) and the information about its association with other data blocks (such as the device ID it belongs to, its position in the original data block sequence, etc.). For example, the timestamp of a data block shows that the deviation between its generation time and the end of the current time window is 1 minute (assuming Δt is 30 seconds and T is 10 minutes), which exceeds Δt but is less than T. At this time, the data block is split into a main data block and a mirror data block. The main data block is written into the current physical node, while the mirror data block carrying the spatio-temporal coordinate metadata is written into the physical node of the adjacent period. This can ensure that data will not be lost or misaligned due to timestamp deviation. At the same time, the coordinate metadata carried by the mirror data block provides clues for subsequent data association. For example, when conducting equipment failure analysis later, through the spatio-temporal coordinate metadata of the mirror data block, data that is actually related but in different time windows can be associated, so as to more comprehensively analyze the equipment operation status, improve the integrity and relevance of the data, solve the data processing problem caused by large but not severely abnormal timestamp deviation, and provide data support for more accurate equipment status monitoring and fault prediction. Equipment clock drift or data transmission delay may cause data blocks to cross the time window boundary. Traditional methods may incorrectly discard or force attribution, resulting in time series breakage. By storing both the main data block and the mirror data block, it not only ensures the integrity of the data in the current window but also retains the cross-window association information.
[0081] If Δτ≥T, it indicates that there may be a serious problem with the device clock. At this time, the data block is determined as an abnormal shard, and the abnormal shard fusing mechanism is triggered. Abnormal shard fusing means taking special processing measures for abnormal data blocks to prevent them from interfering with the normal data processing flow. The specific processing method is to transfer the data block to the central governance node for hierarchical abnormal processing. Hierarchical abnormal processing is divided into two cases: when T≤Δτ<2T, a device clock automatic calibration request (such as NTP synchronization) is triggered, and the data block is temporarily stored in the edge physical node. After calibration, the time window is remapped. This is because in this case, the clock deviation may be temporary, and the device clock can be restored to normal through automatic calibration. After remapping the time window, the data can be correctly processed. For example, if the deviation Δτ is 12 minutes (T is 10 minutes), a device clock calibration request is triggered, and the data block is temporarily stored in the edge physical node during the waiting for calibration. After calibration, the time window to which the data block belongs is re-determined according to the new time information. When Δτ≥2T, it is transferred to the central governance node and marked as "serious clock anomaly", and at the same time, a device offline alarm is triggered. This is because the clock deviation is relatively serious at this time, which may affect the normal operation of the device, and relevant personnel need to be notified in time for processing. Only the extreme scenario of "Δτ≥K×T" (K is an empirical threshold, such as 2) is handled manually. Such a hierarchical processing strategy can reduce manual intervention, improve the automatic processing ability of the system, avoid overloading of the central governance node, and meet the requirements of industrial scenarios for real-time and reliability. In this way, abnormal data is effectively isolated, avoiding interference with the normal data sharding order, ensuring the stability and reliability of the entire data processing system, and providing guarantee for accurately analyzing the device operation status.
[0082] After performing the above dual coding process on the original multi-modal data, device data carrying spatio-temporal feature coding is obtained. The device data carries the following spatio-temporal feature coding to form a structured data unit, as shown in Table 1:
[0083] Table 1 Device Data Carrying Spatio-Temporal Feature Coding
[0084]
[0085] Step S1200 ensures the continuity of data in the time dimension through dynamic time window and deviation classification processing, regardless of whether the device clocks are fully synchronized. The cross-window sharding compensation mechanism preserves the spatio-temporal context of the data, avoiding the association breakage caused by time deviation. Distinguishing repairable and severe anomalies improves the fault tolerance of the system, meeting the high-reliability requirements of industrial scenarios. If step S1210 of time window division is omitted, the data will rely on absolute timestamps for storage, and due to clock differences among different devices, it is difficult to align the time, and when constructing the first-level bitmap later, it is impossible to accurately associate the multi-modal data time windows of the same device, resulting in the failure of cross-modal analysis. If there is no S1220 dynamic sharding strategy, moderately time-deviated data will be misclassified or discarded, forming data holes; unisolated severe anomaly data will contaminate the edge physical nodes, resulting in incorrect construction of the association bitmap, and ultimately affecting the accuracy of the global anomaly propagation map.
[0086] Step S2000, based on the device data carrying spatio-temporal feature encoding, constructs a multi-level association bitmap, where the multi-level association bitmap includes a first-level bitmap, a second-level bitmap, and a third-level bitmap;
[0087] Further, step S2000 includes:
[0088] Step S2100, constructs a first-level bitmap, and the first-level bitmap is a physical association bitmap;
[0089] Further, step S2100 includes:
[0090] Step S2110, creates a two-dimensional bit matrix with M rows and N columns for the device data carrying spatio-temporal feature encoding of each device, denoted as the device association matrix; the rows of the device association matrix represent the sliding time windows, the columns represent the data modalities, the number of rows M corresponds to the number of sliding time windows, and the number of columns N corresponds to the types of data modalities;
[0091] Step S2120, marks the association status at the corresponding row and column positions according to the arrival time window of the device data carrying spatio-temporal feature encoding, forming the first-level bitmap;
[0092] The marking of the association status at the corresponding row and column positions according to the arrival time window of the device data carrying spatio-temporal feature encoding includes:
[0093] Sets the columns in the device association matrix of the row corresponding to the arrival time window of the vibration data block to 1, sets the columns in the device association matrix of the row corresponding to the arrival time window of the thermal imaging data block to 1, and sets the columns in the device association matrix of the row corresponding to the arrival time window of the text data block to 1.
[0094] Specifically, in a distributed data architecture, each device generates multimodal data such as vibration data, thermal imaging data, and text data during operation. These data need to establish a structured association in the time and modality dimensions. Step S2110 creates a two-dimensional bit matrix as the device association matrix DeviceMap for each device based on the spatio-temporal feature encoding generated by S1000 (including the node mapping code with hash spatio-temporal anchoring and the time window sequence number of the dynamic time window slice). The number of rows M of the matrix corresponds to the number of sliding time windows generated during the device operation (for example, when the device runs for 72 hours and the preset period T = 10 minutes, M = 432), and each row uniquely corresponds to a time window (for example, the i-th row corresponds to the time window [iT, (i + 1)T)); the number of columns N corresponds to the number of data modality types (preset N = 3, which are vibration data, thermal imaging data, and text data respectively and can be extended), and each column represents a data modality. The matrix elements are represented by binary bits (0 or 1) to indicate whether there is such modality data in the corresponding time window: all elements are set to 0 in the initial state, and are set to 1 according to their time window and modality type when the data block is written. For example, a wind power sensor device WT-001 collects vibration data and thermal imaging data simultaneously in the 5th time window (10:00 - 10:10), then DeviceMap[5, 1] and DeviceMap[5, 2] are set to 1, and DeviceMap[5, 3] remains 0 (no text log). This matrix is mapped through the node mapping code of S1130 and the sliding time window of S1210 to ensure that the physical storage location (edge physical node) of each data block corresponds one-to-one with the time window and modality type, solving the problem of invisibility of spatio-temporal distribution caused by fragmentation of multimodal data in traditional distributed storage.
[0095] When the device data carrying the spatio-temporal feature encoding is written into the edge physical node in the form of data blocks, the system parses the spatio-temporal feature encoding carried by the data blocks, and extracts the time window serial number W' and the modality type C' (vibration = 1, thermal imaging = 2, text = 3). Locate the matrix element in the W'-th row and C'-th column in the DeviceMap, and set it to 1 to mark the existence of the corresponding modality data within this time window. If the data block generates an image data block through cross-window sharding compensation, both the main data block and the image data block need to be marked in the rows corresponding to their respective time windows (the main data block marks the current time window, and the image data block marks the adjacent time window) to ensure that the existence status of the time window deviation data is traceable. For example, a vibration data block triggers cross-window sharding compensation due to a timestamp deviation of Δτ = 40 seconds (Δt = 30 seconds). The main data block is written into the node corresponding to the time window W' = 10, and the image data block is written into the node corresponding to W' = 11. Then, both DeviceMap[10,1] and DeviceMap[11,1] are set to 1, and the mirror position is recorded in the metadata of the main data block. This marking mechanism realizes efficient data status recording through atomic operations of binary bits, and supports quickly judging the integrity of multi-modal data through bit operations (for example, whether vibration and thermal imaging data exist simultaneously in a certain time window can be directly calculated through DeviceMap[W',1]&DeviceMap[W',2]).
[0096] The management of multi-modal data in traditional distributed storage only stays at the file level or database table level, lacking a unified spatio-temporal dimension correlation index. For example, vibration data is sharded and stored according to the sensor ID, and thermal imaging data is sharded and stored according to the timestamp, resulting in multi-modal data within the same device and the same time window may be distributed on different nodes. When performing cross-modal correlation queries, it is necessary to traverse multiple shards, which takes a long time. In addition, the problem of data time window misalignment caused by device clock drift further exacerbates the fragmentation of data distribution, making it difficult for the analysis system to quickly locate a complete multi-modal data set. Step S2110, through the structured design of the device association matrix, converts the time window and modality type of multi-modal data into computable matrix dimensions, forming a physical existence status index of multi-modal data within the device, enabling the distributed system to quickly obtain the modality integrity information within any time window through matrix row and column positioning, improving the query efficiency compared with traditional unstructured storage. Step S2120, through the dynamic marking of the data arrival status, records the actual storage location of multi-modal data and its association with the time window in real time, solving the problem of "difficult to trace the existence status of cross-modal data". It supports the edge physical node to quickly respond to the data integrity verification request, avoiding the distortion of correlation analysis caused by data block loss or misalignment.
[0097] As the core carrier of physical layer association, the first-level bitmap supports the logical association analysis of the second-level bitmap upward (providing the time window premise for the existence of data), and connects to the spatiotemporal feature coding of S1000 downward (converting the physical storage location of the data block into a computable matrix element), forming the underlying data path of "storage coding-state recording-association analysis". The lightweight bit matrix realizes the explicit expression of the spatiotemporal distribution of multimodal data, so that the fragmented data under the distributed data architecture can build a physically existing "digital twin" with the device as the unit, providing a reliable underlying data source for the subsequent high-level analysis of logical association and topological association, and finally realizing the efficient coordination of edge storage sharding and central association calculation.
[0098] Step S2200, constructing a secondary bitmap, wherein the secondary bitmap is a logical association bitmap;
[0099] Further, step S2200 includes:
[0100] Step S2210, performing logical association determination on the vibration data block and the thermal imaging data block in the same time window;
[0101] Further, step S2210 includes:
[0102] Step S2211, extracting the vibration frequency feature F of the vibration data block vib and the temperature difference feature T of the thermal imaging data block her ;
[0103] Step S2212: when writing a data block, if the vibration frequency characteristic of the vibration data block in the same time window exceeds a preset frequency threshold F th , and the temperature difference feature of the thermal imaging data block exceeds the preset temperature difference threshold T th , it is determined that there is a logical association between the vibration data block and the thermal imaging data block within the time window.
[0104] Specifically, equipment failures are usually manifested as coordinated anomalies of multi-modal data (such as the simultaneous occurrence of abnormal vibration frequency and local temperature rise). Therefore, it is necessary to further explore the logical association of different modal data in the same time window based on the equipment association matrix. Step S2210 implements logical association determination for the vibration data block and the thermal imaging data block through feature extraction and threshold comparison. Perform frequency domain analysis (such as fast Fourier transform) on the vibration data block and extract the main frequency component as the vibration frequency feature F vib , which reflects the periodic vibration intensity of the equipment during operation; for the thermal imaging data block, the difference between the maximum and minimum temperature of the region of interest (ROI) is calculated as the temperature difference feature T her , which reflects the degree of local heating of the equipment. For example, when a motor bearing fails, F vibwill show a peak near the bearing characteristic frequency, and at the same time, the thermal imaging temperature difference T of the bearing housing her significantly increases. The frequency threshold F th (such as the fluctuation range of ±10% of the nominal value of the bearing characteristic frequency) and the temperature difference threshold T th (such as twice the maximum temperature difference during normal operation of the equipment) are determined based on the normal operation parameters of the equipment and historical fault data through data analysis and expert experience. When a data block is written, if F within the same time window vib exceeds F th and T her exceeds T th , it is determined that there is a logical association between the two, indicating that there may be a common fault cause (such as component wear leading to increased vibration and accompanied by heat generation due to friction).
[0105] Step S2220, if the logical association determination result is that there is a logical association, then create a logical association channel in the corresponding time window row of the equipment association matrix;
[0106] Specifically, when it is determined that there is a logical association between the vibration data block and the thermal imaging data block within the same time window, it is necessary to establish a data structure in the equipment association matrix to represent this association relationship, that is, create a logical association channel. In the equipment association matrix, each row represents a sliding time window, and the columns represent different data modalities (such as vibration data, thermal imaging data, text data). Taking a certain equipment as an example, in the 10th time window, if the vibration data block and the thermal imaging data block are determined to be logically associated, a connection is established between the columns corresponding to the vibration data and the thermal imaging data in the 10th row of the equipment association matrix by means of a certain data marker or pointer. An additional data field can be added to the matrix to record this logical association relationship. For example, a flag bit can be set. When the flag bit is 1, it indicates that there is a logical association between these two types of data in this time window; or a pointer can be created to point to other data records associated with these two types of data to clarify their logical connection. By creating a logical association channel, the originally scattered data of different modalities are connected at the logical level, solving the problem that it is difficult to reflect the logical association of cross-modal data, providing an intuitive data structure for subsequent calculation of the logical association strength and construction of the secondary bitmap, enabling quick location and analysis of the logically associated data when analyzing the operating state of the equipment, and improving the efficiency and accuracy of data analysis.
[0107] Step S2230, calculate the spatio-temporal coupling degree C of the logical association channel, and mark the corresponding bit positions of the logical association channel according to the spatio-temporal coupling degree C to generate a secondary bitmap.
[0108] Further, step S2230 includes:
[0109] Step S2231: Obtain the spatio-temporal coupling degree C based on the vibration frequency characteristics, temperature difference characteristics, and deviation Δτ.
[0110] Step S2232: If , then mark the corresponding bit of the logical association channel as 1, and the bit marked as 1 is a strongly associated bit; if , then mark it as 0, and the bit marked as 0 is a weakly associated bit; otherwise, do not mark; where is a preset strong association threshold, is a preset weak association threshold.
[0111] Specifically, in step S2230, by calculating the spatio-temporal coupling degree C of the logical association channel and marking the corresponding bit of the logical association channel according to its value, a secondary bitmap containing rich logical association information is generated. The spatio-temporal coupling degree C is an index that comprehensively considers the vibration frequency characteristics, temperature difference characteristics, and the time deviation Δτ of the data block, and is used to measure the association strength of the logical association channel. It is obtained by multiplying the normalized values of the vibration frequency characteristics, the thermal imaging temperature difference characteristics, and the time deviation respectively. Taking the vibration frequency characteristics as an example, through a specific normalization function, the vibration frequency characteristic F vib is transformed into a unified numerical range, so that the vibration frequency characteristics under different devices or different working conditions are comparable; a similar normalization process is also performed on the thermal imaging temperature difference characteristic T her . The time deviation Δτ is reverse-normalized (the smaller Δτ is, the higher the normalized value, indicating that the data is closer to the ideal time window). Then, these three normalized values are multiplied to obtain the spatio-temporal coupling degree C. The larger this value is, the stronger the abnormal coordination between the vibration and thermal imaging data. This calculation method can comprehensively consider the influence of multiple factors on the logical association strength, and more comprehensively reflect the association degree between data. The actual operating state of the equipment is complex and changeable, and a single characteristic cannot accurately measure the association strength between data. By comprehensively considering multiple factors, the logical association tightness between different modal data in the equipment can be judged more accurately, providing a more reliable basis for subsequent marking and analysis.
[0112] Strong association threshold (such as 0.8) and weak association threshold (such as 0.5) are determined based on the historical operation data and failure cases of the equipment through a large amount of data analysis and experiments. Taking a certain equipment as an example, if the calculated spatio-temporal coupling degree C is greater than or equal to the strong association threshold , it indicates that the association between the vibration data block and the thermal imaging data block represented by this logical association channel is very tight. In the secondary bitmap, mark the corresponding bit of this logical association channel as 1, representing strong association; if C is between the weak association threshold and the strong association threshold If C is less than , no marking is performed. This marking method intuitively reflects the correlation strength of different logical association channels in the secondary bitmap, which facilitates the subsequent analysis of the correlation propagation characteristics of fault symptoms in the equipment, solves the problem of difficulty in quantifying and distinguishing different logical association strengths, provides more accurate data support for local root cause analysis, and can quickly locate key data and logical associations related to faults.
[0113] Traditional distributed data governance only stores the physical existence of multimodal data and lacks explicit expression of logical associations between data. For example, although vibration and thermal imaging data are stored in the same node, the analysis system needs to align timestamps one by one and calculate correlations through complex algorithms, which is time-consuming and has low accuracy. In addition, the time deviation caused by device clock drift will further interfere with the association judgment, increasing the missed judgment rate of fault diagnosis. Step S2210 converts the multimodal signs of equipment failure into computable logical association judgment conditions through feature extraction and threshold comparison, solving the problem of difficult identification of abnormal coordination of cross-modal data. It enables the analysis system to quickly locate time windows with potential fault signs, reducing preprocessing time compared to traditional data block-by-data correlation calculation. Step S2220 establishes a direct association between vibration and thermal imaging data within the time window through explicit marking of the logical association channel, solving the problem of difficult tracing of cross-modal logical association relationships, and supports direct screening of associated time windows through bit operations to improve query efficiency. Step S2230 quantifies the strength of cross-modal associations by calculating spatiotemporal coupling and grading labels, solving the problem of difficulty in quantifying association strength. It provides a weight basis for subsequent abnormal propagation analysis (such as using strong correlation channels as high-weight edges when constructing local abnormal propagation subgraphs), enabling the root cause location algorithm to prioritize the analysis of highly reliable association paths.
[0114] As the core carrier of logical layer association, the secondary bitmap receives the physical existence status of multimodal data in the device (the primary bitmap) and supports the determination of topological associations between devices, forming a middle-level association network of "physical existence-logical association-topological propagation". Through feature threshold determination and coupling quantification, the multimodal signs of equipment failure are converted into computable and traceable logical association tags, so that cross-modal data under the distributed data architecture can build a "digital clue" of logical association centered on the failure, providing a key middle-level index for subsequent system-level abnormal propagation analysis, and finally achieving a core breakthrough from fragmented data storage to semantic association.
[0115] Step S2300, constructing a three-level bitmap, wherein the three-level bitmap is a topological association bitmap;
[0116] Further, step S2300 includes:
[0117] Step S2310: Construct a global spatio-temporal correlation matrix. The rows of the global spatio-temporal correlation matrix represent all devices, and the columns represent all time windows.
[0118] Step S2320: Traverse the secondary bitmaps of each pair of devices. If there is an overlapping area in the absolute time interval of the time windows divided by the two devices in each pair of devices, and strong correlation bit positions appear in the overlapping area, and the spatial distance between the two devices does not exceed the preset distance threshold, then set the same columns in the corresponding rows of the global spatio-temporal correlation matrix for the two devices to 1 to form a cross-device correlation tunnel.
[0119] Step S2330: Calculate the topological correlation strength W of the cross-device correlation tunnel. The topological correlation strength W is inversely proportional to the spatial distance D between the two devices.
[0120] Step S2340: Quantify the topological correlation strength W into L levels, and each level corresponds to a level code; fill the level code after quantization into the corresponding bit positions of the device correlation tunnel to generate a tertiary bitmap.
[0121] Specifically, the global spatio-temporal correlation matrix is a data structure used to describe the correlation relationships of all devices within all time windows. Its rows represent all devices, and its columns represent all time windows. The global spatio-temporal correlation matrix breaks through the intra-device limitation of the device correlation matrix and extends the correlation analysis to between devices. Suppose there are 10 devices, and 50 time windows are divided through Step S1210, then the constructed global spatio-temporal correlation matrix is a 10-row and 50-column matrix. When actually constructing the global spatio-temporal correlation matrix, first traverse all devices and time windows, reserve storage space for the correlation status of each device in each time window, and establish a unified matrix structure to associate devices with time windows, so that the relationships between devices can be analyzed from a global perspective subsequently. In the industrial Internet environment, there are numerous devices and data generation has a time series nature. This matrix structure can effectively integrate and manage the correlation information between devices, provide a basis for subsequent mining of the logical correlations and abnormal propagation paths between devices, solve the problems of scattered correlation information between devices, difficult unified management and analysis, and provide a global device correlation analysis framework, facilitating the overall grasp of the relationships between devices.
[0122] After constructing the global spatio-temporal correlation matrix, it is necessary to determine the logical correlation relationships between different devices, which is achieved by traversing the secondary bitmaps of each pair of devices. The distance threshold can be set according to the maximum functional influence range within the device layout area. For example, in an industrial workshop, the maximum logical influence distance between adjacent functionally related devices usually does not exceed 10 meters, so the distance threshold can be set to 10 meters. When the time windows divided by each pair of devices have an overlapping area in the absolute time interval, and strong correlation bit positions appear in the overlapping area, and the spatial distance between the two devices does not exceed the preset distance threshold, set the same columns in the corresponding rows of the global spatio-temporal correlation matrix for the two devices to 1, thus forming a cross-device correlation tunnel. For example, if device A and device B have an overlapping area in the time windows [10, 20) and [15, 25), and in this overlapping area, strong correlation bit positions appear in the secondary bitmaps of device A and device B, and their spatial distance does not exceed the preset distance threshold, then in the global spatio-temporal correlation matrix, set the positions of the corresponding rows of device A and device B and the columns corresponding to the overlapping time windows to 1. The spatial distance between the two devices is calculated using the Euclidean formula based on the position coordinates of the two devices. This process is based on the device-internal logical correlation information already mined in the secondary bitmap and further extends to the logical correlation analysis between devices. In industrial production, there are often potential logical correlations between adjacent devices or functionally related devices. By this means, these correlation relationships can be accurately identified, solving the problem that it is difficult to discover and locate the logical correlations between devices, revealing the spatial distribution law of the logical correlations between devices, and providing key information for analyzing the mutual influence between devices.
[0123] The cross-device correlation tunnel establishes the logical correlation between devices, but it is also necessary to quantify the strength of this correlation, that is, to calculate the topological correlation strength W. The topological correlation strength W is inversely proportional to the spatial distance D between the two devices, meaning that devices closer to each other are given a higher topological correlation strength. When calculating, first obtain the spatial distance D between the two devices, and then obtain the topological correlation strength W according to this inverse relationship. For example, obtain the spatial distance between device A and device B through the position information of the devices, and calculate their topological correlation strength according to the established inverse calculation rule (such as when the distance is 1, the strength is 10, when the distance is 2, the strength is 5, etc.). This calculation method conforms to the actual situation of the mutual influence between devices in the industrial scenario. Devices closer to each other are more likely to influence each other. It solves the problem that it is difficult to quantify the strength of the logical correlation between devices, can more accurately describe the propagation ability of the logical correlation between devices in the spatial network, and provides a quantitative basis for subsequent analysis of abnormal propagation.
[0124] To store and analyze the topological association strength more effectively, the topological association strength W is quantified into L levels, each level corresponding to a level code. Then, the quantified level codes are filled into the corresponding bit positions of the device association tunnel to generate a three-level bitmap. For example, the value range of the topological association strength W is divided into 5 levels. The strength from 0 to 2 is level 1, corresponding to the code 001; the strength from 3 to 4 is level 2, corresponding to the code 010, etc. Then, these codes are filled into the global spatio-temporal association matrix to form a three-level bitmap. This way of quantization and coding can describe complex spatio-temporal association structures with extremely low storage overhead, solves the problem of difficult storage of complex association structures, reduces the occupation of storage resources while ensuring the integrity of association information, and improves the efficiency of data storage and processing.
[0125] In the traditional distributed data architecture, the topological association analysis between devices is often not deep and comprehensive enough. On the one hand, the full consideration of the spatial position relationship between devices is lacking, resulting in the inability to accurately identify device pairs that are spatially close and have strong data associations. On the other hand, the data association strength between devices is not quantified and coded, making it difficult to manage and analyze effectively. For example, in a large industrial park, there are many devices distributed in different areas. Traditional methods may not be able to accurately identify which devices are associated with each other and the strength of these associations. This makes it impossible to fully utilize the association information between devices during fault diagnosis, performance optimization, etc., resulting in low efficiency and low accuracy. Step S2300 focuses the subsequent topological association analysis on device pairs that are spatially closer by calculating the spatial distance between device pairs and filtering out device pairs whose distance is within a preset threshold, improving the pertinence and effectiveness of the analysis. By constructing a global association matrix and identifying device pairs that both have strong association marks within a common time window, an accurate basis is provided for establishing a device association tunnel, enabling the topological association relationship between devices to be clearly represented. By establishing a device association tunnel and calculating the spatial distance between devices, the data association strength between devices can be quantified by the spatial distance, providing strong support for subsequent analysis and management.
[0126] The single-dimensional limitation of the traditional indexing method makes it difficult to comprehensively and accurately display the association relationship between data, thus affecting the in-depth understanding and effective utilization of data. Step S2000 of the present invention forms a three-dimensional association network of physical-logical-topological by constructing a multi-level association bitmap, successfully breaking through the limitations of traditional indexing. The multi-level association bitmap makes data query more efficient. Such as Figure 4As shown, the multi-level associated bitmap includes a first-level bitmap, a second-level bitmap, and a third-level bitmap. Through the first-level bitmap, the time window and modality where the data is located can be quickly determined. The second-level bitmap can further screen out the logically associated data, and the third-level bitmap can help find other devices associated with a specific device and their data. This hierarchical query method helps reduce the scope of data search and improve the query speed. For example, when querying the relevant data of a certain device within a specific time period, the traditional method may need to traverse a large number of files or database records, while through the multi-level associated bitmap, the relevant data block can be directly located, and the query time is shortened from several minutes to several seconds, greatly improving work efficiency.
[0127] Step S3000, based on the constructed multi-level associated bitmap, implement collaborative computing between the edge physical nodes and the central governance node of the distributed data architecture.
[0128] Furthermore, step S3000 includes:
[0129] Step S3100, at the edge physical node level, based on the first-level bitmap and the node mapping code, establish a device-node routing table;
[0130] Specifically, in the distributed data architecture, the edge physical nodes undertake the task of sharding and storing multi-modal data, and quickly locating the data storage location is the basis for realizing efficient data governance. Step S3100 constructs a device-node routing table by integrating the first-level bitmap and the node mapping code, solving the problem of low data mapping efficiency of traditional edge physical nodes. The device-node routing table is used to store the device ID, the node mapping code, and their corresponding relationships with the physical nodes. The implementation process is as follows: The system traverses the first-level bitmap and extracts the device ID of each device and the corresponding node mapping code information. For example, in an industrial scenario with 100 devices and 10 edge physical nodes, the first-level bitmap records the data storage situation of each device in different time windows and the corresponding node mapping code. Suppose the device ID of device A is "001", and the node mapping code corresponding to its data in multiple time windows is "10101010" (obtained by hash space-time anchoring), and this node mapping code corresponds to edge physical node 3. After the system organizes this information, it constructs a device-node routing table. By solidifying the corresponding relationship between the device ID and the node mapping code in the first-level bitmap into the routing table, the cumbersome process of recalculating the mapping relationship every time device data is accessed is avoided.
[0131] Step S3200, at the central governance node level, based on the second-level bitmap and the third-level bitmap, construct a global exception propagation graph;
[0132] Furthermore, as Figure 5 shown, step S3200 includes:
[0133] Step S3210, construct a local anomaly propagation subgraph: Identify the logical association channels in each device's second-level bitmap, add directed edges between the two end nodes of the logical association channels, with the direction of the edge being the time-increasing direction and the weight of the edge being the spatio-temporal coupling degree C of the logical association channel;
[0134] Step S3220, construct a global anomaly propagation backbone graph: Identify the cross-device association tunnels in the third-level bitmap, add undirected edges between the two end devices of the cross-device association tunnels, and the weight of the edge is the topological association strength W;
[0135] Step S3230, parallelly connect the local anomaly propagation subgraphs of all devices to obtain a complete intra-device anomaly propagation graph, and then superpose and connect the global anomaly propagation backbone graphs of all devices to form a global anomaly propagation graph.
[0136] Specifically, when analyzing anomaly propagation in a distributed system, existing technologies often lack comprehensive and effective methods, making it difficult to accurately grasp the propagation laws of anomalies in different devices and time dimensions. For example, it may only be able to simply analyze the anomalies of a single device and cannot consider how anomalies spread among multiple devices from the overall perspective of the system. As a result, when dealing with complex industrial system failures, it is impossible to quickly locate the root cause of the failure and predict the development trend of the failure. Step S3200 aims to construct a global anomaly propagation graph based on the second-level bitmap and the third-level bitmap to achieve effective analysis and positioning of anomaly propagation in the distributed system, providing a decision-making basis for root cause analysis and predictive maintenance.
[0137] Taking the time window of the device as a node, each node contains the existence state of multi-modal data (from the first-level bitmap) and the logical association mark (from the second-level bitmap) within that time window. Step S3210, from the perspective of a single device, establishes directed edges between successively abnormal state nodes in time sequence, forming a local anomaly propagation subgraph. This subgraph reflects the abnormal development and evolution process within a single device. During the operation of industrial devices, there are often logical connections between abnormal states in different time windows. By establishing directed edges and assigning weights, these connections and association strengths can be intuitively displayed. Traditional time series analysis relies on direct comparison of timestamps. Step S3210 transforms the intra-device anomaly propagation into the edge weights of a directed graph through the explicit marking of logical association channels, enabling the quantitative analysis of the time series evolution of abnormal states. Step S3210 solves the problem that it is difficult to intuitively present and analyze the abnormal development process within a single device, provides an intuitive data structure for local root cause analysis, facilitates in-depth study of the abnormal evolution law within a single device, and helps quickly locate the root cause of internal device failures.
[0138] When constructing the global abnormal propagation backbone graph, it is necessary to identify the cross-device association tunnels in the three-level bitmap. For example, in an industrial production line with multiple devices, there is a cross-device association tunnel between device A and device B in the three-level bitmap, which means that they have a strong logical association within certain time windows and the spatial distance meets the preset distance threshold. The system adds an undirected edge between the two devices at both ends of the cross-device association tunnel, and the weight of the edge is the topological association strength W. Suppose the topological association strength W between device A and device B is 0.6. Then, when constructing the global abnormal propagation backbone graph, an undirected edge with a weight of 0.6 is added between the nodes corresponding to device A and device B. In this way, each cross-device association tunnel is abstracted into an undirected edge of the backbone graph, connecting the device nodes corresponding to both ends of the association tunnel, forming a backbone graph that reflects the global abnormal propagation path. This backbone graph reveals the spatial topology of abnormal diffusion between different devices. Multiple devices are often interconnected, and abnormalities may spread between devices. The global abnormal propagation backbone graph can show the path and intensity of abnormal propagation from the system level, providing a macroscopic perspective for analyzing the fault propagation of the entire system. When an abnormality occurs in a certain device, it is possible to quickly determine which devices may be affected and the degree of influence through the global abnormal propagation backbone graph, providing a basis for taking corresponding preventive and handling measures. Step S3220 solves the problem of difficult to analyze the propagation path and intensity of abnormalities between different devices from the system level, reveals the spatial topology of the global abnormal propagation path, and provides an important basis for system-level fault analysis and prevention.
[0139] After completing the construction of the local abnormal propagation sub-graph and the global abnormal propagation backbone graph, the system parallels the local abnormal propagation sub-graphs of all devices to obtain a complete intra-device abnormal propagation graph. This step integrates the abnormal propagation processes inside each device to form a graph structure covering all the abnormal situations inside the devices. Then, the complete intra-device abnormal propagation graph is superimposed with the global abnormal propagation backbone graph to finally form the global abnormal propagation graph. The global abnormal propagation graph logically reflects the spatio-temporal evolution law of the abnormal state in the distributed system. In this way, the abnormal propagation information inside and between devices is integrated, providing a comprehensive decision-making basis for root cause location and predictive maintenance. In actual industrial scenarios, it is necessary to grasp the overall propagation situation of abnormalities in the entire distributed system. The global abnormal propagation graph can integrate the local and global abnormal propagation information, providing a complete perspective. When an abnormality occurs in the distributed system, the root cause of the abnormality can be quickly located through the global abnormal propagation graph, the development trend of the abnormality can be predicted, corresponding maintenance strategies can be formulated, and the losses caused by equipment failures can be reduced. Step S3230 solves the problem of unable to comprehensively analyze the spatio-temporal evolution law of abnormalities in the distributed system, provides a comprehensive and intuitive decision-making basis for root cause location and predictive maintenance, and greatly improves the system's ability to respond to abnormal situations.
[0140] Step S3300: Based on the global exception propagation graph, implement distributed root cause location through querying in a graph database.
[0141] Specifically, in the distributed data architecture of the industrial Internet, when an exception occurs in the system, accurately finding the root cause of the exception is crucial for quickly solving problems and reducing losses. Step S3300 utilizes the constructed global exception propagation graph and the query function of the graph database to implement distributed root cause location. A graph database is a database specifically designed for processing graph data structures. It is good at storing and querying graph-structured data composed of nodes and edges, and is very suitable for processing the complex device associations and exception propagation relationships in this solution. In the actual implementation process, the system stores the global exception propagation graph in the graph database. Taking a production system with multiple devices as an example, assume that a key device malfunctions. The system inputs information related to this exception, such as the time range of the exception occurrence and the identifier of the exception device, through the query interface of the graph database. The graph database will traverse and analyze starting from the device node where the exception occurred along the edges in the graph (these edges represent the exception association relationships between devices, and the weights of the edges represent the association strengths) according to the structure of the global exception propagation graph and the stored association information. By comparing the association strengths and exception propagation logics on different paths, the source device or event most likely to cause this exception is determined. In the industrial Internet scenario, the number of devices is large and the interconnections are complex. Traditional search methods are difficult to quickly locate the root cause of the exception. The global exception propagation graph integrates the logical associations and exception propagation information between devices, and the graph database can efficiently process this graph-structured data and quickly find the path and root cause of the exception propagation. For example, in an automobile manufacturing plant's production line, numerous devices work together. When an exception occurs in a certain production link, by querying the global exception propagation graph in the graph database, it can quickly be determined that a malfunction of a certain upstream device triggered subsequent chain reactions, rather than blindly checking all devices. This step solves the problem of difficultly quickly and accurately locating the root cause of an exception in a distributed system. The beneficial effect is that it can quickly determine the cause of the exception, provide a basis for subsequent targeted measures, reduce the impact of device failures on production, and improve production efficiency.
[0142] Step S3400: Based on the root cause location result, determine the multi-modal data and logical association channels corresponding to the root cause location result in the first-level bitmap and the second-level bitmap, and converge the data distributed on the edge physical nodes to the central governance node.
[0143] Specifically, after root cause localization, it is necessary to accurately converge associated data to avoid the communication bottleneck caused by the transfer of all data. This step realizes directional data convergence through the spatio-temporal index of the multi-level association bitmap. The primary bitmap records the distribution characteristics of the device's multi-modal data in the time and space dimensions, while the secondary bitmap superimposes the logical association information of cross-modal data. Based on the root cause localization result obtained in step S3300, the system searches in the primary bitmap for the multi-modal data generated by the devices related to the anomaly within different time windows. For example, if the root cause localization determines that the abnormal vibration data of a certain device within a specific time window is one of the reasons for the system problem, the system will find the location information of the vibration data block corresponding to that device within the time window in the primary bitmap. At the same time, in the secondary bitmap, according to the information of the logical association channel, other data blocks that have a logical association with this vibration data block are determined, such as the thermal imaging data block or text data block that may have a logical association. After determining the required data, using the spatio-temporal index information provided by the multi-level association bitmap, the relevant data shards for the query are accurately obtained. Since the data is stored on distributed edge physical nodes, through these index information, the transfer of all data can be avoided, and only the data related to the anomaly is converged to the central governance node. In actual operation, according to the node mapping code and time window information, data is read from the corresponding edge physical node. In the industrial Internet scenario, the data volume is huge and distributed among multiple edge physical nodes. If all data is transferred, it will generate extremely high communication costs and computing overheads, seriously affecting the system performance. By using the spatio-temporal index information of the multi-level association bitmap to converge data targeted, it is possible to reduce resource consumption while ensuring the acquisition of key data. This step solves the problem of excessive communication and computing overheads when obtaining relevant data in a distributed data architecture, reduces the amount of data transmission and waste of computing resources, and improves the data processing efficiency.
[0144] Step S3500, the central governance node receives the converged data and performs unified association calculations.
[0145] Specifically, the central governance node integrates the fragmented data of edge physical nodes and performs complex association calculations using the global anomaly propagation graph, overcoming the computing power limitations of edge physical nodes. Association calculation is a series of graph calculation operations carried out by the central governance node based on the aggregated data, comprehensively using the global anomaly propagation graph. Its purpose is to reveal the association characteristics between devices, estimate the scope of anomaly influence, predict the development trend of anomalies, etc. The central governance node first integrates and preprocesses the aggregated data to ensure data consistency and availability. Then, in combination with the global anomaly propagation graph, the data is analyzed. For example, by analyzing the data changes of different devices before and after the occurrence of an anomaly and their association relationships in the global anomaly propagation graph, the association characteristics between devices are revealed. For estimating the scope of anomaly influence, based on the anomaly propagation path and intensity in the global anomaly propagation graph, combined with the aggregated data, it is judged which devices may be affected by the anomaly and the degree of influence. In terms of predicting the development trend of anomalies, through the analysis of historical data and current anomaly situations, graph calculation methods are used to predict whether the anomaly will further spread and the direction and speed of the spread.
[0146] This way of performing association calculations centrally at the central governance node is because edge physical nodes usually have limited computing power and are difficult to undertake complex association calculation tasks. The central governance node has stronger computing power and can centrally process a large amount of data. In the industrial Internet scenario, it is necessary to analyze the relationships and anomaly situations between devices from a global perspective, and the centralized calculation of the central governance node can better achieve this goal. This step solves the problems that edge physical nodes lack computing power and cannot perform complex association calculations, and it is difficult to analyze device associations and anomaly situations from a global perspective, overcomes the computing power bottleneck of edge physical nodes, realizes the global optimal allocation of computing resources, and can analyze the associations and anomaly situations between devices more comprehensively and deeply, providing more powerful support for decision-making.
[0147] Embodiment 2:
[0148] Based on Embodiment 1, this embodiment provides a data governance system with a distributed data architecture, as Figure 6 shown, including:
[0149] Data encoding module: used to perform spatio-temporal feature encoding on the multi-modal data generated by devices to obtain device data carrying spatio-temporal feature encoding; the spatio-temporal feature encoding includes hash spatio-temporal anchoring and dynamic time window slicing;
[0150] Bitmap construction module: based on the device data carrying spatio-temporal feature encoding, construct a multi-level association bitmap; the multi-level association bitmap includes a first-level bitmap, a second-level bitmap, and a third-level bitmap; the first-level bitmap is a physical association bitmap, the second-level bitmap is a logical association bitmap, and the third-level bitmap is a topological association bitmap;
[0151] Collaborative computing module: Based on the constructed multi-level association bitmap, collaborative computing between edge physical nodes and central governance nodes of the distributed data architecture is realized.
[0152] In the data encoding module, the method for performing the hash time-space anchoring includes:
[0153] Step S1110, obtaining a device ID, and generating a 128-bit hash value according to the device ID, which is recorded as the device hash value;
[0154] Step S1120, intercepting the lower 8 bits of the device hash value as a physical node mapping identifier, recorded as a node mapping code;
[0155] Step S1130, marking the multimodal data generated by the same device within a preset period T as the same node mapping code; the period T is the basic time unit for data governance.
[0156] In the bitmap construction module, the multi-level associated bitmap includes a primary bitmap, a secondary bitmap and a tertiary bitmap;
[0157] The method for constructing the three-level bitmap includes:
[0158] Step S2310, constructing a global spatiotemporal correlation matrix, wherein the rows of the global spatiotemporal correlation matrix represent all devices, and the columns represent all time windows;
[0159] Step S2320, traverse the secondary bitmap of each pair of devices, if the time windows divided by the two devices in each pair of devices have overlapping areas in the absolute time interval, and strong correlation bits appear in the overlapping area, and the spatial distance between the two devices does not exceed the preset distance threshold, then set the same columns of the corresponding rows of the global spatiotemporal correlation matrix of the two devices to 1, and form a cross-device correlation tunnel;
[0160] Step S2330, calculating the topological association strength W of the cross-device association tunnel, where the topological association strength W is inversely proportional to the spatial distance D between the two devices;
[0161] Step S2340, quantize the topology association strength W into L levels, each level corresponds to a level code; fill the quantized level code into the bit corresponding to the device association tunnel to generate a three-level bitmap.
[0162] In the collaborative computing module, the method for implementing collaborative computing between edge physical nodes and central governance nodes of a distributed data architecture based on the constructed multi-level association bitmap includes:
[0163] Step S3100: at the edge physical node level, a device-node routing table is established based on the primary bitmap and the node mapping code;
[0164] Step S3200: At the central governance node level, construct a global exception propagation graph based on the secondary bitmap and the tertiary bitmap.
[0165] Step S3300: Based on the global exception propagation graph, implement distributed root cause location through querying the graph database.
[0166] Step S3400: Based on the root cause location result, determine the multimodal data and logical association channels corresponding to the root cause location result in the primary bitmap and the secondary bitmap, and converge the data distributed at the edge physical nodes to the central governance node.
[0167] Step S3500: The central governance node receives the converged data and performs unified association calculation.
[0168] The methods and systems of the present application can be implemented in many ways. For example, the methods and systems of the present application can be implemented through software, hardware, firmware, or any combination of software, hardware, and firmware. The above order of the steps for the method is only for illustration, and the steps of the method of the present application are not limited to the above specifically described order unless otherwise specifically stated.
[0169] In addition, the parts of the above technical solutions provided in the embodiments of the present application that are consistent with the corresponding technical solutions in the prior art in terms of implementation principles are not described in detail to avoid excessive elaboration.
[0170] As described above in the specific embodiments, the objectives, technical solutions, and beneficial effects of the present invention are further described in detail. It should be understood that the above is only the specific embodiments of the present invention and is not used to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A data governance method for a distributed data architecture, characterized in that, The method includes: Performing spatio-temporal feature encoding on the multi-modal data generated by the device to obtain device data carrying spatio-temporal feature encoding; the spatio-temporal feature encoding includes hash spatio-temporal anchoring and dynamic time window slicing; Based on the device data carrying spatio-temporal feature encoding, constructing a multi-level association bitmap; the multi-level association bitmap includes a first-level bitmap, a second-level bitmap, and a third-level bitmap; the first-level bitmap is a physical association bitmap, the second-level bitmap is a logical association bitmap, and the third-level bitmap is a topological association bitmap; Based on the constructed multi-level association bitmap, realizing collaborative computing between the edge physical nodes and the central governance nodes of the distributed data architecture.
2. The data governance method for a distributed data architecture according to claim 1, wherein, The device data carrying spatio-temporal feature encoding is written in the form of data blocks, and the data blocks at least include vibration data blocks and thermal imaging data blocks.
3. The data governance method for a distributed data architecture according to claim 2, characterized in that, The method for performing the hash spatio-temporal anchoring includes: Obtaining the device ID, generating a 128-bit hash value according to the device ID, denoted as the device hash value; Intercepting the lower 8 bits of the device hash value as the physical node mapping identifier, denoted as the node mapping code; Marking the multi-modal data generated by the same device within a preset period T with the same node mapping code; the period T is the basic time unit for data governance.
4. The data governance method for a distributed data architecture according to claim 3, characterized in that, The method for performing the dynamic time window slicing includes: Taking the device's first activation moment as the time origin, dividing the device operation cycle into multiple sliding time windows according to the preset period T; When the data block generated by the device is written, calculating the deviation Δτ between the time stamp of the data block and the end of the current time window, and judging the time window to which the data block belongs through the deviation Δτ to perform dynamic data sharding.
5. The data governance method of a distributed data architecture according to claim 4, wherein, The method for constructing the first-level bitmap includes: Creating a two-dimensional bit matrix for the device data carrying spatio-temporal feature encoding of each device, denoted as the device association matrix; according to the arrival time window of the device data carrying spatio-temporal feature encoding, marking the association status at the corresponding row and column positions in the device association matrix to form the first-level bitmap.
6. The data governance method for a distributed data architecture according to claim 5, characterized in that, The method for constructing the second-level bitmap includes: Performing logical association determination on the vibration data blocks and thermal imaging data blocks within the same time window; If the result of the logical association determination is that there is a logical association, creating a logical association channel within the corresponding time window row in the device association matrix; Calculating the spatio-temporal coupling degree of the logical association channel, and marking the corresponding bit positions of the logical association channel according to the spatio-temporal coupling degree to generate the second-level bitmap.
7. A data governance method for a distributed data architecture according to claim 6, characterized in that The method for performing logical association determination on the vibration data blocks and thermal imaging data blocks within the same time window includes: Extracting the vibration frequency feature of the vibration data block and the temperature difference feature of the thermal imaging data block; When the data block is written, if the vibration frequency feature of the vibration data block within the same time window exceeds the preset frequency threshold and the temperature difference feature of the thermal imaging data block exceeds the preset temperature difference threshold, it is determined that there is a logical association between the vibration data block and the thermal imaging data block within the time window.
8. A data governance method for a distributed data architecture according to claim 7, characterized in that The method for realizing collaborative computing between the edge physical nodes and the central governance nodes of the distributed data architecture includes: Based on the second-level bitmap and the third-level bitmap, constructing a global anomaly propagation graph; Based on the global anomaly propagation graph, realizing distributed root cause location through graph database query. Based on the root cause localization result, determine the multi-modal data and logical association channels corresponding to the root cause localization result in the first-level bitmap and the second-level bitmap, and converge the data distributed on the edge physical nodes to the central governance node; The central governance node receives the converged data and performs unified association calculation.
9. A data governance system for a distributed data architecture, which is used to implement the data governance method for a distributed data architecture described in any one of claims 1-8, characterized in that, The system includes: Data encoding module: used to perform spatio-temporal feature encoding on the multi-modal data generated by the device to obtain device data carrying spatio-temporal feature encoding; the spatio-temporal feature encoding includes hash spatio-temporal anchoring and dynamic time window slicing; Bitmap construction module: based on the device data carrying spatio-temporal feature encoding, construct a multi-level association bitmap; the multi-level association bitmap includes a first-level bitmap, a second-level bitmap, and a third-level bitmap; the first-level bitmap is a physical association bitmap, the second-level bitmap is a logical association bitmap, and the third-level bitmap is a topological association bitmap; Collaborative computing module: based on the constructed multi-level association bitmap, realize the collaborative computing between the edge physical nodes and the central governance node of the distributed data architecture.
Citation Information
Patent Citations
Method and device for realizing data governance
CN116257509A
Data index construction method for distributed data storage
CN119829551A
Large model interaction method and system based on multi-model collaborative dialogue
CN119557842A
Multi-modal biological characteristic bill anti-counterfeiting identification method
CN119577688A
Industrial equipment real-time monitoring system based on edge computing
CN119644972A
Cited By
Road property state monitoring data analysis method and system applying deep learning
CN120763825A
Data Analysis Method and System for Road Asset Condition Monitoring Using Deep Learning
CN120763825B