Distributed data storage calling method and system supporting Internet of Things equipment
By constructing a distributed storage-call collaboration graph on the IoT device node side, and combining communication mode vectors and multi-dimensional link features, dynamic mapping and adaptive block data fusion are achieved. This solves the problems of uneven resource utilization and unstable path selection in existing technologies, and improves the operational reliability and efficiency of IoT systems.
Patent Information
- Application Number
- CN202511470394.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-15
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2045-10-15
AI Technical Summary
Existing distributed data storage and retrieval methods lack dynamic adaptability and fail to comprehensively schedule data based on real-time communication environment, node energy status, and data requirements. This results in uneven node resource utilization and unstable data access path selection, affecting the reliability and service quality of IoT systems in large-scale and complex environments.
By establishing edge processing channels on the IoT device node side, a distributed storage-call collaboration graph is constructed. Based on communication mode vectors and multi-dimensional link characteristics, dynamic mapping and path selection are performed to achieve adaptive block data fusion and hybrid polymorphic communication, thereby optimizing access paths and resource utilization.
It improved the efficiency of node resource utilization, optimized the stability of access paths, and enhanced the overall operational reliability of large-scale IoT systems.
Smart Images

Figure CN120956744A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data storage and retrieval technology, and in particular to a distributed data storage and retrieval method and system that supports Internet of Things (IoT) devices. Background Technology
[0002] With the rapid development of IoT technology and the continuous expansion of large-scale application scenarios, more and more smart devices are connected to the network, forming a complex system composed of massive nodes. In actual operation, these nodes not only need to collect and process data, but also need to achieve efficient data transmission and retrieval to meet the needs of various scenarios such as smart manufacturing, smart cities, and telemedicine.
[0003] Currently, most existing distributed data storage and retrieval methods rely on fixed storage mapping relationships and static path selection mechanisms, lacking awareness and adaptation to real-time communication environments, node energy states, and dynamic data requirements. For example, many methods only consider the static capacity of nodes when constructing storage mappings, ignoring changes in node energy consumption, leading to some nodes being overloaded while others are idle, ultimately reducing overall efficiency.
[0004] In summary, existing technologies suffer from technical problems due to the lack of dynamic adaptability in distributed data storage and calling mechanisms. These mechanisms fail to comprehensively schedule data based on real-time communication environment, node energy status, and data requirements, resulting in uneven node resource utilization, unstable data access path selection, and decreased overall calling efficiency. This further affects the operational reliability and service quality of IoT systems in large-scale and complex environments. Summary of the Invention
[0005] The purpose of this application is to provide a method and system for distributed data storage and invocation of IoT devices, in order to solve the technical problems in the prior art where the distributed data storage and invocation mechanism lacks dynamic adaptability and fails to combine real-time communication environment, node energy status and data demand for comprehensive scheduling, resulting in uneven node resource utilization, unstable data access path selection and overall call efficiency reduction, which further affects the operational reliability and service quality of IoT systems in large-scale complex environments.
[0006] In view of the above problems, this application provides a method and system for distributed data storage and retrieval that supports IoT devices.
[0007] Firstly, this application provides a method for supporting distributed data storage and invocation of IoT devices, implemented through a distributed data storage and invocation system supporting IoT devices. The method includes: establishing an edge processing channel on the IoT device node side to perform local data acquisition and block processing, and establishing block data; performing self-checks on the IoT device node to establish a communication mode vector, which is constructed based on the real-time environment, node energy status, and data access requirements; constructing a distributed storage-invocation collaboration graph based on the communication mode vector and the block data, wherein the distributed storage-invocation collaboration graph maps data blocks to candidate nodes in a many-to-many manner, and the weight of each mapping edge is dynamically updated based on the communication mode vector; upon receiving an invocation request, identifying a subset of candidate nodes corresponding to the target data block using the distributed storage-invocation collaboration graph, and establishing an access path sequence based on the candidate node subset, mapping edge weights, and transient link indicators; adaptively fusing the data blocks returned by different access paths according to the time axis and block number, matching the adaptively fused block data results with the hybrid polymorphic communication of the upper-level device, and using the matching results for invocation management.
[0008] Preferably, the distributed data storage and invocation method supporting IoT devices further includes: the distributed storage-invocation collaboration graph is a directed multi-attribute graph. ,in, A collection of IoT device nodes. This is a set of edges mapping IoT device nodes to segmented data. This is a dynamic set of edge weights, where the weight of each mapped edge evolves and is constructed as the iteration event is invoked, and is calculated as follows: ;in, Representation at time nodes Next IoT device node For block data Access weight, Representation at time nodes Next IoT device node For block data Access weight, Characterizing IoT device nodes Communication mode vector, Characterizing IoT device nodes The node energy state, Characterizing block data The segmented load status, For block data Historical access patterns The comprehensive score function for visit suitability after normalizing the dimensions of the representation. Characterizing the time decay factor, The attenuation coefficient is... This is the smoothing coefficient.
[0009] Preferably, the method for distributed data storage and invocation supporting IoT devices further includes: obtaining a target data block according to the invocation request; identifying all communication paths of the target data block in the distributed storage-invocation collaboration graph and constructing a subset of candidate nodes; configuring a test signal and using the test signal to perform communication tests on the subset of candidate nodes, establishing transient link indicators, the transient link indicators including link bandwidth, real-time latency, packet loss rate, and link stability entropy; and performing a comprehensive adaptability score based on the mapping edge weights of the subset of candidate nodes and the transient link indicators to establish an access path sequence.
[0010] Preferably, the method for supporting distributed data storage and invocation of IoT devices further includes: obtaining a first access path sequence corresponding to the highest comprehensive adaptability score; after removing access paths corresponding to the comprehensive adaptability score based on the first access path sequence, establishing a second access path sequence based on the highest comprehensive adaptability score retained; and using the second access path sequence as a redundant path of the first access path sequence to complete the establishment of the access path sequence.
[0011] Preferably, the method for supporting distributed data storage and invocation of IoT devices further includes: sorting the data blocks returned by different access paths according to the time axis and block number to establish a multi-source segmented dataset; performing multi-dimensional feature evaluation on the multi-source segmented dataset based on data similarity features, segment importance features, temporal consistency features, and link quality features to establish a multi-dimensional feature evaluation result; and using the multi-dimensional feature evaluation result to perform adaptive segmented data fusion of the multi-source segmented dataset.
[0012] Preferably, the method for supporting distributed data storage and invocation of IoT devices further includes: configuring fusion priorities for key fusion and redundant fusion; using the fusion priorities to perform redundant fusion of high similarity blocks based on multi-dimensional feature evaluation results, and combining and fusion of high importance blocks to establish a first fusion result; and performing adaptive fusion of the remaining multi-source block datasets after removing multi-source block datasets based on the first fusion result.
[0013] Preferably, the method for supporting distributed data storage and invocation of IoT devices further includes: generating a cluster priority identifier based on the highest priority data block of each fusion cluster in the adaptive block data fusion result; when the cluster priority identifier meets a preset threshold, the corresponding fusion cluster is encapsulated into an RS485 standard data frame using a data frame encapsulation unit, and the transmission order is configured based on time division or token mechanism using an RS485 communication channel before the data transmission of the corresponding fusion cluster is executed.
[0014] Preferably, the method for calling distributed data storage for IoT devices further includes: when the cluster priority identifier does not meet a preset threshold, the corresponding fused cluster is transmitted via wireless, Ethernet or industrial bus; when all fused clusters have completed transmission, the calling task of the calling request is terminated.
[0015] Preferably, the method for supporting distributed data storage and retrieval of IoT devices further includes: preprocessing the data acquisition results, wherein the data preprocessing includes noise reduction, filtering, normalization, outlier removal and unit standardization; storing the preprocessed data in the edge processing channel, and then performing adaptive window segmentation to establish segmented data.
[0016] Secondly, this application also provides a distributed data storage and retrieval system for supporting IoT devices, used to execute a distributed data storage and retrieval method for supporting IoT devices as described in the first aspect, comprising: a block data establishment module, used to establish an edge processing channel on the IoT device node side, perform local data acquisition and block processing, and establish block data; a communication mode vector establishment module, used to perform IoT device node self-testing and establish a communication mode vector, wherein the communication mode vector is constructed according to the real-time environment, node energy status, and data access requirements; and a many-to-many mapping module, used to construct a distributed storage-retrieval collaboration graph based on the communication mode vector and the block data. The distributed storage-call collaboration graph maps data blocks to candidate nodes in a many-to-many manner, and the weight of each mapping edge is dynamically updated based on the communication mode vector. The access path sequence establishment module is used to identify the candidate node subset corresponding to the target data block using the distributed storage-call collaboration graph after receiving the call request, and establish an access path sequence based on the candidate node subset, mapping edge weight, and transient link index. The call management module is used to adaptively divide and fuse the data blocks returned by different access paths according to the time axis and block number, match the adaptive divided and fused data with the hybrid polymorphic communication of the upper-level device, and use the matching result for call management.
[0017] The technical solution provided in this application has at least the following technical effects or advantages: by achieving the technical goal of constructing a dynamic distributed storage and calling system based on communication mode vectors and multi-dimensional link characteristics, it achieves the technical effects of improving node resource utilization efficiency, optimizing access path stability, and enhancing the overall operational reliability of large-scale Internet of Things systems.
[0018] The above description is merely an overview of the technical solution of this application. To enable a clearer understanding of the technical means of this application and to facilitate its implementation according to the description, and to make the above and other objects, features, and advantages of this application more apparent, specific embodiments of this application are described below. It should be understood that the content described in this section is not intended to identify key or important features of the embodiments of this application, nor is it intended to limit the scope of this application. Other features of this application will become readily apparent through the following description. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely exemplary. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0020] Figure 1 This is a flowchart illustrating a distributed data storage and invocation method for IoT devices according to this application.
[0021] Figure 2 This is a schematic diagram of the structure of a distributed data storage and retrieval system supporting Internet of Things (IoT) devices according to this application.
[0022] Figure labeling: Block data establishment module 1, communication mode vector establishment module 2, many-to-many mapping module 3, access path sequence establishment module 4, call management module 5. Detailed Implementation
[0023] This application provides a method and system for distributed data storage and retrieval supporting IoT devices. It addresses the technical problems in existing technologies where the distributed data storage and retrieval mechanism lacks dynamic adaptability and fails to comprehensively schedule data based on real-time communication environment, node energy status, and data requirements. This results in uneven node resource utilization, unstable data access path selection, and decreased overall retrieval efficiency, further impacting the operational reliability and service quality of IoT systems in large-scale, complex environments. The application achieves the technical goal of constructing a dynamic distributed storage and retrieval system based on communication mode vectors and multi-dimensional link characteristics, thereby improving node resource utilization efficiency, optimizing access path stability, and enhancing the overall operational reliability of large-scale IoT systems.
[0024] The technical solutions of this application will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. It should be understood that this application is not limited to the exemplary embodiments described herein. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application. It should also be noted that, for ease of description, only the parts related to this application are shown in the accompanying drawings, not all of them.
[0025] Example 1, please refer to the appendix. Figure 1 This application provides a method for distributed data storage and retrieval supporting IoT devices, applied to a system for distributed data storage and retrieval supporting IoT devices, specifically including: S1: Establish an edge processing channel on the IoT device node side to perform local data collection and block processing, and create block data.
[0026] Furthermore, this application also includes: performing data preprocessing on the data acquisition results, the data preprocessing including noise reduction, filtering, normalization, outlier removal and unit standardization; storing the preprocessed data in the edge processing channel, and then performing adaptive window segmentation to establish segmented data.
[0027] Specifically, the IoT device node side refers to the location in an IoT system that is close to or directly part of the front-end device, i.e., the node that collects data. Establishing an edge processing channel on the IoT device node side allows for local data processing and computation before data is transmitted to a remote server, reducing latency and energy consumption during data transmission. For example, in a temperature and humidity sensor node, the edge processing channel can perform preliminary calculations and cache the collected temperature and humidity information.
[0028] Data acquisition results undergo preprocessing, which involves initial correction and optimization. Preprocessing includes denoising, filtering, normalization, outlier removal, and unit standardization. Denoising reduces interference signals generated during sensor acquisition using algorithms; filtering removes high-frequency or low-frequency invalid components using specific functions; normalization adjusts data of different units to the same range for comparison; outlier removal removes extreme data points caused by equipment vibration or sudden interference; and unit standardization converts data from different measurement units into a unified standard.
[0029] The preprocessed data is stored in an edge processing channel, which reduces the load on remote servers and ensures continued operation even when the network is unstable. Then, adaptive windowing is performed, dynamically adjusting the block length or data size based on the data's volatility or rate of change. For example, when the data is stable, a larger window is used to merge more data, while when the data fluctuates drastically, a smaller window is used to subdivide the data, thus creating block data that is easy to store, transmit, or retrieve later.
[0030] S2: Perform self-testing of IoT device nodes and establish a communication mode vector, which is constructed based on the real-time environment, node energy status and data access requirements.
[0031] Specifically, performing self-checks on IoT device nodes means that during the operation of IoT devices, each IoT device node needs to check its own working status. The check includes whether the hardware is normal, whether the sensors are faulty, whether the communication module is available, and whether the storage unit is abnormal. This ensures that the node is in a healthy state before participating in data collection and transmission. For example, a temperature sensor will check the battery level, whether the temperature measuring chip is stable, and whether the wireless communication interface is connected normally before it starts collecting data.
[0032] A communication pattern vector is constructed based on the real-time environment, node energy status, and data access requirements. The real-time environment refers to the physical and network conditions of the node, such as the presence of electromagnetic interference or network congestion. Node energy status refers to the device's battery level or power supply status. Data access requirements represent the frequency and urgency of data access to the node by the user or upstream system. For example, when a node has low power, its communication pattern vector will be more energy-efficient, while when data access requirements are high, it will prioritize transmission speed. Various communication-related attribute parameters of the node at the current moment are represented by mathematical vectors, including signal strength, bandwidth utilization, latency level, and packet loss, facilitating calculation and comparison in subsequent algorithms.
[0033] S3: Construct a distributed storage-call collaboration graph based on the communication mode vector and the block data. The distributed storage-call collaboration graph maps the data blocks to candidate nodes in a many-to-many manner, and the weight of each mapping edge is dynamically updated based on the communication mode vector.
[0034] Furthermore, this application also includes: the distributed storage-call collaboration graph is a directed multi-attribute graph. ,in, A collection of IoT device nodes. This is a set of edges mapping IoT device nodes to segmented data. This is a dynamic set of edge weights, where the weight of each mapped edge evolves and is constructed as the iteration event is invoked, and is calculated as follows: ;in, Representation at time nodes Next IoT device node For block data Access weight, Representation at time nodes Next IoT device node For block data Access weight, Characterizing IoT device nodes Communication mode vector, Characterizing IoT device nodes The node energy state, Characterizing block data The segmented load status, For block data Historical access patterns The comprehensive score function for visit suitability after normalizing the dimensions of the representation. Characterizing the time decay factor, The attenuation coefficient is... This is the smoothing coefficient.
[0035] Specifically, the distributed storage-call collaboration graph is a directed multi-attribute graph. The distributed storage-call collaboration graph is a graph model with direction and multiple attributes. Direction means that the connections from one node to another have an order, such as the mapping from a device node to a data block has a clear direction. Multiple attributes mean that the connections and nodes are accompanied by various parameter information, such as energy, latency, bandwidth, etc., to reflect different influencing factors.
[0036] in, An IoT device node set is a collection of IoT device nodes, where each sensor, actuator, or smart terminal is a member. This is a set of edges mapping IoT device nodes to data blocks. It indicates that in the distributed storage-call collaboration graph, there are not only device nodes, but also the correspondence between device nodes and data blocks. The edge set contains all the connections from nodes to data blocks. A dynamic set of edge weights represents the importance or suitability of a connection as described by weight values. As time or environment changes, the weights are affected by the communication mode, node energy, and data demand. For example, edge weights may increase in a high-energy state and decrease when the link is unstable.
[0037] The weight of each mapped edge is self-evolved and constructed with each iteration event, representing continuous updates and optimization as data is accessed. .in, Representation at time nodes Next IoT device node For block data Access weight, Representation at time nodes Next IoT device node For block data Access weight. Characterizing IoT device nodes The communication mode vector refers to the communication characteristics of a node, such as signal strength, delay, and bandwidth. Characterizing IoT device nodes The node energy status refers to the current remaining power or power supply status. Characterizing block data The block load status indicates the size or complexity of the data block. For block data The historical access pattern indicates the frequency and regularity of access to the data block in the past. The comprehensive scoring function for visit suitability after normalization of dimensions refers to the comprehensive score after uniformly transforming the above factors. The time decay factor refers to the mechanism by which the weight gradually decreases over time. The attenuation coefficient is the specific value of the time decay factor. The smoothing coefficient represents the adjustment parameter used to avoid drastic fluctuations during weight changes.
[0038] The distributed storage call collaboration graph maps data blocks to candidate nodes in a many-to-many manner, meaning that multiple candidate nodes can be mapped simultaneously. A candidate node is a set of devices that can participate in storage or transmission tasks. This many-to-many mapping relationship ensures redundancy and flexibility; a data block can be assigned to multiple nodes, and a single node can store multiple different data blocks.
[0039] Each mapping edge weight is dynamically updated based on the communication pattern vector. This means that the connection between a node and a data block will have a weight value to measure the reliability and adaptability of the mapping. The weight is updated over time according to the communication pattern vector, which is a set of parameters used to describe the communication capabilities of a node, including signal strength, latency, bandwidth, and packet loss rate. The weight will increase when network conditions improve and decrease when energy is insufficient or latency is too high, thus ensuring that data always tends to choose the most suitable node for storage or retrieval.
[0040] S4: After receiving the call request, identify the candidate node subset corresponding to the target data block using the distributed storage-call collaboration graph, and establish an access path sequence based on the candidate node subset, mapping edge weights and transient link indicators.
[0041] Furthermore, this application also includes: obtaining the target data block according to the call request; identifying all communication paths of the target data block in the distributed storage-call collaboration graph and constructing a subset of candidate nodes; configuring a test signal and using the test signal to perform communication tests on the subset of candidate nodes, establishing transient link indicators, the transient link indicators including link bandwidth, real-time latency, packet loss rate and link stability entropy; and performing a comprehensive adaptability score based on the mapping edge weights of the subset of candidate nodes and the transient link indicators to establish an access path sequence.
[0042] Furthermore, this application also includes: obtaining a first access path sequence corresponding to the highest comprehensive adaptability score; after removing access paths corresponding to the comprehensive adaptability score based on the first access path sequence, establishing a second access path sequence based on the highest comprehensive adaptability score retained; and using the second access path sequence as a redundant path of the first access path sequence to complete the establishment of the access path sequence.
[0043] Specifically, a call request is an instruction issued by a user or application to access or manipulate data. The target data block is a specific data unit in the storage system that needs to be accessed or processed. For example, in a distributed database, a user transaction record may be divided into several data blocks and stored on different nodes, so it is necessary to accurately locate the location of the target data block based on the call request.
[0044] Next, the distributed storage-call collaboration graph is a network structure that combines data storage location and call relationships, recording the distribution of data blocks across multiple storage nodes and the call path information. The communication path represents the connection links that data traverses when transmitted between nodes. The candidate node subset is a set of nodes with access potential that are selected by identifying all possible communication paths related to the target data block.
[0045] Next, test signals are configured, and communication tests on a subset of candidate nodes are conducted using these signals. This involves sending detection information to simulate actual communication, thereby verifying the communication capabilities of the candidate node subset and establishing transient link metrics. Transient link metrics are quantitative standards used to characterize the communication quality of a link over a short period. These metrics include link bandwidth, real-time latency, packet loss rate, and link stability entropy. Link bandwidth represents the maximum transmission volume per unit time, real-time latency represents the time consumed from sending to receiving, packet loss rate represents the proportion of data packets lost during transmission, and link stability entropy is a mathematical indicator that measures link fluctuations and uncertainties.
[0046] Obtain the first access path sequence corresponding to the highest overall adaptability score. The highest overall adaptability score is the result of comprehensive calculation of all candidate paths based on indicators such as communication quality, stability, and weight. It represents the path that the system considers optimal, while the first access path sequence is the set of optimal access routes selected according to this score.
[0047] Next, based on the first access path sequence, access paths corresponding to the comprehensive fitness score are eliminated. Access path elimination involves removing paths that have already been selected into the first access path sequence from the candidate set to avoid reuse. The remaining paths are then reordered according to the comprehensive fitness score, and the highest-scoring path forms the second access path sequence. For example, after eliminating the optimal 92-point path, the highest-scoring 88-point path among the remaining 9 paths will become the second access path sequence.
[0048] Then, the second access path sequence is defined as a redundant path of the first access path sequence. When the primary access path fails, latency suddenly increases, or the link is interrupted, it can immediately switch to the backup path to ensure uninterrupted data transmission.
[0049] S5: Adaptively divide and fuse data blocks returned from different access paths according to time axis and block number, match the adaptive data fusion results with the hybrid polymorphic communication of the upper-level device, and use the matching results for call management.
[0050] Furthermore, this application also includes: sorting the data blocks returned by different access paths according to the time axis and block number to establish a multi-source segmented dataset; performing multi-dimensional feature evaluation on the multi-source segmented dataset based on data similarity features, segment importance features, temporal consistency features, and link quality features to establish a multi-dimensional feature evaluation result; and using the multi-dimensional feature evaluation result to perform adaptive segmented data fusion of the multi-source segmented dataset.
[0051] Furthermore, this application also includes: configuring fusion priorities for key fusion and redundant fusion; using the fusion priorities to perform redundant fusion of high similarity blocks based on multi-dimensional feature evaluation results, and performing combined fusion of high importance blocks to establish a first fusion result; after removing multi-source block datasets based on the first fusion result, performing adaptive fusion of the remaining multi-source block datasets.
[0052] Furthermore, this application also includes: generating a cluster priority identifier based on the highest priority data block of each fusion cluster in the adaptive block data fusion result; when the cluster priority identifier meets a preset threshold, the corresponding fusion cluster is encapsulated into an RS485 standard data frame using a data frame encapsulation unit, and the transmission order is configured based on a time division or token mechanism using an RS485 communication channel before the data transmission of the corresponding fusion cluster is executed.
[0053] Furthermore, this application also includes: when the cluster priority identifier does not meet the preset threshold, the corresponding fused cluster is transmitted via wireless, Ethernet or industrial bus; when all fused clusters have completed transmission, the calling task of the calling request is terminated.
[0054] Specifically, the data blocks returned by different access paths are sorted according to the timeline and block number to establish a multi-source partitioned dataset. Different access paths refer to the multiple communication routes that may exist when retrieving data in a distributed storage network. Each path may return the same or different data blocks, while the timeline and block number are used to determine the generation order of the data blocks and their position in the overall data. After arranging the data blocks returned by different access paths according to time and order, a complete multi-source partitioned dataset is formed.
[0055] Next, a multi-dimensional feature evaluation is performed on the multi-source segmented dataset, including data similarity features, segment importance features, temporal consistency features, and link quality features. Data similarity features are used to measure whether the segmented content returned by different paths is consistent. Segment importance features are used to identify the position of a segment in the overall data. Temporal consistency features are used to determine the temporal coordination between segments. Link quality features reflect the reliability of the path in terms of bandwidth, latency, and packet loss rate, thereby establishing a multi-dimensional feature evaluation result.
[0056] Configure the fusion priority for critical fusion and redundant fusion, indicating which fusion blocks must be prioritized when processing multi-source segmented data. Critical fusion refers to integrating highly important or core blocks to ensure data integrity, while redundant fusion merges blocks with similar or duplicate content to reduce redundant data. Fusion priority determines the order in which fusion operations are performed during processing.
[0057] Next, based on the multidimensional feature evaluation results, highly similar blocks are merged to reduce redundant storage and transmission overhead. At the same time, high-importance blocks are combined and fused to integrate multiple key blocks into a group to establish the first fusion result. For example, if five redundant blocks have a similarity of more than 90%, they are merged into two blocks. Meanwhile, three key blocks are combined to form a fusion unit to form the first fusion result.
[0058] Then, after removing the multi-source segmented dataset based on the first fusion result, adaptive fusion of the remaining multi-source segmented dataset is performed. That is, the segments that have been integrated or merged are deleted from the original dataset, and the remaining segments are further fused according to multi-dimensional features and dynamic conditions. Adaptive fusion can flexibly adjust the processing strategy. For some segments with inconsistent timing or poor link quality, segments with reasonable timing and high quality are integrated first to obtain the final fused data.
[0059] When the cluster priority identifier does not meet the preset threshold, it means that after adaptive block data fusion, the score of the highest priority data block of a certain fusion cluster is lower than the standard set by the system. Therefore, the corresponding fusion cluster is considered not to need processing through the high-priority transmission channel, and other communication methods will be used for data transmission. Wireless refers to data transmission via wireless technologies such as Wi-Fi, Bluetooth, or LoRa; Ethernet is transmitted via wired networks; and industrial bus refers to dedicated communication buses commonly used in industrial settings such as CAN, Modbus, or Profibus.
[0060] Next, once all the merged clusters have completed transmission, meaning all the merged data clusters have been sent to the upstream device or target node, whether it is a high-priority cluster transmitted via RS485 channel or a low-priority cluster transmitted via wireless, Ethernet, or industrial bus, the transmission task has been successfully completed. Then, the call task corresponding to the call request can be marked as finished, indicating that the data access or operation process has been completed. For example, in a call with 10 merged clusters, 4 are sent via RS485 and 6 are sent via Ethernet or wireless. When the last cluster is transmitted, the data call task ends.
[0061] Next, when the cluster priority identifier meets a preset threshold, only fused clusters with a priority higher than the preset threshold will be selected for specific transmission operations. The preset threshold can be adjusted according to the application scenario. The corresponding fused cluster is encapsulated into RS485 standard data frames using a data frame encapsulation unit. The data of the fused cluster is packaged according to the frame format of the RS485 communication protocol. The data frame encapsulation unit is the functional module that performs this operation, combining multiple blocks into standard data frames for transmission on the RS485 bus. Subsequently, the data transmission of the corresponding fused cluster is executed after configuring the transmission order based on a time-division or token mechanism using the RS485 communication channel. The RS485 communication channel is the actual transmission line. The time-division mechanism refers to sending data sequentially in different time slices, while the token mechanism controls access permissions through tokens to avoid conflicts and ensure order, ensuring that important data arrives at the receiving end first.
[0062] In summary, the distributed data storage and retrieval method for supporting IoT devices provided in this application has the following technical effects: by achieving the technical goal of constructing a dynamic distributed storage and retrieval system based on communication mode vectors and multi-dimensional link characteristics, it achieves the technical effects of improving node resource utilization efficiency, optimizing access path stability, and enhancing the overall operational reliability of large-scale IoT systems.
[0063] Example 2: Based on the same inventive concept as the distributed data storage and retrieval method for supporting IoT devices in the foregoing examples, this application also provides a distributed data storage and retrieval system for supporting IoT devices. Please refer to the appendix. Figure 2The system includes: a block data establishment module 1, used to establish an edge processing channel on the IoT device node side, perform local data acquisition and block processing, and establish block data; a communication mode vector establishment module 2, used to perform IoT device node self-checks and establish a communication mode vector, which is constructed based on the real-time environment, node energy status, and data access requirements; a many-to-many mapping module 3, used to construct a distributed storage-call collaboration graph based on the communication mode vector and block data, which maps data blocks to candidate nodes in a many-to-many manner, and the weight of each mapping edge is dynamically updated based on the communication mode vector; an access path sequence establishment module 4, used to identify the candidate node subset corresponding to the target data block using the distributed storage-call collaboration graph after receiving a call request, and establish an access path sequence based on the candidate node subset, mapping edge weights, and transient link indicators; and a call management module 5, used to adaptively fuse the data blocks returned by different access paths according to the time axis and block number, match the adaptive block data fusion result with the hybrid polymorphic communication of the upper-level device, and use the matching result for call management.
[0064] Furthermore, the distributed data storage and invocation system supporting IoT devices is further configured such that: the distributed storage-invocation collaboration graph is a directed multi-attribute graph. ,in, A collection of IoT device nodes. This is a set of edges mapping IoT device nodes to segmented data. This is a dynamic set of edge weights, where the weight of each mapped edge evolves and is constructed as the iteration event is invoked, and is calculated as follows: ;in, Representation at time nodes Next IoT device node For block data Access weight, Representation at time nodes Next IoT device node For block data Access weight, Characterizing IoT device nodes Communication mode vector, Characterizing IoT device nodes The node energy state, Characterizing block data The segmented load status, For block data Historical access patterns The comprehensive score function for visit suitability after normalizing the dimensions of the representation. Characterizing the time decay factor, The attenuation coefficient is... This is the smoothing coefficient.
[0065] Furthermore, the distributed data storage and invocation system supporting IoT devices is also used for: obtaining a target data block according to the invocation request; identifying all communication paths of the target data block in the distributed storage-invocation collaboration graph and constructing a subset of candidate nodes; configuring test signals and using the test signals to conduct communication tests on the subset of candidate nodes, establishing transient link indicators, the transient link indicators including link bandwidth, real-time latency, packet loss rate, and link stability entropy; and performing a comprehensive adaptability score based on the mapping edge weights of the subset of candidate nodes and the transient link indicators to establish an access path sequence.
[0066] Furthermore, the distributed data storage and invocation system supporting IoT devices is also used to: obtain a first access path sequence corresponding to the highest comprehensive adaptability score; after removing access paths corresponding to the comprehensive adaptability score based on the first access path sequence, establish a second access path sequence based on the highest comprehensive adaptability score retained; and use the second access path sequence as a redundant path of the first access path sequence to complete the establishment of the access path sequence.
[0067] Furthermore, the distributed data storage and retrieval system supporting IoT devices is also used to: sort the data blocks returned by different access paths according to the time axis and block number to establish a multi-source segmented dataset; perform multi-dimensional feature evaluation on the multi-source segmented dataset using data similarity features, segment importance features, temporal consistency features, and link quality features to establish a multi-dimensional feature evaluation result; and use the multi-dimensional feature evaluation result to perform adaptive segmented data fusion of the multi-source segmented dataset.
[0068] Furthermore, the distributed data storage and invocation system supporting IoT devices is also used for: configuring the fusion priority of key fusion and redundant fusion; using the fusion priority to perform redundant fusion of high similarity blocks based on multi-dimensional feature evaluation results, and combining and fusion of high importance blocks to establish a first fusion result; after removing multi-source block datasets based on the first fusion result, performing adaptive fusion of the remaining multi-source block datasets.
[0069] Furthermore, the distributed data storage and invocation system supporting IoT devices is also used to: generate a cluster priority identifier based on the highest priority data block of each fusion cluster in the adaptive block data fusion result; when the cluster priority identifier meets a preset threshold, the corresponding fusion cluster is encapsulated into an RS485 standard data frame using a data frame encapsulation unit, and the transmission order is configured based on time division or token mechanism using an RS485 communication channel before the data transmission of the corresponding fusion cluster is executed.
[0070] Furthermore, the distributed data storage and invocation system supporting IoT devices is also used to: when the cluster priority identifier does not meet the preset threshold, transmit the corresponding fused cluster through wireless, Ethernet or industrial bus; when all fused clusters have completed transmission, terminate the invocation task of the invocation request.
[0071] Furthermore, the distributed data storage and retrieval system supporting IoT devices is also used for: preprocessing the data collection results, including denoising, filtering, normalization, outlier removal, and unit standardization; storing the preprocessed data to the edge processing channel, and then performing adaptive window segmentation to establish segmented data.
[0072] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The method and specific example of a distributed data storage and calling system for supporting IoT devices in the foregoing embodiment one are also applicable to the distributed data storage and calling system for supporting IoT devices in this embodiment. Through the foregoing detailed description of a distributed data storage and calling method for supporting IoT devices, those skilled in the art can clearly understand the distributed data storage and calling system for supporting IoT devices in this embodiment. Therefore, for the sake of brevity, it will not be described in detail here.
[0073] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0074] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of this application and its equivalents, this application also intends to include such modifications and variations.
Claims
1. A method for distributed data storage and retrieval supporting IoT devices, characterized in that, The method includes: Establish an edge processing channel on the IoT device node side to perform local data collection and block processing, and create block data; Perform self-testing of IoT device nodes and establish a communication mode vector, which is constructed based on the real-time environment, node energy status and data access requirements; Based on the communication mode vector and the block data, a distributed storage-call collaboration graph is constructed. The distributed storage-call collaboration graph maps data blocks to candidate nodes in a many-to-many manner, and the weight of each mapping edge is dynamically updated based on the communication mode vector. After receiving the call request, the candidate node subset corresponding to the target data block is identified using the distributed storage-call collaboration graph, and an access path sequence is established based on the candidate node subset, mapping edge weights and transient link indicators. The data blocks returned from different access paths are adaptively segmented and fused according to the time axis and block number. The adaptive segmented data fusion results are then matched with the hybrid polymorphic communication of the upper-level device, and the matching results are used for call management.
2. The method for distributed data storage and retrieval supporting IoT devices as described in claim 1, characterized in that, Based on the communication mode vector and block data, a distributed storage-recall collaboration graph is constructed, including: The distributed storage-call collaboration graph is a directed multi-attribute graph. ,in, A collection of IoT device nodes. This is a set of edges mapping IoT device nodes to segmented data. This is a dynamic set of edge weights, where the weight of each mapped edge evolves and is constructed as the iteration event is invoked, and is calculated as follows: ; in, Representation at time nodes Next IoT device node For block data Access weight, Representation at time nodes Next IoT device node For block data Access weight, Characterizing IoT device nodes Communication mode vector, Characterizing IoT device nodes The node energy state, Characterizing block data The segmented load status, For block data Historical access patterns The comprehensive score function for visit suitability after normalizing the dimensions of the representation. Characterizing the time decay factor, The attenuation coefficient is... This is the smoothing coefficient.
3. The method for distributed data storage and retrieval supporting IoT devices as described in claim 1, characterized in that, Upon receiving a call request, the distributed storage-call collaboration graph identifies a subset of candidate nodes corresponding to the target data block. Based on this subset of candidate nodes, mapping edge weights, and transient link metrics, an access path sequence is established, including: Based on the aforementioned call request, obtain the target data block; In the distributed storage-call collaboration graph, identify all communication paths of the target data block and construct a subset of candidate nodes; Configure test signals and use the test signals to perform communication tests on a subset of candidate nodes to establish transient link metrics, which include link bandwidth, real-time latency, packet loss rate, and link stability entropy. A comprehensive adaptive score is performed based on the mapping edge weights and transient link metrics of the candidate node subset to establish an access path sequence.
4. The method for distributed data storage and retrieval supporting IoT devices as described in claim 3, characterized in that, A comprehensive adaptive score is performed based on the mapping edge weights and transient link metrics of the candidate node subset to establish an access path sequence, including: Obtain the first access path sequence corresponding to the highest overall adaptability score; After eliminating the access paths corresponding to the comprehensive adaptability score based on the first access path sequence, a second access path sequence is established based on the highest comprehensive adaptability score retained. The second access path sequence is used as a redundant path of the first access path sequence to complete the establishment of the access path sequence.
5. A method for distributed data storage and retrieval supporting IoT devices as described in claim 1, characterized in that, Adaptive data fusion is performed on data blocks returned from different access paths, based on time axis and block sequence number, including: Sort the data blocks returned by different access paths according to the time axis and block number to create a multi-source block dataset; A multi-dimensional feature evaluation was performed on the multi-source segmented dataset, including data similarity features, segment importance features, temporal consistency features, and link quality features, and a multi-dimensional feature evaluation result was established. The multidimensional feature evaluation results are used to perform adaptive block data fusion of multi-source block datasets.
6. The method for distributed data storage and retrieval supporting IoT devices as described in claim 5, characterized in that, Adaptive block data fusion of multi-source block datasets using the multi-dimensional feature evaluation results includes: Configure the fusion priority for critical fusion and redundant fusion; The redundancy fusion of high similarity blocks is performed based on the multidimensional feature evaluation results using the fusion priority, and the combination fusion of high importance blocks is performed to establish the first fusion result; After removing the multi-source block dataset based on the first fusion result, adaptive fusion of the remaining multi-source block dataset is performed.
7. The method for distributed data storage and retrieval supporting IoT devices as described in claim 1, characterized in that, The adaptive block data fusion results are matched with the hybrid polymorphic communication of the upper-level device, and the matching results are used for call management, including: Generate a cluster priority identifier based on the highest priority data block of each fusion cluster in the adaptive block data fusion result; When the cluster priority identifier meets the preset threshold, the corresponding fused cluster is encapsulated into an RS485 standard data frame using a data frame encapsulation unit. After configuring the transmission order based on time division or token mechanism using the RS485 communication channel, the data transmission of the corresponding fused cluster is executed.
8. The method for distributed data storage and retrieval supporting IoT devices as described in claim 7, characterized in that, Generate a cluster priority identifier based on the highest priority data block of each fusion cluster in the adaptive block fusion result, including: If the cluster priority identifier does not meet the preset threshold, the corresponding fused cluster will transmit data via wireless, Ethernet or industrial bus. Once all fusion clusters have completed transmission, the invocation task of the invocation request ends.
9. A method for distributed data storage and retrieval supporting IoT devices as described in claim 1, characterized in that, Perform local data acquisition and block processing to create block data, including: The data acquisition results are preprocessed, including noise reduction, filtering, normalization, outlier removal, and unit standardization. After storing the preprocessed data in the edge processing channel, adaptive window segmentation is performed to create segmented data.
10. A distributed data storage and retrieval system supporting Internet of Things (IoT) devices, characterized in that, The steps for implementing a distributed data storage and retrieval method for supporting Internet of Things (IoT) devices as described in any one of claims 1 to 9 include: The block data creation module is used to establish an edge processing channel on the IoT device node side, perform local data collection and block processing, and create block data. The communication mode vector establishment module is used to perform self-testing of IoT device nodes and establish communication mode vectors. The communication mode vectors are constructed based on the real-time environment, node energy status and data access requirements. The many-to-many mapping module is used to construct a distributed storage-call collaboration graph based on the communication mode vector and the block data. The distributed storage-call collaboration graph maps the data blocks to candidate nodes in a many-to-many manner, and the weight of each mapping edge is dynamically updated based on the communication mode vector. The access path sequence establishment module is used to identify the candidate node subset corresponding to the target data block using the distributed storage-call collaboration graph after receiving the call request, and to establish an access path sequence based on the candidate node subset, mapping edge weights and transient link indicators. The call management module is used to adaptively divide and fuse data blocks returned from different access paths according to the time axis and block number, match the adaptive data fusion results with the hybrid polymorphic communication of the upper-level device, and use the matching results for call management.
Citation Information
Patent Citations
Data allocation strategy in hadoop heterogeneous cluster
CN103218233A
Big data distributed storage and parallel processing cooperation method based on cloud computing
CN120315867A
Enterprise data encryption storage method and system based on strategy optimization
CN120632915A
Network key node identification and optimization method, system and device and storage medium
CN120639590A
Method for dynamically allocating resources in an SDN / NFV network based on load balancing
US20190182169A1