Adaptive Data Management Method Based on Polymorphic Addressing
Through the polymorphic addressing adaptive data management method, the IoT data is preprocessed and multi-dimensionally classified to generate a polymorphic addressing mapping table. Distributed storage and virtual nodes are used to optimize the data storage strategy, solving the problems of low data management efficiency and thread congestion in the IoT and achieving efficient and secure data management.
Patent Information
- Application Number
- CN202510381006.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-03-28
AI Technical Summary
When processing complex and diverse data types in the Internet of Things, existing technologies have low data management efficiency and are prone to thread congestion.
An adaptive data management method based on polymorphic addressing is adopted. By preprocessing and multi-dimensionally classifying multi-scenario data, a polymorphic addressing mapping table is generated. Distributed storage and virtual nodes are used for data storage to optimize data management strategies.
It improves data management efficiency, avoids thread congestion, and ensures the security and reliability of data storage.
Smart Images

Figure CN120263802B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of Internet of Things and big data processing, and more specifically, relates to an adaptive data management method for polymorphic addressing. Background Art
[0002] IoT data is characterized by massive volume, dynamism, polymorphism, and correlation. For example, sensor data may generate tens of thousands of records per second, and device addresses may be dynamically updated as location changes. Data sources include text, coordinates, codes, and other forms, are polymorphic, and must be associated with information such as geographic location and device status. Polymorphic addressing data, as one of the many data types in the IoT, refers to the diverse address forms of IoT devices at different levels (physical, network, logical) and scenarios (static / dynamic). For example, physical identifiers: hardware unique identifiers such as MAC addresses, serial numbers, and IMEI; network addresses: protocol-layer addressing such as IPv4 / IPv6, Zigbee short addresses, and MQTT topics; logical addresses: geo-fence codes and hierarchical structures (such as factory-production line-equipment), and the address formats generated by different devices and protocols vary significantly.
[0003] Currently, when processing complex and diverse data types in the Internet of Things, due to the limitations of traditional single addressing performance, data management efficiency is low when processing complex and diverse data types in the Internet of Things, and in the process of a large amount of different data pouring in, individual threads may also be blocked. Summary of the Invention
[0004] In order to address the deficiencies in the prior art, the present invention aims to solve the above-mentioned defects and further propose an adaptive data management method with polymorphic addressing.
[0005] The present invention adopts the following technical solutions.
[0006] A first aspect of the present invention discloses a method for adaptive data management using polymorphic addressing, the method comprising:
[0007] Acquire multi-scenario data from the Internet of Things, and preprocess the multi-scenario data to construct a standardized data set;
[0008] Performing multi-dimensional classification on the multi-scenario data based on the similarity between key indicators in the standardized data set to obtain a classification table;
[0009] Obtaining a polymorphic addressing scheme corresponding to each type of data in the classification table, and generating a polymorphic addressing mapping table in combination with property parameters of each type of data;
[0010] Storing the polymorphic addressing mapping table in a distributed storage terminal, allocating addressing positions to each storage node in the distributed storage terminal through a network topology structure, and outputting a distributed storage strategy;
[0011] Constructing a virtual node associated with each storage node, and splitting the polymorphic addressing mapping table when the distributed storage strategy is executed to store the polymorphic addressing data with the first scenario label in the virtual node;
[0012] Among them, the property parameters include the prefix, attributes and scene label of each type of data, the polymorphic addressing mapping table has an address segment addressing strategy for address segment retrieval of the polymorphic addressing mapping table, the virtual node is used for distributed storage of the polymorphic addressing mapping table that meets the preset conditions, and the first scene label is the scene label defined in the preset conditions.
[0013] Furthermore, the multi-scenario data from different sources are preprocessed to construct a standardized dataset, including:
[0014] Acquire multi-scenario data from different sources, and perform field mapping and format normalization on the multi-scenario data to obtain standardized data; the standardized data is data that meets the processing requirements of the same parser;
[0015] The standardized data set is constructed based on the standardized data.
[0016] Furthermore, the multi-scenario data is classified in multiple dimensions based on the similarity between the key indicators in the standardized data set to obtain a classification table, including:
[0017] Extracting key indicators from the standardized data set, and constructing key indicator vectors based on the key indicators, calculating similarities between key indicators based on standard deviations of the key indicator vectors, and classifying the key indicators according to the similarities;
[0018] Divide the standard data in the standardized data set into multiple scene labels according to the types of the key indicators, and mark each scene label with a corresponding priority as required, so as to output a classification table containing multiple scene labels and multi-dimensional classification information;
[0019] Among them, the key indicators include the concurrency, write rate, read density and geographical span of multi-scenario data.
[0020] Furthermore, the obtaining of the polymorphic addressing scheme corresponding to each type of data in the classification table and the generation of the polymorphic addressing mapping table in combination with the property parameters of each type of data include:
[0021] Acquire multiple address segments defined based on each type of data in the classification table, and map each type of data in the classification table to a corresponding address segment to set an adjustable access priority for the address segment;
[0022] Constructing an addressing strategy unit with a combination function based on each type of data in the classification table, the polymorphic addressing scheme corresponding to each type of data, and the property parameters;
[0023] Calling the addressing strategy unit to parse the address segment according to the access priority to obtain a calling path and a redundancy strategy when reading and writing the address segment;
[0024] Determine different addressing prefix allocation results according to the key indicators corresponding to each type of scene label in the classification table, and establish an initial mapping record including an address segment-scene label mapping table to determine a write strategy and prefix consistency format in the initial mapping record;
[0025] Acquire a defined addressing attribute matrix according to the addressing prefix allocation result and the initial mapping record;
[0026] Calculating an attribute priority score based on the addressing prefixes and corresponding attributes and attribute weights in the addressing attribute matrix;
[0027] The attribute priority score, the call path and the redundancy strategy are integrated to generate the polymorphic addressing mapping table and the address segment addressing strategy.
[0028] Furthermore, the polymorphic addressing mapping table is stored in a distributed storage terminal, and addressing positions are allocated to each storage node in the distributed storage terminal through a network topology structure, and a distributed storage strategy is output, including:
[0029] Parsing the hardware characteristics and network connection relationships of all current storage nodes to construct a storage node loading table, and storing the polymorphic addressing mapping table in the distributed storage end according to the storage node loading table, wherein the storage node loading table includes resource quotas and read and write performance parameters of each storage node;
[0030] Coupling and allocating the address prefixes in the polymorphic addressing mapping table in the distributed storage end with the storage node loading table, and calculating the network cost of deploying each address segment on different storage nodes, so as to select a storage node or storage node combination with the minimum network cost for each address prefix, and outputting a coupled addressing location allocation strategy;
[0031] Determine a judgment result for indicating whether any two storage nodes are directly connected based on the network topology, and fine-tune the storage strategy for each address segment between the storage nodes to obtain a fine-tuned storage strategy;
[0032] The fine-tuned storage strategy is integrated with the addressing location allocation strategy to obtain the distributed storage strategy of multi-storage node fusion.
[0033] Furthermore, the step of constructing a virtual node associated with each storage node and, when the distributed storage strategy is executed, splitting the polymorphic addressing mapping table to store the polymorphic addressing data having the first scenario label in the virtual node includes:
[0034] Build virtual nodes associated with each storage node;
[0035] Reading scene tag information from the polymorphic addressing mapping table, and splitting the polymorphic addressing data having the first scene tag from the polymorphic addressing mapping table when executing a distributed storage strategy corresponding to the scene tag information;
[0036] Addressing positions are allocated to the virtual nodes through a ring or tree topology structure, so as to store the polymorphic addressed data with the first scene tag in each virtual node.
[0037] The present invention further discloses a polymorphic addressing adaptive data management system for implementing the steps of the polymorphic addressing adaptive data management method described in the first aspect. The system comprises:
[0038] A data preprocessing module is used to obtain multi-scenario data from the Internet of Things and preprocess the multi-scenario data to construct a standardized data set;
[0039] A data classification module, configured to perform multi-dimensional classification on the multi-scenario data based on the similarity between key indicators in the standardized data set to obtain a classification table;
[0040] A mapping table generation module is used to obtain the polymorphic addressing scheme corresponding to each type of data in the classification table, and generate a polymorphic addressing mapping table and an address segment addressing strategy in combination with the property parameters of each type of data;
[0041] A distributed storage module, configured to store the polymorphic addressing mapping table in a distributed storage terminal, allocate addressing positions to storage nodes in the distributed storage terminal through a network topology, and output a distributed storage strategy;
[0042] an addressing data transfer module, configured to construct a virtual node associated with each storage node and, when the distributed storage strategy is executed, split the polymorphic addressing mapping table to store the polymorphic addressing data having the first scenario label in the virtual node;
[0043] Among them, the property parameters include the prefix, attributes and scene label of each type of data, the polymorphic addressing mapping table has an address segment addressing strategy for address segment retrieval of the polymorphic addressing mapping table, the virtual node is used for distributed storage of the polymorphic addressing mapping table that meets the preset conditions, and the first scene label is the scene label defined in the preset conditions.
[0044] A third aspect of the present invention discloses a terminal, comprising a processor and a storage medium, characterized in that:
[0045] The storage medium is used to store instructions;
[0046] The processor is configured to operate according to the instructions to execute the steps of the method of the first aspect.
[0047] A fourth aspect of the present invention discloses a computer-readable storage medium having a computer program stored thereon, wherein the program implements the steps of the method described in the first aspect when executed by a processor.
[0048] The beneficial effects of the present invention are that, compared with the prior art, the present invention has the following advantages:
[0049] (1) The present invention performs feature analysis and multi-dimensional classification on multi-scenario data from different sources, and generates a corresponding polymorphic addressing scheme, a polymorphic addressing mapping table, and a corresponding address segment addressing strategy for each type of data in the classification table. The polymorphic addressing mapping table is stored in a distributed storage terminal, and the addressing position of each storage node in the distributed storage terminal is allocated through the network topology structure, effectively avoiding the problem that a single address is difficult to load complex multi-scenario data. In addition, the present invention performs orderly distributed storage of complex multi-scenario data, and shares complex multi-scenario data through multiple storage nodes, which not only avoids thread congestion when reading and writing massive data, but also improves the efficiency of data management.
[0050] (2) The present invention constructs a virtual node associated with each storage node, and splits the polymorphic addressing mapping table when the distributed storage strategy is executed, so as to distribute the polymorphic addressing data with special set scene labels to the virtual nodes, so that the polymorphic addressing data that may cause thread congestion or abnormal situations (such as addressing data with cross-node shared tags, addressing data with scene labels of "high throughput computing" or "ultra-low latency") can be stored in the virtual nodes in advance to prevent it from affecting each storage node and other threads, further reducing the abnormal situations when reading and writing complex multi-scene data, and improving data management efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 The present invention is a flowchart of a polymorphic addressing adaptive data management method. DETAILED DESCRIPTION
[0052] The present application will be further described below in conjunction with the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and are not intended to limit the scope of protection of the present application.
[0053] like Figure 1 As shown, in one embodiment, a method for adaptive data management with polymorphic addressing includes the following steps:
[0054] Step S110: Acquire multi-scenario data from the Internet of Things and pre-process the multi-scenario data to construct a standardized data set.
[0055] In some embodiments, the polymorphic addressing adaptive data management method provided by the present invention, step S110 specifically includes the following steps:
[0056] Step S111, obtain multi-scenario data from different sources, and perform field mapping and format normalization processing on the multi-scenario data to obtain standardized data. The standardized data is data that meets the processing requirements of the same parser.
[0057] Step S112: construct a standardized data set based on the standardized data.
[0058] Step S120 , performing multi-dimensional classification on the multi-scenario data based on the similarity between the key indicators in the standardized data set to obtain a classification table.
[0059] Among them, each type of data in the classification table has a corresponding scene label.
[0060] In some embodiments, the polymorphic addressing adaptive data management method provided by the present invention, step S120 specifically includes the following steps:
[0061] Step S121 , extracting key indicators from the standardized data set, and constructing key indicator vectors based on the key indicators, calculating the similarity between the key indicators based on the standard deviation of the key indicator vectors, and classifying the key indicators according to the similarity.
[0062] In step S122, the standard data in the standardized data set is divided into multiple scene labels according to the types of key indicators, and each scene label is marked with a corresponding priority as required to output a classification table containing multiple scene labels and multi-dimensional classification information.
[0063] Among them, key indicators include the concurrency, write rate, read density and geographical span of multi-scenario data.
[0064] Step S130 , obtaining the polymorphic addressing scheme corresponding to each type of data in the classification table, and generating a polymorphic addressing mapping table in combination with the property parameters of each type of data.
[0065] The property parameters include a prefix, an attribute, and a scenario tag of each type of data, and the polymorphic addressing mapping table has an address segment addressing strategy for performing address segment retrieval on the polymorphic addressing mapping table.
[0066] In some embodiments, the polymorphic addressing adaptive data management method provided by the present invention, step S130 specifically includes the following steps:
[0067] Step S131 , obtaining multiple address segments defined based on each type of data in the classification table, and mapping each type of data in the classification table to a corresponding address segment, so as to set an adjustable access priority for the address segment.
[0068] Step S132: constructing an addressing strategy unit with a combination function based on each type of data in the classification table, the polymorphic addressing scheme corresponding to each type of data, and the property parameters.
[0069] Step S133 , calling the addressing strategy unit to parse the address segment according to the access priority, and obtaining the calling path and redundancy strategy when reading and writing the address segment.
[0070] Step S134, determine different addressing prefix allocation results according to the key indicators corresponding to each type of scene label in the classification table, and establish an initial mapping record containing an address segment-scene label mapping table to determine the writing strategy and prefix consistency format in the initial mapping record.
[0071] Step S135: Acquire the defined addressing attribute matrix according to the addressing prefix allocation result and the initial mapping record.
[0072] Step S136: Calculate the attribute priority score based on the addressing prefix and the corresponding attribute and attribute weight in the addressing attribute matrix.
[0073] Step S137 : Integrate the attribute priority score, the call path, and the redundancy strategy to generate a polymorphic addressing mapping table and an address segment addressing strategy.
[0074] Step S140: store the polymorphic addressing mapping table in the distributed storage end, allocate addressing positions to each storage node in the distributed storage end through the network topology structure, and output a distributed storage strategy.
[0075] In some embodiments, the polymorphic addressing adaptive data management method provided by the present invention, step S140 specifically includes the following steps:
[0076] Step S141, analyze the hardware characteristics and network connection relationships of all current storage nodes to construct a storage node loading table, and store the polymorphic addressing mapping table in the distributed storage end according to the storage node loading table. The storage node loading table contains the resource quota and read and write performance parameters of each storage node.
[0077] In step S142, the addressing prefixes in the polymorphic addressing mapping table in the distributed storage end are coupled and allocated with the storage node loading table, and the network cost of deploying each address segment on different storage nodes is calculated to select the storage node or storage node combination with the minimum network cost for each addressing prefix, and output the coupled addressing location allocation strategy.
[0078] Step S143 , determining a judgment result for characterizing whether any two nodes are directly connected based on the network topology, and fine-tuning the storage strategy of each address segment between each storage node to obtain a fine-tuned storage strategy.
[0079] Step S144 , integrating the fine-tuned storage strategy with the addressing location allocation strategy to obtain a distributed storage strategy integrating multiple storage nodes.
[0080] Step S150: construct a virtual node associated with each storage node, and when the distributed storage strategy is executed, split the polymorphic addressing mapping table to store the polymorphic addressing data with the first scenario label in the virtual node.
[0081] Among them, the virtual node is a virtual distributed storage end, which is used to distribute the polymorphic addressing mapping table that meets the preset conditions. The first scenario label is the scenario label defined in the preset condition and can be pre-set. For example, for addressing data with a cross-node shared label, a ring or tree topology can be used to allocate the addressing position of the virtual node to avoid the concentration of data copies in the storage node to cause hot spots. For another example, for the scenario label of "high throughput computing" or "ultra-low latency", targeted scheduling is performed and it is scheduled to the constructed virtual node, that is, the core data is copied to the virtual node with higher bandwidth or lower latency to prevent the loss of core data and ensure the security of data management.
[0082] In some embodiments, the polymorphic addressing adaptive data management method provided by the present invention, step S150 specifically includes the following steps:
[0083] Step S151: construct a virtual node associated with each storage node.
[0084] Step S152 : Reading the scene tag information from the polymorphic addressing mapping table, and splitting the polymorphic addressing data with the first scene tag from the polymorphic addressing mapping table when executing the distributed storage strategy corresponding to the scene tag information.
[0085] Step S153 : allocating addressing positions to the virtual nodes through a ring or tree topology structure, so as to store the polymorphic addressing data with the first scene tag in each virtual node.
[0086] In a specific embodiment, the polymorphic addressing adaptive data management method provided by the present invention includes steps 1 to 4:
[0087] Step 1: Multi-scenario data analysis and multi-dimensional classification.
[0088] In order to solve the inefficiency of traditional single addressing when processing complex data types in the Internet of Things, we first analyze the various application scenarios and data access characteristics that may appear in the system to form a multi-dimensional classification result.
[0089] It's important to note that polymorphic addressing refers to the diverse address formats used by IoT devices at different levels (physical, network, logical) and in different scenarios (static / dynamic). Examples include physical identifiers (MAC addresses, serial numbers, IMEI, and other hardware unique identifiers); network addresses (IPv4 / IPv6, Zigbee short addresses, MQTT topics, and other protocol-layer addresses); and logical addresses (geo-fence codes and hierarchical structures, such as factory-production line-equipment). Therefore, polymorphic addressing in the IoT is characterized by diverse scenario representations, diverse scenarios, and diverse address formats, with significant differences in the address formats generated by different devices and protocols.
[0090] Specifically, it includes steps 1.1 to 1.3:
[0091] Step 1.1, read the classification table and establish the processing mapping.
[0092] Collect logs or data from different sources in the IoT, namely multi-scenario data, such as "video surveillance high-frequency nighttime write logs" (timestamp, write rate), "e-commerce flash sale order records" (number of orders, peak QPS), and "IoT (Internet of Things) sensor batch reporting data" (sensor type, geographic location, reporting period). Perform field mapping and format normalization to ensure that all data can be processed by the same parser, and finally output a standardized data set for subsequent analysis.
[0093] In step 1.2, based on the standardized log set, core indicators are extracted to calculate the similarity matrix.
[0094] Key indicators include concurrency C, write rate W, read density R, geographic span G, etc. The specific similarity formula can be calculated using the standard deviation based on the vector composed of key indicators.
[0095] Step 1.3: Scene label division and priority annotation.
[0096] The multi-scenario data records are divided into several scenario tags based on key characteristics such as concurrency, writing, and reading. For example, "high concurrency + high-speed writing + medium geographical distribution", "medium concurrency + extremely high-speed writing + high geographical distribution", and "low concurrency + batch reading + negligible geographical span" are added. Each scenario tag is given a corresponding priority (for example, the highest security requirements are assigned to financial transaction scenarios, and the geographical dimension is strengthened for IoT scenarios). The final output is a classification table containing multiple scenario tags and multi-dimensional classification information.
[0097] Step 2: Construct the polymorphic addressing mapping table.
[0098] Based on the classification table from step 1, a corresponding polymorphic addressing scheme is generated for each type of data, and a polymorphic addressing mapping table is generated within the system. First, various address segments or page attributes are defined, such as "cache page," "streaming mapping segment," and "cross-node shared tag." Each type of data in the classification table is then mapped to a corresponding address segment, and adjustable access priorities are set for it (such as CPU / GPU parallel access and remote cold storage). Next, a composable addressing strategy unit is introduced, which parses address segment tags to determine the call path and redundancy strategy for data reading and writing (for example, enabling near-end cache pass-through for "cache page," adopting prefetching and pipelined write strategies for "streaming mapping segment," and linking the consistency protocol for multi-node synchronization for "shared tag"). By assigning addressing forms to different data according to local conditions, the high concurrency conflicts, latency jitter, and space waste that arise in a single addressing mode are resolved, making the system more scalable.
[0099] Specifically, it includes steps 2.1 to 2.3:
[0100] Step 2.1, read the classification table and assign an addressing prefix.
[0101] Assign different address prefixes (e.g., high-speed segment, cold storage segment, distributed shared segment) based on the performance, security, or geographic dimensions corresponding to each scenario tag in the table (e.g., "high concurrency + strong consistency," "high-volume read-only + cross-region," etc.). Create an initial mapping record to form a list of address segments and scenario tags, indicating the required write strategy or consistency mode.
[0102] For example, if the label "high concurrency + strong consistency" corresponds to the high-speed page prefix "HSP_", then it is written as "HSP_ → high concurrency + strong consistency" in the mapping record.
[0103] Step 2.2, addressing attribute matrix and addressing strategy derivation.
[0104] Specifically, first, define an addressing attribute matrix, which consists of the addressing prefix and its attribute values (write amplification factor, access latency, and fault tolerance level), as well as the performance cost or benefit of each addressing prefix on any attribute. Secondly, calculate the priority score of the corresponding addressing data based on the attribute weight corresponding to each addressing prefix and the corresponding value in the addressing attribute matrix. Among them, addressing prefixes with high priority scores are adapted to high-priority scenarios, and those with low priority scores are adapted to secondary priority scenarios. Based on this, an addressing strategy can be formed that establishes a mapping with "addressing prefix-priority score-scenario label".
[0105] Step 2.3: Strategy mapping integration and output.
[0106] Specifically, the address prefix, attribute score and scene label are unified and integrated to obtain a complete polymorphic address mapping table and address segment addressing strategy. In this process, it is necessary to first determine the write strategy (such as sequential write, block write, parallel copy) and read strategy (cache priority, distributed consistency read, etc.) corresponding to each address prefix, and then the polymorphic address mapping table and address segment addressing strategy
[0107] For example:
[0108] "HSP_ prefix → high concurrency + strong consistency scenario; write strategy = multi-way parallel, read strategy = local cache";
[0109] "CS_ prefix → large-scale read-only + cross-region scenario; write strategy = periodic batch synchronization, read strategy = cross-node load balancing";
[0110] "BK_ prefix → low priority archiving scenario; write strategy = compressed storage, read strategy = delayed access".
[0111] Step 3: Distributed topology adaptation and multi-storage node integration.
[0112] After the polymorphic addressing mapping is generated, it needs to be integrated into the distributed environment to accommodate resource differences between storage nodes and leverage network topology features to accelerate data access. First, a "polymorphic addressing agent" is deployed on each storage node. This agent reads the scenario tag information from the mapping table in step 2 and executes the corresponding addressing strategy. For data with shared tags across nodes, a ring or tree topology can be used to allocate address locations across storage nodes to avoid data replication consolidation and hotspots. Next, node hardware features (such as GPU computing accelerators, FPGA dedicated processing, and NVM storage media) are combined to implement targeted scheduling for scenario tags of "high throughput computing" or "ultra-low latency," placing core data on nodes with higher bandwidth or lower latency. Finally, a distributed index and routing table is established. Once a storage node receives an access request, it determines whether to process it locally or forward it across nodes based on the scenario tag of the address segment in the polymorphic addressing mapping table, improving access efficiency and reducing unnecessary network hops.
[0113] The above-mentioned fusion of multiple storage nodes effectively connects polymorphic addressing to large-scale distributed systems, which can further solve the performance bottlenecks and resource waste of the traditional "single mapping + centralized storage" in heterogeneous clusters.
[0114] Specifically, it includes steps 3.1 to 3.3:
[0115] Step 3.1, node loading and network topology scanning.
[0116] Specifically, the hardware characteristics (such as the number of CPU cores, bandwidth, storage media) and network connection relationships (such as ring or tree topology) of all storage nodes in the current cluster are analyzed, and a storage node loading table is constructed. The table consists of multiple storage nodes and the quota or performance value (such as available bandwidth, maximum number of concurrent connections) that each storage node can provide for any resource. The value can be obtained through system monitoring or node self-test, and the node loading information is output for subsequent distributed adaptation algorithms.
[0117] For example, if storage node X1 has a Gigabit network port and SSD storage, and X2 has a 10Gb network port and NVMe storage, then the storage node loading table records information such as X1 (bandwidth = 1, storage medium = SSD) and X2 (bandwidth = 10, storage medium = NVMe).
[0118] Step 3.2: Addressing and node coupling allocation.
[0119] Specifically, the address prefixes (such as HSP_, CS_, and BK_) in the polymorphic addressing mapping table are coupled with the storage node load table to calculate the overall network cost of deploying each address segment on different storage nodes. This network cost represents the combined cost of deploying any address prefix on any storage node. It is determined by the storage node's weight in the corresponding resource dimension, the demand for that resource dimension, and the corresponding load penalty value, which is used to prevent overloading of individual storage nodes. Finally, the storage node or node combination with the lowest network cost is selected for each address prefix, and the address segments are mapped to the corresponding locations, forming a preliminary distributed layout. This output is the coupled address location allocation decision.
[0120] For example, if the HSP_ prefix has a high network bandwidth requirement (that is, resource dimension requirement = 0.8), and storage node X2 has 10Gb bandwidth (resource dimension = 0.6, load penalty value = 0.1), it can be calculated that the network cost of node X2 is lower than the network cost of node X1, so node X2 is preferred.
[0121] Step 3.3: Strategy fine-tuning and multi-node fusion.
[0122] Specifically, an adjacency matrix is constructed based on the network topology. This adjacency matrix is used to represent whether node a and node b are directly connected. For example, if the adjacency matrix value of node a and node b is 1, it means that node a and node b are directly connected, and if the adjacency matrix value is 0, it means that there is no direct connection between the two. Afterwards, the replication strategy of each address segment between each storage node is fine-tuned to obtain a fine-tuned storage strategy. If the communication cost between adjacent nodes is low, local mirroring or replica synchronization is prioritized to reduce traffic across remote links. Finally, the fine-tuned storage strategy is integrated with the addressing location allocation strategy to output the final multi-node fusion distributed topology adaptation strategy.
[0123] For example, if the adjacency matrix value between node X2 and node X3 in the local area network is 1 and the bandwidth is large, a mirror copy of the HSP_ prefix can be set at node X3 to achieve fast synchronization and fault tolerance.
[0124] Step 4: construct a virtual node associated with each storage node and split the polymorphic addressing data with the first scenario label.
[0125] Specifically, when executing the final distributed storage strategy, the scene tag information is read from the polymorphic addressing mapping table, and the address segment addressing strategy corresponding to the scene tag information is executed. Then, polymorphic addressing data with special scene tags is selected from the polymorphic addressing mapping table. For example, addressing data with a cross-node shared tag (i.e., the first scene tag mentioned above) can be allocated to the virtual nodes using a ring or tree topology to avoid data copies being concentrated in storage nodes and causing hot spots. For another example, targeted scheduling is performed for scenes with the tag "high throughput computing" or "ultra-low latency" (i.e., the first scene tag mentioned above), and it is scheduled to the constructed virtual node, that is, the core data is copied to the virtual node with higher bandwidth or lower latency to prevent the loss of core data and ensure the security of data management.
[0126] The following describes the adaptive data management system for polymorphic addressing provided by the present invention. The adaptive data management system for polymorphic addressing described below and the adaptive data management method for polymorphic addressing described above can be referred to in correspondence with each other.
[0127] In one embodiment, a polymorphic addressing adaptive data management system includes a data preprocessing module, a data classification module, a mapping table generation module, a distributed storage module, and an addressing data transfer module.
[0128] The data preprocessing module is used to obtain multi-scenario data from the Internet of Things and preprocess the multi-scenario data to construct a standardized data set.
[0129] The data classification module is used to perform multi-dimensional classification on multi-scenario data based on the similarity between key indicators in the standardized data set to obtain a classification table.
[0130] The mapping table generation module is used to obtain the polymorphic addressing scheme corresponding to each type of data in the classification table, and generate a polymorphic addressing mapping table in combination with the property parameters of each type of data.
[0131] The distributed storage module is used to store the polymorphic addressing mapping table in the distributed storage end, and allocate addressing positions to each storage node in the distributed storage end through the network topology structure, and output a distributed storage strategy.
[0132] The addressing data transfer module is used to construct a virtual node associated with each storage node, and when the distributed storage strategy is executed, split the polymorphic addressing mapping table to store the polymorphic addressing data with the first scenario label to the virtual node.
[0133] Among them, each type of data in the classification table has a corresponding scene label, the property parameters include the prefix, attribute and scene label of each type of data, the polymorphic addressing mapping table has an address segment addressing strategy for address segment retrieval of the polymorphic addressing mapping table, the virtual node is a virtual distributed storage end, which is used for distributed storage of polymorphic addressing mapping tables that meet preset conditions, and the first scene label is the scene label defined in the preset conditions.
[0134] In this embodiment, the data preprocessing module of the polymorphic addressing adaptive data management system provided by the present invention is specifically used to:
[0135] Acquire multi-scenario data from different sources, perform field mapping and format normalization on the multi-scenario data, and obtain standardized data. Standardized data is data that meets the processing requirements of the same parser.
[0136] Based on the standardized data, a standardized dataset is constructed.
[0137] In this embodiment, the data classification module of the polymorphic addressing adaptive data management system provided by the present invention is specifically used to:
[0138] Key indicators are extracted from the standardized data set, and key indicator vectors are constructed based on the key indicators. The similarity between key indicators is calculated based on the standard deviation of the key indicator vectors, and the key indicators are classified according to the size of the similarity.
[0139] The standard data in the standardized data set is divided into multiple scene labels according to the types of key indicators, and each scene label is marked with the corresponding priority according to the needs to output a classification table containing multiple scene labels and multi-dimensional classification information.
[0140] Among them, key indicators include the concurrency, write rate, read density and geographical span of multi-scenario data.
[0141] In this embodiment, the polymorphic addressing adaptive data management system provided by the present invention, the mapping table generation module is specifically used to:
[0142] A plurality of address segments defined based on each type of data in the classification table are obtained, and each type of data in the classification table is mapped to a corresponding address segment to set an adjustable access priority for the address segment.
[0143] An addressing strategy unit with combination function is constructed based on each type of data in the classification table, the polymorphic addressing scheme and property parameters corresponding to each type of data.
[0144] The call addressing strategy unit parses the address segment according to the access priority to obtain the call path and redundancy strategy when reading and writing the address segment.
[0145] Different addressing prefix allocation results are determined according to the key indicators corresponding to each type of scene label in the classification table, and an initial mapping record containing an address segment-scene label mapping table is established to determine the writing strategy and prefix consistency format in the initial mapping record.
[0146] The defined addressing attribute matrix is obtained according to the addressing prefix allocation result and the initial mapping record.
[0147] The attribute priority score is calculated based on the addressing prefix and the corresponding attribute and attribute weight in the addressing attribute matrix.
[0148] The attribute priority scores, call paths and redundancy strategies are integrated to generate a polymorphic addressing mapping table and address segment addressing strategy.
[0149] In this embodiment, the distributed storage module of the polymorphic addressing adaptive data management system provided by the present invention is specifically used for:
[0150] The hardware characteristics and network connection relationships of all current storage nodes are analyzed, and the polymorphic addressing mapping table is stored in the distributed storage end according to the storage node loading table to construct the storage node loading table. The storage node loading table contains the resource quota and read and write performance parameters of each storage node.
[0151] The addressing prefixes in the polymorphic addressing mapping table in the distributed storage end are coupled and allocated with the storage node loading table, and the network cost of deploying each address segment on different storage nodes is calculated to select the storage node or storage node combination with the minimum network cost for each addressing prefix, and output the coupled addressing location allocation strategy.
[0152] A judgment result for characterizing whether any two storage nodes are directly connected is determined based on the network topology structure, and a storage strategy for each address segment between the storage nodes is fine-tuned to obtain a fine-tuned storage strategy.
[0153] The fine-tuned storage strategy is integrated with the addressing location allocation strategy to obtain a distributed storage strategy that integrates multiple storage nodes.
[0154] In this embodiment, the polymorphic addressing adaptive data management system provided by the present invention, the addressing data transfer module is specifically used to:
[0155] Build a virtual node associated with each storage node.
[0156] The scene tag information is read from the polymorphic addressing mapping table, and when a distributed storage strategy corresponding to the scene tag information is executed, the polymorphic addressing data with the first scene tag is split from the polymorphic addressing mapping table.
[0157] Addressing positions are allocated to the virtual nodes through a ring or tree topology structure, so as to store the polymorphic addressed data with the first scene tag in each virtual node.
[0158] The present disclosure may be a system, method and / or computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.
[0159] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punched card or raised structure in a groove on which instructions are stored, and any suitable combination thereof. As used herein, a computer-readable storage medium is not to be construed as a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse passing through a fiber optic cable), or an electrical signal transmitted through an electrical wire.
[0160] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in the computer-readable storage medium in each computing / processing device.
[0161] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, the state information of the computer-readable program instructions is used to personalize an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), so that the electronic circuit can execute the computer-readable program instructions, thereby implementing various aspects of the present disclosure.
[0162] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0163] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processor of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other equipment to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.
[0164] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more blocks in the flowchart and / or block diagram.
[0165] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple embodiments of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and the part of the module, program segment or instruction contains one or more executable instructions for realizing the prescribed logical function. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the prescribed function or action, or can be implemented by a combination of dedicated hardware and computer instructions.
[0166] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered by the scope of protection of the claims of the present invention.
Claims
1. A method for adaptive data management with polymorphic addressing, characterized in that: The method comprises: Acquire multi-scenario data from the Internet of Things and pre-process the multi-scenario data to construct a standardized dataset; Perform multi-dimensional classification on multi-scenario data based on the similarity between key indicators in the standardized data set to obtain a classification table; Obtain the polymorphic addressing scheme corresponding to each type of data in the classification table, and generate a polymorphic addressing mapping table based on the property parameters of each type of data; The polymorphic address mapping table is stored in the distributed storage end, and the address position of each storage node in the distributed storage end is allocated through the network topology structure, and the distributed storage strategy is output; Constructing a virtual node associated with each storage node, and splitting the polymorphic addressing mapping table when the distributed storage strategy is executed to store the polymorphic addressing data with the first scenario label in the virtual node; The property parameters include a prefix, an attribute, and a scenario tag for each type of data; the polymorphic addressing mapping table has an address segment addressing strategy for performing address segment retrieval on the polymorphic addressing mapping table; the virtual node is used to perform distributed storage on the polymorphic addressing mapping table that meets the preset conditions; and the first scenario tag is a scenario tag defined in the preset conditions. The step of obtaining the polymorphic addressing scheme corresponding to each type of data in the classification table and generating a polymorphic addressing mapping table in combination with the property parameters of each type of data includes: Acquire multiple address segments defined based on each type of data in the classification table, and map each type of data in the classification table to a corresponding address segment to set an adjustable access priority for the address segment; Constructing an addressing strategy unit with a combination function based on each type of data in the classification table, the polymorphic addressing scheme corresponding to each type of data, and the property parameters; Calling the addressing strategy unit to parse the address segment according to the access priority to obtain a calling path and a redundancy strategy when reading and writing the address segment; Determine different addressing prefix allocation results according to the key indicators corresponding to each type of scene label in the classification table, and establish an initial mapping record including an address segment-scene label mapping table to determine a write strategy and prefix consistency format in the initial mapping record; Acquire a defined addressing attribute matrix according to the addressing prefix allocation result and the initial mapping record; Calculating an attribute priority score based on the addressing prefixes and corresponding attributes and attribute weights in the addressing attribute matrix; The attribute priority score, the call path and the redundancy strategy are integrated to generate the polymorphic addressing mapping table with the address segment addressing strategy.
2. The method for adaptive data management with polymorphic addressing according to claim 1, characterized in that: The step of acquiring multi-scenario data from the Internet of Things and preprocessing the multi-scenario data to construct a standardized data set includes: Acquire multi-scenario data from the Internet of Things, and perform field mapping and format normalization on the multi-scenario data to obtain standardized data; the standardized data is data that meets the processing requirements of the same parser; The standardized data set is constructed based on the standardized data.
3. The method for adaptive data management with polymorphic addressing according to claim 2, wherein: The multi-scenario data is classified in multiple dimensions based on the similarity between the key indicators in the standardized data set to obtain a classification table, including: Extracting key indicators from the standardized data set, and constructing key indicator vectors based on the key indicators, calculating similarities between key indicators based on standard deviations of the key indicator vectors, and classifying the key indicators according to the similarities; Divide the standard data in the standardized data set into multiple scene labels according to the types of the key indicators, and mark each scene label with a corresponding priority as required, so as to output a classification table containing multiple scene labels and multi-dimensional classification information; Among them, the key indicators include the concurrency, write rate, read density and geographical span of multi-scenario data.
4. The method for adaptive data management with polymorphic addressing according to claim 1, wherein: The step of storing the polymorphic addressing mapping table in a distributed storage terminal, allocating addressing positions to storage nodes in the distributed storage terminal through a network topology structure, and outputting a distributed storage strategy includes: Parsing the hardware characteristics and network connection relationships of all current storage nodes to construct a storage node loading table, and storing the polymorphic addressing mapping table in the distributed storage end according to the storage node loading table, wherein the storage node loading table includes resource quotas and read and write performance parameters of each storage node; Coupling and allocating the address prefixes in the polymorphic addressing mapping table in the distributed storage end with the storage node loading table, and calculating the network cost of deploying each address segment on different storage nodes, so as to select a storage node or storage node combination with the minimum network cost for each address prefix, and outputting a coupled addressing location allocation strategy; Determine a judgment result for indicating whether any two storage nodes are directly connected based on the network topology, and fine-tune the storage strategy for each address segment between the storage nodes to obtain a fine-tuned storage strategy; The fine-tuned storage strategy is integrated with the addressing location allocation strategy to obtain the distributed storage strategy of multi-storage node fusion.
5. The method for adaptive data management with polymorphic addressing according to claim 4, characterized in that: The step of constructing a virtual node associated with each storage node and splitting the polymorphic addressing mapping table when the distributed storage strategy is executed to store the polymorphic addressing data having the first scenario label in the virtual node includes: Build virtual nodes associated with each storage node; Reading scene tag information from the polymorphic addressing mapping table, and splitting the polymorphic addressing data having the first scene tag from the polymorphic addressing mapping table when executing a distributed storage strategy corresponding to the scene tag information; Addressing positions are allocated to the virtual nodes through a ring or tree topology structure, so as to store the polymorphic addressed data with the first scene tag in each virtual node.
6. A polymorphic addressing adaptive data management system, characterized in that: The system for implementing the method for adaptive data management of polymorphic addressing according to any one of claims 1 to 5 comprises: A data preprocessing module is used to obtain multi-scenario data from the Internet of Things and preprocess the multi-scenario data to construct a standardized data set; A data classification module, configured to perform multi-dimensional classification on the multi-scenario data based on the similarity between key indicators in the standardized data set to obtain a classification table; The mapping table generation module is used to obtain the polymorphic addressing scheme corresponding to each type of data in the classification table, and generate a polymorphic addressing mapping table based on the property parameters of each type of data, specifically including: Acquire multiple address segments defined based on each type of data in the classification table, and map each type of data in the classification table to a corresponding address segment to set an adjustable access priority for the address segment; Constructing an addressing strategy unit with a combination function based on each type of data in the classification table, the polymorphic addressing scheme corresponding to each type of data, and the property parameters; Calling the addressing strategy unit to parse the address segment according to the access priority to obtain a calling path and a redundancy strategy when reading and writing the address segment; Determine different addressing prefix allocation results according to the key indicators corresponding to each type of scene label in the classification table, and establish an initial mapping record including an address segment-scene label mapping table to determine a write strategy and prefix consistency format in the initial mapping record; Acquire a defined addressing attribute matrix according to the addressing prefix allocation result and the initial mapping record; Calculating an attribute priority score based on the addressing prefixes and corresponding attributes and attribute weights in the addressing attribute matrix; Integrating the attribute priority score, the call path and the redundancy strategy to generate a polymorphic addressing mapping table with an address segment addressing strategy; A distributed storage module, configured to store the polymorphic addressing mapping table in a distributed storage terminal, allocate addressing positions to storage nodes in the distributed storage terminal through a network topology, and output a distributed storage strategy; an addressing data transfer module, configured to construct a virtual node associated with each storage node and, when the distributed storage strategy is executed, split the polymorphic addressing mapping table to store the polymorphic addressing data having the first scenario label in the virtual node; Among them, the property parameters include the prefix, attributes and scene label of each type of data, the polymorphic addressing mapping table has an address segment addressing strategy for address segment retrieval of the polymorphic addressing mapping table, the virtual node is used for distributed storage of the polymorphic addressing mapping table that meets the preset conditions, and the first scene label is the scene label defined in the preset conditions.
7. A terminal comprising a processor and a storage medium; characterized in that: The storage medium is used to store instructions; The processor is configured to operate according to the instructions to execute the steps of the method according to any one of claims 1 to 5.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Data storage method and device based on distributed system and relational database
CN112559481A
Data management method for semiconductor storage
CN118760623A