Data storage system and method for multiple data nodes

By building a data storage queue structure in multiple data nodes and changing the key-value pair storage method, the problem of low data storage efficiency in traditional technology is solved, and efficient industrial timing data storage is achieved.

CN120540592AActive Publication Date: 2025-08-26CISDI INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510627746.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-08-26
Estimated Expiration
2045-05-15

AI Technical Summary

Technical Problem

In the high concurrency query scenario, the IO operation of relational databases leads to performance bottlenecks, and the key-value pair storage method is inefficient in industrial timing data storage, which cannot meet the high-throughput and time-efficient data storage requirements.

Method used

By constructing a data structure based on a data storage queue in multiple data nodes, allocating storage nodes according to the original tag, and storing data of the same target data with the same queue number, changing the key-value pair storage method to realize distributed storage and balanced storage computing power.

Benefits of technology

It improves data utilization efficiency, avoids low data reading efficiency caused by the expansion of the number of numerical keys, and meets the storage needs of industrial timing data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120540592A_ABST
    Figure CN120540592A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of distributed storage, and discloses a data storage system and method for multiple data nodes. According to the method, storage nodes are distributed to corresponding original data through a target node according to original labels, meanwhile, a data structure body formed based on data storage queues is constructed in a data node based on the original labels, and data of all dimensions corresponding to the same target data are stored in the corresponding data storage queues with the same queue sequence number. On one hand, distributed storage of data of different data categories is achieved, storage computing power is balanced, on the other hand, data of all dimensions corresponding to the same target data are stored in the corresponding data storage queues with the same queue sequence number, the data of all dimensions in the same target data are made to be in one-to-one correspondence, and the data storage efficiency is improved. A data caching mode based on a key-value pair is changed, and low data reading efficiency caused by number expansion of numerical keys is avoided, so that the data utilization efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of distributed storage technology, and in particular to a data storage system and method for multiple data nodes. Background Art

[0002] In process industry production line scenarios, the real-time data generated by production equipment status monitoring and process parameter collection is characterized by high throughput, high timeliness, and strong time series. Traditional technical solutions primarily rely on relational databases for time series data storage. However, in highly concurrent query scenarios, frequent system interactions and database I / O operations can easily lead to performance bottlenecks, resulting in insufficient real-time control models and a poor user experience. Existing industrial scenarios attempt to address these issues by improving data access performance through caching technology.

[0003] However, since caching technology usually uses key-value pair storage, if fine-grained key values ​​are built based on timestamps, the order of magnitude of numerical keys will expand exponentially, triggering performance degradation. If a block storage strategy with a fixed time period is adopted, although storage order can be maintained, the strong coupling of time intervals cannot flexibly support fast retrieval of data in a dynamic range. Therefore, the key-value pair data caching method results in low data reading efficiency and cannot meet the storage needs of industrial time series data. Summary of the Invention

[0004] In order to provide a basic understanding of some aspects of the disclosed embodiments, a brief summary is given below. The summary is not an extensive review, nor is it intended to identify key / critical elements or delineate the scope of protection of these embodiments, but rather serves as a prelude to the detailed description that follows.

[0005] In view of the above-mentioned shortcomings of the prior art, the present application provides a data storage system and method for multiple data nodes to improve data utilization efficiency by changing the data caching method.

[0006] The present application provides a data storage system for multiple data nodes, taking any data node as a target node, the target node including: an acquisition module, used to acquire original data, an original label corresponding to the original data, and matching from a preset node matching table according to the original label, so as to determine the storage node corresponding to the original data from each data node, and use the original data as the data to be stored corresponding to the storage node, wherein the node matching table stores the correspondence between data nodes and node labels; a configuration module, used to acquire a data structure corresponding to the target label, the data structure including data storage queues corresponding to multiple data dimensions, wherein the data to be stored corresponding to the target node is used as the target data, and the original label corresponding to the target data is used as the target label; a processing module, used to extract data from the target data according to each data dimension corresponding to the target label, obtain multiple dimensional data, and store each dimensional data corresponding to the same target data in the corresponding data storage queue with the same queue sequence number.

[0007] In one embodiment of the present application, the acquisition module is also used to: if the storage node includes a target node, use the data to be stored corresponding to the target node as the target data, and use the original label corresponding to the target data as the target label; if the storage node does not include a target node, send the original data and the original label to the storage node.

[0008] In one embodiment of the present application, the configuration module obtains the data structure corresponding to the target tag in the following manner: if the target node does not have a data structure corresponding to the storage tag, the target tag is used as a new storage tag, and the data structure corresponding to the target tag is generated; if the target node has a data structure corresponding to the storage tag, and the storage tag is the same as the target tag, the data structure corresponding to the storage tag is used as the data structure corresponding to the target tag; if the target node has a data structure corresponding to the storage tag, and the storage tag is different from the target tag, the target tag is used as a new storage tag, and the data structure corresponding to the target tag is generated.

[0009] In one embodiment of the present application, the data structure further includes at least one of the following: an acquisition time extreme value corresponding to an acquisition time queue, wherein the acquisition time extreme value includes the acquisition time maximum value and / or the acquisition time minimum value, the acquisition time queue belongs to a data storage queue, and the acquisition time queue is used to store data acquisition time; a data storage period, wherein the data storage period includes at least one of a period base duration, a period start time, and a period end time; a data valid interval corresponding to a data value queue, wherein the data value queue belongs to a data storage queue, and the data value queue is used to store data values; a data global variable, wherein the data global variable includes one or more of a queue length, a queue sequence number extreme value, a data quantity, and statistical data, wherein the queue sequence number extreme value includes the queue sequence number maximum value and / or the queue sequence number minimum value, and the statistical data is obtained by calculating the data values ​​in the data value queue.

[0010] In one embodiment of the present application, the configuration module is also used for at least one of the following: if the data acquisition time corresponding to the target data is greater than the cycle end time, the cycle end time is calculated according to the cycle benchmark duration to update the cycle end time; in response to the data structure storing new target data, the data global variable is updated.

[0011] In one embodiment of the present application, the processing module is also used for at least one of the following: if the data acquisition time corresponding to the target data is less than or equal to the maximum acquisition time, deleting the target data; if the data value corresponding to the target data is outside the data valid range, deleting the target data.

[0012] In one embodiment of the present application, the target node also includes at least one of the following: a cache module, used to cache the data structure corresponding to the target tag using a preset cache technology; a storage module, used to merge the data structures corresponding to the target tags if the data structures corresponding to the target tags meet the preset storage trigger conditions, obtain merged data, and persist the merged data, wherein the storage trigger conditions include the data volume corresponding to the data structure being greater than or equal to a preset data volume threshold, and / or the data collection time of the data structure being greater than or equal to a preset collection time threshold.

[0013] In one embodiment of the present application, the acquisition module acquires the node configuration table in the following manner: determining the storage efficiency index corresponding to each of the data nodes based on the node storage space and / or the node computing power, and determining the transmission efficiency index between each of the data nodes based on the data transmission rate, wherein the storage efficiency index is positively correlated with the node storage space and the node computing power, and the transmission efficiency index is positively correlated with the data transmission rate; acquiring a data label, and counting the data collection probability corresponding to the data label at each of the data nodes, wherein at least a portion of the data nodes obtain the label data corresponding to the data label through data collection, and the data collection probability is the probability that the data node collects the label data; taking any data node as the current node, calculating according to the storage efficiency index corresponding to the current node, the transmission efficiency index between each of the data nodes and the current node, and the data collection probability corresponding to each of the data nodes, to obtain a node score corresponding to the current node; determining an efficient node from each data node according to the node score corresponding to each of the data nodes, and using the data label as the node label corresponding to the efficient node; storing the correspondence between the efficient node and the node label in the node configuration table.

[0014] In one embodiment of the present application, the acquisition module calculates the node score corresponding to the current node in the following manner: Where, the current node is the mth data node, Score(m) is the node score corresponding to the current node, N is the total number of data nodes, P(i) is the data collection probability corresponding to the i-th data node, S(m) is the storage efficiency index corresponding to the current node, and T(i,m) is the transmission efficiency index from the i-th data node to the current node.

[0015] The present application provides a data storage method for multiple data nodes, taking any data node as a target node and applying it to the target node, the method comprising: obtaining original data and an original label corresponding to the original data, and matching from a preset node matching table according to the original label to determine the storage node corresponding to the original data from each of the data nodes, and using the original data as the data to be stored corresponding to the storage node, wherein the node matching table stores the correspondence between data nodes and node labels; obtaining a data structure corresponding to the target label, the data structure comprising data storage queues corresponding to multiple data dimensions, wherein the data to be stored corresponding to the target node is taken as target data, and the original label corresponding to the target data is taken as target label; extracting data from the target data according to each data dimension corresponding to the target label to obtain multiple dimensional data, and storing each dimensional data corresponding to the same target data in the corresponding data storage queue with the same queue sequence number.

[0016] The present application provides an electronic device, comprising: a processor and a memory; the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory, so that the electronic device performs the above method.

[0017] Beneficial effects of this application:

[0018] The target node allocates storage nodes to the corresponding original data according to the original label. At the same time, a data structure composed of data storage queues is constructed in the data node based on the original label, and each dimension data corresponding to the same target data is stored in the corresponding data storage queue with the same queue sequence number. In this way, on the one hand, storage nodes are allocated to the corresponding original data according to the original label, and a data structure composed of data storage queues is constructed in the data node based on the original label, so that data of different data categories can be distributed and storage computing power can be balanced. On the other hand, each dimension data corresponding to the same target data is stored in the corresponding data storage queue with the same queue sequence number, so that each dimension data in the same target data has a one-to-one correspondence, which changes the data caching method based on key-value pairs, avoids the low data reading efficiency caused by the expansion of the number of numerical keys, and thus improves data utilization efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 This is a schematic diagram of the structure of an application scenario in an embodiment of the present application;

[0020] Figure 2 This is a structural diagram of a data storage system for multiple data nodes in an embodiment of the present application;

[0021] Figure 3 This is a structural diagram of a data storage system for multiple data nodes in an embodiment of the present application;

[0022] Figure 4 This is a flow chart of an industrial data processing method in an embodiment of the present application;

[0023] Figure 5 This is a flow chart of a method for generating a data structure in an embodiment of the present application;

[0024] Figure 6 is a flowchart of a target data processing method in an embodiment of the present application;

[0025] Figure 7 This is a flow chart of a data storage method for multiple data nodes in an embodiment of the present application;

[0026] Figure 8 It is a schematic structural diagram of an electronic device in an embodiment of the present invention. DETAILED DESCRIPTION

[0027] The following describes the embodiments of the present invention through specific examples. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through different specific embodiments. The details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the following embodiments and sub-samples in the embodiments can be combined with each other unless there is a conflict.

[0028] It should be noted that the illustrations provided in the following embodiments are merely schematic illustrations of the basic concept of the present invention. Therefore, the illustrations only show components related to the present invention and are not drawn according to the number, shape, and size of components in actual implementation. In actual implementation, the type, quantity, and proportion of each component may be changed arbitrarily, and the component layout may also be more complex.

[0029] In the following description, numerous details are discussed to provide a more thorough explanation of the embodiments of the present invention. However, it will be apparent to those skilled in the art that the embodiments of the present invention may be practiced without these specific details. In other embodiments, well-known structures and devices are shown in block diagram form rather than in detail to avoid obscuring the embodiments of the present invention.

[0030] The terms "first," "second," and the like in the specification and claims of this application and the accompanying drawings are used to distinguish similar items and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate for the embodiments of the present application described herein. In addition, the terms "including," "having," and any variations thereof are intended to cover non-exclusive inclusions.

[0031] Unless otherwise stated, the term "plurality" means two or more.

[0032] In this application, the character " / " indicates that the preceding and following objects are in an "or" relationship. For example, A / B means: A or B.

[0033] The term "and / or" describes an association between objects, indicating that three relationships can exist. For example, A and / or B means: A or B, or A and B.

[0034] Combine Figure 1 As shown, the present application provides an application scenario for deploying a data storage system for multiple data nodes, wherein the application scenario includes multiple data nodes, and data is transmitted between the data nodes through a network.

[0035] In some embodiments, a data node consists of one or more data servers.

[0036] In some embodiments, the data storage system for multiple data nodes is suitable for high-throughput industrial scenarios, and can classify or retrieve industrial production data through real-time and efficient time-series data caching methods, thereby improving data node performance and reducing database load.

[0037] Combine Figure 2 As shown, the present application provides a data storage system for multiple data nodes, taking any data node as a target node, and the target node includes an acquisition module 201, a configuration module 202 and a processing module 203.

[0038] The acquisition module 201 is used to obtain the original data and the original label corresponding to the original data, and match it from a preset node matching table according to the original label to determine the storage node corresponding to the original data from each data node, and use the original data as the data to be stored corresponding to the storage node, wherein the node matching table stores the correspondence between data nodes and node labels.

[0039] The configuration module 202 is used to obtain a data structure corresponding to a target tag, where the data structure includes data storage queues corresponding to multiple data dimensions, wherein the data to be stored corresponding to the target node is used as the target data, and the original tag corresponding to the target data is used as the target tag.

[0040] The processing module 203 is used to extract data from the target data according to each data dimension corresponding to the target tag to obtain multiple dimension data, and store each dimension data corresponding to the same target data in a corresponding data storage queue with the same queue sequence number.

[0041] The data storage system for multiple data nodes provided by the present application is used to allocate storage nodes to the corresponding original data according to the original labels through the target nodes. At the same time, a data structure composed of data storage queues is constructed based on the original labels in the data nodes, and each dimension data corresponding to the same target data is stored in the corresponding data storage queue with the same queue sequence number. In this way, on the one hand, storage nodes are allocated to the corresponding original data according to the original labels, and a data structure composed of data storage queues is constructed based on the original labels in the data nodes, so that data of different data categories can be distributed and storage computing power can be balanced. On the other hand, each dimension data corresponding to the same target data is stored in the corresponding data storage queue with the same queue sequence number, so that each dimension data in the same target data corresponds to each other one by one, changing the data caching method based on key-value pairs, avoiding low data reading efficiency due to the expansion of the number of numerical keys, and thus improving data utilization efficiency.

[0042] Combine Figure 3As shown, the present application provides a data storage system for multiple data nodes, taking any data node as a target node, and the target node includes an acquisition module 201, a configuration module 202, a processing module 203, a cache module 204 and a storage module 205.

[0043] In some embodiments, the data storage system for multiple data nodes is implemented using one or more programming languages ​​such as Java, C#, C++, and SQL (Structured Query Language).

[0044] The acquisition module 201 is used to obtain the original data and the original label corresponding to the original data, and match it from a preset node matching table according to the original label to determine the storage node corresponding to the original data from each data node, and use the original data as the data to be stored corresponding to the storage node, wherein the node matching table stores the correspondence between data nodes and node labels.

[0045] Optionally, the acquisition module 201 is further used to: if the storage node includes the target node, use the data to be stored corresponding to the target node as the target data, and use the original label corresponding to the target data as the target label; if the storage node does not include the target node, send the original data and the original label to the storage node.

[0046] Combine Figure 4 As shown, the present application provides an industrial data processing method, comprising:

[0047] Step S401: The data node uses the received industrial real-time collected data as raw data;

[0048] Step S402, obtaining the original label corresponding to the original data;

[0049] Step S403, calculating the hash value corresponding to the original tag to obtain the tag hash value;

[0050] Step S404: Match the hash values ​​corresponding to the node labels according to the label hash values, so as to determine the storage node corresponding to the original data from each data node according to the matching result;

[0051] Step S405: determine whether the storage node includes the data node. If so, jump to step S406; if not, jump to step S407;

[0052] Step S406: Use the original data as the target data corresponding to the data node.

[0053] Step S407: forward the original data to the storage node.

[0054] Optionally, the acquisition module 201 acquires the node configuration table in the following manner: determining the storage efficiency index corresponding to each data node according to the node storage space and / or node computing power, and determining the transmission efficiency index between each data node according to the data transmission rate, wherein the storage efficiency index is positively correlated with the node storage space and the node computing power, and the transmission efficiency index is positively correlated with the data transmission rate; acquiring a data label, and calculating the data collection probability corresponding to the data label at each data node, wherein at least a part of the data nodes obtain the label data corresponding to the data label through data collection, and the data collection probability is the probability that the data node collects the label data; taking any data node as the current node, calculating according to the storage efficiency index corresponding to the current node, the transmission efficiency index between each data node and the current node, and the data collection probability corresponding to each data node, to obtain the node score corresponding to the current node; determining the efficient node from each data node according to the node score corresponding to each data node, and using the data label as the node label corresponding to the efficient node; storing the correspondence between the efficient node and the node label in the node configuration table.

[0055] In some embodiments, the data nodes include Node-1, Node-2, and Node-3. By performing weighted calculation on the node storage space and the node computing power, the storage efficiency index of the data node is obtained, and more node scores are allocated to the nodes with higher storage efficiency indexes, thereby maximizing the storage efficiency. The calculation formula corresponding to the storage efficiency index is: The unit of measurement for node storage space is GB (Gigabyte), the weight of node storage space is β = 0.7, the unit of measurement for node computing power is TFLOPS (TeraFloating-point Operations Per Second), the weight of node computing power is γ = 0.3; the data transmission rate is calculated based on the benchmark transmission rate to obtain the transmission efficiency index of the data node, and more node scores are allocated to data nodes with higher transmission efficiency, thereby maximizing the transmission efficiency. The calculation formula for the transmission efficiency index is: The unit of measurement for the benchmark transmission rate is Mbps (Megabits per second), the benchmark transmission rate is 120 Mbps, and the data transmission rate adjustment coefficient α = 0.5. Through calculation, the storage efficiency index and transmission efficiency index of Node-1, Node-2, and Node-3 are shown in Table 1.

[0056] Table 1

[0057] Data Node Node-1 Node-2 Node-3 Node user space (GB) 500 800 300 Node computing power (TFLOPS) 8 6 10 Storage efficiency indicators 0.35 0.56 0.21 Transmission efficiency index to Node-1 1 0.34 0.88 Transmission efficiency index to Node-2 0.42 1 0.23 Transmission efficiency index to Node-3 0.33 0.52 1

[0058] In some embodiments, different data nodes collect different data categories. The data collection probability of the data node for different data tags is determined based on the historical statistical results of the data node, reducing the impact of data nodes with lower data collection probability on node scores.

[0059] Optionally, the acquisition module 201 calculates the node score corresponding to the current node using formula (1):

[0060]

[0061] In formula (1), the current node is the mth data node, Score(m) is the node score corresponding to the current node, N is the total number of data nodes, P(i) is the data collection probability corresponding to the i-th data node, S(m) is the storage efficiency index corresponding to the current node, and T(i,m) is the transmission efficiency index from the i-th data node to the current node.

[0062] In some embodiments, through calculation, the node score of Node-1 is 0.31, the node score of Node-2 is 0.58, and the node score of Node-3 is 0.39. Therefore, the data label is used as the node label of Node-2. When other data nodes extract the original data corresponding to the node label, the original data is sent to Node-2 for storage.

[0063] The configuration module 202 is used to obtain a data structure corresponding to a target tag, where the data structure includes data storage queues corresponding to multiple data dimensions, wherein the data to be stored corresponding to the target node is used as the target data, and the original tag corresponding to the target data is used as the target tag.

[0064] Optionally, the configuration module 202 obtains the data structure corresponding to the target tag in the following manner: if the target node does not have a data structure corresponding to the storage tag, the target tag is used as a new storage tag, and a data structure corresponding to the target tag is generated; if the target node has a data structure corresponding to the storage tag, and the storage tag and the target tag are the same, the data structure corresponding to the storage tag is used as the data structure corresponding to the target tag; if the target node has a data structure corresponding to the storage tag, and the storage tag and the target tag are different, the target tag is used as a new storage tag, and a data structure corresponding to the target tag is generated.

[0065] Combine Figure 5 As shown, the present application provides a data structure generation method, including:

[0066] Step S501, obtaining the storage tag and the data structure corresponding to the storage tag in the target node;

[0067] Step S502, obtaining a target tag;

[0068] Step S503, determine whether the stored tags include the target tag, if so, jump to step S504, if not, jump to step S506;

[0069] Step S504: Load the data configuration information corresponding to the target tag;

[0070] The data configuration information includes at least one of a data storage period, a data valid interval, and a data global variable;

[0071] Step S505 : Use the target tag as a new storage tag, initialize the data structure corresponding to the target tag according to the data configuration information, and allocate cache space to the data structure.

[0072] Step S506: Use the data structure corresponding to the storage tag as the data structure corresponding to the target tag.

[0073] Optionally, the data structure also includes at least one of the following: an acquisition time extreme value corresponding to an acquisition time queue, wherein the acquisition time extreme value includes the acquisition time maximum value and / or the acquisition time minimum value, the acquisition time queue belongs to a data storage queue, and the acquisition time queue is used to store the data acquisition time; a data storage period, wherein the data storage period includes at least one of a period base duration, a period start time, and a period end time; a data valid interval corresponding to a data value queue, wherein the data value queue belongs to a data storage queue, and the data value queue is used to store data values; a data global variable, wherein the data global variable includes one or more of a queue length, a queue sequence number extreme value, a data quantity, and statistical data, the queue sequence number extreme value includes the queue sequence number maximum value and / or the queue sequence number minimum value, and the statistical data is obtained by calculating the data values ​​in the data value queue.

[0074] In some embodiments, at least a portion of the data structure is shown in Table 2.

[0075] Table 2

[0076]

[0077] Optionally, the configuration module 202 is further configured to: if the data collection time corresponding to the target data is greater than the cycle end time, calculate the cycle end time according to the cycle reference duration to update the cycle end time.

[0078] In some embodiments, if the data collection time corresponding to the target data is greater than the cycle end time, it means that the data collection time corresponding to the target data is later than the cycle end time, and the cycle end time automatically increases the cycle base time length.

[0079] Optionally, the configuration module 202 is further configured to update the data global variable in response to the data structure storing new target data.

[0080] In some embodiments, if the data collection time corresponding to the target data is greater than the cycle end time, the data value extreme value, statistical data and data quantity are updated.

[0081] The processing module 203 is used to extract data from the target data according to each data dimension corresponding to the target tag to obtain multiple dimension data, and store each dimension data corresponding to the same target data in a corresponding data storage queue with the same queue sequence number.

[0082] In some embodiments, the data storage queue is a data ring queue.

[0083] Optionally, the processing module 203 is further configured to: delete the target data if the data collection time corresponding to the target data is less than or equal to the maximum collection time.

[0084] In some embodiments, if the data collection time corresponding to the target data is less than or equal to the maximum collection time, it means that the data collection time corresponding to the target data is earlier than the latest data collection time in the collection time queue, the target data is not the latest real-time data, and the target data is deleted.

[0085] Optionally, the processing module 203 is further configured to: delete the target data if the data value corresponding to the target data is outside the data valid range.

[0086] In some embodiments, if the data value corresponding to the target data is outside the data valid range, it means that the target data is abnormal data, and the abnormal data is deleted.

[0087] Combine Figure 6 As shown, the present application provides a target data processing method, comprising:

[0088] Step S601, extracting data collection time and data value from target data;

[0089] Step S602, determine whether the data collection time is greater than the maximum collection time, if so, jump to step S603, if not, jump to step S610;

[0090] Step S603, determine whether the data collection time is greater than the cycle end time, if so, jump to step S604, if not, jump to step S605;

[0091] Step S604: The cycle start time is automatically incremented by the cycle reference time length, and the cycle end time is automatically incremented by the cycle reference time length;

[0092] Step S605, resetting data extreme values, statistical data, and data quantity;

[0093] Step S606, determine whether the data value is within the data valid range, if so, jump to step S607, if not, jump to step S610;

[0094] Step S607, updating the collection time extreme value, data value extreme value, statistical data and data quantity;

[0095] Step S608, calculating the maximum value of the queue sequence number to obtain the storage sequence number corresponding to the target data;

[0096] Step S609: Store the data collection time and data value in the corresponding data storage queue according to the storage sequence number.

[0097] Step S610: Delete target data.

[0098] The cache module 204 is configured to cache the data structure corresponding to the target tag using a preset cache technology.

[0099] In some embodiments, the preset cache technology includes one or more of Caffeine, Guava Cache, Ehcache, Memcache, Redis, etc.

[0100] In this way, multiple data are integrated into the same data structure for storage, and the correspondence of the same data is represented by the queue sequence number, and then the data structure is saved in the form of key-value pairs, thereby reducing the number of key-value pairs and ensuring data utilization efficiency.

[0101] The storage module 205 is used to merge the data structures corresponding to the target tags to obtain merged data and persist the merged data if the data structures corresponding to the target tags meet the preset storage trigger conditions, wherein the storage trigger conditions include that the amount of data corresponding to the data structures is greater than or equal to the preset data amount threshold, and / or that the data collection time of the data structure is greater than or equal to the preset collection time threshold.

[0102] In some embodiments, if the queue length of the data storage queue reaches the queue length requirement, the entire data structure is stored in the hard disk.

[0103] In some embodiments, if the data collection duration between the period start time and the period end time meets the collection duration requirement, the entire data structure is stored in the hard disk to make the data structure persistent.

[0104] Combine Figure 7 As shown, the present application provides a data storage method for multiple data nodes, taking any data node as a target node and applying it to the target node, the method includes:

[0105] Step S701, obtaining original data and original labels corresponding to the original data;

[0106] Step S702: Matching is performed from a preset node matching table according to the original label to determine the storage node corresponding to the original data from each data node, and the original data is used as the data to be stored corresponding to the storage node;

[0107] Among them, the node matching table stores the correspondence between data nodes and node labels;

[0108] Step S703: Obtain a data structure corresponding to the target tag, where the data structure includes data storage queues corresponding to multiple data dimensions.

[0109] The data to be stored corresponding to the target node is used as the target data, and the original label corresponding to the target data is used as the target label;

[0110] Step S704 , extracting data from the target data according to the data dimensions corresponding to the target tag to obtain multiple dimensional data, and storing the dimensional data corresponding to the same target data in the corresponding data storage queues with the same queue sequence number.

[0111] The data storage method for multiple data nodes provided by the present application is adopted, and the target node is used to allocate storage nodes to the corresponding original data according to the original label. At the same time, a data structure composed of data storage queues is constructed in the data node based on the original label, and each dimension data corresponding to the same target data is stored in the corresponding data storage queue with the same queue sequence number. In this way, on the one hand, storage nodes are allocated to the corresponding original data according to the original label, and a data structure composed of data storage queues is constructed in the data node based on the original label, so that data of different data categories can be distributed and storage computing power can be balanced. On the other hand, each dimension data corresponding to the same target data is stored in the corresponding data storage queue with the same queue sequence number, so that each dimension data in the same target data corresponds to each other one by one, which changes the data caching method based on key-value pairs, avoids low data reading efficiency due to the expansion of the number of numerical keys, and thus improves data utilization efficiency.

[0112] The present application also provides an electronic device, comprising: a processor and a memory; the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory, so that the electronic device performs the above method.

[0113] Figure 8 The following is a schematic diagram showing the structure of a computer system suitable for implementing an electronic device according to an embodiment of the present application. Figure 8 The computer system 800 of the electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0114] like Figure 8 As shown, the computer system 800 includes a central processing unit (CPU) 801, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 802 or the program loaded from the storage part 808 into the random access memory (RAM) 803, such as executing the method in the above embodiment. Various programs and data required for system operation are also stored in the RAM 803. The CPU 801, ROM 802 and RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0115] The following components are connected to the I / O interface 805: an input section 806 including a keyboard, a mouse, and the like; an output section 807 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), and a speaker; a storage section 808 including a hard disk and the like; and a communication section 809 including a network interface card such as a LAN (Local Area Network) card or a modem. The communication section 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to the I / O interface 805 as needed. Removable media 811, such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory, is installed in the drive 810 as needed, so that computer programs read therefrom can be installed into the storage section 808 as needed.

[0116] The electronic device disclosed in this embodiment includes a processor, a memory, a transceiver, and a communication interface. The memory and the communication interface are connected to the processor and the transceiver and complete communication with each other. The memory is used to store computer programs, the communication interface is used to communicate, and the processor and the transceiver are used to run the computer program, so that the electronic device executes each step of the above method.

[0117] The above description and the accompanying drawings fully illustrate the embodiments of the present disclosure to enable those skilled in the art to practice them. Other embodiments may include structural, logical, electrical, process and other changes. The embodiments represent only possible variations. Unless expressly required, individual components and functions are optional, and the order of operations may vary. Parts and subsamples of some embodiments may be included in or replace parts and subsamples of other embodiments. Moreover, the terms used in this application are only used to describe the embodiments and are not used to limit the claims. As used in the description of the embodiments and claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to also include plural forms. Similarly, the term "and / or" as used in this application refers to any and all possible combinations of one or more associated listings. In addition, when used in this application, the term "comprise" and its variations "comprises" and / or comprising refer to the presence of a stated subsample, whole, step, operation, element, and / or component, but do not exclude the presence or addition of one or more other subsamples, wholes, steps, operations, elements, components and / or groups of these. In the absence of further restrictions, an element defined by the statement "comprises a..." does not exclude the presence of other identical elements in the process, method or device that includes the element. In this article, each embodiment may focus on the differences from other embodiments, and the same and similar parts between the various embodiments can be referenced to each other. For the methods, products, etc. disclosed in the embodiments, if they correspond to the method part disclosed in the embodiments, then the relevant parts can be found in the description of the method part.

[0118] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software may depend on the specific application and design constraints of the technical solution. Technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application. Technicians can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0119] In the embodiments disclosed herein, the disclosed methods and products (including but not limited to devices, equipment, etc.) can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units can be merely a logical functional division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some sub-samples can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between each other shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, and can be electrical, mechanical or other forms. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the units may be selected according to actual needs to implement this embodiment. In addition, the functional units in this application may be integrated into a processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0120] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of the systems, methods, and computer program products of the present application. In this regard, each block in a flowchart or block diagram may represent a module, program segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in an order different from that marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, or they may sometimes be executed in the opposite order, depending on the functions involved. In the descriptions corresponding to the flowcharts and block diagrams in the accompanying drawings, the operations or steps corresponding to different blocks may also occur in an order different from that disclosed in the description, and sometimes there is no specific order between different operations or steps. For example, two consecutive operations or steps may actually be executed substantially in parallel, or they may sometimes be executed in the opposite order, depending on the functions involved. Each block in the block diagram and / or flowchart, as well as combinations of blocks in the block diagram and / or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or action, or may be implemented using a combination of dedicated hardware and computer instructions.

Claims

1. A data storage system for multiple data nodes, characterized in that: Any data node is used as a target node, and the target node includes: an acquisition module, configured to acquire original data and original labels corresponding to the original data, and perform matching from a preset node matching table based on the original labels to determine the storage node corresponding to the original data from each of the data nodes, and use the original data as the data to be stored corresponding to the storage node, wherein the node matching table stores the correspondence between data nodes and node labels; A configuration module is configured to obtain a data structure corresponding to a target tag, wherein the data structure includes data storage queues corresponding to a plurality of data dimensions, wherein the to-be-stored data corresponding to the target node is used as the target data, and the original tag corresponding to the target data is used as the target tag; The processing module is used to extract data from the target data according to the data dimensions corresponding to the target tag to obtain multiple dimensional data, and store the dimensional data corresponding to the same target data in the corresponding data storage queue with the same queue sequence number.

2. The system according to claim 1, wherein: The acquisition module is further used for: If the storage node includes a target node, the data to be stored corresponding to the target node is used as the target data, and the original label corresponding to the target data is used as the target label; If the storage node does not include the target node, the original data and the original label are sent to the storage node.

3. The system according to claim 1, wherein: The configuration module obtains the data structure corresponding to the target tag in the following way: If the target node does not have a data structure corresponding to the storage tag, use the target tag as a new storage tag and generate a data structure corresponding to the target tag; If the target node has a data structure corresponding to the storage tag, and the storage tag is the same as the target tag, the data structure corresponding to the storage tag is used as the data structure corresponding to the target tag; If the target node has a data structure corresponding to the storage tag, and the storage tag is different from the target tag, the target tag is used as a new storage tag, and a data structure corresponding to the target tag is generated.

4. The system according to claim 1, wherein: The data structure also includes at least one of the following: An acquisition time extreme value corresponding to an acquisition time queue, wherein the acquisition time extreme value includes an acquisition time maximum value and / or an acquisition time minimum value. The acquisition time queue belongs to a data storage queue, and the acquisition time queue is used to store data acquisition time; A data storage period, wherein the data storage period includes at least one of a period base duration, a period start time, and a period end time; a data valid interval corresponding to a data value queue, wherein the data value queue belongs to a data storage queue and is used to store data values; Data global variables, wherein the data global variables include one or more of queue length, queue sequence number extreme value, data quantity, and statistical data, wherein the queue sequence number extreme value includes the queue sequence number maximum value and / or the queue sequence number minimum value, and the statistical data is obtained by calculating the data values ​​in the data value queue.

5. The system according to claim 4, characterized in that The configuration module is further configured to: If the data collection time corresponding to the target data is greater than the cycle end time, the cycle end time is calculated according to the cycle reference duration to update the cycle end time; In response to the data structure storing new target data, the data global variable is updated.

6. The system according to claim 4, characterized in that The processing module is further configured to: If the data collection time corresponding to the target data is less than or equal to the maximum collection time, deleting the target data; If the data value corresponding to the target data is outside the data valid range, the target data is deleted.

7. The system according to any one of claims 1 to 6, characterized in that The target node also includes at least one of the following: A cache module, configured to cache the data structure corresponding to the target tag using a preset cache technology; A storage module is used to merge the data structures corresponding to the target tags to obtain merged data and persist the merged data if the data structures corresponding to the target tags meet the preset storage trigger conditions, wherein the storage trigger conditions include that the amount of data corresponding to the data structures is greater than or equal to a preset data amount threshold, and / or that the data collection time of the data structure is greater than or equal to a preset collection time threshold.

8. The system according to any one of claims 1 to 6, characterized in that The acquisition module obtains the node configuration table in the following manner: Determine a storage efficiency index corresponding to each of the data nodes according to the node storage space and / or the node computing power, and determine a transmission efficiency index between the data nodes according to the data transmission rate, wherein the storage efficiency index is positively correlated with the node storage space and the node computing power, and the transmission efficiency index is positively correlated with the data transmission rate; Obtaining a data tag and calculating a data collection probability corresponding to the data tag at each of the data nodes, wherein at least some of the data nodes obtain label data corresponding to the data tag through data collection, and the data collection probability is a probability that the data node collects the label data; Taking any data node as the current node, calculating based on the storage efficiency index corresponding to the current node, the transmission efficiency index between each data node and the current node, and the data collection probability corresponding to each data node, to obtain the node score corresponding to the current node; Determine an efficient node from each data node according to the node score corresponding to each data node, and use the data label as the node label corresponding to the efficient node; The corresponding relationship between the efficient nodes and the node labels is stored in a node configuration table.

9. The system according to claim 8, characterized in that The acquisition module calculates the node score corresponding to the current node in the following way: Where, the current node is the mth data node, Score(m) is the node score corresponding to the current node, N is the total number of data nodes, P(i) is the data collection probability corresponding to the i-th data node, S(m) is the storage efficiency index corresponding to the current node, and T(i,m) is the transmission efficiency index from the i-th data node to the current node.

10. A data storage method for multiple data nodes, characterized in that: Taking any data node as a target node and applying to the target node, the method includes: Obtaining original data and original labels corresponding to the original data, and matching them from a preset node matching table according to the original labels to determine the storage node corresponding to the original data from each of the data nodes, and using the original data as the data to be stored corresponding to the storage node, wherein the node matching table stores the correspondence between data nodes and node labels; Obtain a data structure corresponding to a target tag, the data structure including data storage queues corresponding to a plurality of data dimensions, wherein the to-be-stored data corresponding to the target node is used as the target data, and the original tag corresponding to the target data is used as the target tag; Data is extracted from the target data according to each data dimension corresponding to the target tag to obtain multiple dimension data, and each dimension data corresponding to the same target data is stored in a corresponding data storage queue with the same queue sequence number.

Citation Information

Patent Citations

  • Data distributing and caching method and data distributing and caching system

    CN102638584A

  • Method for storing multiple cache queues in parallel

    CN107633034A