A data storage node analysis method, device, equipment and storage medium
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA MOBILE (XIONGAN) ICT CO LTD
- Filing Date
- 2026-06-08
- Publication Date
- 2026-08-07
AI Technical Summary
目前常用动态选择策略主要通过对节点的负载与连接数等实时状态来确定节点,这种策略的局限性在于:一是即时计算,无法反映存储节点长期的负载状态,并且实时监控使得管理开销大;二是脱离业务,未考虑业务数据的分布特点
[0007] Fourthly, embodiments of the present invention provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the data storage node analysis method described in any embodiment.
Smart Images

Figure CN122533985A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data storage management technology, and in particular to a data storage node analysis method, apparatus, device, and storage medium. Background Technology
[0002] Data storage node selection: When storing or accessing data in a distributed storage cluster, the system determines which nodes to store data blocks on, or from which nodes to read data, based on specific algorithms and strategies. A good selection strategy can reduce data access latency, increase throughput, prevent single points of failure, ensure continuous service availability, and achieve load balancing, ensuring balanced utilization of storage space, computing resources, and network bandwidth across all nodes in the cluster. Currently, commonly used dynamic selection strategies primarily determine nodes based on real-time status such as node load and connection count. The limitations of this strategy are: firstly, it involves immediate calculation, failing to reflect the long-term load status of storage nodes, and real-time monitoring incurs significant management overhead; secondly, it is detached from business needs and does not consider the distribution characteristics of business data. Summary of the Invention
[0003] This invention provides a data storage node analysis method, apparatus, device, and storage medium. The technical solution of this invention can further improve the rationality of data storage node allocation, thereby improving data query and retrieval efficiency.
[0004] In a first aspect, embodiments of the present invention provide a data storage node analysis method, the method comprising: A network behavior data set is obtained, and data value parameters are determined based on the data density and data information volume characteristics of the network behavior data set; historical load data corresponding to multiple nodes to be analyzed is obtained, and node health is determined for each node to be analyzed based on the historical load data; a target node is determined from the nodes to be analyzed based on the data value parameters and node health, and the network behavior data set is stored in the target node.
[0005] In a second aspect, embodiments of the present invention provide a data storage node analysis device, the device comprising: The data value determination module is used to acquire a set of network behavior data and determine data value parameters based on the data density characteristics and data information content characteristics of the set of network behavior data; the node health determination module is used to acquire historical load data corresponding to multiple nodes to be analyzed and determine the node health of each node to be analyzed based on the historical load data; the storage node selection module is used to determine a target node from the nodes to be analyzed based on the data value parameters and node health, and store the set of network behavior data in the target node.
[0006] Thirdly, embodiments of the present invention provide a computer device, the computer device comprising: One or more processors; Memory, used to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the data storage node analysis method described in any embodiment.
[0007] Fourthly, embodiments of the present invention provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the data storage node analysis method described in any embodiment.
[0008] The technical solution provided by this invention involves acquiring a network behavior data set, determining data value parameters based on the data density and information content characteristics of the network behavior data set, acquiring historical load data corresponding to multiple nodes to be analyzed, determining the node health of each node to be analyzed based on the historical load data, determining a target node from the nodes to be analyzed based on the data value parameters and node health, and storing the network behavior data set in the target node. This technical solution can comprehensively analyze the data value of network behavior data and the historical load characteristics of nodes to determine the storage node for the network behavior data, further improving the rationality of data storage node allocation and thus improving data query and retrieval efficiency. Attached Figure Description
[0009] Figure 1 This is a flowchart of a data storage node analysis method provided in an embodiment of the present invention; Figure 2 This is a flowchart of another data storage node analysis method provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of a data storage node analysis device provided in an embodiment of the present invention; Figure 4 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present invention. Detailed Implementation
[0010] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. The acquisition, storage, use, and processing of data in the technical solutions of the embodiments of the present invention all comply with the relevant provisions of national laws and regulations.
[0011] Figure 1 This is a flowchart of a data storage node analysis method provided by an embodiment of the present invention. The embodiment of the present invention can be applied to scenarios where storage nodes for network behavior data are selected. The method can be executed by a data storage node analysis device, which can be implemented by software and / or hardware.
[0012] like Figure 1 As shown, the data storage node analysis method includes the following steps: S110. Obtain a set of network behavior data, and determine data value parameters based on the data density characteristics and data information content characteristics of the set of network behavior data.
[0013] The network behavior dataset can be a collection of network behavior data that needs to be stored. A network behavior dataset can contain multiple data points. By analyzing the value parameters of these data points, the nodes to be stored can be determined. Data value parameters represent the intrinsic value of the network behavior data. Specifically, for each data point in the dataset, data density and information content characteristics can be extracted, and then the data value parameters are determined based on these characteristics.
[0014] S120. Obtain historical load data for multiple nodes to be analyzed, and determine the node health of each node to be analyzed based on the historical load data.
[0015] The nodes to be analyzed can be nodes that require data storage and analysis. The technical solution of this invention can analyze the historical load of multiple nodes to be analyzed, thereby determining the storage node corresponding to the network behavior data set. Historical load data can be data on the load of the nodes to be analyzed during a historical period. For example, the load data of the nodes to be analyzed within the last 15 days can be used as historical load data. Furthermore, node health can be health parameters of the nodes to be analyzed determined from multiple perspectives. Specifically, load characteristics of various load parameters of the nodes to be analyzed can be extracted from historical load data, and then the node health corresponding to the nodes to be analyzed can be determined by combining the load characteristics of various load parameters.
[0016] S130. Based on data value parameters and node health, determine the target node from the node to be analyzed, and store the network behavior data set in the target node.
[0017] The target node can be a node used to store the network behavior data set. Specifically, for each piece of current behavior data in the network behavior data set, the target node corresponding to each piece of current behavior data can be determined based on the data value parameter of the current behavior data and the node health of the node to be analyzed. Finally, each piece of current behavior data is stored in the corresponding target node.
[0018] The technical solution provided by this invention involves acquiring a network behavior data set, determining data value parameters based on the data density and information content characteristics of the network behavior data set, acquiring historical load data corresponding to multiple nodes to be analyzed, determining the node health of each node to be analyzed based on the historical load data, determining a target node from the nodes to be analyzed based on the data value parameters and node health, and storing the network behavior data set in the target node. This technical solution can comprehensively analyze the data value of network behavior data and the historical load characteristics of nodes to determine the storage node for the network behavior data, further improving the rationality of data storage node allocation and thus improving data query and retrieval efficiency.
[0019] Figure 2 This is a flowchart of another data storage node analysis method provided by the embodiments of the present invention. The embodiments of the present invention can be applied to scenarios in which storage nodes for network behavior data are selected. Based on the above embodiments, this embodiment further explains how to determine data value parameters based on the data density characteristics and data information volume characteristics of the network behavior data set; and how to determine the node health of each node to be analyzed based on historical load data. The device can be implemented by software and / or hardware and integrated into a computer device with application development capabilities.
[0020] like Figure 2 As shown, the data storage node analysis method includes the following steps: S210. Obtain a network behavior data set. For each current behavior data in the network behavior data set, determine the relative data density parameter of the current behavior data based on the current behavior data and a preset number of other behavior data in the network behavior data set that precede the current behavior data.
[0021] The network behavior dataset can be a collection of network behavior data that needs to be stored. This dataset can contain multiple network behavior data points. By analyzing the value parameters of these data points, the nodes to be stored subsequently can be determined. The current behavior data can be the network behavior data currently being analyzed. Specifically, each network behavior data point in the network association dataset can be analyzed sequentially as the current behavior data. Other behavior data can be other data in the network behavior dataset used as reference samples for the value parameters of the current behavior data. Specifically, for each current behavior data point, a predetermined number of network behavior data points preceding it can be considered as other behavior data. Furthermore, the relative data density parameter can be the data density parameter of the current behavior data compared to other behavior data. Specifically, the density parameter of the current behavior data can be compared to the density parameters of other behavior data, and the resulting ratio can be used as the relative data density parameter.
[0022] Optionally, when determining the relative data density parameter of the current behavior data, the ratio of the number of attribute sub-data with values in a preset number of other behavior data to the total number of attribute sub-data in all other behavior data can be used as the average data density corresponding to the other behavior data; the ratio of the number of attribute sub-data with values in the current behavior data to the total number of attribute sub-data in the current behavior data can be used as the current data density corresponding to the current behavior data; and the ratio of the current data density to the average data density can be used as the relative data density parameter of the current behavior data.
[0023] Among these, attribute sub-data can be attribute data from network behavior data. Specifically, network behavior data can contain multiple attribute sub-data. By analyzing the attribute sub-data, the data characteristics of the corresponding network behavior data can be determined. Furthermore, the average data density can be the average density of valuable data in other behavior data. Specifically, the ratio of the number of attribute sub-data with values in a preset number of other behavior data to the total number of attribute sub-data in all other behavior data can be used as the average data density for the other behavior data. Furthermore, the current data density can be the density of valuable data in the current behavior data. Specifically, the ratio of the number of attribute sub-data with values in the current behavior data to the total number of attribute sub-data in the current behavior data can be used. Furthermore, to reflect the relative value of the current behavior data in terms of data density compared to other behavior data, the ratio of the current data density to the average density can be used as the relative data density parameter for the current behavior data.
[0024] For example, a sliding window can be used to sample network behavior data sets, with a window size of n+1 data points and a window step size of 1.
[0025] Let the current behavior data be the first... +1 data, then record the previous Other behavioral data (1, 2, 3, ..., The existence of attribute values for ( ). Before calculation Average data density of other behavioral data: in, For average data density, For the front The total number of values for attribute sub-data in other behavioral data. The number of attribute sub-data. Calculate the number of sub-data. Data density parameter for +1 data point (i.e., current action data): in The data density parameter for the current behavioral data. This represents the number of values in the attribute sub-data within the current behavior data. This represents the number of attribute sub-data points in the current behavioral data. Calculate the relative data density parameter: S220. Based on the current behavior data and a preset number of other behavior data preceding the current behavior data in the network behavior data set, determine the data information quantity parameter of the current behavior data.
[0026] The data information content parameter can be a parameter used to represent the information content of the current behavioral data. Specifically, it can be determined how frequently the data values in the current behavioral data appear in other behavioral data, and then the data information content parameter corresponding to the current behavioral data can be determined based on the frequency of occurrence.
[0027] Optionally, the data information content parameter of the current behavior data is determined, including: for each attribute sub-data in the current behavior data, the probability of the attribute value of each attribute sub-data appearing in other behavior data is determined, and the attribute information entropy corresponding to the attribute sub-data is determined; the attribute information entropy of all attribute sub-data is weighted and summed to obtain the data information content parameter of the current behavior data.
[0028] Attribute information entropy can be the information entropy value corresponding to an attribute sub-data. Specifically, for each attribute sub-data in the current behavioral data, the probability of each attribute value appearing in other behavioral data is calculated. Then, the probability of all attribute sub-data appearing is substituted into the information entropy calculation formula to obtain the attribute information entropy corresponding to the attribute sub-data. Finally, the attribute information entropies of all attribute sub-data can be weighted and summed to obtain the data information content parameter of the current behavioral data.
[0029] For example, network behavior data attributes include both numerical attribute sub-data and categorical attribute sub-data.
[0030] For numerical attributes (such as latitude and longitude), calculations can be performed directly; For categorical attributes, attributes with fewer categories use One-Hot encoding, while attributes with more categories use target encoding. Tag attributes can contain multiple tag values, so a scoring method is used for special processing. Based on the importance of the tag values, they are graded as low, medium, and high, corresponding to 1 point, 2 points, and 3 points respectively. The processed value of the tag attribute is the sum of the scores corresponding to each tag.
[0031] Step 3: Information Entropy Calculation The formula for calculating the attribute information entropy of attribute sub-data is as follows: in, Represents attribute sub-data Value The probability of occurrence in other behavioral data. Represents attribute sub-data The attribute information entropy. Therefore, the data information content parameter for the current behavior data is: in, This parameter represents the amount of data information in the current behavioral data. To better reflect the characteristics of business data, each attribute is weighted. Users can customize this weighting; otherwise, all attributes default to 1. S230. Substitute the relative data density parameter and the data information quantity parameter into the preset Sigmoid function to obtain the data value parameter of the current behavior data.
[0032] The preset function can be a pre-defined function used to determine the value of data. The data value parameter can be a parameter representing the practical value of the current behavioral data. Specifically, the relative data density parameter and the data information content parameter can be substituted into the preset sigmoid function to obtain the data value parameter of the current behavioral data.
[0033] The default Sigmoid function is: in, Indicates the data value parameter, Indicates the amount of data information. This refers to the relative data density parameter; This is the threshold parameter.
[0034] Optional, It can change dynamically and can be adjusted based on the amount of data information in the current behavioral data and other behavioral data. For example, its initial value can be set to... The dynamic adjustment formula is given below: in, This is the smoothing factor, typically set to 0.7; This represents the median amount of information contained in the current behavioral data and other behavioral data.
[0035] S240. For each node to be analyzed, determine the historical peak parameters based on the historical load data of the node to be analyzed.
[0036] The node to be analyzed can be a node that requires data storage analysis. The technical solution of this invention can analyze the historical load of multiple nodes to be analyzed, thereby determining the storage node corresponding to the network behavior data set. Historical load data can be data on the load of the node to be analyzed during a historical period. For example, the load data of the node to be analyzed within the last 15 days can be used as historical load data. Furthermore, historical peak parameters can be peak data of load parameters of the node to be analyzed during a historical period. For example, historical peak parameters include at least two of the following: CPU load peak, memory load peak, disk I / O load peak, and network bandwidth load peak. Specifically, for each load parameter, the maximum value of that load parameter can be determined from the historical load data, and the determined maximum value can be used as the historical peak parameter corresponding to that load parameter. Optionally, a statistical measure of the historical load data can be selected first. This proposal selects P95, which reflects the peak load borne by the server for the vast majority of the time (95%), avoiding interference from occasional spikes and truly reflecting the steady-state pressure level of the server.
[0037] S250. For each historical peak parameter, determine the parameter health level corresponding to the historical peak parameter based on the peak parameter range in which the historical peak parameter is located.
[0038] Parameter health is used to represent the health level of the node to be analyzed, determined from the perspective of a single load parameter. The peak parameter range can be the range where historical peak parameters are located. Specifically, multiple peak parameter ranges can be set for each historical peak parameter. For each historical peak parameter, after determining the peak parameter range it belongs to, a corresponding health evaluation standard can be determined based on the peak parameter range. Then, based on the matched evaluation standard, the health level corresponding to the historical peak parameter is determined, thus obtaining the parameter health level corresponding to the historical peak parameter.
[0039] Optionally, the parameter health of the historical peak parameter is determined based on the peak parameter interval in which the historical peak parameter is located. This includes: matching the peak parameter type of the historical peak parameter with the corresponding peak interval determination table, and determining the target parameter interval in which the historical peak parameter is located from the peak interval determination table; determining the interval health maximum and interval gain coefficient corresponding to the historical peak parameter based on the target parameter interval in which the historical peak parameter is located; multiplying the difference between the historical peak parameter and the interval minimum of the target parameter interval with the interval gain coefficient to obtain the health loss value, and taking the difference between the interval health maximum and the health loss value as the parameter health of the historical peak parameter.
[0040] The peak interval determination table can be a table of criteria for determining the health of load peak parameters. Specifically, each type of historical peak parameter can have its corresponding peak interval determination table, so the corresponding peak interval determination table can be obtained by matching the peak parameter type of each historical peak parameter. The peak interval determination table can contain multiple peak intervals and corresponding health determination formulas. The corresponding health determination formula can be determined based on the peak interval in which the historical peak parameter is located in the peak interval determination table, and then the parameter health of the historical peak parameter can be determined based on this formula. Furthermore, the target parameter interval can be the peak interval in which the historical peak parameter is located in the peak interval determination table. The maximum health value can be the maximum value of the parameter health corresponding to the target parameter interval. The interval gain coefficient can be the coefficient for increasing the health loss value corresponding to the target parameter interval. Specifically, the maximum health value and the interval gain coefficient can be determined based on matching the target parameter interval. The health loss value can represent the degree of loss of the historical peak parameter compared to the ideal health (i.e., the maximum health value) corresponding to the target parameter interval. Specifically, the health loss value can be obtained by multiplying the difference between the historical peak parameter and the minimum value of the target parameter interval by the interval gain coefficient. Finally, the difference between the maximum health value in the interval and the health loss value can be used as the parameter health value corresponding to the historical peak parameter.
[0041] For example, scoring rules can be defined for four metrics: CPU load, memory load, disk I / O load, and network bandwidth load, as follows: Table 1 CPU Load Scoring Table Table 2 Memory Load Scoring Table Table 3 Disk I / O Load Scoring Table Table 4 Network Bandwidth Load Scoring Table S260. The health scores of all historical peak parameters are weighted and summed to obtain the node health score corresponding to the node to be analyzed.
[0042] Node health can be determined from multiple perspectives using health parameters of the node to be analyzed. Specifically, after determining the health value of each historical peak parameter, the health values of all historical peak parameters can be combined for health analysis to determine the node health value of the node to be analyzed. Alternatively, the health values of all historical peak parameters can be weighted and summed to obtain the node health value of the node to be analyzed.
[0043] Optional, the weighting coefficients for various health parameters are shown below: Therefore, the formula for calculating node health is: .
[0044] S270. Based on the data value parameters and node health, determine the target node from the node to be analyzed, and store the network behavior data set in the target node.
[0045] The target node can be a node used to store the network behavior data set. Specifically, for each piece of current behavior data in the network behavior data set, the target node corresponding to each piece of current behavior data can be determined based on the data value parameter of the current behavior data and the node health of the node to be analyzed. Finally, each piece of current behavior data is stored in the corresponding target node.
[0046] Optionally, the amount of data stored in each node to be analyzed and the average value of all stored data (referred to as the node data average value parameter) can be recorded separately. Nodes to be analyzed whose node data average value parameter is less than the data value parameter of the current behavior data can be selected as initial nodes, and the initial node with the highest node health can be selected as the target node corresponding to the current behavior data.
[0047] For example, the amount of data stored in each node to be analyzed can be recorded. Average value parameter of node data Let the data value parameter of the current behavior data be denoted as . Filter out All nodes, and select the node health status from them. The highest-ranking node serves as the final data node for storing the current behavior data.
[0048] The technical solution provided in this invention involves acquiring a network behavior data set, and for each current behavior data point in the set, determining a relative data density parameter based on the current behavior data and a preset number of other behavior data points preceding it; determining a data information content parameter based on the current behavior data and the preset number of other behavior data points preceding it; substituting the relative data density parameter and the data information content parameter into a preset sigmoid function to obtain a data value parameter for the current behavior data; for each node to be analyzed, determining a historical peak parameter based on its historical load data; for each historical peak parameter, determining a parameter health level based on the peak parameter range it falls within; weighted summing of the parameter health levels of all historical peak parameters to obtain the node health level corresponding to the node to be analyzed; determining a target node from the node to be analyzed based on the data value parameter and the node health level, and storing the network behavior data set at the target node. The technical solution of this invention can comprehensively analyze the data value of network behavior data and the historical load characteristics of nodes, thereby determining the storage nodes for network behavior data. This can further improve the rationality of data storage node allocation and thus improve data query and retrieval efficiency.
[0049] The advantages of this application compared to the prior art are as follows: 1. A data value calculation method is proposed. The relative density and information content of the data are obtained through a sliding window method, and the two are combined through an optimized Sigmoid function to obtain a data value score. This method can evaluate the data value of the current data and also reflects the characteristics of the data distribution. 2. A complete method for calculating the load rate of storage nodes is proposed. The comprehensive health score of storage nodes is calculated by the sliding window method, which eliminates the need for real-time monitoring of storage nodes, reduces the performance overhead of computing and management, and reflects the long-term load status of each node. 3. By combining the data value of the current data with the comprehensive health score of the storage nodes, dynamic routing of data storage nodes is performed. This comprehensively considers data characteristics and storage node load, further improving the rationality of data storage node allocation, and thus improving data query and retrieval efficiency.
[0050] Figure 3 This is a schematic diagram of a data storage node analysis device provided in an embodiment of the present invention. The embodiment of the present invention can be applied to scenarios where storage nodes for network behavior data are selected. The device can be implemented by software and / or hardware and integrated into a computer device with application development capabilities.
[0051] like Figure 3 As shown, the data storage node analysis device includes: a data value determination module 310, a node health determination module 320, and a storage node selection module 330.
[0052] The data value determination module 310 is used to acquire a network behavior data set and determine data value parameters based on the data density characteristics and data information content characteristics of the network behavior data set; the node health determination module 320 is used to acquire historical load data corresponding to multiple nodes to be analyzed and determine the node health corresponding to each node to be analyzed based on the historical load data; the storage node selection module 330 is used to determine a target node from the nodes to be analyzed based on the data value parameters and node health, and store the network behavior data set in the target node.
[0053] The technical solution provided by this invention involves acquiring a network behavior data set, determining data value parameters based on the data density and information content characteristics of the network behavior data set, acquiring historical load data corresponding to multiple nodes to be analyzed, determining the node health of each node to be analyzed based on the historical load data, determining a target node from the nodes to be analyzed based on the data value parameters and node health, and storing the network behavior data set in the target node. This technical solution can comprehensively analyze the data value of network behavior data and the historical load characteristics of nodes to determine the storage node for the network behavior data, further improving the rationality of data storage node allocation and thus improving data query and retrieval efficiency.
[0054] In one optional implementation, the data value determination module 310 is specifically configured to: for each piece of current behavior data in the network behavior data set, determine a relative data density parameter of the current behavior data based on the current behavior data and a preset number of other behavior data in the network behavior data set preceding the current behavior data; determine a data information content parameter of the current behavior data based on the current behavior data and a preset number of other behavior data in the network behavior data set preceding the current behavior data; and substitute the relative data density parameter and the data information content parameter into a preset Sigmoid function to obtain a data value parameter of the current behavior data.
[0055] In one optional implementation, the data value determination module 310 includes: a data density parameter determination unit, configured to: use the ratio of the number of attribute sub-data with values in a preset number of other behavioral data to the total number of attribute sub-data in all other behavioral data as the average data density corresponding to the other behavioral data; use the ratio of the number of attribute sub-data with values in the current behavioral data to the total number of attribute sub-data in the current behavioral data as the current data density corresponding to the current behavioral data; and use the ratio of the current data density to the average data density as the relative data density parameter of the current behavioral data.
[0056] In an optional implementation, the data value determination module 310 includes: a data information quantity parameter determination unit, configured to: determine the attribute information entropy corresponding to each attribute sub-data in the current behavior data and the probability of the attribute value of each attribute sub-data appearing in other behavior data; and perform a weighted summation of the attribute information entropies of all attribute sub-data to obtain the data information quantity parameter of the current behavior data.
[0057] In an optional implementation, the node health determination module 320 is specifically used to: for each node to be analyzed, determine historical peak parameters based on the historical load data of the node to be analyzed; wherein, the historical peak parameters include at least two of the following: CPU load peak, memory load peak, disk I / O load peak, and network bandwidth load peak; for each historical peak parameter, determine the parameter health corresponding to the historical peak parameter based on the peak parameter range in which the historical peak parameter is located; and perform a weighted summation of the parameter health of all historical peak parameters to obtain the node health corresponding to the node to be analyzed.
[0058] In an optional implementation, the node health determination module 320 includes: a parameter health determination unit, configured to: match the peak parameter type of the historical peak parameter to a corresponding peak interval determination table, and determine the target parameter interval in which the historical peak parameter is located from the peak interval determination table; determine the interval health maximum and interval gain coefficient corresponding to the historical peak parameter based on the target parameter interval in which the historical peak parameter is located; multiply the difference between the historical peak parameter and the interval minimum of the target parameter interval by the interval gain coefficient to obtain a health loss value, and use the difference between the interval health maximum and the health loss value as the parameter health corresponding to the historical peak parameter.
[0059] In one optional implementation, the preset Sigmoid function is: in, Indicates the data value parameter, Indicates the amount of data information. This refers to the relative data density parameter; This is the threshold parameter.
[0060] The data storage node analysis device provided in this embodiment of the invention can execute the data storage node analysis method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the method execution.
[0061] Figure 4 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present invention. Figure 4 A block diagram of an exemplary computer device 12 suitable for implementing embodiments of the present invention is shown. Figure 4 The computer device 12 shown is merely an example and should not be construed as limiting the functionality or scope of the embodiments of the present invention. The computer device 12 can be any terminal device with computing capabilities and can be configured within a data storage node analysis device.
[0062] like Figure 4 As shown, the computer device 12 is represented in the form of a general-purpose computing device. The components of the computer device 12 may include, but are not limited to: one or more processors or processing units 16, system memory 28, and bus 18 connecting different system components (including system memory 28 and processing unit 16).
[0063] Bus 18 can be one or more of several bus architectures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the various bus architectures. For example, these architectures include, but are not limited to, the Industry Standard Architecture (ISA) bus, the Micro Channel Architecture (MAC) bus, the Enhanced ISA bus, the Video Electronics Standards Association (VESA) local bus, and the Peripheral Component Interconnect (PCI) bus.
[0064] Computer device 12 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by computer device 12, including volatile and non-volatile media, removable and non-removable media.
[0065] System memory 28 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache 32. Computer device 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 34 may be used to read and write non-removable, non-volatile magnetic media (… Figure 4 Not shown; usually referred to as a "hard drive"). Although Figure 4 Not shown, a disk drive for reading and writing to a removable non-volatile disk (e.g., a "floppy disk") and an optical disk drive for reading and writing to a removable non-volatile optical disk (e.g., a CD-ROM, DVD-ROM, or other optical media) may be provided. In these cases, each drive may be connected to bus 18 via one or more data media interfaces. System memory 28 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of the embodiments of the present invention.
[0066] A program / utility 40 having a set (at least one) of program modules 42 may be stored, for example, in system memory 28. Such program modules 42 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment. Program modules 42 typically perform the functions and / or methods described in the embodiments of the present invention.
[0067] Computer device 12 can also communicate with one or more external devices 14 (e.g., keyboard, pointing device, display 24, etc.), and with one or more devices that enable a user to interact with computer device 12, and / or with any device that enables computer device 12 to communicate with one or more other computing devices (e.g., network card, modem, etc.). This communication can be performed through input / output (I / O) interface 22. Furthermore, computer device 12 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 20. Figure 4 As shown, network adapter 20 communicates with other modules of computer device 12 via bus 18. It should be understood that, although... Figure 4 As not shown, it can be used in conjunction with computer device 12 with other hardware and / or software modules, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0068] Processing unit 16 executes various functional applications and data processing by running programs stored in system memory 28, such as implementing the data storage node analysis method provided in this embodiment of the invention, which includes: A network behavior data set is obtained, and data value parameters are determined based on the data density and data information volume characteristics of the network behavior data set; historical load data corresponding to multiple nodes to be analyzed is obtained, and node health is determined for each node to be analyzed based on the historical load data; a target node is determined from the nodes to be analyzed based on the data value parameters and node health, and the network behavior data set is stored in the target node.
[0069] This embodiment provides a computer-readable storage medium storing a computer program thereon. When executed by a processor, the program implements the data storage node analysis method as provided in any embodiment of the present invention, including: A network behavior data set is obtained, and data value parameters are determined based on the data density and data information volume characteristics of the network behavior data set; historical load data corresponding to multiple nodes to be analyzed is obtained, and node health is determined for each node to be analyzed based on the historical load data; a target node is determined from the nodes to be analyzed based on the data value parameters and node health, and the network behavior data set is stored in the target node.
[0070] The computer storage medium of this invention can be any combination of one or more computer-readable media. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0071] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, capable of sending, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device.
[0072] Program code contained on a computer-readable medium may be transmitted using any suitable medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.
[0073] Computer program code for performing the operations of this invention can be written in one or more programming languages or a combination thereof. Programming languages include object-oriented programming languages such as C, Java, Smalltalk, C++, C#, and Python, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., via the Internet using entropy provided by an Internet service).
[0074] Those skilled in the art will understand that the modules or steps of the present invention described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. Optionally, they can be implemented using computer-executable program code, thereby allowing them to be stored in a storage device for execution by a computing device, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, the present invention is not limited to any particular combination of hardware and software.
[0075] Note that the above description is merely a preferred embodiment of the present invention and the technical principles employed. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions can be made without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments, and may include many other equivalent embodiments without departing from the concept of the present invention, the scope of which is determined by the scope of the appended claims.
Claims
1. A data storage node analysis method, characterized in that, include: Obtain a set of network behavior data, and determine data value parameters based on the data density characteristics and data information content characteristics of the set of network behavior data; Obtain historical load data for multiple nodes to be analyzed, and determine the node health of each node to be analyzed based on the historical load data. Based on the data value parameters and node health, a target node is determined from the node to be analyzed, and the network behavior data set is stored in the target node.
2. The method according to claim 1, characterized in that, The determination of data value parameters based on the data density characteristics and data information content characteristics of the network behavior data set includes: For each current behavior data in the network behavior data set, a relative data density parameter for the current behavior data is determined based on the current behavior data and a preset number of other behavior data preceding the current behavior data in the network behavior data set. Based on the current behavior data and a preset number of other behavior data preceding the current behavior data in the network behavior data set, determine the data information quantity parameter of the current behavior data; Substituting the relative data density parameter and the data information content parameter into the preset Sigmoid function, the data value parameter of the current behavior data is obtained.
3. The method according to claim 2, characterized in that, The process of determining the relative data density parameter of the current behavior data based on the current behavior data and a preset number of other behavior data preceding the current behavior data in the network behavior data set includes: The ratio of the number of attribute sub-data with values in a preset number of other behavioral data to the total number of attribute sub-data in all other behavioral data is used as the average data density corresponding to the other behavioral data. The ratio of the number of attribute sub-data with values in the current behavior data to the total number of attribute sub-data in the current behavior data is used as the current data density corresponding to the current behavior data. The ratio of the current data density to the average data density is used as the relative data density parameter of the current behavioral data.
4. The method according to claim 2, characterized in that, The step of determining the data information content parameter of the current behavior data based on the current behavior data and a preset number of other behavior data preceding the current behavior data in the network behavior data set includes: For each attribute sub-data in the current behavior data, and the probability of the occurrence of the attribute value of each attribute sub-data in other behavior data, determine the attribute information entropy corresponding to the attribute sub-data. The data information quantity parameter of the current behavior data is obtained by weighted summing of the attribute information entropy of all attribute sub-data.
5. The method according to claim 1, characterized in that, The process of determining the node health of each node to be analyzed based on the historical load data includes: For each node to be analyzed, historical peak parameters are determined based on the historical load data of the node to be analyzed; wherein, the historical peak parameters include at least two of the following: CPU load peak, memory load peak, disk I / O load peak, and network bandwidth load peak. For each historical peak parameter, the parameter health level corresponding to the historical peak parameter is determined based on the peak parameter range in which the historical peak parameter is located; The node health score of the node to be analyzed is obtained by weighted summation of the health scores of all historical peak parameters.
6. The method according to claim 5, characterized in that, The process of determining the parameter health of the historical peak parameter based on the peak parameter range in which the historical peak parameter is located includes: Based on the peak parameter type matching table corresponding to the historical peak parameter, the target parameter interval in which the historical peak parameter is located is determined from the peak interval determination table. Based on the target parameter range in which the historical peak parameter is located, determine the interval health maximum and interval gain coefficient corresponding to the historical peak parameter; The health loss value is obtained by multiplying the difference between the historical peak parameter and the minimum value of the target parameter interval by the interval gain coefficient, and the difference between the maximum value of the interval health and the health loss value is taken as the parameter health corresponding to the historical peak parameter.
7. The method according to claim 2, characterized in that, The preset Sigmoid function is: in, Indicates the data value parameter, Indicates the amount of data information. This refers to the relative data density parameter; This is the threshold parameter.
8. A data storage node analysis device, characterized in that, The device includes: The data value determination module is used to acquire a set of network behavior data and determine data value parameters based on the data density characteristics and data information content characteristics of the set of network behavior data. The node health determination module is used to acquire historical load data corresponding to multiple nodes to be analyzed, and determine the node health of each node to be analyzed based on the historical load data. The storage node selection module is used to determine the target node from the nodes to be analyzed based on the data value parameters and node health, and to store the network behavior data set in the target node.
9. A computer device, characterized in that, The computer device includes: One or more processors; Memory, used to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the data storage node analysis method as described in any one of claims 1-7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by the processor, the program implements the data storage node analysis method as described in any one of claims 1-7.