A distributed storage method and system based on AI computing power analysis
By building an AI computing power prediction model, optimizing the computing power state management of routing nodes and storage nodes in distributed storage systems, the cluster failure problem of traditional distributed storage under high concurrency conditions is solved, and more efficient storage resource scheduling and low-latency response are achieved.
Patent Information
- Application Number
- CN202510402749.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-04-01
AI Technical Summary
Traditional distributed storage methods are prone to cluster failure under high concurrency conditions, and have poor results in high-precision real-time query, and the upper limit of the performance of the main control node limits concurrency performance.
Build a distributed storage system based on AI computing power analysis. By obtaining the computing power status data of routing nodes and storage nodes, an AI computing power prediction model is built, predicting the computing power status of nodes of different levels, and allocating storage tasks according to business needs to optimize storage resource scheduling.
It improves the stability and adaptability of distributed storage, improves processing efficiency under high concurrency conditions and responds to low-latency business scenarios.
Smart Images

Figure CN119917479B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of storage technology, and particularly to a distributed storage method and system based on AI computing power analysis. Background Art
[0002] Currently, traditional distributed storage methods only perform sharded storage according to the storage capacities of different distributed nodes, and perform scheduling storage of data files through a load balancing method. For example, in the traditional HFDS system (distributed file system), the system includes a naming node and data nodes, where the naming node is the main control node for controlling the storage scheduling of data nodes; however, the traditional HFDS system has high requirements for hardware devices. Since the performance of a single main control node has an upper limit, when performing large-scale data processing and storage, it is necessary to rely on multiple main control nodes for processing, thus limiting the concurrency performance of the HFDS system. In addition, the traditional HFDS system has poor effects on low-latency service scenarios such as high-precision real-time queries. Especially in high-concurrency real-time query scenarios, the traditional HFDS system is very likely to cause the main control node to crash, resulting in the unavailability of the data node cluster connected to the main control node. Therefore, the above-mentioned traditional distributed storage is prone to problems of distributed storage cluster failures under high concurrency conditions due to the single function of nodes and the problem of performance upper limits. Summary of the Invention
[0003] One of the invention objects of the present invention is to provide a distributed storage method and system based on AI computing power analysis. The method and system construct an AI computing power prediction model of a distributed computing power cluster by obtaining dynamic computing power data including routing nodes and storage nodes, convert the data related to routing nodes and storage nodes into computing power feature data, use the AI computing power prediction model to predict the computing power values of different levels of routing nodes and storage nodes, and perform distributed storage in combination with the service requirements of the current storage task. In the present invention, the AI computing power prediction model can better predict the computing power states of different routing nodes and storage nodes at different times, so as to meet the distributed storage under specific storage service rules and improve the service adaptability of storage resource scheduling.
[0004] Another object of the present invention is to provide a distributed storage method and system based on AI computing power analysis. The method and system further include performing heterogeneous computing power analysis on different storage nodes with different computing power resources by using an AI computing power prediction model, and judging the cooperative processing efficiency between different computing power resource devices of the corresponding storage nodes according to the heterogeneous computing power analysis results. That is, the present invention collects the cooperative processing data between the different computing power resource devices, converts the cooperative processing data into heterogeneous computing power feature data, and inputs the heterogeneous computing power feature data into the AI computing power prediction model for training, so that the AI computing power prediction model can predict the computing power state of the heterogeneous computing power analysis model of different storage nodes, so that the present invention can perform distributed storage to meet the heterogeneous computing power storage requirements and improve the heterogeneous processing efficiency of distributed storage tasks.
[0005] Another object of the present invention is to provide a distributed storage method and system based on AI computing power analysis. The method and system perform total computing power analysis and prediction on each routing node of different levels of routing nodes based on the AI computing power prediction model, and perform distributed storage tasks of corresponding storage data packets on the corresponding routing nodes according to the analysis and prediction results and the corresponding storage task allocation rules. The storage task rules include storage task rules with read / write timeliness as the priority, storage task rules with durability as the priority, and storage task rules with structural form as the priority, etc., thereby greatly improving the adaptability of the distributed storage method to different storage task rules and greatly improving the stability of distributed storage.
[0006] In order to achieve at least one of the above object of the invention, the present invention further provides a distributed storage method based on AI computing power analysis, the method comprising:
[0007] S01. Pre-construct a heterogeneous storage resource, where the heterogeneous storage resource includes routing nodes and storage nodes at different levels, and the routing nodes are communicatively connected to at least one storage node or routing node;
[0008] S02. Obtain the computing power state data of each routing node and storage node in the heterogeneous storage resource according to the historical data of the heterogeneous storage resource, and convert the computing power state information into computing power feature data of each routing node and storage node;
[0009] S03. Use the computing power feature data to construct an AI computing power prediction model, and output the computing power state data of each routing node and storage node at different time series according to the AI computing power prediction model;
[0010] S04. Construct a priority rule for distributed storage tasks. According to the computing power status data of the corresponding routing node and storage node and the priority rule for distributed storage tasks, store the corresponding data packets into the corresponding storage nodes through the corresponding routing nodes, and update the storage status data of the storage nodes and routing nodes in real time.
[0011] According to one preferred embodiment of the present invention, the heterogeneous storage resources include: using at least one of a volatile memory, a non-volatile memory, and an object memory as a storage module, and using at least two of a CPU processor, a GPU processor, a DPU processor, a TPU processor, an FPGA module, a DPS processor, an ASIC chip, and an NPU processor as processing modules. Construct the heterogeneous storage resources with the at least one storage module and the at least two processing modules. The heterogeneous storage resources are constructed in the corresponding storage nodes, and a heterogeneous label of a corresponding type is configured for each heterogeneous storage resource, and the heterogeneous label is saved to the routing node maintained by the storage node for scheduling the corresponding storage task priority rule.
[0012] According to another preferred embodiment of the present invention, the method for calculating the computing power status data of the storage node includes: obtaining the computing power status data of the storage node under each historical time series in the storage node according to the construction rule of the time series model, where the computing power status data includes: the remaining capacity H of the storage resources of the corresponding storage node, the average value f of the floating-point computing power Flops of the processor of the corresponding storage node, the read / write speed V1 of the storage node, the routing time t between the routing node and the corresponding storage node, and the cooperative processor speed V2 between heterogeneous processors, and calculate the computing power comparison value P of the corresponding storage node through the following formula: P = , where the computing power comparison value P is the computing power status data for comparing the computing power differences between the current storage node and all storage nodes, where V s1 is the average value of the read / write speeds of all storage nodes obtained by calculation, V s2 is the average value of the cooperative processing speeds between heterogeneous processors of all storage nodes, and f0 is the average value of the floating-point computing power Flops of the processors of all storage nodes.
[0013] According to another preferred embodiment of the present invention, the computing power state data in the historical time series of the storage node is normalized into input data of corresponding types, and the normalized data is used as the cell state data of the LSTM model. The data including the type and size of the input data packet is preprocessed and used as the input data of the LSTM model. The output layer of the LSTM model outputs the predicted computing power state value of the storage node under the corresponding time series. The mean square error is used as the loss function of the LSTM model to perform the regression task of different computing power state data of the storage node, for regression prediction of the computing power state data of the corresponding storage node including the next time series.
[0014] According to another preferred embodiment of the present invention, the computing power state data of the routing node includes: the total storage capacity H of all storage nodes connected to the current routing node z , the average routing speed of the current routing node and all storage nodes and the total parallel routing speed V z , the highest routing speed V of the current routing node and the connected storage nodes 3max and the lowest routing speed V 3min , the routing speed V4 between the current routing node and the upper-level routing node, and the routing speed V5 between the current routing node and the lower-level routing node. The calculation method of the routing speed includes: obtaining the first timestamp when the current routing node receives the data packet, and when the data packet is received by the corresponding-level routing node or storage node, obtaining the second timestamp in real time when it is received. The routing speed is obtained according to the first timestamp, the second timestamp, and the data packet size.
[0015] According to another preferred embodiment of the present invention, the computing power state data in the historical time series of the routing node is normalized into input data of corresponding types, and the normalized data is used as the cell state data of the LSTM model. The data including the type and size of the input data packet is preprocessed and used as the input data of the LSTM model. The output layer of the LSTM model outputs the predicted routing computing power state value under the corresponding time series. The mean square error is used as the loss function of the LSTM model to perform the regression task of different computing power state data of the routing node, for regression prediction of the computing power state data of the corresponding routing node including the next time series.
[0016] According to another preferred embodiment of the present invention, the priority rules for storage tasks include: timeliness priority rule, persistence priority rule, processing method priority rule, redundancy priority rule, and structural form priority rule. Among them, the timeliness priority rule is used to establish storage interaction tasks in low-latency scenarios. Under the timeliness priority rule, the total computing power timeliness M of the computing routing node and the storage node is calculated as: M = α V1 + β V2 + γ V 3max + λ V4 + σ V5, where α, β, γ, λ, and σ are different weight parameters respectively. Under the timeliness priority rule, when the data packet size h is less than the total storage capacity H of the corresponding routing node and the corresponding storage node z at that time, select the predicted routing node with the maximum total computing power timeliness M max as the target routing node to send the data packet, and select the storage node with the maximum computing power comparison value P max as the target storage node for distributed storage of the data packet.
[0017] According to another preferred embodiment of the present invention, judge the data type in the data packet to be stored and the type of processing method required for this data packet, generate a processor type label according to the data packet type and the processing method type, and compare the processor type label with the processor type label of the corresponding storage node stored in the router node, and select the routing node with the same processor type label and meeting the timeliness priority rule to send the data packet.
[0018] To achieve at least one of the above invention purposes, the present invention further provides a distributed storage system based on AI computing power analysis, and the system executes the above-mentioned distributed storage method based on AI computing power analysis.
[0019] The present invention further provides a computer-readable storage medium, and the computer-readable storage medium stores a computer program, and the computer program is executed by a processor to implement the above-mentioned distributed storage method based on AI computing power analysis. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 Shows a flowchart of a distributed storage method based on AI computing power analysis of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0021] The following description is used to disclose the present invention so that those skilled in the art can implement the present invention. The preferred embodiments in the following description are only examples, and those skilled in the art can think of other obvious variations. The basic principles of the present invention defined in the following description can be applied to other embodiments, variants, improvements, equivalent schemes, and other technical solutions without departing from the spirit and scope of the present invention.
[0022] It can be understood that the term "a" should be understood as "at least one" or "one or more". That is, in one embodiment, the number of an element can be one, while in other embodiments, the number of the element can be multiple. The term "a" cannot be understood as a limitation on the quantity.
[0023] Please refer to Figure 1 , the present invention discloses a distributed storage method and system based on AI computing power analysis, and the method includes the following steps:
[0024] S01. Pre-construct a heterogeneous storage resource, where the heterogeneous storage resource includes routing nodes and storage nodes at different levels, and the routing nodes are communicatively connected to at least one storage node or routing node;
[0025] S02. Obtain the computing power status data of each routing node and storage node in the heterogeneous storage resource according to the historical data of the heterogeneous storage resource, and convert the computing power status information into the computing power characteristic data of each routing node and storage node;
[0026] S03. Use the computing power characteristic data to construct an AI computing power prediction model, and output the computing power status data of each routing node and storage node at different time series according to the AI computing power prediction model;
[0027] S04. Construct a distributed storage task priority rule, and store the corresponding data packet into the corresponding storage node through the corresponding routing node according to the computing power status data of the corresponding routing node and storage node and the distributed storage task priority rule, and update the storage status data of the storage node and routing node in real time.
[0028] In the present invention, a distributed storage node cluster needs to be pre-constructed. The distributed storage node cluster includes a routing node and storage nodes, and corresponding resource devices are configured for the corresponding routing nodes and storage nodes. In the present invention, the resource devices of the storage nodes include memory resources and processor resources. Among them, the storage nodes can be configured to include volatile memory or non-volatile memory. The volatile memory can include, but is not limited to, DRAM (Dynamic Random Access Memory) and SRAM (Static Random Access Memory). The volatile memory is generally used for memory storage and is often used in high-performance computing. The non-volatile memory includes, but is not limited to, Flash Memory, SSD (Solid State Drive), MRAM (Magnetic Random Access Memory), and EPROM (Erasable Programmable Read-Only Memory), etc. An appropriate number of the above-mentioned volatile memory and non-volatile memory are pre-configured in the corresponding storage nodes, so that the storage node cluster has good adaptability to diversified storage services. Among them, the processor resources in the corresponding storage nodes can be configured to include, but are not limited to, CPU processors, GPU processors, DPU processors, TPU processors, FPGA modules, DPS processors, ASIC chips, and NPU processors, etc. The above different types of processors can solve different data packet service processor problems. The above processors and memories can be configured at network edge nodes. Since different services require different types of processor resources, the present invention can pre-set at least the above 2 types of processors and at least 1 type of memory as the heterogeneous resource configuration of the corresponding storage nodes, so that the distributed storage nodes in the present invention have improved storage service efficiency and compatibility with different storage service rules. The present invention uses an AI model with time series to construct a prediction model for the computing power status of the storage nodes and routing nodes, and performs distributed storage of corresponding data based on setting relevant storage service rules and the computing power prediction results of different storage nodes and routing nodes. So that the corresponding storage data packets in the present invention can be diversifiedly stored based on different service rules, thereby adapting to service scenarios with different storage resource calls.
[0029] Specifically, the routing nodes can be configured at different levels. For example, the routing nodes can be constructed in a tree structure, and each routing node is communicatively connected to at least one storage node or routing node. The configurations on different routing nodes include routing processors, etc. In another feasible embodiment of the present invention, if a storage cluster with a topological network structure is constructed, corresponding storage resources can be configured on the corresponding routing nodes. So that the routing nodes can be used as storage nodes, and the corresponding storage nodes can be used as routing nodes under specific circumstances. It should be noted that the computing power state data of the routing nodes and storage nodes in the present invention are different. The computing power of the storage node is mainly determined by the overall performance of the computing resources of the current storage node, the read / write performance of the memory, the storage capacity, and the routing time. The overall computing power of the routing node is the sum of the computing powers of the storage nodes of all branches. Moreover, the calculation of the computing power state data of the routing nodes in the present invention also includes the routing time between routing nodes at different levels, providing reference data for the packet routing planning of different routing nodes. It is worth mentioning that the present invention needs to use the historical data of the distributed storage node cluster for AI model training, and use the historical computing power state data of each storage node and routing node in the node cluster as training data for the AI model training. Preferably, the long short-term memory network (LSTM) is used as the AI model carrying time series to train the historical computing power state data of the storage node routing nodes, and a computing power state prediction model for the corresponding storage nodes and routing nodes is obtained. The present invention performs distributed routing storage of corresponding data packets according to the computing power state data of each routing node and each storage node predicted by the AI model and the storage task priority rules through the pre-configured storage task priority rules. Thus, the distributed storage is adapted to various different business scenarios.
[0030] The long short-term memory network (LSTM) includes a cell state C t , an input gate i t , a forget gate f t , an output gate o t , and the input gate data and the forget gate data are input and integrated by element-wise multiplication to update the cell state. In the present invention, the normalization processing of the corresponding type input data of the computing power state data of the routing nodes and storage nodes is performed, and the normalized data is used as the cell state data C of the long short-term memory network (LSTM) model t , and the data including but not limited to the packet type and size obtained after preprocessing the input data packet of the corresponding storage node at the corresponding time series t is used as the input gate i t data, and combined with the forget gate f t to update the cell state C t+1, The long short-term memory (LSTM) model uses the mean squared error as the loss function to calculate the mean squared error values of the different computing power state data predicted by each storage node and routing node, and performs the regression task of the computing power state data of each storage node and routing node, for predicting the computing power state values of the corresponding storage node and routing node computing power state data in a specific time series. Among them, the training method and the optimized training method of the long short-term memory (LSTM) model are prior arts, and the present invention will not elaborate on this in detail.
[0031] It is worth mentioning that in the present invention, the routing and storage of corresponding data packets need to be performed according to the computing power state data predicted by the routing node and the actual storage task priority rules. For example, the storage task priority rules in the present invention include but are not limited to: timeliness priority rule, persistence priority rule, processing method priority rule, redundancy priority rule, and structural form priority rule. Among them, the timeliness priority rule is used for low-latency scenario storage tasks, such as real-time data analysis scenarios, real-time interactive autonomous driving scenarios, and real-time financial transaction scenarios, etc. In the above low-latency scenario storage tasks, DRAM (Dynamic Random Access Memory) can be used as a real-time cache, such as Redis, etc. The persistence priority rule is applicable to data persistence operations. For example, redundant storage of videos can be performed through RAID (Redundant Array of Independent Disks), which can be automatically recovered in case of partial disk failures to improve the security of data storage. For the processing method priority, it is set according to the processing rules of the data packet type. For example, some processing methods require the joint processing operation of heterogeneous processing resources of CPU and GPU. For example, performing fluid dynamics analysis requires parallel analysis and calculation using GPU, and at the same time, it is also necessary to use CPU to perform preprocessing such as setting parameter boundary conditions for data; therefore, it needs to be allocated to a storage node with specific heterogeneous processing resources. Similarly, for the structured priority rule, a storage node with a structured memory needs to be considered first. In the present invention, it is necessary to perform tagging processing on the heterogeneous storage resources and processing resource types in each storage node, obtain all the tags of the heterogeneous storage resources and processing resources in each storage node, and pre-store the tags into the routing node. By performing tag comparison in the routing node, it can be used for the data packet routing operation of the storage node that requires heterogeneous processing resources.
[0032] The method for obtaining and processing the computing power status data of the storage node includes: obtaining the computing power status data of the storage node under each historical time series in the storage node according to the construction rules of the time series model, where the computing power status data includes: the remaining storage capacity H of the corresponding storage node, the average floating-point computing power Flops f of the processor of the corresponding storage node, the memory read and write speed V1 of the storage node, the routing time t between the routing node and the corresponding storage node, and the cooperative processor speed V2 between heterogeneous processors, and calculating the computing power comparison value P of the corresponding storage node through the following formula: P = , where the computing power comparison value P is the computing power status data, which is used to compare the computing power differences between the current storage node and all storage nodes, where V s1 is the average read and write speed of all storage nodes obtained by calculation, V s2 is the average cooperative processing speed between heterogeneous processors of all storage nodes, and f0 is the average floating-point computing power Flops of the processors of all storage nodes. It should be noted that in the present invention, the above computing power comparison value is set to compare the computing power in the storage node, where the computing power comparison of the storage node takes into account the parameters of the storage capacity and processing speed of the storage node itself. Of course, in some other feasible embodiments of the present invention, only the parameters of the storage capacity or processing speed can be selected as the data routing storage parameters according to the specific storage service rules.
[0033] The computing power status data of the routing node includes: the total storage capacity H of all storage nodes connected to the current routing node z , the average routing speed between the current routing node and all storage nodes and the parallel total routing speed V z , the highest routing speed V and the lowest routing speed V between the current routing node and the connected storage node 3max 3min , the routing speed V4 between the current routing node and the upper-level routing node, and the routing speed V5 between the current routing node and the lower-level routing node, where the calculation method of the routing speed includes: obtaining the first timestamp when the current routing node receives the data packet, and when the data packet is received by the corresponding-level routing node or storage node, obtaining the second timestamp in real time when receiving, and obtaining the routing speed according to the first timestamp, the second timestamp and the data packet size.
[0034] In one preferred embodiment of the present invention, in order to better explain the different storage task priority rules, the present invention takes the time-sensitive storage priority rule as an example: under the time-sensitive priority rule, calculating the total computing power time sensitivity M of the routing node and the corresponding storage node: M = α V1 + β V2 + γ V3max +λ V4 + σ V5, where α, β, γ, λ, and σ are different weight parameters respectively. Under the aging priority rule, when the data packet size h is less than the total storage capacity H of the corresponding storage node of the corresponding routing node z select the prediction routing node with the maximum total computing power aging M max as the target routing node for data packet transmission, and select the P with the largest computing power comparison value max as the storage node as the target storage node for distributed data packet storage. In another feasible embodiment of the present invention, specific parameters of the total computing power aging M and the computing power comparison value P of the routing node and the corresponding storage node can be taken into account for weighted processing to obtain the corresponding computing power reference data of the corresponding routing node for routing storage processing of the corresponding data packet. The present invention does not make specific limitations on this
[0035] Embodiments disclosed in the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. Embodiments disclosed in the present invention include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network through a communication part, and / or installed from a removable medium. When the computer program is executed by a central processing unit (CPU), the above functions defined in the methods of the present application are performed. It should be noted that the computer-readable medium in the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection having one or more wire segments, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or combined with an instruction execution system, apparatus, or device. And in the present application, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program codes. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, and the computer-readable medium can send, propagate, or transmit a program for use by or combined with an instruction execution system, apparatus, or device. The program codes contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wireless, wire segments, optical cables, RF, etc., or any suitable combination of the above.
[0036] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that, in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0037] Those skilled in the art should understand that the embodiments of the present invention described above and shown in the accompanying drawings are only examples and do not limit the present invention. The objectives of the present invention have been fully and effectively achieved. The functions and structural principles of the present invention have been demonstrated and illustrated in the embodiments. Without departing from the said principles, the embodiments of the present invention can have any variations or modifications.
Claims
1. A distributed storage method based on AI computing power analysis, characterized in that, The method includes: S01. Pre-construct a heterogeneous storage resource, where the heterogeneous storage resource includes routing nodes and storage nodes at different levels, and where the routing nodes are communicatively connected to at least one storage node or routing node; S02. Obtain the computing power status data of each routing node and storage node in the heterogeneous storage resource according to the historical data of the heterogeneous storage resource, and convert the computing power status information into computing power characteristic data of each routing node and storage node; S03. Use the computing power characteristic data to construct an AI computing power prediction model, and output the computing power status data of each routing node and storage node at different time series according to the AI computing power prediction model; S04. Construct a distributed storage task priority rule, and store the corresponding data packet into the corresponding storage node through the corresponding routing node according to the computing power status data of the corresponding routing node and storage node and the distributed storage task priority rule, and update the storage status data of the storage node and routing node in real time; The heterogeneous storage resource includes: using at least one of a volatile memory, a non-volatile memory, and an object memory as a storage module, and using at least two of a CPU processor, a GPU processor, a DPU processor, a TPU processor, an FPGA module, a DPS processor, an ASIC chip, and an NPU processor as a processing module, constructing the heterogeneous storage resource with the at least one storage module and at least two processing modules, building the heterogeneous storage resource in the corresponding storage node, and configuring a corresponding type of heterogeneous tag for each heterogeneous storage resource, and saving the heterogeneous tag into the routing node maintained by the storage node for the scheduling of the corresponding storage task priority rule; The calculation method of the computing power status data of the storage node includes: obtaining the computing power status data of the storage node under each historical time series in the storage node according to the construction rules of the time series model, where the computing power status data includes: the remaining storage capacity H of the corresponding storage node, the average value f of the floating-point computing power Flops of the processor of the corresponding storage node, the read-write speed V1 of the storage node, the routing time t between the routing node and the corresponding storage node, the cooperative processor speed V2 between heterogeneous processors, and calculating the computing power comparison value P of the corresponding storage node through the following formula: P = , where the computing power comparison value P is the computing power status data, used to compare the computing power differences between the current storage node and all storage nodes, where V s1 is the average value of the read-write speeds of all storage nodes obtained by calculation, V s2 is the average value of the cooperative processing speeds between heterogeneous processors of all storage nodes, and f0 is the average value of the floating-point computing power Flops of the processors of all storage nodes; Perform normalization processing on the computing power status data of the storage node in the historical time series for the corresponding type of input data, and use the normalized data as the cell state data of the LSTM model, and use the data packet data type and size including the input after data preprocessing as the input data of the LSTM model, and the output layer of the LSTM model outputs the predicted storage node computing power status value at the corresponding time series; and use the mean square error as the loss function of the LSTM model to perform the regression task of the different computing power status data of the storage node for regression prediction of the computing power status data of the corresponding storage node including the next time series; The computing power status data of the routing node includes: the total storage capacity H of all storage nodes connected to the current routing node z , the average routing speed of the current routing node and all storage nodes and the total parallel routing speed V z , the highest routing speed V of the current routing node and the connected storage nodes 3max and the lowest routing speed V 3min , the routing speed V4 between the current routing node and the upper-level routing node, and the routing speed V5 between the current routing node and the lower-level routing node, where the calculation method of the routing speed includes: obtaining the first timestamp when the current routing node receives a data packet, and when the data packet is received by the routing node or storage node at the corresponding level, obtaining the second timestamp in real time when it is received, and obtaining the routing speed based on the first timestamp, the second timestamp, and the data packet size; The priority rules for the storage tasks include: timeliness priority rule, persistence priority rule, processing mode priority rule, redundancy priority rule, and structural form priority rule. Among them, the timeliness priority rule is used to establish storage interaction tasks in low-latency scenarios. Under the timeliness priority rule, the total computing power timeliness M of the computing routing node and the storage node is: M = α * V1 + β * V2 + γ * V 3max + λ * V4 + σ * V5, where α, β, γ, λ, and σ are different weight parameters respectively. Under the timeliness priority rule, when the data packet size h is less than the total storage capacity H of the corresponding routing node and the corresponding storage node z select the predicted routing node with the maximum total computing power timeliness M max as the target routing node for data packet transmission, and select the P with the maximum computing power comparison value max as the storage node as the target storage node for distributed data packet storage.
2. The distributed storage method based on AI computing power analysis according to claim 1, wherein, Perform normalization processing on the computing power status data of the routing node in the historical time series for the corresponding type of input data, and use the normalized data as the cell state data of the LSTM model, and use the data packet data type and size including the input after data preprocessing as the input data of the LSTM model, and the output layer of the LSTM model outputs the predicted routing computing power status value at the corresponding time series, and use the mean square error as the loss function of the LSTM model to perform the regression task of the different computing power status data of the routing node for regression prediction of the computing power status data of the corresponding routing node including the next time series.
3. A distributed storage method based on AI computing power analysis according to claim 1, characterized in that, Determine the data type in the data packet to be stored and the type of processing method required for the data packet, generate a processor type label according to the data packet type and the processing method type, and compare the processor type label with the processor type label of the corresponding storage node stored in the router node, and select a routing node with the same processor type label and meeting the aging priority rule to send the data packet.
4. A distributed storage system based on AI computing power analysis, characterized in that, The system executes a distributed storage method based on AI computing power analysis according to any one of claims 1-3.
5. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program is executed by a processor to implement a distributed storage method based on AI computing power analysis according to any one of claims 1-3.
Citation Information
Patent Citations
Task awareness and intelligent scheduling method for convergence network
CN119212105A