Distributed storage system based on large model optimization

Through in-depth analysis of parameter distribution data and dynamic allocation strategies, the problem of uneven load of storage nodes in distributed storage systems is solved, and the efficiency and system adaptability of large-model optimization tasks are improved.

CN120255801AInactive Publication Date: 2025-07-04GUANGDONG OPEN UNIV (GUANGDONG POLYTECHNIC VOCATIONAL COLLEGE)
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510308654.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-07-04
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing distributed storage systems are difficult to effectively utilize hardware resources in large-scale optimization scenarios, resulting in uneven loading of storage nodes, forming I/O bottlenecks, and affecting model optimization efficiency.

Method used

The parameter distribution data of the model training task is obtained through the parameter analysis module, and the access frequency and dependencies are combined for grouping processing, dynamically calculate the allocation weights of the storage nodes, and real-time monitoring of load information, dynamically adjusting the parameter allocation plan to achieve efficient migration and synchronization of parameters.

Benefits of technology

It significantly improves the utilization rate of storage resources and data access efficiency, reduces I/O bottleneck problems, improves the performance of large-model training and inference tasks, and adapts to efficient operation in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120255801A_ABST
    Figure CN120255801A_ABST
Patent Text Reader

Abstract

The invention provides a distributed storage system based on large model optimization, and relates to the technical field of data processing, and the system comprises a parameter analysis module which is used for obtaining parameter distribution data of a model training task; the parameter grouping module is used for performing grouping processing on the parameter distribution data; the distribution weight calculation module is used for calculating the distribution weight of each storage node in combination with the load information and the hardware performance indexes of the storage nodes; the parameter distribution module is used for distributing the parameter grouped data to different storage nodes; the dynamic monitoring module is used for monitoring access requests and load information of the storage nodes in real time in a model training or reasoning process and generating feedback information; the scheme adjusting module is used for dynamically adjusting the parameter distribution scheme according to the feedback information; the data migration module is used for executing migration operation of parameter data on the target storage node and the source storage node and updating the storage state of the parameters; the system processing capacity and efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly to a distributed storage system optimized based on a large model. Background Art

[0002] The distributed storage system is one of the key technologies in the current cloud computing and big data fields. By storing data on multiple physical nodes, it realizes high availability and disaster tolerance of data. In the prior art, the distributed storage system usually adopts the consistent hashing algorithm or the sharding algorithm to distribute data to each storage node, and ensures the reliability of data through the replication mechanism. At the same time, most distributed storage systems use a centralized metadata management node to record the distribution location of data to support fast query and storage operations. However, in order to meet the large-scale data access requirements, the system usually needs to dynamically adjust the data distribution strategy to address the load balancing problem, which is particularly important in the large model optimization task.

[0003] However, the existing distributed storage systems have the problem of being difficult to effectively utilize hardware resources when facing the large model optimization scenario. For example, in deep learning tasks, large model training and inference usually require frequent access to and update of large-scale parameter data. In the prior art, the fixed sharding strategy is used to allocate data to storage nodes, which may cause some hot parameters to be concentrated on a small number of nodes, resulting in uneven load. Specifically, when a large amount of data carried by a storage node is frequently accessed, an I / O bottleneck may occur, thus affecting the model optimization efficiency. Especially in the scenarios of real-time training or online inference, this bottleneck has a significant impact on the overall system performance. Summary of the Invention

[0004] The purpose of the present invention is to provide a distributed storage system optimized based on a large model, aiming to solve the problems mentioned in the background art.

[0005] To solve the above technical problems, the technical solution of the present invention is as follows:

[0006] A distributed storage system optimized based on a large model, the system includes:

[0007] A parameter analysis module, configured to obtain the parameter distribution data of the model training task, and generate a parameter access analysis result based on the parameter distribution data, where the parameter distribution data includes parameter identification information, the original record of access frequency, and the original record of dependency relationship;

[0008] A parameter grouping module, configured to calculate the similarity of parameter access frequency and the degree of dependency of parameter dependency relationship according to the parameter access analysis result, perform grouping processing on the parameter distribution data, and generate parameter grouping data;

[0009] The weight allocation calculation module is used to group parameter data according to parameters, calculate the allocation weights of each storage node in combination with the load information and hardware performance indicators of the storage nodes, and generate an allocation weight set;

[0010] The parameter allocation module is used to allocate the parameter grouped data to different storage nodes according to the allocation weight set, and generate a parameter allocation scheme;

[0011] The dynamic monitoring module is used to monitor the access requests and load information of the storage nodes in real time during the model training or inference process, and generate feedback information when the access request volume or load exceeds the preset range;

[0012] The scheme adjustment module is used to dynamically adjust the parameter allocation scheme according to the feedback information to obtain a dynamic allocation scheme;

[0013] The data migration module is used to perform the migration operation of parameter data on the target storage node and the source storage node according to the dynamic allocation scheme, and update the storage status of the parameters;

[0014] The status synchronization module is used to synchronize the updated storage status of the parameters to all storage nodes after the parameter migration is completed.

[0015] Preferably, the parameter analysis module includes:

[0016] The data acquisition unit is used to receive parameter distribution data from the model training task. The parameter distribution data includes parameter identification information, original access frequency records, and original dependency relationship records;

[0017] The data cleaning unit is used to clean the parameter distribution data, remove redundant records, correct abnormal data, and format it into a standardized data structure to generate cleaned parameter data;

[0018] The frequency analysis unit is used to extract the access frequency information of each parameter from the cleaned parameter data, and generate a frequency distribution result in combination with the frequency change trend. The frequency distribution result includes the classification interval of the parameter access frequency and the access mode characteristics;

[0019] The dependency relationship analysis unit is used to extract the dependency relationship records between each parameter from the cleaned parameter data and construct a dependency relationship network;

[0020] The parameter priority calculation unit is used to calculate the priority score of each parameter according to the frequency distribution result and the dependency relationship network to obtain the parameter priority score. The calculation formula of the parameter priority score is:

[0021] P i =α1×F i +β1×log(1+∑ j∈Dep(i) D ij) + γ1 × T i , where P i is the priority score of parameter i, F i is the access frequency of parameter i, and its normalized value range is [0, 1]. Dep(i) is the set of all other parameters that have a dependency relationship with parameter i, D ij is the strength of the dependency relationship between parameter i and parameter j, T i is the historical access frequency change trend of parameter i, specifically represented as the time function value of parameter i. α1, β1, and γ1 are weight coefficients; where,

[0022] Comm ij is the number of times parameter i and parameter j are accessed simultaneously, Freq ij is the joint access frequency of parameter i and parameter j, C i and C j are the storage data volumes of parameter i and parameter j respectively;

[0023] Freq i,t is the access frequency of parameter i at time point t, T is the total number of time points, and λ1 is a weight coefficient;

[0024] The analysis result generation unit is used to combine the frequency distribution result, the dependency relationship network, and the parameter priority score to form the parameter access analysis result.

[0025] Preferably, the parameter grouping module includes:

[0026] The frequency grouping unit is used to calculate the access frequency similarity of each parameter according to the frequency distribution result and the parameter priority score in the parameter access analysis result, and generate frequency grouping data;

[0027] The dependency grouping unit is used to analyze the strength of the dependency relationship between parameters according to the dependency relationship network and the parameter priority score in the parameter access analysis result, and generate dependency grouping data;

[0028] The grouping quantization unit is used to comprehensively quantify the parameters in the frequency grouping data and the dependency grouping data, and generate parameter grouping data according to the comprehensive quantization result; where, the comprehensive quantization includes:

[0029] Calculate the comprehensive weight value of each parameter in different groups, and the group corresponding to the highest weight value is regarded as the preferred belonging group of the parameter; the calculation formula of the comprehensive weight value is:

[0030] where, W i,k is the comprehensive weight value of parameter i in group k, F i,kLet \(D\) be the access frequency similarity of parameter \(i\) in group \(k\). i,k Let \(L\) be the dependence strength of parameter \(i\) in group \(k\). k Let \(C\) be the data volume of group \(k\). k Let \(Var(F)\) be the variance of \(F\), where \(F\) is the storage capacity of group \(k\), and \(\alpha^2\), \(\beta^2\), \(\gamma^2\), and \(\lambda^2\) are weight coefficients. Let \(F\) be the access frequency similarity of parameter \(i\) in group \(k\). Among them, i,k ) be \(F\) i,k The variance of \(F\), \(\alpha^2\), \(\beta^2\), \(\gamma^2\), and \(\lambda^2\) are weight coefficients, and \(F\) i,k is the access frequency similarity of parameter \(i\) in group \(k\); where,

[0031] Let \(G\) k be the set of parameters in group \(k\), and \(P\) j be the priority score of parameter \(j\).

[0032] Let \(CosSim(Vec i , Vec j ) be the cosine similarity of the access patterns of parameter \(i\) and parameter \(j\), that is, the access frequency similarity of parameter \(i\) and parameter \(j\). \(Vec i and \(Vec j are the access frequency vectors of parameter \(i\) and parameter \(j\) respectively. Let \(Freq j,t be the access frequency of parameter \(j\) at time point \(t\).

[0033] When there are shared parameters between groups, the relevant groups are merged into one group.

[0034] Preferably, the allocation weight calculation module includes:

[0035] A load acquisition unit for acquiring the real-time load information of the storage node. The real-time load information includes the current storage capacity usage and access processing ability of the node.

[0036] A performance evaluation unit for calculating the hardware performance indicators of the storage node.

[0037] A weight calculation unit for calculating the allocation weights of each storage node according to the parameter grouping data, the load information of the storage node, and the node performance evaluation value, and generating an allocation weight set. The allocation weight reflects the preferential bearing capacity of each storage node for different parameter groups. The calculation formula for the allocation weight is:

[0038] where \(P k,j is the allocation weight of group \(k\) on storage node \(l\), \(C l is the remaining capacity of storage node \(l\), \(H l is the hardware performance indicator of storage node \(l\), \(R l is the available bandwidth of storage node \(l\), and \(L lis the current load of storage node l, D k is the storage requirement of group k, T il is the transmission time allocated to storage node l for parameter i, and δ is the weight coefficient; where

[0039] BW l is the network bandwidth of storage node l, IOPS l is the number of input / output operations per second of storage node l, Lat l is the storage access latency of storage node l.

[0040] Preferably, the parameter allocation module includes:

[0041] A group priority determination unit, configured to calculate the priority score of a group according to the comprehensive weight value of each parameter group and the size of the storage requirement, so as to obtain the group priority score; the calculation formula of the group priority score is:

[0042] where, P k is the priority score of group k, Var(P i,k ) is the variance of P i,k , P i,k is the parameter priority score of parameter i in group k, L k is the storage requirement of group k, and σ and η are weight coefficients;

[0043] An adaptability calculation unit, configured to calculate the adaptability score between the storage node and the parameter according to the group priority score, the parameter priority score and the allocation weight value;

[0044] A parameter allocation unit, configured to allocate the corresponding storage node and parameter according to the order of the adaptability scores from large to small, and generate a parameter allocation scheme;

[0045] A first allocation verification unit, configured to verify the rationality of the parameter allocation scheme, including load balancing, localized storage of dependent parameters, and matching of group priority and storage node performance.

[0046] Preferably, the calculation formula of the adaptability score is:

[0047] where, S i,l is the adaptability score of parameter i allocated to storage node l, is the access latency angle between parameter i and storage node l, and ∈ is a constant to prevent division by zero.

[0048] Preferably, the dynamic monitoring module includes:

[0049] An access monitoring unit for real-time monitoring of access request information of storage nodes, including the number of accesses, request types, and target parameters;

[0050] A load monitoring unit for real-time monitoring of load change data of storage nodes to generate a load monitoring result; when the access request volume or load exceeds a preset range, a feedback message is generated

[0051] A feedback generation unit for comprehensively combining request monitoring information and load monitoring results to generate a feedback message, where the feedback message includes the identification of abnormal storage nodes and corresponding parameters.

[0052] Preferably, the scheme adjustment module includes:

[0053] An anomaly extraction unit for extracting the identification of abnormal storage nodes and corresponding parameters according to the feedback message;

[0054] An adjustment strategy generation unit for recalculating the parameter allocation weights of abnormal storage nodes according to the identification of abnormal storage nodes and corresponding parameters, and dynamically adjusting the parameter allocation scheme to generate a dynamic allocation scheme;

[0055] A second allocation verification unit for verifying the rationality of the dynamic allocation scheme, including load balancing, localized storage of dependent parameters, and matching of group priorities with storage node performance.

[0056] Preferably, the data migration module includes:

[0057] A migration path planning unit for determining the migration path of affected parameters according to the dynamic allocation scheme, where the migration path includes the data transfer order between the source storage node and the target storage node;

[0058] A data migration execution unit for extracting parameter data from the source storage node according to the migration path and transmitting it to the target storage node;

[0059] A migration status update unit for updating the migration completion flag after parameter migration is completed and recording the status information of the migration operation.

[0060] Preferably, the status synchronization module includes:

[0061] A synchronization trigger unit for synchronously updating the storage status of storage nodes after parameter migration is completed to obtain the parameter storage status;

[0062] A status update unit for distributing the parameter storage status information to all storage nodes;

[0063] A status consistency verification unit for verifying whether the storage nodes correctly receive and update the parameter storage status.

[0064] The above solution of the present invention has at least the following beneficial effects:

[0065] First, the system obtains the parameter distribution data in the model training task through the parameter analysis module, including access frequency and dependency relationship, so as to accurately analyze the access characteristics of each parameter. Compared with the extensive allocation method using the fixed sharding strategy in the prior art, through in-depth analysis of the access frequency, the system generates the parameter access frequency distribution result and the parameter priority score, can identify the hot parameters with high access frequency, and provides clear guidance for their subsequent optimization grouping and allocation. This fine-grained analysis avoids the phenomenon that hot parameters are concentrated in a small number of storage nodes in the traditional sharding method, and effectively reduces the I / O bottleneck problem of storage nodes.

[0066] Second, with the support of the parameter grouping module, the system groups the parameters according to the access frequency similarity and the dependency relationship strength. By aggregating the parameters with similar access patterns and strong dependencies into the same group, the communication overhead of cross-node access is reduced, and the localization access characteristics of the distributed storage system are improved. At the same time, the grouping quantization unit weights the priority scores of the parameters in the group to generate optimized parameter grouping data, which further provides a solid foundation for the scientific nature of data allocation.

[0067] In the allocation stage, the system dynamically calculates the allocation weight of the storage node through the allocation weight calculation module, considering the load status, hardware performance and grouping requirements of the storage node, to ensure that each parameter group can be allocated to the optimal storage node. Compared with the traditional solution that relies on the centralized metadata node to manage data distribution, the system realizes dynamic adjustment of the distribution, and solves the problem of uneven load of storage nodes in the fixed sharding method. In the parameter allocation module, the system calculates the adaptability score between the parameter and the storage node by using the grouping priority score and the storage node allocation weight, further optimizing the parameter allocation scheme and significantly improving the efficiency of data storage and access.

[0068] In addition, the dynamic monitoring module monitors the access requests and load status of the storage nodes in real time during the operation of the system. When it detects that the access frequency of the hot parameter exceeds the preset range or the load of the storage node is too high, the scheme adjustment module can quickly adjust the allocation strategy and realize the efficient reallocation of the parameter in combination with the data migration module. After the migration is completed, the status synchronization module ensures the data state consistency of all storage nodes through the consistency synchronization mechanism, guaranteeing the high availability and stability of the system. This dynamic adjustment ability significantly enhances the adaptability of the system to real-time training or online inference scenarios, enabling the large model optimization task to run efficiently in complex environments.

[0069] In summary, the system effectively solves the problems of uneven load and hotspot concentration existing in the existing distributed storage system in the scenario of large model optimization. Through the synergistic effect of parameter analysis, grouping, allocation, and dynamic adjustment, not only the utilization rate of storage resources is optimized, but also the system's management ability for large-scale parameter data is significantly improved, thus achieving a significant performance improvement in large model training and inference tasks. This optimization mechanism has important value for practical applications in the fields of cloud computing and big data. Especially in scenarios requiring high concurrency and low latency, it can provide more stable and efficient service support for users. Brief Description of the Drawings

[0070] Figure 1 is an architecture diagram of a distributed storage system based on large model optimization provided by an embodiment of the present invention. Detailed Embodiments

[0071] Hereinafter, exemplary embodiments of the present disclosure will be described in more detail with reference to the drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be completely conveyed to those skilled in the art.

[0072] As Figure 1 shown, an embodiment of the present invention proposes a distributed storage system based on large model optimization, and the system includes:

[0073] A parameter analysis module, configured to obtain parameter distribution data of a model training task, and generate a parameter access analysis result based on the parameter distribution data, where the parameter distribution data includes parameter identification information, an original record of access frequency, and an original record of dependency relationship;

[0074] A parameter grouping module, configured to calculate the similarity of parameter access frequency and the degree of dependency of parameter dependency relationship according to the parameter access analysis result, perform grouping processing on the parameter distribution data, and generate parameter grouping data;

[0075] An allocation weight calculation module, configured to calculate the allocation weights of each storage node according to the parameter grouping data, in combination with the load information and hardware performance indicators of the storage nodes, and generate an allocation weight set;

[0076] A parameter allocation module, configured to allocate the parameter grouping data to different storage nodes according to the allocation weight set, and generate a parameter allocation scheme;

[0077] A dynamic monitoring module, configured to monitor the access requests and load information of the storage nodes in real time during model training or inference, and generate feedback information when the access request volume or load exceeds a preset range;

[0078] A scheme adjustment module, which is used to dynamically adjust the parameter allocation scheme according to the feedback information to obtain a dynamic allocation scheme;

[0079] A data migration module, which is used to perform the migration operation of parameter data on the target storage node and the source storage node according to the dynamic allocation scheme and update the storage status of the parameters;

[0080] A status synchronization module, which is used to synchronize the updated storage status of the parameters to all storage nodes after the parameter migration is completed.

[0081] In the embodiment of the present invention, through the collaborative action of the parameter analysis module, the parameter grouping module, the allocation weight calculation module, the parameter allocation module, the dynamic monitoring module, the scheme adjustment module, the data migration module and the status synchronization module, the system comprehensively optimizes the parameter storage and management process of the large model in the distributed storage environment. In the parameter analysis stage, by accurately obtaining and cleaning the parameter distribution data in the model training task, the system realizes the in-depth analysis and mining of the parameter access frequency and dependency relationship, generates an access analysis result reflecting the parameter characteristics, and provides a clear basis for subsequent grouping optimization and allocation decision-making. The parameter grouping module clusters the model parameters into groups according to their characteristics by combining the access frequency similarity and the dependency relationship strength, significantly reducing the cross-group dependency overhead and improving the access consistency of the parameters within the group.

[0082] In the process of allocating storage resources, the allocation weight calculation module dynamically generates the allocation weight by comprehensively considering the load status, performance evaluation index and grouping requirements of the storage nodes, ensuring that the grouped parameters can be efficiently allocated to the most suitable storage nodes. The parameter allocation module further optimizes the adaptability between the parameters within the group and the storage nodes, and makes accurate allocation based on comprehensive factors such as the priority score and the allocation weight. At the same time, the system has a strong real-time monitoring and dynamic adjustment ability, and can continuously track the access requests and load status of the storage nodes through the dynamic monitoring module during the model training or inference process. Once an abnormal state beyond the preset range is found, the scheme adjustment module can quickly regenerate an optimized allocation scheme to avoid the problems of load overload or performance degradation of the storage nodes.

[0083] To ensure the reliability of the system operation, the data migration module realizes efficient data migration operations through path planning and migration execution in the scheme adjustment stage, and ensures the data state consistency of all storage nodes through the status synchronization module after the migration is completed. The setting of the whole system greatly reduces the waste of storage resources and the overhead of cross-node communication, improves the efficiency of model parameter storage and access, and at the same time ensures the dynamic adaptability and consistency of data allocation. This optimization mechanism significantly improves the overall performance of the large model in the distributed storage environment.

[0084] In a preferred embodiment of the present invention, the parameter analysis module includes:

[0085] A data acquisition unit, configured to receive parameter distribution data from a model training task, where the parameter distribution data includes parameter identification information, original access frequency records, and original dependency records;

[0086] A data cleaning unit, configured to clean the parameter distribution data, remove redundant records, correct abnormal data, and format it into a standardized data structure to generate cleaned parameter data;

[0087] A frequency analysis unit, configured to extract the access frequency information of each parameter from the cleaned parameter data, and generate a frequency distribution result in combination with the frequency change trend. The frequency distribution result includes the classification interval of the parameter access frequency and the access pattern characteristics;

[0088] A dependency relationship analysis unit, configured to extract the dependency relationship records between each parameter from the cleaned parameter data and construct a dependency relationship network;

[0089] A parameter priority calculation unit, configured to calculate the priority score of each parameter according to the frequency distribution result and the dependency relationship network to obtain the parameter priority score. The calculation formula of the parameter priority score is:

[0090] P i = α1×F i + β1×log(1 + ∑ j∈Dep(i) D ij ) + γ1×T i , where P i is the priority score of parameter i, F i is the access frequency of parameter i, and its normalized value range is [0, 1]. Dep(i) is the set of all other parameters that have a dependency relationship with parameter i. D ij is the dependency relationship strength between parameter i and parameter j. T i is the historical access frequency change trend of parameter i, specifically represented as the time function value of parameter i. α1, β1, and γ1 are weight coefficients; where

[0091] Comm ij is the number of times parameter i and parameter j are accessed simultaneously. Freq ij is the joint access frequency of parameter i and parameter j. C i and C j are the storage data amounts of parameter i and parameter j respectively;

[0092] Freq i,t is the access frequency of parameter i at time point t. T is the total number of time points. λ1 is a weight coefficient;

[0093] An analysis result generation unit is configured to combine the frequency distribution result, the dependency network, and the parameter priority score to form a parameter access analysis result.

[0094] In an embodiment of the present invention, through the multi-unit cooperation of the parameter analysis module, the system realizes an efficient transformation of the parameter distribution data in the model training task from disorder to structure. First, the data acquisition unit can accurately obtain the parameter identification information, the original access frequency records, and the original dependency records, providing complete inputs for subsequent analysis. Subsequently, the data cleaning unit performs redundancy removal and anomaly correction on the parameter data, formatting the complex and chaotic data into a unified standard, and generating cleaned parameter data suitable for analysis and processing.

[0095] In the frequency analysis unit, the system extracts and classifies the access frequency of each parameter, generating a frequency distribution result including the access frequency classification intervals and access pattern features, clearly revealing the access rules of each parameter. At the same time, the dependency analysis unit accurately depicts the association characteristics and strength relationships between parameters through the construction of the dependency network. The priority calculation unit further combines the frequency distribution result and the dependency network to generate a priority score for each parameter, providing a clear optimization basis for subsequent grouping and allocation operations. Finally, the analysis result generation unit forms a complete parameter access analysis result by integrating the frequency distribution result, the dependency network, and the priority score.

[0096] The implementation of this module significantly reduces the uncertainty of the distributed storage system's understanding of the parameter access characteristics, provides a basis for the precision of grouping and allocation, and significantly improves the data processing efficiency and analysis quality.

[0097] More specifically, the frequency analysis unit specifically includes:

[0098] Data input: Receive the cleaned parameter data from the data cleaning unit, which contains the access records of each parameter, and each record includes the time point and frequency information of the access.

[0099] Access frequency extraction: Through statistical calculation of the parameter access records, obtain the access frequency of each parameter at multiple time points, such as the number of accesses per hour or per day.

[0100] Frequency distribution result generation: Generate the classification intervals of the access frequency by classifying the access frequency data. For example, divide the access frequency into three categories: "high frequency", "medium frequency", and "low frequency".

[0101] Analyze the change trend of the parameter access frequency and generate parameter access pattern features, such as periodicity and peak periods.

[0102] The output result of the frequency analysis unit can help the system identify hot parameters with high access frequencies and provide a clear basis for subsequent grouping optimization.

[0103] More specifically, the dependency analysis unit specifically includes:

[0104] Dependency extraction: Extract dependency records between parameters from the cleaned parameter data, such as parameter pairs that are accessed simultaneously in the same training task.

[0105] Dependency strength calculation: Such as the calculation formula for the dependency strength D between parameter i and parameter j as described above. ij of the calculation formula.

[0106] Dependency relationship network generation: According to the dependency strength, construct a dependency relationship network reflecting the logical relationship between parameters, where nodes represent parameters and edge weights represent dependency strength.

[0107] The output of the dependency analysis unit enables the system to accurately evaluate the logical relationship between parameters, reduce the communication overhead of cross-node access, and provide guidance for grouping optimization.

[0108] In a preferred embodiment of the present invention, the parameter grouping module includes:

[0109] A frequency grouping unit, configured to calculate the access frequency similarity of each parameter according to the frequency distribution result and parameter priority score in the parameter access analysis result, and generate frequency grouping data;

[0110] A dependency grouping unit, configured to analyze the strength of the dependency relationship between parameters according to the dependency relationship network and parameter priority score in the parameter access analysis result, and generate dependency grouping data;

[0111] A grouping quantization unit, configured to comprehensively quantize the parameters in the frequency grouping data and dependency grouping data, and generate parameter grouping data according to the comprehensive quantization result; wherein, the comprehensive quantization includes:

[0112] Calculate the comprehensive weight value of each parameter in different groups, and the group corresponding to the highest weight value is regarded as the preferred belonging group of the parameter; the calculation formula for the comprehensive weight value is:

[0113] where, W i,k is the comprehensive weight value of parameter i in group k, F i,k is the access frequency similarity of parameter i in group k, D i,k is the dependency strength of parameter i in group k, L k is the data volume of group k, C k is the storage capacity of group k, Var(F i,k ) is F i,kThe variance, where α2, β2, γ2, and λ2 are weight coefficients, and F i,k is the access frequency similarity of parameter i in group k; where

[0114] G k is the set of parameters in group k, and P j is the priority score of parameter j;

[0115] CosSim(Vec i , Vec j ) is the cosine similarity of the access patterns of parameter i and parameter j, that is, the access frequency similarity of parameter i and parameter j. Vec i and Vec j are the access frequency vectors of parameter i and parameter j respectively, Freq j,t is the access frequency of parameter j at time point t;

[0116] When there are shared parameters between groups, the relevant groups are merged into one group.

[0117] In the embodiment of the present invention, the parameter grouping module realizes the optimization of parameter grouping based on access frequency similarity and dependency strength in a distributed storage system. First, the system calculates the access patterns of parameters through the frequency grouping unit, extracts the frequency similarity values based on the access frequency distribution results, and preliminarily divides the parameters with similar frequencies into the same group. Subsequently, the dependency grouping unit calculates the dependency strength between parameters based on the dependency relationship network and merges the parameters with higher dependency degrees into the same group.

[0118] With the support of the grouping quantization unit, the system further quantifies the frequency similarity and dependency strength, combines the group load value and the priority score, dynamically adjusts the parameter grouping result, and generates optimized parameter grouping data. The grouping integration unit solves the possible parameter conflict problem between groups, reasonably allocates the shared parameters, and ensures the high consistency of logical independence between groups and parameter correlation within groups.

[0119] Through the optimization of this module, the system significantly reduces the communication overhead of cross-group dependencies, improves the consistency and localization characteristics of data access within groups, and lays a foundation for the efficiency of the subsequent parameter allocation phase.

[0120] More specifically, the frequency grouping unit specifically includes:

[0121] Access frequency vector construction: Based on the output of the frequency analysis unit, an access frequency vector is constructed for each parameter.

[0122] Frequency similarity calculation: The cosine similarity formula is used to calculate the access frequency similarity between parameters. For example, the cosine similarity CosSim(Vec i , Vec j ) of the access patterns of parameter i and parameter j is calculated as follows.

[0123] Grouping logic: Based on the similarity values, a similarity matrix between parameters is constructed; a clustering algorithm, such as K-Means or distance-based grouping logic, is used to group parameters with similar access frequencies into the same group.

[0124] By focusing on the similarity of parameter access patterns, the frequency grouping unit reduces access conflicts between groups and improves the system access efficiency.

[0125] More specifically, the dependency grouping unit specifically includes:

[0126] Dependency relationship matrix construction: Based on the output of the dependency relationship analysis unit, a dependency relationship matrix is constructed for all parameters, where each element D ij of the matrix represents the strength of the dependency relationship between parameter i and parameter j.

[0127] Grouping logic: A clustering algorithm, such as hierarchical clustering, is used to group parameters with higher dependency strengths into the same group; the grouping results are optimized to reduce the number of edges with cross-group dependencies.

[0128] The design purpose of the dependency grouping unit is to reduce cross-group dependencies and optimize the localized access characteristics of storage resources.

[0129] In a preferred embodiment of the present invention, the allocation weight calculation module includes:

[0130] A load acquisition unit for acquiring real-time load information of the storage node, where the real-time load information includes the current storage capacity usage and access processing capabilities of the node;

[0131] A performance evaluation unit for calculating the hardware performance metrics of the storage node;

[0132] A weight calculation unit for calculating the allocation weights of each storage node based on the parameter grouping data, the load information of the storage node, and the node performance evaluation value, generating an allocation weight set, where the allocation weight reflects the preferred bearing capacity of each storage node for different parameter groups; the calculation formula for the allocation weight is:

[0133] where P k,j is the allocation weight of group k on storage node l, C l is the remaining capacity of storage node l, H l is the hardware performance metric of storage node l, R l is the available bandwidth of storage node l, Ll is the current load of storage node l, D k is the storage requirement of group k, T il is the transmission time allocated to storage node l for parameter i, and δ is the weight coefficient; where

[0134] BW l is the network bandwidth of storage node l, IOPS l is the number of input / output operations per second of storage node l, Lat l is the storage access latency of storage node l.

[0135] In an embodiment of the present invention, the allocation weight calculation module dynamically generates an optimal allocation weight in a distributed storage system through performance evaluation and load analysis of storage nodes. With the support of the load acquisition unit, the system can real-time obtain the storage capacity, access load, and available resource status of storage nodes, providing accurate input for weight calculation. The performance evaluation unit generates performance metrics reflecting the computing power and storage efficiency of nodes through comprehensive evaluation of the hardware performance of storage nodes.

[0136] The weight calculation unit uses load information, performance evaluation values, and group requirements, combined with real-time access patterns, to dynamically generate allocation weights, ensuring that storage nodes can maximize their resource utilization when allocating parameters. The dynamic weight calculation logic of the module effectively solves the problems of performance differences and uneven resource allocation among storage nodes, optimizing the overall performance of the distributed storage system.

[0137] In a preferred embodiment of the present invention, the parameter allocation module includes:

[0138] A group priority determination unit, which is used to calculate the priority score of a group according to the comprehensive weight value of parameters and the storage requirement size of each parameter group, to obtain the group priority score; the calculation formula for the group priority score is:

[0139] where, P k is the priority score of group k, Var(P i,k ) is the variance of P i,k , P i,k is the parameter priority score of parameter i in group k, L k is the storage requirement of group k, and σ and η are weight coefficients;

[0140] An adaptability calculation unit, which is used to calculate the adaptability score between a storage node and a parameter according to the group priority score, the parameter priority score, and the allocation weight; the calculation formula for the adaptability score is:

[0141] where, Si,l The adaptability score assigned to parameter i for storage node l The access latency angle between parameter i and storage node l, and ∈ is a constant to prevent division by zero;

[0142] A parameter allocation unit, which is used to allocate the corresponding storage nodes and parameters according to the order of the adaptability scores from high to low, and generate a parameter allocation scheme;

[0143] A first allocation verification unit, which is used to verify the rationality of the parameter allocation scheme, including load balancing, localized storage of dependent parameters, and matching of group priorities with storage node performance.

[0144] In the embodiment of the present invention, the parameter allocation module completes the precise allocation of parameters through three functional units: group priority score, adaptability calculation, and allocation execution. The group priority determination unit combines the group load value and the priority score to generate the priority score of the group, providing a basis for the allocation decision. The adaptability calculation unit calculates the adaptability score between each parameter and the storage node by comprehensively considering the group priority score, the storage node allocation weight, and the parameter allocation requirement.

[0145] Based on the adaptability score, the allocation execution unit allocates the parameters to the optimal storage nodes in the order of the scores from high to low, generating a complete parameter allocation scheme. The setting of the module ensures the scientificity and efficiency of the parameter allocation, and at the same time, through the dynamic adjustment ability, ensures the stable operation of the system in the environment of load changes and storage node performance differences.

[0146] More specifically, the parameter allocation unit specifically includes:

[0147] Sorting of adaptability scores: Sort the adaptability scores from high to low to ensure that the parameters with high scores are preferentially allocated to the most suitable storage nodes.

[0148] Mapping from parameter to node: Traverse the sorted list of adaptability scores, and select the storage node with the highest score for each parameter for allocation. Check whether the real-time load of the storage node meets the allocation requirement. If not, select the node with the second-highest score.

[0149] Load balancing adjustment: During the parameter allocation process, monitor the load conditions of each storage node in real time. If a certain node has an over-limit load due to too many allocation tasks, suspend the further allocation tasks of this node and instead select a sub-optimal node for the parameter.

[0150] After completing the allocation of all parameters, the parameter allocation unit will generate a complete parameter allocation scheme. This scheme includes the following contents:

[0151] The mapping relationship between parameters and storage nodes: Record the target storage node to which each parameter is allocated.

[0152] Node load status: Statistically analyze the storage load and computing load of each storage node under the current allocation scheme.

[0153] Cross-node dependency distribution: Analyze the quantity and distribution of cross-node dependencies in the parameter allocation scheme to ensure they are within a reasonable range.

[0154] More specifically, the first allocation verification unit specifically includes:

[0155] Load balancing verification: Check whether the load of all storage nodes exceeds the preset range; if it is found that the load of a storage node exceeds the threshold, mark the abnormal allocation.

[0156] Localization verification: Check whether parameters with strong dependency relationships are preferentially allocated to the same storage node; if the cross-node dependency overhead is too high, trigger a scheme adjustment.

[0157] Allocation result feedback: If the verification passes, submit the scheme; if it fails, return it to the adjustment module for optimization.

[0158] The first allocation verification module ensures the rationality of the allocation scheme and avoids performance degradation caused by imbalance or dependency conflicts.

[0159] In a preferred embodiment of the present invention, the dynamic monitoring module includes:

[0160] An access monitoring unit for real-time monitoring of access request information of storage nodes, including the number of accesses, request types, and target parameters;

[0161] A load monitoring unit for real-time monitoring of load change data of storage nodes to generate load monitoring results; when the access request volume or load exceeds the preset range, generate feedback information

[0162] A feedback generation unit for comprehensively combining request monitoring information and load monitoring results to generate feedback information, which includes the identification of abnormal storage nodes and corresponding parameters.

[0163] In the embodiment of the present invention, the dynamic monitoring module realizes the real-time perception of the operating state of the distributed storage system through continuous access monitoring and load tracking. During the access monitoring process, the module dynamically captures changes in parameter access characteristics by recording access requests of storage nodes, providing data support for subsequent adjustments. At the same time, the load monitoring function tracks the resource usage status of storage nodes in real time, including storage load, bandwidth utilization, etc., to ensure that the system can comprehensively control the operating conditions of each node.

[0164] When the detected access request volume or load exceeds the preset range, the feedback generation function can respond quickly, generate feedback information reflecting the abnormal state, and locate the abnormal storage node and its related parameters. This function significantly improves the system's response speed and accuracy to abnormal states, providing a reliable basis for subsequent adjustment and optimization decisions.

[0165] Through the setting of the dynamic monitoring module, the distributed storage system has greatly improved its adaptability to the dynamic environment, can quickly adjust the operation strategy when the data access mode changes, and effectively avoids the problem of performance degradation of storage nodes caused by overload or improper resource allocation.

[0166] In a preferred embodiment of the present invention, the scheme adjustment module includes:

[0167] An abnormal extraction unit for extracting the abnormal storage node identifier and the corresponding parameters according to the feedback information;

[0168] An adjustment strategy generation unit for recalculating the parameter allocation weight of the abnormal storage node according to the abnormal storage node identifier and the corresponding parameters, dynamically adjusting the parameter allocation scheme, and generating a dynamic allocation scheme;

[0169] A second allocation verification unit for verifying the rationality of the dynamic allocation scheme, including load balancing, localized storage of dependent parameters, and matching of group priorities and storage node performance.

[0170] In the embodiment of the present invention, the scheme adjustment module is an important component for the system to achieve dynamic optimization. Its core lies in flexibly adjusting the parameter allocation scheme according to real-time feedback information. After the dynamic monitoring module marks an abnormal storage node, the adjustment module identifies the influence range through feedback parsing, including information such as the current load of the abnormal storage node and the grouping status of related parameters. The adjustment strategy generation function re-evaluates the group priorities and the allocation weights of storage nodes based on these data, and generates a new parameter allocation scheme.

[0171] To ensure that the adjusted scheme can meet the global optimization goal, a verification logic is embedded in the module. By simulating the allocation calculation, the performance indicators of the new scheme are evaluated, including the degree of load balancing and the matching degree of storage node performance. The scheme that passes the verification will be quickly deployed to the execution module to achieve efficient reallocation of parameters.

[0172] The scheme adjustment module greatly enhances the system's dynamic optimization ability, enabling it to flexibly respond to sudden changes in access patterns and fluctuations in resource states, and ensuring the efficient and stable operation of the distributed storage system in a changing environment.

[0173] In a preferred embodiment of the present invention, the data migration module includes:

[0174] The migration path planning unit is used to determine the migration path of the affected parameters according to the dynamic allocation scheme. The migration path includes the data transfer order between the source storage node and the target storage node.

[0175] The data migration execution unit is used to extract parameter data from the source storage node according to the migration path and transfer it to the target storage node.

[0176] The migration status update unit is used to update the migration completion flag after the parameter migration is completed and record the status information of the migration operation.

[0177] In the embodiment of the present invention, the data migration module plays a key role in the parameter reallocation stage. Through efficient path planning and migration execution, it ensures that the parameters can be quickly and safely transferred from the source storage node to the target storage node. The path planning function selects the optimal migration path according to the current network topology structure and the status of the storage nodes in the system, avoiding unnecessary network overhead and transmission delay.

[0178] During the migration execution process, the module combines the priorities of the migration parameters and adopts batch transmission and parallel processing technologies to further improve the migration efficiency. At the same time, during the migration process, the module will record the operation status to ensure that it can quickly roll back to the original state in case of an exception, avoiding data loss or uneven distribution caused by migration failure.

[0179] The setting of this module significantly improves the speed and reliability of parameter migration, provides a solid guarantee for the dynamic adjustment of the distributed storage system, and performs well in reducing network overhead and maintaining data consistency.

[0180] More specifically, the migration path planning unit specifically includes:

[0181] Data input: Obtain the migration plan output by the scheme adjustment module, including the list of parameters to be migrated, the status information of the current storage node, such as load, remaining bandwidth, etc., and the storage performance and current load of the target node.

[0182] Path optimization: Calculate the transmission path of the parameters from the source storage node to the target storage node according to the network topology structure and the real-time status of the storage nodes; calculate the weight according to the bandwidth on the path and the hardware performance indicators of the target storage node, so as to quantify the quality of the path.

[0183] Generate migration path: Select the path with the lowest weight for each parameter to be migrated and generate a migration plan, including the transfer order and specific path information.

[0184] The path planning unit significantly reduces network resource consumption and migration time by optimizing the transmission path of parameter migration, and improves the migration efficiency of the system.

[0185] More specifically, the data migration execution unit specifically includes:

[0186] Migration scheduling: According to the output of the path planning unit, schedule parameter migration tasks and perform batch migration according to priorities.

[0187] Migration operation: Use multi-threading or parallel processing technology to migrate multiple parameters simultaneously to make full use of network bandwidth and storage node resources.

[0188] During the transmission process, monitor the migration status in real time and record the progress of the transmission and possible exceptions.

[0189] Error handling: If a transmission interruption or error occurs during the migration process, the migration execution unit will trigger a rollback mechanism to restore the parameters to the state before migration and reschedule the migration task.

[0190] The migration execution unit ensures the rapid completion of parameter migration through efficient task scheduling and parallel transmission technology, and guarantees the security and reliability of migration through the error handling mechanism.

[0191] More specifically, the migration status update unit specifically includes:

[0192] Status collection: During the parameter migration process, collect various status information of the migration in real time, including the migration progress of the parameters, the current transmission path, the amount of data completed, etc.

[0193] Status storage: Store the collected status information in a log file or a distributed database for subsequent analysis and query.

[0194] Exception recording: If an error or interruption occurs during the migration process, record the detailed information of the exception, including the error type, the occurrence time, and the parameters and storage nodes involved.

[0195] The migration status recording unit provides traceability of the migration process for the system through comprehensive migration status recording, and provides basic data for fault troubleshooting and optimization adjustment.

[0196] In a preferred embodiment of the present invention, the status synchronization module includes:

[0197] A synchronization trigger unit, used to synchronously update the storage status of the storage node after the parameter migration is completed to obtain the parameter storage status;

[0198] A status update unit, used to distribute the parameter storage status information to all storage nodes;

[0199] A status consistency verification unit, used to verify whether the storage node correctly receives and updates the parameter storage status.

[0200] In the embodiments of the present invention, after the parameter migration is completed, the status synchronization module ensures the consistency of the data status of all storage nodes. Through the status update function, the module synchronizes the latest parameter allocation status and storage node information to all nodes in the network. The synchronization operation adopts a multi-threaded broadcast mechanism to complete the global data update at the fastest speed.

[0201] In addition, the module sets up a strict consistency verification mechanism to check the synchronization results one by one, ensuring that the update information received by all nodes is exactly the same, and avoiding access anomalies or data redundancy problems caused by inconsistent status. Once data conflicts or omissions are found during the synchronization process, the module can automatically trigger a supplementary synchronization operation to restore the inconsistent storage nodes to the correct state.

[0202] The implementation of the status synchronization module significantly enhances the coordination of the distributed storage system, ensuring the accuracy and stability of data access in high-concurrency and dynamic adjustment scenarios.

[0203] More specifically, the synchronization trigger unit specifically includes:

[0204] Monitoring the migration completion event: The synchronization trigger unit continuously monitors the task status of the data migration module. Once the parameter migration task is completed, the migration status recording unit will feedback the migration result to the synchronization trigger unit, marking the storage nodes and parameters that need to perform status synchronization.

[0205] Abnormal status trigger: In addition to the migration completion event, the synchronization trigger unit also monitors abnormal events in the system that may cause data inconsistency, such as data loss due to high load of storage nodes, data not being updated in time due to network interruption, etc. Once an abnormality is detected, the synchronization trigger unit will immediately trigger a synchronization operation.

[0206] Trigger mechanism: According to the preset synchronization policy, the synchronization trigger unit can choose full synchronization or incremental synchronization.

[0207] Full synchronization: Suitable for global status update after large-scale parameter migration.

[0208] Incremental synchronization: Suitable for updating the status of a small range of storage nodes to reduce the system burden.

[0209] The synchronization trigger unit significantly improves the system's response speed through the event-driven mechanism, avoiding data inconsistency problems caused by delayed triggering. At the same time, flexibly selecting the full or incremental synchronization strategy optimizes the system's synchronization efficiency.

[0210] More specifically, the status update unit specifically includes:

[0211] Updated data generation: Obtain the migrated parameter distribution information from the data migration module or the exception event report, including the new storage location of the parameters, the associated groups, access priorities, and dependencies.

[0212] Format these data into standard synchronization messages for quick parsing and application by each storage node.

[0213] Distribution of updates: According to the topology of the storage nodes, adopt multiple synchronization strategies for data distribution:

[0214] Broadcast synchronization: Send the updated data to all storage nodes simultaneously, suitable for small-scale storage systems.

[0215] Hierarchical synchronization: First synchronize the core storage nodes, such as the metadata management storage nodes, and then the core storage nodes further synchronize to their associated storage nodes. This is suitable for large-scale distributed storage systems.

[0216] Ensure the integrity and reliability of the synchronization message transmission to prevent data loss due to network interruption or transmission failure.

[0217] Real-time update: After receiving the synchronization message, the storage node modifies the local parameter distribution information through the real-time update mechanism. For example, update the storage location of the parameters, grouping information, and the status of associated storage nodes.

[0218] The status update unit ensures that the updated data can reach all storage nodes quickly through standardized synchronization messages and an efficient distribution mechanism. The hierarchical synchronization strategy further optimizes the synchronization performance in large-scale systems and avoids the network overhead of global broadcasting.

[0219] More specifically, the status consistency verification unit specifically includes:

[0220] Status verification: Collect the synchronization feedback information of all storage nodes, including the status of parameter updates, grouping information, and dependencies. Use the consistency verification algorithm to compare the storage node status to confirm whether all storage nodes have the same data view.

[0221] Data integrity verification: Compare the parameter data integrity of each storage node before and after the update. For example, verify whether data corruption or loss occurs during migration or synchronization by calculating the hash value of the parameter data.

[0222] Conflict detection and handling: When it is found that the status of some storage nodes does not match the global status record, mark the conflict nodes. Trigger a supplementary synchronization operation to update the status of the conflict storage nodes to the correct status again to ensure global data consistency.

[0223] The status consistency verification unit effectively prevents system operation anomalies caused by data synchronization errors through a comprehensive verification and conflict handling mechanism. Its design ensures the high reliability and consistency of the distributed storage system, providing important support for the stable operation of the system in high-concurrency scenarios.

[0224] In a preferred embodiment of the present invention, the system further includes:

[0225] A data security module for encrypting data during parameter migration and status synchronization;

[0226] An exception recovery module for automatically rolling back to the previous parameter allocation scheme and re-executing relevant operations when data migration or status synchronization fails.

[0227] In the embodiment of the present invention, the distributed storage system realizes an efficient, stable, and secure data storage and access process through overall settings. With the coordinated action of parameter analysis, grouping, and allocation, the system can effectively reduce the communication overhead of cross-node access, significantly improve access efficiency and storage resource utilization. At the same time, the combination of dynamic monitoring and adjustment functions enables the system to have the ability of real-time response and dynamic optimization, and can quickly adapt when facing sudden loads or changes in the performance of storage nodes, ensuring the stable operation of the system.

[0228] In terms of security, the system is provided with a complete data encryption and exception recovery mechanism. During migration and synchronization, all transmitted data is encrypted to ensure the security of data during network transmission. The exception recovery module can quickly start the rollback function when migration fails or synchronization is abnormal, restoring the system state to a consistent state and avoiding serious consequences caused by data loss or incorrect allocation.

[0229] This system meets the high requirements of large-scale model training and inference tasks for distributed storage systems with efficient resource management, precise allocation strategies, and complete security mechanisms, demonstrating excellent technical advantages.

[0230] More specifically, the data security module specifically includes:

[0231] Encryption processing: During data migration and synchronization, the module encrypts the transmitted data to ensure the confidentiality of data during network transmission.

[0232] Encryption implementation: The module uses a symmetric encryption algorithm (such as AES) or an asymmetric encryption algorithm (such as RSA) to encrypt the data. After the encrypted data is received by the target storage node, it is restored to the original data by the corresponding decryption algorithm: The distribution of the key is carried out through a secure channel to ensure that the key will not be leaked during transmission.

[0233] Integrity Check: To prevent data from being tampered with during migration or synchronization, the module calculates the hash value of the data before data transmission and verifies the consistency of the hash value at the receiving end.

[0234] The receiving end recalculates the hash value of the received data. If the two hash values are the same, it is confirmed that the data has not been tampered with; otherwise, it is marked as abnormal.

[0235] Enhanced Transmission Security: The data security module also supports the TLS-based secure transmission protocol to further enhance the security during network transmission and prevent man-in-the-middle attacks.

[0236] Through data encryption, integrity check, and secure transmission protocol, the data security module can effectively ensure the confidentiality, integrity, and authenticity of data during migration and synchronization, avoiding system security problems caused by data leakage or tampering.

[0237] More specifically, the exception recovery module specifically includes:

[0238] Exception Detection: The module monitors the task status during migration and synchronization in real time, including issues such as data transmission failure, synchronization timeout, and node unavailability.

[0239] Detection Logic: If the migration task status is not completed within the preset time, it is marked as a timeout exception; if data verification fails in the synchronization task, it is marked as a verification exception.

[0240] Rollback Mechanism: After an exception occurs, the module can roll back to the state before data migration or synchronization to ensure that the system will not cause data loss or incorrect allocation due to exceptions.

[0241] Rollback Logic: Revoke the parameter status of the abnormal migration, restore the parameter storage of the source storage node; update the status information after rollback and notify all relevant storage nodes.

[0242] Re-execution Mechanism: After the rollback is completed, the module re-executes the failed migration or synchronization task by re-planning the path and tasks.

[0243] Re-planning: According to the current network topology and storage node status, select a new migration path or target node to avoid exceptions from occurring again.

[0244] Exception Logging: The module records all exception information in a log file, including the exception type, occurrence time, affected scope, and recovery operation results. The exception log can be used by system administrators to analyze and optimize future task scheduling and resource allocation strategies.

[0245] The exception recovery module ensures that the system can quickly respond to and handle unexpected problems during the migration and synchronization processes through a comprehensive mechanism of detection, rollback, and re-execution, avoiding long-term downtime or data errors caused by exceptions. Meanwhile, detailed exception logs provide data support for system optimization.

[0246] The above are the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A distributed storage system optimized based on a large model, characterized in that, The system includes: A parameter analysis module, which is used to obtain the parameter distribution data of the model training task and generate a parameter access analysis result based on the parameter distribution data. The parameter distribution data includes parameter identification information, original access frequency records, and original dependency relationship records; A parameter grouping module, which is used to calculate the similarity of parameter access frequencies and the degree of dependency of parameter dependency relationships according to the parameter access analysis result, perform grouping processing on the parameter distribution data, and generate parameter grouping data; An allocation weight calculation module, which is used to calculate the allocation weights of each storage node according to the parameter grouping data, combined with the load information and hardware performance indicators of the storage nodes, and generate an allocation weight set; A parameter allocation module, which is used to allocate the parameter grouping data to different storage nodes according to the allocation weight set and generate a parameter allocation scheme; A dynamic monitoring module, which is used to monitor the access requests and load information of the storage nodes in real time during the model training or inference process, and generate feedback information when the access request volume or load exceeds the preset range; A scheme adjustment module, which is used to dynamically adjust the parameter allocation scheme according to the feedback information to obtain a dynamic allocation scheme; A data migration module, which is used to perform a migration operation of parameter data on the target storage node and the source storage node according to the dynamic allocation scheme and update the storage status of the parameters; A status synchronization module, which is used to synchronize the updated storage status of the parameters to all storage nodes after the parameter migration is completed.

2. The distributed storage system optimized based on a large model according to claim 1, wherein The parameter analysis module includes: A data acquisition unit, which is used to receive the parameter distribution data from the model training task. The parameter distribution data includes parameter identification information, original access frequency records, and original dependency relationship records; A data cleaning unit, which is used to clean the parameter distribution data, remove redundant records, correct abnormal data, and format it into a standardized data structure to generate cleaned parameter data; A frequency analysis unit, which is used to extract the access frequency information of each parameter from the cleaned parameter data and generate a frequency distribution result in combination with the frequency change trend. The frequency distribution result includes the classification interval of the parameter access frequency and the access mode characteristics; A dependency relationship analysis unit, which is used to extract the dependency relationship records between each parameter from the cleaned parameter data and construct a dependency relationship network; A parameter priority calculation unit, which is used to calculate the priority score of each parameter according to the frequency distribution result and the dependency relationship network to obtain the parameter priority score. The calculation formula of the parameter priority score is: P i = α1 × F i + β1 × log(1 + ∑ j∈Dep(i) D ij ) + γ1 × T i , where P i is the priority score of parameter i, F i is the access frequency of parameter i, and its normalized value range is [0, 1]. Dep(i) is the set of all other parameters that have a dependency relationship with parameter i. D ij is the dependency strength between parameter i and parameter j, and T i is the historical access frequency change trend of parameter i, which is specifically represented as the time function value of parameter i. α1, β1, and γ1 are weight coefficients; among them, Comm ij is the number of times that parameter i and parameter j are accessed simultaneously, Freq ij is the joint access frequency of parameter i and parameter j, C i and C j are the stored data amounts of parameter i and parameter j respectively; Freq i,t is the access frequency of parameter i at time point t, T is the total number of time points, and λ1 is the weight coefficient; An analysis result generation unit, which is used to combine the frequency distribution result, the dependency relationship network, and the parameter priority score to form a parameter access analysis result.

3. A distributed storage system optimized based on a large model according to claim 2, wherein, The parameter grouping module includes: A frequency grouping unit, which is used to calculate the access frequency similarity of each parameter according to the frequency distribution result and the parameter priority score in the parameter access analysis result and generate frequency grouping data; A dependency grouping unit, which is used to analyze the strength of the dependency relationship between parameters according to the dependency relationship network and the parameter priority score in the parameter access analysis result and generate dependency grouping data; A grouped quantization unit is used to comprehensively quantize the parameters in the frequency grouped data and the dependent grouped data, and generate parameter grouped data according to the comprehensive quantization result; wherein, the comprehensive quantization includes: Calculating the comprehensive weight value of each parameter in different groups, and regarding the group corresponding to the highest weight value as the preferred belonging group of the parameter; the calculation formula of the comprehensive weight value is: Among them, W i,k is the comprehensive weight value of parameter i in group k, F i,k is the access frequency similarity of parameter i in group k, D i,k is the dependence strength of parameter i in group k, L k is the data volume of group k, C k is the storage capacity of group k, Var(F i,k ) is the variance of F i,k , α2, β2, γ2 and λ2 are weight coefficients, F i,k is the access frequency similarity of parameter i in group k; among them, G k is the set of parameters in group k, P j is the priority score of parameter j; CosSim(Vec i , Vec j ) is the cosine similarity of the access patterns of parameter i and parameter j, that is, the similarity of the access frequencies of parameter i and parameter j. Vec i and Vec j are the access frequency vectors of parameter i and parameter j respectively, Freq j,t is the access frequency of parameter j at time point t; When there are shared parameters between groups, merge the relevant groups into one group.

4. A distributed storage system optimized based on a large model according to claim 3, characterized in that, The allocation weight calculation module includes: A load acquisition unit is used to acquire the real-time load information of the storage node, and the real-time load information includes the current storage capacity usage and access processing ability of the node; A performance evaluation unit is used to calculate the hardware performance index of the storage node; A weight calculation unit is used to calculate the allocation weight of each storage node according to the parameter grouped data, the load information of the storage node, and the node performance evaluation value, generate an allocation weight set, and the allocation weight reflects the preferred bearing capacity of each storage node for different parameter groups; the calculation formula of the allocation weight is: Among them, P k,j is the allocation weight of group k on storage node l, C l is the remaining capacity of storage node l, H l is the hardware performance index of storage node l, R l is the available bandwidth of storage node l, L l is the current load of storage node l, D k is the storage requirement of group k, T il is the transmission time for parameter i to be allocated to storage node l, and δ is the weight coefficient; among them, BW l is the network bandwidth of storage node l, IOPS l is the number of input / output operations per second of storage node l, Lat l is the storage access latency of storage node l.

5. A distributed storage system optimized based on a large model according to claim 4, characterized in that, The parameter allocation module includes: A grouped priority determination unit is used to calculate the priority score of the group according to the comprehensive weight value of the parameters in each parameter group and the storage requirement size, and obtain the grouped priority score; the calculation formula of the grouped priority score is: Among them, P k is the priority score of group k, Var(P i,k ) is the variance of P i,k , P i,k is the parameter priority score of parameter i in group k, L k is the storage requirement of group k, and σ and η are weight coefficients; An adaptability calculation unit is used to calculate the adaptability score between the storage node and the parameter according to the grouped priority score, the parameter priority score, and the allocation weight; A parameter allocation unit is used to allocate the corresponding storage node and parameter according to the size order of the adaptability score, and generate a parameter allocation scheme; A first allocation verification unit is used to verify the rationality of the parameter allocation scheme, including load balancing, localized storage of dependent parameters, and matching of grouped priority and storage node performance.

6. A distributed storage system optimized based on a large model according to claim 5, characterized in that, The calculation formula of the adaptability score is: Among them, S i,l is the adaptation score assigned to the storage node l for the parameter i, is the access delay angle between the parameter i and the storage node l, and ∈ is a constant to prevent division by zero.

7. The distributed storage system optimized based on a large model according to claim 6, characterized in that, The dynamic monitoring module includes: An access monitoring unit is used to monitor the access request information of the storage node in real time, including the number of accesses, the request type, and the target parameter; A load monitoring unit is used to monitor the load change data of the storage node in real time, and generate a load monitoring result; when the access request volume or the load exceeds the preset range, generate feedback information A feedback generation unit is used to comprehensively generate feedback information based on the request monitoring information and the load monitoring result, and the feedback information includes the abnormal storage node identifier and the corresponding parameter.

8. A distributed storage system optimized based on a large model according to claim 7, characterized in that The scheme adjustment module includes: An anomaly extraction unit is used to extract the abnormal storage node identifier and the corresponding parameter according to the feedback information; An adjustment strategy generation unit is used to recalculate the parameter allocation weight of the abnormal storage node according to the abnormal storage node identifier and the corresponding parameter, and dynamically adjust the parameter allocation scheme to generate a dynamic allocation scheme; A second allocation verification unit is used to verify the rationality of the dynamic allocation scheme, including load balancing, localized storage of dependent parameters, and matching of grouped priority and storage node performance.

9. A distributed storage system optimized based on a large model according to claim 8, characterized in that, The data migration module includes: A migration path planning unit is used to determine the migration path of the affected parameters according to the dynamic allocation scheme, and the migration path includes the data transmission order between the source storage node and the target storage node; A data migration execution unit, configured to extract parameter data from a source storage node according to a migration path and transmit the parameter data to a target storage node; A migration status update unit, configured to update a migration completion flag after parameter migration is completed and record status information of the migration operation.

10. A distributed storage system optimized based on a large model according to claim 9, characterized in that, The status synchronization module includes: A synchronization trigger unit, configured to synchronously update a storage status of a storage node after parameter migration is completed to obtain a parameter storage status; A status update unit, configured to distribute parameter storage status information to all storage nodes; A status consistency verification unit, configured to verify whether the storage nodes correctly receive and update the parameter storage status.

Citation Information

Cited By

  • Data storage control method and device, storage medium and electronic equipment

    CN120704617A

  • Shell production parameter storage method and system based on Internet of Things

    CN121029762A

  • Data storage device supporting large model training and use method

    CN121326249A