Information processing method and system applied to file all-flash storage
By using the pre-trained storage strategy generation model in an all-flash storage system for dynamic weight fusion and generating a target storage policy parameter set, the problem of inefficient storage resource utilization in traditional storage technology is solved, and dynamic optimization of access paths is realized, improving the efficiency and flexibility of storage management.
Patent Information
- Application Number
- CN202510341217.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-03-21
AI Technical Summary
Traditional file storage technology lacks refined management of data in the storage capacity distribution when processing storage requests, resulting in excessive loading of some storage nodes while idle other nodes, and overall storage resource utilization efficiency is inefficient. At the same time, the existing technology adopts static path planning method in terms of access path optimization, and cannot dynamically adjust according to real-time system load and data access mode changes, resulting in increased data access delay and slower response speed.
By obtaining the storage request information received by the target storage node, multi-dimensional feature extraction and attribute encoding are performed, data block feature vectors and storage attribute feature vectors are generated, and inputting them to the pre-trained storage strategy generation model for dynamic weight fusion, and outputting the target storage strategy parameter set. Based on these parameters, storage node allocation instructions, path priority configuration and compression execution parameters are generated, and the all-flash storage system is driven to perform distributed storage operations.
It realizes the intelligent generation of storage strategy and the efficient execution of storage operations, automatically and accurately matches data characteristics and storage needs, avoids the cumbersome and errors of manual configuration strategies, and greatly improves the efficiency and accuracy of storage management. At the same time, the dynamic weight fusion mechanism makes storage strategies highly flexible and adaptable, and can cope with complex storage needs in different scenarios.
Smart Images

Figure CN120215831A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of storage technology, and more particularly, to an information processing method and system applied to file all-flash storage. Background Art
[0002] With the rapid development of information technology, the amount of data has shown an explosive growth, posing higher requirements for the performance, efficiency, and flexibility of storage systems. As a new type of storage technology, all-flash storage has gradually been widely used in various data storage scenarios due to its high-speed data reading and writing capabilities.
[0003] In traditional file storage technologies, the processing method for storage requests is relatively simple and single. Usually, storage resources are allocated only based on the size of the data volume, lacking refined management of the distribution of data in the storage capacity. For example, data blocks are stored in corresponding positions only according to the order or a fixed storage node allocation mode, without fully considering the possible differences in future access frequencies of different data blocks and the load balancing of different storage nodes. This results in some storage nodes being overloaded while other nodes are idle during actual use, making the overall utilization efficiency of storage resources low.
[0004] In terms of access path optimization, most existing technologies adopt static path planning methods. Once the storage location of the data is determined, the access path is basically fixed and cannot be dynamically adjusted according to real-time system load and changes in data access patterns. In the face of complex and changing business scenarios, such as a large number of users accessing product data simultaneously during an e-commerce promotion event or high-concurrency access caused by the update of popular TV series on a video website, problems such as increased data access latency and slower response speed are likely to occur, seriously affecting the user experience. Summary of the Invention
[0005] In view of the above-mentioned problems, in combination with the first aspect of the present invention, embodiments of the present invention provide an information processing method applied to file all-flash storage, the method comprising: Obtaining storage request information received by a target storage node, the storage request information including a set of data blocks to be processed and a corresponding set of expected storage attributes; wherein, the set of expected storage attributes is used to describe the configuration requirements of the set of data blocks to be processed in terms of storage capacity distribution, access path optimization, and data compression level; Performing multi-dimensional feature extraction on the set of data blocks to be processed to generate a data block feature vector, and performing attribute encoding on the set of expected storage attributes to generate a storage attribute feature vector; Input the data block feature vector and the storage attribute feature vector into a pre-trained storage policy generation model. Through the storage policy generation model, perform dynamic weight fusion on the data block feature vector and the storage attribute feature vector, and output a set of target storage policy parameters; Generate a storage node allocation instruction, a path priority configuration, and compression execution parameters according to the set of target storage policy parameters, and drive an all-flash storage system to perform a distributed storage operation on the set of data blocks to be processed based on the storage node allocation instruction; wherein, the set of target storage policy parameters satisfies the constraint conditions of all configuration requirements in the expected storage attribute set.
[0006] On the other hand, an embodiment of the present invention further provides an information processing system applied to file all-flash storage, including a processor and a machine-readable storage medium. The machine-readable storage medium is connected to the processor. The machine-readable storage medium is used to store programs, instructions, or codes, and the processor is used to execute the programs, instructions, or codes in the machine-readable storage medium to implement the above method.
[0007] Based on the above aspects, the embodiments of the present application achieve the intelligence of storage policy generation and the efficiency of storage operation execution. Specifically, first, obtain the storage request information received by the target storage node, which includes the set of data blocks to be processed and the corresponding set of expected storage attributes. The set of expected storage attributes clearly describes the configuration requirements of the set of data blocks to be processed in terms of storage capacity distribution, access path optimization, and data compression level. On this basis, perform multi-dimensional feature extraction on the set of data blocks to be processed to generate a data block feature vector, and at the same time perform attribute encoding on the set of expected storage attributes to generate a storage attribute feature vector. Input the data block feature vector and the storage attribute feature vector into a pre-trained storage policy generation model. The storage policy generation model analyzes the relationship between the two through a dynamic weight fusion mechanism and outputs a set of target storage policy parameters, which can flexibly adjust the influence degree of each feature in policy generation according to the characteristics of different data blocks and diverse storage requirements, so as to generate a more accurate and effective set of storage policy parameters. Then, generate a storage node allocation instruction, path priority configuration, and compression execution parameters according to the set of target storage policy parameters, and drive the all-flash storage system to perform distributed storage operations on the set of data blocks to be processed based on the storage node allocation instruction. Since the set of target storage policy parameters meets the constraint conditions of all configuration requirements in the set of expected storage attributes, the entire storage operation can be accurately carried out according to the user's expectations. Thus, through the intelligent storage policy generation model, it is possible to automatically and accurately match data characteristics with storage requirements, avoiding the cumbersome and error-prone manual configuration of policies, and greatly improving the efficiency and accuracy of storage management. At the same time, the dynamic weight fusion mechanism makes the storage policy highly flexible and adaptable, capable of coping with complex storage requirements in different scenarios. During the execution stage of distributed storage operations, based on the accurate set of storage policy parameters, it is possible to reasonably allocate storage nodes, optimize access paths, and set appropriate compression parameters, thereby giving full play to the performance advantages of the all-flash storage system, improving the efficiency of data storage and access, and reducing storage costs. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Figure 1 is a schematic flowchart of the execution process of the information processing method applied to file all-flash storage provided by an embodiment of the present invention.
[0009] Figure 2 is a schematic diagram of exemplary hardware and software components of the information processing system applied to file all-flash storage provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0010] The present invention will be specifically described below with reference to the accompanying drawings of the specification. Figure 1 is a flowchart of the information processing method applied to file all-flash storage provided by an embodiment of the present invention. The information processing method applied to file all-flash storage will be introduced in detail below.
[0011] Step S110: Obtain the storage request information received by the target storage node. The storage request information includes a set of data blocks to be processed and a corresponding set of expected storage attributes. Among them, the set of expected storage attributes is used to describe the configuration requirements of the set of data blocks to be processed in terms of storage capacity distribution, access path optimization, and data compression level.
[0012] In this embodiment, a data center of an Internet enterprise uses an all-flash storage system to store massive data, and the target storage node receives storage request information from different business departments within the Internet company.
[0013] For example, the video business department of the company has a large number of video materials that need to be stored. These video materials constitute the set of data blocks to be processed, including video files with different resolutions (such as high definition, standard definition), different durations (from short videos of several seconds to long videos of several hours), and different formats (such as MP4, AVI, etc.). Each video file is regarded as a data block. On this basis, in terms of the corresponding set of expected storage attributes, in terms of storage capacity distribution, due to the huge amount of video data, it is hoped that it can be evenly distributed among each storage node to avoid excessive storage pressure on a certain node. For example, it is expected to distribute the total video data to different storage nodes according to a certain ratio (such as evenly distributing it to three storage nodes, and each node stores approximately one-third of the data volume).
[0014] In terms of access path optimization, for popular videos (determined according to past playback volume statistics, such as recently popular short videos), it is expected to have a shorter access path to quickly respond to the playback requests of the client. For those video materials that are not frequently accessed, the access path can be relatively long, but it is also necessary to ensure that they can be obtained within a certain response time.
[0015] In terms of data compression level, for some high-resolution and high-quality videos, since they have already been compressed to a certain extent and need to maintain a high picture quality, it is expected to use a lower compression level, such as a compression ratio of 2:1; while for some low-resolution videos with low picture quality requirements, a higher compression level can be used, such as a compression ratio of 10:1, to save storage space.
[0016] For another example, the financial department of an enterprise needs to store financial statement data. These statement data, as a set of data blocks to be processed, include monthly, quarterly, and annual financial statements, and the format may be Excel spreadsheets or PDF files, etc. In terms of storage capacity distribution, due to the importance and relevance of financial statement data, it may be desirable to store the statements of the same year on adjacent storage nodes for easy management and query. For access path optimization, the personnel in the financial department may often need to query the statements of the recent few quarters, so the access paths to these statements need to be optimized to ensure fast access. In terms of the data compression level, the financial statement data needs to maintain accuracy, so a relatively low compression level may be adopted, such as a compression ratio of 3:1, to avoid data distortion.
[0017] Step S120: Extract multi-dimensional features from the set of data blocks to be processed to generate data block feature vectors, and perform attribute encoding on the set of expected storage attributes to generate storage attribute feature vectors.
[0018] Specifically, for the set of video data blocks of the video business department, when extracting multi-dimensional features to generate data block feature vectors. Analyze the content distribution pattern of each video data block. For example, the color distribution of the video frame (whether it is a movie frame with rich colors or a documentary frame with relatively single colors), the dynamic degree of the frame (action movie frames have more dynamics, while landscape movie frames are relatively static), etc. Statistic metadata features, such as the data block size distribution (the data block size of high-definition videos is significantly larger than that of standard-definition videos); hot and cold access tags (determine popular videos as hot data blocks according to the playback volume, and infrequently played videos as cold data blocks); associated file types, such as the file types associated with video subtitle files, cover image files, etc. Construct an initial feature set based on these metadata features, and then perform time series correlation analysis. For example, it is found that the access volume of specific types of videos (such as New Year's movies) will increase significantly during certain festivals, then the dynamic weight coefficients of the relevant feature dimensions of these videos will increase during this time period. Perform noise reduction processing on the initial feature set based on the dynamic weight coefficients. For example, if the weight of a metadata feature that has little to do with the video content is lower than a preset threshold, such as the association tag weight of a rarely used video format is low, it will be removed, thus generating an optimized data block feature vector.
[0019] For the expected storage attribute set, if the storage capacity distribution, which is a sub-attribute group, is a numerical indicator, it is assumed to be represented by the percentage of available capacity of the storage node. It is mapped to a preset dimensional feature space through piecewise normalization. For example, the available capacity of 0 - 20% is mapped to one eigenvalue, 20% - 40% is mapped to another eigenvalue, and so on. For the access path optimization, which is a logical indicator, if the access path is divided into three levels: fast, medium, and slow, a discrete feature sequence is generated through binary vectorization. For example, fast is 100, medium is 010, and slow is 001. For the data persistence level, which is a sub-attribute group, if it is divided into three levels: high, medium, and low, similar binary vectorization is also performed. Then, weighted concatenation is carried out according to the indicator priorities (for example, the storage capacity distribution has the highest priority, the data persistence level has the second highest priority, and the access path optimization has the lowest priority) to generate a global storage attribute feature vector.
[0020] For the set of financial statement data blocks in the finance department, analyze the content distribution pattern of each statement data block. For example, analyze the type distribution of data in the statement (whether it mainly contains income and expenditure data or asset and liability data, etc.). Statistic metadata features such as the data block size distribution (annual statement data blocks are larger, monthly statement data blocks are relatively smaller), hot and cold access marks (recent statements are hot data blocks, statements from many years ago are cold data blocks), and associated file types (such as the annotation files of the statements). After constructing the initial feature set, perform time series correlation analysis and find that the access frequency of relevant statements increases during the financial settlement periods at the end of each quarter and at the end of the year, and the dynamic weight coefficients of the corresponding feature dimensions increase. Perform noise reduction processing to eliminate some unimportant metadata features. For example, the weight of a specific format mark in a rarely viewed statement is relatively low, and it is eliminated to generate a data block feature vector.
[0021] For the numerical indicator of storage capacity distribution in the expected storage attribute set (measured by the number of bytes of the remaining space of the storage node), perform piecewise normalization mapping to the preset dimensional feature space. For the logical indicator of access path optimization (such as divided into three levels: emergency access, daily access, and rarely access), perform binary vectorization. For the data persistence level (such as divided into three levels: must be saved long-term, saved short-term, and can be deleted as needed), also perform binary vectorization, and then weighted concatenation is carried out according to the priority to generate a storage attribute feature vector.
[0022] In step S130, input the data block feature vector and the storage attribute feature vector into a pre-trained storage policy generation model. Through the storage policy generation model, perform dynamic weight fusion on the data block feature vector and the storage attribute feature vector, and output a set of target storage policy parameters.
[0023] In this embodiment, taking the video business department as an example, the feature vector of the video data block and the storage attribute feature vector are input into the pre-trained storage policy generation model. The feature interaction layer inside the model will perform interaction processing on the two vectors, such as identifying the relationship between features such as the size of the video data block, cold and hot access tags, and storage attributes such as storage capacity distribution and access path optimization. The policy prediction layer predicts the storage policy parameters based on this relationship. For example, for a popular video with a large data block, it may be allocated to a storage node with a large storage capacity and a fast access speed in terms of storage node allocation; in terms of path priority configuration, a higher priority is given to ensure fast access; in terms of compression execution parameters, they are set according to the determined lower compression ratio (such as 2:1). For a cold data video with a small data block, it may be allocated to a storage node with a relatively small storage capacity and a slightly slower access speed, with a lower path priority, and a higher compression ratio (such as 10:1) is used for the compression execution parameters. The constraint verification layer will verify whether the predicted policy parameters meet the configuration requirements of the expected storage attribute set. For example, verify whether the allocated storage node meets the requirements of the storage capacity distribution, whether the latency corresponding to the predicted path priority meets the latency threshold requirements in access path optimization, and whether the predicted compression execution parameters meet the requirements of the data compression ratio, etc. If not, the model will adjust the weights and re-perform dynamic weight fusion until the target storage policy parameter set that meets all configuration requirements is output.
[0024] For the financial statement data of the finance department, after inputting it into the model, the model identifies the relationship between the features of the financial statement data block (such as data block size, cold and hot access tags, etc.) and the storage attributes (such as the storage capacity distribution hopes to store the same-year statements together, and the access path needs to optimize the recent statements, etc.). The policy prediction layer predicts the storage policy parameters. For important financial statements in the near future (such as annual statements), nodes with high reliability and fast access speed will be selected in terms of storage node allocation, the path priority is set to the highest, and a lower compression ratio (such as 3:1) is used for the compression execution parameters. For less important statements from many years ago, the storage node allocation can be relatively flexible, the path priority is lower, and the compression ratio can be appropriately increased under the premise of ensuring data accuracy for the compression execution parameters. The constraint verification layer also verifies whether the predicted policy parameters meet the configuration requirements of the expected storage attribute set, such as whether the allocation of storage nodes meets the requirements of the storage capacity distribution, whether the access speed corresponding to the path priority meets the needs of the finance department, and whether the compression execution parameters ensure the accuracy of the data, etc., and outputs the target storage policy parameter set through continuous adjustment of weight fusion.
[0025] Step S140: Generate a storage node allocation instruction, a path priority configuration, and compression execution parameters according to the target storage policy parameter set, and drive the all-flash storage system to perform a distributed storage operation on the set of data blocks to be processed based on the storage node allocation instruction. Among them, the target storage policy parameter set satisfies the constraint conditions of all configuration requirements in the expected storage attribute set.
[0026] Continuing with the video business department as an example, generate a storage node allocation instruction according to the target storage policy parameter set. For popular and large-sized video data blocks, according to the storage node allocation parameters, the all-flash storage system distributes these data blocks to storage nodes with large capacity and high read / write speed. In terms of path priority configuration, set the access paths of these popular videos to high priority, and in the network topology of the all-flash storage system, ensure that the path from the client to the storage node is the shortest or passes through the fewest network devices. For the compression execution parameters, compress the video data blocks at a set low compression ratio (such as 2:1) and then store them. In this way, when performing the distributed storage operation, the all-flash storage system can store the video data in the expected manner, meeting both the requirements of storage capacity distribution (avoiding excessive storage pressure on a certain node) and access path optimization (fast access to popular videos) and data compression ratio (low compression ratio for videos with high picture quality requirements).
[0027] For the financial statement data of the finance department, generate a storage node allocation instruction according to the target storage policy parameter set, and allocate recent important statements (such as annual statements) to a dedicated storage node, which has the characteristics of high reliability and fast access. In terms of path priority configuration, set the access path of the recent statements to the highest priority to ensure that financial personnel can query them quickly. Compress the statement data at a low compression ratio (such as 3:1) and then store it according to the compression execution parameters. When the all-flash storage system performs the distributed storage operation, store the financial statement data according to this storage node allocation instruction, meeting all the requirements of storage capacity distribution (storing statements of the same year together), access path optimization (fast access to recent statements), and data compression ratio (ensuring data accuracy).
[0028] Based on the above steps, the embodiments of the present application achieve the intelligence of storage policy generation and the efficiency of storage operation execution. Specifically, first, obtain the storage request information received by the target storage node, which includes the set of data blocks to be processed and the corresponding set of expected storage attributes. The set of expected storage attributes clearly describes the configuration requirements of the set of data blocks to be processed in terms of storage capacity distribution, access path optimization, and data compression level. On this basis, perform multi-dimensional feature extraction on the set of data blocks to be processed to generate a data block feature vector, and at the same time perform attribute encoding on the set of expected storage attributes to generate a storage attribute feature vector. Input the data block feature vector and the storage attribute feature vector into a pre-trained storage policy generation model. The storage policy generation model analyzes the relationship between the two through a dynamic weight fusion mechanism and outputs a set of target storage policy parameters, which can flexibly adjust the influence degree of each feature in policy generation according to the characteristics of different data blocks and diverse storage requirements, so as to generate a more accurate and effective set of storage policy parameters. Then, generate a storage node allocation instruction, path priority configuration, and compression execution parameters according to the set of target storage policy parameters, and drive the all-flash storage system to perform distributed storage operations on the set of data blocks to be processed based on the storage node allocation instruction. Since the set of target storage policy parameters meets the constraint conditions of all configuration requirements in the set of expected storage attributes, the entire storage operation can be accurately carried out according to the user's expectations. Thus, through the intelligent storage policy generation model, it is possible to automatically and accurately match data characteristics with storage requirements, avoiding the cumbersome and error-prone manual configuration of policies, and greatly improving the efficiency and accuracy of storage management. At the same time, the dynamic weight fusion mechanism makes the storage policy highly flexible and adaptable, capable of coping with complex storage requirements in different scenarios. In the stage of executing distributed storage operations, based on the accurate set of storage policy parameters, it is possible to reasonably allocate storage nodes, optimize access paths, and set appropriate compression parameters, thereby giving full play to the performance advantages of the all-flash storage system, improving the efficiency of data storage and access, and reducing storage costs.
[0029] In a possible implementation manner, the set of expected storage attributes includes at least one sub-attribute group, and each sub-attribute group corresponds to a type of storage performance metric. Step S120 includes: Step S121, for each sub-attribute group, select the corresponding encoding rule according to the type of storage performance metric it corresponds to.
[0030] Step S122, if the sub-attribute group is a numerical metric, map it to a preset dimensional feature space through piecewise normalization. If the sub-attribute group is a logical metric, generate a discrete feature sequence through binary vectorization.
[0031] Step S123: Weightedly splice the encoded feature sequences of each sub-attribute group according to the index priority to generate a global storage attribute feature vector. Among them, the types of storage performance indicators include access latency threshold, data persistence level, and storage node load balancing coefficient.
[0032] Taking the video data storage scenario of the video business department as an example. For the expected storage attribute set, it contains multiple sub-attribute groups, and each sub-attribute group corresponds to a type of storage performance indicator.
[0033] Suppose there are three storage nodes in the all-flash storage system, labeled as Node A, Node B, and Node C respectively. The storage node load balancing coefficient is measured by the ratio of the currently used storage capacity of each node to the total storage capacity. Node A has used 30% of the capacity, Node B has used 25%, and Node C has used 20%. Assuming the total storage capacity is 1000GB (only for convenience of calculation and explanation here), then Node A has used 300GB, Node B has used 250GB, and Node C has used 200GB.
[0034] Segmented normalization rules can be set. For example, the situation of 0 - 30% used capacity is mapped to the eigenvalue 1, 30% - 60% is mapped to the eigenvalue 2, and 60% - 100% is mapped to the eigenvalue 3. The eigenvalue corresponding to the load balancing coefficient of Node A is 1, and the same is true for Node B and Node C.
[0035] For the data persistence level, assume it is divided into three levels: high, medium, and low. The high level means that the data must be stored for a long time and have redundant backups. The medium level means that it is stored within a certain period and has a small number of backups. The low level means that it can be deleted or transferred at any time according to the storage resource situation. If the popular video materials of the video business department require a high level of data persistence, perform binary vectorization on it. The high level is represented as 110 (the binary encoding here is custom-defined. 110 represents a coding method for the high level. The first bit may represent whether there is a redundant backup, the second bit represents the classification of the storage period length, and the third bit represents meanings such as whether it can be adjusted according to the resource situation).
[0036] For the access latency threshold, assume it is divided into three levels: fast, medium, and slow. Fast means that when the client requests video playback, the data should start to be transmitted within 1 second. Medium means within 3 seconds, and slow means within 5 seconds. For popular videos, the expected access latency threshold is fast. Perform binary vectorization on it as 100 (also custom-defined encoding. The first bit represents whether it meets the fast requirement within 1 second, and the second and third bits can represent other related latency attributes).
[0037] Among them, the index priorities can be set. Suppose the priority of the storage node load balancing coefficient is 3, the priority of the data persistence level is 2, and the priority of the access latency threshold is 1.
[0038] For the load balancing coefficient eigenvalue of 1, since the priority is 3, the calculated weighted value is 1×3 = 3. For the binary vectorization result 110 of the data persistence level, convert it to a decimal number (calculation process: 1×2² + 1×2¹ + 0×2 0 = 6), and multiply by the priority 2 to get 12. For the binary vectorization result 100 of the access latency threshold, convert it to a decimal number (calculation process: 1×2² + 0×2¹ + 0×2 0 = 4), and multiply by the priority 1 to get 4.
[0039] Then, concatenate these weighted results in order (for example, in the order of the storage node load balancing coefficient, data persistence level, and access latency threshold) to obtain the storage attribute feature vector. Suppose the concatenated result is 3 - 12 - 4 (here, "-" is just used to represent the concatenation order relationship).
[0040] For the financial statement data storage scenario of the finance department, assume that there are three storage nodes in the all-flash storage system for storing financial statement data. Node A stores 40% of the financial statement data (assuming the total data volume is 100GB and Node A stores 40GB), Node B stores 30% (30GB), and Node C stores 30% (30GB). Set the piecewise normalization rule: 0 - 35% is eigenvalue 1, and 35% - 70% is eigenvalue 2. The eigenvalue corresponding to the load balancing coefficient of Node A is 2, and for Node B and Node C it is 1.
[0041] For the data persistence level, assume it is divided into three levels: must be saved long-term, saved short-term, and can be deleted as needed. The annual report of the finance department requires long-term preservation, and its binary vectorization is 100 (custom encoding, the first bit represents whether it is long-term preservation, and the second and third bits represent other related attributes).
[0042] For the access latency threshold, assume it is divided into three levels: urgent access (response within 1 second), daily access (response within 3 seconds), and rarely accessed (response within 5 seconds). For the recent financial statements, the access latency threshold for urgent access is expected, and its binary vectorization is 100.
[0043] The priority of the storage node load balancing coefficient can be set to 3, the priority of the data persistence level to 2, and the priority of the access latency threshold to 1.
[0044] For the load balancing factor, the characteristic value of node A is 2 multiplied by the priority 3 to get 6, and the characteristic value of node B and node C is 1 multiplied by the priority 3 to get 3. For the data persistence level, the binary vectorization result 100 is converted to a decimal number (calculation process: 1×2²+0×2¹+0×2 0 = 4), multiplied by the priority 2 to get 8. The binary vectorization result 100 of the access latency threshold is converted to a decimal number (calculation process: 1×2²+0×2¹+0×2 0 = 4), multiplied by the priority 1 to get 4.
[0045] Therefore, the storage attribute feature vector is obtained by sequentially concatenating, assuming it is 6 - 8 - 4 (weighted value of the load balancing coefficient of node A - weighted value of the data persistence level - weighted value of the access delay threshold).
[0046] Step S124, analyzing the content distribution pattern of each data block in the set of data blocks to be processed, and counting metadata features of each data block, wherein the metadata features include data block size distribution, hot and cold access marks, and associated file types.
[0047] Step S125 , constructing an initial feature set according to the metadata features, and performing a time series correlation analysis on the initial feature set to determine a dynamic weight coefficient of each feature dimension within a historical storage period.
[0048] Step S126, performing noise reduction processing on the initial feature set based on the dynamic weight coefficient, removing feature dimensions whose weights are lower than a preset threshold, and generating an optimized data block feature vector.
[0049] Taking the video data of the video business department as an example, for the video data blocks, the content distribution pattern of each video is analyzed. For example, the color richness of the video screen is determined by analyzing the proportion of different colors in the screen. If the proportion of the main colors such as red, green, and blue in a video screen is relatively uniform, it means that the color richness is high; if a certain color accounts for a large proportion, the color richness is relatively low. At the same time, metadata features are counted, such as the distribution of data block size. The data block size of high-definition video may be between 1GB and 5GB, and the data block size of standard-definition video may be between 500MB and 1GB. The hot and cold access tags are determined according to the number of video playbacks. The hot data blocks are those with more than 1,000 playbacks per week, and the cold data blocks are those with less than 100 playbacks per week. Associate file types, such as subtitle files and cover image files of videos.
[0050] Then, construct an initial feature set based on the above metadata features. For example, the feature set may include <richness of color, data block size, hot / cold access flag, associated file type>. Then, perform temporal correlation analysis with a monthly time period. It is found that the playback volume of certain types of videos, such as action movies, will increase significantly during holidays. Therefore, during the holiday period, the dynamic weight coefficients of the feature dimensions of data blocks related to action movies (such as data block size, hot / cold access flag, etc.) will increase. Assume that during non-holiday periods, the weight coefficient of richness of color is 0.1, data block size is 0.3, hot / cold access flag is 0.4, and associated file type is 0.2; during the holiday period, the weight coefficient of richness of color becomes 0.05, data block size becomes 0.4, hot / cold access flag becomes 0.5, and associated file type becomes 0.05.
[0051] Among them, a preset threshold can be set to 0.1. During the holiday period, the weight coefficient of richness of color, 0.05, is lower than the preset threshold, and it is removed from the feature set. The optimized data block feature vector becomes <data block size, hot / cold access flag, associated file type>, and at the same time, the weight coefficients become data block size 0.44 (the original 0.4 divided by the sum of the weights of the remaining feature dimensions, 0.9), hot / cold access flag 0.56 (the original 0.5 divided by 0.9), and associated file type 0.0 (the original 0.05 has been removed due to being lower than the threshold, and here it is set to 0.0 for the sake of representing a complete vector form).
[0052] For another example, for the financial statement data of the finance department, the content distribution pattern of each statement data block can be parsed, such as the proportion of different types of data such as income and expenditure data, and asset and liability data in the statement. Statistic metadata features. In terms of data block size distribution, the annual statement data block may be between 50MB and 100MB, and the quarterly statement is between 10MB and 50MB. The hot / cold access flag is determined according to the query frequency of the statement. The statement queried frequently in the most recent quarter is a hot data block, and the statement queried infrequently many years ago is a cold data block. Associated file types such as the annotation file of the statement.
[0053] Then, an initial feature set can be constructed, such as <proportion of revenue and expenditure data, data block size, hot and cold access markers, associated file types>. Conducting time series correlation analysis with a quarterly time period, it is found that during the financial settlement periods at the end of each quarter and the end of the year, the query frequency of data blocks related to the balance sheet increases. Therefore, during this time period, the dynamic weight coefficients of the feature dimensions of data blocks related to the balance sheet (such as data block size, hot and cold access markers, etc.) will increase. Assume that during the non-settlement period, the weight coefficient of the proportion of revenue and expenditure data is 0.2, the data block size is 0.3, the hot and cold access marker is 0.4, and the associated file type is 0.1; during the settlement period, the weight coefficient of the proportion of revenue and expenditure data becomes 0.1, the data block size becomes 0.4, the hot and cold access marker becomes 0.5, and the associated file type becomes 0.0.
[0054] Among them, a preset threshold can be set to 0.1. During the settlement period, the weight coefficient of the associated file type, 0.0, is lower than the preset threshold, and it is removed from the feature set. The optimized data block feature vector becomes <proportion of revenue and expenditure data, data block size, hot and cold access markers>, and at the same time, the weight coefficients become 0.125 for the proportion of revenue and expenditure data (the original 0.1 divided by the sum of the weights of the remaining feature dimensions, 0.8), 0.5 for the data block size (the original 0.4 divided by 0.8), and 0.375 for the hot and cold access marker (the original 0.5 divided by 0.8).
[0055] In a possible implementation manner, the training method of the storage policy generation model includes: Step S210, obtaining a historical storage record data set, where the historical storage record data set includes multiple historical data block feature samples, corresponding historical storage attribute samples, and actual storage policy parameter labels.
[0056] Taking the video data storage of the video business department as an example, for the video business department, the historical storage record data set contains multiple historical data block feature samples, corresponding historical storage attribute samples, and actual storage policy parameter labels within a past period of time.
[0057] The historical data block feature samples include various features of previously stored video data blocks. For example, for video data blocks of different types (high-definition, standard-definition, different durations, different themes, etc.), information such as the data block size distribution (such as the frequency distribution of high-definition video data block sizes between 1GB and 5GB), hot and cold access markers (determining the proportion of hot data blocks and cold data blocks based on past play volumes), and associated file types (such as the association relationship between subtitle files, cover image files, etc. and the video).
[0058] The historical storage attribute samples cover the previous storage requirements for these video data. In terms of storage capacity distribution, the proportion of video data allocated to each storage node in different past time periods may be recorded; in terms of access path optimization, the access latency requirements corresponding to different popularity videos (hot data blocks, cold data blocks) are recorded; in terms of data compression level, the compression levels adopted for videos with different picture quality requirements are recorded.
[0059] The actual storage policy parameter tags are the actual storage policy parameters adopted for these video data at that time. For example, the actual storage node allocation (which videos are stored on which node), the actual path priority setting (the actual path arrangement to ensure fast access to popular videos), and the actual compression execution parameters adopted (the actual compression ratio for each video), etc.
[0060] Step S220, construct an initial neural network model, where the initial neural network model includes a feature interaction layer, a policy prediction layer, and a constraint verification layer. Among them, the constraint verification layer is used to verify whether the predicted policy parameters meet the configuration requirements of the historical storage attribute samples.
[0061] The role of the feature interaction layer is to perform interaction processing on the input historical data block features samples and historical storage attribute samples. For example, for the video data block size feature and the storage capacity distribution attribute, the feature interaction layer will analyze the relationship between the two and identify the storage mode of large data blocks under different storage capacity distribution requirements.
[0062] The policy prediction layer predicts the storage policy parameters based on the results of the feature interaction layer. For example, based on the hot and cold access tags of the video and the storage capacity distribution, it predicts which storage node a certain video data block should be allocated to, what the priority of its access path should be set to, and what compression parameters should be adopted.
[0063] The constraint verification layer is used to verify whether the policy parameters predicted by the policy prediction layer meet the configuration requirements of the historical storage attribute samples. For example, for video data, if the access latency threshold requirement for popular videos in the historical storage attribute samples is to start transmission within 1 second (access path optimization requirement), then the constraint verification layer will check whether the predicted latency value corresponding to the path priority parameter predicted by the policy prediction layer meets this upper limit constraint.
[0064] Step S230, input the historical data block feature samples and historical storage attribute samples into the initial neural network model, calculate the first loss value between the predicted storage policy parameters output by the policy prediction layer and the actual storage policy parameter tags, and calculate the second loss value between the constraint satisfaction degree output by the constraint verification layer and the preset standard value.
[0065] Step S240, adjusting the parameters of the initial neural network model according to the weighted sum of the first loss value and the second loss value until the weighted sum is lower than a convergence threshold, thereby obtaining a trained storage strategy generation model.
[0066] In a possible implementation, step S230 includes: Step S231, input the historical data block feature samples and the historical storage attribute samples into the strategy prediction layer of the initial neural network model to generate a set of prediction storage strategy parameters, wherein the set of prediction storage strategy parameters includes prediction node allocation parameters, prediction path priority parameters and prediction compression parameters.
[0067] For example, for a historical video data block, the predicted node allocation parameter may be to allocate it to storage node A, the predicted path priority parameter may be set to high priority (to meet the requirements of fast access to popular videos), and the predicted compression parameter may be to use a compression ratio of 2:1 (assuming that the video quality requirements are high).
[0068] Step S232, performing item-by-item difference calculation on the predicted storage strategy parameter set and the actual storage strategy parameter label to obtain a node allocation deviation value, a path priority deviation value, and a compression parameter deviation value.
[0069] Suppose the actual node allocation is storage node B, then the node allocation deviation value is the conversion cost of allocating node A to node B (which can be quantified according to factors such as the performance difference of storage nodes and data migration costs. For example, the difference between node A and node B in terms of read / write speed and storage capacity. Suppose the read / write speed of storage node A is 100MB / s, the read / write speed of node B is 80MB / s, and the data block size is 1GB. Then the time cost of migrating data from A to B is 1GB / (80MB / s) - 1GB / (100MB / s) = 12.5s - 10s = 2.5s, and this 2.5s can be part of the node allocation deviation value; at the same time, considering the storage capacity difference, if the remaining capacity of A is 500GB and the remaining capacity of B is 300GB, and the risk coefficient brought by the capacity difference is 0.2, then the comprehensive node allocation deviation value is 2.5s + 0.2 = 2.7s). Suppose the actual path priority is medium priority and the predicted path priority is high priority. The path priority deviation value can be calculated according to the delay difference corresponding to different priorities. High priority requires transmission to start within 1 second, medium priority requires transmission to start within 3 seconds, and the deviation value is 2 seconds. Suppose the actual compression ratio is 3:1 and the predicted compression ratio is 2:1. The compression parameter deviation value can be calculated according to the difference in the amount of compressed data. The original data volume is 1GB, compressed to 0.5GB according to 2:1, and compressed to 0.33GB according to 3:1. The deviation value is 0.5GB - 0.33GB = 0.17GB.
[0070] Step S233, according to the configuration requirements in the historical storage attribute samples, assign weight coefficients to the node allocation deviation value, the path priority deviation value, and the compression parameter deviation value, generate a weighted deviation vector, and perform a non-linear normalization process on the weighted deviation vector to generate the first loss value in scalar form.
[0071] Assume that the weight coefficient assigned to the storage node is 0.3 (since the correct allocation of storage nodes has a greater impact on the overall storage architecture), the weight coefficient of the path priority is 0.5 (access speed is crucial for the user experience), and the weight coefficient of the compression parameter is 0.2 (with relatively lower importance compared to the previous two). Then the weighted deviation vector is 〈2.7s × 0.3, 2s × 0.5, 0.17GB × 0.2〉 = 〈0.81s, 1s, 0.034GB〉. Perform a non - linear normalization on the weighted deviation vector. Assume that the sigmoid function is used for normalization (here, the calculation process of the sigmoid function is detailed. For each element x, sigmoid(x) = 1 / (1 + e^(-x)). For 0.81s, calculate sigmoid(0.81) = 1 / (1 + e^(-0.81)) ≈ 0.7. For 1s, sigmoid(1) = 1 / (1 + e^(-1)) ≈ 0.73. For 0.034GB, assume it is first converted into a dimensionless value, such as normalized according to the maximum data volume of the storage system. Assume the maximum data volume is 10GB, 0.034GB / 10GB = 0.0034, sigmoid(0.0034) ≈ 0.5. Sum up these normalized values after weighting, 0.7 × 0.3+0.73 × 0.5 +0.5 × 0.2 = 0.675. This 0.675 is the first loss value in scalar form.
[0072] Step S234, synchronously input the predicted storage policy parameter set and the historical storage attribute sample into the constraint verification layer, and parse the access delay threshold, data persistence level, and load - balancing coefficient in the historical storage attribute sample.
[0073] For example, for popular videos, the access delay threshold in the historical storage attribute sample is to start transmission within 1 second, the data persistence level is high (with redundant backups and long - term storage), and the load - balancing coefficient requires that the usage ratio difference of each storage node does not exceed 10% (assumed).
[0074] Step S235, verify whether the delay prediction value corresponding to the predicted path priority parameter meets the upper - bound constraint based on the access delay threshold, generate the path delay satisfaction degree, and verify whether the number of redundant copies corresponding to the predicted node allocation parameter meets the minimum redundancy constraint based on the data persistence level, generate the redundancy satisfaction degree, and verify whether the node load distribution deviation corresponding to the predicted node allocation parameter is lower than the equilibrium constraint threshold based on the load - balancing coefficient, generate the load - balancing satisfaction degree.
[0075] Suppose the delay prediction value corresponding to the predicted path priority parameter is 0.8 seconds, which meets the requirement within 1 second, and the path delay satisfaction degree is 1 (set to 1 if the requirement is met, and 0 if not). Based on the data persistence level, verify whether the number of redundant copies corresponding to the predicted node allocation parameter meets the minimum redundancy constraint. Suppose the high-level requirement is 3 redundant copies, and the number of redundant copies corresponding to the predicted node allocation parameter is 3, and the redundancy satisfaction degree is 1. Based on the load balancing coefficient, verify whether the node load distribution deviation corresponding to the predicted node allocation parameter is lower than the equilibrium constraint threshold. Suppose the predicted node allocation parameter distributes the video to node A and node B. The used capacity of node A is 30%, and the used capacity of node B is 32%. The difference is 2%, which is lower than the equilibrium constraint threshold of 10%, and the load balancing satisfaction degree is 1.
[0076] Step S236: Aggregate the path delay satisfaction degree, redundancy satisfaction degree, and load balancing satisfaction degree according to a preset ratio to generate an overall constraint satisfaction degree, and perform a logarithmic difference calculation between the overall constraint satisfaction degree and a preset standard value to generate the second loss value.
[0077] Suppose the preset ratio is that the weight of the path delay satisfaction degree is 0.3, the weight of the redundancy satisfaction degree is 0.4, and the weight of the load balancing satisfaction degree is 0.3. The overall constraint satisfaction degree is 1×0.3 + 1×0.4 + 1×0.3 = 1. Suppose the preset standard value is 1. Perform a logarithmic difference calculation between the overall constraint satisfaction degree and the preset standard value. Suppose the logarithmic function log(x) is used, and calculate log(1 / 1) = 0. This 0 is the second loss value.
[0078] Step S237: Input the first loss value and the second loss value into a weight allocator, generate a dynamic weight ratio according to the storage attribute priorities of each sample in the historical storage record dataset, and perform a weighted sum of the first loss value and the second loss value according to the dynamic weight ratio to generate a total loss value.
[0079] Step S238: Backpropagate the total loss value to the parameter update module of the initial neural network model, and adjust the trainable parameters of the policy prediction layer and the constraint verification layer until the total loss value is lower than the convergence threshold.
[0080] Perform a weighted sum of the first loss value 0.675 and the second loss value 0 according to a certain weight ratio. Generate a dynamic weight ratio according to the storage attribute priorities of each sample in the historical storage record dataset. Suppose that since access speed and storage node allocation are relatively important in video storage, the weight ratio of the first loss value is 0.6, and the weight ratio of the second loss value is 0.4. The weighted sum gives a total loss value of 0.675×0.6 + 0×0.4 = 0.405.
[0081] Backpropagate the total loss value of 0.405 to the parameter update module of the initial neural network model. In the policy prediction layer, for example, adjust the neuron connection weights related to node allocation prediction, path priority prediction, and compression parameter prediction. If a neuron is related to predicting the path priority and its weight has a significant impact on the path priority deviation value, then its weight will be adjusted according to the backpropagation algorithm (such as the gradient descent method. Here, the principle of the gradient descent method is detailed. Adjust the weight along the negative gradient direction of the total loss function, that is, the weight update amount = - learning rate × derivative of the loss function with respect to the weight. Assume the learning rate is 0.01. Calculate the derivative of the loss function with respect to this weight and then multiply by - 0.01 to get the weight update amount, thereby adjusting the weight). In the constraint verification layer, adjust the trainable parameters related to verifying the access delay threshold, data persistence level, and load balancing coefficient. Continuously repeat this process, input more historical data block feature samples and historical storage attribute samples, calculate the total loss value and backpropagate to adjust the parameters until the total loss value is lower than the convergence threshold (assume the convergence threshold is 0.01), and at this time, obtain the trained storage policy generation model.
[0082] Further, for the storage of financial statement data in the finance department, the process is similar.
[0083] The historical storage record data set includes past financial statement data block feature samples (such as statement data block size, hot and cold access tags, associated file types, etc.), historical storage attribute samples (requirements such as storage capacity distribution, access path optimization, data compression level, etc.), and actual storage policy parameter tags (actual storage node allocation, path priority setting, compression execution parameters, etc.).
[0084] The constructed initial neural network model also has a feature interaction layer, a policy prediction layer, and a constraint verification layer. When calculating the loss value, for calculating the first loss value, for example, the deviation between the actual storage node allocation and the prediction, the node allocation deviation value may be quantified according to the differences in the security and reliability of the storage nodes and the data migration cost; the path priority deviation value is calculated according to the differences in the query response times corresponding to different priorities; the compression parameter deviation value is calculated according to the differences in the accuracy and data volume of the compressed data, and then weighted and normalized to obtain the first loss value. For the second loss value, according to the access delay threshold (such as the response time requirement for querying recent statements), data persistence level (such as the long-term storage requirement for annual statements), and load balancing coefficient (the requirement for the balanced usage ratio of storage nodes) in the historical storage attribute samples, verify the predicted policy parameters, calculate the satisfaction degree and obtain the second loss value. Finally, sum the total loss value after weighting and backpropagate to adjust the model parameters until convergence.
[0085] In a possible implementation manner, after step S140, the method further includes: Step S150, collect the actual storage performance metrics of the to-be-processed data block set in real time. The actual storage performance metrics include node read / write rate, compression efficiency, and path load balance degree.
[0086] Taking the video data storage of the video business department as an example, for each storage node storing video data, a dedicated monitoring tool is used to measure its read / write rate. For example, on storage node A that stores popular videos, when reading a 1GB popular video data block, record the time from issuing the read instruction to starting data transmission. Suppose this time is 0.5 seconds, then the read rate is 1GB / 0.5 seconds = 2GB / second. The measurement of the write rate is similar. When a new video data block (such as a 2GB new video) is to be written to storage node A, record the time from the start of writing to the completion of writing. Suppose it is 1 second, and the write rate is 2GB / 1 second = 2GB / second.
[0087] Calculate the actual compression efficiency. For example, for a video data block with an original size of 3GB, after being compressed by the storage system, the actual storage size is 1GB. The compression efficiency can be obtained by calculating the ratio of the data amounts before and after compression. The compression ratio is the original size / the compressed size = 3GB / 1GB = 3:1.
[0088] Monitor the load conditions of video data on different access paths. Suppose a full-flash storage system has three main access paths, namely Path 1, Path 2, and Path 3. Statistically analyze the traffic of accessing video data through each path within a certain period of time (such as 1 hour). If the traffic through Path 1 is 500GB, the traffic through Path 2 is 300GB, and the traffic through Path 3 is 200GB, the calculation idea of variance can be used to calculate the path load balance degree. First, calculate the average traffic as (500GB + 300GB + 200GB) / 3 = 333.33GB. Then calculate the variance. For Path 1, the deviation is 500GB - 333.33GB = 166.67GB; for Path 2, the deviation is 300GB - 333.33GB = -33.33GB; for Path 3, the deviation is 200GB - 333.33GB = -133.33GB. The variance is [(166.67GB)² + (-33.33GB)² + (-133.33GB)²] / 3 ≈ 11111.11GB² (here is just to illustrate the calculation process, and in practice, a more appropriate load balance degree calculation method may be adopted according to the system characteristics). The smaller the value, the more balanced the load.
[0089] Step S160, compare the actual storage performance metrics with the configuration requirements in the expected storage property set to generate a performance deviation report.
[0090] In the set of expected storage attributes, for the read and write rates of popular video storage nodes, the requirements may be that the read rate is not less than 3 GB / second and the write rate is not less than 2.5 GB / second. However, the actual read rate of storage node A is 2 GB / second and the write rate is 2 GB / second, both of which are lower than the expected read rate, resulting in a deviation. For the compression efficiency, the expected compression ratio may be 2:1, while the actual ratio is 3:1, also showing a deviation. For the path load balance degree, the expected variance does not exceed 5000 GB², while the actual value is 11111.11 GB², also having a deviation. Organize these deviation situations into a performance deviation report, which details the actual values, expected requirements, and deviation situations of each indicator.
[0091] Step S170, if there are indicator items in the performance deviation report that exceed the tolerance threshold, trigger the online update mechanism of the storage policy generation model, and input the actual storage performance indicators and the corresponding storage request information as incremental training data into the model to readjust the model parameters.
[0092] Suppose the set tolerance thresholds are: the read rate deviation does not exceed 0.5 GB / second, the write rate deviation does not exceed 0.3 GB / second, the compression efficiency deviation does not exceed 0.5:1, and the path load balance degree variance deviation does not exceed 3000 GB². Since the deviations of the actual read rate, write rate, compression efficiency, and path load balance degree all exceed the tolerance thresholds, the online update mechanism of the storage policy generation model is triggered. Input the actual storage performance indicators (node read and write rates, compression efficiency, path load balance degree) and the corresponding storage request information (such as the type, size, hot and cold access marks of video data blocks) as incremental training data into the model.
[0093] In a possible implementation manner, the execution steps of the online update mechanism include: Step S171, extract the incremental storage operation records within a preset time window from the log database of the all-flash storage system, generate a set of incremental storage operation records, and map the storage request information in the set of incremental storage operation records to the actual storage performance indicators to generate a set of incremental training data samples.
[0094] For example, extract the incremental storage operation records within a preset time window (e.g., the past 24 hours) from the log database of the all-flash storage system. These incremental storage operation records contain the storage operation conditions of video data within these 24 hours, such as which new video data blocks are stored, which video data blocks are read or modified, etc. Map the storage request information (such as the feature information of video data blocks) in these incremental storage operation records to the actual storage performance metrics (the node read / write rate, compression efficiency, path load balancing degree, etc. collected previously). For example, map the video data block size field in the storage request information to the compression efficiency because the data block size affects the compression efficiency; map the hot / cold access flag of the video to the node read / write rate because the read / write frequency of popular videos is high and will affect the node read / write rate. Through this mapping relationship, generate an incremental training data sample set, and each sample contains a set of relevant storage request information and actual storage performance metrics.
[0095] Step S172: Freeze the network parameters of the feature interaction layer in the storage policy generation model to generate an initial fine-tuning model in the frozen parameter state.
[0096] The feature interaction layer has learned some basic relationships between the video data block features and storage attributes in the previous training. To avoid destroying these existing relationships during the fine-tuning process, its parameters are frozen. In this way, an initial fine-tuning model in the frozen parameter state is generated, and this initial fine-tuning model mainly adjusts the parameters in the policy prediction layer.
[0097] Step S173: Generate a sorted list of the severity of metric items according to the deviation magnitude of each metric item in the performance deviation report exceeding the tolerance threshold, and assign a gradually decreasing learning rate weight to the network parameters of the policy prediction layer based on the sorted list of the severity of metric items to generate dynamic learning rate configuration parameters.
[0098] For example, the read rate deviation is 1 GB / s (3 GB / s - 2 GB / s), the write rate deviation is 0.5 GB / s (2.5 GB / s - 2 GB / s), the compression efficiency deviation is 1:1 (2:1 - 3:1), and the path load balance variance deviation is 6111.11 GB² (11111.11 GB² - 5000 GB²). Sorting by the magnitude of the deviation, assume the order is path load balance variance deviation, read rate deviation, compression efficiency deviation, and write rate deviation. Based on this sorted list of the severity of the metric items, the learning rate weights for the network parameter allocation of the policy prediction layer decrease layer by layer. Assume the policy prediction layer has a three-layer network structure. For the topmost layer network, since the path load balance variance deviation is the most severe, the learning rate weight is assigned as 0.01; for the middle layer network, because the read rate deviation is the second most severe, the learning rate weight is assigned as 0.005; for the bottommost layer network, since the write rate deviation is relatively small, the learning rate weight is assigned as 0.001. In this way, the dynamic learning rate configuration parameters are generated.
[0099] Step S174: Input the incremental training data sample set into the initial fine-tuning model, perform backpropagation training according to the dynamic learning rate configuration parameters, generate the updated network parameters of the policy prediction layer, and splice the updated network parameters of the policy prediction layer with the network parameters of the feature interaction layer in the frozen parameter state to generate a candidate updated model parameter set.
[0100] During the training process, based on the input of each sample (storage request information and actual storage performance metrics), the model predicts the policy parameters, and then calculates the error between the predicted parameters and the parameters that should actually be achieved (according to the expected storage attribute set). According to the error and the learning rate weight, the network parameters of the policy prediction layer are adjusted. For example, for a video data block in a sample, if the predicted storage node allocation is unreasonable and the read and write rates do not meet the requirements, then according to the backpropagation algorithm, the network parameters related to the storage node allocation are adjusted along the direction of reducing the error. After multiple rounds of training, the updated network parameters of the policy prediction layer are generated.
[0101] Step S175: Load the candidate updated model parameter set into the verification environment of the storage policy generation model, input the verification set storage request information, generate a verification set predicted policy parameter set, calculate the matching degree between the simulated storage performance metrics corresponding to the verification set predicted policy parameter set and the expected storage attribute set in the verification set storage request information, and generate a verification set predicted accuracy improvement value.
[0102] The verification set storage request information is a part selected from the historical storage request information (for example, 10% of the historical video storage request information is selected as the verification set). According to the candidate updated model parameter set, the model generates a verification set prediction strategy parameter set. For example, for a video data block storage request in the verification set, the model predicts strategy parameters such as storage node allocation, path priority setting, and compression execution parameters. Then, calculate the simulated storage performance metrics corresponding to these predicted strategy parameter sets, such as the predicted node read / write rate, compression efficiency, and path load balance degree. Calculate the matching degree between these simulated storage performance metrics and the expected storage attribute set in the verification set storage request information. For example, calculate the matching degree between the simulated node read / write rate and the expected node read / write rate. If the simulated value meets the expected requirement, the matching degree is 1; otherwise, calculate the matching degree according to the deviation size (the larger the deviation, the lower the matching degree). Average the matching degrees of all verification samples to obtain the verification set prediction accuracy improvement value.
[0103] Step S176, when the verification set prediction accuracy improvement value exceeds the set percentage, mark the candidate updated model parameter set as the effective updated parameter set, encapsulate the effective updated parameter set into a model parameter update instruction, and generate a synchronization data packet carrying a timestamp and a version identifier.
[0104] Step S177, traverse all target storage nodes of the all-flash storage system, detect the difference between the model parameter version of the storage policy currently running on each target storage node and the version identifier in the synchronization data packet. If there is a target storage node whose running parameter version is earlier than the version identifier in the synchronization data packet, send a forced synchronization signal to this target storage node to trigger the target storage node to interrupt the current storage operation and load the effective updated parameter set.
[0105] When the verification set prediction accuracy improvement value exceeds the set percentage (assumed to be 10%), mark the candidate updated model parameter set as the effective updated parameter set. Encapsulate the effective updated parameter set into a model parameter update instruction, and generate a synchronization data packet carrying a timestamp (recording the update time) and a version identifier (indicating that this is a new model version). Traverse all target storage nodes of the all-flash storage system, and detect the difference between the model parameter version of the storage policy currently running on each target storage node and the version identifier in the synchronization data packet. If there is a target storage node whose running parameter version is earlier than the version identifier in the synchronization data packet, send a forced synchronization signal to this target storage node. After receiving the signal, the target storage node interrupts the current storage operation (such as the ongoing video data writing or reading operation) and loads the effective updated parameter set.
[0106] Step S178, after the parameter synchronization of all target storage nodes is completed, clear the expired data records in the incremental training data sample set and release the storage space of the log database.
[0107] For example, delete the data related to the incremental storage operation records 24 hours ago from the incremental training data sample set and release the storage space of the log database to provide space for subsequent incremental training data storage.
[0108] For the financial statement data storage scenario of the finance department, the process is similar. When collecting the actual storage performance metrics, the node read and write rates are measured according to the read and write operations on the storage nodes for storing financial statement data; the compression efficiency is calculated based on the data volume before and after the compression of the financial statement data; the path load balance degree is calculated according to the traffic conditions of different paths for accessing the financial statement data. When generating the performance deviation report, these actual metrics are compared with the requirements in the expected storage attribute set (such as the read and write rate requirements corresponding to the fast access to recent statements, the compression efficiency corresponding to the data accuracy requirements, etc.). If there is an indicator item exceeding the tolerance threshold, trigger the online update mechanism and perform operations according to similar steps, such as generating an incremental training data sample set (mapping the financial statement storage request information with the actual storage performance metrics), freezing the network parameters of the feature interaction layer, generating the dynamic learning rate configuration parameters, updating the network parameters of the policy prediction layer, verifying the candidate update model parameter set, updating the model parameters to the target storage nodes, and cleaning the incremental training data sample set, etc.
[0109] In a possible implementation manner, when the storage policy generation model is deployed on the edge computing node of the all-flash storage system, the method further includes: Step S310, locally cache the frequently used data block feature templates and storage attribute combination patterns on the edge computing node.
[0110] In this embodiment, taking the video data storage scenario of the video service department as an example, at the edge computing node of the all-flash storage system in the video service department, local caching of the feature templates of frequently used data blocks and the storage attribute combination patterns is started. For video data, the data block feature templates may include feature information such as the resolution of the video (such as high definition, standard definition, etc.), the duration range of the video (such as shorter than 5 minutes, 5 - 30 minutes, longer than 30 minutes, etc.), the format of the video (such as MP4, AVI, etc.), and the corresponding storage attribute combination patterns, such as storage capacity distribution (such as allocated to a specific storage node or storage node group), access path optimization (such as setting a specific access priority), and data compression level (such as using a specific compression ratio), etc. For example, for a popular video with high definition, a duration of 5 - 30 minutes, and a format of MP4 that is frequently accessed, the corresponding storage attribute combination pattern may be to allocate it to storage node A with fast read and write speeds, set a high access priority, and use a lower compression ratio (such as 2:1). The edge computing node will cache this video data block feature template and the storage attribute combination pattern.
[0111] Step S320, when receiving new storage request information, first perform a similarity match between the data block feature vector and the data block feature templates in the cache. If the similarity match result indicates that the match degree is higher than the fast response threshold, directly call the pre-generated policy parameters in the cache for the storage operation, bypassing the calculation process of the storage policy generation model.
[0112] When the edge computing node receives new storage request information, for example, there is new video data to be stored. First, perform a similarity match between the data block feature vector of the new video data and the data block feature templates in the cache. Suppose the new video is a high-definition video with a duration of 10 minutes and a format of MP4, and compare it with the above-mentioned high-definition, 5 - 30 minutes, MP4 video data block feature template in the cache. When calculating the similarity, for the resolution feature, if it is exactly the same, a certain score (such as 0.3) is added to the similarity; for the duration range, since the new video is within the duration range of the cache template, another certain score (such as 0.3) is added to the similarity; for the same format, another certain score (such as 0.3) is added to the similarity, and the comprehensive similarity is 0.9. Suppose the fast response threshold is set to 0.8. Since 0.9 is higher than 0.8, it indicates a high match degree. At this time, directly call the pre-generated policy parameters in the cache for the storage operation, that is, store the new video according to the previously cached storage attribute combination pattern, that is, allocate it to storage node A, set a high access priority, and use a compression ratio of 2:1, thereby bypassing the calculation process of the storage policy generation model and greatly improving the efficiency of the storage operation.
[0113] In a possible implementation manner, the method for updating the frequently used data block feature templates includes: Step S410, monitor the policy invocation records of the edge computing nodes, and obtain the invocation frequency of each data block feature template and the policy effectiveness generated in the corresponding storage operation, where the policy effectiveness is calculated by the matching ratio between the actual performance metrics after the storage operation and the set of expected storage attributes.
[0114] For example, for the above-mentioned video data block feature template of high definition, 5 - 30 minutes, and MP4 format, count the number of times it is invoked within a period of time (such as within a day), assuming it is 50 times, which is its invocation frequency. At the same time, calculate the policy effectiveness generated in the corresponding storage operation. The policy effectiveness is calculated by the matching ratio between the actual performance metrics after the storage operation and the set of expected storage attributes. For example, after storing this type of video, the actual storage performance metrics are the read and write rates of storage node A, the actual compression efficiency, and the actual access path load balancing degree, etc. The set of expected storage attributes requires that the read rate of storage node A is not less than 3 GB / s, the write rate is not less than 2.5 GB / s, the compression ratio is 2:1, and the variance of the path load balancing degree does not exceed 5000 GB². The actual measurement shows that the read rate of storage node A is 3 GB / s, the write rate is 2.5 GB / s, the compression ratio is 2:1, and the variance of the path load balancing degree is 4000 GB². Each index meets the requirements, so the policy effectiveness is 1 (because all indexes match, and the matching ratio is 100%).
[0115] Step S420, generate dynamic screening conditions according to the invocation frequency and the policy effectiveness, and screen out the data block feature templates with an invocation frequency higher than the active threshold and a policy effectiveness higher than the effective threshold to form a candidate template set.
[0116] Suppose the active threshold is set to 30 times / day, and the effective threshold is set to 0.8. For the video data block feature template of high definition, 5 - 30 minutes, and MP4 format, its invocation frequency is 50 times / day, which is higher than 30 times / day, and the policy effectiveness is 1, which is higher than 0.8, meeting the conditions. Screen all the data block feature templates, and select those data block feature templates with an invocation frequency higher than the active threshold and a policy effectiveness higher than the effective threshold to form a candidate template set. For example, in addition to the above video data block feature template, there is also a video data block feature template of standard definition, shorter than 5 minutes, and AVI format, with an invocation frequency of 40 times / day and a policy effectiveness of 0.9, which also meets the conditions and is selected into the candidate template set.
[0117] Step S430, perform clustering analysis on the data block feature templates in the candidate template set, calculate the similarity between any two data block feature templates, and assign a weight coefficient to each similarity according to the policy effectiveness to generate a weighted similarity matrix.
[0118] For example, for the video data block feature templates of high definition, 5 - 30 minutes, and MP4 format, and those of standard definition, shorter than 5 minutes, and AVI format. When calculating the similarity, for the resolution feature, since high definition and standard definition are different, a certain score (such as -0.3) is subtracted from the similarity; for the duration range, since they are different, a certain score (such as -0.2) is subtracted from the similarity; for the different formats, a certain score (such as -0.3) is subtracted from the similarity, and the comprehensive similarity is 0.2. Then, according to the policy effectiveness, a weight coefficient is assigned to each similarity. Assume that the policy effectiveness of the video data block feature template of high definition, 5 - 30 minutes, and MP4 format is 1, and that of the video data block feature template of standard definition, shorter than 5 minutes, and AVI format is 0.9. Then the weight coefficient for the high-definition template is 1 / (1 + 0.9) ≈ 0.53, and the weight coefficient for the standard-definition template is 0.9 / (1 + 0.9) ≈ 0.47. For this similarity of 0.2, the weighted similarities are generated as 0.2 × 0.53 = 0.106 (for the high-definition template) and 0.2 × 0.47 = 0.094 (for the standard-definition template). Such calculations are performed for all candidate template pairs to generate a weighted similarity matrix.
[0119] Step S440, traverse the candidate template set based on the weighted similarity matrix, and perform feature mean merging on the data block feature templates with a similarity higher than the merging threshold and a policy effectiveness difference lower than the tolerance range to generate an optimized merged template set.
[0120] Assume that the merging threshold is set to 0.5. If it is found that the similarity between two data block feature templates is higher than the merging threshold and the policy effectiveness difference is lower than the tolerance range. For example, there are two high-definition video data block feature templates, one with a duration range of 5 - 15 minutes and the other with a duration range of 15 - 30 minutes. Their similarity is 0.6 (higher than 0.5), and their policy effectiveness values are 0.95 and 0.9 respectively, with a difference of 0.05 lower than the tolerance range (assuming the tolerance range is 0.1). Then these two templates are subjected to feature mean merging. For the resolution feature, since both are high definition, it remains unchanged; for the duration range, after merging, it becomes 5 - 30 minutes; for other storage attribute combination modes, such as storage capacity distribution, access path optimization, and data compression level, if they are the same, they remain unchanged, and if they are different, they are weighted averaged according to the policy effectiveness. For example, if the storage node allocation of one is storage node A and that of the other is storage node B, and the policy effectiveness corresponding to storage node A is 0.95 and that corresponding to storage node B is 0.9, then the merged storage node allocation is more inclined to storage node A (calculated according to weighted average). In this way, an optimized merged template set is generated.
[0121] Step S450: Periodically perform two-way synchronization between the merged template set and the global template library of the central server, and receive the template update instruction sent by the server. The template update instruction includes the identifier of the expired template to be replaced and the corresponding new template in the merged template set.
[0122] Step S460: According to the template update instruction, clear the data block feature templates in the local cache of the edge computing node that match the identifier of the expired template, insert the new template into the head of the cache queue, and update the values of the active threshold and the effective threshold according to the call frequency distribution of the candidate template set in the most recent synchronization period.
[0123] For example, synchronization is performed once a day. During the synchronization process, the template update instruction sent by the server is received. The template update instruction includes the identifier of the expired template to be replaced and the corresponding new template in the merged template set. Suppose the server detects that a certain old video data block feature template (such as a video template with low resolution, a specific duration range, and a specific format) is no longer applicable to the current storage requirements, and marks it as an expired template. The edge computing node, according to the template update instruction, clears the data block feature templates in the local cache that match the identifier of the expired template, and inserts the new template into the head of the cache queue, so that the new template will be preferentially matched and used. At the same time, the values of the active threshold and the effective threshold are updated according to the call frequency distribution of the candidate template set in the most recent synchronization period. For example, if it is found that the call frequency of high-resolution videos has increased recently while the call frequency of low-resolution videos has decreased, the active threshold may be appropriately increased (because the call frequency of high-resolution videos is generally higher), and the effective threshold is adjusted according to the overall policy effectiveness to better meet the video storage requirements of the video business department.
[0124] For the financial statement data storage scenario of the finance department, the process is similar. When caching the frequently used data block feature template and the storage attribute combination mode, the data block feature template may include feature information such as the type of the financial statement (such as annual statement, quarterly statement, etc.), the size range of the financial statement, the data type in the financial statement (such as revenue and expenditure data, balance sheet data, etc.), and the corresponding storage attribute combination mode (such as storage node allocation, access path priority, compression level, etc.). When calculating the call frequency and policy effectiveness of the policy call record, the policy effectiveness is calculated according to the matching ratio between the actual performance indicators (such as the read and write rates of the storage node, compression efficiency, access path load balancing degree, etc.) after the financial statement is stored and the expected storage attribute set (such as the access speed requirements for different financial statements, the compression level corresponding to the data accuracy requirements, etc.). Subsequent operations such as screening the candidate template set, clustering analysis, merging templates, two-way synchronization, and template update are also performed similarly according to the characteristics of the financial statement data to optimize the storage policy of the financial statement data in the edge computing node of the all-flash storage system.
[0125] Figure 2 FIG. shows a schematic diagram of exemplary hardware and software components of an information processing system 100 for file all-flash storage that can implement the idea of the present application provided by some embodiments of the present application. For example, the processor 120 can be used on the information processing system 100 for file all-flash storage and is used to execute the functions in the present application.
[0126] The information processing system 100 for file all-flash storage can be a general-purpose server or a special-purpose server, both of which can be used to implement the information processing method for file all-flash storage of the present application. Although only one server is shown in the present application, for convenience, the functions described in the present application can be implemented in a distributed manner on multiple similar platforms to balance the processing load.
[0127] For example, the information processing system 100 for file all-flash storage can include a network port 110 connected to a network, one or more processors 120 for executing program instructions, a communication bus 130, and different forms of storage media 140, such as disks, ROMs, or RAMs, or any combination thereof. Exemplarily, the information processing system 100 for file all-flash storage can also include program instructions stored in ROM, RAM, or other types of non-transitory storage media, or any combination thereof. According to these program instructions, the method of the present application can be implemented. The information processing system 100 for file all-flash storage also includes an input / output (I / O) interface 150 between the computer and other input / output devices.
[0128] For ease of explanation, only one processor is described in the information processing system 100 for file all-flash storage. However, it should be noted that the information processing system 100 for file all-flash storage in the present application can also include multiple processors. Therefore, the steps executed by one processor described in the present application can also be jointly executed or separately executed by multiple processors. For example, if the processor of the information processing system 100 for file all-flash storage executes step A and step B, it should be understood that step A and step B can also be jointly executed by two different processors or separately executed in one processor. For example, the first processor executes step A, the second processor executes step B, or the first processor and the second processor jointly execute steps A and B.
[0129] In addition, an embodiment of the present invention also provides a readable storage medium, in which computer-executable instructions are preset. When the processor executes the computer-executable instructions, the information processing method for file all-flash storage as described above is implemented.
[0130] It should be noted that, in order to simplify the description of the present invention disclosure and thus assist in the understanding of one or more embodiments of the invention, in the foregoing description of the embodiments of the present invention, sometimes a plurality of features are incorporated into one embodiment, drawing or description thereof.
Claims
1. An information processing method applied to all-flash file storage, characterized in that: The method comprises: Acquire storage request information received by the target storage node, wherein the storage request information includes a set of data blocks to be processed and a corresponding set of expected storage attributes; wherein the set of expected storage attributes is used to describe configuration requirements of the set of data blocks to be processed in terms of storage capacity distribution, access path optimization, and data compression level; Performing multi-dimensional feature extraction on the set of data blocks to be processed to generate a data block feature vector, and performing attribute encoding on the set of expected storage attributes to generate a storage attribute feature vector; Inputting the data block feature vector and the storage attribute feature vector into a pre-trained storage strategy generation model, dynamically weighting the data block feature vector and the storage attribute feature vector through the storage strategy generation model, and outputting a target storage strategy parameter set; Generate storage node allocation instructions, path priority configuration and compression execution parameters according to the target storage policy parameter set, and drive the all-flash storage system to perform distributed storage operations on the set of data blocks to be processed based on the storage node allocation instructions; wherein the target storage policy parameter set satisfies the constraints of all configuration requirements in the expected storage attribute set.
2. The information processing method applied to all-flash file storage according to claim 1, characterized in that: The expected storage attribute set includes at least one sub-attribute group, each sub-attribute group corresponds to a storage performance indicator type; the attribute encoding of the expected storage attribute set to generate a storage attribute feature vector includes: For each sub-attribute group, select a corresponding encoding rule according to its corresponding storage performance indicator type; If the sub-attribute group is a numerical indicator, it is mapped to a preset dimensional feature space through piecewise normalization; if the sub-attribute group is a logical indicator, a discrete feature sequence is generated through binary vectorization; The feature sequences encoded by each sub-attribute group are weighted and spliced according to the indicator priority to generate a global storage attribute feature vector; wherein the storage performance indicator types include access delay threshold, data persistence level and storage node load balancing coefficient.
3. The information processing method applied to all-flash file storage according to claim 1, characterized in that: The step of extracting multi-dimensional features from the set of data blocks to be processed to generate data block feature vectors includes: Analyze the content distribution pattern of each data block in the set of data blocks to be processed, and count the metadata features of each data block, wherein the metadata features include data block size distribution, hot and cold access marks, and associated file types; Constructing an initial feature set according to the metadata features, and performing a time series correlation analysis on the initial feature set to determine a dynamic weight coefficient of each feature dimension within a historical storage period; The initial feature set is subjected to noise reduction processing based on the dynamic weight coefficient, feature dimensions whose weights are lower than a preset threshold are eliminated, and an optimized data block feature vector is generated.
4. The information processing method applied to all-flash file storage according to claim 3, characterized in that: The training method of the storage strategy generation model includes: Acquire a historical storage record data set, wherein the historical storage record data set includes a plurality of historical data block feature samples, corresponding historical storage attribute samples and actual storage strategy parameter labels; Constructing an initial neural network model, the initial neural network model comprises a feature interaction layer, a strategy prediction layer and a constraint verification layer; wherein the constraint verification layer is used to verify whether the prediction strategy parameters meet the configuration requirements of the historical storage attribute samples; Input the historical data block feature samples and the historical storage attribute samples into the initial neural network model, calculate the first loss value between the predicted storage policy parameter output by the policy prediction layer and the actual storage policy parameter label, and calculate the second loss value between the constraint satisfaction output by the constraint verification layer and the preset standard value; The parameters of the initial neural network model are adjusted according to the weighted sum of the first loss value and the second loss value until the weighted sum is lower than a convergence threshold, thereby obtaining a trained storage strategy generation model.
5. The information processing method applied to all-flash file storage according to claim 4, characterized in that: The step of inputting the historical data block feature samples and the historical storage attribute samples into the initial neural network model, calculating a first loss value between the prediction parameter output by the strategy prediction layer and the actual storage strategy parameter label, and calculating a second loss value between the constraint satisfaction output by the constraint verification layer and a preset standard value includes: Inputting the historical data block feature samples and the historical storage attribute samples into the strategy prediction layer of the initial neural network model to generate a set of prediction storage strategy parameters, wherein the set of prediction storage strategy parameters includes prediction node allocation parameters, prediction path priority parameters and prediction compression parameters; Calculate the difference between the predicted storage strategy parameter set and the actual storage strategy parameter label item by item to obtain a node allocation deviation value, a path priority deviation value, and a compression parameter deviation value; According to the configuration requirements in the historical storage attribute sample, a deviation value, a path priority deviation value, and a compression parameter deviation value are assigned a weight coefficient to generate a weighted deviation vector, and the weighted deviation vector is subjected to nonlinear normalization processing to generate the first loss value in scalar form; Synchronously inputting the predicted storage strategy parameter set and the historical storage attribute sample into the constraint verification layer, parsing the access delay threshold, data persistence level and load balancing coefficient in the historical storage attribute sample; Verifying whether the delay prediction value corresponding to the predicted path priority parameter satisfies the upper limit constraint based on the access delay threshold, and generating path delay satisfaction, and verifying whether the number of redundant copies corresponding to the predicted node allocation parameter satisfies the minimum redundancy constraint based on the data persistence level, and generating redundancy satisfaction, and verifying whether the node load distribution deviation corresponding to the predicted node allocation parameter is lower than the balance constraint threshold based on the load balancing coefficient, and generating load balancing satisfaction; Aggregating the path delay satisfaction, redundancy satisfaction and load balancing satisfaction according to a preset ratio to generate an overall constraint satisfaction, and performing a logarithmic difference calculation between the overall constraint satisfaction and a preset standard value to generate the second loss value; Inputting the first loss value and the second loss value into a weight allocator, generating a dynamic weight ratio according to the storage attribute priority of each sample in the historical storage record data set, and performing weighted summation of the first loss value and the second loss value according to the dynamic weight ratio to generate a total loss value; The total loss value is back-propagated to the parameter update module of the initial neural network model, and the trainable parameters of the strategy prediction layer and the constraint verification layer are adjusted until the total loss value is lower than the convergence threshold.
6. The information processing method applied to all-flash file storage according to claim 5, characterized in that: After driving the all-flash storage system to perform a distributed storage operation on the set of data blocks to be processed based on the storage node allocation instruction, the method further includes: Collecting actual storage performance indicators of the set of data blocks to be processed in real time, wherein the actual storage performance indicators include node read and write rate, compression efficiency and path load balance; Compare the actual storage performance indicator with the configuration requirements in the expected storage attribute set to generate a performance deviation report; If there are indicator items in the performance deviation report that exceed the tolerance threshold, the online update mechanism of the storage policy generation model is triggered, and the actual storage performance indicator and the corresponding storage request information are input into the model as incremental training data to readjust the model parameters.
7. The information processing method applied to all-flash file storage according to claim 6, characterized in that: The execution steps of the online update mechanism include: Extracting incremental storage operation records within a preset time window from the log database of the all-flash storage system to generate an incremental storage operation record set, and performing field mapping between the storage request information in the incremental storage operation record set and the actual storage performance indicator to generate an incremental training data sample set; Freezing the network parameters of the feature interaction layer in the storage strategy generation model to generate an initial fine-tuning model in a frozen parameter state; Generate a ranking list of severity of the indicator items according to the deviation amplitude of each indicator item in the performance deviation report that exceeds the tolerance threshold, and assign layer-by-layer decreasing learning rate weights to the network parameters of the strategy prediction layer based on the ranking list of severity of the indicator items to generate dynamic learning rate configuration parameters; Input the incremental training data sample set into the initial fine-tuning model, perform back propagation training according to the dynamic learning rate configuration parameters, generate updated policy prediction layer network parameters, and splice the updated policy prediction layer network parameters with the feature interaction layer network parameters in the frozen parameter state to generate a candidate update model parameter set; The candidate update model parameter set is loaded into the verification environment of the storage strategy generation model, the verification set storage request information is input, a verification set prediction strategy parameter set is generated, the simulated storage performance index corresponding to the verification set prediction strategy parameter set is matched with the expected storage attribute set in the verification set storage request information, and a verification set prediction accuracy improvement value is generated; When the prediction accuracy improvement value of the validation set exceeds a set percentage, the candidate update model parameter set is marked as a valid update parameter set, the valid update parameter set is encapsulated as a model parameter update instruction, and a synchronization data packet carrying a timestamp and a version identifier is generated; Traversing all target storage nodes of the all-flash storage system, detecting the difference between the storage policy generation model parameter version currently running on each target storage node and the version identifier in the synchronization data packet, and if there is a target storage node running a parameter version earlier than the version identifier in the synchronization data packet, sending a forced synchronization signal to the target storage node, triggering the target storage node to interrupt the current storage operation and load the valid update parameter set; After completing the parameter synchronization of all target storage nodes, the expired data records in the incremental training data sample set are cleared to release the storage space of the log database.
8. The information processing method applied to all-flash file storage according to claim 1, characterized in that: When the storage strategy generation model is deployed on an edge computing node of an all-flash storage system, the method further includes: Cache frequently used data block feature templates and storage attribute combination patterns locally at edge computing nodes; When new storage request information is received, the data block feature vector is preferentially matched with the data block feature template in the cache for similarity. If the similarity matching result indicates that the matching degree is higher than the fast response threshold, the pre-generated strategy parameters in the cache are directly called to perform the storage operation, bypassing the calculation process of the storage strategy generation model.
9. The information processing method applied to all-flash file storage according to claim 8, characterized in that: The method for updating the frequently used data block feature template includes: Monitor the policy call records of edge computing nodes, obtain the call frequency of each data block feature template and the policy effectiveness generated in the corresponding storage operation, where the policy effectiveness is calculated by the matching ratio of the actual performance index after the storage operation and the expected storage attribute set; Generate dynamic screening conditions according to the call frequency and policy effectiveness, screen out data block feature templates whose call frequency is higher than the active threshold and whose policy effectiveness is higher than the effective threshold, and form a candidate template set; Performing cluster analysis on the data block feature templates in the candidate template set, calculating the similarity between any two data block feature templates, and assigning a weight coefficient to each similarity according to the effectiveness of the strategy to generate a weighted similarity matrix; Based on the weighted similarity matrix, the candidate template set is traversed, and the feature mean of the data block feature templates whose similarity is higher than the merging threshold and whose strategy effectiveness difference is lower than the tolerance range is merged to generate an optimized merged template set; Periodically bidirectionally synchronizing the merged template set with the global template library of the central server, and receiving a template update instruction issued by the server, wherein the template update instruction includes an expired template identifier to be replaced and a corresponding new template in the merged template set; According to the template update instruction, the data block feature template matching the expired template identifier in the local cache of the edge computing node is cleared, and the new template is inserted into the head of the cache queue. At the same time, the values of the active threshold and the effective threshold are updated according to the call frequency distribution of the candidate template set in the most recent synchronization cycle.
10. An information processing system applied to all-flash file storage, characterized in that: The information processing system applied to all-flash file storage includes a processor and a memory, the memory is connected to the processor, the memory is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the memory to implement the information processing method applied to all-flash file storage as described in any one of claims 1 to 9 above.
Citation Information
Patent Citations
All-flash storage server system data management method and related components
CN112000289A
Caching method and device for deduplication metadata of full-flash storage system and medium
CN112148217A
Data storage method and system of all-flash storage system, medium and program product
CN119002820A
Data storage method based on large model
CN119292524A
Cited By
Storage information verification system based on consensus mechanism
CN120602228A
A storage information verification system based on consensus mechanism
CN120602228B
Data loading optimization method based on distributed cache
CN120804163A
File caching performance improving method and system based on distributed storage
CN121722325A