A data storage method and system based on an artificial intelligence accelerator
By introducing artificial intelligence accelerators into data storage, analyzing feature indicators and generating data sequence tables, the performance bottlenecks and inefficient resource utilization in traditional data storage methods are solved when processing large-scale data, and efficient and intelligent data storage allocation is achieved.
Patent Information
- Application Number
- CN202510022970.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-07
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-01-07
AI Technical Summary
Traditional data storage methods have performance bottlenecks when processing large-scale data, which is difficult to meet the needs of real-time and efficientness, and lack intelligent management and optimization mechanisms, resulting in inefficient storage resource utilization.
The data storage method based on artificial intelligence accelerator is adopted to analyze feature indicators, generate a sequence table of data to be stored, and allocate data based on this table to optimize the data storage allocation strategy.
Effectively identify data feature performance, improve the accuracy and efficiency of feature extraction, accurately evaluate the storage priority and access characteristics of data, generate efficient data order tables to be stored, and optimize data storage allocation strategies.
Smart Images

Figure CN119415975B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of information technology, and specifically relates to a data storage method and system based on an artificial intelligence accelerator. Background Art
[0002] With the advent of the big data era, the demand for data storage has increased sharply, posing higher requirements for storage efficiency and accuracy. Traditional data storage methods have performance bottlenecks when dealing with large-scale data and are difficult to meet the requirements of real-time and high efficiency. At the same time, the existing storage systems often lack intelligent management and optimization mechanisms, resulting in low utilization efficiency of storage resources and increasing business risks.
[0003] In the prior art, relying on traditional monitoring tools and methods, which focus on the basic monitoring of storage hardware and clusters, but in the above prior art, there is a lack of real-time acquisition and analysis of the response data of storage units, unable to more accurately identify performance anomalies, and in the above prior art, it is impossible to accurately evaluate the storage priority and access characteristics of data, unable to generate an efficient order list of data to be stored, and unable to optimize the data storage allocation strategy.
[0004] Therefore, the present invention provides a data storage method and system based on an artificial intelligence accelerator. Summary of the Invention
[0005] In order to make up for the deficiencies of the prior art and solve at least one of the technical problems proposed in the background art.
[0006] The technical solution adopted by the present invention to solve its technical problems is: a data storage method based on an artificial intelligence accelerator, including:
[0007] Analyze the characteristic indicators to obtain the secondary values of the characteristic indicators, further analyze the secondary values of the characteristic indicators to obtain the characteristic performance values, and based on the comparison between the characteristic performance values and the characteristic performance thresholds, further analyze the data to be stored;
[0008] Through the analysis and processing of the information entropy value, obtain the number value of key data blocks and the information entropy overrun value, and perform a weighted ratio process on the number value of key data blocks and the information entropy overrun value to obtain the key performance value;
[0009] Screen the data to be stored in the dataset according to the access frequency to obtain the proportion of the number of high-frequency data blocks and the access frequency deviation ratio, and perform a weighted sum on the proportion of the number of high-frequency data blocks and the access frequency deviation ratio to obtain the data access value;
[0010] Combine and analyze the key performance value and the data access value to generate an order list of data to be stored, and allocate the data to be stored based on the order list of data to be stored;
[0011] As a further technical solution of the present invention: The acquisition method of the characteristic performance value is as follows:
[0012] Count the number of unqualified times of feature extraction, and process the ratio of the number of unqualified times of feature extraction to the total number of feature extractions to obtain the unqualified ratio of feature extraction;
[0013] Sum up all the secondary difference ratios of feature indicators and take the average value to obtain the secondary deviation value of feature indicators;
[0014] Mark the unqualified ratio of feature extraction as BHG, and mark the secondary deviation value of feature indicators as YPC;
[0015] Perform data processing on the unqualified ratio of feature extraction BHG and the secondary deviation value of feature indicators YPC, and through the formula Calculate to obtain the characteristic performance value BX, where s1 and s2 are both preset proportionality coefficients;
[0016] As a further technical solution of the present invention: The acquisition method of the unqualified ratio of feature extraction BHG is as follows:
[0017] Compare the secondary value of the feature indicator with the secondary threshold of the feature indicator. The specific comparison process is as follows:
[0018] If the secondary value of the feature indicator is greater than the secondary threshold of the feature indicator, the feature extraction is unqualified;
[0019] If the secondary value of the feature indicator is less than or equal to the secondary threshold of the feature indicator, the feature extraction is qualified;
[0020] During the data storage period, count the number of unqualified times of feature extraction, and process the ratio of the number of unqualified times of feature extraction to the total number of feature extractions to obtain the unqualified ratio of feature extraction;
[0021] As a further technical solution of the present invention: The acquisition method of the secondary deviation value of feature indicators YPC is as follows:
[0022] Based on the unqualified feature indicators, perform a difference process on the secondary value of the feature indicator and the secondary threshold of the feature indicator, and take the absolute value to obtain the secondary difference of the feature indicator. Perform a ratio process on the secondary difference of the feature indicator and the secondary threshold of the feature indicator to obtain the secondary difference ratio of the feature indicator. Sum up all the secondary difference ratios of the feature indicators and take the average value to obtain the secondary deviation value of the feature indicator;
[0023] As a further technical solution of the present invention: The acquisition method of the key performance value GJ is as follows:
[0024] Perform a weighted ratio process on the key data block quantity value and the information entropy overrun value to obtain the key performance value, and mark it as GJ;
[0025] As a further technical solution of the present invention: the obtaining method of the key data block quantity value and the information entropy overrun value is as follows:
[0026] Sum all the information entropy values and take the average to obtain the information entropy average value. Calculate the difference between the information entropy value and the information entropy average value and take the absolute value to obtain the information entropy difference;
[0027] Compare the information entropy difference with the preset information entropy difference threshold. The specific comparison process is as follows:
[0028] If the information entropy difference is less than the information entropy difference threshold, mark this data block as a key data block;
[0029] If the information entropy deviation is greater than or equal to the information entropy deviation threshold, mark this data block as a non-key data block;
[0030] Count the number of key data blocks and calculate the ratio with the number of data blocks to obtain the key data block quantity value;
[0031] Based on any one key information block;
[0032] Calculate the difference between the information entropy difference and the information entropy difference threshold to obtain the information entropy deviation. Calculate the ratio of the information entropy deviation to the information entropy average value, process to obtain the information entropy deviation ratio. Sum all the information entropy deviation ratios and take the average to obtain the information entropy overrun value;
[0033] As a further technical solution of the present invention: the obtaining method of the data access value FW is as follows:
[0034] Perform a weighted sum of the high-frequency data block quantity ratio and the access frequency deviation ratio to obtain the data access value, and mark it as FW;
[0035] As a further technical solution of the present invention: the obtaining methods of the high-frequency data block quantity ratio and the access frequency deviation ratio are as follows:
[0036] Count the number of data blocks in the access frequency curve that exceed the baseline and perform a ratio process with the number of data blocks to obtain the high-frequency data block quantity ratio;
[0037] In the access frequency curve, select the data points that exceed the baseline and connect them to obtain the high-frequency data curve;
[0038] Based on any one data point of the high-frequency data curve;
[0039] Perform a difference process on the access frequency and the access frequency standard value to obtain the access frequency difference. Perform a ratio process on the access frequency difference and the access frequency standard value to obtain the access frequency difference ratio;
[0040] Sum all the access frequency difference ratios and take the average to obtain the access frequency deviation ratio;
[0041] As a further technical solution of the present invention: the acquisition method of the storage characterization value CH is:
[0042] Perform data processing on the obtained key performance value GJ and the data access value FW, and calculate the storage characterization value CH through the formula where a1 and a2 are both preset proportionality coefficients;
[0043] As a further technical solution of the present invention: Feature analysis module: preprocess the data to be stored, extract features based on the preprocessed data to be stored to obtain feature indicators, analyze the feature indicators to obtain secondary feature indicator values, further analyze the secondary feature indicator values to obtain feature performance values, compare the feature performance values with the feature performance threshold, and based on the comparison result, if the feature performance value is less than the feature performance threshold, further analyze the data to be stored;
[0044] Analysis signal processing module: based on the analysis and processing of the information entropy value, obtain the key data block quantity value and the information entropy overrun value, perform a weighted ratio process on the key data block quantity value and the information entropy overrun value to obtain the key performance value, screen the data to be stored in the dataset according to the access frequency to obtain the high-frequency data block quantity ratio and the access frequency deviation ratio, and perform a weighted sum on the high-frequency data block quantity ratio and the access frequency deviation ratio to obtain the data access value;
[0045] Storage allocation module: combine and analyze the key performance value and the data access value to generate an order table of the data to be stored, and allocate the data to be stored based on the order table of the data to be stored.
[0046] The beneficial effects of the present invention are as follows:
[0047] 1. Preprocess the data to be stored, extract features based on the preprocessed data to be stored to obtain feature indicators, analyze the feature indicators to obtain secondary feature indicator values, further analyze the secondary feature indicator values to obtain feature performance values, compare the feature performance values with the feature performance threshold, and based on the comparison result, if the feature performance value is less than the feature performance threshold, further analyze the data to be stored. The present invention can effectively identify the data feature performance, improve the accuracy and efficiency of feature extraction, and facilitate timely discovery and handling of problems;
[0048] 2. Generate the information entropy value of the data set according to the information entropy method. Based on the analysis and processing of the information entropy value, obtain the key data block quantity value and the information entropy overrun value. Perform a weighted ratio process on the key data block quantity value and the information entropy overrun value to obtain the key performance value. Screen the data to be stored in the data set according to the access frequency. Analyze and process the access frequency of the data subset according to the screening result to obtain the proportion of the number of high-frequency data blocks and the access frequency deviation ratio. Perform a weighted sum on the proportion of the number of high-frequency data blocks and the access frequency deviation ratio to obtain the data access value. Combine and analyze the key performance value and the data access value to generate an order table for the data to be stored. Based on the order table of the data to be stored, allocate the data to be stored. By combining and analyzing the key performance value and the data access value, the present invention can accurately evaluate the storage priority and access characteristics of the data, thereby generating an efficient order table for the data to be stored and optimizing the data storage allocation strategy. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] The present invention will be further described below with reference to the accompanying drawings.
[0050] Figure 1 is a flowchart of the steps of a data storage method based on an artificial intelligence accelerator according to Embodiment 1 of the present invention;
[0051] Figure 2 is a flowchart of the steps of a data storage method based on an artificial intelligence accelerator according to Embodiment 2 of the present invention;
[0052] Figure 3 is a program block diagram of a data storage system based on an artificial intelligence accelerator according to Embodiment 4 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0053] In order to make the technical means, creative features, achieved purposes and effects of the present invention easy to understand, the present invention will be further described below in conjunction with specific embodiments.
[0054] Embodiment 1
[0055] As Figure 1 shown, a data storage method based on an artificial intelligence accelerator according to an embodiment of the present invention includes:
[0056] Step 1: Preprocess the data to be stored, extract features based on the preprocessed data to be stored to obtain feature indicators, analyze the feature indicators to obtain the secondary values of the feature indicators, and further analyze the secondary values of the feature indicators to obtain the feature performance values;
[0057] It should be noted that the feature indicators include access frequency and importance;
[0058] In some embodiments, the data to be stored is merged to obtain a data set, the data set is cleaned to remove duplicate data therein, and a recurrent neural network (RNN) and its variants (such as LSTM) are used to extract features from the data set to obtain feature metrics, where the feature metrics include the access frequency and importance of the data to be stored;
[0059] Specifically, based on any one of the feature metrics;
[0060] The feature metric is compared with the standard value of the feature metric, and the specific comparison process is as follows:
[0061] If the feature metric is greater than or equal to the standard value of the feature metric, no processing is performed;
[0062] If the feature metric is less than the standard value of the feature metric, it is marked as an unimportant feature metric;
[0063] The difference between the feature metric and the standard value of the feature metric is calculated and the absolute value is taken to obtain the feature metric difference;
[0064] It should be noted that the standard value of the feature metric is set by those skilled in the art by traversing a large amount of historical data, and the purpose is to determine whether the feature metric in the data storage process meets the requirements;
[0065] The number of secondary metrics in the feature metrics is counted, and a ratio is taken with the number of feature metrics to obtain the secondary metric quantity ratio;
[0066] Based on any one of the secondary metrics;
[0067] The feature metric difference is calculated as a ratio with the standard value of the feature metric, and the feature metric difference value is obtained through processing. The sum of all feature metric difference values is calculated and the average value is taken to obtain the feature metric difference ratio;
[0068] The secondary metric quantity ratio and the feature metric difference ratio are summed to obtain the feature metric secondary value;
[0069] The feature metric secondary value is compared with the feature metric secondary threshold, and the specific comparison process is as follows:
[0070] If the feature metric secondary value is greater than the feature metric secondary threshold, the feature extraction is unqualified;
[0071] If the feature metric secondary value is less than or equal to the feature metric secondary threshold, the feature extraction is qualified;
[0072] During the data storage period, the number of unqualified feature extraction times is counted, and a ratio is taken with the total number of feature extraction times to obtain the unqualified feature extraction ratio;
[0073] Based on the unqualified characteristic indicators, the difference between the secondary value of the characteristic indicator and the secondary threshold of the characteristic indicator is processed, and the absolute value is taken to obtain the secondary difference of the characteristic indicator. The secondary difference of the characteristic indicator is processed by ratio with the secondary threshold of the characteristic indicator to obtain the secondary difference ratio of the characteristic indicator. The secondary difference ratios of all characteristic indicators are summed up and averaged to obtain the secondary deviation value of the characteristic indicator;
[0074] Mark the unqualified ratio of feature extraction as BHG and mark the secondary deviation value of the characteristic indicator as YPC;
[0075] Perform data processing on the unqualified ratio of feature extraction BHG and the secondary deviation value of the characteristic indicator YPC. Through the formula Calculate to obtain the feature performance value BX, where s1 and s2 are both preset proportionality coefficients;
[0076] It should be noted that the feature performance value BX reflects the quantitative standard of the secondary degree of the characteristic indicator. The larger the feature performance value BX, the smaller the unqualified ratio of feature extraction BHG and the secondary deviation value of the characteristic indicator YPC, and the smaller the secondary degree of the characteristic indicator. On the contrary, the smaller the feature performance value BX, the larger the unqualified ratio of feature extraction BHG and the secondary deviation value of the characteristic indicator YPC, and the larger the secondary degree of the characteristic indicator;
[0077] Step 2: Compare the feature performance value with the feature performance threshold. Based on the comparison result, if the feature performance value is less than the feature performance threshold, further analyze the data to be stored;
[0078] Specifically, compare the feature performance value with the feature performance threshold;
[0079] If the feature performance value is greater than or equal to the feature performance threshold, further analyze the data to be stored and generate an analysis signal;
[0080] If the feature performance value is less than the feature performance threshold, mark this part of the data as secondary data;
[0081] The technical solution of the embodiment of the present invention is as follows: preprocess the data to be stored, perform feature extraction based on the preprocessed data to be stored to obtain characteristic indicators, analyze the characteristic indicators to obtain the secondary values of the characteristic indicators, further analyze the secondary values of the characteristic indicators to obtain the feature performance value, compare the feature performance value with the feature performance threshold. Based on the comparison result, if the feature performance value is greater than or equal to the feature performance threshold, further analyze the data to be stored. The present invention can effectively identify the data feature performance, improve the accuracy and efficiency of feature extraction, and facilitate timely discovery and handling of problems.
[0082] Embodiment 2
[0083] Such as Figure 2As shown, based on Embodiment 1, a data storage method based on an artificial intelligence accelerator according to an embodiment of the present invention includes:
[0084] Step 3: Based on the analysis signal, generate the information entropy value of the data set according to the information entropy method. Based on the analysis and processing of the information entropy value, obtain the key data block quantity value and the information entropy overrun value, and perform a weighted ratio process on the key data block quantity value and the information entropy overrun value to obtain the key performance value;
[0085] It should be noted that the information entropy value reflects the uncertainty or chaos degree of the data. The larger the information entropy value, the more uncertain or chaotic the data is. On the contrary, the smaller the information entropy value, the more certain or ordered the data is;
[0086] In some embodiments, obtain the information entropy value of each data block in the data set, and classify the data set based on the information entropy value;
[0087] Among them, the data block represents the data in the data set;
[0088] Sum all the information entropy values and take the average to obtain the information entropy average value. Calculate the difference between the information entropy value and the information entropy average value, and take the absolute value to obtain the information entropy difference;
[0089] Compare the information entropy difference with a preset information entropy difference threshold. The specific comparison process is as follows:
[0090] If the information entropy difference is less than the information entropy difference threshold, mark this data block as a key data block;
[0091] If the information entropy deviation is greater than or equal to the information entropy deviation threshold, mark this data block as a non-key data block;
[0092] Count the number of key data blocks, and calculate the ratio with the number of data blocks to obtain the key data block quantity value;
[0093] Based on any one key information block;
[0094] Calculate the difference between the information entropy difference and the information entropy difference threshold to obtain the information entropy deviation. Calculate the ratio of the information entropy deviation to the information entropy average value, process to obtain the information entropy deviation ratio. Sum all the information entropy deviation ratios and take the average to obtain the information entropy overrun value;
[0095] Perform a weighted ratio process on the key data block quantity value and the information entropy overrun value to obtain the key performance value, and mark it as GJ;
[0096] It should be noted that the meaning represented by the key performance value is as follows: the key performance value reflects the quantitative standard of the importance degree of the data set. The larger the key performance value, the higher the importance degree of the data set. On the contrary, the smaller the key performance value, the lower the importance degree of the data set;
[0097] Step 4: Screen the data to be stored in the data set according to the access frequency, analyze and process the access frequency of the data subset according to the screening result, obtain the proportion of the number of high-frequency data blocks and the access frequency deviation ratio, and perform weighted summation on the proportion of the number of high-frequency data blocks and the access frequency deviation ratio to obtain the data access value;
[0098] Divide the data set into several data subsets;
[0099] Based on any one of the data subsets;
[0100] In some embodiments, an X-Y coordinate system is established, where the X-axis is the data block order and the Y-axis is the access frequency. Mark the access frequency values in the coordinate system and connect the data points to obtain an access frequency curve;
[0101] In the access frequency curve, mark the access frequency standard value as a reference point and draw a reference line;
[0102] Count the number of data blocks in the access frequency curve that exceed the reference line, and perform a ratio process with the number of data blocks to obtain the proportion of the number of high-frequency data blocks;
[0103] In the access frequency curve, select the data points that exceed the reference line and connect them to obtain a high-frequency data curve;
[0104] Based on any data point of the high-frequency data curve;
[0105] Perform a difference process on the access frequency and the access frequency standard value to obtain an access frequency difference, and perform a ratio process on the access frequency difference and the access frequency standard value to obtain an access frequency difference ratio;
[0106] Sum up all the access frequency difference ratios and take the average value to obtain the access frequency deviation ratio;
[0107] Perform weighted summation on the proportion of the number of high-frequency data blocks and the access frequency deviation ratio to obtain a data access value, and mark it as FW;
[0108] It should be noted that the data access value reflects the quantitative standard of the access frequency of the data set. The higher the data access value, the higher the access frequency of the data set. On the contrary, the lower the data access value, the lower the access frequency of the data set;
[0109] Compare the data access value with the data access threshold. The specific comparison process is as follows:
[0110] If the data access value is greater than or equal to the data access threshold, the data subset with high-frequency access is marked as hot data;
[0111] If the data access value is less than the data access threshold, the data subset with low-frequency access is marked as cold data;
[0112] The additional classification criteria for hot data and cold data are as follows:
[0113] From the time dimension, the data within the specified time range is marked as hot data, and vice versa as cold data. The specified time range includes three months, five months, and twelve months;
[0114] Step Five: Combine and analyze the key performance value and the data access value to generate an order table for the data to be stored, and allocate the data to be stored based on the order table of the data to be stored;
[0115] Specifically, perform data processing on the obtained key performance value GJ and the data access value FW, and calculate the storage representation value CH through the formula where a1 and a2 are both preset proportionality coefficients;
[0116] In some embodiments, sort the data to be stored in descending order according to the storage representation value CH to generate a storage order table;
[0117] Divide the data storage unit into a secondary storage unit and a primary storage unit;
[0118] According to the comparison result of the characteristic performance value, store the secondary data in the secondary storage unit, and store the data in the primary storage unit in sequence according to the storage order table;
[0119] Step Six: Obtain the access frequency value of the data to be stored in real time, re-obtain the characteristic performance value according to the real-time access frequency value, and detect the data to be stored according to the comparison result between the characteristic performance value and the characteristic performance threshold;
[0120] In some embodiments, if the data to be stored is stored back and forth between the primary storage unit and the secondary storage unit, increase the characteristic performance threshold.
[0121] The technical solution of the embodiment of the present invention is as follows: Generate the information entropy value of the data set according to the information entropy method. Based on the analysis and processing of the information entropy value, obtain the key data block quantity value and the information entropy overrun value. Perform weighted ratio processing on the key data block quantity value and the information entropy overrun value to obtain the key performance value. Screen the data to be stored in the data set according to the access frequency. Analyze and process the access frequency of the data subset according to the screening result to obtain the proportion of the number of high-frequency data blocks and the access frequency deviation ratio. Perform weighted summation on the proportion of the number of high-frequency data blocks and the access frequency deviation ratio to obtain the data access value. Combine and analyze the key performance value and the data access value to generate the order table of the data to be stored. Based on the order table of the data to be stored, perform the allocation of the data to be stored. By combining and analyzing the key performance value and the data access value, the present invention can accurately evaluate the storage priority and access characteristics of the data, thereby generating an efficient order table of the data to be stored and optimizing the data storage allocation strategy.
[0122] Embodiment 3
[0123] A data storage method based on an artificial intelligence accelerator according to an embodiment of the present invention includes:
[0124] Classify the data to be stored through the artificial intelligence accelerator, classify the data to be stored in terms of importance, and divide the data storage unit based on the importance level;
[0125] The specific process for the artificial intelligence accelerator to distinguish and store data is as follows:
[0126] S1: Use a data cleaning tool (such as Pandas, etc.) to preprocess the data to be stored, extract information such as the data access frequency, modification time, data size, data type, etc. of the data, as one of the characteristics of data importance, and clean, denoise, and format the data to be stored to ensure the quality and consistency of the data;
[0127] S2: Use a feature selection algorithm to screen out the most valuable features, reduce the complexity and computational amount of the model, screen out the features most valuable for evaluating data importance, and perform normalization, standardization, etc. on the features to ensure that different features have the same weight in the model, so as to improve the stability and accuracy of the model;
[0128] S3: Select a suitable machine learning or deep learning model (such as decision tree, random forest, neural network, etc.), use historical data and features for training, and configure an artificial intelligence accelerator (such as GPU) to accelerate the training process to obtain a trained model;
[0129] S4: Input the data to be evaluated into the trained model to obtain the importance score of the data, sort the data according to the scoring result, and determine the order of storing the data;
[0130] S5: Continuously optimize and adjust the model according to the actual application scenarios and feedback, utilize the computing power of the accelerator to process the importance evaluation task of the stored data in real time, and ensure the timeliness and accuracy of the data;
[0131] Divide the data storage unit into a core sub-unit and an edge sub-unit, and store the data in sequence according to the importance of the data to be stored.
[0132] Embodiment 4
[0133] As Figure 3 shown, a data storage system based on an artificial intelligence accelerator according to an embodiment of the present invention includes:
[0134] Feature analysis module: Preprocess the data to be stored, extract features based on the preprocessed data to be stored to obtain feature indicators, analyze the feature indicators to obtain secondary feature indicator values, further analyze the secondary feature indicator values to obtain feature performance values, compare the feature performance values with a feature performance threshold, and based on the comparison result, if the feature performance value is less than the feature performance threshold, further analyze the data to be stored;
[0135] Analysis signal processing module: Based on the analysis and processing of the information entropy value, obtain the number value of key data blocks and the information entropy overrun value, perform a weighted ratio process on the number value of key data blocks and the information entropy overrun value to obtain a key performance value, screen the data to be stored in the data set according to the access frequency to obtain the proportion of the number of high-frequency data blocks and the access frequency deviation ratio, and perform a weighted sum on the proportion of the number of high-frequency data blocks and the access frequency deviation ratio to obtain a data access value;
[0136] Storage allocation module: Combine and analyze the key performance value and the data access value to generate an order table of the data to be stored, and allocate the data to be stored based on the order table of the data to be stored.
[0137] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art of this industry should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.
Claims
1. A data storage method based on an artificial intelligence accelerator, characterized in that: include: Analyze the characteristic index, calculate the difference between the characteristic index and the standard value of the characteristic index, and take the absolute value to obtain the characteristic index difference; The number of secondary indicators in the characteristic indicators is counted, and the ratio is processed with the number of characteristic indicators to obtain the ratio of the number of secondary indicators; the characteristic indicator difference is calculated by ratio with the characteristic indicator standard value, and the characteristic indicator difference is obtained by processing; all the characteristic indicator differences are summed and the average is taken to obtain the characteristic indicator difference ratio; the secondary indicator quantity ratio and the characteristic indicator difference ratio are summed to obtain the characteristic indicator secondary value; the characteristic indicator secondary value is further analyzed to obtain the characteristic performance value; based on the comparison between the characteristic performance value and the characteristic performance threshold, the data to be stored is further analyzed; By analyzing and processing the information entropy value, the number of key data blocks and the information entropy excess value are obtained, and the number of key data blocks and the information entropy excess value are processed by weighted ratio to obtain the key performance value; The data to be stored in the data set is screened according to the access frequency to obtain the proportion of high-frequency data blocks and the access frequency deviation ratio, and the proportion of high-frequency data blocks and the access frequency deviation ratio are weighted and summed to obtain the data access value; Combine and analyze the key performance value and the data access value to generate a sequence table of data to be stored, and allocate the data to be stored based on the sequence table of data to be stored; The obtained key performance value GJ and data access value FW are processed by the formula The storage characterization value CH is calculated, where a1 and a2 are both preset proportional coefficients; Sort the stored data in descending order according to the storage characterization value CH to generate a storage order table; dividing the data storage unit into a secondary storage unit and a primary storage unit; According to the comparison result of the characteristic performance value, the secondary data is stored in the secondary storage unit, and according to the storage order table, the data is sequentially stored in the primary storage unit; If the data to be stored is stored back and forth between the primary storage unit and the secondary storage unit, the feature expression threshold is adjusted higher.
2. The data storage method based on artificial intelligence accelerator according to claim 1, characterized in that: The method for obtaining the characteristic performance value is as follows: Count the number of unqualified feature extractions, and perform ratio processing on the number of unqualified feature extractions and the total number of feature extractions to obtain the unqualified feature extraction ratio; Sum up all the minor difference ratios of the characteristic indicators and take the average to obtain the minor deviation value of the characteristic indicator; The feature extraction failure ratio is marked as BHG, and the feature index minor deviation value is marked as YPC; The feature extraction unqualified ratio BHG and the feature index minor deviation value YPC are processed by data, and the formula The characteristic performance value BX is calculated, where s1 and s2 are both preset proportional coefficients.
3. The data storage method based on artificial intelligence accelerator according to claim 2 is characterized in that: The method for obtaining the feature extraction failure ratio BHG is as follows: The feature index secondary value is compared with the feature index secondary threshold. The specific comparison process is as follows: If the feature index secondary value is greater than the feature index secondary threshold, the feature extraction fails; If the feature index secondary value is less than or equal to the feature index secondary threshold, the feature extraction is qualified; During the data storage period, the number of unqualified feature extractions is counted, and the ratio of the number of unqualified feature extractions to the total number of feature extractions is processed to obtain the unqualified feature extraction ratio.
4. The data storage method based on artificial intelligence accelerator according to claim 3 is characterized in that: The method for obtaining the secondary deviation value YPC of the characteristic index is as follows: Based on the failure of the characteristic indicator, the secondary value of the characteristic indicator is processed with the secondary threshold of the characteristic indicator, and the absolute value is taken to obtain the secondary difference of the characteristic indicator. The secondary difference of the characteristic indicator is processed with the secondary threshold of the characteristic indicator to obtain the secondary difference ratio of the characteristic indicator. All the secondary difference ratios of the characteristic indicators are summed up and the average is taken to obtain the secondary deviation value of the characteristic indicator.
5. The data storage method based on artificial intelligence accelerator according to claim 3 is characterized in that: The key performance value GJ is obtained as follows: The key data block quantity value and the information entropy excess value are weighted ratio processed to obtain the key performance value, which is marked as GJ.
6. The data storage method based on artificial intelligence accelerator according to claim 1, characterized in that: The method for obtaining the number of key data blocks and the information entropy excess value is as follows: Sum all information entropy values and take the average to get the information entropy mean. Calculate the difference between the information entropy value and the information entropy mean and take the absolute value to get the information entropy difference. The information entropy difference is compared with the preset information entropy difference threshold. The specific comparison process is as follows: If the information entropy difference is less than the information entropy difference threshold, the data block is marked as a key data block; If the information entropy deviation is greater than or equal to the information entropy deviation threshold, the data block is marked as a non-critical data block; Count the number of key data blocks and calculate the ratio with the number of data blocks to obtain the value of the number of key data blocks; Based on any key information block; The information entropy difference is calculated with the information entropy difference threshold to obtain the information entropy deviation, the information entropy deviation is calculated with the information entropy mean to obtain the information entropy deviation ratio, all the information entropy deviation ratios are summed up and the average is taken to obtain the information entropy excess value.
7. The data storage method based on artificial intelligence accelerator according to claim 1 is characterized in that: The data access value FW is obtained in the following manner: The weighted sum of the proportion of high-frequency data blocks and the access frequency deviation ratio is taken to obtain the data access value, which is marked as FW.
8. The data storage method based on artificial intelligence accelerator according to claim 1 is characterized in that: The method for obtaining the proportion of the number of high-frequency data blocks and the access frequency deviation ratio is as follows: The number of data blocks exceeding the baseline in the access frequency curve is counted, and the ratio is processed with the number of data blocks to obtain the proportion of high-frequency data blocks; In the access frequency curve, data points exceeding the baseline are selected and connected to obtain a high-frequency data curve; Based on any data point of the high-frequency data curve; Perform difference processing on the access frequency and the access frequency standard value to obtain the access frequency difference, perform ratio processing on the access frequency difference and the access frequency standard value to obtain the access frequency difference ratio; All access frequency difference ratios are summed up and the average is taken to obtain the access frequency deviation ratio.
9. A data storage system based on an artificial intelligence accelerator, the data storage system being used to execute the data storage method according to any one of claims 1 to 8, characterized in that: Feature analysis module: preprocessing the data to be stored, extracting features based on the preprocessed data to be stored to obtain feature indicators, analyzing the feature indicators to obtain feature indicator secondary values, further analyzing the feature indicator secondary values to obtain feature performance values, comparing the feature performance values with the feature performance threshold, and based on the comparison result, if the feature performance value is less than the feature performance threshold, further analyzing the data to be stored; Analysis signal processing module: Based on the analysis and processing of the information entropy value, the number of key data blocks and the information entropy excess value are obtained, and the number of key data blocks and the information entropy excess value are weighted ratio processed to obtain the key performance value, and the data to be stored in the data set are screened according to the access frequency to obtain the proportion of high-frequency data blocks and the access frequency deviation ratio, and the proportion of high-frequency data blocks and the access frequency deviation ratio are weighted summed to obtain the data access value; Storage allocation module: combines and analyzes the key performance value and the data access value, generates a sequence table of data to be stored, and allocates the data to be stored based on the sequence table of data to be stored; The obtained key performance value GJ and data access value FW are processed by the formula The storage characterization value CH is calculated, where a1 and a2 are both preset proportional coefficients; Sort the stored data in descending order according to the storage characterization value CH to generate a storage order table; dividing the data storage unit into a secondary storage unit and a primary storage unit; According to the comparison result of the characteristic performance value, the secondary data is stored in the secondary storage unit, and according to the storage order table, the data is sequentially stored in the primary storage unit; If the data to be stored is stored back and forth between the primary storage unit and the secondary storage unit, the feature expression threshold is adjusted higher.
Citation Information
Patent Citations
System and method for tracing product quality
CN118569744A
Cache data management and control method and system based on distributed encrypted storage
CN119025566A
Data storage method and system based on large model
CN119148946A