Intelligent file storage allocation method and system

By constructing a joint knowledge base for archival storage strategies and a deep optimization reasoning model, the problem that existing archival storage management systems cannot dynamically respond to equipment failures or environmental changes has been solved. This enables intelligent generation and adaptive adjustment of archival storage strategies, improving storage efficiency and security.

CN121412236BActive Publication Date: 2026-06-23TIBET YANRUI INFORMATION SECURITY TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TIBET YANRUI INFORMATION SECURITY TECHNOLOGY CO LTD
Filing Date
2025-10-31
Publication Date
2026-06-23

AI Technical Summary

Technical Problem

Existing archive storage management systems lack in-depth integrated analysis of the risks to the physical value of archives and the security of the storage environment. They cannot be linked with the real-time changes in the physical resource status, resulting in uneven utilization of storage resources. The storage security of important archives cannot be intelligently and precisely guaranteed, and they cannot dynamically respond to equipment failures or environmental changes.

Method used

A joint knowledge base for archival storage strategies is constructed, including a sub-base of standard archival storage processes, a sub-base of historical storage scenario information, and a sub-base of storage device resource information. Deep retrieval is performed through a deep optimization inference model and a fast association index to assess the security level of the storage environment and the probability of archival risks in real time, generate optimized storage allocation strategies, and make dynamic adjustments based on the real-time device status.

Benefits of technology

It enables intelligent generation and adaptive adjustment of file storage strategies, improving storage efficiency and security, and enhancing resource utilization and system response speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121412236B_ABST
    Figure CN121412236B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of file storage, and particularly relates to a file intelligent storage allocation method and system, which comprises the following steps: constructing a joint knowledge base composed of a file standard storage process sub-base, a historical storage scene information sub-base and a storage device resource information sub-base; constructing a demand depth optimization reasoning model, and performing deep retrieval on the knowledge base by using a fast association index; acquiring file and scene data in real time, and calculating the current environment safety level and file damage risk probability by using a safety evaluation model; based on the above evaluation results, generating an initial evaluation score and a storage strategy by using a comprehensive evaluation model; performing secondary matching retrieval by using the reasoning model to obtain an optimized storage allocation strategy; and finally dynamically adjusting storage resource allocation according to the optimized storage allocation strategy and real-time device state; the application realizes intelligent generation and adaptive adjustment of file storage strategies, and improves storage efficiency and safety.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of archival storage technology, and in particular relates to an intelligent archival storage and allocation method and system. Background Technology

[0002] Existing archival storage management systems mostly employ automated systems based on fixed rules or simple database queries. Their storage strategies are often pre-set statically, failing to adapt to complex and ever-changing real-world storage environments. These systems generally lack in-depth analysis integrating the inherent value and risks of the archival entities with the security of the storage environment. They also fail to link with real-time changes in physical resource status, such as equipment health, warehouse temperature and humidity, and shelf availability, resulting in a gap between storage decisions and resource allocation. The consequence is often uneven utilization of storage resources, difficulty in intelligently and precisely ensuring the security of important archives, and an inability to make rapid and optimal adjustments in response to equipment failures or the addition of new archives. Therefore, there is an urgent need for an intelligent allocation system that deeply integrates knowledge reasoning, risk assessment, and dynamic resource scheduling to solve these problems.

[0003] The existing technologies mentioned above have the following problems: Traditional archive storage management systems mainly rely on manual experience or fixed rules for storage allocation, which has problems such as rigid and monotonous strategies, inability to accurately assess archive storage risks, and disconnection from real-time physical resource status. This results in low efficiency of storage resource allocation, inability to dynamically respond to equipment failures or environmental changes, and insufficient overall intelligence and security. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention proposes an intelligent archival storage allocation method and system. The method includes: constructing a joint knowledge base consisting of a standard archival storage process sub-base, a historical storage scenario information sub-base, and a storage device resource information sub-base; constructing a demand-based deep optimization reasoning model and performing deep retrieval of the knowledge base using a fast association index; acquiring archival and scenario data in real time and calculating the current environmental security level and the probability of archival damage risk using a security assessment model; generating an initial assessment score and storage strategy based on the above assessment results using a comprehensive assessment model; performing secondary matching retrieval through a reasoning model to obtain an optimized storage allocation strategy; and finally dynamically adjusting storage resource allocation based on the optimized storage allocation strategy and real-time device status. This invention achieves intelligent generation and adaptive adjustment of archival storage strategies, improving storage efficiency and security.

[0005] To achieve the above objectives, the present invention provides the following technical solution:

[0006] Intelligent storage allocation methods for archives include:

[0007] S1. Construct a joint knowledge base for archival storage strategies; the joint knowledge base for archival storage strategies includes a sub-base of standard archival storage procedures, a sub-base of historical storage scenario information, and a sub-base of storage device resource information.

[0008] S2. Construct and train a demand-based deep optimization reasoning model, configure it in the joint knowledge base of file storage strategies, and combine it with a preset fast association index to perform deep association retrieval of each sub-base of the corresponding knowledge base; the fast association index is constructed from the historical index of the storage allocation strategy corresponding to the file storage strategy with the lowest risk probability and the lowest latency under the same scenario;

[0009] S3. Obtain the data of the archive to be stored and the current storage scenario data. Use the configured security assessment model to assess the security level of the storage environment and the risk of archive damage, and obtain the security level of the current storage scenario and the probability of archive risk.

[0010] S4. Based on the security level of the current storage scenario and the probability of file risk, use a comprehensive evaluation model to conduct a comprehensive evaluation of the file entity and generate an initial evaluation score and an initial storage allocation strategy.

[0011] S5. Based on the initial evaluation score and the initial storage allocation strategy, a second deep matching retrieval is performed from the joint knowledge base of archive storage strategies through the demand deep optimization reasoning model to obtain the optimized storage allocation strategy for archives under the current storage scenario.

[0012] S6. Based on the storage device resource configuration and actual device status in the optimized storage allocation strategy, determine the resource configuration differences and dynamically adjust the storage resources in combination with real-time device information.

[0013] Specifically, the construction steps of the standard archive storage process sub-library include:

[0014] S101. Based on the acquired data on different file types, the corresponding standard storage procedure data, and the frequency data of standard storage procedure calls, construct the input sequence of the standard storage procedure sub-library. ,in This represents the type identifier of the i-th type of file. This represents the basic metadata information of the i-th type of file. This represents the j-th operation step in the standard storage process of the i-th type of archive. This indicates the historical frequency of the i-th standard storage procedure being retrieved or invoked;

[0015] S102. Input the standard storage process sub-library input sequence into the entity extraction algorithm to obtain the archive standard storage process triplet;

[0016] S103. Based on the standard storage process triplet, the archive type is taken as the first-level node, the call frequency corresponding to each standard storage process is taken as the second-level node, and the basic metadata of the archive and the standard storage process are taken as the third-level nodes.

[0017] S104. Based on the operation steps of the standard storage process for each type of archive, divide them into n sub-steps according to the order of execution. Store the divided sub-steps in order to each sub-node under the corresponding third-level node of the corresponding type of archive. Each sub-node is hierarchically mounted according to the operation order.

[0018] S105. Based on the operation steps stored in each hierarchical mounting sub-node, the similarity generality of the operation steps stored in each sub-node under the third-level node of different file types is calculated by a similarity detection algorithm. Sub-nodes under different file types with similarity generality greater than a set threshold are linked according to the similarity generality between each pair of nodes to obtain a similar operation general chain. The similar operation general chain is then used to connect the corresponding sub-nodes under the third-level node.

[0019] S106. Utilize the frequency of calls to the corresponding standard archive storage process to construct a priority retrieval factor for the corresponding standard archive storage process, and embed the priority retrieval factor into the third-level node of the corresponding standard archive storage process.

[0020] Specifically, the construction steps of the historical storage scenario information sub-database include:

[0021] S111. Based on the storage strategies used in different historical storage scenarios, storage results and problem records, storage efficiency and completion, security level of the corresponding storage scenario, deviation of the corresponding storage strategy from the standard process, and storage device resource configuration status in the corresponding scenario, construct the scenario information sub-database input, and obtain the corresponding scenario information triplet through the same process in S102.

[0022] S112. Based on the scenario information triplet, the corresponding type of file is used as a first-level node, the different storage scenarios experienced by the corresponding type of file are used as a second-level node set, and the security level of the corresponding scenario is stored in the corresponding second-level child node in the second-level node set.

[0023] S113. Treat the storage technology process as a third-level node under the second-level sub-node of the corresponding scenario, and perform hierarchical mounting of storage through the same process as S104, storing the deviation between the storage strategy and the standard process in the corresponding hierarchical mounting storage sub-node.

[0024] S114. Store the storage results and problem records, storage efficiency and completion rate, and storage device resource configuration status in the corresponding scenario to the last-level node.

[0025] S115. Based on the deviations of the storage strategy and standard process, as well as the storage efficiency and completion rate, construct the first link factor between the historical storage scenario information sub-database and the standard storage process sub-database; construct the second link factor between the historical storage scenario information sub-database and the storage device resource information sub-database through storage device resource configuration information and storage efficiency and completion rate.

[0026] Specifically, the steps for constructing the storage device resource information sub-database include:

[0027] S121. Based on the acquired resource configuration data, the device grouping information of each type of file in different storage scenarios, the basic information of the storage devices in each group, the adaptability of each device to the standard storage process of various types of files, the number and efficiency of the device in processing professional type files, the number and efficiency of the device in processing non-professional type files, and the current operating status information of the device, construct the input sequence of the device resource information sub-database.

[0028] S122. Obtain the corresponding device information triplet through the same process as S102. Save the device grouping information of each type of file in the storage scenario to the first-level node. Save the basic information of each device to the parallel sub-nodes in the second-level node set. Save the device's adaptability to the standard storage process of various types of files, the number of times and efficiency of processing professional and / or non-professional types of files, and the current operating status information of the device to the third-level node under the corresponding second-level sub-node.

[0029] S123. Calculate the degree of job substitution among all devices based on the number of times and efficiency of the devices in processing non-professional type files.

[0030] S124. Set a job substitution threshold. When the job substitution degree between two devices is greater than the job substitution threshold, establish a link between the corresponding secondary sub-nodes of the two devices using the corresponding job substitution degree.

[0031] Specifically, the steps for constructing the joint knowledge base for archive storage strategies include:

[0032] S131. Based on the first link factor and the priority retrieval factors stored in each third-level node of the archive standard storage process sub-library, construct the first deep search connection between nodes in the archive standard storage process sub-library and the historical storage scenario information sub-library.

[0033] S132. Using the second link factor and the degree of device work substitution, construct a second deep search connection between nodes in the storage device resource information sub-database and the historical storage scenario information sub-database.

[0034] S133. Based on the first deep search connection and the second deep search connection, and combined with the topology algorithm, perform deep linking on the archive standard storage process sub-database, the historical storage scenario information sub-database, and the storage device resource information sub-database to obtain a joint knowledge base for archive storage strategies.

[0035] Specifically, the steps for obtaining the security level and file risk probability of the current storage scenario include:

[0036] S301. Receive digitized metadata of the archive to be stored through the system interface, including the physical type of the carrier, the security level setting, and the planned access frequency; and simultaneously receive real-time continuous monitoring data streams from temperature and humidity sensors, power monitoring modules, network probes, and access control systems.

[0037] S302. Initialize the hidden Markov model structure, define the hidden state as discrete security level labels, and define the observation state as a set of preprocessed environmental monitoring event codes.

[0038] S303. Access the historical storage scenario information sub-database, traverse all historical state records, and count the number of direct transitions between hidden states; for each hidden state, calculate the frequency ratio of its transitions to other states, and generate a state transition probability matrix.

[0039] S304. Traverse the historical scene data, and for each hidden security level, count the number of times each environmental monitoring event occurs under the corresponding hidden security level; calculate the proportion of the number of occurrences of each environmental monitoring event to the total number of all events under the corresponding hidden security level, and generate the observation event probability matrix.

[0040] S305. Using the encoded sequence of environmental monitoring events within the current time window as input, call the Viterbi decoding algorithm to calculate the hidden state sequence;

[0041] S306. Extract the state label corresponding to the latest time point from the hidden state sequence output by the Viterbi algorithm, and assign the corresponding state label value to the current environmental security level.

[0042] S307. Read the transition probability values ​​related to the current and previous states in the state transition probability matrix, obtain the basic probability of file damage corresponding to the hidden security level in historical statistics, use the transition probability values ​​as dynamic weighting factors, perform weighted calculations on the basic probability of file damage, and obtain the corrected risk probability value.

[0043] S308. Write the final determined security level label value and the corrected risk probability value into the system evaluation result data table, and transmit them to the required depth optimization inference model to make real-time adjustments to the depth matching retrieval under the corresponding storage scenarios at different time points.

[0044] Specifically, the steps for dynamically adjusting storage resources in S6 include:

[0045] Based on the storage device resource configuration information and actual available device status information in the optimized storage allocation strategy, if the current number of actual available devices is less than the number required by the strategy, then by using the demand depth optimization inference model and the work substitution degree of the device configured in the storage device resource information sub-database, candidate storage devices that meet the work substitution degree threshold conditions and their corresponding work substitution degrees are retrieved from other device groups not specified by the current strategy.

[0046] If the number of supplementary devices is 1, the candidate device with the largest product of job substitution degree and storage operation adaptability is selected for allocation; if the number of supplementary devices is greater than 1, multiple candidate devices are selected in descending order of the product of job substitution degree and storage operation adaptability for allocation, and the adjusted storage device resource configuration is finally obtained.

[0047] Specifically, the training steps for a demand-deep optimization inference model include:

[0048] S201. Construct a demand depth optimization reasoning model based on the pre-trained Bayesian network reasoning model, and construct the input sequence of the demand depth optimization reasoning model according to the file storage requirements under the current storage scenario, the information saved by each sub-node in the file storage strategy joint knowledge base, the first depth search connection and the second depth search connection.

[0049] S202. Utilize the storage efficiency and data integrity metrics corresponding to the optimized storage allocation strategy under each storage scenario to construct a fine-tuned pre-training loss function for the demand-deep optimization inference model.

[0050] S203. Set the fine-tuning pre-training threshold, input the input sequence and the fine-tuning pre-training loss function into the demand depth optimization inference model for fine-tuning training, and obtain the fine-tuned demand depth optimization inference model.

[0051] S204. Real-time acquisition of the execution effect data of file storage strategies in various scenarios, including actual storage efficiency and data integrity indicators, and feeding the real-time execution effect data back to the requirement-deep optimization inference model for online adaptive fine-tuning and optimization of the model.

[0052] The intelligent archive storage and allocation system includes: a knowledge base module and a deep linking module;

[0053] The knowledge base module includes a data processing unit and a knowledge base construction unit;

[0054] The data processing unit is used to acquire and preprocess different types of archive data, corresponding standard storage process data, and storage strategies, operation processes and equipment resource configuration data under different historical storage scenarios.

[0055] The knowledge base construction unit is used to construct a joint knowledge base for archive storage strategies using preprocessed data and knowledge graph algorithms.

[0056] The deep linking module is used to construct and train a demand-deep optimization reasoning model based on the knowledge base node reasoning algorithm, configure the demand-deep optimization reasoning model into the joint knowledge base of the archive storage strategy, and perform deep association retrieval of each sub-base through the storage information distribution node status.

[0057] Specifically, the intelligent archive storage allocation system also includes an initial evaluation module and a deep search configuration module;

[0058] The initial assessment module includes a scenario assessment unit and an initial file assessment unit;

[0059] The scenario assessment unit is used to acquire scenario data to be stored, and to assess the security level of the current storage environment and the probability of file damage risk through the configured security assessment model, and output the security level of the current storage scenario and the probability of file risk.

[0060] The initial file assessment unit is used to comprehensively assess the current file to be stored through a comprehensive assessment model under the security level and file risk probability conditions corresponding to the current storage scenario, and generate an initial assessment score and an initial storage allocation strategy.

[0061] The deep search configuration module includes a deep search unit and a resource adjustment unit;

[0062] The deep search unit is used to perform a second deep matching retrieval of the storage strategy in the current scenario from the joint knowledge base of archive storage strategies, based on the initial evaluation score and initial storage allocation strategy of the archive under the current storage scenario security level and archive risk probability conditions, through the demand deep optimization reasoning model, to obtain the optimized storage allocation strategy of the archive in the corresponding scenario.

[0063] The resource adjustment unit is used to determine the resource configuration differences based on the device resource configuration and actual device status in the optimized storage allocation strategy, and to dynamically allocate and adjust storage resources in combination with real-time device information.

[0064] Compared with the prior art, the beneficial effects of the present invention are:

[0065] This invention addresses the shortcomings of existing technologies by constructing a joint knowledge base for archival storage strategies that integrates multi-dimensional information and incorporates a demand-based deep optimization reasoning model. This enables the intelligent generation and dynamic optimization of storage strategies. The system can automatically generate initial storage strategies based on real-time environmental security assessments and archival risk probabilities, and then use a secondary deep search to match the lowest-risk and most efficient optimized allocation scheme. Finally, it completes precise scheduling and adaptive adjustment of resources based on real-time device status, significantly improving the intelligence level of archival storage management. While ensuring the security and integrity of archives, it greatly improves the utilization efficiency of storage resources and the system response speed. Attached Figure Description

[0066] Figure 1 This is a flowchart of the intelligent file storage allocation method according to Embodiment 1 of the present invention;

[0067] Figure 2 This is a block diagram of the intelligent archive storage and allocation system according to Embodiment 2 of the present invention. Detailed Implementation

[0068] Example 1:

[0069] Please see Figure 1 The present invention provides an embodiment of an intelligent archive storage allocation method, the steps of which include:

[0070] S1. Construct a joint knowledge base for archival storage strategies; the joint knowledge base for archival storage strategies includes a sub-base of standard archival storage procedures, a sub-base of historical storage scenario information, and a sub-base of storage device resource information.

[0071] S2. Construct and train a demand-based deep optimization reasoning model, configure it in the joint knowledge base of file storage strategies, and combine it with a preset fast association index to perform deep association retrieval of each sub-base of the corresponding knowledge base; the fast association index is constructed from the historical index of the storage allocation strategy corresponding to the file storage strategy with the lowest risk probability and the lowest latency under the same scenario;

[0072] It should be further explained that the construction process of the fast association index in this embodiment includes:

[0073] S211. Access the historical storage scenario information sub-database and extract all historical storage task records according to the preset time range. Each record should include: storage policy identifier, file type code, storage scenario feature vector, record archiving timestamp, total storage operation delay time, and file corruption identifier.

[0074] S212. Based on the archive type code and storage scenario feature vector, perform cluster analysis to divide the historical storage task records into several groups with the same features or similarity greater than the preset similarity threshold.

[0075] S213. For all records in each group, classify them according to their storage strategy identifier, and calculate the arithmetic mean of storage operation latency and the file corruption rate of all records under the same strategy; the file corruption rate is the ratio of the number of corrupted records to the total number of records under the strategy.

[0076] S214. For each storage strategy within each group, normalize the arithmetic mean of its storage operation latency and the file corruption rate, and calculate the comprehensive score of each strategy using a weighted geometric mean.

[0077] S215. Select the storage strategy with the highest comprehensive score within the same group and determine it as the optimal storage allocation strategy under this characteristic condition.

[0078] S216. Extract the complete policy index identifier and storage path corresponding to the optimal storage allocation policy in the archive standard storage process sub-database or the historical storage scenario information sub-database.

[0079] S217. Concatenate and hash the file type encoding and storage scene feature vector to generate a unique grouping feature key;

[0080] S218. Associate the grouping feature key with the corresponding optimal storage strategy index identifier and storage path to construct a key-value pair and store it in the hash mapping table of the fast association index.

[0081] S219. To quickly associate indexes, a cache management mechanism based on the Least Recently Used algorithm is established, and index items are periodically traversed and updated. The optimal storage allocation strategy association is recalculated and updated based on newly added historical scenario data.

[0082] S3. Obtain the data of the archive to be stored and the current storage scenario data. Use the configured security assessment model to assess the security level of the storage environment and the risk of archive damage, and obtain the security level of the current storage scenario and the probability of archive risk.

[0083] S4. Based on the security level of the current storage scenario and the probability of file risk, use a comprehensive evaluation model to conduct a comprehensive evaluation of the file entity and generate an initial evaluation score and an initial storage allocation strategy.

[0084] S5. Based on the initial evaluation score and the initial storage allocation strategy, a second deep matching retrieval is performed from the joint knowledge base of archive storage strategies through the demand deep optimization reasoning model to obtain the optimized storage allocation strategy for archives under the current storage scenario.

[0085] S6. Based on the storage device resource configuration and actual device status in the optimized storage allocation strategy, determine the resource configuration differences and dynamically adjust the storage resources in combination with real-time device information.

[0086] It should be further explained that the construction steps of the standard archive storage process sub-library in this embodiment include:

[0087] S101. Based on the acquired data of different file types, the corresponding standard storage procedure data, and the frequency data of the standard storage procedure being called, construct the input sequence of the standard storage procedure sub-library. ,in This represents the type identifier of the i-th type of file. This represents the basic metadata information of the i-th type of file. This represents the j-th operation step in the standard storage process of the i-th type of archive. This indicates the historical frequency of the i-th standard storage procedure being retrieved or invoked;

[0088] S102. Input the standard storage process sub-library input sequence into the entity extraction algorithm to obtain the archive standard storage process triplet;

[0089] S103. Based on the standard storage process triplet, the archive type is taken as the first-level node, the call frequency corresponding to each standard storage process is taken as the second-level node, and the basic metadata of the archive and the standard storage process are taken as the third-level nodes.

[0090] S104. Based on the specific operation steps of the standard storage process for each type of archive, divide it into n sub-steps according to the execution order. Store the divided sub-steps in order to each sub-node under the corresponding third-level node of the corresponding type of archive. Each sub-node is hierarchically mounted according to the operation order.

[0091] S105. Based on the operation steps stored in each hierarchical mounting sub-node, the similarity generality of the operation steps stored in each sub-node under the third-level node of different file types is calculated by a similarity detection algorithm. Sub-nodes under different file types with similarity generality greater than a set threshold are linked according to the similarity generality between each pair of nodes to obtain a similar operation general chain. The similar operation general chain is then used to connect the corresponding sub-nodes under the third-level node.

[0092] S106. Utilize the frequency of calls to the corresponding standard archive storage process to construct a priority retrieval factor for the corresponding standard archive storage process, and embed the priority retrieval factor into the third-level node of the corresponding standard archive storage process.

[0093] It should be further explained that the construction steps of the historical storage scenario information sub-database in this embodiment include:

[0094] S111. Based on the specific storage strategies used in different historical storage scenarios, storage results and problem records, storage efficiency and completion, security level of the corresponding storage scenario, deviation between the specific strategies and standard processes, and storage device resource configuration status in the corresponding scenario, construct the scenario information sub-database input, and obtain the corresponding scenario information triplet through the same process in S102.

[0095] S112. Based on the scenario information triplet, the corresponding type of file is used as a first-level node, the different storage scenarios experienced by the corresponding type of file are used as a second-level node set, and the security level of the corresponding scenario is stored in the corresponding second-level child node in the second-level node set.

[0096] S113. The specific storage technology process is used as a third-level node under the second-level sub-node of the corresponding scenario, and the storage is mounted hierarchically through the same process as S104. The deviation between the specific storage strategy and the standard process is stored in the corresponding hierarchical mounting storage sub-node.

[0097] S114. Store the storage results and problem records, storage efficiency and completion rate, and storage device resource configuration status in the corresponding scenario to the last-level node.

[0098] S115. Based on specific storage strategies, deviations from standard processes, and storage efficiency and completion rate, construct the first link factor between the historical storage scenario information sub-database and the standard storage process sub-database; construct the second link factor between the historical storage scenario information sub-database and the storage device resource information sub-database through storage device resource configuration information and storage efficiency and completion rate.

[0099] It should be further explained that the construction steps of the storage device resource information sub-database in this embodiment include:

[0100] S121. Based on the acquired resource configuration data, the device grouping information of each type of file in different storage scenarios, the basic information of the storage devices in each group, the adaptability of each device to the standard storage process of various types of files, the number and efficiency of the device in processing professional type files, the number and efficiency of the device in processing non-professional type files, and the current operating status information of the device, construct the input sequence of the device resource information sub-database.

[0101] S122. Obtain the corresponding device information triplet through the same process as S102. Save the device grouping information of each type of file in the storage scenario to the first-level node. Save the basic information of each device to the parallel sub-nodes in the second-level node set. Save the device's adaptability to the standard storage process of various types of files, the number of times and efficiency of processing professional and / or non-professional types of files, and the current operating status information of the device to the third-level node under the corresponding second-level sub-node.

[0102] S123. Calculate the degree of job substitution among all devices based on the number of times and efficiency of the devices in processing non-professional type files.

[0103] S124. Set a job substitution threshold. When the job substitution degree between two devices is greater than the job substitution threshold, establish a link between the corresponding secondary sub-nodes of the two devices using the corresponding job substitution degree.

[0104] It should be further explained that the construction steps of the joint knowledge base for archive storage strategy in this embodiment include:

[0105] S131. Based on the first link factor and the priority retrieval factors stored in each third-level node of the archive standard storage process sub-library, construct the first deep search connection between nodes in the archive standard storage process sub-library and the historical storage scenario information sub-library.

[0106] It should be further explained that the specific construction process of the first depth search connection in this embodiment includes:

[0107] S1311. Read the first link factor stored in the historical storage scenario information sub-database. The first link factor includes the standard process deviation coefficient and the storage efficiency weight value.

[0108] S1312. Obtain the priority retrieval factor stored in each third-level node of the standard archive storage process sub-library. This factor is a value calculated based on the process call frequency.

[0109] S1313. Multiply the standard process deviation coefficient in the first linking factor with the priority retrieval factor of the corresponding standard storage process node by weight to obtain the correlation strength value between nodes.

[0110] S1314. Set the association strength threshold, and link the archive standard storage process sub-database node with the association strength value greater than the association strength threshold to the historical storage scenario information sub-database node in both directions to form the first deep search connection.

[0111] S132. Using the second link factor and the degree of device work substitution, construct a second deep search connection between nodes in the storage device resource information sub-database and the historical storage scenario information sub-database.

[0112] It should be further explained that the specific construction process of the second depth search connection in this embodiment includes:

[0113] S1321. Extract the second link factor from the specified field of the historical storage scenario information sub-database. The factor is a structured data object containing device resource configuration status code and storage completion index.

[0114] S1322. Query the device node attribute table of the storage device resource information sub-database to obtain the work substitution degree numerical field;

[0115] S1323. Input the storage completion index and the device work substitution degree value into the weighted calculation unit, and use the formula: Device-scenario correlation degree = storage completion index × weight C + device work substitution degree × weight D, where weight C and weight D are preset empirical coefficients;

[0116] S1324. Compare the calculated device-scene correlation value with a preset threshold. When the correlation value is greater than the threshold, establish a second deep search connection between the storage device resource information sub-database node and the historical storage scene information sub-database node.

[0117] S133. Based on the first deep search connection and the second deep search connection, and combined with the topology algorithm, perform deep linking on the archive standard storage process sub-database, the historical storage scenario information sub-database, and the storage device resource information sub-database to obtain a joint knowledge base for archive storage strategies.

[0118] It should be further explained that the construction and implementation process of the joint knowledge base for archive storage strategy in this embodiment includes:

[0119] S1331. Traverse all node link relationship tables and collect all node link records between the first depth search connection and the second depth search connection.

[0120] S1332. Starting from any node in any sub-database, perform a breadth-first traversal algorithm, recording the sequence of visited nodes and the link relationships between nodes during the traversal.

[0121] S1333. Create a unified addressing mapping table. Each record in the table contains: a globally unique identifier for the node, the identifier of the sub-library to which it belongs, the local identifier of the node in the atom, and a list of global identifiers of associated nodes.

[0122] S1334. Construct a hash index structure based on the unified addressing mapping table, where the key is the globally unique identifier of the node and the value is a list of globally unique identifiers of all cross-sub-database nodes associated with the node.

[0123] S1335. Serialize the unified addressing mapping table and hash index structure into binary data format and store them in a distributed file system to complete the construction of the joint knowledge base for archive storage strategies.

[0124] This process achieves unified knowledge management and deep correlation analysis of multi-source heterogeneous data by constructing a joint knowledge base for archival storage strategies. It establishes a three-layer knowledge system comprising a standard archival storage process sub-base, a historical storage scenario information sub-base, and a storage device resource information sub-base, and employs entity extraction and tripletization techniques to transform various data types into structured knowledge nodes. Deep search connections between nodes are built using priority retrieval factors and linking factors, combined with topology algorithms to form a cross-sub-base association network, enabling the system to quickly retrieve the optimal storage strategy. This knowledge representation method not only achieves standardized management of storage processes but also continuously optimizes the decision-making process through historical scenario data. An evaluation mechanism based on device substitutability and collaborative work efficiency ensures the rationality and efficiency of storage resource allocation. By updating the knowledge base in real time and dynamically adjusting strategies, the system can adapt to constantly changing storage needs and environmental conditions, significantly improving the security, reliability, and resource utilization of archival storage.

[0125] It should be further explained that the training steps for the depth-optimized inference model in this embodiment include:

[0126] S201. Construct a demand depth optimization reasoning model based on the pre-trained Bayesian network reasoning model, and construct the input sequence of the demand depth optimization reasoning model according to the file storage requirements under the current storage scenario, the information saved by each sub-node in the file storage strategy joint knowledge base, the first depth search connection and the second depth search connection.

[0127] It should be further explained that the process of constructing the input sequence of the demand depth optimization inference model in this embodiment includes:

[0128] S2011. Load the structure definition file and parameter file of the pre-trained Bayesian network inference model from persistent storage, and use the deep learning framework to initialize the network weights of the inference model to optimize the required depth.

[0129] S2012. Receive the current storage request through the system data interface, parse the file type, size, security level, and expected retention time attributes contained therein, perform one-hot encoding and normalization on these attributes, and combine them to generate a fixed-dimensional file storage requirement feature vector.

[0130] S2013. Access the graph database interface of the joint knowledge base of archive storage strategy, traverse all nodes in the archive standard storage process sub-base, historical storage scenario information sub-base and storage device resource information sub-base, and use graph embedding algorithm to convert the attribute information of each node into a 128-dimensional feature vector representation.

[0131] S2014. Retrieve the relationship data between the first depth search connection and the second depth search connection through the query interface of the knowledge base, extract the connection weights and relationship types between the associated nodes, and construct the association weight matrix.

[0132] S2015. The archive storage requirement feature vector, the 128-dimensional feature vector representation of all nodes, and the associated weight matrix data are concatenated along the feature dimension to generate a high-dimensional comprehensive feature vector.

[0133] S2016. Perform Z-score standardization on the comprehensive feature vector to eliminate differences in feature dimensions and generate an input sequence that the model can process.

[0134] S202. Utilize the storage efficiency and data integrity metrics corresponding to the optimized storage allocation strategy under each storage scenario to construct a fine-tuned pre-training loss function for the demand-deep optimization inference model.

[0135] It should be further explained that the construction process of the pre-training loss function in this embodiment includes:

[0136] S2021. Batch read the storage policy execution records of the most recent year from the historical storage scenario information sub-database.

[0137] S2022. For each record, calculate the quantitative value of the storage efficiency index, including data transfer rate, storage operation completion time and resource utilization, and weight these sub-indicators to form a single efficiency score.

[0138] S2023. For each record, calculate the quantitative value of the data integrity index, including the verification and validation pass rate, data block integrity rate and error recovery success rate, and weight these sub-indicators to form a single integrity score.

[0139] S2024. Combine the storage efficiency score and the data integrity score according to the preset weight ratio to construct a multi-objective loss function, and calculate the overall loss value by weighted summation.

[0140] S203. Set the fine-tuning pre-training threshold, input the input sequence and the fine-tuning pre-training loss function into the demand depth optimization inference model for fine-tuning training, and obtain the fine-tuned demand depth optimization inference model.

[0141] It should be further explained that the fine-tuning process of the demand depth optimization inference model in this embodiment includes:

[0142] S2031. Set the iteration stopping threshold for fine-tuning training to 0.001 based on the model complexity, which means that training will stop when the decrease in the loss function value is less than this threshold.

[0143] S2032. Input the standardized input sequence into the forward propagation network of the demand depth optimization inference model, and calculate the multi-objective loss function value.

[0144] S2033. Use the Adam optimization algorithm to perform backpropagation, calculate the gradient of the loss function with respect to the weights of each layer of the model, and update the model parameters.

[0145] S2034. After each training cycle, compare the difference between the current loss value and the loss value of the previous cycle. Terminate the training process when the difference is less than 0.001.

[0146] S2035. Serialize the fine-tuned model weight parameters and network structure definition into Protocol Buffer format and persist them to the model repository.

[0147] S204. Real-time acquisition of the execution effect data of file storage strategies in various scenarios, including actual storage efficiency and data integrity indicators, and feeding the real-time execution effect data back to the requirement-deep optimization inference model for online adaptive fine-tuning and optimization of the model.

[0148] It should be further explained that the steps for online adaptive fine-tuning and optimization of the model in this embodiment include:

[0149] S2041. Deploy a RESTful API interface to collect and store policy execution data in real time, including system monitoring metrics and business metrics;

[0150] S2042. Perform missing value imputation, outlier removal, and standardization on the collected raw data to make the data meet the model input requirements;

[0151] S2043. Set up a data cache queue to automatically trigger the online learning process when the amount of collected data reaches 1000 records.

[0152] S2044. The Mini-batch gradient descent algorithm is used to incrementally train the model with a batch size of 32 and the learning rate is set to 0.0001.

[0153] S2045. Use hot update technology to update model parameters without interrupting service and keep the system running continuously.

[0154] This process achieves intelligent generation and dynamic optimization of archival storage strategies by constructing a deep-optimization inference model based on demand and employing a multi-stage training optimization strategy. Initialization is performed using a pre-trained Bayesian network model, combined with deep fusion of archival storage demand features and knowledge base node information. High-quality input sequences are generated through feature vector concatenation and standardization. A multi-objective loss function is used to simultaneously optimize storage efficiency and data integrity metrics, enabling the model to learn the optimal storage strategy decision-making pattern. An adaptive fine-tuning training mechanism is employed to continuously optimize model parameters based on real-time feedback data, ensuring the system can quickly adapt to changing storage environments. Online learning and hot-update technologies are used to continuously improve model performance without service interruption. This deep learning-based inference model not only improves the accuracy and efficiency of storage strategy formulation but also significantly enhances the intelligence level and decision-making quality of the archival storage system by continuously learning and optimizing from historical data, ultimately achieving efficient utilization of storage resources and comprehensive protection of archival security.

[0155] It should be further explained that the steps for obtaining the security level of the current storage scenario and the probability of file risk in this embodiment include:

[0156] S301. Receive digitized metadata of the archive to be stored through the system interface, including the physical type of the carrier, the security level setting, and the planned access frequency; and simultaneously receive real-time continuous monitoring data streams from temperature and humidity sensors, power monitoring modules, network probes, and access control systems.

[0157] S302. Initialize the Hidden Markov Model structure, defining the hidden states as discrete security level labels and the observed states as a preprocessed set of environmental monitoring event codes. It should be further noted that in this embodiment, during the initialization of the Hidden Markov Model, the hidden state set is defined as a finite number of discrete security level labels, for example, d1 = low risk, d2 = medium risk, and d3 = high risk, used to represent the internal true security level of the system that cannot be directly observed. The observed state set is defined as all preprocessed and numerically encoded environmental monitoring events, for example, V1 = temperature over-limit alarm, V2 = normal humidity, and V3 = network interruption event, used to describe visible data that can be directly collected by sensors and input into the model. The hidden states are used to characterize the evolution of the system's internal security posture, while the observed states are used to transform the raw environmental data into feature inputs that the model can process. Together, they constitute the core framework of the model, providing a computational foundation for subsequent dynamic security inference based on sequence data.

[0158] S303. Access the historical storage scenario information sub-database, traverse all historical state records, and count the number of direct transitions between hidden states; for each hidden state, calculate the frequency ratio of its transitions to other states, and generate a state transition probability matrix.

[0159] S304. Traverse the historical scene data, and for each hidden security level, count the number of times each environmental monitoring event occurs under the corresponding hidden security level; calculate the proportion of the number of occurrences of each environmental monitoring event to the total number of all events under the corresponding hidden security level, and generate the observation event probability matrix.

[0160] S305. Using the encoded sequence of environmental monitoring events within the current time window as input, call the Viterbi decoding algorithm to calculate the hidden state sequence most likely to produce the observation sequence under the given model parameters.

[0161] S306. Extract the state label corresponding to the latest time point from the hidden state sequence output by the Viterbi algorithm, and assign the corresponding state label value to the current environmental security level.

[0162] S307. Read the transition probability values ​​related to the current and previous states in the state transition probability matrix; obtain the basic probability of file damage corresponding to the hidden security level in historical statistics; use the transition probability values ​​as dynamic weighting factors to perform weighted calculations on the basic probability of file damage to obtain the corrected risk probability value.

[0163] S308. Write the final determined security level label value and the corrected risk probability value into the system evaluation result data table, and transmit it to the required depth optimization inference model to make real-time adjustments to the depth matching retrieval under the corresponding storage scenarios at different time points.

[0164] This process, by receiving digitized metadata of the archives to be stored and multi-dimensional real-time monitoring data streams, achieves a comprehensive perception of the basic attributes and dynamic environment of the storage scenario, providing rich and timely input for security assessment. By introducing a Hidden Markov Model (HMM), using discrete security levels as hidden states and pre-processed environmental monitoring events as observed states, it accurately captures the correlation between the inherent security posture that cannot be directly observed and the perceptible environmental data, aligning with the dynamic evolution of the storage scenario's security status. The state transition probability matrix and observed event probability matrix generated based on historical data provide a solid statistical foundation for the model, making the quantitative analysis of state transitions and event correlations more reliable. Using the Viterbi decoding algorithm, the most likely hidden state sequence can be inferred from the time-series environmental event sequence, ensuring that the determination of the corresponding hidden security level conforms to the actual evolutionary pattern. Dynamically weighting and correcting the basic probability of archive damage through state transition probabilities ensures that risk probability calculation is based on historical statistical patterns while responding to state changes in real time, improving the accuracy and timeliness of risk assessment. Applying the results to the demand-based deep optimization inference model allows for real-time adjustment of the deep matching retrieval of storage scenarios, ultimately achieving dynamic and precise control over the security of archive storage and enhancing the adaptability of the storage system to environmental changes and the reliability of archive protection.

[0165] It should be further explained that the process of generating the initial evaluation score and the initial storage allocation strategy in this embodiment includes:

[0166] S401. Obtain the security level quantification value and file damage risk probability value of the current storage scenario, and read the metadata feature vector of the file to be stored, including file security level, media type, access frequency and expected retention period.

[0167] S402. Standardize and preprocess the input parameters, map the security level quantification value to the numerical range between 0 and 1 through linear transformation, transform the archive damage risk probability value using a logarithmic function and then perform minimum-maximum normalization processing, and perform one-hot encoding and feature scaling processing on the archive metadata feature vector.

[0168] S403. The initial evaluation score is calculated using a weighted linear combination algorithm. The weight coefficients for security level, archive damage risk probability, and archive value characteristics are set to negative values. After normalization, the sum of all weight coefficients is one.

[0169] S404. Based on the initial evaluation score, divide the storage priority range, set a first threshold and a second threshold, classify the files with an initial evaluation score higher than the first threshold as high priority storage category, classify the files with an initial evaluation score lower than the second threshold as low priority storage category, and classify the files with an initial evaluation score between the two thresholds as medium priority storage category.

[0170] S405: Assign storage devices with the highest security certification level to high-priority files, configure a real-time dual backup storage strategy, set a daily data integrity verification cycle, and allocate dedicated high-speed storage channels and priority network bandwidth.

[0171] S406. Assign storage devices with standard security levels to medium-priority archives, configure differential incremental backup policies, set weekly data integrity verification cycles, and allocate shared storage channels and standard network bandwidth.

[0172] S407. Assign storage devices with basic security levels to low-priority files, configure a monthly full backup policy, set a monthly data integrity verification cycle, and allocate shared storage channels and basic network bandwidth.

[0173] S408. Query the storage device resource information sub-database to obtain the real-time load indicators of each storage device, perform load balancing verification on the initially allocated storage devices, and automatically reallocate some files to storage devices with lower loads when the device load exceeds the preset load threshold.

[0174] S409. Generate a structured storage policy document, recording configuration information such as storage device identifier, physical storage path, backup policy parameters, and data verification cycle, and associate the initial evaluation score with the policy document and store it in the system decision database.

[0175] This process achieves comprehensive coverage of multi-dimensional parameters for storage assessment by acquiring the storage scenario security level, the probability value of archive damage risk, and the feature vector of archive metadata, providing a complete input foundation for subsequent decision-making. Standardized preprocessing of input parameters eliminates the dimensional differences between different types of data, enabling the security level, risk probability, and metadata features to be calculated on a unified scale, thus improving the scientific rigor of the assessment. Employing a weighted linear combination algorithm with appropriately set weight directions organically combines security risk factors with archive value characteristics, reflecting both the constraint of security risks on storage priority and highlighting the positive impact of archive value, making the initial assessment score more aligned with actual storage needs. The three-tiered storage priority based on the initial assessment score provides a clear basis for differentiated resource allocation: high-priority archives utilize high-security equipment and real-time dual backup strategies to strengthen the security of core archives; while the tiered backup and verification strategies for medium- and low-priority archives achieve precise resource allocation, avoiding waste caused by over-protection. Load balancing verification dynamically adjusts storage allocation to prevent equipment overload and ensure system stability. The generation and associated storage of structured storage strategy documents achieve traceability and standardized management of the decision-making process. Through precise assessment, tiered measures, and dynamic optimization, the overall process ensures the security of archives while achieving efficient utilization of storage resources, thereby improving the system's adaptability and management efficiency.

[0176] It should be further explained that the step of dynamically adjusting storage resources in S6 of this embodiment includes:

[0177] Based on the storage device resource configuration information and actual available device status information in the optimized storage allocation strategy, if the current number of actual available devices is less than the number required by the strategy, then by using the demand depth optimization inference model and the work substitution degree of the device configured in the storage device resource information sub-database, candidate storage devices that meet the work substitution degree threshold conditions and their corresponding work substitution degrees are retrieved from other device groups not specified by the current strategy.

[0178] If the number of devices to be added is 1, the candidate device with the largest product of job substitution degree and storage operation adaptability is selected for allocation; if the number of devices to be added is greater than 1, multiple candidate devices are selected in descending order of the product of job substitution degree and storage operation adaptability for allocation, and the adjusted storage device resource configuration is finally obtained.

[0179] It should be further explained that the more detailed implementation process of dynamically adjusting storage resources in this embodiment includes:

[0180] S601. Based on the resource configuration information such as storage device type, quantity, and performance indicators required in the optimized storage allocation strategy, compare it with the real-time device status information obtained through the device monitoring module to determine whether the currently available resources meet the strategy requirements; performance indicators include but are not limited to IOPS, bandwidth, and remaining capacity; real-time device status information includes but is not limited to device online or offline status, health, current load, and allocated task queue.

[0181] If the number or capacity of currently available devices is less than the number required by the policy, a resource replenishment process is triggered. Through a demand-based deep optimization inference model, a deep search is performed on the storage device resource information sub-database to query non-policy-specified device groups with the same or similar functions (e.g., all are cold storage servers or all are SSD cache devices). Based on the pre-set device replacement index in the storage device resource information sub-database, a list of candidate storage devices with a replacement degree higher than a set threshold is selected. At the same time, their specific replacement degree values ​​and real-time status parameters are obtained. The device replacement index is the degree of work substitution between devices.

[0182] S602. If the resource difference analysis results indicate that only one additional device is needed to meet the strategy requirements, the system enters the precise replacement process. This involves calculating the comprehensive priority score of each candidate device, specifically the product of the device's substitutability and its operational adaptability to the current storage process. A real-time load factor can be introduced for weighted adjustment, where the load factor is the reciprocal of the current load assessment score. The candidate device with the highest comprehensive priority score is selected and added to the resource allocation pool, and a device allocation instruction is generated. Simultaneously, the software environment or data index required by the target storage process is preloaded for this device to reduce switching latency.

[0183] S603. If the number of devices to be added is greater than 1, the system enters the batch allocation mode. Similarly, the comprehensive priority score of all candidate devices is calculated, and then they are sorted in descending order according to the score. According to the specific number required by the strategy, multiple candidate devices are selected sequentially from the top of the sorted list. At the same time, the collaborative working efficiency between devices needs to be considered. For example, avoid selecting device combinations that are across different data centers or have high network latency, so as to form the optimal supplementary device group and generate a batch allocation scheme.

[0184] S604. After determining the candidate devices, check whether these devices have been reserved or occupied by other high-priority tasks. If a resource allocation conflict occurs, arbitration is carried out based on the task priority and the weighted score of device substitutability. For occupied devices, try to negotiate with the task scheduling system to release resources or migrate tasks. If they cannot be released immediately, backtrack to step S602 or S603 and select the next candidate device in the priority list until the required number of devices are successfully allocated.

[0185] S605. Send control commands to the selected candidate storage devices to execute specific allocation operations, such as mounting specific volumes, allocating storage pool space, and configuring network paths. After confirming successful resource allocation, update the current status (e.g., marked as "allocated"), current load, and task association information of the relevant devices in the storage device resource information sub-database. At the same time, store the decision logic of this resource adjustment, the final selected device, and the execution effect (e.g., whether the policy requirements were successfully met) as a new historical scenario information in the historical storage scenario information sub-database for optimizing future reasoning and decision-making.

[0186] S606. After the device starts executing storage tasks, continuously monitor its actual performance indicators, such as read / write speed and latency, to ensure they meet standards. If the actual performance of a newly added device does not meet expectations, or if it causes new bottlenecks during operation, such as network congestion, this is used as a feedback signal. Dynamically trigger the demand-driven deep optimization inference model for small-scale re-optimization, fine-tune device configurations, or initiate a new round of resource adjustment processes, forming a closed-loop continuous optimization system.

[0187] S607. If resource adjustments involve the migration of stored files, such as from a faulty device to a newly allocated device, the system automatically triggers a data migration process. First, the integrity and consistency of the file data between the source and target devices are verified. Then, based on a strategy to minimize I / O interference, the file data blocks are migrated in batches. During the migration process, metadata updates and access routes are kept synchronized to ensure that the business is unaware of the migration. After the migration is completed, the storage association of the original device with the relevant files is removed.

[0188] S608. After allocation, the system will continuously collect the performance data of the newly allocated device in the storage task and compare it with the historical performance baseline of the device under the storage policy in the knowledge base. If the performance index continues to deviate from the baseline and exceeds the allowable threshold, a performance anomaly warning will be generated, triggering the root cause analysis process.

[0189] S609: The system integrates an energy efficiency monitoring module to calculate the energy consumption per unit storage capacity and the overall estimated cost of newly allocated equipment combinations in real time. This data is compared with the energy efficiency targets required by the strategy. If the data does not meet the requirements, the energy efficiency cost factor is incorporated into the next round of optimization decisions, and equipment that meets the performance requirements and has better energy efficiency is given priority.

[0190] S610. Before executing physical resource allocation, the optimized storage allocation strategy can be pre-verified in a simulation environment. By simulating real load, the performance, stability and risk probability of the strategy under expected pressure can be verified, and the strategy can be fine-tuned based on the simulation results to form the final executable solution.

[0191] S611. Use the full-link data of this resource adjustment (including decision input, execution process, and final performance results) as training samples to incrementally train the demand depth optimization inference model, optimize the accuracy of parameters such as equipment substitutability and operation adaptability, so that the model has the ability to continuously evolve and better adapt to future scenarios.

[0192] S612. In a multi-tenant environment, verify whether this resource allocation affects the preset resource quotas or service level agreements of other tenants; if there is a conflict, coordinate the allocation plan or negotiate with the relevant tenant systems according to the predefined business priority and fairness strategy to ensure that the resource allocation complies with the global management strategy.

[0193] This embodiment constructs a joint knowledge base for archival storage strategies that integrates multi-dimensional information, achieving structured storage and associated management of standard archival processes, historical scene information, and equipment resource status. Based on a pre-trained demand-driven deep optimization reasoning model, the system can perform deep association retrieval of the knowledge base, quickly matching the optimized storage allocation strategy with the lowest risk probability and least latency. Through a Hidden Markov Model, the system dynamically assesses the security level of the storage environment and the risk of archival damage, and dynamically adjusts resources based on real-time equipment status, significantly improving the intelligence level and decision-making accuracy of archival storage management. The system uses a multi-objective loss function to continuously optimize the model and introduces an online learning mechanism, enabling the model to adapt to environmental changes. Furthermore, by establishing equipment substitutability indicators and collaborative work efficiency evaluation mechanisms, precise scheduling and fault-tolerant management of storage resources are achieved. This system not only effectively ensures the security and integrity of archival storage but also significantly improves the utilization efficiency of storage resources and system response speed, providing a reliable coordination mechanism for resource allocation in a multi-tenant environment. Simultaneously, simulation pre-verification and performance baseline comparison functions further enhance the system's stability and reliability. Ultimately, the system forms a closed-loop continuous optimization system that can continuously learn and optimize from historical data to adapt to increasingly complex archival storage needs.

[0194] Example 2:

[0195] Please see Figure 2 Another embodiment of the present invention provides an intelligent archive storage and allocation system, comprising: a knowledge base module, a deep linking module, an initial evaluation module, and a deep search configuration module;

[0196] The knowledge base module is used to construct a joint knowledge base for archival storage strategies; the joint knowledge base for archival storage strategies includes a sub-base of standard archival storage procedures, a sub-base of historical storage scenario information, and a sub-base of storage device resource information.

[0197] The knowledge base module includes a data processing unit and a knowledge base construction unit;

[0198] The data processing unit is used to acquire and preprocess different types of archive data, corresponding standard storage process data, and storage strategies, operation processes and equipment resource configuration data under different historical storage scenarios.

[0199] The knowledge base construction unit is used to construct the sub-bases in the joint knowledge base of the archive storage strategy using preprocessed data and knowledge graph algorithms.

[0200] The deep linking module is used to construct and train a demand-deep optimization reasoning model based on the knowledge base node reasoning algorithm, configure the demand-deep optimization reasoning model into the joint knowledge base of the archive storage strategy, and perform deep association retrieval of each sub-base through the storage information distribution node status.

[0201] The initial assessment module performs a preliminary, multi-factor comprehensive analysis of current storage needs and the environment, and generates a basic storage strategy. Its core objective is to quickly respond to storage requests and make initial decisions based on currently known static and dynamic data.

[0202] The initial assessment module includes a scenario assessment unit and an initial file assessment unit;

[0203] The scenario assessment unit is used to acquire scenario data to be stored, and to assess the security level of the current storage environment and the probability of file damage risk through the configured security assessment model, and output the security level of the current storage scenario and the probability of file risk.

[0204] The initial file assessment unit is used to comprehensively assess the current file to be stored through a comprehensive assessment model under the security level and file risk probability conditions corresponding to the current storage scenario, and generate an initial assessment score and an initial storage allocation strategy.

[0205] The deep search configuration module is used to deeply optimize and accurately match the basic storage strategy output by the initial evaluation module, ensuring that the optimized strategy can be accurately implemented under the current actual physical resource conditions. Its core purpose is to pursue the optimal solution for storage efficiency, security, and resource utilization.

[0206] The deep search configuration module includes a deep search unit and a resource adjustment unit;

[0207] The deep search unit is used to perform a second deep matching retrieval of the storage strategy in the current scenario from the joint knowledge base of archive storage strategies, based on the initial evaluation score and initial storage allocation strategy of the archive under the current storage scenario security level and archive risk probability conditions, through the demand deep optimization reasoning model, to obtain the optimized storage allocation strategy of the archive in the corresponding scenario.

[0208] The resource adjustment unit is used to determine the resource configuration differences based on the device resource configuration and actual device status in the optimized storage allocation strategy, and to dynamically allocate and adjust storage resources in combination with real-time device information.

[0209] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments under the guidance of the present invention without departing from the spirit and scope of the claims. All of these variations are within the protection scope of the present invention.

Claims

1. An intelligent storage and allocation method for archives, characterized in that, include: S1. Construct a joint knowledge base for archive storage strategies; The joint knowledge base for archival storage strategies includes a sub-base of standard archival storage procedures, a sub-base of historical storage scenario information, and a sub-base of storage device resource information. S2. Construct and train a demand-based deep optimization reasoning model, configure it in the joint knowledge base of file storage strategies, and combine it with a preset fast association index to perform deep association retrieval of each sub-base of the corresponding knowledge base; the fast association index is constructed from the historical index of the storage allocation strategy corresponding to the file storage strategy with the lowest risk probability and the lowest latency under the same scenario; S3. Obtain the data of the archive to be stored and the current storage scenario data. Use the configured security assessment model to assess the security level of the storage environment and the risk of archive damage, and obtain the security level of the current storage scenario and the probability of archive risk. S4. Based on the security level of the current storage scenario and the probability of file risk, use a comprehensive evaluation model to conduct a comprehensive evaluation of the file entity and generate an initial evaluation score and an initial storage allocation strategy. S5. Based on the initial evaluation score and the initial storage allocation strategy, a second deep matching retrieval is performed from the joint knowledge base of archive storage strategies through the demand deep optimization reasoning model to obtain the optimized storage allocation strategy for archives under the current storage scenario. S6. Based on the storage device resource configuration and actual device status in the optimized storage allocation strategy, determine the resource configuration differences and dynamically adjust the storage resources in combination with real-time device information.

2. The intelligent archive storage and allocation method as described in claim 1, characterized in that, The construction steps of the standard archive storage process sub-library include: S101. Based on the acquired data on different file types, the corresponding standard storage procedure data, and the frequency data of standard storage procedure calls, construct the input sequence of the standard storage procedure sub-library. ,in This represents the type identifier of the i-th type of file. This represents the basic metadata information of the i-th type of file. This represents the j-th operation step in the standard storage process of the i-th type of archive. This indicates the historical frequency of the i-th standard storage procedure being retrieved or invoked; S102. Input the standard storage process sub-library input sequence into the entity extraction algorithm to obtain the archive standard storage process triplet; S103. Based on the standard storage process triplet, the archive type is taken as the first-level node, the call frequency corresponding to each standard storage process is taken as the second-level node, and the basic metadata of the archive and the standard storage process are taken as the third-level nodes. S104. Based on the operation steps of the standard storage process for each type of archive, divide them into n sub-steps according to the order of execution. Store the divided sub-steps in order to each sub-node under the corresponding third-level node of the corresponding type of archive. Each sub-node is hierarchically mounted according to the operation order. S105. Based on the operation steps stored in each hierarchical mounting sub-node, the similarity generality of the operation steps stored in each sub-node under the third-level node of different file types is calculated by a similarity detection algorithm. Sub-nodes under different file types with similarity generality greater than a set threshold are linked according to the similarity generality between each pair of nodes to obtain a similar operation general chain. The similar operation general chain is then used to connect the corresponding sub-nodes under the third-level node. S106. Utilize the frequency of calls to the corresponding standard archive storage process to construct a priority retrieval factor for the corresponding standard archive storage process, and embed the priority retrieval factor into the third-level node of the corresponding standard archive storage process.

3. The intelligent archive storage and allocation method as described in claim 2, characterized in that, The construction steps of the historical storage scenario information sub-database include: S111. Based on the storage strategies used in different historical storage scenarios, storage results and problem records, storage efficiency and completion, security level of the corresponding storage scenario, deviation of the corresponding storage strategy from the standard process, and storage device resource configuration status in the corresponding scenario, construct the scenario information sub-database input, and obtain the corresponding scenario information triplet through the same process in S102. S112. Based on the scenario information triplet, the corresponding type of file is used as a first-level node, the different storage scenarios experienced by the corresponding type of file are used as a second-level node set, and the security level of the corresponding scenario is stored in the corresponding second-level child node in the second-level node set. S113. Treat the storage technology process as a third-level node under the second-level sub-node of the corresponding scenario, and perform hierarchical mounting of storage through the same process as S104, storing the deviation between the storage strategy and the standard process in the corresponding hierarchical mounting storage sub-node. S114. Store the storage results and problem records, storage efficiency and completion rate, and storage device resource configuration status in the corresponding scenario to the last-level node. S115. Based on the deviations of the storage strategy and standard process, as well as the storage efficiency and completion rate, construct the first link factor between the historical storage scenario information sub-database and the standard storage process sub-database; construct the second link factor between the historical storage scenario information sub-database and the storage device resource information sub-database through storage device resource configuration information and storage efficiency and completion rate.

4. The intelligent archive storage and allocation method as described in claim 3, characterized in that, The steps for constructing the storage device resource information sub-database include: S121. Based on the acquired resource configuration data, the device grouping information of each type of file in different storage scenarios, the basic information of the storage devices in each group, the adaptability of each device to the standard storage process of various types of files, the number and efficiency of the device in processing professional type files, the number and efficiency of the device in processing non-professional type files, and the current operating status information of the device, construct the input sequence of the device resource information sub-database. S122. Obtain the corresponding device information triplet through the same process as S102. Save the device grouping information of each type of file in the storage scenario to the first-level node. Save the basic information of each device to the parallel sub-nodes in the second-level node set. Save the device's adaptability to the standard storage process of various types of files, the number of times and efficiency of processing professional and / or non-professional types of files, and the current operating status information of the device to the third-level node under the corresponding second-level sub-node. S123. Calculate the degree of job substitution among all devices based on the number of times and efficiency of the devices in processing non-professional type files. S124. Set a job substitution threshold. When the job substitution degree between two devices is greater than the job substitution threshold, establish a link between the corresponding secondary sub-nodes of the two devices using the corresponding job substitution degree.

5. The intelligent archive storage and allocation method as described in claim 4, characterized in that, The steps for constructing the joint knowledge base for the archive storage strategy include: S131. Based on the first link factor and the priority retrieval factors stored in each third-level node of the archive standard storage process sub-library, construct the first deep search connection between nodes in the archive standard storage process sub-library and the historical storage scenario information sub-library. S132. Using the second link factor and the degree of device work substitution, construct a second deep search connection between nodes in the storage device resource information sub-database and the historical storage scenario information sub-database. S133. Based on the first deep search connection and the second deep search connection, and combined with the topology algorithm, perform deep linking on the archive standard storage process sub-database, the historical storage scenario information sub-database, and the storage device resource information sub-database to obtain a joint knowledge base for archive storage strategies.

6. The intelligent archive storage and allocation method as described in claim 5, characterized in that, The steps for obtaining the security level and file risk probability of the current storage scenario include: S301. Receive digitized metadata of the archive to be stored through the system interface, including the physical type of the carrier, the security level setting, and the planned access frequency; and simultaneously receive real-time continuous monitoring data streams from temperature and humidity sensors, power monitoring modules, network probes, and access control systems. S302. Initialize the hidden Markov model structure, define the hidden state as discrete security level labels, and define the observation state as a set of preprocessed environmental monitoring event codes. S303. Access the historical storage scenario information sub-database, traverse all historical state records, and count the number of direct transitions between hidden states; for each hidden state, calculate the frequency ratio of its transitions to other states, and generate a state transition probability matrix. S304. Traverse the historical scene data, and for each hidden security level, count the number of times each environmental monitoring event occurs under the corresponding hidden security level; calculate the proportion of the number of occurrences of each environmental monitoring event to the total number of all events under the corresponding hidden security level, and generate the observation event probability matrix. S305. Using the encoded sequence of environmental monitoring events within the current time window as input, call the Viterbi decoding algorithm to calculate the hidden state sequence; S306. Extract the state label corresponding to the latest time point from the hidden state sequence output by the Viterbi algorithm, and assign the corresponding state label value to the current environmental security level. S307. Read the transition probability values ​​related to the current and previous states in the state transition probability matrix, obtain the basic probability of file damage corresponding to the hidden security level in historical statistics, use the transition probability values ​​as dynamic weighting factors, perform weighted calculations on the basic probability of file damage, and obtain the corrected risk probability value. S308. Write the final determined security level label value and the corrected risk probability value into the system evaluation result data table, and transmit them to the required depth optimization inference model to make real-time adjustments to the depth matching retrieval under the corresponding storage scenarios at different time points.

7. The intelligent archive storage and allocation method as described in claim 6, characterized in that, The steps for dynamically adjusting storage resources in S6 include: Based on the storage device resource configuration information and actual available device status information in the optimized storage allocation strategy, if the current number of actual available devices is less than the number required by the strategy, then by using the demand depth optimization inference model and the work substitution degree of the device configured in the storage device resource information sub-database, candidate storage devices that meet the work substitution degree threshold conditions and their corresponding work substitution degrees are retrieved from other device groups not specified by the current strategy. If the number of supplementary devices is 1, the candidate device with the largest product of job substitution degree and storage operation adaptability is selected for allocation; if the number of supplementary devices is greater than 1, multiple candidate devices are selected in descending order of the product of job substitution degree and storage operation adaptability for allocation, and the adjusted storage device resource configuration is finally obtained.

8. The intelligent archive storage and allocation method as described in claim 7, characterized in that, The training steps for the demand depth optimization inference model include: S201. Construct a demand depth optimization reasoning model based on the pre-trained Bayesian network reasoning model, and construct the input sequence of the demand depth optimization reasoning model according to the file storage requirements under the current storage scenario, the information saved by each sub-node in the file storage strategy joint knowledge base, the first depth search connection and the second depth search connection. S202. Utilize the storage efficiency and data integrity metrics corresponding to the optimized storage allocation strategy under each storage scenario to construct a fine-tuned pre-training loss function for the demand-deep optimization inference model. S203. Set the fine-tuning pre-training threshold, input the input sequence and the fine-tuning pre-training loss function into the demand depth optimization inference model for fine-tuning training, and obtain the fine-tuned demand depth optimization inference model. S204. Real-time acquisition of the execution effect data of file storage strategies in various scenarios, including actual storage efficiency and data integrity indicators, and feeding the real-time execution effect data back to the requirement-deep optimization inference model for online adaptive fine-tuning and optimization of the model.

9. An intelligent archive storage and allocation system, used to implement the intelligent archive storage and allocation method according to any one of claims 1-8, characterized in that, include: Knowledge base module and deep linking module; The knowledge base module includes a data processing unit and a knowledge base construction unit; The data processing unit is used to acquire and preprocess different types of archive data, corresponding standard storage process data, and storage strategies, operation processes and equipment resource configuration data under different historical storage scenarios. The knowledge base construction unit is used to construct a joint knowledge base for archive storage strategies using preprocessed data and knowledge graph algorithms. The deep linking module is used to construct and train a demand-deep optimization reasoning model based on the knowledge base node reasoning algorithm, configure the demand-deep optimization reasoning model into the joint knowledge base of the archive storage strategy, and perform deep association retrieval of each sub-base through the storage information distribution node status.

10. The intelligent archive storage and allocation system as described in claim 9, characterized in that, The intelligent archive storage and allocation system also includes an initial evaluation module and a deep search configuration module; The initial assessment module includes a scenario assessment unit and an initial file assessment unit; The scenario assessment unit is used to acquire scenario data to be stored, and to assess the security level of the current storage environment and the probability of file damage risk through the configured security assessment model, and output the security level of the current storage scenario and the probability of file risk. The initial file assessment unit is used to comprehensively assess the current file to be stored through a comprehensive assessment model under the security level and file risk probability conditions corresponding to the current storage scenario, and generate an initial assessment score and an initial storage allocation strategy. The deep search configuration module includes a deep search unit and a resource adjustment unit; The deep search unit is used to perform a second deep matching retrieval of the storage strategy in the current scenario from the joint knowledge base of archive storage strategies, based on the initial evaluation score and initial storage allocation strategy of the archive under the current storage scenario security level and archive risk probability conditions, through the demand deep optimization reasoning model, to obtain the optimized storage allocation strategy of the archive in the corresponding scenario. The resource adjustment unit is used to determine the resource configuration differences based on the device resource configuration and actual device status in the optimized storage allocation strategy, and to dynamically allocate and adjust storage resources in combination with real-time device information.

Citation Information

Patent Citations

  • CN119357239A

  • WO2025177077A1