Archive management system based on artificial intelligence
By introducing variable-scale Levi flight and improving the Gray Wolf optimization algorithm in the archive management system, the filling accuracy and adaptability problems of archival data corruption and missing archive data are solved, and efficient and accurate archival data repair and filling are achieved, meeting the needs of modern archive management.
Patent Information
- Application Number
- CN202510274400.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-03-10
AI Technical Summary
When handling data corruption and missing, existing archive management technologies have problems such as insufficient accuracy, limited adaptability and high computing resources, which are difficult to meet the needs of modern archive management.
An archive management system based on artificial intelligence is proposed, which adopts modules such as data preprocessing, missing area detection, global search optimization, local optimization, data fusion and consistency matching, filling quality evaluation and optimization, and efficient filling and repair of archive data through variable-scale Levi flight algorithm and improved gray wolf optimization algorithm.
It improves the overall consistency and rationality of archive filling, improves the accuracy and adaptability of filling data, reduces the consumption of computing resources, and meets the actual needs of modern archive management.
Smart Images

Figure CN120216747A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of file management, and in particular, to a file management system based on artificial intelligence. Background Art
[0002] As an important carrier for information storage and historical records, files play an indispensable role in various enterprises, institutions, government agencies and research institutions. Over time, files in the process of file management may have problems such as data loss, content damage and information loss due to physical damage, storage medium aging, improper data migration or external environmental factors, which not only affect the integrity and readability of files, but also have a serious impact on historical research, legal evidence preservation and enterprise operation.
[0003] Currently, the repair methods for damaged file data mainly include manual repair, filling methods based on statistical models, and intelligent repair technologies based on deep learning. However, the existing methods still have many deficiencies in practical applications.
[0004] First of all, the manual repair method mainly relies on the experience and subjective judgment of professional file repair personnel. For text files, experts can fill in the missing content by manual comparison and speculation. For the repair process of image or multimedia files, it is more complex and requires superb repair skills. However, the existing methods are not only time-consuming and laborious, but also due to the uncontrollability of human factors, it is easy to lead to unstable repair quality. In addition, in the face of large-scale file data, the efficiency of manual repair is extremely low, and it is difficult to meet the requirements of modern file management for repairing a large number of data.
[0005] Secondly, filling methods based on statistical models, such as interpolation methods, regression models and Markov models, usually predict missing data by calculating the probability distribution of existing data. The existing methods are effective in dealing with small-scale regularized data, but for content such as file data that is complex and has non-linear characteristics, traditional statistical methods are difficult to accurately restore the semantic information and context logic of the data, often resulting in insufficient accuracy of the filled content. In addition, statistical models are difficult to adapt to different types of file data, and the applicable range is relatively limited.
[0006] In recent years, with the development of artificial intelligence technology, deep learning methods have gradually been applied to the field of archival restoration. Restoration methods based on convolutional neural networks, generative adversarial networks, or variational autoencoder technology have achieved certain results in filling in image and video archives, but there are still certain limitations: on the one hand, deep learning methods usually require a large amount of high-quality training data, and in practical applications, it is often very difficult to obtain and annotate archival data; on the other hand, when existing deep learning models process archival data with relatively severe missing information, the filling results are prone to blurring, distortion, and even semantic errors, and it is difficult to ensure that the restored archival content conforms to the logic and style of the original archives. In addition, some deep learning-based restoration methods have a high computational complexity and consume a large amount of computing resources, which is not conducive to the restoration application of large-scale archival data.
[0007] Therefore, there is an urgent need for an artificial intelligence-based archival management system to improve the accuracy and adaptability of archival data filling and meet the actual needs of modern archival management. Summary of the Invention
[0008] An object of the present invention is to provide an artificial intelligence-based archival management system, which improves the overall consistency and rationality of archival filling in the process of archival management.
[0009] An artificial intelligence-based archival management system according to an embodiment of the present invention includes the following modules:
[0010] A data preprocessing module, which is used to obtain the original archival data, and perform denoising, removing invalid information, format unification, and standardization processing on the original archival data. The data preprocessing module analyzes different types of archival data based on the archival data storage format and encoding method;
[0011] A missing area detection module, which is used to detect and extract the features of the missing area in the preprocessed archival data, determine and calibrate the boundaries and context features of the damaged area and the complete area in the archival data. The missing area detection module calculates the integrity score of the damaged area, sets the missing area detection threshold, filters the archival data with the integrity score lower than the threshold, and generates missing area description information based on the boundary information, texture features, and context semantic information of the archival data;
[0012] A global search optimization module, which is used to perform a global search using the variable-scale Levy flight algorithm based on the missing area description information, adaptively adjust the search step size during the global search process, and identify the global search candidate data that matches the missing area;
[0013] A local optimization module, which is used to search for candidate data globally and optimize the details of local missing areas by using an improved grey wolf optimization algorithm. During the local optimization process, the optimization parameters are dynamically adjusted according to the global search results to achieve local data completion and detail restoration;
[0014] A data fusion and consistency matching module, which is used to fuse the filled data after local optimization with the preprocessed archive data and perform semantic, style, and logical consistency matching on the fused archive data;
[0015] A filling quality evaluation and optimization module, which is used to perform error feedback evaluation on the filled archive data, calculate the overall matching degree of the filled data, and screen the data with filling quality higher than the threshold based on the global evaluation results. According to the global matching evaluation results, the filled data is optimized twice to make it achieve the best match with the original archive in terms of content logic, format layout, and semantic consistency, and finally the corrected archive data is archived.
[0016] An archive management method based on artificial intelligence, which is used to execute an archive management system based on artificial intelligence, including the following steps:
[0017] S1. Preprocess the original archive data to obtain the preprocessed archive data;
[0018] S2. Detect and extract the missing areas of the preprocessed archive data, determine and calibrate the boundaries and context features of the damaged areas and complete areas in the archive data, and form missing area description information;
[0019] S3. Based on the missing area description information, use the variable-scale Lévy flight algorithm to globally search the preprocessed archive data and generate global search candidate data that matches the missing areas;
[0020] S4. Use the improved grey wolf optimization algorithm to optimize the details of the local missing areas, and dynamically adjust the optimization parameters according to the global search results during the local optimization process to generate the filled data after local optimization;
[0021] S5. Fuse the filled data after local optimization with the preprocessed archive data, perform semantic, style, and logical consistency matching on the fused archive data, and form the preliminarily corrected archive data;
[0022] S6. Implement error feedback evaluation on the preliminarily corrected archive data, compare the filling results with the expected features through the error evaluation module, dynamically adjust the global search and local optimization parameters, and re-optimize the preliminarily corrected archive data to generate the final completed data.
[0023] Optionally, the S1 includes the following steps:
[0024] S11. Obtain the original archive data, perform format recognition on the original archive data based on the archive data storage format and encoding method, construct a format mapping relationship, and define an archive data set:
[0025] D raw ={d1,d2,…,d N};
[0026] Among them, D raw represents the original archive data set, d i represents the i-th archive data instance, and N is the total amount of original archive data;
[0027] S12. Denoise the original archive data set, filter out invalid information from the denoised archive data, and remove data items containing low-correlation information to form a valid archive data set;
[0028] S13. Perform format unification processing on the valid archive data set, standardize the format-unified archive data set D formatted , and perform normalization and data alignment operations on the data according to the distribution characteristics of the archive content to generate a preprocessed archive data set D processed .
[0029] Optionally, the S2 includes the following steps:
[0030] S21. Perform integrity detection on the preprocessed archive data set D processed , calculate the integrity score C(d i ) of each data instance d i :
[0031]
[0032] Among them, C(d i ) represents the integrity of the data instance d i , |D missing,i | represents the number of missing contents in the data instance d i ;
[0033] S22. Set a missing area detection threshold based on the integrity score C(d i ), screen out data instances with integrity scores lower than the threshold, and form a missing data subset D missing ;
[0034] S23. Use a structured analysis method to segment the damaged areas of each data instance in the missing data subset, construct a damaged area determination function M(r i,j ), and finally form a damaged area set R missing ;
[0035] For each potential damaged area r in the data instance i,j Calculate the comprehensive damage score Λ i,j :
[0036] Λ i,j = α1·D(r i,j ) + β·S(r i,j ) + γ·C(r i,j );
[0037] Wherein, D(r i,j ) represents a deviation index for measuring the deviation of the internal data from the expected value of r i,j , S(r i,j ) represents an index reflecting the degree of discontinuity of the internal structure of r i,j , C(r i,j ) represents a correlation index for evaluating the context consistency between r i,j and its adjacent areas, and α1, β, and γ are weight coefficients;
[0038] Define the damaged area determination function as:
[0039]
[0040] Wherein, δ th is the damage threshold determined based on the experience of historical archive repair data;
[0041] S24. For each damaged area r i,j Extract its boundary information to form a boundary information set B missing , and define the boundary extraction function B(r i,j ):
[0042]
[0043] Wherein, represents the second-order gradient operator, W(r i,j ) is a weight function dynamically adjusted according to the size and shape of r i,j , and b i,j represents the boundary information of the damaged area r i,j ;
[0044] S25. Extract the features of the context of each damaged area r i,j to form a context feature set F context , and construct a context feature vector F context,i,j :
[0045] F context,i,j = [ξ1·E(r i,j ), ξ2·T(r i,j ), ξ3·L(r i,j )],
[0046] Among them, E(r i,j ) represents the edge continuity measure, reflecting the edge ductility around the damaged area r i,j ; T(r i,j ) represents the texture similarity measure, reflecting the texture consistency between the damaged area r i,j and the adjacent area; L(r i,j ) represents the layout consistency index, and ξ1, ξ2, and ξ3 are weight coefficients;
[0047] S26. Fuse the damaged area R missing , the boundary information set B missing and the context feature vector F context to construct the missing area description information D desc :
[0048] D desc,i,j = (r i,j , b i,j , F context,i,j , P(r i,j ));
[0049] Among them, P(r i,j ) is the recovery priority prediction function:
[0050]
[0051] Among them, f k is the k-th feature component in F context,i,j , M represents the dimension of the context feature vector, and λ is the balance coefficient;
[0052] Finally, form the missing area description information set:
[0053]
[0054] Optionally, the S3 includes the following steps:
[0055] S31. Extract the damaged area R desc and the context feature set F missing from the missing area description information set D context , and construct the global search space S search according to the spatial distribution of the missing area:
[0056]
[0057] Among them, S search represents the search range of the damaged area of the archive data in the two-dimensional space, and (x i,j , y i,j ) is the damaged area r i,jThe set of coordinate points in this search space;
[0058] S32. In the global search space S search , the variable-scale Levy flight algorithm is used to perform global search. According to the context feature set F context , calculate the entropy value of the damaged area and dynamically adjust the search step size S t :
[0059]
[0060] where S t represents the step size at the t-th iteration in the search process, S max is the initial maximum step size, H(F context,i,j ) represents the information entropy of the damaged area r i,j in the context feature set, which is used to measure the complexity of this area. The higher the information entropy value, the more serious the information loss in this area. λ1 is a control factor;
[0061] S33. Calculate the search jump step size L α according to the variable-scale Levy flight model, and update the next search point (x t+1 , y t+1 ) based on the position of the current search point:
[0062] (x t+1 , y t+1 ) = (x t , y t ) + S t ·L α ·(cosθ, sinθ);
[0063] where cosθ and sinθ represent the horizontal and vertical components of the jump direction respectively, and L α obeys the Levy distribution:
[0064]
[0065] where Γ(α) is the gamma function, which is used for normalizing the calculation of the jump step size. θ is a uniformly random angle, representing the direction of the search jump, and s is a random variable that controls the change of the jump step size;
[0066] S34. Calculate the matching degree S(x t+1 , y t+1 ) between the updated search point (x t+1 , y t+1 ) and the boundary of the corresponding damaged area, and determine whether to enter the local optimization mode:
[0067]
[0068] where Bmissing is the boundary set of the damaged area, b k are the boundary point coordinates, and γ1 is the matching degree control factor, which determines the weight of the influence of the matching degree on the search points;
[0069] The matching degree reflects whether the search point enters the high - correlation area. When the matching degree S(x t+1 ,y t+1 ) > τ, the local optimization mode is started, and the step size is dynamically adjusted:
[0070] S local,t+1 = S max ·(1 - S(x t+1 ,y t+1 ));
[0071] where τ is the local optimization trigger threshold, and S local,t+1 is the local optimization step size, and the step size tends to 0 when the matching degree is high;
[0072] S35. Introduce the direction weight vector W t in the local optimization mode to optimize the search direction and calculate the search direction adjustment parameter:
[0073]
[0074] where w x and w y are the matching degrees in the x - direction and y - direction respectively;
[0075] The updated formula for the optimized search point:
[0076] (x t+2 ,y t+2 )=(x t+1 ,y t+1 ) + S local,t+1 ·L α ·W t ;
[0077] where L α has its value range adjusted to 1.5 < α < 2 in the local search mode;
[0078] S36. Store the local optimal points during the local optimization process. If the optimal value has not been updated for K jump consecutive rounds of searches, perform a small - range jump to prevent local convergence;
[0079] S37. Calculate the objective fitness function F(x t ,y t ) based on the spatial distribution of the search points and their matching degree with the damaged area, and use it to screen the optimal search path:
[0080]
[0081] Among them, S(x t+2 , y t+2 ) is the matching degree between the optimized search point and the boundary of the damaged area, and G(x t+2 , y t+2 ) is the coverage degree of the optimized search point inside the damaged area. The normalized convergence rate, and w1, w2, w3 are weight parameters for adjusting various influencing factors;
[0082] S38. According to the convergence situation of the search trajectory, screen the optimal search path and construct a global search candidate dataset D that matches the missing area global :
[0083]
[0084] Optionally, the S4 includes the following steps:
[0085] S41. Obtain the global search candidate dataset D global , and construct an adaptive multi-layer local optimization population P for the complexity of different file damaged areas multi . Different levels of optimization populations undertake different optimization tasks during the search process:
[0086]
[0087] Among them, P multi represents the multi-layer optimization population. The individuals of all optimization layers form a complete search space. P m represents the m-th layer optimization population. The optimization population divides the damaged area according to the complexity of the damaged area. represents the k-th search individual in the m-th layer optimization population, that is, the local optimization candidate point, and M is the number of optimization layers;
[0088] Calculate the complexity C of the damaged area damage . When the complexity C of the damaged area damage is higher, it means that the structural complexity of the damaged area is greater, and the grey wolf optimization algorithm increases the number of optimization layers M:
[0089]
[0090] Among them, σ G (r i,j ) represents the standard deviation of the gradient inside the damaged area r i,j , measuring the edge complexity within the area, is the degree of regional texture mutation calculated by the Laplace operator, describing the structural discontinuity of the damaged area;
[0091] S42. During the optimization process, different layers of the optimization population adopt different fitness functions to adapt to different damage characteristics of the archive data, and a multi-layer adaptive fitness function is defined:
[0092]
[0093] Among them, is the matching degree between the local search point and the boundary of the damaged area, making the filling points consistent with the morphological characteristics of the damaged area, is the context consistency, measuring whether the filled data conforms to the logic and content structure of the original archive data, is the texture similarity, measuring the similarity of the texture features between the filling points and the adjacent restored data:
[0094]
[0095] Among them, G(x, y) is the local gradient feature, and N k is the filled data points within the neighborhood range of the search point;
[0096] is the local data smoothness, making the continuity of the filled data in the local area:
[0097]
[0098] Among them, I(x, y) represents the pixel value or text feature value of the archive data;
[0099] is the information entropy increment, measuring the contribution of the filling points to the overall information integrity:
[0100]
[0101] Among them, H before and H after are the entropy values of the damaged area before and after filling respectively;
[0102] S43. In the archive filling task, the search direction is guided by combining the boundary information of the archive data, and texture gradient guidance is introduced on the basis of the traditional grey wolf optimization search to construct a texture-guided grey wolf optimization search:
[0103]
[0104] Among them, is the search position of the k-th grey wolf in the m-th layer of the optimization population at the (t + 1)-th iteration, indicating the position change of the filling candidate point during the optimization process, represent the positions of the three grey wolves, the optimal solution, the sub-optimal solution, and the third-optimal solution, in this optimization population respectively. A mis the coefficient for controlling the search step size and convergence speed of gray wolf individuals, respectively represent the distances between the gray wolf individual and the optimal solution, the second-best solution, and the third-best solution positions, B m is the gradient guidance term:
[0105]
[0106] Among them, represents the local gradient direction, guiding the filling point to converge towards the optimal texture matching direction, and γ2 is the gradient guidance weight;
[0107] S44. In the archival data repair task, adopt a hybrid step size strategy to optimize the search of gray wolf individuals, dynamically adjust the search step size to balance accuracy and convergence speed:
[0108]
[0109] Among them, 2·(1 - t / T max ) controls the global convergence trend, and λ m ·N(0,1) enhances search diversity through hierarchical perturbation;
[0110] S45. Introduce a local convergence detection mechanism. If a certain search point does not improve the fitness for K stuck consecutive rounds of iteration, then perform a jump:
[0111]
[0112] S46. After the gray wolf optimization algorithm converges, screen the optimal filling points to finally form the locally optimized filling data:
[0113]
[0114] Among them, are the coordinates of k + 2 search individuals in the m-th optimized population after optimization, and F m (p k+2 ) is the optimized multi-level adaptive fitness function.
[0115] Optionally, the S5 includes the following steps:
[0116] S51. Obtain the locally optimized filling data and the preprocessed archival data. According to the data format, storage structure, and encoding method, perform normalization processing on the filling data, and perform preliminary fusion on the normalized filling data and the preprocessed archival data to construct an initial data fusion set;
[0117] S52. Analyze the local features of the area where the filled data is located, extract the local feature vectors of the damaged area, and calculate the local similarity between the filled data and the adjacent areas, including the boundary shape similarity, local structure consistency, and content feature matching degree. Determine whether the filled data conforms to the local style features of the original file based on the calculation results of the local similarity. When the local similarity is lower than the preset threshold, adjust the filled data.
[0118] S53. Perform style, semantic, and logical consistency matching on the fused file data based on the pattern features of the historical file data. Analyze the style adaptability of the filled data according to the language features, layout styles, and data structure rules of the file content. For the filled data with a matching degree lower than the threshold, correct it through a style adjustment function to make the filled data consistent with the original file in terms of font format, text semantics, structural logic, and context coherence.
[0119] S54. After completing the local similarity adjustment and style matching, conduct a global evaluation of the overall filling result, calculate the matching degree between the filled data and the entire file content. Based on the global matching result, screen out the data with a filling quality higher than the threshold, and perform secondary adjustment on the filled data with a low matching degree to make the filled data achieve the best match with the entire file in terms of content logic, format layout, and semantic consistency.
[0120] S55. After completing the local and global consistency matching, file the finally screened filled data to generate the preliminarily corrected file data.
[0121] The beneficial effects of the present invention are as follows:
[0122] (1) The present invention conducts global search through the variable-scale Lévy flight algorithm. During the search process, the search step size is adaptively adjusted to enable the search to jump within a wide range, and at the same time, fine-grained search is performed in the high-matching area. The variable-scale Lévy flight can combine large-step global exploration with small-step local refinement, improve the search efficiency and avoid falling into local optima. At the same time, it can search for the optimal filled data within a larger range, improving the overall consistency and rationality of the filled data.
[0123] (2) The present invention introduces an improved grey wolf optimization algorithm to optimize and adjust the local area. The grey wolf optimization algorithm dynamically adjusts the boundary features, texture consistency, and context logic of the filled data by simulating the hunting behavior of wolf packs, ensuring that the filled area is consistent with the semantics, style, and structure of the original data. Through the guided search of the grey wolf optimization algorithm, the filled data can be more naturally integrated with the surrounding areas.
[0124] (3) After the filling of the present invention is completed, an error feedback mechanism is adopted to automatically evaluate the repair result, and the global search and local optimization parameters are dynamically adjusted based on the feedback result, so as to achieve adaptive optimization. The matching degree between the filled data and the original data is calculated by the error evaluation module, including semantic consistency, style adaptability and structural rationality, and the search step size and optimization parameters are adjusted according to the feedback result to ensure that the filled data continuously approaches the optimal solution. BRIEF DESCRIPTION OF THE DRAWINGS
[0125] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention, and do not constitute a limitation to the present invention. In the drawings:
[0126] Figure 1 is a flowchart of an archive management system based on artificial intelligence proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0127] Now, the present invention will be further described in detail with reference to the drawings. These drawings are all simplified schematic diagrams, only showing the basic structure of the present invention in a schematic way, so they only show the components related to the present invention.
[0128] Refer to Figure 1 , an archive management system based on artificial intelligence, includes the following modules:
[0129] A data preprocessing module for obtaining the original archive data and performing denoising, removing invalid information, format unification and standardization processing on the original archive data. The data preprocessing module analyzes different types of archive data based on the storage format and encoding method of the archive data;
[0130] A missing area detection module for detecting and extracting the missing area of the preprocessed archive data, determining and calibrating the boundary and context features of the damaged area and the complete area in the archive data. The missing area detection module calculates the integrity score of the damaged area, sets the missing area detection threshold, filters the archive data with the integrity score lower than the threshold, and generates missing area description information based on the boundary information, texture features and context semantic information of the archive data;
[0131] A global search and optimization module for performing global search using the variable scale Levy flight algorithm based on the missing area description information, adaptively adjusting the search step size during the global search, and identifying the global search candidate data matching the missing area;
[0132] A local optimization module for performing detailed optimization on the local missing area based on the global search candidate data, dynamically adjusting the optimization parameters according to the global search result during the local optimization, and realizing local data completion and detail restoration;
[0133] A data fusion and consistency matching module, which is used to fuse the locally optimized filled data with the preprocessed archive data, and perform semantic, style, and logical consistency matching on the fused archive data;
[0134] A filling quality evaluation and optimization module, which is used to perform error feedback evaluation on the filled archive data, calculate the overall matching degree of the filled data, and screen the data with filling quality higher than the threshold based on the global evaluation result, and perform secondary optimization on the filled data according to the global matching evaluation result to make it achieve the best matching with the original archive in terms of content logic, format layout, and semantic consistency, and finally archive the corrected archive data.
[0135] An archive management method based on artificial intelligence, which is used to execute an archive management system based on artificial intelligence, including the following steps:
[0136] S1. Preprocess the original archive data to obtain the preprocessed archive data;
[0137] S2. Detect and extract the missing areas of the preprocessed archive data, determine and calibrate the boundaries and context features of the damaged areas and intact areas in the archive data, and form missing area description information;
[0138] S3. Based on the missing area description information, use the variable-scale Lévy flight algorithm to perform a global search on the preprocessed archive data, and generate global search candidate data that matches the missing area;
[0139] S4. Use the improved grey wolf optimization algorithm to perform detailed optimization on the local missing areas, and dynamically adjust the optimization parameters according to the global search results during the local optimization process to generate locally optimized filled data;
[0140] S5. Fuse the locally optimized filled data with the preprocessed archive data, and perform semantic, style, and logical consistency matching on the fused archive data to form preliminarily corrected archive data;
[0141] S6. Implement error feedback evaluation on the preliminarily corrected archive data, compare the filling results with the expected features through the error evaluation module, dynamically adjust the global search and local optimization parameters, and perform re-optimization on the preliminarily corrected archive data to generate the final completed data.
[0142] In this embodiment, S1 includes the following steps:
[0143] S11. Obtain the original archive data, perform format recognition on the original archive data based on the archive data storage format and encoding method, construct a format mapping relationship, and define an archive data set:
[0144] D raw = {d1, d2, …, d N};
[0145] Among them, D raw represents the original archive data set, d i represents the i-th archive data instance, and N is the total amount of original archive data;
[0146] S12. Denoise the original archive data set, filter out invalid information from the denoised archive data, and eliminate data items containing low-correlation information to form a valid archive data set;
[0147] S13. Uniformly process the format of the valid archive data set, standardize the uniformly formatted archive data set D formatted , and perform normalization and data alignment operations on the data according to the distribution characteristics of the archive content to generate the preprocessed archive data set D processed .
[0148] In this embodiment, S2 includes the following steps:
[0149] S21. Detect the integrity of the preprocessed archive data set D processed , calculate the integrity score C(d i ) of each data instance d i :
[0150]
[0151] Among them, C(d i ) represents the integrity of the data instance d i , |D missing,i | represents the number of missing contents in the data instance d i ;
[0152] S22. Set the missing area detection threshold based on the integrity score C(d i ), screen out the data instances with integrity scores lower than the threshold, and form a missing data subset D missing ;
[0153] S23. Use a structured analysis method to segment the damaged areas for each data instance in the missing data subset, construct a damaged area determination function M(r i,j ), and finally form a damaged area set R missing ;
[0154] For each potential damaged area r i,j in the data instance, calculate the comprehensive damage score Λ i,j :
[0155] Λi,j = α1·D(r i,j ) + β·S(r i,j ) + γ·C(r i,j );
[0156] Wherein, D(r i,j ) represents a deviation index for measuring the deviation of r i,j from the expected value of the internal data, S(r i,j ) represents an index reflecting the degree of discontinuity of the internal structure of r i,j , C(r i,j ) represents a correlation index for evaluating the context consistency of r i,j with its adjacent regions, and α1, β, and γ are weight coefficients;
[0157] Define the damaged area determination function as:
[0158]
[0159] Wherein, δ th is a damage threshold determined based on the experience of historical file repair data;
[0160] S24. For each damaged area r i,j Extract its boundary information to form a boundary information set B missing , and define the boundary extraction function B(r i,j ):
[0161]
[0162] Wherein, represents the second-order gradient operator, W(r i,j ) is a weight function dynamically adjusted according to the size and shape of r i,j , and b i,j represents the boundary information of the damaged area r i,j ;
[0163] S25. Extract the features of the context of each damaged area r i,j to form a context feature set F context , and construct a context feature vector F context,i,j :
[0164] F context,i,j = [ξ1·E(r i,j ), ξ2·T(r i,j ), ξ3·L(r i,j )],
[0165] Wherein, E(r i,j ) represents the edge continuity measure, reflecting the damaged area r i,jPeripheral edge ductility, T(r i,j ) represents the texture similarity measure, reflecting the damaged area r i,j and the texture consistency between the damaged area and the adjacent area, L(r i,j ) represents the layout consistency index, and ξ1, ξ2, and ξ3 are weight coefficients;
[0166] S26. Fuse the damaged area R missing , the boundary information set B missing and the context feature vector F context to construct the missing area description information D desc :
[0167] D desc,i,j =(r i,j , b i,j , F context,i,j , P(r i,j ));
[0168] Among them, P(r i,j ) is the recovery priority prediction function:
[0169]
[0170] Among them, f k is the k-th feature component in F context,i,j , M represents the dimension of the context feature vector, and λ is the balance coefficient;
[0171] Finally, form the missing area description information set:
[0172]
[0173] In this embodiment, S3 includes the following steps:
[0174] S31. Extract the damaged area R desc and the context feature set F missing from the missing area description information set D context , and construct the global search space S search according to the spatial distribution of the missing area:
[0175]
[0176] Among them, S search represents the search range of the damaged area of the archive data in the two-dimensional space, and (x i,j , y i,j ) is the coordinate point set of the damaged area r i,j in this search space;
[0177] S32. In the global search space S searchIn it, the variable-scale Levy flight algorithm is adopted to perform global search, and based on the context feature set F context Calculate the entropy value of the damaged area and dynamically adjust the search step size S t :
[0178]
[0179] Among them, S t represents the step size of the t-th iteration in the search process, S max is the initial maximum step size, H(F context,i,j ) represents the information entropy of the damaged area r i,j in the context feature set, which is used to measure the complexity of this area. The higher the information entropy value, the more serious the information loss in this area. λ1 is a control factor;
[0180] S33. Calculate the search jump step size L according to the variable-scale Levy flight model α , and update the next search point (x t+1 , y t+1 ) according to the position of the current search point:
[0181] (x t+1 , y t+1 ) = (x t , y t ) + S t ·L α ·(cosθ, sinθ);
[0182] Among them, cosθ and sinθ respectively represent the horizontal and vertical components of the jump direction, and L α obeys the Levy distribution:
[0183]
[0184] Among them, Γ(α) is the gamma function, which is used for normalizing the calculation of the jump step size. θ is a uniformly random angle, representing the direction of the search jump, and s is a random variable, controlling the change of the jump step size;
[0185] S34. Calculate the matching degree S(x t+1 , y t+1 ) between the updated search point (x t+1 , y t+1 ) and the boundary of the corresponding damaged area, and judge whether to enter the local optimization mode:
[0186]
[0187] Among them, B missing is the boundary set of the damaged area, and b kis the boundary point coordinate, and γ1 is the matching degree control factor, which determines the weight of the influence of the matching degree on the search point;
[0188] The matching degree reflects whether the search point enters the high-correlation area. When the matching degree S(x t+1 ,y t+1 ) of the damaged area boundary is greater than τ, the local optimization mode is started, and the step size is dynamically adjusted:
[0189] S local,t+1 = S max ·(1 - S(x t+1 ,y t+1 ));
[0190] where τ is the local optimization trigger threshold, and S local,t+1 is the local optimization step size, and the step size tends to 0 when the matching degree is high;
[0191] S35. Introduce the direction weight vector W t in the local optimization mode to optimize the search direction and calculate the search direction adjustment parameter:
[0192]
[0193] where w x and w y are the matching degrees in the x and y directions respectively;
[0194] The updated formula for the optimized search point:
[0195] (x t+2 ,y t+2 ) = (x t+1 ,y t+1 ) + S local,t+1 ·L α ·W t ;
[0196] where L α has a value range adjustment of 1.5 < α < 2 in the local search mode;
[0197] S36. Store the local optimal point during the local optimization process. If the optimal value has not been updated for K jump consecutive rounds of searches, perform a small-range jump to prevent local convergence;
[0198] S37. Calculate the objective fitness function F(x t ,y t ) based on the spatial distribution of the search points and their matching degree with the damaged area, and use it to screen the optimal search path:
[0199]
[0200] Among them, S(x t+2 , y t+2 ) is the matching degree between the optimized search point and the boundary of the damaged area, and G(x t+2 , y t+2 ) is the coverage degree of the optimized search point inside the damaged area. The normalized convergence rate, and w1, w2, w3 are weight parameters for adjusting various influencing factors;
[0201] S38. According to the convergence situation of the search trajectory, screen the optimal search path and construct a global search candidate dataset D that matches the missing area global :
[0202]
[0203] In this embodiment, S4 includes the following steps:
[0204] S41. Obtain the global search candidate dataset D global , and construct an adaptive multi-layer local optimization population P for the complexity of different file damaged areas multi . Different levels of optimization populations undertake different optimization tasks during the search process:
[0205]
[0206] Among them, P multi represents the multi-layer optimization population, and the individuals of all optimization layers form a complete search space. P m represents the m-th layer optimization population. The optimization population divides the damaged area according to the complexity of the damaged area. represents the k-th search individual in the m-th layer optimization population, that is, the local optimization candidate point, and M is the number of optimization layers;
[0207] Calculate the complexity C of the damaged area damage . When the complexity C of the damaged area damage is higher, it means that the structural complexity of the damaged area is greater, and the grey wolf optimization algorithm increases the number of optimization layers M:
[0208]
[0209] Among them, σ G (r i,j ) represents the standard deviation of the gradient within the damaged area r i,j , which measures the edge complexity within the area, is the degree of regional texture mutation calculated by the Laplace operator, describing the structural discontinuity of the damaged area;
[0210] S42. During the optimization process, different layers of the optimization population adopt differential fitness functions to adapt to different damage characteristics of the archival data, and a multi-layer adaptive fitness function is defined:
[0211]
[0212] Among them, is the matching degree between the local search point and the boundary of the damaged area, making the filling points consistent with the morphological characteristics of the damaged area, is the context consistency, measuring whether the filled data conforms to the logic and content structure of the original archival data, is the texture similarity, measuring the texture feature similarity between the filling points and the adjacent restored data:
[0213]
[0214] Among them, G(x, y) is the local gradient feature, and N k is the filled data points within the neighborhood range of the search point;
[0215] is the local data smoothness, making the continuity of the filled data in the local area:
[0216]
[0217] Among them, I(x, y) represents the pixel value or text feature value of the archival data;
[0218] is the information entropy increment, measuring the contribution of the filling points to the overall information integrity:
[0219]
[0220] Among them, H before and H after are the entropy values of the damaged area before and after filling respectively;
[0221] S43. In the archival filling task, the search direction is guided by combining the boundary information of the archival data, and texture gradient guidance is introduced on the basis of the traditional grey wolf optimization search to construct a texture-guided grey wolf optimization search:
[0222]
[0223] Among them, is the search position of the k-th grey wolf in the m-th layer of the optimization population at the (t + 1)-th iteration, indicating the position change of the filling candidate point during the optimization process, represent the positions of the three grey wolves, namely the optimal solution, the sub-optimal solution, and the third-optimal solution, in the optimization population respectively, and A mis the coefficient for controlling the search step size and convergence speed of the gray wolf individuals. respectively represent the distances between the gray wolf individuals and the positions of the optimal solution, the second-best solution, and the third-best solution. B m is the gradient guiding term:
[0224]
[0225] Among them, represents the local gradient direction, guiding the filling point to converge towards the optimal texture matching direction, and γ2 is the gradient guiding weight;
[0226] S44. In the archival data repair task, adopt a hybrid step size strategy to optimize the search of gray wolf individuals, dynamically adjust the search step size to balance accuracy and convergence speed:
[0227]
[0228] Among them, 2·(1 - t / T max ) controls the global convergence trend, and λ m ·N(0,1) enhances the search diversity through hierarchical perturbation;
[0229] S45. Introduce a local convergence detection mechanism. If a certain search point does not improve the fitness for K stuck consecutive iterations, then perform a jump:
[0230]
[0231] S46. After the gray wolf optimization algorithm converges, screen the optimal filling points to finally form the locally optimized filling data:
[0232]
[0233] Among them, are the coordinates of k + 2 search individuals in the m-th optimized population after optimization, and F m (p k+2 ) is the optimized multi-level adaptive fitness function.
[0234] In this embodiment, S5 includes the following steps:
[0235] S51. Obtain the locally optimized filling data and the preprocessed archival data. According to the data format, storage structure, and encoding method, perform normalization processing on the filling data, and perform preliminary fusion on the normalized filling data and the preprocessed archival data to construct an initial set of data fusion;
[0236] S52. Analyze the local features of the area where the filled data is located, extract the local feature vectors of the damaged area, and calculate the local similarity between the filled data and the adjacent areas, including the boundary shape similarity, local structure consistency, and content feature matching degree. Determine whether the filled data conforms to the local style features of the original file based on the calculation results of the local similarity. When the local similarity is lower than the preset threshold, adjust the filled data.
[0237] S53. Perform style, semantic, and logical consistency matching on the fused file data based on the pattern features of the historical file data. Analyze the style adaptability of the filled data according to the language features, layout styles, and data structure rules of the file content. For the filled data with a matching degree lower than the threshold, correct it through the style adjustment function to make the filled data consistent with the original file in terms of font format, text semantics, structural logic, and context coherence.
[0238] S54. After completing the local similarity adjustment and style matching, conduct a global evaluation of the overall filling result, calculate the matching degree between the filled data and the entire file content. Based on the global matching result, screen out the data with a filling quality higher than the threshold, and perform secondary adjustment on the filled data with a relatively low matching degree to make the filled data achieve the best match with the entire file in terms of content logic, format layout, and semantic consistency.
[0239] S55. After completing the local and global consistency matching, file the finally screened filled data to generate the preliminarily corrected file data.
[0240] Example 1:
[0241] In July 2023, in the digital file management project of an archive in a certain province in East China, when the staff were checking the digital storage of the archived files in the collection, they found that there were abnormalities in the electronic file documents of some historical official documents. When randomly checking the collection database, the technical personnel found that there were serious omissions in the content of the file document numbered AG1952 - 001. Some texts were damaged due to the aging of the storage medium, and the integrity score of the file was lower than the set security threshold. Further inspection found that about 18% of the content in the lower right corner areas of pages 2 and 5 of the scanned image of this file was missing, and some content of the document could not be automatically recognized due to blurring and noise interference.
[0242] To repair this file document, the technical team decided to use the file damage content filling method combining variable - scale Lévy flight and grey wolf optimization for repair, and compare the repair results with traditional methods to verify the effectiveness of the method of the present invention.
[0243] Technicians first input the scanned image of the file numbered AG1952-001 into the system. The system automatically detects the damaged areas of the document and calculates the integrity score as 0.82 (lower than the threshold of 0.9). By comparison, it is found that 5 lines of text in the lower right corner of page 2 are completely missing, and 2 paragraphs on the right side of page 5 are unreadable due to blurring. The system automatically denoises, enhances the contrast, and binarizes the file image to improve the clarity of the text.
[0244] After the preprocessing is completed, the system starts to detect the missing areas, calculates the text damage score as 0.24 (the lower the value, the more serious the damage), extracts the context features of the damaged areas, and records the coordinate information of the damaged areas to prepare for subsequent filling.
[0245] After entering the filling stage, the system calls the variable-scale Lévy flight algorithm to perform a global search. The system analyzes the official documents in the historical file database from 1951 to 1955, calculates the text fragments most similar to the missing part. During the search process, the system adaptively adjusts the search step size and finds content fragments with a similarity greater than 85% within the global scope. The system screens out three possible matching candidates:
[0246] 1. File number AG1952-002: The similarity of the official document content is 86.3%, but there are deviations in the layout structure;
[0247] 2. File number AG1951-007: The similarity of the official document content is 82.1%, but the language style is relatively old;
[0248] 3. File number AG1953-004: The similarity of the official document content is 91.5%, and the layout and semantics are both well-matched.
[0249] The system finally selects the matching fragment of AG1953-004 as the basis for filling data, fills it into the missing part, and at the same time passes the filling data to the local optimization module.
[0250] After the filling is completed, the system uses the grey wolf optimization algorithm for local adjustment. The system analyzes the boundary consistency score of the filled text as 0.76 (the target value is above 0.85), indicating that the boundary transition is not natural enough. Therefore, it enters the optimization process:
[0251] 1. Adjust font alignment: Based on the comparison between the filled text and the original text, adjust the line spacing and character spacing to increase the format matching degree of the filled part by 12.7%;
[0252] 2. Semantic logic optimization: Adopt context consistency analysis, compare the semantics of the filled content with the context before and after, eliminate 1 grammar error, and correct 2 phenomena where the subject-predicate relationship does not match;
[0253] 3. Text style adjustment: Adjust the wording of the filled parts to make it more in line with the official document language of the 1950s, ensuring consistency in style for the repaired content;
[0254] Finally, the boundary consistency score of the locally optimized filled text was increased to 0.89, and the semantic matching degree reached 96.4%, indicating the completion of the repair.
[0255] To verify the repair effect of the method of the present invention, the technical team simultaneously used the traditional OCR + rule matching filling method for repair and compared the repair results of the two methods. The experimental data is as follows:
[0256]
[0257] The test results show that the method of the present invention exhibits higher filling accuracy, more consistent semantic logic, more natural boundary transition, and higher repair efficiency in the repair of all file types. Especially in the repair of ancient books and documents, the effect improvement is the most obvious.
[0258] After this experiment, the method of the present invention was officially applied by the archives to the digital repair project of the historical documents in the collection. During the period from September to December 2023, 3,127 historical archives were repaired, with a cumulative repair of more than 215,000 pages of archives, greatly improving the efficiency and quality of archive repair. The repaired archives were re-deposited into the collection database for researchers and the public to consult.
[0259] This embodiment details the application of the method for filling damaged content of archives based on the combination of variable-scale Levy flight and grey wolf optimization through an actual archive repair case of an archive in a certain province in East China. The experimental results show that the method of the present invention is significantly superior to the traditional method in terms of filling accuracy, semantic consistency, and repair efficiency, and can effectively improve the quality and automation level of archive repair, having important practical application value.
[0260] The present invention conducts a global search through the variable-scale Levy flight algorithm, adaptively adjusting the search step size during the search process to enable the search to jump within a wide range, and simultaneously performing fine-grained search in the high-matching region. The variable-scale Levy flight can combine large-step global exploration and small-step local refinement, improving the search efficiency and avoiding falling into local optima. At the same time, it can search for the optimal filling data within a larger range, improving the overall consistency and rationality of the filling data.
[0261] The present invention introduces an improved grey wolf optimization algorithm to optimize and adjust the local area. The grey wolf optimization algorithm dynamically adjusts the boundary features, texture consistency, and context logic of the filling data by simulating the hunting behavior of wolf packs, ensuring that the filled area is consistent with the semantics, style, and structure of the original data. Through the guided search of the grey wolf optimization algorithm, the filled data can be more naturally integrated with the surrounding area.
[0262] After the filling of the present invention is completed, an error feedback mechanism is adopted to automatically evaluate the repair result, and the global search and local optimization parameters are dynamically adjusted based on the feedback result, so as to achieve adaptive optimization. The matching degree between the filled data and the original data is calculated by the error evaluation module, including semantic consistency, style adaptability and structural rationality, and the search step size and optimization parameters are adjusted according to the feedback result to ensure that the filled data continuously approaches the optimal solution.
[0263] As mentioned above, it is only the preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and all should be covered within the protection scope of the present invention.
Claims
1. An artificial intelligence-based archive management system, characterized in that: Includes the following modules: The data preprocessing module is used to obtain the original archival data, and to perform denoising, invalid information removal, format unification and standardization on the original archival data. The data preprocessing module parses different types of archival data based on the archival data storage format and encoding method; The missing region detection module is used to detect missing regions and extract features from the preprocessed archive data, determine and calibrate the boundaries and context features of damaged regions and intact regions in the archive data. The missing region detection module calculates the integrity score of the damaged region, sets a missing region detection threshold, filters archive data with integrity scores below the threshold, and generates missing region description information based on the boundary information, texture features and context semantic information of the archive data; A global search optimization module is used to perform a global search based on the missing region description information using a variable-scale Levy flight algorithm, adaptively adjust the search step size during the global search process, and identify global search candidate data that matches the missing region; The local optimization module is used to optimize the details of the local missing areas based on the global search for candidate data using the improved gray wolf optimization algorithm. During the local optimization process, the optimization parameters are dynamically adjusted according to the global search results to achieve local data completion and detail recovery. The data fusion and consistency matching module is used to fuse the locally optimized filled data with the preprocessed archive data, and to match the semantics, style and logical consistency of the fused archive data; The filling quality assessment and optimization module is used to perform error feedback assessment on the filled archival data, calculate the overall matching degree of the filled data, and screen the data with filling quality higher than the threshold based on the global evaluation results. The filled data is optimized again according to the global matching evaluation results to achieve the best match with the original archive in terms of content logic, format layout and semantic consistency, and finally archive the corrected archival data.
2. An artificial intelligence-based archive management method, used to implement the artificial intelligence-based archive management system according to claim 1, characterized in that: The following steps are involved: S1. Preprocessing the original archival data to obtain preprocessed archival data; S2. Perform missing region detection and feature extraction on the preprocessed archive data, determine and calibrate the boundaries and context features of the damaged area and the intact area in the archive data, and form missing region description information; S3. Based on the missing region description information, the variable scale Levy flight algorithm is used to perform a global search on the preprocessed archive data, and generate global search candidate data matching the missing region; S4. Use the improved grey wolf optimization algorithm to optimize the details of the local missing area, dynamically adjust the optimization parameters according to the global search results during the local optimization process, and generate the locally optimized filling data; S5. Fusing the locally optimized filled data with the preprocessed archive data, matching the semantics, style and logical consistency of the fused archive data to form preliminary revised archive data; S6. Implement error feedback evaluation on the preliminary corrected archival data, compare the filling results with the expected features through the error evaluation module, dynamically adjust the global search and local optimization parameters, and optimize the preliminary corrected archival data again to generate the final completed data.
3. The artificial intelligence-based archive management method according to claim 2, characterized in that: The S1 comprises the following steps: S11. Obtain the original archival data, identify the format of the original archival data based on the archival data storage format and encoding method, and build a format mapping relationship to define the archival data set: D raw ={d1,d2,…,d N }; Among them, D raw represents the original archive dataset, d i represents the i-th archival data instance, and N is the total amount of original archival data; S12. De-noising the original archive data set, filtering invalid information on the de-noised archive data, and removing data items containing low-relevance information to form a valid archive data set; S13. Unify the format of valid archive data sets and process the unified archive data sets D formatted Perform standardization processing, normalize and align the data according to the distribution characteristics of the archive content, and generate the preprocessed archive data set D processed .
4. The artificial intelligence-based archive management method according to claim 2, characterized in that: The S2 comprises the following steps: S21. Preprocessed archive data set D processed Perform integrity check and calculate each data instance d i The completeness score C(d i ): Among them, C(d i ) represents the data instance d i The completeness of |D missing,i | represents data instance d i The number of missing contents in S22. Based on the completeness score C(d i ) Set the missing region detection threshold, filter the data instances whose completeness scores are lower than the threshold, and form the missing data subset D missing ; S23. For each data instance in the missing data subset, a structured analysis method is used to segment the damaged area and construct a damaged area determination function M(r i,j ), and finally form the damaged area set R missing ; For each potentially damaged region r in the data instance i,j Calculate the comprehensive damage score Λ i,j : L i,j =α1·D(r i,j )+β·S(r i,j )+γ·C(r i,j ); Among them, D(r i,j ) represents the value used to measure r i,j The deviation index of internal data from the expected value, S(r i,j ) indicates the reflection r i,j An indicator of the degree of internal structural discontinuity, C(r i,j ) indicates the value used to evaluate r i,j The correlation index with the contextual consistency of its adjacent regions, α1, β and γ are weight coefficients; The damage area determination function is defined as: Among them, δ th The damage threshold is determined based on experience with historical archive restoration data; S24. For each damaged region r i,j Extract its boundary information to form a boundary information set B missing , define the boundary extraction function B(r i,j ): in, represents the second-order gradient operator, W(r i,j ) is based on r i,j The weight function of the size and shape of the dynamic adjustment, b i,j The damaged area r i,j Boundary information; S25. For each damaged area r i,j The context is used to extract features and form a context feature set F context , construct the context feature vector F context,i,j : F context,i,j =[ξ1·E(r i,j ),ξ2·T(r i,j ),ξ3·L(r i,j )], Among them, E(r i,j ) represents the edge continuity measure, reflecting the damaged area r i,j Peripheral edge ductility, T(r i,j ) represents the texture similarity measure, reflecting the damaged area r i,j The texture consistency between the adjacent regions, L(r i,j ) represents the layout consistency index, ξ1, ξ2 and ξ3 are weight coefficients; S26. The damaged area R missing , boundary information set B missing and the context feature vector F context To merge.
5. The artificial intelligence-based archive management method according to claim 4, characterized in that: The construction of missing region description information D desc Specifically include: D desc,i,j =(r i,j ,b i,j ,F context,i,j ,P(r i,j )); Among them, P(r i,j ) is the recovery priority prediction function: Among them, f k F context,i,j The kth feature component in , M represents the dimension of the context feature vector, and λ is the balance coefficient; Finally, the missing area description information set is formed:
6. The artificial intelligence-based archive management method according to claim 2, characterized in that: The S3 comprises the following steps: S31. Describing information set D from missing regions desc Extract the damaged area R missing and context feature set F context , construct a global search space S according to the spatial distribution of missing regions search : Among them, S search represents the search range of the damaged area of the archive data in two-dimensional space, (x i,j ,y i,j ) is the damaged area r i,j A set of coordinate points in the search space; S32. In the global search space S search In the process, the variable scale Levy flight algorithm is used to perform global search, based on the context feature set F context Calculate the entropy value of the damaged area and dynamically adjust the search step size S t : Among them, S t represents the step size of the tth iteration in the search process, S max is the initial maximum step length, H(F context,i,j ) represents the damaged area r i,j The information entropy in the context feature set is used to measure the complexity of the region. The higher the information entropy value, the more serious the information loss in the region. λ1 is the control factor. S33. Calculate the search jump step length L based on the variable scale Levy flight model α , and update the next search point (x t+1 ,y t+1 ): (x t+1 ,y t+1 )=(x t ,y t )+S t ·L α ·(cosθ,sinθ); Among them, cosθ and sinθ represent the horizontal component and vertical component of the jumping direction respectively, L α To follow the Levy distribution: Among them, Γ(α) is the gamma function, which is used to normalize the jump step size, θ is a uniform random angle, which indicates the direction of the search jump, and s is a random variable that controls the change of the jump step size; S34. Calculate the updated search point (x t+1 ,y t+1 ) and the matching degree S(x t+1 ,y t+1 ), determine whether to enter the local optimization mode, and obtain the local optimization step length S local,t+1 ; S35. Introducing the direction weight vector W in the local optimization mode t Optimize the search direction and calculate the search direction adjustment parameters: Among them, w x and w y are the matching degrees in the x direction and y direction respectively; Optimized search point update formula: (x t+2 ,y t+2 )=(x t+1 ,y t+1 )+S local,t+1 ·L α ·W t ; Among them, L α In the local search mode, the value range is adjusted to 1.5<α<2; S36. Store the local optimal point during the local optimization process. If K consecutive jump If the round search fails to update the optimal value, a small range jump is performed to prevent local convergence; S37. Calculate the target fitness function F(x t ,y t ), used to filter the optimal search path: Among them, S(x t+2 ,y t+2 ) is the matching degree between the optimized search point and the damaged area boundary, G(x t+2 ,y t+2 ) is the coverage of the optimized search points inside the damaged area, The convergence rate after normalization, w1, w2, w3 are the weight parameters for adjusting various influencing factors; S38. Based on the convergence of the search trajectory, select the optimal search path and construct a global search candidate dataset D that matches the missing area. global :
7. The artificial intelligence-based archive management method according to claim 6, characterized in that: The matching degree S(x) corresponding to the damaged area boundary in S34 t+1 ,y t+1 )for: Among them, B missing is the boundary set of the damaged area, b k is the coordinate of the boundary point, γ1 is the matching degree control factor, which determines the weight of the matching degree on the search point; The matching degree reflects whether the search point enters the high correlation area. When the matching degree S(x t+1 ,y t+1 )>τ, the local optimization mode is started and the step size is adjusted dynamically: S local,t+1 =S max ·(1-S(x t+1 ,y t+1 )); Among them, τ is the local optimization trigger threshold, S local,t+1 It is a local optimization step size, and the step size tends to 0 when the matching degree is high.
8. The artificial intelligence-based archive management method according to claim 2, characterized in that: The S4 comprises the following steps: S41. Obtain global search candidate data set D global , construct an adaptive multi-layer local optimization population P according to the complexity of different archive damage areas multi , optimization populations at different levels undertake different optimization tasks during the search process: Among them, P multi Represents a multi-layer optimization population, and the individuals in all optimization layers constitute a complete search space, P m represents the mth layer optimization group, which divides the damaged area according to the complexity of the damaged area. represents the kth search individual in the mth layer optimization group, that is, the local optimization candidate point, and M is the number of optimization layers; Calculate the complexity of the damaged area C damage , when the damage area complexity C damage The higher it is, the greater the structural complexity of the damaged area, and the Gray Wolf optimization algorithm increases the number of optimization layers M: Among them, σ G (r i,j ) represents the damaged area r i,j The standard deviation of the gradient within the region measures the edge complexity within the region. The degree of regional texture mutation calculated by the Laplacian operator describes the structural discontinuity of the damaged area; S42. During the optimization process, different layers of optimization groups adopt differentiated fitness functions to adapt to different damage characteristics of archive data, and define a multi-layer adaptive fitness function: in, is the matching degree between the local search point and the boundary of the damaged area, so that the morphological characteristics of the filling point and the damaged area are consistent. For contextual consistency, measure whether the infill data conforms to the logic and content structure of the original archival data. is the texture similarity, which measures the similarity of the texture features of the filling point and the adjacent restored data. To smooth the local data, make the filling data continuous in the local area. It is the information entropy increment, which measures the contribution of the filling point to the overall information completeness; S43. In the archive filling task, the search direction is guided by combining the boundary information of the archive data, and the texture gradient guidance is introduced on the basis of the traditional gray wolf optimization search to construct the gray wolf optimization search based on texture guidance: in, is the search position of the kth gray wolf in the mth layer optimization group at the t+1th iteration, indicating the position change of the filling candidate point during the optimization process. Represent the positions of the three gray wolves in the optimization group, namely the optimal solution, the second optimal solution, and the third optimal solution. m To control the coefficient of the search step size and convergence speed of the gray wolf individual, Respectively represent the distance between the gray wolf individual and the optimal solution, the second optimal solution, and the third optimal solution, B m For the gradient guide item: in, Represents the local gradient direction, guiding the filling point to converge to the optimal texture matching direction, γ2 is the gradient guidance weight; S44. In the archival data restoration task, a hybrid step-size strategy is used to optimize the search for individual gray wolves, and the search step-size is dynamically adjusted to balance accuracy and convergence speed: Among them, 2·(1-t / T max ) controls the global convergence trend, λ m N(0,1) enhances search diversity through hierarchical perturbations; S45. Introduce a local convergence detection mechanism. If a search point is K consecutive stuck If the fitness is not improved after a round of iterations, a jump is performed: S46. After the Grey Wolf Optimization Algorithm converges, the optimal filling point is selected to finally form the locally optimized filling data: in, is the coordinates of k+2 search individuals in the optimized m-th layer of the population, F m (p k+2 ) is the optimized multi-level adaptive fitness function.
9. The artificial intelligence-based archive management method according to claim 8, characterized in that: The texture similarity The calculation is as follows: Among them, G(x,y) is the local gradient feature, N k is the filled data point in the neighborhood of the search point; The local data smoothness The calculation is as follows: Where I(x,y) represents the pixel value or text feature value of the archive data; The information entropy increment The calculation is as follows: Among them, H before and H after are the entropy values of the damaged area before and after filling, respectively.
10. The artificial intelligence-based archive management method according to claim 2, characterized in that: The S5 comprises the following steps: S51. Obtain the locally optimized filled data and the preprocessed archive data, normalize the filled data according to the data format, storage structure and encoding method, perform preliminary fusion on the normalized filled data and the preprocessed archive data, and construct an initial data fusion set; S52. Analyze the local features of the area where the filling data is located, extract the local feature vector of the damaged area, and calculate the local similarity between the filling data and the adjacent area, including the boundary morphology similarity, local structure consistency and content feature matching degree, and judge whether the filling data conforms to the local style characteristics of the original file based on the local similarity calculation result. When the local similarity is lower than the preset threshold, adjust the filling data; S53. Match the style, semantics and logical consistency of the fused archive data based on the pattern characteristics of the historical archive data, analyze the style adaptability of the filling data according to the language characteristics, typesetting style and data structure rules of the archive content, and correct the filling data with a matching degree lower than the threshold through the style adjustment function, so that the filling data is consistent with the original archive in terms of font format, text semantics, structural logic and contextual coherence; S54. After completing the local similarity adjustment and style matching, the overall filling result is evaluated globally, the matching degree between the filling data and the entire archive content is calculated, and based on the global matching result, the data with filling quality higher than the threshold is screened, and the filling data with low matching degree is adjusted again, so that the filling data and the entire archive are optimally matched in terms of content logic, format layout, and semantic consistency; S55. After completing the local and global consistency matching, the final screened filled data is archived to generate preliminary corrected archival data.
Citation Information
Patent Citations
Code completion method based on artificial intelligence
CN118276913A
Wind power plant SCADA (Supervisory Control And Data Acquisition) data restoration method based on multiple correlation learning
CN118445556A
File digital quality detection method based on real estate
CN119110024A
XGBoost low-voltage transformer area missing voltage completion method
CN119180456A
Interactive context-based text completions
US20180101599A1
Cited By
File electronic information data recovery method
CN120631871A
A method for recovering electronic information data from archives
CN120631871B
Notebook computer backlight source bottom plate structure optimization method
CN120805331A
PDF intelligent retrieval and generation method and system based on RAG
CN120994845A
RAG-based pdf intelligent retrieval and generation method and system
CN120994845B