Archival digitization methods and systems based on intelligent image enhancement and automatic classification
Through technologies such as intelligent image enhancement, intelligent text recognition, semantic understanding and dynamic multi-dimensional classification, problems such as image processing, text recognition, classification and knowledge extraction in archive digitization are solved, and efficient and accurate archive digitization and management are achieved.
Patent Information
- Application Number
- CN202411481796.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-23
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-10-23
AI Technical Summary
The existing technology has problems in the process of archive digitization, low accuracy in OCR technology when processing handwritten and ancient characters, errors in archive classification, limited knowledge extraction to superficial text information, low efficiency in traditional methods when processing high-dimensional feature vectors, and lack of cross-border knowledge correlation capabilities in archive digitization systems.
The methods of intelligent image enhancement and automatic classification are adopted, including preprocessing, adaptive multi-scale image enhancement, intelligent text recognition and layout analysis, semantic understanding and knowledge extraction, and dynamic multi-dimensional archive classification and index construction.
It realizes all-round archive processing from pixel level to semantic level, improves the accuracy of text recognition and the efficiency of processing flow. Dynamic classification and multi-dimensional indexing technology enable the archive management system to adapt to the ever-changing user needs and archive structure. The construction of knowledge graphs provides possibilities for in-depth mining of archives and cross-domain applications.
Smart Images

Figure CN119049066B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the field of archive digitization, in particular to an archive digitization method and system based on intelligent image enhancement and automatic classification. Background Art
[0002] At present, research on archive digitization has made some progress. In terms of image enhancement, traditional histogram equalization and Retinex algorithms have been widely used to improve the quality of archive images. For text recognition, optical character recognition (OCR) technology based on convolutional neural network (CNN) has improved recognition accuracy. In document classification, machine learning algorithms such as support vector machine (SVM) and naive Bayes are used for automatic classification. Some studies have also tried to use deep learning models such as BERT for text semantic analysis and topic extraction. In terms of archive retrieval, technologies such as inverted index and vector space model are widely used in full-text retrieval systems. However, these technologies still have many limitations when facing complex and diverse archival materials.
[0003] Despite the progress made in research, there are still many technical challenges in practical applications. First, in terms of image enhancement, the existing algorithms do not work well for historical archive images that are severely faded or partially damaged, and it is difficult to effectively balance detail enhancement and noise suppression. Second, the accuracy of current OCR technology drops significantly when processing handwritten, ancient characters or documents with complex layouts, especially for documents with sloppy handwriting or traces of erasure, which are almost impossible to recognize. Furthermore, the archive classification system often misclassifies when dealing with complex archives across disciplines and themes, and it is difficult to capture the implicit semantic associations between documents. In addition, existing knowledge extraction technologies are often limited to surface text information and it is difficult to understand the deep semantics and contextual relationships in archives. In terms of index construction for large-scale archives, traditional methods are inefficient in processing high-dimensional feature vectors and are difficult to support complex multi-dimensional queries. Finally, most of the existing archive digitization systems are isolated and lack effective cross-library and cross-domain knowledge association capabilities, which limits the in-depth utilization and value mining of archive resources. The existence of these technical difficulties has restricted the progress and application effect of archive digitization, and innovative solutions are urgently needed. Summary of the invention
[0004] The purpose of the invention is to provide an archive digitization method and system based on intelligent image enhancement and automatic classification to solve the above-mentioned problems existing in the prior art.
[0005] The technical solution is a method for digitalizing archives based on intelligent image enhancement and automatic classification, including the following steps:
[0006] S1. Obtaining the original archive image and performing preprocessing to obtain a high-quality image after preprocessing; wherein the preprocessing includes contrast enhancement, noise reduction, tilt correction and boundary clipping;
[0007] S2. Based on the pre-processed high-quality image, an adaptive multi-scale image enhancement process is performed to obtain an enhanced image with improved quality; wherein the adaptive multi-scale image enhancement process includes wavelet transform, coefficient adjustment, nonlinear sharpening and local contrast enhancement;
[0008] S3. Based on the enhanced image, intelligent text recognition and layout analysis are performed to obtain the text content and logical structure of the document; wherein the intelligent text recognition and layout analysis include text area extraction, adaptive binarization, character segmentation, feature extraction, character recognition and document structure analysis;
[0009] S4. Based on the text content and logical structure of the document, semantic understanding and knowledge extraction are performed to obtain the core semantic information of the document; semantic understanding and knowledge extraction include named entity recognition, keyword extraction, concept map construction, importance calculation, semantic annotation and relationship extraction;
[0010] S5. Based on the logical structure and core semantic information of the document, dynamic multi-dimensional archive classification and index construction are performed to obtain the classification index structure of the archive; wherein the dynamic multi-dimensional archive classification and index construction includes feature vector construction, dimensionality reduction processing, dynamic clustering, classification tree construction and multi-dimensional index construction.
[0011] The archive digitization system based on intelligent image enhancement and automatic classification includes:
[0012] at least one processor; and,
[0013] a memory communicatively connected to at least one of the processors; wherein,
[0014] The memory stores instructions that can be executed by the processor, and the instructions are used to be executed by the processor to implement the archive digitization method based on intelligent image enhancement and automatic classification.
[0015] Beneficial effects: The present invention realizes all-round archive processing from pixel level to semantic level, making historical archives easy to understand and use; it not only improves the accuracy of text recognition, but also improves the efficiency of the entire processing flow; the application of dynamic classification and multi-dimensional indexing technology enables the archive management system to adapt to the ever-changing user needs and archive structure. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 It is a flow chart of the present invention.
[0017] Figure 2This is a flow chart of step S1 of the present invention.
[0018] Figure 3 This is a flow chart of step S2 of the present invention.
[0019] Figure 4 This is a flow chart of step S3 of the present invention.
[0020] Figure 5 This is a flow chart of step S4 of the present invention.
[0021] Figure 6 This is a flow chart of step S5 of the present invention. DETAILED DESCRIPTION
[0022] like Figure 1 As shown, the present application proposes a method for digitizing archives based on intelligent image enhancement and automatic classification, comprising the following steps:
[0023] S1. Obtaining the original archive image and performing preprocessing to obtain a high-quality image after preprocessing; wherein the preprocessing includes contrast enhancement, noise reduction, tilt correction and boundary clipping;
[0024] S2. Based on the pre-processed high-quality image, an adaptive multi-scale image enhancement process is performed to obtain an enhanced image with improved quality; wherein the adaptive multi-scale image enhancement process includes wavelet transform, coefficient adjustment, nonlinear sharpening and local contrast enhancement;
[0025] S3. Based on the enhanced image, intelligent text recognition and layout analysis are performed to obtain the text content and logical structure of the document; wherein the intelligent text recognition and layout analysis include text area extraction, adaptive binarization, character segmentation, feature extraction, character recognition and document structure analysis;
[0026] S4. Based on the text content and logical structure of the document, semantic understanding and knowledge extraction are performed to obtain the core semantic information of the document; semantic understanding and knowledge extraction include named entity recognition, keyword extraction, concept map construction, importance calculation, semantic annotation and relationship extraction;
[0027] S5. Based on the logical structure and core semantic information of the document, dynamic multi-dimensional archive classification and index construction are performed to obtain the classification index structure of the archive; wherein the dynamic multi-dimensional archive classification and index construction includes feature vector construction, dimensionality reduction processing, dynamic clustering, classification tree construction and multi-dimensional index construction.
[0028] like Figure 2 As shown, according to one aspect of the present application, step S1 is further:
[0029] S11. Scan the original file using a high-resolution scanner to obtain an original file image with a resolution not less than a preset threshold;
[0030] S12, dividing the original archive image into a predetermined number of small areas of equal size; calculating the histogram of the pixel grayscale value for each small area; based on the histogram of each small area, using a histogram equalization algorithm, recombining all the small areas to obtain contrast-enhanced image data;
[0031] S13, calculating the local mean and variance of the contrast-enhanced image data as local statistical characteristics; constructing a Wiener filter based on the local statistical characteristics; using the Wiener filter to process the contrast-enhanced image data to obtain denoised image data;
[0032] S14, based on the image data after noise reduction, using the Hough transform algorithm to detect the straight line features in the image; calculating the angle between the straight line features and the horizontal line to obtain the tilt angle of the image; based on the tilt angle, using the bilinear interpolation algorithm to perform rotation correction to obtain the corrected image data;
[0033] S15. Based on the corrected image data, use the Canny edge detection algorithm to extract the image edge; based on the image edge, perform contour analysis to identify the boundary information; based on the boundary information, crop the corrected image data to obtain the final pre-processed high-quality image.
[0034] In one embodiment of the present application, a high-resolution scanner is used to scan the original file to obtain original image data with a resolution of no less than 600 dots per inch. The obtained original image data is stored as a digital image file, denoted as I raw . Digital image file I raw The image is divided into multiple small areas of equal size. The histogram of the pixel grayscale value is calculated for each small area, and the histogram equalization algorithm is applied. All the processed small areas are recombined to obtain the contrast-enhanced image data, which is recorded as I contrast . Calculate image data I contrast The local mean and variance of the image data I are used as local statistical characteristics. Based on the calculated local statistical characteristics, a Wiener filter is constructed. The constructed Wiener filter is used to contrast After processing, the image data after noise reduction is obtained, which is recorded as I denoised .
[0035] For the denoised image data I denoised The Hough transform algorithm is used to detect the straight line features in the image. The angle between the detected straight line and the horizontal line is calculated, and the angle value with the highest frequency is taken as the overall tilt angle θ of the image. Based on the obtained tilt angle θ, the bilinear interpolation algorithm is used to de-noise the image data I denoisedPerform rotation correction to obtain the corrected image data, denoted as I corrected . For the corrected image data I corrected Apply Gaussian filter for smoothing to obtain the filtered image I smoothed . Calculate image I smoothed The gradient in the x and y directions is obtained as the gradient image I gradient For the gradient image I gradient Perform non-maximum suppression to obtain the edge candidate image I candidate . For edge candidate image I candidate Apply the double threshold algorithm to obtain the preliminary edge image I edge . For edge image I edge Perform edge tracking, connect broken edges, and obtain a complete edge image I edgecomplete For the complete edge image I edgecomplete Perform morphological processing, including dilation and thinning operations, to obtain the optimized edge image I edgeoptimized . Use the contour tracking algorithm to process the optimized edge image I edgeoptimized , extract the outer boundary of the document and obtain the boundary coordinate set B. Based on the boundary coordinate set B, calculate the minimum bounding rectangle R. Use rectangle R to correct the original image I corrected Crop to obtain the final preprocessed image data, denoted as I preprocessed .
[0036] In another embodiment of the present application, Wiener filter optimization is performed as follows: using the image edge region to estimate the noise variance σ² n ; Estimate the signal variance σ² using the entire image s ; Adaptive window size selection: the initial window size is 3x3; if the local σ² s / σ² n >10, increase the window to 5x5; if the local σ² s / σ² n <2, reduce the window to 1x1; filter using adaptive window size and estimated variance.
[0037] This embodiment improves the quality and readability of archival images through image preprocessing technology, laying the foundation for subsequent text recognition and content analysis. Specifically, the use of high-resolution scanning technology ensures that the details of the original image are retained, and the adaptive histogram equalization algorithm effectively enhances the image contrast, especially for archival images that are faded or lack contrast due to age. The application of Wiener filter not only effectively reduces image noise, but also retains the edge and detail information of the image, which is crucial for subsequent text recognition. The introduction of Hough transform for tilt detection and correction solves the common image tilt problem in the archival scanning process, ensures the horizontal alignment of text lines, and improves the accuracy of subsequent OCR. The application of edge detection and contour analysis technology accurately extracts the document boundary and removes redundant background information, which not only reduces the calculation amount of subsequent processing, but also improves the accuracy of text area positioning. The comprehensive application of this series of preprocessing steps allows even archival images of poor quality and age to be effectively processed, expands the scope of archives that can be digitized, and provides strong technical support for archival digitization.
[0038] According to one aspect of the present application, step S12 is further:
[0039] S121, using a sliding window method to divide the original archive image into a predetermined number of overlapping local areas; calculating a histogram of each local area; calculating a cumulative distribution function of each histogram; performing contrast limitation processing on each cumulative distribution function to obtain a processed cumulative distribution function;
[0040] S122, remapping based on the processed cumulative distribution function to obtain a mapping list; based on the mapping list, remapping the gray value of each pixel in the original archive image to obtain a new image matrix;
[0041] S123, performing global contrast stretching on the new image matrix to obtain contrast-enhanced image data.
[0042] In one embodiment of the present application, the original image data I obtained in step S11 is read raw , use the sliding window method to divide the image into multiple overlapping local areas: set the window size to w×w pixels, the step size to s pixels, start from the upper left corner of the image, slide the window along the row and column directions, and obtain a series of overlapping local areas. Store these local area information in the area list R in the memory. Read the area list R, and calculate its histogram Hi for each local area Ri: traverse each pixel in Ri, count the number of pixels with different gray levels, and generate a 256-level grayscale histogram. Store the histogram information of all local areas in the histogram list H.
[0043] Read the histogram list H and calculate the cumulative distribution function (CDF) for each histogram Hi: Starting from gray level 0, accumulate the number of pixels for each gray level and divide by the total number of pixels to obtain the CDF value. Store the calculated CDF information in the CDF list C. Read the CDF list C and perform contrast limiting processing on each CDF: Set the contrast limiting threshold t. If the slope of a certain gray level in the CDF exceeds t, evenly distribute the excess to other gray levels. Repeat this process until the slopes of all gray levels do not exceed t. Update and store the processed CDF information in C. Read the updated CDF list C and remap each CDF: Map the original gray levels to new gray levels through the CDF so that the processed image has a more uniform gray distribution. Store the mapping relationship in the mapping list M.
[0044] Read the region list R and the mapping list M, and remap the gray values of each pixel in the original image data I raw : For each pixel in the image, find all the overlapping regions it belongs to, calculate the weighted average of the remapped gray values of these regions, where the weight is inversely proportional to the distance from the pixel to the center of the region. Store the processed pixel values in a new image matrix. Read the new image matrix and perform global contrast stretching on it: Calculate the global minimum and maximum gray values of the image and linearly map them to the range of 0 - 255. This process obtains contrast-enhanced image data, denoted as I contrast . Save the contrast-enhanced image data I contrast as a file or store it in memory.
[0045] In another embodiment of the present application, perform adaptive histogram equalization (AHE) optimization, specifically: Divide the image into 8x8 small blocks, each block being 1 / 64 of the original image size; Introduce a contrast limiting parameter α with an initial value set to 0.01; Adaptive parameter adjustment: Calculate the local variance σ² of each block. If σ² > T (T is a preset threshold, with an initial value of 100), then α = α * 0.9; If σ² < T / 2, then α = α * 1.1; Use the updated α value to perform contrast-limited adaptive histogram equalization on each block; Use bilinear interpolation to smooth the transition between blocks.
[0046] This embodiment improves the contrast and clarity of the archival image through the adaptive histogram equalization processing technology. Specifically, the application of the sliding window method realizes the local analysis of the image, so that the enhancement process can adapt to the characteristics of different areas of the image. The introduction of the overlapping area processing mechanism effectively avoids the block effect that is easily caused by the traditional block processing method, and ensures the visual continuity of the enhanced image. The adaptive histogram calculation and cumulative distribution function (CDF) processing are performed on each local area, so that the enhancement effect can be dynamically adjusted according to the local content. The introduction of contrast limitation technology effectively prevents excessive amplification of noise and maintains the overall quality of the image. The application of the multi-region weighted average strategy realizes a smooth transition of the local enhancement effect and avoids the appearance of hard boundaries. The post-processing step of global contrast stretching further optimizes the overall dynamic range of the image. This embodiment not only improves the readability of the image, but also maintains the texture features and detail information of the original image. For the digitalization of archives, this means that various types of historical archive images can be effectively processed, including handwritten documents, printed materials, photos, etc., even those archives that have faded or damaged due to poor storage conditions can be improved. This expands the scope of archives that can be digitized, improves the success rate of subsequent text recognition, and lays a solid foundation for the effective extraction and utilization of archive content. Especially for large-scale archive digitization projects, this adaptive image enhancement technology can improve processing efficiency and reduce manual intervention while ensuring the consistency and high quality of processing results.
[0047] According to another aspect of the present application, the noise reduction image data I denoised The specific application of Hough transform algorithm is as follows: Read the noise reduction image data I denoised , use Sobel operator for edge detection: apply Sobel operator in horizontal and vertical directions to the denoised image data I denoised Convolution is performed to obtain a horizontal gradient map Gx and a vertical gradient map Gy. The horizontal gradient map Gx and the vertical gradient map Gy are stored in memory.
[0048] Calculate the gradient magnitude and direction of each pixel. For each pixel, use the formula G=sqrt(Gx 2 +Gy 2 ) calculates the gradient magnitude, and uses the formula θ=arctan(Gy / Gx) to calculate the gradient direction. The calculated gradient magnitude map G and gradient direction map θ are stored in memory. Based on the gradient magnitude map G and gradient direction map θ, non-maximum suppression is performed: the gradient magnitude of each pixel is compared with the gradient magnitude of its neighboring pixels along the gradient direction, and if it is not a local maximum, it is set to zero. The processed edge map E is stored in memory.
[0049] Construct a Hough transform accumulator: For each non-zero pixel (x, y) in the edge map E, calculate the parameters (ρ, θ) of all possible lines it may pass through, where ρ = xcos(θ) + ysin(θ). Increase the count at the corresponding position of the accumulator. Store the constructed accumulator A in memory. Read accumulator A and perform peak detection: Use a sliding window to find local maxima in A. These maxima correspond to the most significant straight lines in the image. Store the detected peak list P in memory.
[0050] Calculate the main direction: For each peak in the peak list P, convert its θ value into an angle, and calculate the weighted average of these angles, with the weight being the size of the peak. The average angle obtained is the main tilt angle θ of the image main Based on the main tilt angle θ main , perform angle optimization: at the main tilt angle θ main A fine search is performed in a small area nearby, using the original image I denoised Calculate the variance of each angle and select the angle with the smallest variance as the final tilt angle θ final The final tilt angle θ final Save as a file or store in memory for use in subsequent steps.
[0051] This embodiment uses Hough transform and image analysis technology to accurately detect and correct the tilt angle of the archive image, solving the common image tilt problem in the archive scanning process. Specifically, the application of the Sobel operator realizes efficient edge detection and provides reliable input for subsequent line detection. The introduction of gradient direction information not only improves the accuracy of the Hough transform, but also reduces the computational overhead. The application of non-maximum suppression technology effectively filters redundant edge information and improves the accuracy of line detection. The peak detection in the accumulator space adopts a sliding window strategy, which can better adapt to images with different degrees of tilt. The calculation of the main direction takes into account the contribution of multiple peaks, enhancing the adaptability of the algorithm to complex layouts. The introduction of the angle optimization step based on variance minimization further improves the accuracy of the tilt angle estimation. The comprehensive application of this series of technologies enables the tilt correction process to be accurate to 0.1 degrees, surpassing the accuracy of traditional methods. For archive digitization work, it can not only process standard printed documents, but also effectively correct the tilt problem of non-standardized archives such as handwritten manuscripts and historical documents. Accurate tilt correction directly improves the accuracy of subsequent text recognition, especially for structured content such as tables and multi-column texts, the effect of tilt correction is particularly obvious. In addition, this embodiment can also identify and handle local tilt problems, which is particularly effective for handling wrinkled and curled old archives. It also shows adaptability to complex archives containing non-text elements such as pictures and seals, ensuring the correct alignment of the entire image, laying the foundation for subsequent layout analysis and content extraction.
[0052] In another embodiment of the present application, the Canny edge detection algorithm is applied to the corrected image data to extract the image edge, specifically: the image is smoothed using a Gaussian filter to reduce the influence of noise. The two-dimensional convolution kernel of the Gaussian filter is defined as: G(x,y) = (1 / (2πσ 2 )) * e -(x2+y2) / (2σ2) , where σ is the standard deviation of the Gaussian distribution, usually 1.4. Calculate the gradient magnitude and direction of the image. Use the Sobel operator to calculate the gradients in the x and y directions respectively: Gx = [-1 0 1; -2 0 2; -1 0 1] * I; Gy = [-1 -2 -1; 0 0 0; 1 2 1] * I; where I is the input image and * indicates the convolution operation. Then calculate the gradient magnitude and direction: G = sqrt(Gx 2 + Gy 2);θ= arctan(Gy / Gx). Perform non-maximum suppression on the gradient magnitude image. Quantize the gradient direction θ into four directions: 0°, 45°, 90°, and 135°. For each pixel, compare the gradient magnitude of the current pixel with its neighbors on both sides along the gradient direction. If the current pixel is not a local maximum, suppress it to 0.
[0053] Apply the dual threshold algorithm for edge detection. Select two thresholds T low and T high , usually T high = 2 * T low If the pixel gradient value is greater than T high , it is marked as a strong edge point; if the pixel gradient value is less than T low , then suppress; if the pixel gradient value is between T low and T high If the weak edge point is between , it is marked as a weak edge point. Starting from the strong edge point, recursively track the weak edge point along the edge direction. If the weak edge point is connected to the strong edge point, it is retained; otherwise, it is suppressed.
[0054] The obtained edge image is subjected to morphological processing, such as dilation and thinning, to obtain a continuous edge contour. The outer boundary of the document is extracted using a contour tracking algorithm (such as the Moore-Neighbor tracking algorithm). The algorithm starts from any edge pixel in the edge image and traverses the adjacent pixels in a clockwise or counterclockwise direction until it returns to the starting point to form a closed contour. Based on the extracted outer boundary information, the minimum enclosing rectangle is calculated and used to crop the original image to remove the redundant background area. The cropped image is saved as the final preprocessed image data.
[0055] like Figure 3 As shown, according to one aspect of the present application, step S2 is further:
[0056] S21, performing a one-dimensional wavelet transform on each row of the preprocessed high-quality image to obtain a row transform result; performing a one-dimensional wavelet transform on each column of the row transform result to obtain a two-dimensional wavelet transform result; decomposing the two-dimensional wavelet transform result into a predetermined number of sub-bands, including low-frequency approximate coefficients, horizontal high-frequency coefficients, vertical high-frequency coefficients, and diagonal high-frequency coefficients, to form a wavelet coefficient set;
[0057] S22, based on the wavelet coefficient set, calculating the local statistical characteristics of each sub-band, including the local mean and local variance of each sub-band; based on the local statistical characteristics, adjusting the sub-band coefficients in the wavelet coefficient set to obtain an adjusted coefficient set;
[0058] S23, based on the adjusted coefficient set, extracting high-frequency subband coefficients and low-frequency subband coefficients; applying an improved nonlinear sharpening operator to the high-frequency subband coefficients to obtain sharpened high-frequency subband coefficients; combining the sharpened high-frequency subband coefficients with the extracted low-frequency subband coefficients to form a new coefficient set;
[0059] S24, performing a one-dimensional inverse wavelet transform on each column in the new coefficient set to obtain a column inverse transform result; performing a one-dimensional inverse wavelet transform on each row of the column inverse transform result to obtain a final reconstructed image;
[0060] S25, dividing the final reconstructed image into a predetermined number of small blocks of equal size; performing histogram equalization on each small block to obtain processed small blocks; and using a bilinear interpolation method to recombine the processed small blocks to obtain an enhanced image with improved quality.
[0061] In one embodiment of the present application, the pre-processed image data I is read preprocessed . Preprocess image data I preprocessed Apply discrete wavelet transform: First, perform a one-dimensional wavelet transform on each row of the image to obtain the row transform result I rowtransformed ; Then transform the result I rowtransformed Perform a one-dimensional wavelet transform on each column of to obtain the final two-dimensional wavelet transform result. Decompose the two-dimensional wavelet transform result into four sub-bands: low-frequency approximate coefficient LL, horizontal high-frequency coefficient LH, vertical high-frequency coefficient HL and diagonal high-frequency coefficient HH. The four sub-band coefficient sets are denoted as C wavelet = {LL, LH, HL,HH}.
[0062] For the subband coefficient set C wavelet Each subband coefficient matrix in is processed separately. For each subband, its local statistical characteristics are calculated, including the local mean μ and the local variance σ². Each coefficient C is adjusted using the formula C'= C* (1+k * log(1+σ²)), where k is a preset adjustable parameter. After all subband coefficients are adjusted, the adjusted coefficient set C is obtained. adjusted = {LL', LH', HL', HH'}.
[0063] Extract coefficient set C adjustedThe high-frequency subband coefficients LH', HL', HH' in . Apply the improved nonlinear sharpening operator to these high-frequency coefficients. The specific operation is: for each high-frequency coefficient x, apply the function S(x)=x+α*sign(x)*(1-exp(-β|x|)) for processing, where α and β are preset adjustable parameters, sign( ) represents the sign function, and exp( ) represents the exponential function. After the processing is completed, the sharpened high-frequency subband coefficients LH'', HL'', HH'' are obtained. These sharpened high-frequency coefficients together with the subband coefficient LL' form a new coefficient set C sharpened = {LL', LH'', HL'', HH''}.
[0064] For the new coefficient set C sharpened Apply inverse discrete wavelet transform to reconstruct the image. First, the new coefficient set C sharpened Perform one-dimensional inverse wavelet transform on each column in and obtain the column inverse transform result I colreconstructed Then the inverse column transform result I colreconstructed Each row of is transformed by one-dimensional inverse wavelet transform, which converts the coefficients in wavelet domain back to spatial domain to obtain the final reconstructed image I reconstructed . For the final reconstructed image I reconstructed Apply the local contrast limited adaptive histogram equalization (CLAHE) algorithm: reconstruct the image I reconstructed The image is divided into multiple small blocks of equal size. Each small block is subjected to histogram equalization, and the amplitude of contrast enhancement is limited to avoid excessive enhancement of noise. The processed small blocks are reassembled using the bilinear interpolation method to obtain the final enhanced image data I enhanced . The enhanced image data I enhanced Save as a file or store in memory for use in subsequent steps.
[0065] This embodiment improves the quality and readability of archival images through adaptive multi-scale image enhancement technology. Specifically, the application of wavelet transform realizes multi-scale decomposition of images, so that enhancement processing can be performed separately in different frequency domains. An adaptive coefficient adjustment method based on local statistical features is introduced to dynamically adjust the enhancement strength according to the local mean and variance of the image, effectively avoiding the problem of over-enhancement or under-enhancement that is easily caused by traditional global enhancement methods. The application of nonlinear sharpening operators, especially the edge protection mechanism therein, effectively suppresses noise amplification while enhancing image details, which is crucial to improving the clarity and readability of text. The introduction of local contrast limited adaptive histogram equalization (CLAHE) further optimizes the local contrast of the image, making the separation of text and background more obvious. This embodiment not only improves the overall quality of the image, but also particularly enhances the clarity and recognizability of the text area, providing high-quality input for subsequent text recognition and content analysis. For archive digitization work, this means that more archives that were originally difficult to digitize due to poor image quality can be processed, which improves the coverage and success rate of archive digitization.
[0066] According to one aspect of the present application, step S22 is further:
[0067] S221, read the wavelet coefficient set {LL, LH, HL, HH}, and calculate the global statistical features for each sub-band. Calculate the mean μ of each sub-band global and standard deviation σ global , store these global statistical features in the feature list F in memory global middle.
[0068] S222, read wavelet coefficient set and global feature list F global , use the sliding window method to calculate local statistical features. For each subband, use a window of size w×w and calculate the local mean μ within the window local and the local standard deviation σ local , these local statistical features are stored in the feature map F local middle.
[0069] S223, read characteristic graph F local and F global , calculate the adaptive gain factor for each coefficient. Use the formula g=(σ global / σ local ) α *(μ local / μ global ) β Calculate the gain factors, where α and β are adjustable parameters. Store the calculated gain factor map G in memory.
[0070] S224, read the gain factor graph G and the original wavelet coefficient set, and adjust the coefficients. For each coefficient C of each subband, use the formula C' = C * (1+ k* log(1 +g)) for adjustment, where k is a preset adjustable parameter and g is the gain factor of the corresponding position. Store the adjusted coefficients in a new coefficient set {LL', LH', HL', HH'}.
[0071] S225, read the adjusted coefficient set {LL', LH', HL', HH'}, and perform edge protection processing. Detect the significant coefficients in each high-frequency subband. If the absolute value of a coefficient is greater than t times the mean of its neighborhood (t is a threshold), mark it as an edge point. Perform additional enhancement on the coefficients of these edge points. Update and store the processed coefficient set in {LL', LH', HL', HH'}.
[0072] S226, read the updated coefficient set {LL', LH', HL', HH'}, and perform coefficient truncation. Set a threshold Th, and set the high-frequency coefficients whose absolute values are less than Th to zero to suppress noise. Update the processed coefficient set again and store it in {LL', LH', HL', HH'}, and save it as a file or store it in memory for use in subsequent steps.
[0073] This embodiment uses adaptive coefficient adjustment technology to achieve fine enhancement of archival images in the wavelet domain, improve the clarity and readability of the image, and effectively suppress noise. Specifically, the combined use of global and local statistical features enables the enhancement process to consider both the overall characteristics and local details of the image. The introduction of a local feature calculation method based on a sliding window enables the coefficient adjustment to adapt to the characteristics of different regions of the image, especially for complex archival images containing multiple elements such as text, pictures, and tables, and performs well. The design of the adaptive gain factor realizes dynamic adjustment of different frequency components by considering the ratio of local and global statistical features, effectively enhancing the edges of the text while suppressing background noise. The application of nonlinear mapping functions, especially the introduction of logarithmic functions, makes the enhancement process more sensitive to low-contrast areas, which is particularly effective for processing historical archives with faded or insufficient contrast. The application of the edge protection mechanism effectively retains the sharp edges of text and images by identifying and specially processing significant coefficients, thereby improving the readability of the text. The introduction of coefficient truncation technology effectively reduces the noise level by suppressing low-amplitude high-frequency coefficients while retaining important image details. This embodiment makes the archival image enhancement process adaptive and robust. For archive digitization, this means being able to process archive images of various qualities and types, from clear modern documents to blurred historical manuscripts, and improving their quality. Especially for those precious archives that have faded or been damaged due to age, it is possible to restore and enhance the image content without introducing artifacts, providing high-quality input for subsequent text recognition and content analysis.
[0074] According to one aspect of the present application, step S23 is further:
[0075] S231, read the adjusted high frequency subband coefficients LH', HL', HH', and calculate the energy of each subband. For each subband, calculate the square sum of all coefficients to obtain the energy value E LH , E HL , E HH . Store these energy values in the energy list E in memory.
[0076] S232, read the energy list E, and calculate the adaptive parameters α and β. Use the formula α=k1*(E LH +E HL +E HH ) / (3*E LL ) to calculate α, and use the formula β = k2 / (1+exp(-k3*α)) to calculate β, where k1, k2, k3 are preset constants, E LL is the energy of the low frequency subband. The calculated α and β are stored in the memory.
[0077] S233, read the parameters α and β, and the high-frequency subband coefficients, and apply a nonlinear function to each high-frequency coefficient x. Use the formula S(x) =x+α* sign(x) * (1 - exp(-β|x|)) to process each coefficient. Store the processed coefficients in the new high-frequency subbands LH'', HL'', HH''.
[0078] S234, read the processed high frequency sub-bands LH'', HL'', HH'', and adjust the coefficients. Calculate the standard deviation σ of each sub-band, and further adjust each coefficient using the formula x' = x * (1 +γ* log(1 +σ)), where γ is an adjustable parameter. Update and store the adjusted coefficients in LH'', HL'', HH''.
[0079] S235, read the updated high-frequency subband LH'', HL'', HH'' and low-frequency subband LL', and perform subband fusion. Use the weighted average method to fuse the subbands of adjacent scales, and the weights are dynamically adjusted according to the energy of the subbands. Store the fused subband coefficients in the new coefficient set {LL', LH''', HL''', HH'''}.
[0080] S236, read the fused coefficient set {LL', LH''', HL''', HH'''}, and perform edge suppression. Detect isolated large coefficients in the high-frequency subband, and if there is no other large coefficient support around a large coefficient, attenuate it to avoid artifacts. Update the processed coefficient set and store it in {LL', LH''', HL''', HH'''}, and save it as a file or store it in memory for use in subsequent steps.
[0081] This embodiment achieves fine enhancement of archival images, especially the enhancement of text edges and details, through nonlinear sharpening processing technology, while effectively suppressing the generation of noise and artifacts. Specifically, the adaptive parameter calculation method based on wavelet subband energy enables the sharpening intensity to be dynamically adjusted according to the frequency characteristics of the image, which is particularly effective for processing archival images of different types and qualities. The constructed nonlinear mapping function achieves differentiated processing of different amplitude coefficients by introducing adjustable α and β parameters, effectively enhancing edge and texture information while avoiding over-sharpening. The application of subband fusion technology, by considering adjacent scale information, realizes the organic combination of multi-scale features and improves the naturalness and coherence of the enhancement effect. The introduction of the edge suppression mechanism effectively prevents the generation of ringing effects and other sharpening artifacts by identifying and specially processing isolated large coefficients. The comprehensive application of this series of technologies enables the sharpening process of archival images to both improve image clarity and maintain the natural appearance of the image. For archive digitization, it can effectively process various types of archive images, including printed documents, handwritten manuscripts, photos, charts, etc., making originally unclear text clear and legible, and fully displaying detailed information. In terms of text recognition, clear character edges directly improve the accuracy of OCR. For archives containing pictures or charts, it can also effectively enhance image details and improve the interpretability of images.
[0082] like Figure 4 As shown, according to one aspect of the present application, step S3 is further:
[0083] S31, traversing all pixels of the enhanced image, analyzing grayscale changes, and identifying stable extreme value regions; calculating the features of each extreme value region, including area, perimeter, and aspect ratio; based on the features of the extreme value region and a preset threshold, screening the extreme value region to obtain a potential text region set;
[0084] S32, based on the text region set, calculating the grayscale histogram of each text region; applying the Otsu method to the grayscale histogram to calculate the optimal binarization threshold; using the optimal binarization threshold to perform binarization processing on each text region to obtain a binarized text image set;
[0085] S33, based on the binary text image set, scanning each pixel in the binary text image, identifying connected regions composed of adjacent black pixels; calculating the features of each connected region, including area, aspect ratio and density; based on the features of the connected regions and preset morphological rules, screening the connected regions to obtain screened connected regions; based on the screened connected regions, forming a character image set;
[0086] S34, based on the character image set, calculating the Zernike moment of each character image to obtain a geometric feature vector; based on the geometric feature vector, using a local binary pattern algorithm to obtain a texture feature vector; concatenating the geometric feature vector and the texture feature vector to form a mixed feature vector; combining all the mixed feature vectors to obtain a feature vector set;
[0087] S35, using a pre-trained support vector machine classifier to classify each feature vector in the feature vector set to obtain a recognition result string; based on the recognition result string, forming the text content of the document;
[0088] S36. Analyze the spatial relationship between the text regions in the text region set and identify the structural elements; map the text content of the document to the structural elements and construct a tree-like data structure, which is the logical structure of the document.
[0089] In one embodiment of the present application, the final enhanced image data I is read enhanced . For the final enhanced image data I enhanced The improved Maximally Stable Extremal Regions (MSER) algorithm is applied. enhanced , analyze the grayscale changes and identify stable extreme regions. For each identified region, calculate its area, perimeter, aspect ratio and other features. Filter these regions according to the preset threshold to obtain the potential text region set T ={T1,T2,..., n}, where each Ti represents an independent text region.
[0090] Each text region Ti in the text region set T is processed. First, the grayscale histogram Hi of each text region Ti is calculated. Then the Otsu method is applied to the grayscale histogram Hi to calculate the optimal binarization threshold θi. Each text region Ti is binarized using the threshold θi: pixels greater than θi are set to 1 (white), and pixels less than or equal to θi are set to 0 (black). After processing all text regions Ti in the text region set T, a binary text image set B ={B1, B2, ..., Bn} is obtained.
[0091] The improved connected component analysis algorithm is applied to each binary image Bi in the binary text image set B. The algorithm scans each pixel in each binary image Bi and identifies the connected regions composed of adjacent black pixels. For each connected region, its area, aspect ratio, density and other features are calculated. These connected regions are screened according to the preset morphological rules to remove noise and non-character regions. Each screened connected region is regarded as a separate character. After processing all the binary images Bi in the binary text image set B, the character image set C = {C1, C2, ..., Cm} is obtained.
[0092] Feature extraction is performed on each character image Ci in the character image set C. First, the Zernike moment of each character image Ci is calculated to obtain the geometric feature vector Gi. Then, the Local Binary Patterns (LBP) algorithm is applied to the geometric feature vector Ci to obtain the texture feature vector Ti. Each character image Gi and the texture feature vector Ti are concatenated to form a mixed feature vector Fi. After processing all character images Ci in the character image set C, the feature vector set F = {F1, F2, ..., Fm} is obtained.
[0093] Use a pre-trained support vector machine (SVM) classifier to classify each feature vector in the feature vector set F. The SVM classifier maps the feature vector to the corresponding character category based on its distribution. After classifying all feature vectors in the feature vector set F, the recognition result string S is obtained. Arrange the characters in the recognition result string S in the order of their positions in the original image to form a complete recognition result text.
[0094] Analyze the spatial relationship of each text region in the text region set T and identify structural elements such as titles, paragraphs, and lists. Map the text content in the recognition result string S to these structural elements. Based on the identified structural elements and text content, construct a tree data structure to represent the logical hierarchical relationship of the document. The constructed logical structure tree is denoted as L.
[0095] This embodiment realizes high-precision extraction and structured representation of text content in archive images through text recognition and layout analysis technology. Specifically, the improved MSER algorithm performs well in text area extraction, especially for texts with complex backgrounds and different font sizes, and can accurately locate them. The application of adaptive binarization technology effectively solves the problem of insufficient contrast between text and background caused by aging and fading of archives, and provides clear input for subsequent character segmentation. The introduction of a hybrid feature extraction method based on Zernike moments and LBP improves the accuracy of character recognition, especially for archive files mixed with handwriting and print, the recognition effect is obvious. The application of support vector machine (SVM) classifier, combined with the high-quality features extracted above, makes the accuracy of character recognition reach a new height. The application of layout analysis technology not only recognizes the text content, but also accurately understands the logical structure of the document, such as titles, paragraphs, tables, etc., which is crucial for subsequent content understanding and knowledge extraction. The comprehensive application of this series of technologies makes the text recognition and layout analysis in the process of archive digitization more accurate and efficient, and improves the quality and speed of archive digitization. It demonstrates adaptability and robustness for historical archives containing complex layouts, multiple fonts, and even partial damage, expanding the scope of archives that can be digitized.
[0096] According to one aspect of the present application, step S31 is further:
[0097] S311, read the final enhanced image data and perform image grayscale processing. For each pixel, calculate the weighted average of its R, G, and B channel values, with weights of 0.299, 0.587, and 0.114 respectively. Store the obtained grayscale image data as I gray .
[0098] S312, read grayscale image I gray , construct an image pyramid. Repeatedly perform the grayscale image I gray Perform 2×2 mean downsampling until the minimum side length of the image is less than 64 pixels. The resulting multi-scale image sequence is stored as an image pyramid P = {I0, I1, ..., I n}, where I0 is the original image I gray .
[0099] S313, read the image pyramid P, for each scale image I i Perform MSER region detection. Starting from the minimum grayscale value, gradually increase the threshold and maintain the connected region set. When the area change rate of a connected region is less than the preset threshold, it is marked as a candidate MSER region. The candidate region set detected at all scales is stored as R = {R0, R1, ..., R n}.
[0100] S314, read the candidate region set R, and perform multi-scale region merging. Project the regions detected at different scales to the original image scale, and calculate the overlap between the regions. If the overlap exceeds a preset threshold, merge these regions. Store the merged region set as R merged .
[0101] S315: Read the merged region set R merged , calculate the shape features of each region. Calculate the area, perimeter, aspect ratio, density and other features of the region. Store the calculated feature vector set as F = {f1, f2, ..., f m}, where m is the merged region set R merged The number of regions in .
[0102] S316, read the feature vector set F, and use the pre-trained SVM classifier to classify each region as text / non-text. Input each feature vector into the SVM classifier to obtain the probability that the region is text. Store the classification results as a probability list P = {p1, p2, ..., p m}.
[0103] S317, read probability list P and merged region set R merged , perform post-processing. Delete areas with probability lower than the threshold, merge highly overlapping text areas, and remove areas that are too small or too large. Store the processed text area set as T = {T1, T2, ..., T k}. Save T to a file or store it in memory for use in subsequent steps.
[0104] This embodiment achieves high-precision extraction of text areas in archive images, especially when processing archives with complex backgrounds, diverse fonts and layouts. Specifically, the construction of a multi-scale image pyramid enables the algorithm to process texts of different sizes at the same time, effectively solving the problem that the traditional MSER algorithm is sensitive to font size. The introduced region merging strategy effectively merges small areas belonging to the same text line or paragraph by considering the overlap of detection results of different scales, thereby improving the integrity of the text area. The calculation of shape features and the application of SVM-based classifiers improve the accuracy of text area recognition and effectively filter non-text areas such as pictures and table borders. The probability threshold and region merging technology introduced in the post-processing step further optimize the boundaries of the text area and improve the accuracy of the extraction results. For archive digitization work, it can effectively process various types of archives, including printed documents, handwritten manuscripts, newspapers, magazines, etc., and accurately locate text areas even in the case of complex backgrounds, diverse fonts, and irregular layouts. This directly improves the accuracy and efficiency of subsequent OCR, because the OCR engine can focus on processing real text areas and avoid invalid processing of non-text areas. For complex archives containing non-text elements such as pictures, charts, and seals, the ability to accurately locate text is particularly important, as it lays the foundation for subsequent layout analysis and content understanding. When processing historical archives, this embodiment shows strong adaptability and can effectively identify text areas with blurred handwriting and deformed pages due to age. This embodiment provides key technical support for the digitization and intelligent management of archives, improving the overall quality of archive digitization and the possibility of subsequent utilization.
[0105] According to one aspect of the present application, step S34 is further:
[0106] S341, based on the character image set, normalizing the size of each character image to obtain a normalized character image set;
[0107] S342, based on the normalized character image set, calculating the Zernike moment feature of each character image;
[0108] S343, based on the normalized character image set, calculating the local binary pattern feature of each character image;
[0109] S344, based on the normalized character image set, calculating the directional gradient histogram feature of each character image;
[0110] S345, performing feature fusion on the Zernike moment feature, the local binary pattern feature and the oriented gradient histogram feature to obtain a fused feature vector set;
[0111] S346. Based on the fused feature vector set, use the principal component analysis method to perform dimensionality reduction to obtain a feature vector set after dimensionality reduction; based on the feature vector set after dimensionality reduction, perform feature normalization to obtain a normalized feature vector set.
[0112] In one embodiment of the present application, a character image set C is read, and for each character image C i Normalize the size. Adjust each character image to a fixed size (such as 28×28 pixels) and scale it using the bicubic interpolation algorithm. Store the normalized character image set as C norm = {C norm1 , C norm2 , ..., C normm}.
[0113] Read the normalized character image set C norm , calculate the Zernike moments of each character image. Map the image coordinates to the unit circle, calculate Zernike polynomials of different orders and repetitions, and then calculate the projection of the image onto these polynomials. Store the resulting set of Zernike moment feature vectors as Z = {Z1, Z2, ..., Z m}. Read the normalized character image set C norm , calculate the LBP features of each character image. For each pixel, compare its size relationship with the surrounding 8 pixels and generate an 8-bit binary code. Count the LBP code distribution of the entire image. Store the obtained LBP feature vector set as L ={L1, L2, ..., L m}. Read the normalized character image set C norm , calculate the directional gradient histogram (HOG) feature of each character image. Calculate the gradient magnitude and direction of the image, divide the image into several cells, and count the histogram of the gradient direction in each cell. Store the obtained HOG feature vector set as H = {H1, H2, ..., H m}.
[0114] Read the Zernike moment feature Z, LBP feature L and HOG feature H, and perform feature fusion. For each character, concatenate its Zernike moment feature, LBP feature and HOG feature into a long vector. Store the fused feature vector set as F = {F1, F2, ..., F m}. Use principal component analysis (PCA) to reduce the dimension of the fused feature vector set F. Calculate the covariance matrix of the feature vector, solve its eigenvalues and eigenvectors, and select the eigenvectors corresponding to the first k largest eigenvalues as the projection matrix. Project each eigenvector in F to this k-dimensional space. Store the reduced eigenvector set as Freduced ={F reduced1 , F reduced2 , ..., F reducedm}.
[0115] Read the reduced dimension feature vector set F reduced , perform feature normalization. Calculate the mean and standard deviation of each feature dimension, and then normalize each feature vector to zero mean and unit variance. Store the normalized feature vector set as F normalized = {F normalized1 , F normalized2 , ..., F normalizedm}. The feature vector set F normalized Save to file or store in memory.
[0116] This embodiment achieves high-precision recognition of characters in archives through a hybrid feature extraction method. Specifically, the size normalization preprocessing of the character image ensures the consistency of feature extraction and improves the stability of subsequent recognition. The application of Zernike moments captures the overall geometric features of the characters and has good invariance to the rotation and scaling of the characters, which is particularly effective for processing tilted or unequal-sized characters. The introduction of LBP features effectively describes the local texture information of the characters and enhances the system's adaptability to font changes and local deformations. The calculation of the histogram of directional gradients (HOG) features accurately captures the edge and contour information of the characters, which is particularly important for distinguishing characters with similar glyphs. The fusion of these three features forms a multi-dimensional, highly discriminative feature vector, which improves the accuracy of character recognition. The introduction of principal component analysis (PCA) for dimensionality reduction not only reduces the computational complexity, but also effectively removes redundant information, improving the efficiency and robustness of recognition. The application of feature normalization ensures the equal contribution of different types of features in the fusion process, further optimizing the recognition effect. For archive digitization work, this embodiment can effectively process various types of archival text, including printed text from different periods, handwritten text of various styles, and even partially blurred or damaged characters. When processing low-quality images, such as old archives that are severely faded or documents with poor copy quality, this hybrid feature method shows robustness. For archives containing non-standard text elements such as special symbols, seals, mathematical formulas, etc., good recognition capabilities are also demonstrated. In large-scale archive digitization projects, the accuracy and efficiency of text recognition can be improved, the workload of manual proofreading can be reduced, and the consistency and high quality of the recognition results can be guaranteed.
[0117] In another embodiment of the present application, the improved MSER algorithm is applied specifically as follows: Gaussian blur (σ=0.8) is applied for preprocessing to reduce the impact of noise; MSER is run at different scales (original image, 1 / 2 scale, 1 / 4 scale); stability parameters are selected: the area change rate threshold Δ is initially set to 0.5, if the number of detected regions is <100, then Δ=Δ*0.9; if the number of detected regions is >1000, then Δ=Δ*1.1; hierarchical clustering is used to merge regions with overlap >0.7; and text lines are constructed based on the spatial relationship of the detected regions.
[0118] In another embodiment of the present application, the SVM classifier is optimized, specifically: using principal component analysis (PCA) to select the top K features with the greatest contribution (K is initially set to 20); comparing linear kernels, RBF kernels, and polynomial kernels, and selecting the kernel function with the best performance on the validation set; using grid search and 5-fold cross-validation to optimize the character image set C and γ parameters; using the optimized parameters and features to train the SVM model; evaluating the model on an independent test set, if the accuracy is <95%, increase the K value and repeat the above steps.
[0119] like Figure 5 As shown, according to one aspect of the present application, step S4 is further:
[0120] S41, based on the text content of the document, use the pre-trained word segmentation model to perform word segmentation processing to obtain a first word sequence; based on each word in the first word sequence, extract word features, including the word itself, part of speech, and context information, to form a feature vector; combine all feature vectors into a feature sequence, use the pre-trained conditional random field model to process the feature sequence, and obtain a corresponding label sequence; according to the label sequence, identify named entities in the text to form a named entity set;
[0121] S42, based on the text content of the document, perform word segmentation and part-of-speech tagging to obtain a second word sequence and a corresponding part-of-speech sequence; based on the second word sequence and the corresponding part-of-speech sequence, construct a word graph; iteratively calculate the text ranking score for each node in the word graph until convergence or reaching a preset number of iterations, to obtain a final text ranking score; based on the final text ranking score, select a predetermined number of keywords to form a keyword set;
[0122] S43, constructing a concept graph based on the named entity set and the keyword set; using the PageRank algorithm to calculate the importance scores of nodes in the concept graph; based on the importance scores, selecting a predetermined number of nodes to form a core concept set; based on the core concept set, searching for the most similar ontology concepts in a predefined ontology library; using the found ontology concepts as semantic tags to obtain a semantic tag set, which is the core semantic information of the document;
[0123] S44. Identify node pairs with labels in the semantic label set in the concept map; based on a preset path threshold, use a depth-first search algorithm to find all valid paths between each pair of node pairs; based on the characteristics of the nodes and edges on each valid path, use a predefined pattern or machine learning model to obtain the semantic relationship of the valid path; summarize the semantic relationships of all valid paths to form an entity relationship set.
[0124] In one embodiment of the present application, the recognition result string S is read. S is segmented using a pre-trained word segmentation model to obtain a word sequence W = {W1, W2, ..., Wk}. Features are extracted for each word Wi in the word sequence W, including the word itself, part of speech, context information, etc., to form a feature vector Xi. All feature vectors are combined into a sequence X = {X1, X2, ..., Xk}. The feature sequence X is processed using a pre-trained conditional random field (CRF) model to obtain a corresponding label sequence Y = {Y1, Y2, ..., Yk}. Named entities in the text are identified based on the label sequence Y to form a named entity set E = {E1, E2, ..., Ej}.
[0125] Read the recognition result string S. Perform word segmentation and part-of-speech tagging on the recognition result string S to obtain the word sequence W' and the corresponding part-of-speech sequence P. Construct a word graph G, where the nodes are candidate keywords in W' (usually nouns and adjectives), and the edge weights are based on the co-occurrence relationship of the words. Iterate and calculate the text ranking (TextRank) score for each node in the word graph G until convergence or the preset number of iterations is reached. According to the final TextRank score, select the words with the highest scores as keywords to form a keyword set K = {K1, K2, ..., Kl}.
[0126] Read the entity set E and the keyword set K. Use the elements in the entity set E and the keyword set K as nodes to construct the concept graph G'. Traverse the recognition result string S and count the co-occurrence of elements in the entity set E and the keyword set K. If two elements co-occur within a certain window size in the recognition result string S, add an edge between them in the concept graph G', and the weight of the edge is the co-occurrence frequency. Get the concept graph G' of the document.
[0127] Initialize the scores of all nodes in the concept graph G' to be equal. Iterate and update the score of each node. The new score is obtained by weighting the scores of other nodes pointing to the node according to the edge weight. Repeat this process until the score converges or the preset number of iterations is reached. Sort the nodes according to the final score, select the N nodes with the highest scores, and form the core concept set C core .
[0128] Load the predefined ontology library O. For the core concept set C core For each concept in , find the most similar ontology concept in the ontology library O. The similarity is calculated by string matching, synonym expansion, vector space model and other methods. The found ontology concept is assigned as a semantic label to the original concept. The semantic label set S is obtained tag .
[0129] Identify the set of semantic labels S in the concept graph G' tag Node pairs with labels in . For each pair of nodes, use the depth-first search algorithm to find all paths between them, and the path length does not exceed the preset threshold. For each path, based on the characteristics of the nodes and edges on the path, use a predefined pattern or machine learning model to determine whether it represents a certain semantic relationship. Summarize the relationships discovered by all valid paths to form an entity relationship set R = {R1, R2, ..., Rp}.
[0130] This embodiment achieves a deep understanding of the archive content and knowledge extraction through natural language processing and knowledge graph construction technology. Specifically, the named entity recognition technology based on conditional random field (CRF) can accurately identify key entities in the archive, such as names, place names, organization names, etc., which lays the foundation for the subsequent knowledge graph construction. The improved TextRank algorithm performs well in keyword extraction, especially for long documents and highly professional archive content, and can accurately capture the core theme of the document. The document representation method based on the graph is introduced to convert the document into a concept graph, which not only retains the relationship between entities, but also reflects the overall semantic structure of the document. The improved PageRank algorithm considers the entity type and relationship strength when calculating the importance of the node, making the extracted core concepts more accurate and representative. The application of semantic annotation technology maps the extracted concepts to the predefined ontology library, realizes the docking of archive content with the standardized knowledge system, and is conducive to the knowledge association and retrieval across archives. The application of the path mining algorithm can discover the implicit relationship between entities and enrich the content of the knowledge graph. The comprehensive application of this series of technologies makes the archive content no longer a simple text collection, but is transformed into a structured knowledge network. For the mining and utilization of a large number of historical archives, this embodiment can help researchers quickly discover the knowledge connections hidden in massive texts, thereby improving the efficiency and depth of archive utilization.
[0131] According to one aspect of the present application, step S42 is further:
[0132] S421, read the recognition result string S, and perform text preprocessing. Segment the text into sentences, remove punctuation marks and stop words, and convert all words to lowercase. Store the processed sentence list as S preprocessed ={s1,s2,...,sn}.
[0133] S422, read preprocessed sentence list S preprocessed , construct a word graph. Set the sliding window size w, and for each word in the sentence, establish edges connecting it with other words in the window. The weight of the edge is initialized to the number of co-occurrences. Store the constructed word graph as G = (V, E), where V is the vertex set (words) and E is the edge set.
[0134] S423, read the word graph G, and calculate the part-of-speech information of each word. Use the pre-trained part-of-speech tagging model to assign a part-of-speech tag to each word in the word graph G. Add the part-of-speech information to the node attributes of the graph G, and store the updated graph as G. pos .
[0135] S424, read the word graph G with part-of-speech information pos , calculate the new weight of the edge. Different weight coefficients are assigned to the edge according to the part-of-speech combination of the connecting words. For example, the weight of a noun-noun connection may be higher than that of an adjective-adverb connection. Store the graph with updated weights as G weighted .
[0136] S425, read weighted word graph G weighted , execute the improved PageRank algorithm. Initialize the rank value of each node to 1 / |V|, and then iteratively update the rank value of each node. The calculation formula is: PR(Vi)=(1-d)+d*Σ(w ji *PR(Vj) / Σw jk ), where d is the damping coefficient, w ji is the edge weight. Vi and Vj represent the nodes in the vertex set. Iterate until the rank value converges or the maximum number of iterations is reached. The final rank value is stored in the attribute of each node, and the updated graph is stored as G ranked .
[0137] S426, read graph G with rank value ranked , extract keywords. Sort all words in descending order of rank value, and select the first k words as candidate keywords. At the same time, consider the part of speech information and give priority to nouns, verbs and adjectives. Store the selected keyword list as K = {k1, k2, ..., k m}.
[0138] S427, read the keyword list K and the original text S, and perform keyword expansion. For each keyword, search for its co-occurring words in the original text. If the rank value of the co-occurring word is also high, add it to the keyword list. Store the expanded keyword list as K extended = {k1, k2, ..., kn}. The keyword list K extended Save to file or store in memory.
[0139] This embodiment achieves high-quality keyword extraction of archival text by improving the TextRank algorithm. Specifically, the sentence segmentation and stop word removal introduced in the text preprocessing step effectively improve the accuracy and efficiency of subsequent processing. The constructed word graph not only considers the co-occurrence relationship of words, but also introduces a sliding window mechanism, which can capture a wider range of semantic associations. The introduction of part-of-speech information increases the possibility of important words such as nouns and verbs being selected as keywords by assigning different weights to different parts of speech. The improvement of the edge weight calculation method considers the part-of-speech combination of words, so that semantically more related word pairs obtain higher connection strength. The improved PageRank algorithm introduces dynamic adjustment of the damping coefficient in the calculation of node importance, which improves the convergence speed and stability of the algorithm. In the process of keyword selection, the rank value and part-of-speech information of the word are combined to ensure that the extraction result not only reflects the core content of the document, but also has good grammatical characteristics. The introduction of the keyword expansion step effectively supplements important words that may be ignored by considering the rank value of the co-occurring words, and enhances the integrity and representativeness of the keyword set. The comprehensive application of this series of technologies makes the keyword extraction process have strong adaptability and high accuracy. For archive management and utilization, it can effectively process various types of archive texts, including administrative documents, technical reports, historical documents, etc., and even in the face of long documents with strong professionalism and complex writing, core keywords can be accurately extracted. This directly improves the retrieval efficiency and accuracy of archives, and users can quickly locate the required information through these high-quality keywords. When building an archive knowledge base and a subject classification system, these accurate keywords provide a reliable foundation for automatic classification and subject modeling. In cross-language archive management, these keywords can be used as an important basis for translation and cross-language retrieval. The present embodiment provides key technical support for the intelligent management and in-depth utilization of archives, improves the overall quality and value of archive digitization, and lays a solid foundation for efficient retrieval, knowledge discovery and decision support of archive information.
[0140] According to one aspect of the present application, the importance score of the node in the concept graph is calculated using the PageRank algorithm in step S43 as follows:
[0141] S431, read the concept graph G, and initialize the node importance scores. Assign an initial importance score r0(vi) = 1 / N to each node vi in the concept graph G, where N is the total number of nodes in the graph. Store the initialized graph as G init .
[0142] S432, read initialization graph G init, calculate the out-degree and in-degree of the node. Traverse each edge in the graph and count the number of outgoing edges (out-degree) and incoming edges (in-degree) of each node. Store the calculation results in the attributes of the node, and the updated graph is stored as G degree .
[0143] S433, read graph G degree , calculate the weight of the edge. For the edge eij connecting nodes vi and vj, calculate its weight w(eij)=(f(vi)+f(vj)) / (d(vi)*d(vj)), where f(v) is the frequency of node v and d(v) is the degree of node v. Store the calculated weight in the attributes of the edge, and the updated graph is stored as G weighted .
[0144] S434, read weighted graph G weighted , perform the improved PageRank algorithm iteration. For each node vi, update its importance score: r(vi) =(1-d) / N +d*Σ(w(eji) *r(vj)), where d is the damping factor, usually set to 0.85. Repeat this process until the score change of all nodes is less than the preset threshold ε or the maximum number of iterations is reached. The final importance score is stored in the node's attributes, and the updated graph is stored as G ranked .
[0145] S435, read the sorting graph G ranked , apply node importance correction. Assign different weight coefficients according to the type of node (such as entity, keyword), and multiply the importance score of the node by the corresponding weight coefficient. Update the corrected importance score to the node attribute to obtain graph G adjusted .
[0146] S436, read the adjusted graph G adjusted , perform node clustering. Use the spectral clustering algorithm to divide the nodes into several clusters based on the similarity of the connection strength and importance scores between nodes. Store the clustering results in the attributes of the nodes, and store the updated graph as G clustered .
[0147] S437, read cluster graph G clustered , select the core concepts. Select the N nodes with the highest importance scores in each cluster as the representatives of the cluster. Merge the representative nodes of all clusters to form the core concept set C core . The core concept set C core Save to file or store in memory.
[0148] This embodiment achieves accurate calculation of node importance in archive knowledge graph by improving PageRank algorithm, especially when dealing with complex archive semantic network. Specifically, the type and attribute of the node are considered in the process of node importance initialization, so that the algorithm can distinguish the importance of different types of nodes such as entities and keywords from the beginning. The introduction of node out-degree and in-degree information not only considers the number of connections of the node, but also reflects the influence and popularity of the node in the network. The improvement of edge weight calculation method more accurately reflects the semantic strength of the connection by considering the ratio of node frequency and degree. The dynamic adjustment mechanism of damping factor is introduced in the improved iterative formula to improve the convergence speed and stability of the algorithm in complex networks. The introduction of node importance correction step makes the calculation result more consistent with the professional knowledge in the archive field by giving different weight coefficients to different types of nodes. The application of spectral clustering algorithm for node clustering not only considers the direct connection between nodes, but also captures more complex structural relationships. In the process of core concept selection, by selecting the most important node in each cluster, it is ensured that the selection result reflects the overall importance and maintains diversity. The comprehensive application of this series of technologies makes the node importance calculation process have strong adaptability and high accuracy. This embodiment can effectively process various complex archival knowledge networks, including large-scale semantic networks across fields, multi-level concept systems, etc. It directly improves the quality and practicality of the knowledge graph, so that the system can accurately identify the core concepts, key entities and important relationships in the archives. In the intelligent retrieval system, these importance scores can be used to optimize the sorting of search results to ensure that the most relevant and important information is presented to users first. In archival knowledge discovery and association analysis, important nodes can be used as entry points to help researchers quickly grasp the core content and key clues of documents. For the knowledge extraction and concentration of large-scale archives, this embodiment can automatically generate high-quality summaries and knowledge graphs, improving the readability and utilization efficiency of information. In cross-archival and cross-domain knowledge associations, these important nodes often become bridges between different knowledge domains, promoting the integration and innovation of knowledge.
[0149] like Figure 6 As shown, according to one aspect of the present application, step S5 is further:
[0150] S51. Based on the logical structure of the document, extract structural features, including title levels and the number of paragraphs, to form a structural feature vector; convert the semantic tags in the core semantic information into one-hot encoding to form a semantic feature vector; represent the entity relationships in the entity relationship set as numerical pairs of relationship type and relationship strength to form a relationship feature vector; concatenate the structural feature vector, the semantic feature vector and the relationship feature vector to obtain a document feature vector;
[0151] S52, performing an inner product between the document feature vector and a pre-generated random projection vector to obtain an inner product result; forming a binary code based on the positive or negative value of the inner product result; concatenating all binary codes to obtain a local sensitive hash code of the document feature vector; repeating until a predetermined number of iterations is reached to obtain a predetermined number of local sensitive hash codes to form a final low-dimensional representation of the document feature vector;
[0152] S53. Based on the final low-dimensional representation, an improved density clustering algorithm is applied to perform dynamic clustering to obtain a dynamic category set; based on the dynamic category set, the center point of each category is calculated; a non-weighted group average algorithm is used to hierarchically cluster the center points to obtain a tree structure; a semantic label is assigned to each node in the tree structure to obtain a multi-level classification tree;
[0153] S54. Select a predetermined dimension as the index key in the final low-dimensional representation; construct a root node of a B+ tree, insert the index key into the root node of the B+ tree, and obtain a multidimensional index structure; combine the multi-level classification tree and the multidimensional index structure to obtain a classified index structure of the archive.
[0154] In one embodiment of the present application, the document logical structure tree L and the semantic tag set S are read. tag And entity relationship set R. Extract structural features from the document logical structure tree L, such as title level, number of paragraphs, etc., to form a structural feature vector V struct . The semantic label set S tag The semantic labels in are converted into one-hot encoding to form a semantic feature vector V semantic The entity relationships in the entity relationship set R are represented as numerical pairs of relationship type and relationship strength to form a relationship feature vector V relation . The feature vector V struct 、V semantic and V relation Concatenate and get the document feature vector V.
[0155] Generate a set of random projection vectors {P1, P2, ..., Pk}. Take the inner product of the document feature vector V and each random projection vector Pi, and get a binary code bi according to the sign of the inner product result. Concatenate all binary codes bi to get the local sensitive hash code H of the document feature vector V. Repeat this process multiple times to get multiple hash codes, which constitute the final low-dimensional representation V of V. low .
[0156] Calculate the final low-dimensional representation V lowThe distances from all other vectors are obtained to obtain the distance matrix D. For each vector, count the number of points in its ε neighborhood. If the number exceeds the preset threshold, the vector is marked as a core point. Starting from any unvisited core point, recursively add all points in its ε neighborhood to the current cluster. Repeat this process until all points are visited. Assign noise points to the nearest cluster or form a new cluster. Get the dynamic category set D = {D1, D2, ..., Dq}.
[0157] Calculate the center point Ci of each category Di in the dynamic category set D. Use the Unweighted PairGroup Method with Arithmetic Mean (UPGMA) algorithm to hierarchically cluster these center points. Get a tree structure T. Take the original category as the leaf node of the tree structure T, and the clustering result as the non-leaf node. Assign a semantic label to each node in the tree structure T. The label can use the common features of the documents contained in the node or the most frequent semantic label. Get a multi-level classification tree T class .
[0158] Select a low-dimensional feature vector V low Several dimensions in the tree are used as index keys K = {K1, K2, ..., Km}. Create the root node Root of the B+ tree. Insert the index key and document ID of each document into the B+ tree in turn. If the node is full, split it and choose the split point that can maximize the query efficiency. Adjust the structure of the tree to keep it balanced. Get the multidimensional index structure I multi .
[0159] The multidimensional index I multi Each leaf node in the classification tree T class In the classification tree T class Add a pointer to the multidimensional index I on each non-leaf node multi Pointers to the relevant parts. Create metadata structure M to store the classification tree T class and multidimensional index I multi The root node information and other necessary global parameters of the classification tree T class , Multidimensional Index I multi and metadata structure M into a unified data structure. The final archive classification index structure I is obtained final .
[0160] The present embodiment realizes efficient organization and rapid retrieval of massive archives through multidimensional indexing and dynamic classification technology. Specifically, the application of local sensitive hashing (LSH) algorithm realizes efficient dimensionality reduction while maintaining data similarity, which lays the foundation for subsequent clustering and indexing. The improved DBSCAN dynamic clustering algorithm can not only process clusters of irregular shapes, but also adjust clustering parameters adaptively, which is particularly effective for processing diversified archive content. The introduction of fast neighbor search technology based on kd tree improves the efficiency of clustering process. The construction of multi-level classification tree combines bottom-up and top-down classification strategies, which not only ensures the accuracy of classification, but also provides flexible classification levels to adapt to archive management needs of different granularities. The application of improved B+ tree multidimensional index structure, combined with Z-order curve mapping technology, realizes efficient multidimensional data indexing and improves the response speed of complex queries. The introduction of auxiliary index structure further optimizes the performance of range query and multi-condition combination query. The comprehensive application of this series of technologies makes the organization and retrieval of massive archives efficient and flexible. For archive management systems, this means being able to complete complex multi-dimensional queries within milliseconds, supporting flexible dynamic adjustment of classification systems, and realizing intelligent association and recommendation of archive content. Especially for large archives or cross-institutional archive management systems, it can effectively process PB-level archive data, support high-concurrency user access, and dynamically optimize index structures according to archive usage patterns, improving the efficiency of archive utilization and user experience.
[0161] In another embodiment of the present application, in step S53, an improved density-based clustering (DBSCAN) algorithm is applied to perform dynamic clustering, further comprising:
[0162] S531, read low-dimensional feature vector V low , construct a kd tree index structure. Use the low-dimensional feature vector V low Each vector in is regarded as a k-dimensional point, and the space is recursively divided to construct a balanced kd tree. The constructed kd tree is stored as T kd .
[0163] S532, read kd tree T kd , calculate the adaptive neighborhood radius. Randomly select m sample points and use the kd tree T kd Quickly find the k nearest neighbors of each sample point and calculate the average of these distances as the initial neighborhood radius ε. Store the calculated ε in memory.
[0164] S533, read the neighborhood radius ε and kd tree T kd , perform improved DBSCAN clustering. Traverse the low-dimensional feature vector V low For each point p in, use the kd tree T kdFind the points in the ε-neighborhood of p. If the number of points in the neighborhood is greater than the minimum number of points (MinPts), mark p as a core point and add the points in its neighborhood to the same cluster. Repeat this process until all points have been visited. Store the clustering results as a list of clusters C = {C1, C2, ..., Ck}.
[0165] S534, read the cluster list C, and calculate the features of each cluster. For each cluster Ci, calculate its center point, radius, density and other features. Store the calculated cluster features as F = {F1, F2, ..., Fk}, where Fi corresponds to the features of cluster Ci.
[0166] S535, read cluster feature F, merge and split clusters. Calculate the distance matrix D between clusters. If the distance between two clusters is less than the threshold T merge , then merge the two clusters. If the radius of a cluster is greater than the threshold T split , then use the k-means algorithm to split the cluster into two sub-clusters. Update the processed cluster list to C' = {C'1, C'2, ..., C'n}.
[0167] S536, read the updated cluster list C', and redistribute the points in each cluster. Calculate the distance from each point to the center of all clusters, and assign the point to the cluster with the closest distance. If the distance between a point and the nearest cluster is greater than 2ε, mark it as a noise point. Update the redistributed cluster list to C'' = {C''1, C''2, ..., C''m}.
[0168] S537, read the final cluster list C'', and construct a dynamic category set. Assign a unique category identifier to each cluster C''i, and map the points in the cluster to the category. Obtain a dynamic category set D = {D1, D2, ..., Dm}, where Di corresponds to the category of cluster C''i. Save D as a file or store it in memory for use in subsequent steps.
[0169] This embodiment realizes efficient and adaptive clustering of archive feature vectors by improving the DBSCAN dynamic clustering algorithm, especially when processing large-scale, high-dimensional, irregular-shaped clusters of archive data. Specifically, the construction of the kd tree index structure improves the efficiency of neighbor search, which is crucial for processing massive archive data. The introduction of the adaptive neighborhood radius calculation method enables the algorithm to dynamically adjust parameters according to the local density characteristics of the data, thereby improving the adaptability and accuracy of clustering. In the improved DBSCAN core process, the kd tree is used for efficient neighborhood query, which reduces the computational complexity and enables the algorithm to process larger data sets. The introduction of the cluster feature calculation step not only provides a basis for subsequent cluster merging and splitting, but also lays a foundation for the semantic understanding of the cluster. The constructed cluster merging and splitting mechanism effectively solves the problem of the sensitivity of the traditional DBSCAN algorithm to parameters by dynamically adjusting the structure of the cluster. The application of the redistribution step further optimizes the clustering results and improves the compactness and separability of the cluster. The comprehensive application of this series of technologies makes the archive clustering process have strong adaptability, efficiency and accuracy. This embodiment can effectively process various types of archival data, including text, images, multimedia, etc., and can obtain high-quality clustering results even in the face of highly heterogeneous and unevenly distributed data sets. This directly improves the organizational efficiency and retrieval accuracy of archives. The system can automatically discover the inherent structure and relationship in the archives to form meaningful categories and themes. In the construction and maintenance of the archive classification system, this dynamic clustering technology can automatically identify new themes and categories, and help administrators update the classification system in a timely manner. For the arrangement and research of massive historical archives, potential themes and trends can be automatically discovered, providing new perspectives and clues for historical research. In cross-domain and cross-language archive management, it can help discover potential connections between archives in different fields and languages. For personalized archive recommendation systems, clustering results can be used to build user interest models and provide more accurate content recommendations.
[0170] According to one aspect of the present application, in step S54, a predetermined dimension is selected as an index key in the final low-dimensional representation; a root node of a B+ tree is constructed, and the index key is inserted into the root node of the B+ tree, so that the multidimensional index structure is obtained as follows:
[0171] S541. Calculate the variance of each dimension in the final low-dimensional representation, and select k dimensions with the largest variance as index keys; where k is a natural number greater than 1;
[0172] S542, based on the index key, perform dimensionality reduction processing on the final low-dimensional representation to obtain a vector set after dimensionality reduction; based on the vector set after dimensionality reduction, construct a Z-order curve mapping to obtain a Z value list; based on the Z value list, use a quick sort algorithm to perform ascending sorting to obtain a sorted Z value list;
[0173] S543, based on the sorted Z value list, construct leaf nodes of the B+ tree to obtain a leaf node list; based on the leaf node list, construct the root node of the B+ tree from bottom to top; based on the leaf nodes and the root node, form a B+ tree;
[0174] S544, optimizing the B+ tree to obtain an optimized B+ tree; constructing an auxiliary index structure based on the optimized B+ tree; combining the auxiliary index structure and the optimized B+ tree to form a multidimensional index structure.
[0175] In one embodiment of the present application, the low-dimensional feature vector V is read low , determine the index key dimension. Calculate the low-dimensional feature vector V low The variance of each dimension in , select the k dimensions with the largest variance as index keys. Store the selected index key dimensions as K = {k1, k2, ..., kk}. For the low-dimensional feature vector V low Perform dimensionality reduction processing, retaining only the dimensions specified in the index key dimension K, and obtain the reduced dimensionality vector set V index . Construct Z-order curve mapping: Map each k-dimensional vector to a one-dimensional space and generate Z values using the bit-interleaving method. Store the mapping result as a key-value pair list Z = {(z1, id1), (z2,id2), ..., (zn, idn)}, where zi is the Z value and idi is the corresponding document ID.
[0176] Read the Z value list and sort it. Use the quick sort algorithm to sort the list in ascending order according to the Z value. Store the sorted list as Z sorted . Construct the leaf nodes of the B+ tree: List Z sorted Divide into blocks of size B, each block forms a leaf node. The leaf node stores a (z, id) pair and a pointer to the next leaf node. Store the constructed leaf node list as L = {L1, L2, ..., Lm}.
[0177] Read the leaf node list L and construct the internal nodes of the B+ tree. Build from the bottom up, build a parent node for every B child nodes, and the parent node stores the minimum Z value of the child node as the separation key. Repeat this process until the root node is generated. Store the constructed B+ tree structure as T bplus .
[0178] For B+ tree T bplus Optimize: Analyze the balance and node utilization of the tree. If the utilization of a node is lower than the threshold, merge or reallocate entries with adjacent nodes. Repeat this process until the tree structure is stable. Update the optimized B+ tree structure and store it as T bplusoptFor each dimension ki, create a sorted list Si, where the list elements are (v, ptr) pairs, where v is the value of the dimension and ptr is a pointer to the child node of the B+ tree. Store these k sorted lists as S = {S1, S2,..., Sk}.
[0179] The B+ tree T bplusopt Combined with the auxiliary index structure S, the final multidimensional index structure I is formed multi . Multidimensional index structure I multi Save as a file or store in a database for subsequent efficient multi-dimensional query.
[0180] This embodiment achieves efficient storage and fast retrieval of archive feature vectors by improving the B+ tree multidimensional index construction technology, especially when processing high-dimensional and large-scale archive data. Specifically, the intelligent selection method of the index key dimension ensures that the selected dimension has the greatest discrimination by analyzing the variance of each dimension, thereby improving the efficiency and accuracy of the index. The Z-order curve mapping technology is introduced to map the high-dimensional vector to a one-dimensional space, which effectively solves the difficulties of the traditional B+ tree in processing multidimensional data. The calculation of the Z value adopts the bit crossover method, which not only maintains the locality of the data, but also improves the calculation efficiency. In the construction process of the B+ tree, a node splitting strategy is constructed, and the query efficiency is maximized by selecting the optimal splitting point. The introduction of the tree structure optimization step improves the space efficiency and query performance of the index by dynamically adjusting the node utilization. The creation of the auxiliary index structure constructs an independent sorting list for each dimension, which improves the efficiency of range query and multidimensional query. The comprehensive application of this series of technologies makes the archive index construction and query process efficient, scalable and flexible. This embodiment can effectively process various types of archival data, including text feature vectors, image features, multimedia features, etc., and can guarantee millisecond query response time even in the face of TB-level archives. This directly improves the efficiency and user experience of archive retrieval. The system can quickly locate similar archives, perform content-based retrieval, and support complex multi-condition combination queries. In archive version management and change tracking, this efficient index structure can quickly locate and compare archives of different versions. For the management of spatiotemporal archival data, time series analysis and spatial location query can be efficiently supported. In archive association analysis and knowledge discovery, efficient multidimensional indexing enables the system to quickly identify potential associations and patterns. For application scenarios with high real-time requirements, such as emergency management archive systems, fast decision support can be supported. In large-scale distributed archive systems, the improved B+ tree structure has good distributed characteristics and supports efficient distributed queries and load balancing. This embodiment provides key technical support for efficient management and rapid retrieval of archives, improves the efficiency and flexibility of archive utilization, and lays the foundation for deep mining, knowledge association and intelligent application of archive information.
[0181] In another embodiment of the present application, the parameters of the DBSCAN algorithm are optimized, specifically: using the K-distance graph method, K is initially set to 0.2% of the number of samples; setting MinPts to log (number of samples), rounded up; performing density adaptation: calculating the global average density ρ avg ; For density ρ<ρ avg / 2, ε=ε*1.2; for density ρ>ρ avg 2, ε=ε*0.8; update ε and MinPts every 1000 samples; mark the points that are >3std (intra-cluster distance) from the nearest cluster center as noise.
[0182] In another embodiment of the present application, the B+ tree index is optimized: the initial order m = ⌈sqrt(N)⌉, where N is the total number of documents; adaptive adjustment is performed: if the average query time > 1ms, increase m: m = m*1.2; if the average insertion time > 5ms, reduce m: m = m*0.8; the splitting strategy uses a splitting point selection algorithm that minimizes variance; implements LRU cache, and the cache size is twice the tree height; implements a sorted batch insertion algorithm, processing sqrt(N) documents each time.
[0183] According to one aspect of the present application, an archive digitization system based on intelligent image enhancement and automatic classification includes:
[0184] at least one processor; and,
[0185] a memory communicatively connected to at least one of the processors; wherein,
[0186] The memory stores instructions that can be executed by the processor, and the instructions are used to be executed by the processor to implement the archive digitization method based on intelligent image enhancement and automatic classification described in any of the above embodiments.
[0187] In another embodiment of the present application, step S0 is further included, namely, identifying the file type, which is specifically:
[0188] S0a, color analysis: calculate the color histogram of the image to determine whether it is a color image.
[0189] S0b. Text density analysis: Calculate the text area ratio of the image to determine whether it is a text-intensive document.
[0190] S0c. Handwriting analysis: Use Hough transform to detect handwriting features and determine whether it is a handwritten document.
[0191] S0d. Layout complexity analysis: calculate the entropy value of the image and determine the complexity of the layout.
[0192] S0e. Type classification: Based on the above features, the files are classified into: ordinary text, color images, handwritten documents, and complex layout documents.
[0193] In another embodiment of the present application, step S1 also includes special processing of the color image: converting the RGB image into the LAB color space; applying the processing flow of step S1 to the L channel, and processing the A and B channels to maintain color information; using an adaptive color enhancement algorithm to improve image saturation and contrast; applying a white balance algorithm to correct the overall hue; and recombining the processed LAB channels and converting them back to the RGB space.
[0194] In another embodiment of the present application, step S1 also includes special processing of handwritten documents: using morphological operations to enhance handwriting contours; applying adaptive thresholding methods to suppress uneven backgrounds; using Radon transforms to detect and correct overall tilt angles; using projection contour methods to extract text lines; and applying thinning algorithms and stroke connection algorithms to improve handwriting continuity.
[0195] In another embodiment of the present application, step S3 also includes special processing for complex layout documents: applying the processing flow of step S3 at multiple scales and merging the results; using the XY cutting algorithm to perform preliminary layout segmentation; classifying the content type of each segmented area (such as text, table, picture, etc.); constructing a logical structure tree of the document based on the segmentation and classification results; identifying and processing cross-references and page jumps within the document.
[0196] In another embodiment of the present application, a method for digitizing archives based on intelligent image enhancement and automatic classification includes the following steps:
[0197] S1. High-fidelity archival image acquisition and preprocessing.
[0198] S11. Use high-resolution scanner to obtain original archive images I raw , the resolution shall not be less than 600dpi.
[0199] S12, applying adaptive histogram equalization (AHE) to process the original image I raw , and obtain the contrast-enhanced image I contrast .
[0200] S13, use Wiener filter to image I contrast Perform denoising to obtain the denoised image I denoised .
[0201] S14, detecting denoised image I based on Hough transform denoisedThe tilt angle θ in the image is used to perform image rotation correction through bilinear interpolation to obtain the corrected image I corrected .
[0202] S15, using edge detection and contour analysis techniques from the corrected image I corrected Extract the document boundary and crop to get the final preprocessed image I preprocessed .
[0203] S2. Adaptive multi-scale image enhancement.
[0204] S21, final preprocessed image I preprocessed Applying wavelet transform, we obtain the multi-scale decomposition coefficients {LL, LH, HL, HH}.
[0205] S22. Use local statistical features (mean μ and variance σ²) to adaptively adjust each subband coefficient to obtain the enhancement coefficient {LL', LH', HL', HH'}. Enhancement function: C'=f(C,μ,σ²)=C*(1+ k*log(1+σ²)), where C is the original coefficient and k is an adjustable parameter.
[0206] S23. Apply the improved nonlinear sharpening operator to the high frequency subbands LH', HL', HH' to enhance edge details. Sharpening function: S(x) = x+α*sign(x)*(1 -exp(-β|x|)), where α and β are adjustable parameters.
[0207] S24, reconstruct the image using inverse wavelet transform to obtain the enhanced image I enhanced .
[0208] S25, applying local contrast limited adaptive histogram equalization (CLAHE) to the enhanced image I enhanced , and get the final enhanced image I final .
[0209] S3. Intelligent text recognition and layout analysis.
[0210] S31, using the improved MSER algorithm from the final enhanced image I final Extract the text region set T={T1,T2,...,Tn}.
[0211] S32. Apply adaptive binarization processing to each text region Ti to obtain a binarized text image set B={B1, B2, ..., Bn}.
[0212] S33. Use the improved connected component analysis algorithm to perform character segmentation on each Bi to obtain a character image set C={C1,C2,...,Cm}.
[0213] S34. Apply a hybrid feature extraction method based on Zernike moments and LBP to process each character image Ci to obtain a feature vector set F={F1, F2, ..., Fm}.
[0214] S35. Use a support vector machine (SVM) classifier to classify each feature vector in the feature vector set F to obtain a recognition result character string S.
[0215] S36. Based on the spatial relationship of the text region T and in combination with the recognition result S, a logical structure tree L of the document is constructed.
[0216] S4. Semantic understanding and knowledge extraction.
[0217] S41. Use a sequence labeling model based on conditional random fields (CRF) to perform named entity recognition on the string S and obtain an entity set E={E1,E2,...,Ek}.
[0218] S42. Apply the improved TextRank algorithm to extract the document keyword set K={K1,K2,...,Kj}.
[0219] S43. Construct a concept graph G of the document based on the entity set E and the document keyword set K, where the nodes are entities and keywords and the edges represent co-occurrence relationships.
[0220] S44. Use the PageRank algorithm to calculate the importance scores of the nodes in the concept graph G, and select the top-N nodes as the core concept set C. core .
[0221] S45. Use the predefined ontology library O to set the core concept C core Mapped to the corresponding ontology concepts to form a semantic label set S tag .
[0222] S46, based on the concept graph G and the semantic label set S tag , use the path mining algorithm to extract the relationship R={R1,R2,...,Rp} between entities.
[0223] S5. Dynamic multi-dimensional archive classification and index construction.
[0224] S51, document logical structure L, semantic label S tag and entity relationship R into document feature vector V.
[0225] S52, using the local sensitive hashing (LSH) algorithm to reduce the dimension of the document feature vector V to obtain a low-dimensional feature vector V low .
[0226] S53, apply dynamic clustering algorithm (such as improved DBSCAN) to the low-dimensional feature vector V low Clustering is performed to obtain the dynamic category set D={D1,D2,...,Dq}.
[0227] S54. Construct a multi-level classification tree structure T based on the dynamic category set D class .
[0228] S55, using the improved B+ tree algorithm, with low-dimensional feature vector V low As the key, the document ID is the value, and a multidimensional index I is constructed. multi .
[0229] S56, the classification tree structure T class and multidimensional index I multi Combined to form the final archive classification index structure I final .
[0230] The present invention realizes the optimization of the whole process from archival image acquisition to knowledge extraction and intelligent retrieval through the organic combination of a series of technologies, thereby improving the efficiency, quality and availability of archive digitization. Specifically, high-fidelity image acquisition and preprocessing technology provide high-quality input for subsequent processing, effectively solving the problems of fading and defacement of historical archives. Adaptive multi-scale image enhancement technology further improves the image quality, especially the enhancement of text areas, laying the foundation for text recognition. Intelligent text recognition and layout analysis technology not only realizes high-precision text extraction, but also retains the structural information of the document, which is crucial for subsequent semantic understanding. Semantic understanding and knowledge extraction technology converts unstructured text into structured knowledge graphs, improving the value and availability of archives. Dynamic multi-dimensional archive classification and indexing technology provides efficient organization and retrieval solutions for massive archives. The present invention realizes all-round archive processing from pixel level to semantic level, making historical archives easy to understand and use; the seamless connection between each step and the optimization of data flow improve the efficiency of the entire processing flow, making large-scale archive digitization possible; the application of dynamic classification and multidimensional indexing technology enables the archive management system to adapt to the ever-changing user needs and archive structure; the construction of knowledge graphs opens up new possibilities for deep mining and cross-domain applications of archives. For archive management agencies, it can not only accelerate the process of archive digitization, but also provide smarter and more accurate archive services. The present invention not only solves many technical challenges currently faced by archive digitization, but also opens up new paths for intelligent management and deep utilization of archives.
[0231] The preferred embodiments of the present invention are described in detail above; however, the present invention is not limited to the specific details in the above embodiments. Within the technical concept of the present invention, various equivalent transformations can be made to the technical solutions of the present invention, and these equivalent transformations all belong to the protection scope of the present invention.
Claims
1. The archive digitization method based on intelligent image enhancement and automatic classification is characterized by: The steps include: S1. Obtaining the original archive image and performing preprocessing to obtain a high-quality image after preprocessing; wherein the preprocessing includes contrast enhancement, noise reduction, tilt correction and boundary clipping; S2, based on the preprocessed high-quality image, performing adaptive multi-scale image enhancement processing to obtain an enhanced image with improved quality; The adaptive multi-scale image enhancement processing includes wavelet transform, coefficient adjustment, nonlinear sharpening and local contrast enhancement; S3. Based on the enhanced image, intelligent text recognition and layout analysis are performed to obtain the text content and logical structure of the document; wherein the intelligent text recognition and layout analysis include text area extraction, adaptive binarization, character segmentation, feature extraction, character recognition and document structure analysis; S4. Based on the text content and logical structure of the document, semantic understanding and knowledge extraction are performed to obtain the core semantic information of the document; semantic understanding and knowledge extraction include named entity recognition, keyword extraction, concept map construction, importance calculation, semantic annotation and relationship extraction; S5. Based on the logical structure and core semantic information of the document, dynamic multi-dimensional archive classification and index construction are performed to obtain the classified index structure of the archive; The dynamic multi-dimensional archive classification and index construction includes feature vector construction, dimensionality reduction processing, dynamic clustering, classification tree construction and multi-dimensional index construction; Step S4 is further as follows: S41, based on the text content of the document, using a pre-trained word segmentation model to perform word segmentation processing to obtain a first word sequence; based on each word in the first word sequence, extract word features, including the word itself, part of speech, and context information, to form a feature vector; All feature vectors are combined into a feature sequence, and the feature sequence is processed using a pre-trained conditional random field model to obtain a corresponding label sequence; based on the label sequence, the named entities in the text are identified to form a named entity set; S42, based on the text content of the document, perform word segmentation and part-of-speech tagging to obtain a second word sequence and a corresponding part-of-speech sequence; based on the second word sequence and the corresponding part-of-speech sequence, construct a word graph; iteratively calculate the text ranking score for each node in the word graph until convergence or reaching a preset number of iterations, to obtain a final text ranking score; Based on the final text ranking score, a predetermined number of keywords are selected to form a keyword set; S43, constructing a concept graph based on the named entity set and the keyword set; using the PageRank algorithm to calculate the importance scores of nodes in the concept graph; based on the importance scores, selecting a predetermined number of nodes to form a core concept set; based on the core concept set, searching for the most similar ontology concepts in a predefined ontology library; using the found ontology concepts as semantic tags to obtain a semantic tag set, which is the core semantic information of the document; S44, identifying node pairs with labels in the semantic label set in the concept graph; based on a preset path threshold, using a depth-first search algorithm to find all valid paths between each pair of nodes; based on the characteristics of nodes and edges on each valid path, using a predefined pattern or machine learning model, obtaining the semantic relationship of the valid path; summarizing the semantic relationships of all valid paths to form an entity relationship set; Step S5 is further as follows: S51. Based on the logical structure of the document, extract structural features, including title levels and the number of paragraphs, to form a structural feature vector; convert the semantic tags in the core semantic information into one-hot encoding to form a semantic feature vector; represent the entity relationships in the entity relationship set as numerical pairs of relationship type and relationship strength to form a relationship feature vector; concatenate the structural feature vector, the semantic feature vector and the relationship feature vector to obtain a document feature vector; S52, performing an inner product between the document feature vector and a pre-generated random projection vector to obtain an inner product result; forming a binary code based on the positive or negative value of the inner product result; concatenating all binary codes to obtain a local sensitive hash code of the document feature vector; repeating until a predetermined number of iterations is reached to obtain a predetermined number of local sensitive hash codes to form a final low-dimensional representation of the document feature vector; S53. Based on the final low-dimensional representation, an improved density clustering algorithm is applied to perform dynamic clustering to obtain a dynamic category set; based on the dynamic category set, the center point of each category is calculated; a non-weighted group average algorithm is used to hierarchically cluster the center points to obtain a tree structure; a semantic label is assigned to each node in the tree structure to obtain a multi-level classification tree; S54, selecting a predetermined dimension in the final low-dimensional representation as an index key; constructing a root node of a B+ tree, inserting the index key into the root node of the B+ tree, and obtaining a multidimensional index structure; The multi-level classification tree and multi-dimensional index structure are combined to obtain the classification index structure of the archives.
2. The method for digitalizing archives based on intelligent image enhancement and automatic classification according to claim 1, characterized in that: Step S1 is further as follows: S11. Scan the original file using a high-resolution scanner to obtain an original file image with a resolution not less than a preset threshold; S12, dividing the original archive image into a predetermined number of small areas of equal size; calculating a histogram of pixel grayscale values for each small area; Based on the histogram of each small area, a histogram equalization algorithm is used to recombine all the small areas to obtain contrast-enhanced image data; S13, calculating the local mean and variance of the contrast-enhanced image data as local statistical characteristics; and constructing a Wiener filter based on the local statistical characteristics; The contrast-enhanced image data is processed using a Wiener filter to obtain the denoised image data; S14, based on the image data after noise reduction, using Hough transform algorithm to detect straight line features in the image; Calculate the angle between the straight line feature and the horizontal line to obtain the tilt angle of the image; based on the tilt angle, use the bilinear interpolation algorithm to perform rotation correction to obtain the corrected image data; S15. Based on the corrected image data, use the Canny edge detection algorithm to extract the image edge; based on the image edge, perform contour analysis to identify the boundary information; based on the boundary information, crop the corrected image data to obtain the final pre-processed high-quality image.
3. The method for digitalizing archives based on intelligent image enhancement and automatic classification according to claim 2, characterized in that: Step S2 is further as follows: S21, performing a one-dimensional wavelet transform on each row of the preprocessed high-quality image to obtain a row transform result; performing a one-dimensional wavelet transform on each column of the row transform result to obtain a two-dimensional wavelet transform result; decomposing the two-dimensional wavelet transform result into a predetermined number of sub-bands, including low-frequency approximate coefficients, horizontal high-frequency coefficients, vertical high-frequency coefficients, and diagonal high-frequency coefficients, to form a wavelet coefficient set; S22, based on the wavelet coefficient set, calculating the local statistical characteristics of each sub-band, including the local mean and local variance of each sub-band; based on the local statistical characteristics, adjusting the sub-band coefficients in the wavelet coefficient set to obtain an adjusted coefficient set; S23, based on the adjusted coefficient set, extracting high-frequency subband coefficients and low-frequency subband coefficients; applying an improved nonlinear sharpening operator to the high-frequency subband coefficients to obtain sharpened high-frequency subband coefficients; combining the sharpened high-frequency subband coefficients with the extracted low-frequency subband coefficients to form a new coefficient set; S24, performing a one-dimensional inverse wavelet transform on each column in the new coefficient set to obtain a column inverse transform result; performing a one-dimensional inverse wavelet transform on each row of the column inverse transform result to obtain a final reconstructed image; S25, dividing the final reconstructed image into a predetermined number of small blocks of equal size; performing histogram equalization on each small block to obtain processed small blocks; and using a bilinear interpolation method to recombine the processed small blocks to obtain an enhanced image with improved quality.
4. The method for digitalizing archives based on intelligent image enhancement and automatic classification according to claim 3 is characterized in that: Step S3 is further as follows: S31, traversing all pixels of the enhanced image, analyzing grayscale changes, and identifying stable extreme value regions; calculating the features of each extreme value region, including area, perimeter, and aspect ratio; based on the features of the extreme value region and a preset threshold, screening the extreme value region to obtain a potential text region set; S32, based on the text region set, calculating the grayscale histogram of each text region; Apply the Otsu method to the grayscale histogram to calculate the optimal binarization threshold; Use the optimal binarization threshold to perform binarization on each text area to obtain a set of binarized text images; S33, based on the binary text image set, scanning each pixel in the binary text image, identifying a connected region composed of adjacent black pixels; calculating features of each connected region, including area, aspect ratio and density; Based on the features of the connected regions and the preset morphological rules, the connected regions are screened to obtain screened connected regions; based on the screened connected regions, a character image set is formed; S34, based on the character image set, calculating the Zernike moment of each character image to obtain a geometric feature vector; based on the geometric feature vector, using a local binary pattern algorithm to obtain a texture feature vector; concatenating the geometric feature vector and the texture feature vector to form a mixed feature vector; combining all the mixed feature vectors to obtain a feature vector set; S35, using a pre-trained support vector machine classifier to classify each feature vector in the feature vector set to obtain a recognition result string; based on the recognition result string, forming the text content of the document; S36. Analyze the spatial relationship between the text regions in the text region set and identify the structural elements; map the text content of the document to the structural elements and construct a tree-like data structure, which is the logical structure of the document.
5. The method for digitalizing archives based on intelligent image enhancement and automatic classification according to claim 4 is characterized in that: In step S54, a predetermined dimension is selected as an index key in the final low-dimensional representation; a root node of a B+ tree is constructed, and the index key is inserted into the root node of the B+ tree, so that the multidimensional index structure is obtained as follows: S541. Calculate the variance of each dimension in the final low-dimensional representation, and select k dimensions with the largest variance as index keys; where k is a natural number greater than 1; S542, based on the index key, perform dimensionality reduction processing on the final low-dimensional representation to obtain a vector set after dimensionality reduction; based on the vector set after dimensionality reduction, construct a Z-order curve mapping to obtain a Z value list; Based on the Z value list, use the quick sort algorithm to sort in ascending order to obtain the sorted Z value list; S543, based on the sorted Z value list, construct leaf nodes of the B+ tree to obtain a leaf node list; based on the leaf node list, construct the root node of the B+ tree from bottom to top; based on the leaf nodes and the root node, form a B+ tree; S544, optimizing the B+ tree to obtain an optimized B+ tree; Build auxiliary index structure based on optimized B+ tree; The auxiliary index structure and the optimized B+ tree are combined to form a multidimensional index structure.
6. The method for digitalizing archives based on intelligent image enhancement and automatic classification according to claim 4, characterized in that: Step S12 is further as follows: S121, using a sliding window method to divide the original archive image into a predetermined number of overlapping local areas; calculating a histogram of each local area; calculating a cumulative distribution function of each histogram; performing contrast limitation processing on each cumulative distribution function to obtain a processed cumulative distribution function; S122, remapping based on the processed cumulative distribution function to obtain a mapping list; Based on the mapping list, the gray value of each pixel in the original archive image is remapped to obtain a new image matrix; S123, performing global contrast stretching on the new image matrix to obtain contrast-enhanced image data.
7. The method for digitalizing archives based on intelligent image enhancement and automatic classification according to claim 4 is characterized in that: Step S34 is further as follows: S341, based on the character image set, normalizing the size of each character image to obtain a normalized character image set; S342, based on the normalized character image set, calculating the Zernike moment feature of each character image; S343, based on the normalized character image set, calculating the local binary pattern feature of each character image; S344, based on the normalized character image set, calculating the directional gradient histogram feature of each character image; S345, performing feature fusion on the Zernike moment feature, the local binary pattern feature and the oriented gradient histogram feature to obtain a fused feature vector set; S346. Based on the fused feature vector set, use the principal component analysis method to perform dimensionality reduction to obtain a feature vector set after dimensionality reduction; based on the feature vector set after dimensionality reduction, perform feature normalization to obtain a normalized feature vector set.
8. The archive digitization system based on intelligent image enhancement and automatic classification is characterized by: include: at least one processor; as well as, a memory communicatively connected to at least one of the processors; wherein, The memory stores instructions that can be executed by the processor, and the instructions are used to be executed by the processor to implement the archive digitization method based on intelligent image enhancement and automatic classification as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Digital management method and system based on enterprise archives and storage medium
CN116663549A
Document AI system based on deep learning
CN118470730A