ConvNeXt V2-based file document quality evaluation method and system
By using an improved ConvNeXt V2 model and an OCR text-image dual-branch structure, and dynamically adjusting the weights, accurate classification and adaptive quality assessment of archival document images were achieved, improving classification accuracy and the reasonableness of the assessment results.
Patent Information
- Application Number
- CN202511165252.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-20
- Publication Date
- 2025-11-14
AI Technical Summary
Existing technologies do not adequately extract text density features in archival document classification, resulting in low classification accuracy, especially in complex backgrounds. Furthermore, the quality assessment methods cannot adaptively adjust feature weights, leading to inaccurate evaluation results.
An improved ConvNeXt V2 model combined with an OCR text-image dual-branch structure is adopted. The ConvNeXt V2 model is used to classify text density, and the ViT model and ResNet are used to extract visual features. The weights of the OCR text and image branches are dynamically adjusted to achieve accurate classification and quality evaluation.
It improves the classification accuracy and quality evaluation of archival document images, adapts to the characteristics of documents with different text densities, and solves the evaluation bias problem caused by fixed weights in traditional methods.
Smart Images

Figure CN120954025A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of computer vision and document processing technology, and in particular relates to a method and system for evaluating the quality of archival documents based on ConvNeXt V2. Background Technology
[0002] In the process of digitizing archival management, it is necessary to classify and evaluate the quality of massive amounts of archival document images. Traditional document classification methods often lack the ability to capture local texture features such as text density, making it difficult to accurately distinguish between archival documents with low, medium, and high text density. Regarding quality evaluation, existing technologies typically use fixed weights to fuse different modal features, failing to adaptively adjust according to the characteristics of the document itself, resulting in inaccurate evaluation results.
[0003] Specifically, existing technologies have the following shortcomings: The document classification model does not fully extract the text density features of archival document images, and the classification accuracy is low, especially when dealing with archives with complex backgrounds and signs of aging.
[0004] In quality assessment methods, fixed-weight strategies cannot adapt to the feature differences of different document types. For example, when processing high-text-density documents such as contracts, the text error rate of OCR recognition is a key indicator, but existing systems still fuse visual features at a fixed ratio, resulting in errors not being effectively amplified; while for low-text-density documents such as historical photos, key indicators such as visual clarity are weakened due to insufficient weight.
[0005] Existing methods cannot dynamically adjust the focus of feature selection based on document type. In high-text-density documents, OCR results such as text integrity and grammatical errors should be the core evaluation criteria; while in low-text-density documents, visual features such as image resolution and color reproduction should be the primary focus. However, traditional systems lack dynamic weighting mechanisms, making it difficult to achieve accurate quality evaluation and classification. Summary of the Invention
[0006] The purpose of this invention is to provide a method and system for evaluating the quality of archival documents based on ConvNeXt V2, in order to solve the problems of inaccurate classification and unreasonable quality evaluation of archival documents in the prior art, and to achieve accurate classification and adaptive quality evaluation of archival document images.
[0007] To achieve the above-mentioned objectives, the present invention specifically adopts the following technical solution: In a first aspect, the present invention provides a method for evaluating the quality of archival documents based on ConvNeXt V2, which includes the following steps: S1. Obtain the preprocessed archive document image and input it into the trained improved ConvNeXt V2 model. After classifying the text density of the image, the improved ConvNeXt V2 model outputs the probability distribution corresponding to different levels of text density. S2. A trained quality assessment model is used to evaluate the quality of the preprocessed archival document images. The quality assessment model adopts an OCR text-image dual-branch structure: In the OCR text branch, the ViT model is used to extract primary visual features from the preprocessed archival document images, and then the primary visual features are fed into the transposed attention module to form high-level visual features, which are then input into the score weighting module to output the document stream score; In the image branch, a convolutional neural network is used as the backbone network to extract primary image features from the preprocessed archival document images, and ResNet is used to extract features from the primary image features to form high-level image features, which are then residually connected with the primary image features to form global multi-scale features, which are then input into a quality regression network composed of fully connected layers to output the image stream score; Finally, the image stream score and the document stream score are weighted and fused based on dual-stream dynamic weights to obtain the quality assessment score of the archival document images. The dual-stream dynamic weights consist of OCR text branch weights and image branch weights that sum to 1, and the OCR text branch weights are obtained by weighted summation of the probability distributions of text densities at different levels.
[0008] Based on the above scheme, each step can be implemented in the following preferred manner.
[0009] As a preferred embodiment of the first aspect mentioned above, in step S1, the improved ConvNeXt V2 model has four stages. In each ConvNeXt V2 Block of the first stage, the output features of the 7×7 deep convolutional layer are first processed by a spatial attention module, and then the spatial attention weight map generated by the spatial attention module is layer normalized. In each ConvNeXt V2 Block of the third stage, the output features of the 7×7 deep convolutional layer are first processed by a global self-attention layer, and then the global layout feature map generated by the global self-attention layer is layer normalized. In each ConvNeXt V2 Block of the fourth stage, the output of the original ConvNeXt V2 Block is processed by parallel pooling with three branches of 1×1, 2×2 and 4×4 respectively, and the pooled features are expanded and concatenated as the final output. Each ConvNeXt V2 Block of the second stage has the same structure as the ConvNeXt V2 Block of the ConvNeXt V2 model.
[0010] As a preferred embodiment of the first aspect, in the OCR text branch of step S2, the preprocessed archival document image is first divided into a group of image blocks of equal size. Then, the image blocks are converted into low-dimensional vector representations as input to the ViT model. The primary visual features output by the ViT model pass through two cascaded transposed attention modules in sequence. The second transposed attention module outputs high-level visual features, which are then input to the score weighting module. The score weighting module includes a weight calculation branch and a score prediction branch. The weight calculation branch outputs the weight of each image block, and the score prediction branch outputs the score of each image block. Then, the weight obtained by each image block in the weight calculation branch is multiplied by the score obtained by the score prediction branch to obtain the document flow score of that image block. Finally, the document flow scores of all image blocks are added together to obtain the document flow score of the entire archival document image.
[0011] As a preferred embodiment of the first aspect mentioned above, the weight calculation branch is composed of a linear layer, a ReLU activation function, another linear layer, and a Sigmoid function cascaded in sequence.
[0012] As a preferred embodiment of the first aspect mentioned above, the score prediction branch is composed of a linear layer, a ReLU activation function, another linear layer, and another ReLU activation function cascaded in sequence.
[0013] As a preferred embodiment of the first aspect mentioned above, in step S2, the probability distribution of the output document image of the improved ConvNeXt V2 model belonging to low, medium, and high text density is then multiplied by a preset weight coefficient for each of the low, medium, and high text density probability distributions, and the multiplication result is used as the OCR text branch weight. Based on the OCR text branch weight, the image branch weight is calculated so that the sum of the two weights is 1.
[0014] Secondly, the present invention provides an archival document quality evaluation system based on ConvNeXt V2, which includes: The classification module is used to acquire the preprocessed archival document image and input it into the trained improved ConvNeXt V2 model. After classifying the text density of the image, the improved ConvNeXt V2 model outputs the probability distribution corresponding to different levels of text density. The quality assessment module is used to evaluate the quality of preprocessed archival document images using a trained quality assessment model. This model employs a dual-branch OCR text-image structure: In the OCR text branch, a ViT model is used to extract primary visual features from the preprocessed archival document images. These primary visual features are then fed into a transposed attention module to form advanced visual features, which are then input into a score weighting module to output a document stream score. In the image branch, a convolutional neural network is used as the backbone to extract primary image features from the preprocessed archival document images. ResNet is then used to extract features from these primary image features, forming advanced image features. These advanced features are then residually connected to the primary image features to form global multi-scale features, which are then input into a quality regression network composed of fully connected layers to output an image stream score. Finally, the image stream score and document stream score are weighted and fused based on dual-stream dynamic weights to obtain the archival document image's quality assessment score. The dual-stream dynamic weights consist of OCR text branch weights and image branch weights that sum to 1. The OCR text branch weights are obtained by weighted summation of the probability distributions of different levels of text density.
[0015] Thirdly, the present invention provides a computer program product, including a computer program / instruction, which, when executed by a processor, can implement the archival document quality evaluation method based on ConvNeXt V2 as described in any of the solutions of the first aspect above.
[0016] Fourthly, the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the archival document quality evaluation method based on ConvNeXt V2 as described in any of the solutions of the first aspect above.
[0017] Fifthly, the present invention provides a computer electronic device, which includes a memory and a processor; The memory is used to store computer programs; The processor is configured to, when executing the computer program, implement the archival document quality evaluation method based on ConvNeXt V2 as described in any of the first aspects above.
[0018] Compared with the prior art, the present invention has the following advantages: This invention achieves accurate classification and adaptive quality evaluation of archival document images by combining an improved ConvNeXt V2 model with a dual-stream dynamic weight evaluation algorithm. ConvNeXt V2 significantly improves the classification accuracy of documents with different text densities by embedding an attention mechanism and multi-scale pooling. Based on this, this invention dynamically adjusts the evaluation weights of the OCR text stream (ViT) and image stream (ResNet) according to the classification results, so that high text density documents focus on text quality and low text density documents focus on image quality. This solves the evaluation bias problem caused by fixed weights in traditional methods and improves the overall intelligence and accuracy of archival digitization processing. Attached Figure Description
[0019] Figure 1 This is a schematic diagram of the overall process of the method of the present invention; Figure 2 This is a detailed flowchart illustrating the improved ConvNeXt V2 model in this invention; Figure 3 This is a detailed flowchart illustrating the process of quality score fusion in the method of the present invention; Figure 4 This is a schematic diagram of the internal structure of the original ConvNeXt V2 Block in an embodiment of the present invention; Figure 5 This is a schematic diagram of the improved internal structure of ConvNeXt V2 Stage1 in an embodiment of the present invention; Figure 6 This is a schematic diagram of the improved internal structure of ConvNeXt V2 Stage3 in an embodiment of the present invention; Figure 7 This is a schematic diagram of the improved internal structure of ConvNeXt V2 Stage4 in an embodiment of the present invention; Figure 8 This is a schematic diagram of the internal structure of the two branches of the fraction weighting module in an embodiment of the present invention; Figure 9 This is a system block diagram of the present invention; Figure 10 This is a schematic diagram of the composition of a computer electronic device according to the present invention. Detailed Implementation
[0020] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Many specific details are set forth in the following description to provide a thorough understanding of the present invention. However, the present invention can be practiced in many other ways different from those described herein, and those skilled in the art can make similar modifications without departing from the spirit of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below. Technical features in the various embodiments of the present invention can be combined accordingly without mutual conflict.
[0021] like Figure 1 As shown, in a preferred embodiment of the present invention, the above-mentioned document quality evaluation method based on ConvNeXt V2 includes the following steps S1 to S2. The specific implementation process of each step will be described in detail below.
[0022] S1. Obtain the preprocessed archive document image and input it into the trained improved ConvNeXt V2 model. After classifying the text density of the image, the improved ConvNeXt V2 model outputs the probability distribution corresponding to different levels of text density.
[0023] It should be noted that the preprocessing operations performed on the archive document image in step S1 of this embodiment mainly include the following aspects: 1) Image acquisition and format standardization Collect archival document images (supporting JPG and PNG formats), and convert them to a unified RGB format using an image decoding tool to eliminate feature extraction bias caused by format differences.
[0024] 2) Noise Removal and Enhancement Adaptive median filtering is used to remove salt-and-pepper noise generated during the scanning process, and Gamma correction is used to adjust the brightness and contrast of the image to ensure the distinction between text areas and background in aged and faded archives.
[0025] 3) Tilt correction and edge trimming The Hough transform is used to detect straight lines at the edges of the document, calculate the tilt angle and perform rotation correction (correction accuracy ≤ 0.5°), crop redundant edge areas (such as scan borders and finger shadows), and retain the effective document content area.
[0026] It should be noted that in step S1 of this invention, the improved ConvNeXt V2 model has four stages. In each ConvNeXt V2 Block of the first stage, the output features of the 7×7 deep convolutional layer are first processed by a spatial attention module, and then the spatial attention weight map generated by the spatial attention module is layer normalized. In each ConvNeXt V2 Block of the third stage, the output features of the 7×7 deep convolutional layer are first processed by a global self-attention layer, and then the global layout feature map generated by the global self-attention layer is layer normalized. In each ConvNeXt V2 Block of the fourth stage, the output of the original ConvNeXt V2 Block is respectively processed by 1×1, 2×2 and 4×4 three-branch parallel pooling, and the pooled features are expanded and concatenated as the final output. Each ConvNeXt V2 Block of the second stage has the same structure as the ConvNeXt V2 Block of the ConvNeXt V2 model.
[0027] In step S1 of this embodiment, the structure of one of the blocks in the original ConvNeXt V2 model is as follows: Figure 4 As shown. This invention improves the original ConvNeXt V2 model, designing an improved ConvNeXt V2 model for feature extraction from input archival document images. In the improved ConvNeXt V2 model, as... Figure 2 As shown, the image size of the archive document is first adjusted to 384×384, and then downsampled to 96×96×96 through 4×4 convolution. While compressing the data volume, key local details such as text edges and stroke outlines are preserved, laying the foundation for subsequent feature extraction. The improved ConvNeXt V2 model has four stages, and the specific improvements are as follows: Figure 5 As shown, a spatial attention module is embedded in each ConvNeXt V2 Block of Stage 1 of the original ConvNeXt V2 model. This module generates a spatial attention weight map through 3×3 convolution to enhance the feature response of text regions; as shown... Figure 6 As shown, a global self-attention layer is added to each ConvNeXt V2Block of the original ConvNeXt V2 model Stage 3 to capture the global layout features of text paragraphs in the document (such as row and column distribution, spacing patterns); for example... Figure 7As shown, 1×1, 2×2, and 4×4 multi-scale pooling is performed on the Stage 4 features of the original ConvNeXt V2 model. 1×1 pooling focuses on the details of local dense text regions, while 2×2 and 4×4 pooling statistically analyze the global text density distribution. This multi-scale fusion achieves feature complementarity between "local details" and "global statistics." Furthermore, the improved ConvNeXt V2 model designed in this invention retains the architecture of the original ConvNeXt V2 model. Each stage of the ConvNeXt V2 Block contains a 7×7 depthwise convolution (enhancing local feature extraction capabilities), LayerNorm (for stabilizing the training process), a 1×1 convolution (dimensionality transformation), a GELU activation function (introducing non-linear features), a 1×1 convolution (feature mapping), LayerScale (to alleviate gradient vanishing), and residual connections + DropPath (to prevent overfitting). These components work together to improve the robustness of feature extraction. Then, a linear layer is used as the classification head to process the high-dimensional features extracted by the last ConvNeXt V2 Block, mapping them to three target categories (low, medium, and high text density). The classification result is then output, which is the probability distribution of low text density. Probability distribution of Chinese character density and probability distribution of high text density The probability values range from [0,1], and the sum of the three is 1, providing a basis for the dynamic weight calculation of subsequent quality evaluation.
[0028] It should also be noted that the implementation methods of the spatial attention module and the global self-attention layer in step S1 of this invention are both existing technologies, and therefore will not be described in detail.
[0029] S2. A trained quality assessment model is used to evaluate the quality of the preprocessed archival document images. The quality assessment model adopts an OCR text-image dual-branch structure: In the OCR text branch, the ViT model is used to extract primary visual features from the preprocessed archival document images, and then the primary visual features are fed into the transposed attention module to form high-level visual features, which are then input into the score weighting module to output the document stream score; In the image branch, a convolutional neural network (CNN) is used as the backbone network to extract primary image features from the preprocessed archival document images, and ResNet is used to extract features from the primary image features to form high-level image features, which are then residually connected with the primary image features to form global multi-scale features, which are then input into a quality regression network composed of fully connected layers to output the image stream score; Finally, the image stream score and the document stream score are weighted and fused based on dual-stream dynamic weights to obtain the quality assessment score of the archival document images. The dual-stream dynamic weights consist of OCR text branch weights and image branch weights that sum to 1, and the OCR text branch weights are obtained by weighted summation of the probability distributions of text densities at different levels.
[0030] It should be noted that in step S2 of this invention, the improved ConvNeXt V2 model outputs a probability distribution of low, medium and high text density in the document image. Then, the probability distributions of low, medium and high text density are multiplied by preset weight coefficients, and the multiplication result is used as the OCR text branch weight. Based on the OCR text branch weight, the image branch weight is calculated so that the sum of the two weights is 1.
[0031] It should be noted that in the OCR text branch of step S2 of this invention, the preprocessed archival document image is first split into a group of image blocks of equal size. Then, the image blocks are converted into low-dimensional vector representations as input to the ViT model. The primary visual features output by the ViT model pass through two cascaded transposed attention modules in sequence. The second transposed attention module outputs high-level visual features, which are then input to the score weighting module. The score weighting module includes a weight calculation branch and a score prediction branch. The weight calculation branch outputs the weight of each image block, and the score prediction branch outputs the score of each image block. Then, the weight obtained by each image block in the weight calculation branch is multiplied by the score obtained by the score prediction branch to obtain the document flow score of that image block. Finally, the document flow scores of all image blocks are added together to obtain the document flow score of the entire archival document image.
[0032] It should be noted that, as Figure 8 As shown, in the score weighting module, the weight calculation branch is composed of a linear layer, a ReLU activation function, a linear layer, and a Sigmoid function cascaded in sequence, while the score prediction branch is composed of a linear layer, a ReLU activation function, a linear layer, and a ReLU activation function cascaded in sequence.
[0033] In this embodiment, for the image branch, a CNN is used as the backbone network to extract primary image features, and ResNet is used as the feature extractor for the natural image stream to further extract high-level image features. To map the extracted high-level and low-level image features to the image stream score, this invention constructs a quality regression network composed of fully connected layers. This network takes the global multi-scale features formed by concatenating high-level and low-level image features as input, and ReLU as the activation function. After propagation through the aforementioned fully connected layers, the final quality evaluation score of the archival document image as a natural image, i.e., the image stream score, is obtained.
[0034] In this embodiment, for the OCR text branch, the ViT model (Vision Transformer) is used to extract features from the archival document image. During this process, such as... Figure 3 As shown, the archival document image is first divided into a set of image patches of equal size, with each patch treated as an element in the sequence. These patches are then transformed into low-dimensional vector representations, which serve as input to the ViT model. Next, through a multi-head attention mechanism, ViT can simultaneously consider the correlations between different locations in the image, as well as the correlations between different image patches, thereby capturing global information. This allows the model to gain a more comprehensive understanding of the archival document image and extract more representative features.
[0035] The features output by the ViT model are fed into two cascaded transpose attention modules. Each transpose attention module first reshapes the feature matrix of size H×W×C and then maps the new query matrix using a fully connected layer. Key matrix Value matrix ,Will After transpose and The dot product yields the transpose attention matrix of size C×C, which is then compared with... The result of multiplication and the original input Add them together to get the output. The process is defined by the following formula: in, This represents the weight matrix, used to perform a linear transformation on the transposed attention output; This represents the Softmax function, used to convert the dot product result into a probability distribution; This represents the scaling factor, used to control the scaling degree of the Softmax function; Yes and Scale the dot product result (divide by) Then, apply the Softmax function to obtain the transposed attention matrix; This indicates that the transposed attention matrix and the value matrix are... The transposed attention output is obtained after multiplication; H, W, and C are the height, width, and number of channels of the feature matrix, respectively.
[0036] Finally, the features output from the second transposed attention module are input into the score weighting module, which consists of two main branches: a weight calculation branch and a score prediction branch. The input features are processed in these two branches respectively. In the weight calculation branch, the model focuses on the text information in the image, while in the score prediction branch, the model predicts the score for each image patch. Then, the weight obtained for each image patch in the weight calculation branch is multiplied by the score obtained in the score prediction branch to obtain the quality score of that image patch. Finally, the scores of all image patches are summed to obtain the final document quality score, or document flow score, for the entire archival document image.
[0037] After the OCR text and image branches are processed as described above, a quality score is fused. First, dual-stream dynamic weights are calculated. Specifically, based on the probability distribution of low, medium, and high text density in the archival document image output by the improved ConvNeXt V2 model, the weights of the OCR text and image branches are calculated using preset weight coefficients. Furthermore, when designing the weights, the following conditions must be met simultaneously: the image branch corresponding to the low text density probability distribution has a larger weight, and the OCR text branch corresponding to the high text density probability distribution has a larger weight. Then, based on the calculated dual-stream dynamic weights, the document stream is scored. Image stream score Weighted fusion is performed to obtain the final quality evaluation score of the archival document images. : in, , , These are preset weighting coefficients; Weights for OCR text branches; For image branch weights.
[0038] It should also be noted that the implementation methods of the ViT model and the transposed attention module in step S2 of this invention are both existing technologies, and therefore will not be described in detail.
[0039] The present invention will now demonstrate the application effect of the ConvNeXt V2-based archival document quality assessment method described in S1~S2 of the above embodiments on a specific dataset through a specific example, so as to facilitate understanding of the essence of the present invention.
[0040] Example The specific implementation process of the archival document quality evaluation method based on ConvNeXt V2 used in this embodiment is as described above and will not be repeated here. The processing process of this embodiment is briefly introduced below.
[0041] In this embodiment, a total of 15,000 archival document images were selected from the dataset (5,000 images with low text density, such as photos and hand-drawn illustrations; 5,000 images with Chinese text density, such as reports, resumes, and tables; and 5,000 images with high text density, such as contracts and applications). The data sources were public archival datasets (such as the ICDAR document image database), scanned data from real archives (after anonymization), and online images. Additionally, the Tobacco3482 public document dataset, containing 35,000 archival document images across 16 categories (including resumes, reports, and hand-drawn illustrations), was used, and these 16 categories were manually divided into 3 subcategories.
[0042] The dataset was then preprocessed, including the following steps: 1) Format unification: JPG and PNG formats were converted to a unified RGB format using image decoding tools to eliminate feature extraction bias caused by format differences. 2) Noise removal and enhancement: Adaptive median filtering was used to remove salt-and-pepper noise, and Gamma correction was used to adjust image brightness and contrast, improving the distinction between text areas and background in aged and faded documents. 3) Tilt correction and edge cropping: Hough transform was used to detect straight lines at document edges, the tilt angle was calculated and rotated for correction (correction accuracy ≤ 0.5 degrees), and redundant edge areas (such as scan borders and finger shadows) were cropped, retaining the effective document content area.
[0043] After preprocessing, the preprocessed dataset is divided into training and validation sets in an 8:2 ratio. The training set, comprising 80% of the total samples, is used for model parameter learning and covers archival document images with low, medium, and high text density, including a variety of samples to ensure that the model learns comprehensive text density features and quality evaluation rules. The validation set, comprising 20% of the total samples, is used to monitor model performance during training, avoid overfitting, and guide model parameter adjustment.
[0044] Next, the improved ConvNeXt V2 model was trained. The AdamW optimizer was used, with an initial learning rate of 5e-4. A cosine annealing scheduling strategy was employed to dynamically adjust the learning rate, avoiding parameter oscillations in the later stages of training. Simultaneously, training was stopped and the current model parameters were saved when the validation set classification accuracy fluctuated by ≤0.5% for five consecutive rounds to prevent overfitting.
[0045] Then, the quality assessment model is trained, and the weights of the OCR branch and the image branch are dynamically adjusted during training to adapt the model to the quality assessment requirements of documents with different text densities. Every 10 rounds, SRCC and mean RMSE are calculated on the validation set, and the model with the highest SRCC and the lowest RMSE is saved as the optimal quality assessment model.
[0046] The experimental results of the method of the present invention on the Tobacco3482 document dataset are shown in Table 1, and the experimental results of the method of the present invention on the self-constructed Chinese document dataset are shown in Table 2.
[0047] Table 1. Experimental results on the Tobacco3482 document dataset. Table 2. Experimental results on a self-constructed Chinese document dataset. To further verify the effectiveness of the dual-branch structure proposed in this invention, an ablation experiment was also conducted in this embodiment. The weights of both the OCR text branch and the image branch were set to 0.5, without calculating weights based on probability distribution. For simplicity, this fixed-weight method is referred to as Method 1, the method using only the OCR text branch as Method 2, and the method using only the image branch as Method 3. The quality evaluation results of the method of this invention and Methods 1, 2, and 3 on the SRCC index are shown in Table 3. The quality evaluation results of the method of this invention and Methods 1, 2, and 3 on the RMSE index are shown in Table 4. In Tables 3 and 4, L represents low text density accuracy, M represents Chinese text density accuracy, and H represents high text density accuracy.
[0048] Table 3. Ablation experiments on SRCC index Table 4. Ablation experiments on RMSE index In summary, the method of this invention addresses the problem of insufficient local texture feature capture in original models for archival document images, enhancing the ability to extract text features against complex backgrounds. Adding a global self-attention layer to Stage 3 of ConvNeXt V2 allows for the capture of the overall structural patterns of text paragraphs across regions, avoiding layout misjudgments caused by the original model's reliance on only local features, and improving the accuracy of capturing global document layout features. Simultaneously, the use of a "local details + global statistics" fusion mechanism enables the model to adapt to both the sparse features of low-text-density documents and the dense layout of high-text-density documents, solving the problem of incomplete feature extraction at a single scale in traditional models.
[0049] In the quality assessment phase, the improved method, based on text density classification results (low / medium / high probability distribution), adaptively adjusts the weights of the OCR and image branches through a dynamic weight function: higher weights are assigned to the visual branch (emphasizing clarity and color reproduction) for low text density documents (such as historical photos), while higher weights are assigned to the OCR branch (emphasizing text error rate and completeness) for high text density documents (such as contracts). Compared to a fixed weight strategy, this approach better aligns with the core needs of quality assessment for different document types, improving the reasonableness of the assessment results.
[0050] It should also be noted that the archival document quality assessment method based on ConvNeXt V2 in the above embodiments can essentially be executed by a computer program or module. Therefore, similarly, based on the same inventive concept, another preferred embodiment of the present invention also provides an archival document quality assessment system based on ConvNeXt V2, corresponding to the archival document quality assessment method based on ConvNeXt V2 provided in the above embodiments, such as... Figure 9 As shown, it includes: The classification module is used to acquire the preprocessed archival document image and input it into the trained improved ConvNeXt V2 model. After classifying the text density of the image, the improved ConvNeXt V2 model outputs the probability distribution corresponding to different levels of text density. The quality assessment module is used to evaluate the quality of preprocessed archival document images using a trained quality assessment model. This model employs a dual-branch OCR text-image structure: In the OCR text branch, a ViT model is used to extract primary visual features from the preprocessed archival document images. These primary visual features are then fed into a transposed attention module to form advanced visual features, which are then input into a score weighting module to output a document stream score. In the image branch, a convolutional neural network is used as the backbone to extract primary image features from the preprocessed archival document images. ResNet is then used to extract features from these primary image features, forming advanced image features. These advanced features are then residually connected to the primary image features to form global multi-scale features, which are then input into a quality regression network composed of fully connected layers to output an image stream score. Finally, the image stream score and document stream score are weighted and fused based on dual-stream dynamic weights to obtain the archival document image's quality assessment score. The dual-stream dynamic weights consist of OCR text branch weights and image branch weights that sum to 1. The OCR text branch weights are obtained by weighted summation of the probability distributions of different levels of text density.
[0051] It is understood that the archival document quality assessment method based on ConvNeXt V2 described in S1~S2 above can essentially be implemented by a computer program. Therefore, based on the same inventive concept, another preferred embodiment of the present invention also provides a computer program product corresponding to the archival document quality assessment method based on ConvNeXt V2 provided in the above embodiments, which includes a computer program / instructions. When executed by a processor, the computer program / instructions can implement the archival document quality assessment method based on ConvNeXt V2 as described in the above embodiments.
[0052] Similarly, based on the same inventive concept, another preferred embodiment of the present invention also provides a computer electronic device corresponding to the archival document quality evaluation method based on ConvNeXt V2 provided in the above embodiments, such as... Figure 10 As shown, it includes a memory and a processor; The memory is used to store computer programs; The processor is configured to implement the ConvNeXtV2-based archival document quality evaluation method in the above embodiments when executing the computer program.
[0053] Furthermore, the logical instructions in the aforementioned memory can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.
[0054] Therefore, based on the same inventive concept, another preferred embodiment of the present invention also provides a computer-readable storage medium corresponding to the archival document quality evaluation method based on ConvNeXt V2 provided in the above embodiments. The storage medium stores a computer program, which, when executed by a processor, can implement the archival document quality evaluation method based on ConvNeXt V2 in the above embodiments.
[0055] It is understood that the aforementioned storage media may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Furthermore, the storage media may also be various media capable of storing program code, such as USB flash drives, external hard drives, magnetic disks, or optical discs.
[0056] It is understood that the processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0057] It should also be noted that those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the system described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here. In the embodiments provided in this application, the division of steps or modules in the system and method is merely a logical functional division, and there may be other division methods in actual implementation. For example, multiple modules or steps may be combined or integrated together, and a module or step may also be split.
[0058] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the invention. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the invention. Therefore, all technical solutions obtained through equivalent substitution or transformation fall within the protection scope of the present invention.
Claims
1. A method for evaluating the quality of archival documents based on ConvNeXt V2, characterized in that, Includes the following steps: S1. Obtain the preprocessed archive document image and input it into the trained improved ConvNeXt V2 model. After classifying the text density of the image, the improved ConvNeXt V2 model outputs the probability distribution corresponding to different levels of text density. S2. A trained quality assessment model is used to evaluate the quality of the preprocessed archival document images. The quality assessment model adopts an OCR text-image dual-branch structure: In the OCR text branch, the ViT model is used to extract primary visual features from the preprocessed archival document images, and then the primary visual features are fed into the transposed attention module to form high-level visual features, which are then input into the score weighting module to output the document stream score; In the image branch, a convolutional neural network is used as the backbone network to extract primary image features from the preprocessed archival document images, and ResNet is used to extract features from the primary image features to form high-level image features, which are then residually connected with the primary image features to form global multi-scale features, which are then input into a quality regression network composed of fully connected layers to output the image stream score; Finally, the image stream score and the document stream score are weighted and fused based on dual-stream dynamic weights to obtain the quality assessment score of the archival document images. The dual-stream dynamic weights consist of OCR text branch weights and image branch weights that sum to 1, and the OCR text branch weights are obtained by weighted summation of the probability distributions of text densities at different levels.
2. The method for evaluating the quality of archival documents based on ConvNeXt V2 as described in claim 1, characterized in that, In step S1, the improved ConvNeXt V2 model has four stages. In each ConvNeXt V2 Block of the first stage, the output features of the 7×7 deep convolutional layer are first processed by a spatial attention module, and then the spatial attention weight map generated by the spatial attention module is normalized. In each ConvNeXt V2 Block of the third stage, the output features of the 7×7 deep convolutional layer are first processed by a global self-attention layer, and then the global layout feature map generated by the global self-attention layer is normalized. In each ConvNeXt V2 Block of the fourth stage, the output of the original ConvNeXt V2Block is processed by parallel pooling with three branches of 1×1, 2×2 and 4×4 respectively, and the pooled features are unfolded and concatenated as the final output. Each ConvNeXt V2 Block of the second stage has the same structure as the ConvNeXt V2Block of the ConvNeXt V2 model.
3. The method for evaluating the quality of archival documents based on ConvNeXt V2 as described in claim 1, characterized in that, In the OCR text branch of step S2, the preprocessed archival document image is first split into a group of image blocks of equal size. Then, the image blocks are converted into low-dimensional vector representations as input to the ViT model. The primary visual features output by the ViT model pass through two cascaded transposed attention modules. The second transposed attention module outputs high-level visual features, which are then input to the score weighting module. The score weighting module includes a weight calculation branch and a score prediction branch. The weight calculation branch outputs the weight of each image block, and the score prediction branch outputs the score of each image block. Then, the weight obtained by each image block in the weight calculation branch is multiplied by the score obtained by the score prediction branch to obtain the document flow score of that image block. Finally, the document flow scores of all image blocks are added together to obtain the document flow score of the entire archival document image.
4. The method for evaluating the quality of archival documents based on ConvNeXt V2 as described in claim 3, characterized in that, The weight calculation branch is composed of a linear layer, a ReLU activation function, another linear layer, and a Sigmoid function cascaded in sequence.
5. The method for evaluating the quality of archival documents based on ConvNeXt V2 as described in claim 3, characterized in that, The score prediction branch is composed of a linear layer, a ReLU activation function, a linear layer, and a ReLU activation function cascaded in sequence.
6. The method for evaluating the quality of archival documents based on ConvNeXt V2 as described in claim 1, characterized in that, In step S2, the improved ConvNeXt V2 model outputs the probability distribution of low, medium, and high text density in the document image. Then, the probability distributions of low, medium, and high text density are multiplied by preset weight coefficients, and the multiplication result is used as the OCR text branch weight. Based on the OCR text branch weight, the image branch weight is calculated so that the sum of the two weights is 1.
7. An archival document quality evaluation system based on ConvNeXt V2, characterized in that, include: The classification module is used to acquire the preprocessed archival document image and input it into the trained improved ConvNeXtV2 model. After classifying the text density of the image, the improved ConvNeXt V2 model outputs the probability distribution corresponding to different levels of text density. The quality assessment module is used to evaluate the quality of preprocessed archival document images using a trained quality assessment model. This model employs a dual-branch OCR text-image structure: In the OCR text branch, a ViT model is used to extract primary visual features from the preprocessed archival document images. These primary visual features are then fed into a transposed attention module to form advanced visual features, which are then input into a score weighting module to output a document stream score. In the image branch, a convolutional neural network is used as the backbone to extract primary image features from the preprocessed archival document images. ResNet is then used to extract features from these primary image features, forming advanced image features. These advanced features are then residually connected to the primary image features to form global multi-scale features, which are then input into a quality regression network composed of fully connected layers to output an image stream score. Finally, the image stream score and document stream score are weighted and fused based on dual-stream dynamic weights to obtain the archival document image's quality assessment score. The dual-stream dynamic weights consist of OCR text branch weights and image branch weights that sum to 1. The OCR text branch weights are obtained by weighted summation of the probability distributions of different levels of text density.
8. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instruction is executed by the processor, it can implement the archival document quality evaluation method based on ConvNeXt V2 as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, implements the archival document quality evaluation method based on ConvNeXt V2 as described in any one of claims 1 to 6.
10. A computer electronic device, characterized in that, Including memory and processor; The memory is used to store computer programs; The processor is configured to, when executing the computer program, implement the archival document quality evaluation method based on ConvNeXt V2 as described in any one of claims 1 to 6.
Citation Information
Cited By
CT image pulmonary embolism segmentation and classification method combined with quality evaluation
CN121121129A
A CT image pulmonary embolism segmentation and classification method combined with quality evaluation
CN121121129B