Text image quality detection method, system, equipment and medium
By adopting a detection model without reference images in text image quality evaluation, and using a hybrid framework composed of convolutional neural network and Transformer, the problems of insufficient accuracy and poor robustness of text image quality evaluation in the prior art are solved, and high-precision and real-time text image quality evaluation are achieved.
Patent Information
- Application Number
- CN202411881127.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-19
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-12-19
AI Technical Summary
The prior art has problems of insufficient accuracy and poor robustness in text image quality evaluation, especially in the absence of reference images, which are difficult to effectively evaluate.
Using a text image quality detection model without reference images, a hybrid framework composed of a convolutional neural network and a Transformer-based encoder is used to extract the local and global features of the text image and perform quality detection. This model does not require a fixed reference image and can be more flexible and practical in practical application scenarios.
It improves the detection accuracy and accuracy of text image quality evaluation, enhances the robustness of the model, and achieves accurate and real-time text image quality evaluation.
Smart Images

Figure CN120047381A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image quality assessment, and particularly relates to a method, system, device and medium for text image quality detection. Background Art
[0002] Most existing image quality assessment methods are usually used to evaluate natural scene images, but perform poorly on text images. The quality assessment of natural scene images often depends on features such as color, texture and structural information. The content of text images is mainly characters and detailed edges, and the color and texture information it contains is limited. The main effective information of text images is concentrated in the high-frequency region and edges, which results in most image quality assessment methods being unable to extract key features from text images or identify quality defects. In addition, it is very difficult to obtain a high-quality reference image during the text image quality assessment process. For example, in OCR applications, there may be problems such as noise, blur, and distortion between the scanned image and the original image, and even a high-quality reference image may not perfectly match the characteristics of each different type of text image. Therefore, the effectiveness of most quality assessment methods is limited.
[0003] Traditional text image quality assessment mainly relies on manual visual inspection, which has the problems of strong subjectivity, low efficiency and poor consistency. With the continuous development of computer vision and image processing technologies, it has become more feasible to use automated methods to assess the quality of text images. Image quality assessment is divided into subjective assessment and objective assessment, and objective assessment is further divided into full reference (FR-IQA), reduced reference (RR-IQA) and no reference (NR-IQA). The full reference image quality assessment method mainly evaluates the quality of a distorted image by using the difference between the distorted image and the ideal reference image. The reduced reference image quality assessment method extracts a small amount of feature information of the ideal reference image using prior knowledge, and compares it with the feature information of the distorted image to complete the quality assessment of the distorted image. The no reference image quality assessment method NR-IQA evaluates the quality of a distorted image without a reference image at all, and is more suitable for actual application scenarios. For text image quality assessment,
[0004] Traditional NR-IQA methods utilize the perceptual characteristics of the human visual system (HVS) to obtain useful features from the image to be evaluated for modeling, and map the degree of image distortion to a quality score; for example, Liu et al. proposed the SNP-NIQE algorithm, which extracts the natural statistical characteristics of the distorted image from three aspects: structural, natural and perceptual, and combines unsupervised learning for image quality assessment; Li et al. proposed using histograms to represent the brightness statistical features and structural statistical features to perceive image quality changes, called NRSL, and used rotation-invariant local binary patterns to extract structural features.
[0005] For text image quality assessment, it is difficult to obtain high-quality reference images, and the FR-IQA method and RR-IQA method are not applicable. However, traditional NR-IQA methods have certain limitations, which easily lead to inaccurate evaluation results. Summary of the Invention
[0006] To solve the deficiencies and defects in the prior art, the present invention provides a method, system, device, and medium for text image quality detection. A text image quality detection model without a reference image is used to detect the quality of text images, improving the detection accuracy of text image quality assessment, as well as the accuracy and robustness of text image quality detection, and achieving accurate and real-time text image quality assessment.
[0007] To solve the above problems, a first aspect of the present invention provides a method for text image quality detection, including:
[0008] Obtain a text image and divide the text image into a training set and a test set;
[0009] Construct a text image quality detection model and input the training set into the text image quality detection model for training;
[0010] Input the test set into the trained text image quality detection model to obtain a detection result, and evaluate and optimize the text image quality detection model.
[0011] Further, the evaluation of the text image quality detection model includes:
[0012] Obtain the subjective score of the text image, and combine the subjective score with the output detection result to evaluate the optimized linear correlation coefficient (OLCC) between the subjective score and the detection result:
[0013]
[0014] Where N represents the number of batch images, y i and respectively represent the subjective score of the i-th image and the output detection result, which are obtained through the subjective score of the actual human eye annotation data and the detection of a single training picture by the text image quality detection model, respectively; and respectively represent the average subjective score and the average detection result, and m i is the weight corresponding to the i-th image, which is used to adjust the weight ratio of the improved OLCC.
[0015] Further, the evaluation of the text image quality detection model further includes:
[0016] Obtain the subjective score of the text image and the ranking between the subjective scores;
[0017] Obtain the detected results of the output and the sorting among the detected results;
[0018] Based on the subjective score sorting and the detected result sorting, evaluate the rank correlation OSROCC between the subjective score and the detected results:
[0019]
[0020] where, v i and p i respectively represent the sorting positions of y i and in the subjective score sequence and the detected result sequence, and w 3 is the corresponding weight, used to adjust the amplitude of the overall monotonicity.
[0021] Furthermore, for the construction of the text image quality detection model, training the training set by inputting it into the text image quality detection model includes:
[0022] Input the text image into the first feature extraction module to obtain the first feature image;
[0023] Input the first feature image into the second feature extraction module to obtain the second feature image;
[0024] Concatenate the first feature image and the second feature image through a fully connected layer, and map and output the image quality detection result.
[0025] Furthermore, for the construction of the text image quality detection model, training the training set by inputting it into the text image quality detection model further includes:
[0026] Generate position encoding information based on the first feature image;
[0027] Combine the position encoding information with the first feature image to obtain the first feature image with position information, and input it into the second feature extraction module.
[0028] Furthermore, for the construction of the text image quality detection model, training the training set by inputting it into the text image quality detection model further includes:
[0029] Train the text image quality detection model by minimizing the regression loss function:
[0030]
[0031] where, n represents the number of encoder layers, q i is the detected result of the i-th image, s i is the subjective score of the i-th image, w 1 and w 2They are the weights corresponding to the detection results and subjective scores respectively.
[0032] Furthermore, the first feature extraction module is a convolutional neural network, and the second feature extraction module is a Transformer-based encoder.
[0033] The second aspect of the present invention provides a text image quality detection system for implementing the above-mentioned text image quality detection method, including a text image quality detection model and a model evaluation module. The text image quality detection model includes:
[0034] A first feature extraction module, configured to extract features from the input text image and output a first feature image;
[0035] A position encoding module, configured to generate position encoding information for the first feature image output by the first feature extraction module to form a first feature image with position information;
[0036] A second feature extraction module, configured to extract features from the first feature image with position information and output a second feature image;
[0037] A fully connected layer, configured to splice the first feature image and the second feature image and map and output the detection result;
[0038] The model evaluation module is configured to evaluate and optimize the detection accuracy of the text image quality detection model in combination with the subjective score of the text image.
[0039] The third aspect of the present invention provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the above-mentioned method are implemented.
[0040] The fourth aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned method are implemented.
[0041] Compared with the prior art, the present invention has the following beneficial effects:
[0042] The present invention discloses a method, system, device and medium for text image quality detection. The method includes obtaining a text image, dividing the text image into a training set and a test set; constructing a text image quality detection model, and inputting the training set into the text image quality detection model for training; inputting the test set into the trained text image quality detection model to obtain a detection result, and evaluating and optimizing the text image quality detection model; using a text image quality detection model without a reference image to detect the quality of the text image, which does not require a fixed reference image, is more suitable for actual application scenarios, is more flexible and practical, improves the detection accuracy of text image quality evaluation, and improves the accuracy and robustness of text image quality detection, so as to achieve accurate and real-time text image quality evaluation. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] The following further elaborates in detail the specific embodiments of the present invention with reference to the accompanying drawings, wherein:
[0044] Figure 1 It is a schematic structural diagram of the text image quality detection system described in the embodiment;
[0045] Figure 2 It is a flowchart of the text image quality detection method described in the embodiment;
[0046] Figure 3 It is a flowchart of training the text image quality detection model of the text image quality detection method described in the embodiment;
[0047] Figure 4 It is a schematic structural diagram of the computer device described in the embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0048] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Usually, the components of the embodiments of the present invention described and illustrated herein can be arranged and designed in various different configurations.
[0049] Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0050] It should be noted that similar reference numerals and letters denote similar items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. At the same time, in the description of the present invention, terms such as "first" and "second" are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
[0051] It should be noted that, in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprise", "include" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device.
[0052] In the description of the present invention, it should also be noted that unless otherwise clearly specified and limited, the terms "arrange" and "connect" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0053] The embodiment of the present application discloses a text image quality detection system. The system structure diagram is as Figure 1 shown, which includes a text image quality detection model and a model evaluation module. The text image quality detection model includes a first feature extraction module, a position encoding module, a second feature extraction module, and a fully connected layer. The first feature extraction module is used to extract features from the input text image and output a first feature image; the position encoding module is used to generate position encoding information for the first feature image output by the first feature extraction module to form a first feature image with position information; the second feature extraction module is used to extract features from the first feature image with position information and output a second feature image; the fully connected layer is used to splice the first feature image and the second feature image and map and output a detection result. The model evaluation module is used to evaluate and optimize the detection accuracy of the text image quality detection model in combination with the subjective score of the text image.
[0054] In one embodiment, the first feature extraction module is a convolutional neural network, and the second feature extraction module is an encoder based on a transformer. The convolutional neural network includes four convolutional blocks 1, 2, 3, and 4, and performs four downsampling operations on the input text image to obtain four downsampling results:
[0055] x 1 ∈ B 1 × C 1 × H 1 × W 1
[0056] x 2 ∈ B 1 × C 2 × H 2 × W 2
[0057] x 3 ∈ B 1 × C 3 × H 3 × W 3
[0058] x 4 ∈ B 1 × C 4 × H 4 × W 4
[0059] Among them, B i represents the batch size of the current batch i, C i represents the number of channels of the output of the i-th layer, H i and W i represent the image output by the i-th layer.
[0060] The convolutional neural network extracts multi-scale local features from the input image through four convolutional blocks. Specifically, a normalization layer, a pooling layer, and dropout processing are also set after the convolutional block. The pooling layer performs pooling on x 1 、x 2 and x 3 to align their tensors to
[0061] x 1 、x 2 、x 3 ∈ B 4 × C 4 × H 4 × W 4
[0062] The features after dropout processing are added by tensor addition in a concatenated manner to obtain
[0063] x 5 ∈ B 4 × C 4 × H 4 × W 4
[0064] and input it into a Transformer-based encoder.
[0065] The position encoding module generates position encoding information based on the output x 5 to form a first feature image with position information. By setting the position encoding module, the permutation-invariant property of the Transformer is processed. The Transformer-based encoder performs attention weighting on the first feature image with position information and outputs a second feature image:
[0066]
[0067] where B i represents the batch size of the current batch i, k i represents the weight corresponding to the output of the i-th layer, C i represents the number of channels of the output of the i-th layer, W 4 and H 4 represent the image output by the fourth layer.
[0068] The second feature image is concatenated and added to the x 4 tensor through a pooling layer and a fully connected layer, and then input into the fully connected layer to map and obtain the text image quality detection result.
[0069] For text image quality detection, local features at different positions often have correlations, such as the relationship between edge and texture quality and text. Therefore, long-distance capture is beneficial for providing global context information, thereby improving the accuracy of text image quality detection. In addition, Transformer can adaptively assign weights to each feature element through self-attention weights, enabling the model to pay more attention to important quality features, such as text font blur, alignment, and image noise.
[0070] The present invention introduces a hybrid framework composed of CNN and Transformer to model the global and local information of text image quality detection. Considering that it is difficult to obtain high-quality text reference images, a no-reference text image quality detection model is proposed. CNN is good at extracting local features of images, while Transformer can effectively capture long-distance dependencies between global images through the self-attention mechanism. Combining the two can better fuse local and global information and improve the understanding of image content to obtain better IQA performance.
[0071] Based on the above text image quality detection system, an embodiment of the present application also discloses a text image quality detection method, such as Figure 2 , including:
[0072] S1. Obtain a text image and divide the text image into a training set and a test set.
[0073] S2. Build a text image quality detection model and input the training set into the text image quality detection model for training.
[0074] In one embodiment, step S2 includes:
[0075] The text image is input into the first feature extraction module to obtain a first feature image.
[0076] Based on the first feature image, position encoding information is generated.
[0077] The position coding information is combined with the first feature image to obtain the first feature image with position information.
[0078] The first feature image with position information is input into the second feature extraction module to obtain the second feature image.
[0079] The first feature image and the second feature image are concatenated through a fully connected layer to map and output the image quality detection result.
[0080] Specifically, the first feature extraction module is a convolutional neural network, and the second feature extraction module is a transformer-based encoder.
[0081] In one embodiment, step S2 further includes:
[0082] The text image quality detection model is trained by minimizing the regression loss function:
[0083]
[0084] Where n represents the number of encoder layers, q i is the detection result of the i-th image, s i is the subjective score of the i-th image, w 1 and w 2 are the weights corresponding to the detection results and subjective scores respectively. By using their corresponding weights, the detection results of the text image quality detection model are made more inclined to the subjective scores, and a more robust detection result is obtained.
[0085] S3. Input the test set into the trained text image quality detection model to obtain the detection results, and evaluate and optimize the text image quality detection model. The evaluation indicators of the text image quality detection model include optimized linear correlation OLCC (Optimized Linear Correlation Coefficient) and optimized rank correlation OSROCC (Optimized Spearman Rank-Order Correlation Coefficient).
[0086] The OLCC describes the linear correlation between the subjective score and the objective detection result. By measuring the metric of the linear correlation between the subjective score (scoring the text image quality by human eyes) and the objective detection result (the image quality score given by the text image quality detection model), the best linear transformation is found to make the objective detection result closer to the subjective score.
[0087] In one embodiment, the evaluation of the text image quality detection model includes:
[0088] Obtain the subjective score of the text image, combine the subjective score with the output detection result, and evaluate the optimized linear correlation OLCC between the subjective score and the detection result:
[0089]
[0090] where N represents the number of batch images, y i and respectively represent the subjective score of the i-th image and the output detection result, which are obtained through the subjective score of the actual human eye annotation data and the text image quality detection model detecting a single training picture respectively; and respectively represent the average value of the subjective scores and the average value of the detection results, m i is the weight corresponding to the i-th image, which is used to adjust the weight ratio of the improved OLCC.
[0091] Use OSROCC to measure the monotonicity of the text image quality detection model, calculate the rank correlation between the subjective score and the objective detection result, focus on the consistency in sorting between the two, and do not consider the absolute value of the score difference. The Spearman rank correlation coefficient is a non-parametric statistical method, which calculates the relationship between the subjective evaluation and the objective evaluation based on the ranking to adjust the objective score to achieve the best consistency in ranking with the subjective score.
[0092] In one embodiment, the evaluation of the text image quality detection model further includes:
[0093] Obtain the subjective score of the text image and the sorting between the subjective scores.
[0094] Obtain the output detection result and the sorting between the detection results.
[0095] Based on the subjective score sorting and the detection result sorting, evaluate the rank correlation OSROCC between the subjective score and the detection result:
[0096]
[0097] where v i and pi respectively represent y i and the sorting positions in the subjective scoring sequence and the detection result sequence, w 3 is the corresponding weight, used to adjust the amplitude of the overall monotonicity. When the training set images enter the text image quality detection model, pre-training segmentation is performed, so that the training set images are split on the same batch, v i represents the subjective scoring part, that is, the corresponding part sorting with the detection result part; p i represents the sequential position after the text image quality detection model detects and sorts the split parts of the training set images. Through OSROCC, it is possible to effectively perform partition dynamic attention allocation and evaluation on pictures for different local weightings.
[0098] The training process of the text image quality detection model is as Figure 3 , obtain text images, and divide the text images into a training set and a test set. Randomly select 100 blocks of size 512×512 pixels from each training image in the training set, and at the same time perform model preprocessing and configuration: use the Adam optimizer, and the weight decay is 5×10 -3 , train the text image quality detection model for at most 10 epochs, and the batch size is 48. The learning rate of the text image quality detection model is initially set to 3×10 -5 , and the learning rate is reduced by 10 times after each epoch.
[0099] Specifically, use ResNet50 as the CNN backbone network and initialize the weights with Imagenet. The number of encoder layers used in the Transformer is 2.
[0100] The output feature F corresponding to the encoder based on the transformer T is:
[0101]
[0102] Among them, B represents the batch size, k i represents the weight corresponding to the output of the i-th layer, C i represents the number of channels of the output of the i-th layer, W 4 and H 4 represent the image of the output of the 4th layer. Specifically, the value of k i can be expressed as 0.6 + 0.1×i.
[0103] For each image batch, the text image quality detection model is trained by minimizing the regression loss:
[0104]
[0105] Among them, n represents the number of encoder layers, and q i is the detection result of the i-th image, and s i is the subjective score of the i-th image, and w 1 and w 2 are the weights corresponding to the detection result and the subjective score respectively. Using their corresponding weights, specifically, w 1 is 0.9, and w 2 is 0.7.
[0106] In the test phase, 100 blocks of 512×512 pixels are randomly sampled from the test images in the test set and input into the trained text image quality detection model, and average pooling is performed on their detection result scores to obtain the final quality score.
[0107] Evaluate the optimized linear correlation OLCC between the subjective score and the detection result:
[0108]
[0109] Among them, N represents the number of batch images, y i and represent the subjective score and the output detection result of the i-th image respectively, and represent the average value of the subjective score and the average value of the detection result respectively, and m i is the weight corresponding to the i-th image. Specifically, m i is
[0110]
[0111] Use OSROCC to measure the monotonicity of the text image quality detection model. The calculation formula is:
[0112]
[0113] Among them, v i represents the sorting position of y i in the subjective score sequence, and p i represents in the detection result sequence, and w 3 is the corresponding weight used to adjust the amplitude of the overall monotonicity. Specifically, w 3 is 0.2.
[0114] This application uses a text image quality detection model without a reference image to detect the quality of text images. It does not require a fixed reference image, is more suitable for actual application scenarios, is more flexible and practical, improves the detection accuracy of text image quality evaluation, as well as the accuracy and robustness of text image quality detection, and realizes accurate and real-time text image quality evaluation.
[0115] Each module in the above text image quality detection system can be implemented in whole or in part by software, hardware, and their combination. Each of the above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each of the above modules.
[0116] In one embodiment, a computer device is provided. The computer device can be a server or a terminal integrated with a scheduler, and its internal structure diagram can be as Figure 4 shown. The computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. When the computer program is executed by the processor, it realizes a text image quality detection method.
[0117] In one embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the following steps are realized:
[0118] Obtain a text image, and divide the text image into a training set and a test set;
[0119] Build a text image quality detection model, and input the training set into the text image quality detection model for training;
[0120] Input the test set into the trained text image quality detection model to obtain a detection result, and evaluate and optimize the text image quality detection model.
[0121] In one embodiment, when the processor executes the computer program, the following steps are also realized:
[0122] Obtain the subjective score of the text image, combine the subjective score with the output detection result, and evaluate the optimized linear correlation coefficient OLCC between the subjective score and the detection result:
[0123]
[0124] Where N represents the number of batch images, yi and respectively represent the subjective score of the i-th image and the output detection result, which are obtained by the subjective score of the actual human eye annotation data and the text image quality detection model detecting a single training picture respectively; and respectively represent the average subjective score and the average detection result, and m i is the weight corresponding to the i-th image, which is used to adjust the weight ratio of the improved OLCC.
[0125] In one embodiment, when the processor executes the computer program, the following steps are further implemented:
[0126] Obtain the subjective scores of the text images and the ranking among the subjective scores;
[0127] Obtain the output detection results and the ranking among the detection results;
[0128] Based on the subjective score ranking and the detection result ranking, evaluate the rank correlation OSROCC between the subjective score and the detection result:
[0129]
[0130] wherein, v i and p i respectively represent the ranking positions of y i and in the subjective score sequence and the detection result sequence, and w 3 is the corresponding weight, which is used to adjust the amplitude of the overall monotonicity.
[0131] In one embodiment, when the processor executes the computer program, the following steps are further implemented:
[0132] Input the text image into the first feature extraction module to obtain the first feature image;
[0133] Input the first feature image into the second feature extraction module to obtain the second feature image;
[0134] Concatenate the first feature image and the second feature image through a fully connected layer, and map and output the text image quality detection result.
[0135] In one embodiment, when the processor executes the computer program, the following steps are further implemented:
[0136] Generate position encoding information based on the first feature image;
[0137] Combine the position encoding information with the first feature image to obtain the first feature image with position information, and input it into the second feature extraction module.
[0138] In one embodiment, when the processor executes the computer program, the following steps are further implemented:
[0139] Train the text image quality detection model using a minimized regression loss function:
[0140]
[0141] where n represents the number of encoder layers, q i is the detection result of the i-th image, s i is the subjective score of the i-th image, w 1 and w 2 are the weights corresponding to the detection result and the subjective score respectively.
[0142] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0143] Obtain a text image, and divide the text image into a training set and a test set;
[0144] Construct a text image quality detection model, and input the training set into the text image quality detection model for training;
[0145] Input the test set into the trained text image quality detection model to obtain a detection result, and evaluate and optimize the text image quality detection model.
[0146] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:
[0147] Obtain the subjective score of the text image, combine the subjective score with the output detection result, and evaluate the optimized linear correlation coefficient OLCC between the subjective score and the detection result:
[0148]
[0149] where N represents the number of batch images, y i and respectively represent the subjective score of the i-th image and the output detection result, which are obtained through the subjective score of the actual human eye annotation data and the detection of a single training picture by the text image quality detection model respectively; and respectively represent the average value of the subjective score and the average value of the detection result, m i is the weight corresponding to the i-th image, which is used to adjust the weight ratio of the improved OLCC.
[0150] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:
[0151] Obtain the subjective scores of the text images and the ranking among the subjective scores;
[0152] Obtain the output detection results and the ranking among the detection results;
[0153] Based on the subjective score ranking and the detection result ranking, evaluate the rank correlation OSROCC between the subjective score and the detection result:
[0154]
[0155] where v i and p i respectively represent the ranking positions of y i and in the subjective score sequence and the detection result sequence, and w 3 is the corresponding weight, used to adjust the amplitude of the overall monotonicity.
[0156] In one embodiment, when the computer program is executed by the processor, the following steps are further implemented:
[0157] Input the text image into the first feature extraction module to obtain the first feature image;
[0158] Input the first feature image into the second feature extraction module to obtain the second feature image;
[0159] Concatenate the first feature image and the second feature image through a fully connected layer, and map and output the image quality detection result.
[0160] In one embodiment, when the computer program is executed by the processor, the following steps are further implemented:
[0161] Generate position encoding information based on the first feature image;
[0162] Combine the position encoding information with the first feature image to obtain the first feature image with position information, and input it into the second feature extraction module.
[0163] In one embodiment, when the computer program is executed by the processor, the following steps are further implemented:
[0164] Train the text image quality detection model by minimizing the regression loss function:
[0165]
[0166] where n represents the number of encoder layers, q i is the detection result of the i-th image, s i is the subjective score of the i-th image, w 1 and w 2 are the weights corresponding to the detection result and the subjective score respectively.
[0167] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memories can include read-only memory (ROM), magnetic tapes, floppy disks, flash memories, optical memories, high-density embedded non-volatile memories, resistive random access memories (ReRAM), magnetoresistive random access memories (MRAM), ferroelectric random access memories (FRAM), phase change memories (PCM), graphene memories, etc. Volatile memories can include random access memory (RAM) or external cache memories, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logics, data processing logics based on quantum computing, etc., without limitation.
[0168] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0169] As described above, it is only a preferred embodiment of the present invention and does not impose any form of limitation on the present invention. Therefore, any modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention still fall within the scope of the technical solution of the present invention.
Claims
1. A text image quality detection method, characterized in that: include: Obtain text images and divide the text images into training sets and test sets; Build a text image quality detection model, and input the training set into the text image quality detection model for training; The test set is input into the trained text image quality detection model to obtain the detection results, and the text image quality detection model is evaluated and optimized.
2. The text image quality detection method according to claim 1, characterized in that: The evaluation of the text image quality detection model includes: Obtain the subjective score of the text image, combine the subjective score with the output detection result, and evaluate the optimized linear correlation OLCC between the subjective score and the detection result: Where N is the number of batch images, y i and They represent the subjective score of the i-th image and the output test result, which are obtained by the subjective score of the actual human eye annotation data and the text image quality detection model testing a single training image respectively; and Represent the average subjective score and the average test result, m i is the weight corresponding to the i-th image, which is used to adjust the weight ratio of the improved OLCC.
3. The text image quality detection method according to claim 1, characterized in that: The evaluating of the text image quality detection model further includes: Obtaining subjective scores of text images and rankings between subjective scores; Obtain the output test results and the order of the test results; Based on the subjective score ranking and the test result ranking, the rank correlation OSROCC between the subjective score and the test result is evaluated: Among them, v i and p i Respectively represent y i and In the ranking position in the subjective score sequence and the test result sequence, w3 is the corresponding weight, which is used to adjust the amplitude of the overall monotonicity.
4. The text image quality detection method according to claim 1, characterized in that: The step of constructing a text image quality detection model and inputting a training set into the text image quality detection model for training includes: Inputting the text image into a first feature extraction module to obtain a first feature image; Inputting the first feature image into a second feature extraction module to obtain a second feature image; The first feature image and the second feature image are concatenated through a fully connected layer to output the image quality detection result.
5. The text image quality detection method according to claim 4, characterized in that: The step of constructing a text image quality detection model and inputting a training set into the text image quality detection model for training further includes: Based on the first feature image, generating position coding information; The position coding information is combined with the first feature image to obtain the first feature image with position information, and the first feature image is input into the second feature extraction module.
6. The text image quality detection method according to claim 1, characterized in that: The step of constructing a text image quality detection model and inputting a training set into the text image quality detection model for training further includes: The text image quality detection model is trained by minimizing the regression loss function: Where n represents the number of encoder layers, q i is the detection result of the i-th image, s i is the subjective score of the i-th image, w1 and w2 are the weights corresponding to the detection results and subjective scores, respectively.
7. The text image quality detection method according to claim 1, characterized in that: The first feature extraction module is a convolutional neural network, and the second feature extraction module is a transformer-based encoder.
8. A text image quality detection system, used to implement the text image quality detection method according to any one of claims 1 to 7, characterized in that: It includes a text image quality detection model and a model evaluation module, and the text image quality detection model includes: A first feature extraction module, used for extracting features from an input text image and outputting a first feature image; A position coding module, used for generating position coding information for the first feature image output by the first feature extraction module, so as to form a first feature image with position information; A second feature extraction module, used for extracting features from the first feature image with position information and outputting a second feature image; A fully connected layer, used to concatenate the first feature image and the second feature image, and map and output the detection result; The model evaluation module is used to evaluate and optimize the detection accuracy of the text image quality detection model in combination with the subjective score of the text image.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Document image quality evaluation method
CN111583259A
Full-reference high-dynamic image quality evaluation method based on multi-feature fusion
CN111768362A
Screen content image quality evaluation method
CN113610862A
Blind image quality evaluation method and device based on attention and multiple tasks, equipment and medium
CN118014962A
AIGC image quality evaluation method based on text image encoder
CN118537711A