A method for preventing see-through in intelligent correction, a storage medium and an apparatus
Patent Information
- Application Number
- CN202611093669.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-22
- Publication Date
- 2026-09-25
AI Technical Summary
传统图像处理技术,如简单的去噪或模糊处理,难以区分并有效消除这种因光学透射而产生的特异性干扰
1.本发明实现了对透印的智能、高精度识别与自动化去除。通过获取正反两面扫描图像并协同分析,利用一个先进的深度学习网络,能够精准区分并自动消除掉因背面文字透印造成的干扰字迹,显著提升了后续识别和评分的准确性。
Smart Images

Figure QLYQS_2 
Figure QLYQS_3
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent education technology, specifically to a method, storage medium, and device for suppressing overprinting in intelligent grading. Background Technology
[0002] In the digital transformation of paper-based assignments in education, intelligent grading is a crucial step. However, in practical applications, when scanning double-sided printed assignments, the thinness of the paper often leads to "transparency," where text or images printed or written on the back of the paper show through and interfere with the image on the front. This interference significantly impacts the accuracy of subsequent optical character recognition (OCR) and semantic-based answer analysis and automatic scoring. Traditional image processing techniques, such as simple denoising or blurring, struggle to distinguish and effectively eliminate this specific interference caused by optical transmission. Manual correction methods are not only costly and inefficient but also unsuitable for large-scale assignment grading scenarios. Therefore, in the preprocessing stage of intelligent grading, accurately and automatically eliminating the interference of transparency on the front-side answer content becomes a key technical bottleneck for improving the overall system usability and accuracy. Summary of the Invention
[0003] To address the shortcomings of existing technologies, this invention aims to provide a method, storage medium, and device for suppressing overprinting in intelligent correction.
[0004] To achieve the above objectives, the present invention adopts the following technical solution: A method for suppressing print leakage in intelligent correction includes the following steps: S1. Obtain images of both sides of the same paper worksheet to construct a double-sided image pair; S2. Perform geometric registration and pixel-level alignment on the two sides of the double-sided image to obtain the spatially corresponding transparent printing reference relationship; S3. The image of the side to be corrected is taken as the front homework image, and the scanned image of the other side is taken as the back homework image; based on the pre-built individual student handwriting feature model, the student-specific writing pattern is extracted from the front homework image. It should be noted that in this embodiment, the front side represents the target side that needs to be corrected and protected from bleed-through, while the other side is the back side, serving as an auxiliary side to provide a reference for bleed-through interference. When both sides of the same paper worksheet need to be corrected, the bleed-through suppression process is performed on each side as the front side before correction.
[0005] S4. Using a pre-constructed bleed separation neural network, the back-side work image is used as a supervision signal, and the student-specific writing pattern extracted in step S3 is used as a priori reference to separate and remove the interference components caused by bleed from the front-side work image, and output the front-side work image after bleed removal. S5. Perform text recognition and semantic analysis on the front-side image after de-bleeding to complete the correction.
[0006] Furthermore, in step S3, each student has their own unique individual handwriting feature model. The specific construction process of the individual student handwriting feature model is as follows: A1. Data Acquisition and Preprocessing: Collect scanned images and / or handwritten trajectory data of students' past N assignments; denoise and binarize the collected scanned images and then crop them into single-character or stroke images; smooth and normalize the collected handwritten trajectory data. A2. Network Model Construction: Build a network model based on the type of input data; if the input data is a single character or stroke image, construct a convolutional neural network to extract spatial visual features; if the input data is handwritten trajectory data, construct a Transformer encoder to capture temporal and dynamic features. A3. Model Training and Consolidation: Input the preprocessed image or handwritten trajectory data from step A1 into the network model built in step A2. Iteratively train the network model through comparative learning or classification tasks so that the network model can fully learn and remember the student's unique writing pattern, including stroke width distribution, writing angle, frequency of connecting strokes, ink density gradient and pen acceleration, etc., to obtain the student's individual handwriting feature model. A4. Incremental Updates and Optimization: We continuously collect new homework data and teacher review and correction results from the student, and make incremental fine-tuning to the student's individual handwriting feature model to continuously adapt to the natural changes in the student's writing habits.
[0007] Further, in step S4, the through-print separation neural network is a generative adversarial network (GAN), which includes a generator and a discriminator. The generator receives the front-side job image and the back-side job image as input and outputs the front-side job image after de-through-printing. During the training phase of the through-print separation neural network, the discriminator simultaneously receives the original front-side job image and the front-side job image after de-through-printing, and combines the through-printing prior information provided by the back-side job image to determine the authenticity of the generated front-side job image after de-through-printing.
[0008] Furthermore, the construction process of the transparent separation neural network is as follows: B1. Collect basic data: acquire a large number of clean front-side operation images without bleed-through and the corresponding back-side operation images, and perform geometric registration between the back-side operation images and the corresponding clean front-side operation images; B2. Calculate the transmission intensity: Using the optical attenuation physical model, based on the grayscale value of the back image, the equivalent transmission distance mapping, and the correction coefficients of ambient light and scanner gain, calculate the transmission intensity distribution of the text on the back image transmitted to the front. The formula for the physical model of optical attenuation is: I t(x,y) =α[I b(x,y) exp( β D(x,y))]; Among them, I t(x,y) This represents the estimated light penetration value at coordinates (x, y) of the front-facing image; I b(x,y) α is the gray value at the corresponding coordinates of the back image after geometric registration; D(x,y) is the equivalent transmission distance mapping preset according to the paper type or estimated by the local contrast of the image; α is the correction coefficient for ambient light and scanner gain; β is the inherent attenuation coefficient of the paper. B3. Synthesize the simulated image: Overlay the bleed intensity distribution calculated in step B2 onto a clean front-facing work image with reasonable transparency to synthesize a front-facing work image with realistic bleed interference. B4. Constructing the dataset: Pack the clean front-side work images collected in step B1, the corresponding back-side work images, and the front-side work images with realistic see-through interference generated in step B3 into sample pairs, which ultimately constitute the dataset for end-to-end training of the see-through separation neural network. B5. Input the dataset constructed in step B4 into the transparent printing separation neural network for training; The data processing procedure of the generator in the pass-through separation neural network is as follows: B51. Dual-path feature extraction: The generator simultaneously receives front and back images of the completed assignment with bleed-through. It extracts the mixed content features of the front image and the bleed-through prior features of the back image through convolutional layers. The mixed content features are extracted from the front image, including real handwriting and answer details, as well as bleed-through interference components. Real handwriting and answer details include the student's actual writing, while bleed-through interference components include bleed-through interference characters. The bleed-through prior features are extracted from the back image, mainly including the original handwriting and spatial information on the back, as well as bleed-through intensity and guidance information. The original handwriting and spatial information on the back includes the actual writing content of the back image and its corresponding spatial position, providing a reference for the source of the bleed-through. The bleed-through intensity and guidance information includes the characteristics of the bleed-through intensity distribution. B52. Transparency Prediction and Separation: Using the prior features of transparency provided by the back-side work image as a guiding signal, combined with the transparency perception mechanism, and introducing the student's unique writing pattern identified in step S3 as a prior reference, by comparing the physical differences between the student's actual writing content and the transparency interference handwriting, the transparency interference components caused by back-side transparency in the front-side work image are accurately predicted. B53. Preservation of Authentic Handwriting: Remove or suppress predicted bleed interference components from the mixed content features of the front-facing assignment image. In this process, the student's exclusive writing mode is used to constrain and verify the preserved handwriting features to ensure that the student's authentic handwriting is not lost while denoising, and that the details of the original answer are not lost. B54. Image Reconstruction and Output: The denoised clean features obtained in step B53 are used to reconstruct the image through deconvolution or upsampling layers, and finally a clear, transparent, and interference-free frontal working image is generated and output.
[0009] Furthermore, during the training phase, the discriminator and generator of the transillumination separation neural network engage in adversarial gameplay to optimize the network parameters; the specific process by which the discriminator determines the authenticity of the generated image is as follows: (C1) Multi-source information input and fusion: The discriminator stitches or fuses features of the front working image without bleed-through, the original front working image with bleed-through, and the back working image output by the generator in the channel dimension. (C2) Feature extraction and conditional constraint comparison: The features fused in step (C1) are extracted through the convolutional layer; the discriminator combines the back prior information to check whether there are still traces of back traces or light ink marks in the front work image after removing the back trace; at the same time, the original front work image is compared to confirm whether the student's real handwriting and edge details are completely preserved and not distorted. (C3) Authenticity Probability Output: Based on the comparison results of step (C2), the discriminator outputs a probability score to determine whether the front-side image of the de-bleeding operation comes from a real clean operation image in the dataset or is an image "forged" by the generator. (C4) Adversarial feedback optimization: The discriminator backpropagates the error of the judgment result obtained in step (C3), updates the discriminator's own parameters to improve the discrimination power, guides the generator to adjust its strategy, and forces it to generate an image with more thorough bleed removal and more perfect preservation of real handwriting until the two reach a balance.
[0010] Furthermore, step S4 also includes a content repair step: the front-facing work image after piercing is depierced, output by the piercing separation neural network, is detected. If content is missing or blurred in the key answer area, the front-facing work image after piercing is depierced is subjected to semantic consistency-guided content repair: the domain-adapted language model is called, and the most likely answer fragment is predicted by combining the knowledge context of the question type corresponding to the key answer area to be repaired; then the prediction result is embedded into the corresponding position of the front-facing work image after piercing is depierced, output by the piercing separation neural network, to form semantically coherent and subject-logical complete text, which is used for the scoring decision in step S5.
[0011] Furthermore, in the content restoration process, the process of locating the key answer areas is as follows: D1. Based on the optical attenuation physical model, calculate the tracing intensity value of each pixel position in the front and back working images after geometric registration and pixel-level alignment, and generate a tracing intensity heat map. D2. Based on the heat map of bleed intensity generated in step D1, an adaptive threshold segmentation algorithm is used to locate the key answer areas with severe bleed interference. The adaptive threshold T... adapt The calculation formula is: T adapt = μ + k * σ; Where μ is the global mean of the heatmap of bleed intensity, σ is the standard deviation of the heatmap of bleed intensity, and k is the sensitivity coefficient set according to the average light transmittance of the paper. Will I t(x,y) >T adapt The areas identified as key answer areas with high transparency interference were only subject to semantic consistency-guided content repair in these key answer areas.
[0012] Furthermore, the sensitivity coefficient k is directly calculated based on the average light transmittance of the paper using a linear mapping formula: k = kmin + (kmax) kmin)×τ τ represents the average light transmittance of the paper; Kmin represents the lower limit of the sensitivity coefficient; Kmax represents the upper limit of the sensitivity coefficient.
[0013] The present invention also provides a computer-readable storage medium, characterized in that the computer-readable storage medium stores a computer program, which, when executed by a processor, implements the above-described method.
[0014] The present invention also provides a computer device, including a processor and a memory, wherein the memory is used to store a computer program; and the processor is used to execute the computer program to implement the above-described method.
[0015] The beneficial effects of this invention are as follows: 1. This invention achieves intelligent, high-precision recognition and automated removal of bleed-through printing. By acquiring and collaboratively analyzing scanned images from both sides, and utilizing an advanced deep learning network, it can accurately distinguish and automatically eliminate interfering text caused by bleed-through printing on the back, significantly improving the accuracy of subsequent recognition and scoring.
[0016] 2. This invention possesses targeted intelligent content repair capabilities based on prior judgment. Instead of applying a "one-size-fits-all" approach to all image areas, this invention first accurately locates the critical answer areas with the most severe bleed-through, and then, combining subject knowledge and the contextual logic of the question, intelligently reasoning and supplementing missing or ambiguous answer content, ensuring that the repaired result is both clear and reasonable. Detailed Implementation
[0017] The present invention will be further described below. It should be noted that this embodiment is based on the present technical solution and provides detailed implementation methods and specific operation processes, but the protection scope of the present invention is not limited to this embodiment.
[0018] Example 1 This embodiment provides a method for suppressing print overprint in intelligent correction, including the following steps: S1. Obtain images of both sides of the same paper worksheet to construct a double-sided image pair; S2. Perform geometric registration and pixel-level alignment on the two sides of the double-sided image to obtain the spatially corresponding transparent printing reference relationship; S3. The image of the side to be corrected is taken as the front homework image, and the scanned image of the other side is taken as the back homework image; based on the pre-built individual student handwriting feature model, the student-specific writing pattern is extracted from the front homework image. It should be noted that in this embodiment, the front side represents the target side that needs to be corrected and protected from bleed-through, while the other side is the back side, serving as an auxiliary side to provide a reference for bleed-through interference. When both sides of the same paper worksheet need to be corrected, the bleed-through suppression process is performed on each side as the front side before correction.
[0019] S4. Using a pre-constructed bleed separation neural network, the back-side work image is used as a supervision signal, and the student-specific writing pattern extracted in step S3 is used as a priori reference to separate and remove the interference components caused by bleed from the front-side work image, and output the front-side work image after bleed removal. S5. Perform text recognition and semantic analysis on the front-side image after de-bleeding to complete the correction.
[0020] In this embodiment, in step S3, each student has their own unique individual handwriting feature model. The specific construction process of the individual handwriting feature model is as follows: A1. Data Acquisition and Preprocessing: Collect scanned images and / or handwritten trajectory data from students' past N assignments (collected while students are answering on the handwriting tablet or touchscreen). Denoise and binarize the collected scanned images, then crop them into single-character or stroke images; smooth and normalize the collected handwritten trajectory data.
[0021] A2. Network Model Construction: Build a network model based on the type of input data; if the input data is a single character or stroke image, construct a convolutional neural network (CNN) to extract spatial visual features; if the input data is handwritten trajectory data, construct a Transformer encoder to capture temporal and dynamic features.
[0022] A3. Model Training and Consolidation: Input the preprocessed images or handwritten trajectory data from step A1 into the network model built in step A2. Iteratively train the model through comparative learning or classification tasks (e.g., bring the network model closer to the handwriting features of the same student at different times and push away the handwriting features of different students). This allows the network model to fully learn and remember the student's unique writing pattern, including stroke width distribution, writing angle, frequency of connecting strokes, ink density gradient, and pen acceleration, thus obtaining the student's individual handwriting feature model.
[0023] A4. Incremental Updates and Optimization: In subsequent practical applications, new homework data and teacher review and correction results of the student will be continuously collected to incrementally fine-tune the student's individual handwriting feature model, so that it can continuously adapt to the natural changes in the student's writing habits.
[0024] In this embodiment, in step S4, the through-print separation neural network is a generative adversarial network (GAN), which includes a generator and a discriminator. The generator receives the front and back images of the work as input and outputs the front image of the work after the through-print is removed. During the training phase of the through-print separation neural network, the discriminator simultaneously receives the original front image of the work and the front image of the work after the through-print is removed, and combines the through-print prior information provided by the back image of the work to determine the authenticity of the generated front image of the work after the through-print is removed.
[0025] In this embodiment, the construction process of the transparent separation neural network is as follows: B1. Collect basic data: Obtain a large number of clean front-facing images without bleed-through (as the ground truth for subsequent training) and the corresponding back-facing images, and perform geometric registration between the back-facing images and the corresponding clean front-facing images.
[0026] It should be noted that a clean, non-bleed-through front-side work image refers to an original front-side work image that has no bleed-through (or has very little bleed-through that can be ignored) (e.g., thicker paper for paper-based work).
[0027] B2. Calculate the transmission intensity: Using the optical attenuation physical model, based on the grayscale value of the back image, the equivalent transmission distance mapping, and the correction coefficients for ambient light and scanner gain, calculate the transmission intensity distribution of the text on the back image transmitted to the front.
[0028] The formula for the physical model of optical attenuation is: I t(x,y) =α[I b(x,y) exp( β D(x,y))]; Among them, I t(x,y) This represents the estimated light penetration value at coordinates (x, y) of the front-facing image; I b(x,y) α is the gray value at the corresponding coordinates of the back image after geometric registration; D(x,y) is the equivalent transmission distance mapping preset according to the paper type or estimated by the local contrast of the image; α is the correction coefficient for ambient light and scanner gain; β is the inherent attenuation coefficient of the paper. Equivalent transmission distance mapping refers to a two-dimensional distribution map of the effective physical thickness or optical transmission path length that light travels from the back to the front of a piece of paper at each pixel location. Due to uneven fiber distribution, wrinkles, or local indentations caused by pen pressure during writing, the actual light transmission thickness of the paper is not absolutely uniform throughout, thus it is a mapping matrix that varies with spatial position (x, y). Thick / high-density paper (such as cardstock, high-grammage printing paper) has a larger overall equivalent transmission distance mapping value, with greater light attenuation and weaker light-through. Thin / low-density paper (such as thin draft paper, newsprint) has a smaller overall equivalent transmission distance mapping value, with less light attenuation and more severe light-through. Different paper materials have different fluctuations in local thickness; paper with a rough texture or uneven fibers has a greater local difference (variance) in its equivalent transmission distance mapping value.
[0029] In this embodiment, the equivalent transmission distance mapping can be preset based on the paper type or estimated through local image contrast: (1) Based on paper type preset (static acquisition): When the workbook or paper specifications (such as standard 70g or 80g A4 paper) are known, the average physical thickness or standard thickness distribution template of each type of paper is directly pre-calibrated, and the average physical thickness or standard thickness distribution template of the corresponding type is used as the value of D(x,y).
[0030] (2) Image Local Contrast Estimation (Dynamic Back-Calculation): For unknown specific paper types, dynamic calculation is performed by analyzing the visual characteristics of the scanned image itself. By utilizing the local contrast, gray-level gradient, or shadow features of the image, the paper indentation area (where the paper is compressed and the transmission distance is reduced) or paper wrinkle area caused by pen tip pressure is identified. Then, combined with a pre-calibrated optical scattering inversion algorithm, the changes in the local contrast, gray-level gradient, or shadow features of the image are converted into relative thickness change values, thereby estimating and generating the equivalent transmission distance mapping map of the entire sheet of paper.
[0031] In this embodiment, the correction coefficients for ambient light and scanner gain are obtained in the following ways: 1) Direct reading of hardware parameters (device preset): Through the scanner driver or built-in sensor, the current scanning exposure time, hardware gain settings and ambient light intensity are directly read. These parameters are then substituted into the calibration formula or mapping table preset at the factory to directly calculate the correction coefficients for ambient light and scanner gain.
[0032] 2) Standard White Board Calibration (Static Calibration): Before scanning the paper worksheet, scan a blank calibration board with known standard reflectance (such as a standard white board or gray board). Compare the grayscale value of the scanned image with the theoretical grayscale value of the standard board, and calculate the comprehensive compensation ratio of ambient light and hardware gain according to α = (G / G_ref) × [(L_s+E) / L_s]. Use this as the correction coefficient α; G represents the actual hardware gain value of the current scanner, G_ref represents the preset standard (reference) hardware gain value, L_s is the standard luminous intensity of the scanner's internal light source, and E represents the measured intensity of the current external ambient light.
[0033] Hardware gain compensation ratio (G / G_ref): The scanner's hardware gain directly and linearly amplifies the image signal (including bleed-through interference). The ratio of the current gain to the standard gain is the hardware compensation ratio. Ambient light compensation ratio (L_s+E) / L_s: Ambient light (E) is superimposed on the scanner's built-in light source (L_s), increasing the total amount of light penetrating the paper, thus making bleed-through more obvious. This ratio reflects the additional light transmission increase brought by ambient light. Multiplying the two yields the comprehensive compensation ratio, used to accurately correct the estimated value of bleed-through intensity, ensuring the robustness of the model under different scanning devices and lighting conditions.
[0034] 3) Dynamic estimation based on blank areas in the image (dynamic back-calculation): In the actual scanned image, pure blank background areas without text (pure paper areas without text or bleed-through) are automatically identified and extracted. The average gray value of these blank areas is calculated, and the impact of the current ambient light and scanner gain is dynamically back-calculated to obtain the comprehensive correction coefficient α for the current ambient light and scanner gain. The calculation formula is as follows: α = I_bg / I_ref I_bg represents the actual average grayscale value of the pure blank background area in the currently scanned image. I_ref represents the preset standard background grayscale reference value under standard ambient light and standard gain.
[0035] During the scanning and imaging process, the grayscale value of the blank area in the image is linearly positively correlated with the ambient light intensity and hardware gain. Therefore, without relying on external physical sensors, by simply extracting the actual grayscale value (I_bg) of the pure blank background area of the currently scanned image and calculating its ratio to the standard grayscale value (I_ref), the overall impact of the current ambient light and gain on imaging can be directly and comprehensively quantified, thereby enabling real-time dynamic calculation and updating of the correction coefficient.
[0036] It should be noted that the inherent attenuation coefficient of the paper refers to the paper's ability to absorb and scatter light, reflecting its intrinsic physical property of blocking light penetration. It is mainly determined by the paper's fiber composition, density, thickness, and filler. A larger attenuation coefficient indicates that the paper is more light-blocking, and the less the text on the back is transmitted to the front; conversely, a smaller attenuation coefficient indicates that the paper is more translucent, and the more severe the light transmission phenomenon. In this embodiment, the inherent attenuation coefficient of the paper can be obtained in the following three ways: (a) Offline optical experiment calibration (instrument measurement): Use a transmittance tester or spectrophotometer to measure the change in light intensity before and after light passes through the paper sample. Calculate the attenuation coefficient of the paper accurately using the Beer-Lambert Law in optics. Based on this, establish a mapping database of each type of paper material (such as common 70g or 80g copy paper) and attenuation coefficient.
[0037] (ii) Dynamic estimation based on image features (online back-calculation): When scanning images of both sides of the same worksheet, the system automatically finds a reference area in the images of both sides where the front is blank and the back has dark text. By comparing the original grayscale of the text on the back with the actual grayscale after it is transmitted to the front, and combining the paper thickness data, the system uses an optical attenuation physical model to solve for the actual attenuation coefficient of the currently scanned paper, thereby achieving dynamic adaptation.
[0038] B3. Synthesize the simulated image: Overlay the bleed intensity distribution calculated in step B2 onto a clean front-facing work image with reasonable transparency to synthesize a front-facing work image with realistic bleed interference.
[0039] It should be noted that in step B3, using the collected back-side work images from step B2, the bleed intensity distribution caused by the back text being projected onto the front is simulated and calculated using an optical physical model. This calculated bleed intensity distribution can be superimposed onto a clean front-side work image to artificially synthesize a simulated image with realistic bleed interference. This allows the construction of a training dataset, where the synthesized front-side work image with bleed interference serves as input, and the collected clean front-side work images serve as the ground truth, thereby training the bleed separation neural network to learn how to remove bleed.
[0040] More specifically, in this embodiment, the superimposed transparency is not determined using a single fixed value, but rather dynamically through a combination of base intensity mapping and severity adjustment, while simultaneously satisfying real optical physical constraints. The specific process for determining transparency is as follows: (I) Basic transparency mapping: First, the transparency intensity I of each pixel calculated in step B2 is... t The values (x, y) are normalized and mapped to the interval [0, 1], serving as the base transparency for each pixel. Areas with higher calculated transparency have greater base transparency, making the superimposed text on the back more visible.
[0041] (II) Introduction of a degree adjustment coefficient (data augmentation): In order for the trained bleed separation neural network to cope with different degrees of bleed in reality (such as slight, moderate and severe bleed), a global transparency adjustment coefficient λ (usually randomly selected between 0.2 and 0.8) is introduced to control the overall bleed depth.
[0042] (III) Final transparency calculation: The final overlay transparency A(x,y) of each pixel is determined as: A(x,y) = λ × Normalize(I t (x,y)), where Normalize represents the normalization function.
[0043] To ensure a reasonable and realistic overlay effect, the following constraints can be added during the overlay process: Conforms to optical blending modes: Instead of simply adding pixel values, it uses image blending modes such as Multiply that conform to the physical laws of light passing through paper.
[0044] Protect genuine handwriting on the front: In areas where there is already genuine student handwriting (dark pixels) in the front-side assignment image, automatically reduce or mask the transparency of the overlay of genuine student handwriting to prevent synthetic overlay from damaging or obscuring the details of genuine handwriting on the front.
[0045] Prevent grayscale overflow: Ensure that the pixel values of the superimposed and synthesized image are truncated and limited to the effective image range (such as 0-255) to avoid distorted pure black or pure white dead zones.
[0046] B4. Constructing the dataset: Pack the clean front-side work images collected in step B1, the corresponding back-side work images, and the front-side work images with realistic see-through interference generated in step B3 into sample pairs, which ultimately constitute the dataset for end-to-end training of the see-through separation neural network. B5. Input the dataset constructed in step B4 into the transparent separation neural network for training.
[0047] In this embodiment, the data processing procedure of the generator of the transparent separation neural network is as follows: B51, Dual-path feature extraction: The generator simultaneously receives front-side and back-side operation images with bleed-through, and extracts the mixed content features of the front-side operation image and the bleed-through prior features of the back-side operation image through convolutional layers.
[0048] The mixed content features were extracted from the front-facing homework image. These features included genuine handwriting and answer details, as well as spectral interference. The genuine handwriting and answer details comprised the student's actual writing, while the spectral interference consisted of spectral-interference characters, which were characterized by lighter ink and blurred edges. It should be noted that the extracted genuine handwriting, answer details, and spectral interference were mixed together; the spectral interference components within the mixed content features were not predicted.
[0049] The prior features of tracing are extracted from the back of the work image, mainly including the original handwriting and spatial information on the back, as well as the tracing intensity and guidance information. The original handwriting and spatial information on the back includes the actual written content of the back work image and its corresponding spatial position, providing a source reference for tracing. The tracing intensity and guidance information includes the characteristics of the tracing intensity distribution.
[0050] B52. Transparency Prediction and Separation: Using the prior features of transparency provided by the back-side work image as a guiding signal, combined with a transparency perception mechanism (such as an attention mechanism), and introducing the student's unique writing pattern identified in step S3 as a prior reference, by comparing the physical differences between the student's actual writing content and the transparency interference handwriting, the transparency interference components caused by back-side transparency in the front-side work image are accurately predicted.
[0051] B53. Preservation of Authentic Handwriting: Eliminate or suppress predicted bleed interference components from the mixed content features of the front-facing assignment image. In this process, use the student's exclusive writing mode to constrain and verify the preserved handwriting features to ensure that the student's authentic handwriting is not lost while denoising, and that the details of the original answer are not lost.
[0052] B54. Image Reconstruction and Output: The denoised clean features obtained in step B53 are used to reconstruct the image through deconvolution or upsampling layers, and finally a clear, transparent, and interference-free frontal working image is generated and output.
[0053] During the training phase, the discriminator and generator of the bleed separation neural network engage in adversarial competition to optimize network parameters. In actual grading applications, the trained generator is simply used to process and output the debleed-off front-facing image. Based on the image inpainting principle of Generative Adversarial Networks (GANs), in this embodiment, the specific process by which the discriminator determines the authenticity of the generated image is as follows: (C1) Multi-source information input and fusion: The discriminator splices or fuses the front working image (the target to be discriminated) output by the generator (without bleed-through), the original front working image (with bleed-through) (providing the original context reference), and the back working image (providing bleed-through prior conditions) in the channel dimension.
[0054] (C2) Feature Extraction and Conditional Constraint Comparison: The features fused in step (C1) are extracted through the convolutional layer. The discriminator, combined with the back prior information, checks whether there are still traces of back bleed texture or light ink marks in the front work image after bleed-through; at the same time, it compares with the original front work image to confirm whether the student's real handwriting and edge details are completely preserved and not distorted.
[0055] (C3) True / False Probability Output: Based on the comparison results of step (C2), the discriminator outputs a probability score to determine whether the front-side image of the de-bleeding operation comes from a real clean operation image in the dataset (true) or an image "forged" by the generator (false).
[0056] (C4) Adversarial Feedback Optimization: The discriminator backpropagates the error of the judgment result obtained in step (C3). This is not only used to update the discriminator's own parameters to improve its discrimination power, but more importantly, it guides the generator to adjust its strategy, forcing it to generate images with more thorough bleed removal and more perfect preservation of real handwriting, until the two reach a balance (model convergence).
[0057] It should be noted that the Transparent Imprint Separation Neural Network (TI-GAN) uses a dataset with simulated transparent imprints for end-to-end optimization during the training phase, and can complete denoising without labels during the inference phase.
[0058] In this embodiment, in step S5, during the text recognition stage, combining the student's personal writing habits and handwriting characteristics can help the OCR engine more accurately adapt to the student's unique cursive or personalized handwriting, thereby improving the accuracy of answer extraction and automated scoring.
[0059] In this embodiment, step S4 further includes a content repair step: the front-facing work image after de-bleeding output by the bleed separation neural network is detected. If content is missing or blurred in the key answer area, the front-facing work image after de-bleeding is repaired with semantic consistency guidance: the domain-adapted language model is called, and the knowledge context of the question type corresponding to the key answer area to be repaired (such as mathematical formula rules, physical laws, and Chinese grammar structure) is combined to predict the most likely answer fragment; then the prediction result is embedded into the corresponding position of the front-facing work image after de-bleeding output by the bleed separation neural network to form a semantically coherent and subject-logical complete text for grading in step S5.
[0060] Furthermore, in this embodiment, the process of locating the key answer area in the content repair step is as follows: D1. Based on the optical attenuation physical model, calculate the tracing intensity value of each pixel position in the front and back working images after geometric registration and pixel-level alignment, and generate a tracing intensity heat map. D2. Based on the heat map of bleed intensity generated in step D1, an adaptive threshold segmentation algorithm is used to locate the key answer areas with severe bleed interference. The adaptive threshold T... adapt The calculation formula is: T adapt = μ + k * σ; Where μ is the global mean of the heatmap of bleed intensity, σ is the standard deviation of the heatmap of bleed intensity, and k is the sensitivity coefficient set according to the average light transmittance of the paper. Will I t(x,y) >T adapt The areas identified as key answer areas with high transparency interference were only subject to semantic consistency-guided content repair in these key answer areas.
[0061] In standard deviation-based adaptive thresholding algorithms, the sensitivity coefficient k typically ranges from 0.5 to 3.0. The average light transmittance of the paper is positively correlated with the sensitivity coefficient k. When the average light transmittance is high (thin paper, easily translucent), overall light transmission is prevalent and background noise is significant. To avoid misjudging minor light transmission as "serious interference," the sensitivity coefficient k needs to be increased (increasing the threshold, decreasing the sensitivity), thus accurately filtering out only extremely severe abnormal areas of light transmission. When the average light transmittance is low (thick paper, good light blocking), overall light transmission is weaker. In this case, the k value needs to be decreased (decreasing the threshold, increasing the sensitivity), so that even minor light transmission interference can be detected more sensitively.
[0062] Specifically, the sensitivity coefficient k can be directly calculated using a linear mapping formula based on the average light transmittance of the paper: k = kmin + (kmax) kmin)×τ τ represents the average light transmittance of the paper (normalized to between 0 and 1, where 0 represents complete opacity and 1 represents complete light transmittance). Kmin represents the lower limit of the sensitivity coefficient (e.g., set to 0.5 for thick paper with high sensitivity). Kmax represents the upper limit of the sensitivity coefficient (e.g., set to 2.5 for thin paper with low sensitivity).
[0063] For example, if we set kmin = 0.5 and kmax = 2.5, when using thinner paper (average light transmittance τ = 0.8), we calculate k = 0.5 + (2.5) / 2.5. 0.5)×0.8=2.1; When using thicker paper (average light transmittance τ=0.2), the calculated value is k=0.5+2.0×0.2=0.9.
[0064] Furthermore, in this embodiment, in step S5, the answer recognition confidence of each identified question is calculated. The answer recognition confidence is calculated by comprehensively considering three indicators: the heat map of transparency intensity, the character integrity score, and the semantic matching probability. When the answer recognition confidence is lower than a preset threshold, the question is marked as "requires manual review" and pushed to the teacher review interface.
[0065] Specifically, the formula for calculating the confidence score of answer recognition is as follows: C=w1×(1 I t )+w2×Schar+w3×Psem I t This represents the average spectral intensity of the answer area on the front-facing image (normalized to 0-1). The more severe the spectral interference, the higher the intensity of the spectral interference. (1) The lower the value of (It), the better. Schar represents the character integrity score (between 0 and 1). Psem represents the semantic matching probability (between 0 and 1). w1, w2, and w3 are the weight coefficients of the three indicators, and w1 + w2 + w3 = 1, which can be dynamically adjusted according to the specific subject and question type.
[0066] In this embodiment, the character integrity score (Schar) is mainly used to evaluate the visual clarity and stroke completeness of the handwriting in the answer area, and is calculated synchronously by the OCR recognition engine when extracting text from the answer area using any of the following methods: (E1) Visual morphology analysis: The OCR recognition engine performs connected component analysis and edge detection on the recognized character images, and counts whether there are breaks, adhesions, missing parts or severe blurring in the strokes. The more complete the morphology, the higher the character integrity score.
[0067] (E2) Confidence Mapping: Utilizing the character-level Softmax probability distribution output by the underlying deep learning model of OCR (such as CRNN), the maximum probability value of the predicted character is taken as the character integrity score, or the information entropy of the probability distribution is calculated as the character integrity score. The more confident the model is in extracting the glyph features, the higher the character integrity score.
[0068] In this embodiment, the semantic matching probability (Psem) is mainly used to evaluate the reasonableness of the identified text in the subject logic and question context, and there are two main calculation methods: ① Contextual Language Model Evaluation: The identified answer text sequence is input into a pre-trained domain language model (such as a subject-specific model based on Transformer). Combined with the question stem and known context, the conditional probability of the answer text sequence is calculated as the semantic matching probability.
[0069] ② Semantic similarity calculation: The identified answer text sequence is converted into a vector representation, compared with the reasonable answer nodes in the standard answer database or subject knowledge graph, and the cosine similarity is calculated as the semantic matching probability.
[0070] The more logically sound and consistent with common sense in the subject matter, the higher the probability of semantic matching.
[0071] For those skilled in the art, various corresponding changes and modifications can be made based on the above technical solutions and concepts, and all such changes and modifications should be included within the protection scope of the claims of this invention.
Claims
1. A method for suppressing print leakage in intelligent correction, characterized in that, Includes the following steps: S1. Obtain images of both sides of the same paper worksheet to construct a double-sided image pair; S2. Perform geometric registration and pixel-level alignment on the two sides of the double-sided image to obtain the spatially corresponding transparent printing reference relationship; S3. The image of the side to be corrected is taken as the front homework image, and the scanned image of the other side is taken as the back homework image; based on the pre-built individual student handwriting feature model, the student-specific writing pattern is extracted from the front homework image. S4. Using a pre-constructed bleed separation neural network, the back-side work image is used as a supervision signal, and the student-specific writing pattern extracted in step S3 is used as a priori reference to separate and remove the interference components caused by bleed from the front-side work image, and output the front-side work image after bleed removal. S5. Perform text recognition and semantic analysis on the front-side image after de-bleeding to complete the correction.
2. The method according to claim 1, characterized in that, In step S3, each student has their own unique handwriting feature model. The specific construction process of the student's individual handwriting feature model is as follows: A1. Data Acquisition and Preprocessing: Collect scanned images and / or handwritten trajectory data of students' past N assignments; after denoising and binarizing the collected scanned images, crop them into single-character or stroke images; The collected handwritten trajectory data was smoothed and normalized. A2. Network Model Construction: Build a network model based on the data type of the input; if the input data is a single character or stroke image, construct a convolutional neural network to extract spatial visual features; If the input data is handwritten trajectory data, a Transformer encoder is constructed to capture temporal and dynamic features; A3. Model Training and Consolidation: Input the preprocessed image or handwritten trajectory data from step A1 into the network model built in step A2. Iteratively train the network model through comparative learning or classification tasks so that the network model can fully learn and remember the student's unique writing pattern, including stroke width distribution, writing angle, frequency of connecting strokes, ink density gradient and pen acceleration, etc., to obtain the student's individual handwriting feature model. A4. Incremental Updates and Optimization: We continuously collect new homework data and teacher review and correction results from the student, and make incremental fine-tuning to the student's individual handwriting feature model to continuously adapt to the natural changes in the student's writing habits.
3. The method according to claim 1, characterized in that, In step S4, the bleed separation neural network is a generative adversarial network (GAN), which includes a generator and a discriminator. The generator receives the front and back images of the work as input and outputs the front image of the work after bleed separation. During the training phase of the bleed separation neural network, the discriminator simultaneously receives the original front-facing work image and the debleed-out front-facing work image, and combines the bleed prior information provided by the back-facing work image to determine the authenticity of the generated debleed-out front-facing work image.
4. The method according to claim 3, characterized in that, The construction process of the transparent separation neural network is as follows: B1. Collect basic data: acquire a large number of clean front-side operation images without bleed-through and the corresponding back-side operation images, and perform geometric registration between the back-side operation images and the corresponding clean front-side operation images; B2. Calculate the transmission intensity: Using the optical attenuation physical model, based on the grayscale value of the back image, the equivalent transmission distance mapping, and the correction coefficients of ambient light and scanner gain, calculate the transmission intensity distribution of the text on the back image transmitted to the front. The formula for the physical model of optical attenuation is: I t(x,y) =α[I b(x,y) exp( β D(x,y))]; Among them, I t(x,y) This represents the estimated light penetration value at coordinates (x, y) of the front-facing image; I b(x,y) α is the gray value at the corresponding coordinates of the back image after geometric registration; D(x,y) is the equivalent transmission distance mapping preset according to the paper type or estimated by the local contrast of the image; α is the correction coefficient for ambient light and scanner gain; β is the inherent attenuation coefficient of the paper. B3. Synthesize the simulated image: Overlay the bleed intensity distribution calculated in step B2 onto a clean front-facing work image with reasonable transparency to synthesize a front-facing work image with realistic bleed interference. B4. Constructing the dataset: Pack the clean front-side work images collected in step B1, the corresponding back-side work images, and the front-side work images with realistic see-through interference generated in step B3 into sample pairs, which ultimately constitute the dataset for end-to-end training of the see-through separation neural network. B5. Input the dataset constructed in step B4 into the transparent printing separation neural network for training; The data processing procedure of the generator in the pass-through separation neural network is as follows: B51. Dual-path feature extraction: The generator simultaneously receives front and back images of the completed assignment with bleed-through. It extracts the mixed content features of the front image and the bleed-through prior features of the back image through convolutional layers. The mixed content features are extracted from the front image, including real handwriting and answer details, as well as bleed-through interference components. Real handwriting and answer details include the student's actual writing, while bleed-through interference components include bleed-through interference characters. The bleed-through prior features are extracted from the back image, mainly including the original handwriting and spatial information on the back, as well as bleed-through intensity and guidance information. The original handwriting and spatial information on the back includes the actual writing content of the back image and its corresponding spatial position, providing a reference for the source of the bleed-through. The bleed-through intensity and guidance information includes the characteristics of the bleed-through intensity distribution. B52. Transparency Prediction and Separation: Using the prior features of transparency provided by the back-side work image as a guiding signal, combined with the transparency perception mechanism, and introducing the student's unique writing pattern identified in step S3 as a prior reference, by comparing the physical differences between the student's actual writing content and the transparency interference handwriting, the transparency interference components caused by back-side transparency in the front-side work image are accurately predicted. B53. Preservation of Authentic Handwriting: Remove or suppress predicted bleed interference components from the mixed content features of the front-facing assignment image. In this process, the student's exclusive writing mode is used to constrain and verify the preserved handwriting features to ensure that the student's authentic handwriting is not lost while denoising, and that the details of the original answer are not lost. B54. Image Reconstruction and Output: The denoised clean features obtained in step B53 are used to reconstruct the image through deconvolution or upsampling layers, and finally a clear, transparent, and interference-free frontal working image is generated and output.
5. The method according to claim 4, characterized in that, During the training phase, the discriminator and generator of the pass-through separation neural network engage in adversarial gameplay to optimize the network parameters; the specific process by which the discriminator determines the authenticity of the generated image is as follows: (C1) Multi-source information input and fusion: The discriminator stitches or fuses features of the front working image without bleed-through, the original front working image with bleed-through, and the back working image output by the generator in the channel dimension. (C2) Feature extraction and conditional constraint comparison: The features fused in step (C1) are extracted through the convolutional layer; the discriminator combines the back prior information to check whether there are still traces of back traces or light ink marks in the front work image after removing the back trace; at the same time, the original front work image is compared to confirm whether the student's real handwriting and edge details are completely preserved and not distorted. (C3) Authenticity Probability Output: Based on the comparison results of step (C2), the discriminator outputs a probability score to determine whether the front-side image of the de-bleeding operation comes from a real clean operation image in the dataset or is an image "forged" by the generator. (C4) Adversarial feedback optimization: The discriminator backpropagates the error of the judgment result obtained in step (C3), updates the discriminator's own parameters to improve the discrimination power, guides the generator to adjust its strategy, and forces it to generate an image with more thorough bleed removal and more perfect preservation of real handwriting until the two reach a balance.
6. The method according to claim 1, characterized in that, Step S4 also includes a content repair step: the front-facing work image after piercing is depierced, output by the piercing separation neural network, is detected. If content is missing or blurred in the key answer area, the front-facing work image after piercing is depierced is subjected to semantic consistency-guided content repair: the domain-adapted language model is called, and the most likely answer fragment is predicted by combining the knowledge context of the question type corresponding to the key answer area to be repaired; then the prediction result is embedded into the corresponding position of the front-facing work image after piercing is depierced, output by the piercing separation neural network, to form semantically coherent and subject-logical complete text, which is used for the scoring decision in step S5.
7. The method according to claim 6, characterized in that, In the content restoration process, the process of locating the key answer areas is as follows: D1. Based on the optical attenuation physical model, calculate the tracing intensity value of each pixel position in the front and back working images after geometric registration and pixel-level alignment, and generate a tracing intensity heat map. D2. Based on the heat map of bleed intensity generated in step D1, an adaptive threshold segmentation algorithm is used to locate the key answer areas with severe bleed interference. The adaptive threshold T... adapt The calculation formula is: T adapt = μ + k * σ; Where μ is the global mean of the heatmap of bleed intensity, σ is the standard deviation of the heatmap of bleed intensity, and k is the sensitivity coefficient set according to the average light transmittance of the paper. Will I t(x,y) >T adapt The areas identified as key answer areas with high transparency interference were only subject to semantic consistency-guided content repair in these key answer areas.
8. The method according to claim 7, characterized in that, The sensitivity coefficient k can be directly calculated based on the average light transmittance of the paper using the linear mapping formula: k=kmin+(kmax kmin)×τ τ represents the average light transmittance of the paper; Kmin represents the lower limit of the sensitivity coefficient; Kmax represents the upper limit of the sensitivity coefficient.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method described in any one of claims 1-8.
10. A computer device, characterized in that, It includes a processor and a memory, the memory being used to store a computer program; the processor being used to execute the computer program to implement the method of any one of claims 1-8.