Image correction and paper recognition method and system
Patent Information
- Application Number
- CN202610153151.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-02-03
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2046-02-03
AI Technical Summary
功能单一:多数方案侧重于几何校正,缺乏对纸张边界的识别,而教育类图像处理应用需完成校正+纸张识别的功能
本发明实施例还提供了一种图像校正与纸张识别方法及系统,通过获得至少两帧拍摄时间相邻的原始图像,对原始图像进行预处理,得到预处理图像;其中原始图像是反射成像图像;初始校正模块,用于对预处理图像进行初始校正,获得第一校正图像;联动校正模块,用于联动第一校正图像的空间信息和时间信息,自适应对第一校正图像进行二次校正,获得第二校正图像;图像恢复模块,用于基于第二校正图像对原始图像进行透视恢复,获得恢复纸张图像。
Smart Images

Figure CN121999498B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer vision technology, and more specifically, to an image correction and paper recognition method and system. Background Technology
[0002] With the widespread adoption of smart learning devices such as learning machines and student tablets, functions like "photo-based question search," "homework correction," "finger-based learning," and "finger-based reading" have become common applications. To allow users to take photos without moving their devices, some learning machines (electronic devices used to assist learning) are equipped with front-facing cameras. These cameras use a reflector attached to the camera to capture images of books and papers on a table. However, due to the unavoidable tilt between the reflector and the table, the camera introduces perspective distortion, positional shifts, and localized stretching during the shooting process, affecting the accuracy of subsequent homework correction and text recognition. Furthermore, for easier post-processing, it's necessary to accurately identify the area of the paper and remove unnecessary background information from the image.
[0003] To address this challenge, the following technical solutions are the main approaches: 1. Rear-facing shooting solution: The user holds up the device and directly uses the rear camera of the learning machine to take pictures to avoid distortion. Although this solution produces high image quality, the learning device is relatively large and heavy. Frequently holding it up is not only cumbersome but also puts extra strain on the student, causing arm pain and physical fatigue. It is difficult to apply to long-term learning scenarios requiring continuous shooting; furthermore, this method does not provide image semantic information about the paper boundaries.
[0004] 2. Traditional image geometric transformation schemes: These schemes use techniques such as perspective transformation, solving for graphic transformation matrices, or edge detection to restore the original shape of the image plane. However, these methods are highly dependent on stable, clear, and regular paper boundaries. In complex real-world learning environments, arbitrary paper placement, uneven lighting, and interference from clutter on the desktop can make these techniques extremely unreliable, thus affecting the algorithm's robustness and making it difficult to generalize.
[0005] 3. 3D reconstruction technology based on binocular cameras: The image transformation matrix is obtained by estimating the camera pose and tilt angle using the binocular parallax of a stereo camera. Its advantages include good real-time performance and accurate correction. However, its disadvantages are equally significant: high hardware costs, which not only hinder product adoption but also increase the burden on users. Furthermore, it can only reconstruct the entire image and cannot identify the semantic information of image boundaries, resulting in the camera's inability to effectively distinguish between paper and background or identify the left and right boundaries of the paper.
[0006] 4. Transformers-based 3D reconstruction model: This method utilizes a deep learning model with a massive number of parameters, directly inputting distorted images and corrected images during training to learn a correction mapping. Its advantage is its ability to handle a certain degree of page quality issues, including but not limited to creases, bends, and blurriness. However, its disadvantages make it difficult to implement on edge devices: firstly, the model's massive parameter count and the high computational demands of inference make it difficult to meet the requirements of small size, precision, and real-time performance needed by edge devices; secondly, the specialized data required often necessitates significant effort in data preparation.
[0007] In summary, the existing technology has the following shortcomings: Limited functionality: Most solutions focus on geometric correction and lack the ability to recognize paper boundaries, while educational image processing applications need to perform both correction and paper recognition.
[0008] Poor environmental robustness: Traditional geometric methods are easily affected by environmental factors such as blurred boundaries and changes in lighting, making it difficult to work reliably in complex and realistic scenarios.
[0009] Cost and functional limitations: Binocular camera solutions have high hardware costs and cannot complete paper recognition tasks. Large model solutions cannot take advantage of edge computing power, resulting in high computational costs and very high costs for the required orientation data.
[0010] The root cause of these shortcomings lies in the fact that the front-facing imaging of a learning machine is a complex imaging process involving refraction through a mirror, resulting in distortions that combine perspective, reflection, and nonlinearity. Traditional methods rely on ideal assumptions that are difficult to guarantee in real-world scenarios; while advanced depth vision solutions are either limited by hardware costs or constrained by enormous computational overhead. This makes it difficult to implement existing solutions on terminals (learning machines), which is a pressing technical problem that needs to be solved. Summary of the Invention
[0011] The purpose of this invention is to provide an image correction and paper recognition method and system to solve the aforementioned problems existing in the prior art.
[0012] In a first aspect, embodiments of the present invention provide an image correction and paper recognition method, comprising: Step S1: Obtain at least two original images with adjacent shooting times, preprocess the original images to obtain preprocessed images; wherein the original images are reflection imaging images; Step S2: Perform initial correction on the preprocessed image to obtain the first corrected image; Step S3: Link the spatial and temporal information of the first corrected image, and adaptively perform secondary correction on the first corrected image to obtain the second corrected image; Step S4: Perform perspective restoration on the original image based on the second corrected image to obtain the restored paper image.
[0013] Optionally, step S2 includes: The preprocessed image is subjected to binary segmentation to obtain a segmented image; the segmented image includes multiple first segmentation masks. Perform connected component analysis on multiple first segmentation masks to determine the first segmentation mask with the largest area as the main body region of the paper. Morphological opening operations are performed on the main area of the paper to remove isolated noise, and closing operations are performed on the main area of the paper to fill gaps, making the boundary of the main area of the paper smoother and more continuous, thereby obtaining a morphologically optimized paper area. Edge correction and contour fitting are performed on the morphologically optimized paper region. Specifically, contour extraction is performed in the morphologically optimized paper region to obtain the edge of the first paper region; edge smoothing is performed on the edge of the first paper region to remove local sharp areas to obtain the edge of the second paper region; the image with the second paper edge region marked is used as the first corrected image.
[0014] Optionally, at least two original images are processed through step S2 above to obtain at least two first corrected images; step S3 includes: The latest original image from at least two original images is used as the current frame; the original image preceding the current frame is used as the previous frame. Center point offset monitoring calculation: Obtain the center point offset between the edge of the second paper region in the current frame image and the edge of the second paper region in the previous frame image; If the offset is less than the set value, it is assumed that the paper has not actually moved, and the edge detection result of the previous frame is used as the current edge detection result. Region overlap estimation: Obtain the intersection-union ratio (CIU) between the edge of the second paper region in the current frame image and the edge of the second paper region in the previous frame image; evaluate whether the paper shape is consistent between the two frames based on the CIU; if the paper shape is consistent between the two frames, determine that the two frames are consecutive observation images; Historical result reuse and state memory: Under the condition that the two frames are consecutively observed images, the edge detection result of the previous frame is used as the current edge detection result; the current edge detection result includes the currently detected paper mask area; Extract the contour feature points from the current edge detection results. Optionally, extracting the contour feature points of the current edge detection result includes: Mask candidate region selection and geometric smoothing: Morphological processing is performed on the current edge detection results to obtain a smooth boundary image; contour extraction is performed in the smooth boundary image to obtain the main boundary of the paper; convex hull is calculated based on the main boundary, and the influence of local depressions and noise is reduced through geometric convergence to obtain a stable main boundary; Extract feature points from the stable principal boundary to obtain contour feature points.
[0015] Optionally, the adaptive correction of the contour feature points includes: Fill the region formed by the contour feature points into a binary region to obtain the initial contour binary region; Obtain the difference between the initial contour binary region and the currently detected paper mask region; if the difference is not empty, identify the pixels in the difference as uncovered pixels; If the proportion of the difference set to the area of the current detected paper mask region exceeds the tolerance threshold, then the initial contour binary region is locally corrected based on the edge normal offset algorithm, specifically: Offset: Each edge of the initial contour binary region is translated along its normal direction by a preset step size to obtain the translated edge; Correction: Obtain the intersection of adjacent translation edges to update the vertices of the initial contour binary region, and obtain the corrected contour binary region; Repeat the above translation-correction process until the area ratio between the corrected contour binary region and the currently detected paper mask region is lower than the threshold or the preset number of iterations is reached. Stop the adaptive correction and obtain the final contour feature points.
[0016] Optionally, the step of performing binary segmentation on the preprocessed image to obtain a segmented image is implemented using a lightweight model. Specifically, the preprocessed image is input into the lightweight model, and the lightweight model performs binary segmentation on the preprocessed image.
[0017] Optionally, the lightweight model includes an adaptive processing module, an encoder, a decoder, and an edge-weighted loss optimization module; The input to the adaptive processing module is the preprocessed image in step S1, the input to the encoder is the output of the adaptive processing module, the input to the decoder is the output of the encoder, the input to the edge-weighted loss optimization module is the output of the decoder, and the decoder outputs a segmented image, which includes one or more first segmentation masks. The adaptive processing module is used to adaptively process the preprocessed image to obtain the fused image, specifically including: The preprocessed image is converted to grayscale to obtain a grayscale image; Set a block kernel with a size of 8×8 pixels, and use the block kernel to perform non-overlapping block division on the grayscale image; Calculate the average gray value of each block. and global mean ; Adaptive contrast enhancement processing is performed on the grayscale image to obtain the enhanced image, specifically as follows: in, For adaptive contrast enhancement factor, Represents pixels in a grayscale image pixel values, Represents pixels in an enhanced image The pixel values, i and j represent the x and y coordinates of the pixel, respectively; Indicates to The values are then truncated / trimmed. Specifically, if... ,but ;like ,but ;like ,but ; Indicates taking and The maximum of the two; The enhanced image is fused using bilinear interpolation to obtain the fused image. The encoder is used to perform deep convolutional spatial feature extraction and ordinary point convolutional cross-channel fusion on the fused image to obtain the encoded image; The decoder is used to perform nearest neighbor interpolation and edge feature compensation on the encoded image to obtain a binarized segmented image. The edge-weighted loss optimization module is used to calculate the loss value of the segmented image to train a lightweight model.
[0018] Optionally, the loss value calculation method for the edge-weighted loss optimization module is as follows: in, This represents the loss value of the edge-weighted loss optimization module. Indicates edge weights, , The Sobel edge strength values of the true mask are used to enhance the edge loss. ; These are balancing parameters. ; This represents the cross-entropy loss value between the segmented image and the ground truth image, which is provided beforehand. , Indicates the degree of regional overlap coefficient. , , This represents the edge weight of the k-th pixel in the segmented image. , This represents the edge intensity value of the k-th pixel. , This represents the pixel value of the k-th pixel in the segmented image, and N represents the number of pixels in the segmented image. This represents the pixel value of the k-th pixel in the actual label image.
[0019] Optionally, step S4 includes: The projective transformation matrix T is calculated based on the feature points in the second corrected image. The original image is then subjected to perspective correction based on the projective transformation matrix T to obtain the restored paper image.
[0020] Secondly, embodiments of the present invention provide an image correction and paper recognition system, comprising: The preprocessing module is used to obtain at least two frames of raw images that are captured at adjacent times, and to preprocess the raw images to obtain a preprocessed image; wherein the raw image is a reflection imaging image; The initial correction module is used to perform initial correction on the preprocessed image to obtain the first corrected image; The linkage correction module is used to link the spatial and temporal information of the first correction image and adaptively perform secondary correction on the first correction image to obtain the second correction image. The image restoration module is used to perform perspective restoration on the original image based on the second corrected image to obtain a restored paper image.
[0021] Compared with the prior art, the embodiments of the present invention achieve the following beneficial effects: This invention also provides an image correction and paper recognition method and system, which obtains at least two original images captured at adjacent times, preprocesses the original images to obtain a preprocessed image; wherein the original image is a reflection imaging image; an initial correction module is used to perform initial correction on the preprocessed image to obtain a first corrected image; a linkage correction module is used to link the spatial and temporal information of the first corrected image to adaptively perform secondary correction on the first corrected image to obtain a second corrected image; and an image restoration module is used to perform perspective restoration on the original image based on the second corrected image to obtain a restored paper image.
[0022] The image correction and paper recognition method and system provided in this application address the core pain points of existing intelligent learning device image processing technologies, such as limited functionality, poor environmental robustness, difficulty in terminal adaptation, and high cost. Through multi-stage technological innovation, it achieves multi-dimensional beneficial effects, as detailed below: I. Integrating image correction and paper recognition functions, breaking through the limitations of "single function". Existing technologies often focus on a single step in geometric correction or paper recognition, failing to meet the integrated "correction + recognition" requirements of educational image processing. This application's technical solution employs a full-process design of "preprocessing - initial correction - secondary correction - perspective restoration," achieving both image distortion correction and accurate paper region recognition simultaneously: Step S2 uses binary segmentation and connected component analysis to filter the main paper region, combines morphological opening / closing operations to remove noise and fill gaps, and then uses edge correction and contour fitting to lock the paper boundary; Step S4 calculates the projective transformation matrix based on the corrected feature points, ultimately outputting a restored image with "distortion-free + clearly defined paper boundaries." This design requires no additional algorithms or modules to simultaneously complete two core tasks, perfectly matching the functional requirements of educational scenarios such as "homework correction" and "text recognition" in learning machines, avoiding process redundancy caused by multiple solutions.
[0023] II. Strong environmental robustness, adaptable to complex real-world learning scenarios To address the issues of traditional geometric transformation schemes being susceptible to uneven lighting, haphazard paper placement, and desktop clutter, this invention enhances robustness through a multi-layered technical design: Firstly, the preprocessing stage employs a lightweight adaptive processing module that uses an 8×8 block kernel to statistically analyze the grayscale mean, combined with an adaptive contrast enhancement coefficient (…). First, the dynamic optimization of image contrast effectively counteracts the effects of lighting fluctuations. Second, the morphological operations (opening operations to remove isolated noise and closing operations to fill gaps) and edge smoothing processing in the initial correction stage can filter out desktop clutter and avoid boundary misjudgment caused by sharp local areas. Third, the spatiotemporal linkage mechanism in step S3 reuses historical detection results from consecutive frames through center point offset monitoring and regional overlap estimation. Even if the paper moves slightly or the boundary is locally blurred, it can still maintain stable edge detection accuracy, completely solving the drawback of the traditional method "relying on the ideal boundary assumption" and can be reliably applied to real learning scenarios.
[0024] III. Low power consumption and high efficiency adaptability to terminals, solving the problem of "difficulty in implementation". In existing technologies, binocular camera solutions have high hardware costs, and large-model solutions based on Transformers have high computational overhead and complex data preparation, making them difficult to implement on terminal devices such as learning machines. This invention overcomes this bottleneck through two major innovations: First, it uses a lightweight model to achieve binary segmentation, with a structure (adaptive processing module + lightweight encoder / decoder) that has a small number of parameters, and the edge-weighted loss optimization module (… This approach can improve segmentation accuracy while reducing training and inference costs, without relying on high-performance hardware. Furthermore, the "historical result reuse and state memory" mechanism in step S3 avoids repeated edge detection by judging the offset and intersection-over-union ratio of consecutive frames, significantly reducing computational overhead and meeting the "small and precise" and real-time requirements of terminal devices. In addition, this solution requires no additional hardware (such as a binocular camera), and can be implemented using only the existing front-facing camera and reflector of the learning machine, significantly reducing product costs and user burden, and facilitating the large-scale deployment of the technology on terminal devices.
[0025] IV. Improve calibration accuracy and ensure the effectiveness of subsequent applications. Existing reflective imaging methods suffer from perspective distortion and positional shifts, directly impacting the accuracy of subsequent text recognition and homework grading. This invention minimizes image distortion through precise processing using a "secondary correction + perspective restoration" approach: Step S2's initial correction locks the core area of the paper through contour fitting; Step S3's adaptive secondary correction (local correction based on an edge normal offset algorithm) dynamically adjusts the contour until it matches the paper's mask area; Step S4, perspective restoration based on the projective transformation matrix, accurately restores the original planar shape of the paper, completely eliminating the visual distortion introduced by reflective imaging. The restored paper image processed by this method has clear boundaries and is distortion-free, providing high-quality image input for subsequent text recognition and homework grading, significantly improving the accuracy and reliability of educational applications.
[0026] In summary, this invention comprehensively addresses the shortcomings of existing technologies from four dimensions: functional integration, environmental adaptability, terminal compatibility, and accuracy assurance. It provides an efficient, low-cost, and easily implementable technical solution for image processing in intelligent learning devices, demonstrating significant technical value and application prospects. Attached Figure Description
[0027] Figure 1 This is a flowchart of an image correction and paper recognition method provided in an embodiment of the present invention.
[0028] Figure 2 This is a schematic diagram of a lightweight model structure provided in an embodiment of the present invention.
[0029] Figures 3-6 This is an experimental diagram of an image correction and paper recognition method provided in an embodiment of the present invention. Detailed Implementation
[0030] The present invention will now be described in detail with reference to the accompanying drawings.
[0031] Example Figure 1As shown, this embodiment of the invention provides an image correction and paper recognition method applicable to learning machines that use a front-facing camera and reflector to photograph desktop workbooks. This method can automatically correct and recognize perspective distortion, local stretching, and positional shifts caused by reflective imaging on the learning machine terminal without requiring manual angle adjustment or additional calibration, thereby obtaining standardized document images. In this application, the reflective imaging unit is formed by a front-facing camera and a reflector attached to the front of the camera to facilitate photographing images of desktop paper.
[0032] The overall structural flow of this method includes: 1. Image processing: Since the imaging of the mirror is mirrored and the learning machine has limited computing power, the image needs to be compressed.
[0033] 2. Data Augmentation: Based on the training requirements of deep learning models, perform multi-dimensional data augmentation on images.
[0034] 3. Lightweight deep learning feature point prediction and recognition: Predicts and recognizes the range of paper areas through a trained deep neural network.
[0035] 4. Inter-frame stability processing: In conjunction with the target tracking strategy, the processing results of consecutive frames are correlated within a fixed time window to ensure the continuity and stability of the output in terms of timing.
[0036] 5. Adaptive coordinate extraction: Extract the range of feature points on the paper by calculating the output of the prediction and recognition module.
[0037] 6. Perspective Restoration: Generate a perspective transformation matrix based on the corrected feature points to restore the distorted image to a top-down view of the paper.
[0038] 7. Image Output: Outputs the final corrected standard paper image for subsequent recognition tasks.
[0039] Specifically, this invention provides an image correction and paper recognition method, including the following steps: Step S1: Obtain at least two original images that are captured at adjacent times, and preprocess the original images to obtain preprocessed images.
[0040] The original image is a reflection image. The learning machine's front-facing camera acquires the original image through a mirror. Due to the mirror's mirror-like imaging and the learning machine's limited computing power, normalization preprocessing needs to be performed on the input image to ensure the stability and real-time performance of the subsequent instance segmentation model. Specifically: 1. Horizontally mirror the input image (original image) and scale it to the target size along the long side (this method sets it to 640 pixels, which can be adjusted according to computing power). Fill the short side with black pixels to the target size to ensure consistency between model training and prediction.
[0041] 2. Convert the input image to RGB format and normalize the pixels to stabilize the model's inference results.
[0042] This ensures the stability and real-time performance of the preprocessed image as input data in subsequent lightweight segmentation and correction.
[0043] Step S2: Perform initial correction on the preprocessed image to obtain the first corrected image.
[0044] In this embodiment of the application, step S2 includes: The preprocessed image is subjected to binary segmentation to obtain a segmented image. The segmented image includes multiple first segmentation masks.
[0045] Alternatively, region growing algorithms or Transformer segmentation algorithms can be used to directly segment the features of the preprocessed image.
[0046] Optionally, in this embodiment of the application, the step of performing binary segmentation on the preprocessed image to obtain the segmented image is implemented by a lightweight model. Specifically, the preprocessed image is input into the lightweight model, and the lightweight model performs binary segmentation on the preprocessed image.
[0047] The lightweight model provided in this application includes an adaptive processing module, an encoder, a decoder, and an edge-weighted loss optimization module. The structure of the lightweight model is as follows: Figure 2 As shown. The input to the adaptive processing module is the preprocessed image from step S1, the input to the encoder is the output of the adaptive processing module, the input to the decoder is the output of the encoder, the input to the edge-weighted loss optimization module is the output of the decoder, and the decoder outputs a segmented image, which includes one or more first segmentation masks.
[0048] The adaptive processing module is used to adaptively process the preprocessed image to obtain the fused image, specifically including: The preprocessed image is converted to grayscale to obtain a grayscale image; Set a block kernel with a size of 8×8 pixels, and use the block kernel to perform non-overlapping block division on the grayscale image; Calculate the average gray value of each block. and global mean ; Adaptive contrast enhancement processing is performed on the grayscale image to obtain the enhanced image, specifically as follows: in, For adaptive contrast enhancement factor, Represents pixels in a grayscale image pixel values, Represents pixels in an enhanced image The pixel values, i and j represent the x and y coordinates of the pixel, respectively; Indicates to The values are then truncated / trimmed. Specifically, if... ,but ;like ,but ;like ,but ; Indicates taking and The maximum of the two.
[0049] The above scheme reduces the computational load by 90% by designing block mean statistics to address the binary nature of binary segmentation.
[0050] After obtaining the enhanced image, bilinear interpolation is performed to fuse the enhanced image to obtain a fused image. In this embodiment, the grayscale image is divided into non-overlapping blocks using a block kernel to obtain multiple non-overlapping blocks. These non-overlapping blocks are then stitched together to obtain the enhanced image. Since direct stitching of blocks will result in block boundary effects (abrupt changes in grayscale values between blocks, forming grid-like artifacts), it is necessary to eliminate artifacts. Therefore, bilinear interpolation is performed to fuse the enhanced image. Specifically, the value of each pixel in the stitched image (enhanced image) is recalculated so that its value is obtained by weighted average of the grayscale values of the four nearest pixels (the weight is inversely proportional to the distance between pixels), thus making the grayscale values at the block boundaries "gradual" rather than "abrupt," smoothing out grid artifacts and avoiding boundary effects. After eliminating artifacts, the image needs to be normalized to [0,1], that is, the processed image pixel values are mapped to the 0~1 range to adapt to the input requirements of the subsequent neural network.
[0051] The encoder is used to perform deep convolutional spatial feature extraction and ordinary point convolutional cross-channel fusion on the fused image to obtain the encoded image. Specifically: spatial feature extraction is performed using deep convolution (DW), assigning an independent K×K convolution kernel to each input channel, and performing spatial convolution only on that channel (without crossing channels), with the number of output channels equal to the number of input channels. In this embodiment, the value of K is 3. Then, cross-channel fusion is performed using point convolution (PW), specifically using a 1×1 convolution kernel to fuse the features output by DW across channels, adjusting the number of output channels to the target dimension.
[0052] The above scheme focuses on spatial features (edges and textures) and channel fusion, while PW focuses on channel fusion. After decoupling, it is more suitable for the spatial edge + channel feature requirements of binary segmentation. It retains the advantages of lightweight and cross-channel fusion of ordinary point convolution, and solves the pain point of edge accuracy loss in lightweight models, so that feature extraction is not lost. On this basis, edge adaptive weights are added (in conjunction with subsequent steps), which not only retains the advantages of lightweight, but also solves the problem of edge accuracy loss in lightweight models, and improves the accuracy of feature extraction.
[0053] The decoder performs nearest-neighbor interpolation and edge feature compensation on the encoded image to obtain a binarized segmented image. Specifically, nearest-neighbor interpolation is performed on the encoded image. Nearest-neighbor interpolation has the advantages of low computational cost and minimal redundancy, but it also suffers from block artifacts (abnormal edge pixels). In contrast, the encoder's skip-connection features are high-resolution features that retain all edges (e.g., the encoder's 3-layer output is H / 4×W / 4, which is higher resolution than the decoder's input H / 8×W / 8). Here, H represents the image length (number of pixels), and W represents the image width (number of pixels). Then, edge features are extracted, and the interpolated features are weighted and enhanced to compensate for lost details. The specific compensation method is calculated using the formula shown below: in, The feature (pixel value) is the difference between the nearest neighbors. The edge response value of the encoder skip connection feature. , For compensation coefficient, It can keep the model lightweight and avoid over-enhancement. The compensated features (pixel values).
[0054] In summary, the low-resolution features of the decoder are used, then nearest-neighbor interpolation upsampling is performed, then the encoder skip-connection feature channels are aligned, then the edge response of the skip-connection features is extracted, then the interpolation feature is multiplied by (1+γ×edge response) (compensation), and finally the compensation feature and skip-connection feature are fused to output high-resolution features.
[0055] In this way, the advantages of the nearest neighbor interpolation in achieving extreme lightweighting are retained, while the shortcomings of detail loss are made up for by the high-resolution edge features of the encoder. This is the core of the decoder to achieve lightweight + high-precision edge segmentation.
[0056] The edge-weighted loss optimization module is used to calculate the loss value of the segmented image to train a lightweight model.
[0057] In this embodiment of the application, the loss value of the edge-weighted loss optimization module is calculated as follows: in, This represents the loss value of the edge-weighted loss optimization module. Indicates edge weights, , The Sobel edge strength values of the true mask are used to enhance the edge loss. ; These are balancing parameters. ; This represents the cross-entropy loss value between the segmented image and the ground truth image, which is provided beforehand. , Indicates the degree of regional overlap coefficient. , , This represents the edge weight of the k-th pixel in the segmented image. , This represents the edge intensity value of the k-th pixel. , This represents the pixel value of the k-th pixel in the segmented image, and N represents the number of pixels in the segmented image. This represents the pixel value of the k-th pixel in the actual label image.
[0058] In this application, B is responsible for "pixel-wise probability fitting": ensuring that the probability value predicted by the model is close to the real label and that the training converges stably; D is responsible for "region-level overlap optimization": solving the problem of foreground-background imbalance and enhancing the segmentation accuracy of edges / small targets; edge weights We: acting on both B and D, allowing the model to prioritize the optimization of the core of binary segmentation - the edge region.
[0059] When the loss value of the edge-weighted loss optimization module converges, the lightweight model is considered to be well trained. At this point, the lightweight model outputs a first corrected image containing multiple first segmentation masks.
[0060] The lightweight model described above has the following three characteristics: 1. Extremely low parameter size (≤30MB): Meets the computing power limitations of embedded inference.
[0061] 2. High segmentation accuracy: It can remove edge effects and improve segmentation accuracy.
[0062] 3. Short prediction time: The model's prediction speed can achieve real-time performance in video inference tasks.
[0063] After obtaining multiple first segmentation masks, connected component analysis is performed on the multiple first segmentation masks to determine the first segmentation mask with the largest area as the main body area of the paper.
[0064] Morphological opening operations are performed on the main area of the paper to remove isolated noise, and closing operations are performed on the main area of the paper to fill gaps, making the boundaries of the main area of the paper smoother and more continuous, thereby obtaining a morphologically optimized paper area.
[0065] Edge correction and contour fitting are performed on the morphologically optimized paper region. Specifically, contour extraction is performed in the morphologically optimized paper region to obtain the edge of the first paper region; edge smoothing is performed on the edge of the first paper region to remove local sharp areas to obtain the edge of the second paper region; the image with the second paper edge region marked is used as the first corrected image.
[0066] The above scheme employs a connected component selection process. Connected component analysis is performed on the segmented mask, automatically selecting the largest connected region as the main body of the paper, while eliminating small false detection areas caused by reflections, hands, background textures, etc. Mask smoothing and morphological operations apply morphological opening operations to the selected region to remove isolated noise points, followed by closing operations to fill tiny gaps, making the boundaries smoother and more continuous, thus providing stable results for subsequent contour fitting. Edge correction and contour fitting: The contour structure is re-extracted based on the processed mask, and a smoothing strategy is used to remove local sharp areas, making the paper boundaries more closely match the actual paper geometry and reducing the number of iterations for corner fitting to some extent.
[0067] Step S3: Link the spatial and temporal information of the first corrected image, and adaptively perform secondary correction on the first corrected image to obtain the second corrected image.
[0068] Step S3 includes: The latest original image from at least two original images is used as the current frame image; the previous original image is used as the previous frame image. In this embodiment, at least two original images are processed through step S2 to obtain at least two first corrected images. For example, the current frame image corresponds to the current frame first corrected image, and the previous frame image is processed through step S2 to obtain the previous frame first corrected image. In the previous frame first corrected image, there are multiple second segmentation masks, each corresponding to a multiple first segmentation mask, and each first segmentation mask corresponds to a multiple second paper region edge. The previous frame image also contains one second paper region edge. In the current frame first corrected image, there are multiple second segmentation masks, each corresponding to a multiple first segmentation mask, and each first segmentation mask corresponds to a multiple second paper region edge. The current frame image also contains one second paper region edge.
[0069] Center point offset monitoring calculation: Obtain the center point offset between the edge of the second paper region in the current frame image and the edge of the second paper region in the previous frame image. Optionally, the Euclidean distance between the center points of the second paper region edges in the current frame image and the previous frame image can be used as the offset.
[0070] If the offset is less than a set value, it is considered that the paper has not actually moved, and the edge detection result of the previous frame is used as the current edge detection result. In this embodiment, the set value is a distance of 0-3 pixels. Using the edge detection result of the previous frame as the current edge detection result means that the edge of the second paper region in the previous frame is used as the edge of the second paper region in the current frame, and the second segmentation mask in the previous frame is used as the second segmentation mask in the current frame.
[0071] Region overlap estimation: Obtain the cross-union ratio (CUP) between the edge of the second paper region in the current frame image and the edge of the second paper region in the previous frame image; evaluate whether the paper shape is consistent between the two frames based on the CUP; if the paper shape is consistent between the two frames, determine that the two frames are consecutive observation images.
[0072] In this embodiment of the application, the cross-union ratio of the mask can be used as the cross-union ratio between the second paper area of the current frame image and the second paper area of the previous frame image.
[0073] Historical result reuse and state memory: Under the condition that the two frames are continuously observed images, the edge detection result of the previous frame is used as the current edge detection result; the current edge detection result includes the currently detected paper mask region; the currently detected paper mask region can be the region formed by the edge of the second paper region of the current frame image.
[0074] Extract the contour feature points from the current edge detection results.
[0075] Alternatively, contour feature points can be extracted directly using the simple Canny operator.
[0076] As another optional implementation, the extraction of contour feature points from the current edge detection result includes: Mask candidate region selection and geometric smoothing: Morphological processing (opening / closing operation) is performed on the current edge detection results to obtain a smooth boundary image; contour extraction is performed in the smooth boundary image to obtain the main boundary of the paper; convex hull is calculated based on the main boundary, and the influence of local depressions and noise is reduced through geometric convergence to obtain a stable main boundary.
[0077] Extract feature points from the stable principal boundary to obtain contour feature points. In this step, the Canny operator is used to extract contour feature points again.
[0078] In this step, if the paper is polygonal, a polygon approximation algorithm can be used, employing the golden ratio and a fine-grained search strategy to perform a stepwise search in the parameter space, prioritizing the search for 4-point structures (trapezoidal or quadrilateral) as the initial feature points of the paper. If the geometry is complex, a polygon point set can be generated as the initial sample estimate.
[0079] If the region formed by the contour feature points does not completely cover the current detection paper mask region (there are uncovered pixels), then the contour feature points are adaptively corrected; until the region formed by the contour feature points completely covers the current detection paper mask region, the region formed by the contour feature points is used as the second corrected image.
[0080] In this embodiment of the application, the adaptive correction of contour feature points includes: Fill the region formed by the contour feature points into a binary region to obtain the initial contour binary region; Obtain the difference between the initial binary contour region and the currently detected paper mask region; if the difference is not empty, identify the pixels in the difference as uncovered pixels. If there are different pixels between the initial binary contour region and the currently detected paper mask region, use the different pixels as elements of the difference to construct the difference.
[0081] If the proportion of the difference set to the area of the current detected paper mask region exceeds the tolerance threshold (in this application, the tolerance threshold can be 0.1, 0.2, or 0.3), then the initial contour binary region is locally corrected based on the edge normal offset algorithm, specifically as follows: Offset: Each edge of the initial contour binary region is translated along its normal direction by a preset step size (the preset step size can be 1, 2, or 3 pixels in length) to obtain the translated edge; Correction: Obtain the intersection points of the translation edges, and use the intersection points of the translation edges as the vertices of the correction contour binary region to update the vertices of the initial contour binary region and obtain the correction contour binary region. Repeat the above translation-correction process until the area ratio between the corrected contour binary region and the currently detected paper mask region is lower than the threshold or the preset number of iterations (between 100 and 1000) is reached. Stop the adaptive correction and obtain the final contour feature points.
[0082] This step ultimately yields a stable, geometrically sound set of feature points that is fully consistent with the mask.
[0083] Step S4: Perform perspective restoration on the original image based on the second corrected image to obtain the restored paper image.
[0084] As an optional implementation, step S4 includes: The projective transformation matrix T is calculated based on the feature points in the second corrected image. The original image is then subjected to perspective correction based on the projective transformation matrix T to obtain the restored paper image.
[0085] Specifically, at least four feature points are first obtained from the second corrected image. Feature point matching is then performed between the second corrected image and the original image to obtain at least four pairs of matching feature points. Each pair includes a first feature point and a second feature point. The first feature point is a feature point in the original image, and the second feature point is a feature point in the second corrected image that matches the first feature point. These at least four pairs of feature points are used to construct four sets of transformation equations. Eight transformation parameters are solved from these four sets of equations. Finally, the projective transformation matrix T is constructed from the eight matrix parameters and a normalized 1. For a detailed explanation of the principle, please refer to the principle of homography transformation. In short: Substitute the pixels from the four feature point pairs into the homogeneous equation: ,in For feature points in the second corrected image, Let T be the homography matrix (projective transformation matrix) and T be the feature points in the original image. Eight transformation parameters are obtained by solving four sets of equations, thus yielding T. Then, the image is reconstructed according to the mapping relationship matrix to obtain the restored paper image. in It restores the pixels in the paper image. These are the pixels in the original image.
[0086] This step eliminates perspective distortion caused by the difference between the lens imaging plane and the document plane, thereby outputting a document image that is approximately viewed from the front, providing higher quality visual input for subsequent evaluation tasks.
[0087] Following step S4, this embodiment further includes the following steps: Step S5: Image Output: Output the perspective-restored document image to the display screen or downstream algorithm module.
[0088] Step S6: Data Augmentation Training System: To adapt to the complex conditions in real teaching environments, this invention constructs an augmentation system for the "mirror-instance segmentation" task during the training phase, including: 1. Random brightness perturbation: Enhances the robustness of the model to small changes in lighting conditions.
[0089] 2. Affine / rotational perturbation: Simulates a scenario where students place paper at arbitrary angles.
[0090] 3. Local perspective perturbation: Simulates nonlinear distortion caused by changes in the angle of the reflecting mirror.
[0091] 4. Random stretching and scaling: Simulates different object distances and different paper bending shapes.
[0092] 5. Background replacement and partial occlusion: Simulate real-world interference such as desktop clutter and hands obscuring the view.
[0093] Based on the above-described image correction and paper recognition method, this invention also provides an image correction and paper recognition system, which includes a preprocessing module, an initial correction module, a linkage correction module, and an image restoration module.
[0094] The preprocessing module is used to obtain at least two frames of raw images that are captured at adjacent times, and to preprocess the raw images to obtain a preprocessed image; wherein the raw image is a reflection imaging image; The initial correction module is used to perform initial correction on the preprocessed image to obtain the first corrected image; The linkage correction module is used to link the spatial and temporal information of the first correction image and adaptively perform secondary correction on the first correction image to obtain the second correction image. The image restoration module is used to perform perspective restoration on the original image based on the second corrected image to obtain a restored paper image.
[0095] The modules of the above system are used to execute the above image correction and paper recognition methods. For the specific implementation methods, please refer to the above image correction and paper recognition methods, which will not be repeated here.
[0096] Based on the aforementioned solution, the image correction and paper recognition method and system provided in this application require no additional hardware, no additional calibration, and no manual adjustment of the device angle by the user. The entire system has been deeply optimized and can stably achieve a processing performance of 3fps on mainstream learning machine chips, meeting the requirements of real-time interaction. This solution can improve the accuracy of image correction and paper recognition. Specifically, this is reflected in the following three aspects: (I) Achieving breakthrough applications of deep learning technology on low-computing-power learning machine platforms Traditional transformer solutions require significant computational power, making them difficult to deploy on learning machine platforms. This application has been approved. A single-stage image segmentation algorithm is proposed using lightweight deep learning network design techniques.
[0097] While maintaining accuracy, the computational load of the model was reduced to 10% of that of traditional solutions, and real-time deep learning inference at 3fps was successfully achieved on a learning machine platform with limited computing power.
[0098] Technical results: Under the same hardware conditions, the detection accuracy is improved by more than 50% compared with traditional image processing methods, while meeting the real-time requirements.
[0099] (II) Constructing a comprehensive data augmentation system based on deep learning To address the data distribution discrepancies caused by different learning machines, scenarios, and user habits, this invention transfers mature data augmentation techniques from the deep learning field to the learning machine scenario: Enhanced geometric transformations: Mirror flip (supports smart label mapping), random rotation, multi-scale scaling Optical property enhancement: random brightness perturbation, noise injection, and Gaussian blur simulation.
[0100] Enhanced structural deformation: local perspective perturbation, affine transformation, and random cropping.
[0101] Scene simulation enhancement: background replacement, partial occlusion, and composite distortion generation.
[0102] Technical results: The model's generalization ability is significantly improved, and the accuracy fluctuation under different devices and usage scenarios is reduced from ±30% of the traditional method to ±8%.
[0103] (III) Introducing a multi-layered stability guarantee mechanism Traditional solutions are prone to detection shake during continuous shooting, affecting user experience. This application uses target tracking technology to establish a complete stability assurance system: • Inter-frame correlation analysis: Establish a time continuity model through centroid trajectory and deformation features; • Adaptive state fusion: Intelligently balances current detection results with historical states; • Geometric constraint optimization: Ensure the rationality and stability of the quadrilateral structure.
[0104] Technical results: In video stream processing scenarios, the inter-frame stability of the detection results is improved by more than 70%, and the user experience is significantly improved.
[0105] The experimental results of the above scheme are shown in the figure. Figure 3-6 As shown.
[0106] Numerous specific details are set forth in the specification provided herein. However, it will be understood that embodiments of the invention may be practiced without these specific details. In some instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this specification.
[0107] Similarly, it should be understood that, in order to simplify this disclosure and aid in understanding one or more of the various aspects of the invention, in the above description of exemplary embodiments of the invention, various features of the invention are sometimes grouped together in a single embodiment, figure, or description thereof. However, this method of disclosure should not be construed as reflecting an intention that the claimed invention requires more features than are expressly recited in each claim. Rather, as reflected in the following claims, inventive aspects lie in fewer than all features of a single foregoing disclosed embodiment. Therefore, the claims following the detailed description are hereby expressly incorporated into this detailed description, wherein each claim itself is a separate embodiment of the invention.
[0108] Those skilled in the art will understand that modules in the device of the embodiments can be adaptively changed and placed in one or more devices different from that embodiment. Modules, units, or components in the embodiments can be combined into a single module, unit, or component, and further, they can be divided into multiple sub-modules, sub-units, or sub-components. Except where at least some of such features and / or processes or units are mutually exclusive, any combination can be used to combine all features disclosed in this specification (including the accompanying claims, abstract, and drawings) and all processes or units of any method or device so disclosed. Unless expressly stated otherwise, each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) may be replaced by an alternative feature that serves the same, equivalent, or similar purpose.
[0109] Furthermore, those skilled in the art will understand that although some embodiments herein include certain features included in other embodiments but not others, combinations of features from different embodiments are intended to be within the scope of the invention and form different embodiments. For example, in the following claims, any of the claimed embodiments can be used in any combination.
[0110] The various component embodiments of the present invention can be implemented in hardware, or as software modules running on one or more processors, or a combination thereof. Those skilled in the art will understand that microprocessors or digital signal processors (DSPs) can be used in practice to implement some or all of the functions of some or all of the components in the apparatus according to embodiments of the present invention. The present invention can also be implemented as a device or apparatus program (e.g., a computer program and computer program product) for performing part or all of the methods described herein. Such programs implementing the present invention can be stored on a computer-readable medium or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, provided on a carrier signal, or provided in any other form.
[0111] It should be noted that the above embodiments are illustrative of the invention and not restrictive, and that those skilled in the art can devise alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses should not be construed as limiting the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The invention can be implemented by means of hardware comprising several different elements and by means of a suitably programmed computer. In the unit claims enumerating several means, several of these means may be embodied by the same item of hardware. The use of the words first, second, and third, etc., does not indicate any order. These words can be interpreted as names.
Claims
1. An image correction and paper recognition method, characterized in that, include: Step S1: Obtain at least two original images with adjacent shooting times, preprocess the original images to obtain preprocessed images; wherein the original images are reflection imaging images; Step S2: Perform initial correction on the preprocessed image to obtain the first corrected image; Step S3: Link the spatial and temporal information of the first corrected image, and adaptively perform secondary correction on the first corrected image to obtain the second corrected image; Step S4: Perform perspective restoration on the original image based on the second corrected image to obtain the restored paper image; Step S2 includes: The preprocessed image is subjected to binary segmentation to obtain a segmented image; the segmented image includes multiple first segmentation masks. Perform connected component analysis on multiple first segmentation masks to determine the first segmentation mask with the largest area as the main body region of the paper. Morphological opening operations are performed on the main area of the paper to remove isolated noise, and closing operations are performed on the main area of the paper to fill gaps, making the boundary of the main area of the paper smoother and more continuous, thereby obtaining a morphologically optimized paper area. Edge correction and contour fitting are performed on the morphologically optimized paper region. Specifically, contour extraction is performed in the morphologically optimized paper region to obtain the edge of the first paper region; edge smoothing is performed on the edge of the first paper region to remove local sharp areas to obtain the edge of the second paper region; the image with the second paper edge region marked is used as the first correction image. At least two original images are processed through step S2 described above to obtain at least two first corrected images; step S3 includes: The latest original image from at least two original images is used as the current frame; the original image preceding the current frame is used as the previous frame. Center point offset monitoring calculation: Obtain the center point offset between the edge of the second paper region in the current frame image and the edge of the second paper region in the previous frame image; If the offset is less than the set value, it is assumed that the paper has not actually moved, and the edge detection result of the previous frame is used as the current edge detection result. Region overlap estimation: Obtain the intersection-union ratio (CIU) between the edge of the second paper region in the current frame image and the edge of the second paper region in the previous frame image; evaluate whether the paper shape is consistent between the two frames based on the CIU; if the paper shape is consistent between the two frames, determine that the two frames are consecutive observation images; Historical result reuse and state memory: Under the condition that the two frames are consecutively observed images, the edge detection result of the previous frame is used as the current edge detection result; the current edge detection result includes the currently detected paper mask area; Extract the contour feature points from the current edge detection results; The extraction of contour feature points from the current edge detection result includes: Mask candidate region selection and geometric smoothing: Morphological processing is performed on the current edge detection results to obtain a smooth boundary image; contour extraction is performed in the smooth boundary image to obtain the main boundary of the paper; convex hull is calculated based on the main boundary, and the influence of local depressions and noise is reduced through geometric convergence to obtain a stable main boundary; Extract feature points from the stable principal boundary to obtain contour feature points; Adaptive correction of contour feature points includes: Fill the region formed by the contour feature points into a binary region to obtain the initial contour binary region; Obtain the difference between the initial contour binary region and the currently detected paper mask region; if the difference is not empty, identify the pixels in the difference as uncovered pixels; If the proportion of the difference set to the area of the current detected paper mask region exceeds the tolerance threshold, then the initial contour binary region is locally corrected based on the edge normal offset algorithm, specifically: Offset: Each edge of the initial contour binary region is translated along its normal direction by a preset step size to obtain the translated edge; Correction: Obtain the intersection of adjacent translation edges to update the vertices of the initial contour binary region, and obtain the corrected contour binary region; Repeat the above translation-correction process until the area ratio between the corrected contour binary region and the currently detected paper mask region is lower than the threshold or the preset number of iterations is reached. Stop the adaptive correction and obtain the final contour feature points. Use the region formed by the contour feature points as the second corrected image.
2. The image correction and paper recognition method according to claim 1, characterized in that, The step of performing binary segmentation on the preprocessed image to obtain the segmented image is implemented through a lightweight model. Specifically, the preprocessed image is input into the lightweight model, and the lightweight model performs binary segmentation on the preprocessed image.
3. The image correction and paper recognition method according to claim 2, characterized in that, The lightweight model includes an adaptive processing module, an encoder, a decoder, and an edge-weighted loss optimization module; The input to the adaptive processing module is the preprocessed image in step S1, the input to the encoder is the output of the adaptive processing module, the input to the decoder is the output of the encoder, the input to the edge-weighted loss optimization module is the output of the decoder, and the decoder outputs a segmented image, which includes one or more first segmentation masks. The adaptive processing module is used to adaptively process the preprocessed image to obtain the fused image, specifically including: The preprocessed image is converted to grayscale to obtain a grayscale image; Set a block kernel with a size of 8×8 pixels, and use the block kernel to perform non-overlapping block division on the grayscale image; Calculate the average gray value of each block. and global mean ; Adaptive contrast enhancement processing is performed on the grayscale image to obtain the enhanced image, specifically as follows: in, For adaptive contrast enhancement factor, Represents pixels in a grayscale image pixel values, Represents pixels in an enhanced image The pixel values, i and j represent the x and y coordinates of the pixel, respectively; Indicates to The values are then truncated / trimmed. Specifically, if... ,but ;like ,but ;like ,but ; Indicates taking and The maximum of the two; The enhanced image is fused using bilinear interpolation to obtain the fused image. The encoder is used to perform deep convolutional spatial feature extraction and ordinary point convolutional cross-channel fusion on the fused image to obtain the encoded image; The decoder is used to perform nearest neighbor interpolation and edge feature compensation on the encoded image to obtain a binarized segmented image. The edge-weighted loss optimization module is used to calculate the loss value of the segmented image to train a lightweight model.
4. The image correction and paper recognition method according to claim 3, characterized in that, The loss value of the edge-weighted loss optimization module is calculated as follows: in, This represents the loss value of the edge-weighted loss optimization module. Indicates edge weights, , The Sobel edge strength values of the true mask are used to enhance the edge loss. ; These are balancing parameters. ; This represents the cross-entropy loss value between the segmented image and the ground truth image, which is provided beforehand. , Indicates the degree of regional overlap coefficient. , , This represents the edge weight of the k-th pixel in the segmented image. , This represents the edge intensity value of the k-th pixel. , This represents the pixel value of the k-th pixel in the segmented image, and N represents the number of pixels in the segmented image. This represents the pixel value of the k-th pixel in the actual label image.
5. The image correction and paper recognition method according to claim 3, characterized in that, Step S4 includes: The projective transformation matrix T is calculated based on the feature points in the second corrected image. The original image is then subjected to perspective correction based on the projective transformation matrix T to obtain the restored paper image.
6. An image correction and paper recognition system, characterized in that, include: The preprocessing module is used to obtain at least two frames of raw images that are captured at adjacent times, and to preprocess the raw images to obtain a preprocessed image. The original image is a reflection imaging image; The initial correction module is used to perform initial correction on the preprocessed image to obtain the first corrected image; The linkage correction module is used to link the spatial and temporal information of the first correction image and adaptively perform secondary correction on the first correction image to obtain the second correction image. The image restoration module is used to perform perspective restoration on the original image based on the second corrected image to obtain a restored paper image; Step S2 includes: The preprocessed image is subjected to binary segmentation to obtain a segmented image; the segmented image includes multiple first segmentation masks. Perform connected component analysis on multiple first segmentation masks to determine the first segmentation mask with the largest area as the main body region of the paper. Morphological opening operations are performed on the main area of the paper to remove isolated noise, and closing operations are performed on the main area of the paper to fill gaps, making the boundary of the main area of the paper smoother and more continuous, thereby obtaining a morphologically optimized paper area. Edge correction and contour fitting are performed on the morphologically optimized paper region. Specifically, contour extraction is performed in the morphologically optimized paper region to obtain the edge of the first paper region; edge smoothing is performed on the edge of the first paper region to remove local sharp areas to obtain the edge of the second paper region; the image with the second paper edge region marked is used as the first correction image. At least two original images are processed through step S2 described above to obtain at least two first corrected images; step S3 includes: The latest original image from at least two original images is used as the current frame; the original image preceding the current frame is used as the previous frame. Center point offset monitoring calculation: Obtain the center point offset between the edge of the second paper region in the current frame image and the edge of the second paper region in the previous frame image; If the offset is less than the set value, it is assumed that the paper has not actually moved, and the edge detection result of the previous frame is used as the current edge detection result. Region overlap estimation: Obtain the intersection-union ratio (CIU) between the edge of the second paper region in the current frame image and the edge of the second paper region in the previous frame image; evaluate whether the paper shape is consistent between the two frames based on the CIU; if the paper shape is consistent between the two frames, determine that the two frames are consecutive observation images; Historical result reuse and state memory: Under the condition that the two frames are consecutively observed images, the edge detection result of the previous frame is used as the current edge detection result; the current edge detection result includes the currently detected paper mask area; Extract the contour feature points from the current edge detection results; The extraction of contour feature points from the current edge detection result includes: Mask candidate region selection and geometric smoothing: Morphological processing is performed on the current edge detection results to obtain a smooth boundary image; contour extraction is performed in the smooth boundary image to obtain the main boundary of the paper; convex hull is calculated based on the main boundary, and the influence of local depressions and noise is reduced through geometric convergence to obtain a stable main boundary; Extract feature points from the stable principal boundary to obtain contour feature points; Adaptive correction of contour feature points includes: Fill the region formed by the contour feature points into a binary region to obtain the initial contour binary region; Obtain the difference between the initial contour binary region and the currently detected paper mask region; if the difference is not empty, identify the pixels in the difference as uncovered pixels; If the proportion of the difference set to the area of the current detected paper mask region exceeds the tolerance threshold, then the initial contour binary region is locally corrected based on the edge normal offset algorithm, specifically: Offset: Each edge of the initial contour binary region is translated along its normal direction by a preset step size to obtain the translated edge; Correction: Obtain the intersection of adjacent translation edges to update the vertices of the initial contour binary region, and obtain the corrected contour binary region; Repeat the above translation-correction process until the area ratio between the corrected contour binary region and the currently detected paper mask region is lower than the threshold or the preset number of iterations is reached. Stop the adaptive correction and obtain the final contour feature points. Use the region formed by the contour feature points as the second corrected image.
Citation Information
Patent Citations
Document image geometric correction method, system and device and medium
CN114418869A
Automatic book checking identification method and system for smart library
CN120997588A