Image correction method, homework correction method, related device, equipment and medium

By extracting image features and performing global modeling and local enhancement, the problem of poor adaptability to nonlinear deformation in existing technologies is solved, and the balance and applicability of image correction are improved, thereby enhancing the accuracy of image recognition and correction.

CN120976077AActive Publication Date: 2025-11-18IFLYTEK CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510923795.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-04
Publication Date
2025-11-18
Estimated Expiration
2045-07-04

AI Technical Summary

Technical Problem

Existing image correction methods are poorly adapted to nonlinear deformations and have uneven correction effects, making it difficult to effectively handle deformation and bending distortion of paper documents in image data.

Method used

By extracting image features from the image to be corrected, modeling the dependencies between pixels based on image features to obtain global features, and performing local feature enhancement, the image is corrected using a deformation field, including feature extraction, global modeling, local enhancement, and correction processing.

Benefits of technology

It improves the uniformity of image correction and its applicability to different deformations, effectively handles linear and nonlinear deformations, and improves the accuracy of image recognition and correction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120976077A_ABST
    Figure CN120976077A_ABST
Patent Text Reader

Abstract

The invention discloses an image correction method, a homework correction method, a related device, equipment and a medium, and the method comprises the steps: extracting the image features of a to-be-corrected image; modeling a dependency relationship between pixels in the to-be-corrected image based on the image features to obtain global features of the to-be-corrected image; performing local feature enhancement based on the global feature to obtain an enhanced feature; performing prediction based on the enhanced features to obtain a deformation field; and correcting the to-be-corrected image based on the deformation field to obtain a corrected image. According to the scheme, the image correction balance and the applicability to different deformations can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image processing, in particular to an image correction method, a homework correction method and related devices, equipment and media. BACKGROUND

[0002] With the continuous promotion of digital transformation, the image of paper documents is widely used in many fields. However, due to improper shooting angle or uneven lighting and other factors, the paper document often has distortion, bending and other distortions in the image data, which affects subsequent tasks such as image analysis.

[0003] The existing image correction method is mainly based on geometric modeling, neural network and the like, but the former is mainly suitable for regular deformation and has poor adaptability to nonlinear deformation, and the latter has uneven correction effect. Therefore, how to improve the uniformity of image correction and the adaptability to different deformations has become a problem to be solved. SUMMARY

[0004] The technical problem solved by the present application is to provide an image correction method, a homework correction method and related devices, equipment and media, which can improve the uniformity of image correction and the adaptability to different deformations.

[0005] In order to solve the above technical problem, the first aspect of the present application provides an image correction method, comprising: extracting image features of a to-be-corrected image; modeling the dependency relationship between pixels in the to-be-corrected image based on the image features to obtain global features of the to-be-corrected image; enhancing local features based on the global features to obtain enhanced features; predicting based on the enhanced features to obtain a deformation field; correcting the to-be-corrected image based on the deformation field to obtain a corrected image.

[0006] In order to solve the above technical problem, the second aspect of the present application provides a homework correction method, comprising: obtaining a homework image as a to-be-corrected image; performing correction processing based on the to-be-corrected image to obtain a corrected image; wherein the corrected image is obtained by the image correction method in the first aspect; identifying based on the corrected image to obtain homework data; correcting based on the homework data to obtain a correction result.

[0007] To solve the above technical problems, the third aspect of the present application provides an image correction device, comprising: a feature extraction module, a global modeling module, a local enhancement module, a deformation prediction module, and a correction processing module. The feature extraction module is configured to extract image features of a to-be-corrected image. The global modeling module is configured to model the dependency relationship between pixels in the to-be-corrected image based on the image features to obtain global features of the to-be-corrected image. The local enhancement module is configured to enhance local features based on the global features to obtain enhanced features. The deformation prediction module is configured to predict based on the enhanced features to obtain a deformation field. The correction processing module is configured to correct the to-be-corrected image based on the deformation field to obtain a corrected image.

[0008] To solve the above technical problems, the fourth aspect of the present application provides a work correction device, comprising: an image acquisition module, an image correction module, an image recognition module, and a data processing module. The image acquisition module is configured to acquire a work image as a to-be-corrected image. The image correction module is configured to perform correction processing based on the to-be-corrected image to obtain a corrected image. The corrected image is obtained by the image correction device of the third aspect described above. The image recognition module is configured to recognize based on the corrected image to obtain work data. The data processing module is configured to correct based on the work data to obtain a correction result.

[0009] To solve the above technical problems, the fifth aspect of the present application provides an electronic device, comprising at least a memory and a processor coupled to each other. The memory stores at least program instructions. The processor is configured to execute the program instructions to implement the image correction method of the first aspect described above, or the work correction method of the second aspect described above.

[0010] To solve the above technical problems, the sixth aspect of the present application provides a computer-readable storage medium, which stores program instructions capable of being executed by a processor. The program instructions are configured to implement the image correction method of the first aspect described above, or the work correction method of the second aspect described above.

[0011] The scheme extracts image features of the image to be corrected, models a dependency relationship between pixels in the image to be corrected based on the image features, obtains global features of the image to be corrected, performs local feature enhancement based on the global features to obtain enhanced features, and then performs prediction based on the enhanced features to obtain a deformation field, and corrects the image to be corrected based on the deformation field to obtain a corrected image. On the one hand, by extracting global features and performing local feature enhancement in sequence, the local details can be further enhanced under the premise of capturing the global deformation mode of the image, which helps to improve the balance of image correction. On the other hand, since the global features are extracted to capture the global deformation mode in the image correction process, compared with the traditional geometric-based method, not only linear deformation can be processed, but also non-linear deformation can be processed, which helps to improve the applicability to different deformations. Therefore, the balance of image correction and the applicability to different deformations can be improved. BRIEF DESCRIPTION OF DRAWINGS

[0012] Figure 1 is a flowchart of an embodiment of the image correction method of the present application; Figure 2a is a process diagram of an embodiment of the feature extraction of the image to be corrected of the present application; Figure 2b is a process diagram of an embodiment of the global feature extraction of the present application; Figure 2c is a process diagram of an embodiment of the local feature enhancement of the present application; Figure 2d is an effect diagram of an embodiment of the image correction based on the deformation field of the present application; Figure 2e is a process diagram of an embodiment of the image correction method of the present application; Figure 2f is an effect diagram of an embodiment of the image correction of the present application; Figure 2g is an effect diagram of another embodiment of the image correction of the present application; Figure 2h is an effect diagram of still another embodiment of the image correction of the present application; Figure 2i is an effect diagram of still another embodiment of the image correction of the present application; Figure 2j is an effect diagram of still another embodiment of the image correction of the present application; Figure 2k is an effect diagram of still another embodiment of the image correction of the present application; Figure 3 is a flowchart of an embodiment of the homework correction method of the present application; Figure 4 is a framework diagram of an embodiment of the image correction device of the present application; Figure 5 is a schematic diagram of a framework of an embodiment of the job correction device of the present application; Figure 6 is a schematic diagram of a framework of an embodiment of the electronic device of the present application; Figure 7 is a schematic diagram of a framework of an embodiment of the computer-readable storage medium of the present application. DETAILED DESCRIPTION

[0013] The scheme of the embodiments of the present application will be described in detail below with reference to the accompanying drawings of the specification.

[0014] In the following description, specific details are set forth in connection with the particular structures, interfaces, techniques, etc., to provide a thorough understanding of the present application. It should be understood, however, that this description is intended to be illustrative, and not restrictive.

[0015] The terms "system" and "network" are often used interchangeably herein. The term "and / or" herein is merely used to represent an associated relationship between associated objects, and can represent three relationships, for example, A and / or B, which can represent three cases of A alone, A and B together, and B alone. In addition, the segment " / " herein generally represents an "or" relationship between the front and rear associated objects. In addition, "multiple" herein means two or more than two.

[0016] Please refer to Figure 1 , Figure 1 is a flowchart of an embodiment of the image correction method of the present application. Specifically, it can include the following steps: Step S11: Extract the image features of the image to be corrected.

[0017] In one implementation scenario, the image to be corrected can be a photographed image of a paper document. In addition, the specific content of the image to be corrected can be different depending on the application scenario. For example, in a teaching application scenario, the image to be corrected can be a photographed image of a paper document such as a test paper, a workbook, etc.; or, for another example, in a medical application scenario, the image to be corrected can be a photographed image of a paper document such as a medical record, a test sheet, etc.; or, for yet another example, in an industrial application scenario, the image to be corrected can be a photographed image of a paper document such as an instruction manual, a test report, etc. Of course, the above examples are only a few possible examples of the image to be corrected in actual application, and the specific content of the image to be corrected is not limited herein, nor will it be exemplified one by one.

[0018] In one implementation scenario, when extracting features from the image to be corrected, the image to be corrected can be extracted based on a feature extraction model such as a convolutional neural network to obtain the image features of the image to be corrected. It should be noted that when the feature extraction model is constructed based on a convolutional neural network, the number of convolutional layers in the feature extraction model can not be limited.

[0019] In another implementation scenario, please refer to Figure 2a , Figure 2a is a process schematic diagram of an embodiment of the present application for feature extraction on a to-be-corrected image. As shown in Figure 2a , unlike the foregoing implementation, in order to reduce the computational complexity of subsequent processes such as global feature modeling as much as possible, when performing feature extraction on the to-be-corrected image, preliminary dimension reduction can also be performed on the to-be-corrected image to obtain initial features of the to-be-corrected image, and the initial features can be down-sampled to obtain to-be-processed features, and then the to-be-processed features can be extracted by a cross-stage local network to obtain first features, and the first features can be processed by a fusion attention mechanism to obtain second features, and the fusion attention mechanism can include channel attention and spatial attention, so that the second features can be extracted by the cross-stage local network to obtain third features, and then the third features can be subjected to spatial pyramid pooling to obtain image features. The above-mentioned method, by sequentially performing preliminary dimension reduction, down-sampling, feature extraction based on a cross-stage local network, fusion attention mechanism, feature extraction based on a cross-stage local network, and spatial pyramid pooling, can help to reduce the computational complexity of subsequent processes such as global feature modeling.

[0020] In a specific implementation scenario, for ease of description, the to-be-corrected image can be denoted as H*W*C, where H represents the height, W represents the width, and C represents the channel. For example, for an RGB image, the channel C can be 3. In addition, the resolution H*W of the to-be-corrected image can be set according to actual application needs, such as 496*496, and the resolution of the to-be-corrected image is not limited herein and will not be exemplified one by one. It should be noted that the paper document in the to-be-corrected image can have structural distortion due to shooting angle, paper bending or other deformation factors, and the corrected image obtained by the image correction process in the embodiment of the present disclosure can have removed the above-mentioned structural distortion. Of course, the to-be-corrected image can also not have structural distortion, and the corrected image obtained by the image correction process in the embodiment of the present disclosure can be substantially the to-be-corrected image itself.

[0021] In a specific implementation scenario, as shown in Figure 2a , the preliminary dimension reduction can be implemented by a convolution operation. For example, a 5*5 convolution kernel can be used to convert the features of the to-be-corrected image to achieve preliminary dimension reduction and retain key information.

[0022] In a specific implementation scenario, as shown in Figure 2a , after obtaining the initial features, the initial features can be down-sampled by a factor of two (i.e. Figure 2aThe initial features are down-sampled to obtain the to-be-processed features (s=2). Of course, the above example is only one possible example of the down-sampling rate, and other possible cases are not limited here, such as the down-sampling rate can also be set to 4, 8, etc., and the down-sampling rate will not be exemplified one by one here.

[0023] In one specific implementation scenario, as shown in Figure 2a After obtaining the to-be-processed features, the cross-stage partial network can be used to extract features from the to-be-processed features to obtain first features. It should be noted that the cross-stage partial network performs feature branch fusion at different levels to improve gradient flow and reduce computational redundancy and enhance feature expression capability. In addition, the specific process of feature extraction by the cross-stage partial network can refer to the technical details of the cross-stage partial (CSP) network, which will not be repeated here.

[0024] In one specific implementation scenario, as shown in Figure 2a After obtaining the first features, the first features can be further processed based on the fusion attention mechanism to obtain second features. It should be noted that by combining channel attention and spatial attention, the network can pay more attention to the key regions of the image, and the perception ability of the complex deformation region can be improved. In addition, the specific process of the fusion attention mechanism can refer to the technical details of the spatial attention and the channel attention, which will not be repeated here.

[0025] In one specific implementation scenario, as shown in Figure 2a After processing the first features based on the fusion attention mechanism to obtain the second features, and before extracting features from the second features based on the cross-stage partial network, it can be detected whether the total number of down-sampling cumulative executions is not higher than a preset number. It should be noted that the preset number can be set according to actual application needs, such as in the case of needing to fully perceive the deformation region, the preset number can be set to be appropriately larger, or in the case of needing to appropriately reduce the computational amount of the feature extraction stage, the preset number can be set to be appropriately smaller. As shown in Figure 2aAs shown, for ease of description, the preset number of times can be denoted as N. On this basis, in response to the total number of times of cumulative execution of downsampling being not higher than the preset number of times N, the second feature can be selected as a new initial feature. Exemplarily, the latest second feature can be selected as the new initial feature. Further, the new initial feature can be returned to perform the step of performing downsampling based on the initial feature to obtain the to-be-processed feature until the total number of times of cumulative execution of downsampling is equal to the preset number of times, at which time the latest second feature can be subjected to feature extraction by the cross-stage partial network to obtain the third feature and subsequent steps to obtain the image feature of the to-be-corrected image. It should be noted that the specific process of feature extraction by the cross-stage partial network can refer to the technical details of the Cross Stage Partial (CSP) network, which will not be described here. The above manner, after processing the first feature based on the fusion attention mechanism to obtain the second feature and before feature extraction of the second feature based on the cross-stage partial network, detects whether the total number of times of cumulative execution of downsampling is not higher than the preset number of times, so as to select the second feature as a new initial feature in the case of not higher than the preset number of times, and return to perform the step of performing downsampling based on the initial feature to the new initial feature for iteration until the total number of times of cumulative execution of downsampling is equal to the preset number of times, so as to significantly improve the perception of the deformed structure through multiple iterations.

[0026] In a specific implementation scenario, please refer to Figure 2a After obtaining the third feature, spatial pyramid pooling is performed based on the third feature to obtain the image feature of the to-be-corrected image. It should be noted that through different scale pooling operations, the understanding ability of the model to the global information of the image can be enhanced, so as to as far as possible ensure that the model can effectively cope with the deformation of complex scenes. In addition, the specific process of spatial pyramid pooling can refer to the technical details of Spatial Pyramid Pooling (SPP), which will not be described here.

[0027] Step S12: modeling the dependency relationship between pixels in the to-be-corrected image based on the image feature to obtain the global feature of the to-be-corrected image.

[0028] In one implementation scenario, as one possible implementation example, in order to model the dependency relationship between pixels in the image to be corrected, the sub-features at each position in the image feature can be processed based on an attention mechanism to model the dependency relationship between each position in the image feature, and then a global feature is obtained. It should be noted that since each position in the image feature corresponds to a different pixel region in the image to be corrected, the dependency relationship between pixels in the image to be corrected can be modeled accordingly. For ease of understanding, the image feature can be denoted as H’*W’*C’, and the sub-features at each position can be denoted as 1*1*C’, i.e., there are a total of H’*W’ sub-features at different positions. Based on this, the H’*W’ sub-features at different positions in the image feature can be processed based on a self-attention mechanism to obtain a global feature. Of course, the specific process of processing the sub-features at each position in the image feature based on the attention mechanism can refer to the technical details of the self-attention mechanism, and will not be described here.

[0029] In another implementation scenario, please refer to Figure 2b , Figure 2b is a process schematic diagram of an embodiment of the global feature extraction of the present application. As shown in Figure 2b , unlike the foregoing embodiments, as another possible implementation example, in order to model the dependency relationship between pixels in the image to be corrected, position encoding can also be performed on the image feature to obtain a fourth feature, and the fourth feature can be processed based on a multi-head self-attention mechanism to obtain a fifth feature, and then a non-linear transformation can be performed on the fifth feature to obtain a sixth feature, and finally the fourth feature and the sixth feature can be fused to obtain a global feature. The above-mentioned method can effectively model the dependency relationship between pixels in the image to be corrected by sequentially performing steps such as position encoding, multi-head self-attention mechanism, non-linear transformation, and feature fusion, which is helpful to effectively understand the global deformation of the image, and is particularly suitable for processing large-scale distortion or perspective deformation in the image.

[0030] In one specific implementation scenario, as shown in Figure 2b , before position encoding, the image feature can be converted in the channel. It can be understood that since a fixed number of channels is usually required in the subsequent processing process such as multi-head self-attention mechanism, the number of channels of the image feature can be adjusted by, for example, a 1*1 convolution layer, so as to adapt to the subsequent calculation requirements. For example, in the case where the number of channels required to adapt to the subsequent steps such as multi-head self-attention mechanism is K (such as 16, 32, 64, etc.), the number of channels of the image feature can be adjusted to K by 1*1 convolution.

[0031] In a specific implementation scenario, still taking the image feature represented as H' * W' * C' as an example, the sub-features at different positions can be respectively denoted as 1 * 1 * C', that is, there are a total of H' * W' sub-features at different positions. On this basis, the position encoding can be performed on each sub-feature at the position to ensure that the spatial relationship between the sub-features can be understood in the subsequent self-attention processing as much as possible. It should be noted that the specific manner of position encoding can refer to the technical details of position encoding (PE), which will not be described here.

[0032] In a specific implementation scenario, after the position encoding of the fourth feature, the fourth feature can be processed based on the multi-head self-attention mechanism to obtain a fifth feature. Specifically, the fourth feature can be divided according to the number of attention heads to obtain sub-features to be processed by each attention head. On this basis, in the process of being processed by the attention head, each sub-feature can be first projected by the query projection parameter, the key projection parameter and the value projection parameter to obtain the query feature, the key feature and the value feature, and then the attention score can be obtained based on the query feature and the key feature to establish the long-distance dependency relationship of different regions of the image, and the value feature is weighted by using the attention score to obtain the output feature of the attention head. Finally, the output features of the attention heads can be fused to obtain the fifth feature of the multi-head self-attention mechanism. Of course, the above description is only an exemplary description of the multi-head self-attention mechanism, and the technical details of the multi-head self-attention mechanism can be referred to, which will not be described here.

[0033] In a specific implementation scenario, after obtaining the fifth feature, a nonlinear transformation can be performed on the fifth feature to obtain a sixth feature. For example, the fifth feature can be nonlinearly transformed based on a feedforward network to obtain the sixth feature. Of course, the above example is only one possible example of nonlinear transformation, and other possible implementation manners are not limited here, and will not be exemplified one by one.

[0034] In a specific implementation scenario, after obtaining the sixth feature, the fourth feature and the sixth feature can be fused to obtain a global feature. For example, the fourth feature and the sixth feature can be spliced to obtain the global feature. Of course, the above example is only one possible implementation manner of fusing the fourth feature and the sixth feature, and other possible implementation manners are not limited here, and will not be exemplified one by one.

[0035] Step 13: Perform local feature enhancement based on the global feature to obtain an enhanced feature.

[0036] In one implementation scenario, the global feature can be processed by a feature pyramid network to implement local feature enhancement on the global feature, so as to obtain the enhanced feature. It should be noted that the feature pyramid network fuses high-level semantic features and low-level detail features through a top-down path, so that the shallow feature map obtains stronger semantic information and the deep feature map retains more local details. For specific processes of the local feature enhancement, reference can be made to technical details of the feature pyramid, which will not be described here.

[0037] In another implementation scenario, please refer to Figure 2c , Figure 2c is a process schematic diagram of an embodiment of the local feature enhancement of the present application. As shown in Figure 2c , different from the foregoing implementation, as another possible implementation example, the global feature can also be respectively processed by a plurality of dilated convolutions based on a plurality of dilated coefficients to obtain a plurality of output features of the plurality of dilated convolutions, and the plurality of dilated convolutions have different dilated coefficients. On this basis, the enhanced feature can be obtained by fusing the plurality of output features of the plurality of dilated convolutions. Through the convolution processing of the global feature by the dilated convolutions with different dilated coefficients and the fusion after the processing, the above-mentioned manner can more accurately restore the image details, especially in the local deformation area.

[0038] In one specific implementation scenario, the set of the dilated coefficients of the plurality of dilated convolutions can cover at least two of 1, 2, 3, 4, and 5, so as to restore the local details of the image, effectively expand the receptive field, capture the local information of the image from different scales, and be particularly suitable for processing the local deformation. Exemplarily, the plurality of dilated convolutions can be a total of 5, and the dilated coefficients of each dilated convolution can be 1, 2, 3, 4, and 5 respectively. Of course, the above-mentioned example is only one possible example of the plurality of dilated convolutions and the dilated coefficients thereof in the actual application process, and other possible cases are not limited here and will not be exemplified one by one.

[0039] In one specific implementation scenario, please refer to Figure 2c , after obtaining the output features of each dilated convolution, the output features of each dilated convolution can be spliced, and then the feature dimension can be reduced through a convolution layer such as 5*5 to obtain the enhanced feature.

[0040] Step S14: performing prediction based on the enhanced feature to obtain the deformation field.

[0041] Specifically, after obtaining the enhanced feature, the deformation field can be predicted based on the enhanced feature. For example, the decoder can decode the enhanced feature to obtain the deformation field. It should be noted that each element in the deformation field is used to describe the displacement of the pixels in the image (e.g., the offset in the X and Y directions). For example, still taking the to-be-corrected image represented as H*W*C as an example, as described above, the size of the feature tensor such as the image feature, the global feature, the local feature, etc. can be represented as H’*W’*C’, in which case the output size of the deformation field can be represented as 2*H’*W’. That is, each feature position can correspond to an offset of size 2 (offsets in the X and Y directions, respectively). In addition, the decoder can include but is not limited to a network structure such as a Transformer, and the network structure of the decoder is not limited herein.

[0042] Step S15: correcting the to-be-corrected image based on the deformation field to obtain a corrected image.

[0043] In one implementation scenario, the deformation field can be up-sampled to the resolution of the to-be-corrected image, and then the to-be-corrected image is corrected based on the up-sampled deformation field to obtain a corrected image. Please refer to Figure 2d Figure 2d is an effect diagram of an embodiment of the image correction method based on the deformation field. As shown in Figure 2d , Figure 2d the three images in the figure respectively represent the to-be-corrected image, the control points (red dots) based on the deformation field, and the corrected image. Taking the resolution of the to-be-corrected image as 496*496 as an example, if the resolution of the deformation field is 31*31, each 16*16 pixel region in the to-be-corrected image can correspond to a pixel point (control point) in the deformation field, and then the deformation field can be up-sampled, for example, up-sampled by 16 times in this example to obtain a 496*496 deformation field. That is, each pixel point in the to-be-corrected image can correspond to an offset of size 2 (offsets in the X and Y directions, respectively), and then the to-be-corrected image can be corrected based on the up-sampled deformation field to obtain a corrected image, so that the image geometry can be restored.

[0044] In one implementation scenario, please refer to Figure 2e Figure 2e is a process diagram of an embodiment of the image correction method. As shown in Figure 2e , the to-be-corrected image is first subjected to image feature extraction to obtain the image feature of the to-be-corrected image, and then the image feature is subjected to global feature modeling to obtain the global feature. The global feature is then subjected to local feature enhancement to obtain the enhanced feature, and the enhanced feature is further subjected to deformation field prediction to obtain the deformation field. Finally, the deformation field can act on the to-be-corrected image to obtain a corrected image. Please refer to Figure 2f ,​​Figure 2f is a schematic diagram of the effect of an embodiment of the image correction of the present application. As shown in Figure 2f , Figure 2f the left image in the figure is a to-be-corrected image, and the right image is a corrected image after the to-be-corrected image is subjected to image correction by the embodiment of the present disclosure. For a to-be-corrected image that has perspective deformation, after the image correction by the embodiment of the present disclosure, on the one hand, the perspective deformation can be corrected, and on the other hand, the overall global and local details of the paper document in the to-be-corrected image can be highlighted. Please refer to Figure 2g , Figure 2g is a schematic diagram of the effect of another embodiment of the image correction of the present application. As shown in Figure 2g , Figure 2g the left image in the figure is a to-be-corrected image, and the right image is a corrected image after the to-be-corrected image is subjected to image correction by the embodiment of the present disclosure. For a to-be-corrected image that has curl deformation when a book is turned, after the image correction by the embodiment of the present disclosure, on the one hand, the curl deformation can be corrected, and on the other hand, the overall global and local details of the paper document in the to-be-corrected image can be highlighted. Please refer to Figure 2h , Figure 2h is a schematic diagram of the effect of still another embodiment of the image correction of the present application. As shown in Figure 2h , Figure 2h the left image in the figure is a to-be-corrected image, and the right image is a corrected image after the to-be-corrected image is subjected to image correction by the embodiment of the present disclosure. For a to-be-corrected image that has curl deformation under a normal view angle, after the image correction by the embodiment of the present disclosure, on the one hand, the curl deformation can be corrected, and on the other hand, the overall global and local details of the paper document in the to-be-corrected image can be highlighted. Please refer to Figure 2i , Figure 2i is a schematic diagram of the effect of still another embodiment of the image correction of the present application. As shown in Figure 2i , Figure 2i the left image in the figure is a to-be-corrected image, and the right image is a corrected image after the to-be-corrected image is subjected to image correction by the embodiment of the present disclosure. For a to-be-corrected image that has curl deformation under a top view angle, after the image correction by the embodiment of the present disclosure, on the one hand, the curl deformation can be corrected, and on the other hand, the overall global and local details of the paper document in the to-be-corrected image can be highlighted. Please refer to Figure 2j , Figure 2j is a schematic diagram of the effect of still another embodiment of the image correction of the present application. As shown in Figure 2j , Figure 2j the left image in the figure is a to-be-corrected image, and the right image is a corrected image after the to-be-corrected image is subjected to image correction by the embodiment of the present disclosure. For a to-be-corrected image that has curl deformation under a folded view angle, after the image correction by the embodiment of the present disclosure, on the one hand, the curl deformation can be corrected, and on the other hand, the overall global and local details of the paper document in the to-be-corrected image can be highlighted. Please refer to Figure 2k ,Figure 2k is a schematic diagram of the effect of another embodiment of the image correction method of the present application. As shown in Figure 2k , Figure 2k the left image in the figure is the image to be corrected, and the right image is the corrected image after the image correction by the embodiment of the present application. For the image to be corrected with wrinkles caused by kneading, after the image correction by the embodiment of the present application, on the one hand, the wrinkles can be corrected, and on the other hand, the overall global and local details of the paper document in the image to be corrected can be highlighted. As can be seen from the above examples, the embodiment of the present application can effectively correct various complex deformations. Of course, Figure 2f to Figure 2g the images to be corrected and their corrected images shown in the figure are only several possible examples in the actual application, and other possible cases are not limited herein, nor are they exemplified one by one.

[0045] The above scheme extracts the image features of the image to be corrected, models the dependency relationship between the pixels in the image to be corrected based on the image features, obtains the global features of the image to be corrected, enhances the local features based on the global features, obtains the enhanced features, and then performs prediction based on the enhanced features to obtain the deformation field, and further corrects the image to be corrected based on the deformation field to obtain the corrected image. Therefore, on the one hand, by extracting the global features and enhancing the local features in sequence, the local details can be further enhanced under the premise of capturing the global deformation mode of the image, which helps to improve the balance of image correction. On the other hand, since the global features are extracted to capture the global deformation mode in the image correction process, compared with the traditional geometric-based method, not only linear deformation can be processed, but also non-linear deformation can be processed, which helps to improve the applicability to different deformations. Therefore, the balance of image correction and the applicability to different deformations can be improved.

[0046] Please refer to Figure 3 , Figure 3 is a flowchart of an embodiment of the homework correction method of the present application. Specifically, it can include the following steps: Step S31: Obtain the homework image as the image to be corrected.

[0047] In one implementation scenario, the homework image can be obtained by an electronic device such as a smartphone or a learning machine by photographing a paper document (such as a test paper or a workbook), and directly processed by the electronic device as the image to be corrected.

[0048] In another implementation scenario, the homework image can also be obtained by an electronic device such as a smartphone or a learning machine by photographing a paper document (such as a test paper or a workbook), and then uploaded to a homework correction system by the electronic device for subsequent processing by the homework correction system.

[0049] Step S32: Perform correction processing based on the image to be corrected to obtain the corrected image.

[0050] In the embodiments of the present disclosure, the corrected image is obtained through the process steps in the above-mentioned image correction method embodiments, and can be specifically referred to the above-mentioned image correction method embodiments, which will not be repeated here.

[0051] Step S33: identifying based on the corrected image to obtain the work data.

[0052] Specifically, the recognition technology such as OCR (Optical Character Recognition) can be used to identify the corrected image to obtain the work data. It should be noted that, since the work image has been subjected to correction processing as the to-be-corrected image in the above-mentioned step, the corrected image has corrected the deformation structure that may originally exist in the to-be-corrected image as much as possible, and highlighted the global and local details of the document, so the accuracy of image recognition can be improved as much as possible on this basis.

[0053] Step S34: correcting based on the work data to obtain the correction result.

[0054] Specifically, the standard answer of the work data can be obtained in advance, and then the pre-trained language model such as BERT can be used to understand the semantics and give the correction result based on the work data and its standard answer; or, a large language model can also be used to correct the work data, such as constructing a prompt instruction based on the work data, and the prompt instruction is used to instruct the large language model to correct the work data, so that the output result of the large language model in response to the prompt instruction can be obtained as the correction result of the work data. Of course, the above examples are only several possible implementation modes of correcting the work data in actual application, and the correction mode of the work data is not limited here, and will not be exemplified one by one.

[0055] The above scheme obtains the work image as the to-be-corrected image, performs correction processing based on the to-be-corrected image to obtain the corrected image, and the corrected image is obtained through the process steps in the above-mentioned image correction method embodiments, so that the work data is obtained by identifying the corrected image, and the correction result is obtained based on the work data. Since the corrected image is obtained through the process steps in the above-mentioned image correction method embodiments, the balance of image correction and the applicability to different deformations can be improved, and based on this, the accuracy of the work data can be improved, and accordingly the accuracy of the work correction can also be improved.

[0056] Please refer to Figure 4 , Figure 4Fig. 1 is a schematic diagram of a framework of an embodiment of the image correction device. The image correction device 40 comprises a feature extraction module 41, a global modeling module 42, a local enhancement module 43, a deformation prediction module 44, and a correction processing module 45. The feature extraction module 41 is configured to extract image features of an image to be corrected. The global modeling module 42 is configured to model a dependency relationship between pixels in the image to be corrected based on the image features, to obtain global features of the image to be corrected. The local enhancement module 43 is configured to perform local feature enhancement based on the global features, to obtain enhanced features. The deformation prediction module 44 is configured to perform prediction based on the enhanced features, to obtain a deformation field. The correction processing module 45 is configured to perform correction on the image to be corrected based on the deformation field, to obtain a corrected image.

[0057] The above scheme extracts image features of an image to be corrected, models a dependency relationship between pixels in the image to be corrected based on the image features, to obtain global features of the image to be corrected, performs local feature enhancement based on the global features, to obtain enhanced features, and performs prediction based on the enhanced features, to obtain a deformation field, and performs correction on the image to be corrected based on the deformation field, to obtain a corrected image. On the one hand, by extracting global features and performing local feature enhancement in sequence, local details can be further enhanced under the premise of capturing global deformation patterns, which helps to improve the balance of image correction. On the other hand, since global features are extracted to capture global deformation patterns in the image correction process, compared with traditional geometric-based methods, not only linear deformations can be handled, but also non-linear deformations can be more easily handled, which helps to improve the applicability to different deformations. Therefore, the balance of image correction and the applicability to different deformations can be improved.

[0058] In some disclosed embodiments, the feature extraction module 41 comprises a preliminary dimension reduction submodule configured to perform preliminary dimension reduction on the image to be corrected, to obtain initial features of the image to be corrected. The feature extraction module 41 comprises a down-sampling submodule configured to perform down-sampling on the initial features, to obtain processed features. The feature extraction module 41 comprises a first extraction submodule configured to perform feature extraction on the processed features based on a cross-stage local network, to obtain first features. The feature extraction module 41 comprises a fusion attention submodule configured to process the first features based on a fusion attention mechanism, to obtain second features. The fusion attention mechanism comprises channel attention and spatial attention. The feature extraction module 41 comprises a second extraction submodule configured to perform feature extraction on the second features based on the cross-stage local network, to obtain third features. The feature extraction module 41 comprises a pyramid pooling submodule configured to perform spatial pyramid pooling on the third features, to obtain the image features.

[0059] In some disclosed embodiments, the feature extraction module 41 comprises a number of detection submodule for detecting whether the total number of times of performing down-sampling accumulation is not higher than a preset number of times; the feature extraction module 41 comprises a loop iteration submodule for, in response to the total number of times of performing down-sampling accumulation being not higher than the preset number of times, selecting a second feature as a new initial feature, and returning to performing the step of performing down-sampling based on the initial feature to obtain a to-be-processed feature on the new initial feature until the total number of times of performing down-sampling accumulation is not higher than the preset number of times.

[0060] In some disclosed embodiments, the global modeling module 42 comprises a position encoding submodule for performing position encoding based on the image feature to obtain a fourth feature; the global modeling module 42 comprises a multi-head self-attention submodule for processing the fourth feature based on a multi-head self-attention mechanism to obtain a fifth feature; the global modeling module 42 comprises a nonlinear transformation submodule for performing nonlinear transformation based on the fifth feature to obtain a sixth feature; the global modeling module 42 comprises a residual connection submodule for fusing based on the fourth feature and the sixth feature to obtain a global feature.

[0061] In some disclosed embodiments, the local enhancement module 43 comprises a cavity convolution submodule for respectively performing convolution processing on the global feature based on a plurality of cavity convolutions to obtain a plurality of output features of the plurality of cavity convolutions; wherein the plurality of cavity convolutions have different cavity coefficients; the local enhancement module 43 comprises a feature fusion submodule for fusing based on the plurality of output features of the plurality of cavity convolutions to obtain an enhanced feature.

[0062] In some disclosed embodiments, the set of cavity coefficients of the plurality of cavity convolutions respectively covers at least two cavity coefficients of 1, 2, 3, 4, and 5.

[0063] In some disclosed embodiments, the rectification processing module 45 comprises an up-sampling submodule for up-sampling based on the deformation field to the resolution of the to-be-rectified image; the rectification processing module 45 comprises a rectification submodule for rectifying the to-be-rectified image based on the deformation field after up-sampling to obtain a rectified image.

[0064] Please refer to Figure 5 , Figure 5 is a frame schematic diagram of an embodiment of the work correction device of the present application. The work correction device 50 comprises an image acquisition module 51, an image rectification module 52, an image recognition module 53, and a data processing module 54. The image acquisition module 51 is configured to acquire a work image as a to-be-rectified image. The image rectification module 52 is configured to perform rectification processing based on the to-be-rectified image to obtain a rectified image. The rectified image is obtained by the image rectification device described above. The image recognition module 53 is configured to perform recognition based on the rectified image to obtain work data. The data processing module 54 is configured to perform correction based on the work data to obtain a correction result.

[0065] The above scheme, the work correction device 50 acquires a work image as a to-be-corrected image, performs correction processing based on the to-be-corrected image to obtain a corrected image, and the corrected image is obtained by the above image correction device, so as to perform recognition on the corrected image to obtain work data, and then performs correction based on the work data to obtain a correction result. Since the corrected image is obtained by the above image correction method embodiment, the uniformity of image correction and the applicability to different deformations can be improved. Based on this, the accuracy of the work data can be improved, and accordingly the accuracy of the work correction can also be improved.

[0066] Please refer to Figure 6 , Figure 6 is a schematic diagram of the framework of an embodiment of the electronic device. The electronic device 60 at least includes a memory 61 and a processor 62 coupled with each other. The memory 61 at least stores program instructions. The processor 62 is configured to execute the program instructions to implement the steps in any of the above image correction method embodiments or the steps in any of the above work correction method embodiments. For details, please refer to the foregoing disclosed embodiments, which will not be repeated here. As a possible example, the electronic device 60 can include but is not limited to a mobile phone, a tablet computer, a learning machine, a smart large screen, a server, and the like. The specific type of the electronic device 60 is not limited here.

[0067] Specifically, the processor 62 is configured to control itself and the memory 61 to implement the steps in any of the above image correction method embodiments or the steps in any of the above work correction method embodiments. The processor 62 can also be referred to as a CPU (Central Processing Unit). The processor 62 can be an integrated circuit chip with a signal processing capability. The processor 62 can also be a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field-Programmable Gate Array) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor. In addition, the processor 62 can be implemented by an integrated circuit chip together.

[0068] In the above scheme, the electronic device 60 extracts the image features of the image to be corrected, models the dependency relationship between pixels in the image to be corrected based on the image features, obtains the global features of the image to be corrected, performs local feature enhancement based on the global features to obtain enhanced features, and then performs prediction based on the enhanced features to obtain the deformation field, and further corrects the image to be corrected based on the deformation field to obtain the corrected image. On the one hand, by extracting the global features and performing local feature enhancement in sequence, the local details can be further enhanced under the premise of capturing the global deformation mode of the image, which helps to improve the balance of image correction. On the other hand, since the global features are extracted to capture the global deformation mode in the image correction process, compared with the traditional geometric-based method, it can not only handle linear deformation but also be more conducive to handling nonlinear deformation, which helps to improve the adaptability to different deformations. Therefore, the balance of image correction and the adaptability to different deformations can be improved. In addition, the work image is obtained as the image to be corrected, the correction processing is performed based on the image to be corrected to obtain the corrected image, and the corrected image is obtained by the process steps in the above image correction method embodiment. Then, the work data is obtained by identifying the corrected image, and the grading result is obtained based on the work data. Since the corrected image is obtained by the process steps in the above image correction method embodiment, the balance of image correction and the adaptability to different deformations can be improved. Based on this, the image recognition can be performed, the accuracy of the work data can be improved, and the accuracy of the work grading can also be improved accordingly.

[0069] Please refer to Figure 7 , Figure 7 is a frame diagram of an embodiment of the computer readable storage medium of the present application. The computer readable storage medium 70 stores program instructions 71 capable of being executed by the processor, and the program instructions 71 are used to implement the steps in any of the above image correction method embodiments or the steps in any of the above work grading method embodiments.

[0070] The above scheme, the computer readable storage medium 70 extracts the image features of the image to be corrected, and models the dependency relationship between the pixels in the image to be corrected based on the image features to obtain the global features of the image to be corrected, and then performs local feature enhancement based on the global features to obtain enhanced features, and then performs prediction based on the enhanced features to obtain the deformation field, and then corrects the image to be corrected based on the deformation field to obtain the corrected image. On the one hand, by extracting the global features and performing local feature enhancement in sequence, the local details can be further enhanced under the premise of capturing the global deformation mode of the image, which helps to improve the balance of image correction. On the other hand, since the global features are extracted to capture the global deformation mode in the image correction process, compared with the traditional geometric-based method, it can not only handle linear deformation but also be more conducive to handling nonlinear deformation, which helps to improve the adaptability to different deformations. Therefore, the balance of image correction and the adaptability to different deformations can be improved. In addition, the work image is obtained as the image to be corrected, the correction processing is performed based on the image to be corrected to obtain the corrected image, and the corrected image is obtained by the process steps in the above image correction method embodiment. Then, the work data is obtained by identifying the corrected image, and the correction result is obtained based on the work data. Since the corrected image is obtained by the process steps in the above image correction method embodiment, the balance of image correction and the adaptability to different deformations can be improved. Based on this, image recognition can be performed, the accuracy of the work data can be improved, and the accuracy of the work correction can also be improved accordingly.

[0071] In some embodiments, the device provided by the embodiments of the present disclosure has functions or includes modules that can be used to perform the methods described in the above method embodiments, and the specific implementation can refer to the description of the above method embodiments. For brevity, it will not be repeated here.

[0072] The above description of various embodiments tends to emphasize the differences between various embodiments, and the same or similar parts can be mutually referred to. For brevity, it will not be repeated here.

[0073] In several embodiments provided in the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the above-described device implementation is only schematic. For example, the division of modules or units is only a logical function division, and there can be another division manner in actual implementation. For example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the shown or discussed mutual elements can be indirect coupling or communication connection through some interface, device or unit, which can be electrical, mechanical or other form.

[0074] The units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, i.e., may be located in one place, or may be distributed to multiple network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the present embodiment.

[0075] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0076] The integrated unit, if realized in the form of a software functional unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor (processor) to execute all or part of the steps of the methods of the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.

[0077] If the technical solutions of the present application involve personal information, the product applying the technical solutions of the present application has been explicitly informed of the personal information processing rules before processing the personal information, and has obtained the personal independent consent. If the technical solutions of the present application involve sensitive personal information, the product applying the technical solutions of the present application has obtained the personal independent consent before processing the sensitive personal information, and at the same time meets the requirement of "explicit consent". For example, at the personal information collection device such as camera, a clear and prominent mark is set to inform that the personal information collection range has been entered, and the personal information will be collected. If the person voluntarily enters the collection range, it is regarded as agreeing to collect the personal information; or on the device for processing personal information, the personal information processing rules are informed by using obvious marks / information, and the personal authorization is obtained by means of pop-up information or asking the person to upload his / her personal information, etc. The personal information processing rules can include personal information processor, personal information processing purpose, processing method, and personal information type, etc.

Claims

1. An image correction method, characterized in that, include: Extract image features from the image to be corrected; Based on the image features, the dependencies between pixels in the image to be corrected are modeled to obtain the global features of the image to be corrected; Based on the global features, local feature enhancement is performed to obtain enhanced features; Based on the enhanced features, a prediction is made to obtain the deformation field; The image to be corrected is corrected based on the deformation field to obtain a corrected image.

2. The method according to claim 1, characterized in that, The extraction of image features from the image to be corrected includes: Preliminary dimensionality reduction is performed on the image to be corrected to obtain the initial features of the image to be corrected; Based on the initial features, downsampling is performed to obtain the features to be processed; The first feature is obtained by extracting features from the features to be processed based on a cross-stage local network. The first feature is processed based on a fusion attention mechanism to obtain the second feature; wherein, the fusion attention mechanism includes channel attention and spatial attention; Based on the cross-stage local network, feature extraction is performed on the second feature to obtain the third feature; The image features are obtained by performing spatial pyramid pooling based on the third feature.

3. The method according to claim 2, characterized in that, After processing the first feature based on the fusion attention mechanism to obtain the second feature, and before performing feature extraction on the second feature based on the cross-stage local network, the method further includes: Check whether the total number of times the downsampling has been performed is not higher than a preset number; In response to the fact that the total number of downsampling executions is not higher than the preset number, the second feature is selected as the new initial feature, and the step of downsampling based on the initial feature to obtain the feature to be processed is returned for the new initial feature, until the total number of downsampling executions is equal to the preset number.

4. The method according to claim 1, characterized in that, The step of modeling the dependencies between pixels in the image to be corrected based on the image features to obtain the global features of the image to be corrected includes: Based on the image features, position encoding is performed to obtain the fourth feature; The fifth feature is obtained by processing the fourth feature based on the multi-head self-attention mechanism; Based on the fifth feature, a nonlinear transformation is performed to obtain the sixth feature; The global feature is obtained by fusing the fourth feature and the sixth feature.

5. The method according to claim 1, characterized in that, The process of enhancing local features based on the global features to obtain enhanced features includes: The global features are processed by several dilated convolutions to obtain the output features of each dilated convolution; wherein the several dilated convolutions have different dilation coefficients. The enhanced features are obtained by fusing the output features of the several dilated convolutions.

6. The method according to claim 5, characterized in that, The set of dilation coefficients for each of the plurality of dilated convolutions covers at least two of the dilation coefficients in 1, 2, 3, 4, and 5.

7. The method according to claim 1, characterized in that, The step of correcting the image to be corrected based on the deformation field to obtain a corrected image includes: Upsample to the resolution of the image to be corrected based on the deformation field; The image to be corrected is corrected based on the deformation field after upsampling to obtain the corrected image.

8. A method for grading homework, characterized in that, include: Obtain the image of the work as the image to be corrected; A corrected image is obtained by performing a correction process on the image to be corrected; wherein the corrected image is obtained by the image correction method according to any one of claims 1 to 7; Based on the corrected image, the operation data is obtained; The work data is graded to obtain the grading results.

9. An image correction device, characterized in that, include: The feature extraction module is used to extract image features from the image to be corrected. A global modeling module is used to model the dependencies between pixels in the image to be corrected based on the image features, so as to obtain the global features of the image to be corrected. The local enhancement module is used to enhance local features based on the global features to obtain enhanced features; The deformation prediction module is used to predict the deformation field based on the enhanced features; The correction processing module is used to correct the image to be corrected based on the deformation field to obtain a corrected image.

10. A homework correction device, characterized in that, include: The image acquisition module is used to acquire the work image as the image to be corrected. An image correction module is used to perform correction processing on the image to be corrected to obtain a corrected image; wherein the corrected image is obtained by the image correction device according to claim 9; The image recognition module is used to perform recognition based on the corrected image to obtain job data; The data processing module is used to revise the work data and obtain the revision results.

11. An electronic device, characterized in that, It includes at least a memory and a processor, wherein the memory stores at least program instructions, and the processor is used to execute the program instructions to implement the image correction method according to any one of claims 1 to 7, or the job correction method according to claim 8.

12. A computer-readable storage medium, characterized in that, The device stores program instructions that can be executed by a processor, the program instructions being used to implement the image correction method according to any one of claims 1 to 7, or the job correction method according to claim 8.

Citation Information

Patent Citations

  • Exercise correction method and system based on question recognition and intelligent family education learning machine

    CN113673405A

  • Semantic segmentation method and device for multi-scale feature fusion and storage medium

    CN115272677A

  • Warped document correction method based on multi-scale hybrid encoder and feature fusion

    CN118334675A

  • Mobile phone image distortion correction method and device based on deep learning

    CN118396903A

  • QR two-dimensional code correction and repair method and system

    CN118940782A