Image processing method and device, electronic equipment and storage medium

By determining the schema mapped area and target coordinate points in the input image, combining the perspective transformation matrix and deformation field, the problem of backward mapping cannot be converted into the coordinate system is solved, and efficient image correction and functional service support are achieved.

CN120298209APending Publication Date: 2025-07-11深圳市星桐科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510295076.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In the prior art, the backward mapping method cannot realize the conversion of the input image coordinate system to the corrected image coordinate system, resulting in some functional services being unable to be realized, and forward mapping may lead to the missing and overlap of pixel points in the corrected image.

Method used

By obtaining the original coordinate points of the input image, the schemap area is determined using the backward mapping relationship, and the target coordinate points are selected according to the distance between each source pixel point and the original coordinate point, and combining the perspective transformation matrix and deformation field, the forward mapping relationship is determined.

Benefits of technology

Without adding too much execution overhead, pixel points are missing and overlapped, and the conversion from the input image coordinate system to the corrected image coordinate system is realized, supporting the implementation of more functional services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298209A_ABST
    Figure CN120298209A_ABST
Patent Text Reader

Abstract

The invention provides an image processing method and device, electronic equipment and a storage medium. According to the embodiment of the invention, for each original coordinate point, a plurality of to-be-mapped coordinate points corresponding to the original coordinate point under the coordinate system of the corrected image are determined, and then the preset number of to-be-mapped coordinate points with the minimum offset distance in the to-be-mapped coordinate points are selected as the to-be-mapped coordinate points. And taking the coordinate point as a target coordinate point corresponding to the original coordinate point in the coordinate system of the corrected image. According to the implementation mode, the problem that the mapping mode of backward mapping cannot support conversion from the input image coordinate system to the corrected image coordinate system is solved, and on the premise of not excessively increasing the method execution overhead, the phenomena of pixel missing and pixel overlapping of the corrected image caused by forward mapping are avoided; and the conversion from the coordinates in the input image coordinate system to the coordinates in the corrected image coordinate system is realized, and the realization of more functional services can be supported.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of image processing technologies, and in particular, to an image processing method, apparatus, electronic device, and storage medium. Background Art

[0002] When photographing document materials, the images captured by a camera may undergo various distortions. When inputting an image into an algorithm corresponding to certain services, it is necessary to map an input image with relatively severe distortions into a corrected image with relatively mild distortions, and then input the corrected image into the service algorithm to ensure the normal implementation of the service. The above mapping is usually divided into two mapping methods: forward mapping and backward mapping. Since the forward mapping method often causes the phenomenon of pixel point loss and pixel point overlap in the corrected image, the backward mapping method is mostly used.

[0003] In related technologies, the backward mapping method lacks the mapping relationship between pixel points in the input image and pixel points in the corrected image, so the conversion of coordinates in the input image coordinate system to coordinates in the corrected image coordinate system cannot be achieved, which hinders the implementation of some functional services. Summary of the Invention

[0004] To overcome the problems existing in related technologies, the present disclosure provides an image processing method, apparatus, electronic device, and storage medium.

[0005] The first aspect of the present disclosure provides an image processing method, and the method includes:

[0006] Obtain an original coordinate point in the input image, where the original coordinate point represents the imaging position of a first target object in the input image;

[0007] According to the feature information of the input image and the feature information of the corrected image obtained after the input image is corrected, determine a quasi-mapping area of the original coordinate point in the corrected image;

[0008] According to the obtained backward mapping relationship, determine a source pixel point in the input image corresponding to each pixel point in the quasi-mapping area, where the backward mapping relationship is the mapping relationship from a pixel point in the corrected image to a pixel point in the input image;

[0009] According to the distance between each source pixel point and the original coordinate point, select a target coordinate point corresponding to the original coordinate point in the quasi-mapping area, where the target coordinate point is used to determine the forward mapping relationship corresponding to the original coordinate point, and the forward mapping relationship is the mapping relationship from a pixel point in the input image to a pixel point in the corrected image.

[0010] Optionally, determining a quasi-mapping region of the original coordinate point in the corrected image according to the feature information of the input image and the feature information of the corrected image obtained after the input image is corrected includes:

[0011] Mapping the original coordinate point to the corrected image according to the size of the input image and the size of the corrected image obtained after the input image is corrected to obtain a quasi-mapping coordinate point;

[0012] Taking a neighboring region with a preset size around the quasi-mapping coordinate point in the corrected image as the quasi-mapping region of the original coordinate point.

[0013] Optionally, the method further includes:

[0014] Determining an imaging region of a second target object in the input image and determining a perspective transformation matrix corresponding to the imaging region;

[0015] Performing perspective transformation correction on the imaging region based on the perspective transformation matrix to obtain a preprocessed image;

[0016] Determining a deformation field corresponding to the preprocessed image, where the deformation field represents a mapping relationship from pixel points in the corrected image to pixel points in the input image;

[0017] Processing the deformation field corresponding to the preprocessed image according to the perspective transformation matrix to obtain the backward mapping relationship.

[0018] Optionally, processing the deformation field corresponding to the preprocessed image according to the perspective transformation matrix to obtain the backward mapping relationship includes:

[0019] Converting the coordinates of each point in the deformation field corresponding to the preprocessed image in the coordinate system of the preprocessed image into homogeneous coordinates, and updating the deformation field according to the result of multiplying each homogeneous coordinate by the inverse matrix of the perspective transformation matrix to obtain the deformation field corresponding to the input image.

[0020] Optionally, selecting a target coordinate point corresponding to the original coordinate point in the quasi-mapping region according to the distance between each source pixel point and the original coordinate point includes:

[0021] Sorting the source pixel points in the input image according to the distance between each source pixel point and the original coordinate point from small to large to obtain a source pixel point sequence;

[0022] Selecting the Nth source pixel point from the source pixel point sequence, and determining the pixel point corresponding to the Nth source pixel point in the corrected image as the target coordinate point corresponding to the original coordinate point.

[0023] Optionally, obtaining the original coordinate points in the input image includes:

[0024] Determining the imaging position of the first target object in the input image based on an object detection algorithm;

[0025] Determining at least one original coordinate point in the input image according to the imaging position of each first target object in the input image.

[0026] Optionally, the method further includes:

[0027] Obtaining the image within a preset range around the target coordinate point in the corrected image, performing preset processing on the image within the preset range, and presenting the processing result to the user;

[0028] And / or, determining the quasi-mapping area of the original coordinate point in the corrected image according to the feature information of the input image and the corrected image includes:

[0029] If the forward mapping relationship corresponding to the original coordinate point has not been obtained yet, determining the quasi-mapping area of the original coordinate point in the corrected image according to the feature information of the input image and the corrected image.

[0030] A second aspect of the present disclosure provides an image processing apparatus, the apparatus includes:

[0031] A coordinate acquisition module, configured to acquire the original coordinate points in the input image, where the original coordinate points represent the imaging position of the first target object in the input image;

[0032] A quasi-mapping module, determining the quasi-mapping area of the original coordinate point in the corrected image according to the feature information of the input image and the feature information of the corrected image obtained after the input image is corrected;

[0033] A pixel traversal module, configured to determine the source pixel points in the input image of each pixel point in the quasi-mapping area according to the obtained backward mapping relationship, where the backward mapping relationship is the mapping relationship from the pixel points in the corrected image to the pixel points in the input image;

[0034] A forward mapping module, configured to select the target coordinate points corresponding to the original coordinate points in the quasi-mapping area according to the distance between each source pixel point and the original coordinate points, where the target coordinate points are used to determine the forward mapping relationship corresponding to the original coordinate points, and the forward mapping relationship is the mapping relationship from the pixel points in the input image to the pixel points in the corrected image.

[0035] Optionally, when the quasi-mapping module is used to determine the quasi-mapping area of the original coordinate point in the corrected image according to the feature information of the input image and the feature information of the corrected image obtained after the input image is corrected, it is specifically used for:

[0036] Map the original coordinate point to the corrected image according to the size of the input image and the size of the corrected image obtained after the input image is corrected, so as to obtain a quasi-mapped coordinate point;

[0037] Take the adjacent area with a preset size around the quasi-mapped coordinate point in the corrected image as the quasi-mapping area of the original coordinate point.

[0038] Optionally, the device further includes a distortion correction module, which is used to perform the following steps:

[0039] Determine the imaging area of the second target object in the input image, and determine the perspective transformation matrix corresponding to the imaging area;

[0040] Perform perspective transformation correction on the imaging area based on the perspective transformation matrix to obtain a preprocessed image;

[0041] Determine the deformation field corresponding to the preprocessed image, where the deformation field represents the mapping relationship from the pixel points in the corrected image to the pixel points in the input image;

[0042] Process the deformation field corresponding to the preprocessed image according to the perspective transformation matrix to obtain the backward mapping relationship.

[0043] Optionally, when the distortion correction module is used to process the deformation field corresponding to the preprocessed image according to the perspective transformation matrix to obtain the backward mapping relationship, it is specifically used for:

[0044] Convert the coordinates of each point in the deformation field corresponding to the preprocessed image in the coordinate system of the preprocessed image into homogeneous coordinates, and update the deformation field according to the result of multiplying each homogeneous coordinate by the inverse matrix of the perspective transformation matrix, so as to obtain the deformation field corresponding to the input image.

[0045] Optionally, when the forward mapping module is used to select the target coordinate point corresponding to the original coordinate point in the quasi-mapping area according to the distance between each source pixel point and the original coordinate point, it is specifically used for:

[0046] Sort the source pixel points in the input image in ascending order of the distance between each source pixel point and the original coordinate point to obtain a sequence of source pixel points;

[0047] Select the Nth source pixel point from the sequence of source pixel points, and determine the pixel point corresponding to the Nth source pixel point in the obtained corrected image as the target coordinate point corresponding to the original coordinate point.

[0048] Optionally, when the coordinate acquisition module is used to acquire the original coordinate points in the input image, it is specifically used for:

[0049] Determine the imaging position of the first target object in the input image based on the target detection algorithm;

[0050] Determine at least one original coordinate point in the input image according to the imaging positions of each first target object in the input image.

[0051] Optionally, the device further includes:

[0052] A processing module, configured to obtain the image within a preset range around the target coordinate point in the corrected image, perform preset processing on the image within the preset range, and present the processing result to the user;

[0053] And / or, when the quasi-mapping module is used to determine the quasi-mapping area of the original coordinate point in the corrected image according to the feature information of the input image and the corrected image, it is specifically used for:

[0054] If the forward mapping relationship corresponding to the original coordinate point has not been obtained yet, determine the quasi-mapping area of the original coordinate point in the corrected image according to the feature information of the input image and the corrected image.

[0055] A third aspect of the present disclosure provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the computer program, the method described in the first aspect is implemented.

[0056] A fourth aspect of the present disclosure provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the method described in the first aspect is implemented.

[0057] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects:

[0058] In the embodiments of the present disclosure, for each original coordinate point, a plurality of quasi-mapped coordinate points corresponding to the original coordinate point in the coordinate system of the corrected image are determined, and then the first preset number of quasi-mapped coordinate points with the smallest offset distance among these quasi-mapped coordinate points are used as the target coordinate points corresponding to the original coordinate point in the coordinate system of the corrected image. This implementation method solves the problem that the backward mapping method cannot support the conversion from the input image coordinate system to the corrected image coordinate system. Without significantly increasing the execution overhead of the method, it not only avoids the phenomena of pixel loss and pixel overlap caused by forward mapping, but also realizes the conversion of coordinates from the input image coordinate system to the corrected image coordinate system, and can support the implementation of more functional services.

[0059] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] The accompanying drawings herein are incorporated into the specification and constitute a part of the present disclosure, showing embodiments consistent with the present disclosure and, together with the specification, are used to explain the principles of the present disclosure.

[0061] Figure 1 It is a flowchart of an image processing method shown in some exemplary embodiments.

[0062] Figure 2 It is a schematic diagram of an image processing method shown in some exemplary embodiments.

[0063] Figure 3 It is a schematic diagram of another image processing method shown in some exemplary embodiments.

[0064] Figure 4 It is a schematic diagram of yet another image processing method shown in some exemplary embodiments.

[0065] Figure 5 It is a block diagram of an image processing apparatus shown in some exemplary embodiments.

[0066] Figure 6 It is a hardware structure diagram of an electronic device shown in some exemplary embodiments. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0067] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0068] As described in the background art, when taking pictures of document materials, various distortions may occur in the images captured by the camera. If it is desired to improve the accuracy of the subsequent optical character recognition process (or for other requirements), this kind of distortion can be corrected to obtain a corrected image with less distortion, and then delivered for subsequent processing.

[0069] The present disclosure proposes a new technical scenario. That is, an electronic device can be equipped or connected with a camera. A user can place physical materials (such as printed or handwritten exercise books, textbooks, workbooks, and mobile phones, tablets, e-books with display functions, etc.) within the field of view of the camera, and indicate one or more positions (hereinafter taking one as an example) on the physical materials with a finger (or a knuckle, a point-reading pen, a pencil, etc.). And the electronic device can correct the captured image and present the position indicated by the user in the corrected image (this position should be consistent with the user's intuition. In other words, if in the input image the user's right index finger points to the word "photo", even if the coordinates of the "photo" word in the whole image change after image correction, the coordinates of the user's right index finger should still point to this word), or based on the corrected image and the position indicated by the user, to realize functions such as "point when you don't understand, and there will be an analysis at the pointed place" during the user's learning process.

[0070] A special feature of the above process is that the user's finger position is often abstracted as one or several coordinate points in the input image coordinate system and does not need to participate in the image correction process. Therefore, the process of "mapping the user's finger position to the corrected image" is often carried out after or simultaneously with the image correction process. And in the process of image correction, the distortion can generally be divided into two categories: linear perspective distortion and non-linear geometric distortion.

[0071] Perspective distortion usually refers to the change in the shape and size ratio of a three-dimensional object when it is projected onto a two-dimensional plane. In the above scenario, this kind of transformation is usually caused by the tilt of the paper document or the tilt of the camera angle. Since this transformation is linear, perspective distortion can usually be corrected by a perspective transformation matrix (i.e., Perspective Transformation Matrix, hereinafter represented by the letter M) (hereinafter referred to as "perspective correction"). For example, after taking the value of a pixel point and converting it into homogeneous form, multiplying it by M can obtain the value of the corrected pixel point.

[0072] Geometric Distortion is a non - linear distortion in camera imaging caused by the limitations of the camera's own physical characteristics (for example, the imaging process of the camera may cause barrel distortion and pincushion distortion), or the geometric shape of the object being photographed is distorted (such as a piece of paper being curled or bent, and this kind of distortion is usually three - dimensional). For example, after the paper is bent, the text printed on it undergoes an arched distortion. Since it involves a non - linear deformation recovery process, the correction performed for geometric distortion (hereinafter referred to as "geometric correction") usually cannot be completed by a single perspective transformation matrix. Instead, a preset number of pixel points (usually a single pixel point) need to be used as a unit of a transformation relationship. That is, different pixel points may correspond to different offset methods, rather than corresponding to a unified transformation matrix.

[0073] In geometric correction, the transformation relationships corresponding to each pixel point in the image can be combined into a deformation field (i.e., Deformation Field, hereinafter represented by the letter D). The deformation field can be a vector matrix. Each position in this vector matrix corresponds to a pixel point in the image, and the vector at this position can characterize the deformation method of this pixel point. In other words, the deformation field can characterize the mapping relationship between different images. Through this mapping relationship, the corrected image (i.e., the image after correction) can be obtained from the input image (i.e., the image before correction, which can be an image captured by a camera). The optional mapping methods to achieve this mapping relationship usually include forward mapping (i.e., Forward Mapping, also known as forward mapping) and backward mapping (i.e., Backward Mapping, or Inverse Mapping, also known as reverse mapping, inverse mapping, image filling mapping). For the convenience of elaboration, hereinafter the input image will be abbreviated as α, the corrected image will be abbreviated as β. If P is used to represent a coordinate point, then P α indicates that this coordinate point is the coordinate point in the coordinate system of the input image, and P β indicates that this coordinate point is the coordinate point in the coordinate system of the output image, and P α and P β can be the same coordinate point in different coordinate systems. For example, if P α refers to the position of the user's right index finger in the input image, then P β can refer to the position of the user's right index finger in the corrected image.

[0074] In "forward mapping", each position of the deformation field can respectively correspond to a pixel point in α, and the vector at that position can be used to record the offset method of that pixel point (for example, the offset direction and offset amount can be recorded. After moving the coordinates of that pixel point according to this, it can be filled into the corresponding coordinate position in β), or directly record which coordinate position in β that pixel point should be filled into. In "backward mapping", each position of the deformation field can respectively correspond to a pixel point in β, and the vector at that position can be used to record from which coordinate position in α the pixel value at that position in β can be taken (the recording method can be the same as above).

[0075] It can be seen that the mapping method of "backward mapping" can avoid the problems that some pixel points in β cannot be filled and some pixel points overlap in "forward mapping", thereby avoiding information loss and visual perception decline caused during the interpolation algorithm and blurring process. However, in the mapping method of "backward mapping", the deformation field only has a mapping relationship from β to α (this relationship is not strictly one-to-one at the pixel point level), lacking a mapping relationship from α to β (the mapping relationship from β to α is based on β. Even if all existing β→α relationship pairs are reversed, it is far from covering all pixel points in α). But the positions indicated by the user may appear at any position in the input image, and related technologies cannot achieve this mapping.

[0076] In view of this, the present disclosure provides an image processing method, apparatus, electronic device, and storage medium. Next, the embodiments of the present disclosure will be described in detail.

[0077] The first aspect of the present disclosure provides an image processing method. Please refer to Figure 1 which may include steps S101 to S104.

[0078] Step S101, obtain the original coordinate points in the input image, where the original coordinate points represent the imaging positions of the first target object in the input image.

[0079] The original coordinate points can be understood as the coordinate points in the input image that have the "need to be mapped to the corrected image". This mapping is a mapping of coordinate values and does not necessarily involve the mapping of pixel values. The corrected image is obtained after the input image undergoes image processing. When the method is executed, the corrected image does not necessarily have been generated completely, but its coordinate system, and the pixel point mapping relationship from the corrected image to the input image can be known (which will be introduced later).

[0080] The original coordinate points can be the coordinate points indicated by any algorithm or application software, or the coordinate points indicated by the user. The method for the user to indicate the coordinate points can be by using a body part (such as a fingertip), or by using a pen tip (including a point reading pen) and any other object. Exemplarily, obtaining the original coordinate points in the input image includes: determining the imaging position of the first target object in the input image based on a target detection algorithm; and determining at least one original coordinate point in the input image according to the imaging position of each first target object in the input image.

[0081] Among them, the target detection algorithm can be the YOLO (You Only Look Once) target detection algorithm or the SSD (Single Shot MultiBox Detector) algorithm, etc. When there are multiple original coordinate points, the method can be executed separately for each coordinate point, without affecting the original technical effect of the method.

[0082] Further, if the forward mapping relationship corresponding to the original coordinate points (the forward mapping relationship is the mapping relationship from the pixel points in the input image to the pixel points in the corrected image) is not obtained currently (does not exist currently), then steps S102 to S104 can be executed to obtain the forward mapping relationship corresponding to the original coordinate points (this mapping relationship can be saved or directly applied), and the purpose of the method can be to map the original coordinate points to the corrected image (the mapped coordinate points are called target coordinate points in the present disclosure).

[0083] Step S102: Determine the quasi-mapping region of the original coordinate points in the corrected image according to the feature information of the input image and the feature information of the corrected image obtained after correcting the input image.

[0084] Among them, the quasi-mapping region can be any region containing multiple pixel points, and this region can be continuous or discontinuous (being several discrete pixel points). A part of the pixel points can be randomly selected in the corrected image (or all pixel points can also be selected) to form the quasi-mapping region; or this region can be given by an algorithm (such as the algorithm for perspective transformation correction in the following text) or a large model (such as the large model for outputting the deformation field in the following text); or the quasi-mapping region can also be determined according to certain logical assumptions.

[0085] For example, the "feature information" in S102 can be the size ratio (or scaling ratio) of the input image to the corrected image, a known coordinate system offset (for example, the input image uses the lower left corner as the coordinate origin, but the corrected image uses the center of the image as the coordinate origin), etc. (hereinafter, the resolution ratio of the input image and the corrected image is 1:1, and the coordinate system has not been offset. This is because, since the corrected image is obtained by image processing the input image, it can be assumed that in most of the actual application processes of the user, the curling deformation of the book photographed by the camera module is usually limited. In other words, the extreme case of "the pixel point in the lower right corner of the input image is corrected to the upper left corner of the corrected image" is very rare. Therefore, it can be considered that the optimal mapping position of a pixel point in the input image in the corrected image should not deviate too far from its original coordinates. Therefore, according to the feature information of the input image and the feature information of the corrected image obtained after the input image is corrected, the proposed mapping area of ​​the original coordinate point is determined in the corrected image, including: according to the size of the input image and the size of the corrected image obtained after the input image is corrected (that is, according to the size ratio between the input image and the corrected image), the original coordinate point is mapped to the corrected image to obtain the proposed mapping coordinate point.

[0086] The unit of "size" can be the number of pixels, and the size ratio is the resolution ratio, which also reflects the scaling ratio (scaling means changing the resolution of the image without changing the aspect ratio); and the "adjacent area" can be the neighborhood of the coordinate point to be mapped, or it can be any area covering the coordinate point to be mapped (the shape and size can be preset or determined in real time, for example, it can be given by the algorithm or model as above), especially when the coordinate point to be mapped is located at the edge of the image, this area can be slightly offset toward the center of the image to avoid the situation where the target coordinate point cannot be found as much as possible. In other words, when the resolution ratio of the input image to the corrected image is 1:1 and the coordinate system is not offset, please refer to Figure 2 , if the original coordinate point is P α1 (100,100), then we can take the direct mapping P in the rectified image coordinate system β1 The coordinate point (100,100) is used as the pseudo-mapping coordinate point of the original coordinate point, and then a circle with a radius of 10 pixels and the pseudo-mapping coordinate point as the center is used as the pseudo-mapping area (the default coordinate system unit here is also the number of pixels).

[0087] It should be understood that the quasi-mapping area is a logical concept. In the actual execution process, this "area" can be skipped and the pixel points therein can be directly determined. For example, in the above example, the distance between each pixel point and the quasi-mapping coordinate point can be directly judged, and the pixel points with a distance less than 10 pixel points are used as the pixel points in step S103.

[0088] Step S103: According to the obtained backward mapping relationship, determine the source pixel points of each pixel point in the above quasi-mapping area in the above input image, where the backward mapping relationship is the mapping relationship from the pixel points in the above corrected image to the pixel points in the above input image.

[0089] As mentioned above, the mapping relationship of backward mapping (the deformation field D mentioned above can be used to characterize this mapping relationship, and this will be used as an example for elaboration later) is a mapping method with better effect in the process of correcting the input image to the corrected image. Before performing the above steps, this mapping method has been determined. That is to say, it is known which pixel point in the input image should fill each pixel point in the corrected image (if P in the corrected image β2 is filled by P in the input image α2 , then P α2 will be called the source pixel point of P β2 later).

[0090] The following introduces an exemplary acquisition method of the deformation field of backward mapping. First, the deformation field can be obtained based on algorithms such as cubic polynomial deformation technology, bilinear interpolation method, or models such as DocUNet (Document Unwrapping Network) and STN (Spatial Transformer Network). These algorithms and large models can be offline, so as to better protect user privacy information. The deformation field can be obtained by directly inputting the input image into the algorithm or large model, or by first preprocessing the input image to obtain a better preprocessed image and then inputting it into the algorithm or large model.

[0091] For example, the above method may further include: determining the imaging area of the second target object in the above input image, and determining the perspective transformation matrix corresponding to the imaging area; performing perspective transformation correction on the imaging area based on the perspective transformation matrix to obtain a preprocessed image; determining the deformation field corresponding to the preprocessed image, where the deformation field represents the mapping relationship from the pixel points in the above corrected image to the pixel points in the above input image; and processing the deformation field corresponding to the preprocessed image according to the perspective transformation matrix to obtain the above backward mapping relationship.

[0092] Among them, the second target object can be printed materials such as documents (e.g., textbooks, workbooks), or it can be a tablet computer, etc. Specifically, it can depend on the application scenario of the product and the specific settings of the user. For example, the user can specify what the second target object is, and then the method can call the algorithm or model corresponding to the second target object. Please refer to Figure 3 , in the input image shown therein, the second target object is a printed material, and the main body area (i.e., the imaging area of the second target object) is as Figure 3 shown. This main body area can be determined by algorithms such as YOLO, SSD, or models such as MobileNetV3, EfficientDet. For example Figure 3 , for higher processing efficiency, after inputting the input image into an offline MobileNetV3 model, the model can output 4 vertices of the main body area (i.e., the main body vertices; if the device computing power permits, the number of vertices can be more), and then based on these 4 vertices, the imaging area in the input image can be determined (the polygon formed by the 4 vertices can be used as the main body area).

[0093] Next, perspective correction can be performed on the main body area. For example, the 4 vertices of the input image and the 4 main body vertices of the main body area (a total of 4 pairs of coordinate points) can be input into the perspective transformation algorithm of OpenCV (Open Source Computer Vision Library) or the HomographyNet (Homography Neural Network) model, so as to obtain the perspective transformation matrix M corresponding to the main body area. Next, an empty preprocessed image can be set, and then each pixel point coordinate of the preprocessed image is multiplied by M (the specific method is not limited), so as to obtain the pixel point that can be used to fill this pixel point in the input image (the mapping method shown here is the reverse mapping method, and the process of determining the preprocessed image using the forward mapping method is also feasible, which will not be elaborated here), and finally the preprocessed image (i.e., the image after cropping and perspective transformation correction, as Figure 3 shown) is obtained. It should be noted that more restrictions can be added to this process. For example, it can be restricted to only use the pixel points in the main body area to fill the pixel points in the preprocessed image. If the result after multiplication exceeds the range of the main body area, an interpolation algorithm can be used to determine the pixel point value, or if the pixel point is located at the edge of the image, the pixel point value can be directly made vacant (or filled with a preset pixel point value), which will not be elaborated here.

[0094] The preprocessed image is generated based on the main body region, and basically completes the cropping of the main body region (for example, invalid regions such as the desktop usually no longer exist in the preprocessed image at this time) and the perspective correction of the main body region. However, there is usually still geometric distortion in the preprocessed image at this time. Therefore, the preprocessed image can be continuously input into the above-mentioned algorithm or model to obtain the deformation field corresponding to the preprocessed image. Similarly, precisely because the preprocessed image has basically completed the cropping and perspective correction work, the model is naturally no longer easily affected by invalid regions and perspective deformations. Therefore, the accuracy and correction effect of the deformation field will far exceed the effect of directly inputting the input image into the geometric correction model.

[0095] However, since the image input into the model is actually the preprocessed image, the mapping relationship in the deformation field output by the model is actually from the corrected image to the preprocessed image. However, the processing process from the input image to the preprocessed image may very likely result in the loss of valid pixel points due to the existence of various errors (for example, please refer to Figure 3 , in its main body region, obviously a part of the pixel points exceed the range of the preprocessed image, so these pixel points are not included in the preprocessed image, but these pixel points actually should belong to the main body region). Therefore, although the corrected image can be obtained by directly applying this deformation field, the accuracy is not high enough. Therefore, in the above exemplary embodiment, according to M or the inverse matrix M -1 of M, the deformation field is further processed to correct the coordinates in the preprocessed image coordinate system back to the coordinates in the input image coordinate system, and at this time the required backward mapping relationship is obtained.

[0096] Specifically, if M -1Characterize the conversion relationship between the pixel points in the preprocessed image and the pixel points in the input image. Then, the deformation field corresponding to the preprocessed image can be processed according to the above perspective transformation matrix to obtain the above backward mapping relationship, which may include: converting the coordinates of each point in the deformation field corresponding to the preprocessed image under the coordinate system of the preprocessed image into homogeneous coordinates, and updating the deformation field according to the result of multiplying each homogeneous coordinate by the inverse matrix of the above perspective transformation matrix to obtain the deformation field corresponding to the input image (for example, replacing the value at the corresponding position in the deformation field with the directly obtained result). If M characterizes the conversion relationship between the pixel points in the preprocessed image and the pixel points in the input image, then the deformation field can be updated according to the result of multiplying each homogeneous coordinate by M. It can be seen that this method not only benefits from the advantages of less interference and less distortion in the preprocessed image but also takes into account the advantage of rich information in the input image. Even if some pixel points are cropped in the preprocessed image (at this time, the pixel values or coordinate values of the pixels input to the deformation field may be negative, indicating that the pixel points are invalid), the pixel information of these pixel points can still be used normally in the deformation field obtained in the above steps (after correction, these negative values may no longer be negative but become valid coordinate values in the input image).

[0097] Step S104: Select, according to the distance between each source pixel point and the above original coordinate point, the target coordinate point corresponding to the above original coordinate point in the above quasi-mapping area, where the above target coordinate point is used to determine the forward mapping relationship corresponding to the above original coordinate point.

[0098] After obtaining the quasi-mapping area, the pixel points in this area can be traversed, and the source pixel point of each pixel point can be determined. Please refer to Figure 4 , which takes the pixel points in the neighborhood of the quasi-source pixel point P β as the quasi-mapping area (where the shaded pixel points are P β ). At this time, the source pixel point of each pixel point (i.e., P β and pixel points 1 to 8) in this area can be determined. If the source pixel point of a certain pixel point is exactly the original coordinate point, then this pixel point can be determined as the target pixel point of the original coordinate point (if there are multiple, one can be selected, or the average point can be taken, or they can all be used as the target coordinate points). If the source pixel point of no pixel point is the original coordinate point, then the target coordinate point can be determined according to the distance between the source pixel point and the original coordinate point.

[0099] For example, selecting the target coordinate point corresponding to the original coordinate point in the pseudo-mapping area according to the distance between each source pixel point and the original coordinate point may include: sorting the source pixel points in the input image according to the order of the distances between each source pixel point and the original coordinate point from small to large to obtain a source pixel point sequence; selecting the Nth source pixel point from the source pixel point sequence, and determining the pixel point corresponding to the Nth source pixel point in the corrected image as the target coordinate point corresponding to the original coordinate point (the sequence can be sorted from small to large or from large to small, as long as it is ensured that the finally selected source pixel is not the one with the farthest distance from the original coordinate point). In other words, the pixel point corresponding to the Nth (N is a positive integer and N is less than the total number of pixel points in the pseudo-mapping area) source pixel point with the closest distance to the original coordinate point can be determined as the target coordinate point corresponding to the original coordinate point. When N = 1, the pixel point with the closest distance from the source pixel point to the original coordinate point is taken as the target coordinate point of the original coordinate point.

[0100] Logically speaking, the target coordinate point is used to determine the forward mapping relationship corresponding to the original coordinate point. Of course, this is only to describe the relationship between the target coordinate point and the original coordinate point by defining its use, and it is not necessarily an actual execution step; if the business does not need to record the forward mapping relationship of the current frame image (i.e., the input image), but only needs to find the target coordinate points corresponding to one or more original coordinate points in the corrected image, then step S104 has achieved this goal.

[0101] After obtaining the target coordinate point corresponding to the original coordinate point, the target coordinate point can be used. For example, the above method may further include: obtaining the image within a preset range around the target coordinate point in the corrected image, performing a preset process on the image within the preset range, and presenting the processing result to the user.

[0102] Among them, the preset range can be obtained by identifying the main object around the target coordinate point based on any of the above algorithms (identifying what the main object here is, where the boundary or vertex is), or it can be an adjacent area with a preset size around the target coordinate point (defined as above). Or, the main object can be a preset (or user-specified) third target object, such as a word (English word, Chinese idiom, etc.), a question (an exercise question), etc. And the preset processing can be to increase the brightness, draw a bounding box (so that children, parents or teachers can more easily see the area indicated by the current fingertip position); it can also be to send the area of the main object at this position to another client; it can also be to read out the pronunciation of the main object at this position; it can also be to perform an answering process for the main object at this position, such as translating a foreign language word in the area of the main object, giving an explanation of the word in the area of the main object (such as the origin, meaning, and usage of a Chinese idiom), giving the answering idea or the correct answer to the question in the area of the main object.

[0103] It can be seen that based on the above method, not only can "fingertip dynamic answering" be implemented on educational products such as learning machines, but also the accuracy is extremely high, the tilted and curled printed materials can be corrected, and in the corrected image, the position of the user's finger can still be known. It should be noted that the user's finger itself does not belong to the printed material, and the distortion correction is not performed on the user's finger. Therefore, the recognition work for the user's finger should be for the input image (rather than the corrected image) to ensure higher accuracy and higher compatibility; while functions such as text recognition should be based on the corrected image (rather than the input image); the solution provided by the present disclosure realizes putting the results obtained from the above two processes into the same coordinate system (i.e., the corrected image coordinate system), making them have the value of mutual reference and call. In addition to the learning machine scenario, it has applications in many image processing methods.

[0104] In summary, for each original coordinate point in the embodiments of the present disclosure, a plurality of quasi-mapped coordinate points corresponding to the original coordinate point in the coordinate system of the above-mentioned corrected image are determined, and then the preset number of quasi-mapped coordinate points with the smallest offset distance among these quasi-mapped coordinate points are used as the target coordinate points corresponding to the original coordinate point in the coordinate system of the above-mentioned corrected image. This implementation method solves the problem that the backward mapping method cannot support the conversion from the input image coordinate system to the corrected image coordinate system. Without significantly increasing the execution overhead of the method, it not only avoids the phenomenon of missing pixel points and overlapping pixel points in the corrected image caused by forward mapping, but also realizes the conversion of the coordinates in the input image coordinate system to the coordinates in the corrected image coordinate system, and can support the implementation of more functional services.

[0105] Corresponding to the embodiments of the foregoing method, the present disclosure also provides embodiments of a device and a terminal to which the device is applied.

[0106] In a second aspect of the present disclosure, an image processing apparatus is provided. Please refer to Figure 5 The above-mentioned apparatus includes:

[0107] A coordinate acquisition module 501, configured to acquire original coordinate points in an input image, where the original coordinate points represent the imaging position of a first target object in the input image;

[0108] An approximate mapping module 502, configured to determine an approximate mapping area of the original coordinate points in the corrected image according to the feature information of the input image and the feature information of the corrected image obtained after the input image is corrected;

[0109] A pixel traversal module 503, configured to determine the source pixel points in the input image of each pixel point in the approximate mapping area according to the obtained backward mapping relationship, where the backward mapping relationship is the mapping relationship from the pixel points in the corrected image to the pixel points in the input image;

[0110] A forward mapping module 504, configured to select target coordinate points corresponding to the original coordinate points in the approximate mapping area according to the distance between each source pixel point and the original coordinate points, where the target coordinate points are used to determine the forward mapping relationship corresponding to the original coordinate points, and the forward mapping relationship is the mapping relationship from the pixel points in the input image to the pixel points in the corrected image.

[0111] Optionally, when the approximate mapping module is configured to determine the approximate mapping area of the original coordinate points in the corrected image according to the feature information of the input image and the feature information of the corrected image obtained after the input image is corrected, it is specifically configured to:

[0112] Map the original coordinate points to the corrected image according to the size of the input image and the size of the corrected image obtained after the input image is corrected, to obtain approximate mapping coordinate points;

[0113] Use the adjacent area with a preset size around the approximate mapping coordinate points in the corrected image as the approximate mapping area of the original coordinate points.

[0114] Optionally, the above-mentioned apparatus further includes a distortion correction module, configured to perform the following steps:

[0115] Determine the imaging area of a second target object in the input image, and determine the perspective transformation matrix corresponding to the imaging area;

[0116] Perform perspective transformation correction on the imaging area based on the perspective transformation matrix to obtain a preprocessed image;

[0117] Determine the deformation field corresponding to the preprocessed image, where the deformation field represents the mapping relationship from the pixel points in the corrected image to the pixel points in the input image;

[0118] Process the deformation field corresponding to the preprocessed image according to the perspective transformation matrix to obtain the backward mapping relationship.

[0119] Optionally, when the distortion correction module is used to process the deformation field corresponding to the preprocessed image according to the perspective transformation matrix to obtain the backward mapping relationship, it is specifically used for:

[0120] Convert the coordinates of each point in the deformation field corresponding to the preprocessed image in the coordinate system of the preprocessed image into homogeneous coordinates, and update the deformation field according to the result of multiplying each homogeneous coordinate by the inverse matrix of the perspective transformation matrix to obtain the deformation field corresponding to the input image.

[0121] Optionally, when the forward mapping module is used to select the target coordinate point corresponding to the original coordinate point in the quasi-mapping area according to the distance between each source pixel point and the original coordinate point, it is specifically used for:

[0122] Sort the source pixel points in the input image according to the distance between each source pixel point and the original coordinate point from small to large to obtain a sequence of source pixel points;

[0123] Select the Nth source pixel point from the sequence of source pixel points, and determine the pixel point corresponding to the Nth source pixel point in the obtained corrected image as the target coordinate point corresponding to the original coordinate point.

[0124] Optionally, when the coordinate acquisition module is used to acquire the original coordinate point in the input image, it is specifically used for:

[0125] Determine the imaging position of the first target object in the input image based on the target detection algorithm;

[0126] Determine at least one original coordinate point in the input image according to the imaging position of each first target object in the input image.

[0127] Optionally, the device further includes:

[0128] A processing module, configured to obtain the image within a preset range around the target coordinate point in the corrected image, perform preset processing on the image within the preset range, and present the processing result to the user.

[0129] The implementation processes of the functions and roles of each module in the device are specifically described in the implementation processes of the corresponding steps in the above method, and will not be elaborated here;

[0130] And / or, when the above-mentioned quasi-mapping module is used to determine the quasi-mapping area of the above-mentioned original coordinate point in the above-mentioned corrected image according to the feature information of the above-mentioned input image and the corrected image, it is specifically used for:

[0131] If the forward mapping relationship corresponding to the above-mentioned original coordinate point has not been obtained currently, then according to the feature information of the above-mentioned input image and the corrected image, determine the quasi-mapping area of the above-mentioned original coordinate point in the above-mentioned corrected image.

[0132] Adaptively, the present disclosure also provides a computer program product, including computer programs / instructions, and when the computer programs / instructions are executed by a processor, the methods provided in the foregoing embodiments are implemented.

[0133] For the device embodiments and the computer program product embodiments, since they basically correspond to the method embodiments, the relevant parts can refer to the partial descriptions of the method embodiments. In addition, the device embodiments described above are only illustrative. The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be located in one place, or may be distributed to multiple network modules.

[0134] Adaptively, some embodiments of the present disclosure provide an electronic device, please refer to Figure 6 , which shows the structure of the electronic device. The electronic device includes a memory and a processor. The memory is used to store computer instructions that can run on the processor, and the processor is used to implement the methods shown in any of the foregoing embodiments when executing the computer instructions.

[0135] Adaptively, the present disclosure also provides a non-transitory computer-readable storage medium including instructions, such as a memory including instructions. The above-mentioned instructions can be executed by an electronic device or a processor of the electronic device to complete the methods shown in any of the foregoing embodiments. For example, the non-transitory computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0136] It should be understood that in some cases, the actions or steps recited in the claims may be performed in a different order from those in the embodiments and still achieve the desired results. Also, the embodiments provided by the present disclosure can be applied independently, or combined and comprehensively applied with each other. The present disclosure aims to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include the common general knowledge or conventional technical means in the technical field not claimed by the present disclosure. In addition, the content provided above is only a preferred embodiment of the present disclosure and is not intended to limit the present disclosure. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present disclosure shall be included within the scope of protection of the present disclosure.

Claims

1. An image processing method, characterized in that, The method includes: Obtaining original coordinate points in the input image, where the original coordinate points represent the imaging positions of a first target object in the input image; Determining a quasi-mapping region of the original coordinate points in the corrected image according to the feature information of the input image and the feature information of the corrected image obtained after the input image is corrected; Determining source pixel points in the input image corresponding to each pixel point in the quasi-mapping region according to the obtained backward mapping relationship, where the backward mapping relationship is the mapping relationship from pixel points in the corrected image to pixel points in the input image; Selecting target coordinate points corresponding to the original coordinate points in the quasi-mapping region according to the distances between each source pixel point and the original coordinate points, where the target coordinate points are used to determine the forward mapping relationship corresponding to the original coordinate points, and the forward mapping relationship is the mapping relationship from pixel points in the input image to pixel points in the corrected image.

2. The image processing method according to claim 1, wherein The determining the quasi-mapping region of the original coordinate points in the corrected image according to the feature information of the input image and the feature information of the corrected image obtained after the input image is corrected includes: Mapping the original coordinate points to the corrected image according to the size of the input image and the size of the corrected image obtained after the input image is corrected to obtain quasi-mapping coordinate points; Taking the adjacent region with a preset size around the quasi-mapping coordinate points in the corrected image as the quasi-mapping region of the original coordinate points.

3. The image processing method according to claim 1, characterized in that, The method further includes: Determining the imaging region of a second target object in the input image and determining the perspective transformation matrix corresponding to the imaging region; Performing perspective transformation correction on the imaging region based on the perspective transformation matrix to obtain a preprocessed image; Determining the deformation field corresponding to the preprocessed image, where the deformation field represents the mapping relationship from pixel points in the corrected image to pixel points in the input image; Processing the deformation field corresponding to the preprocessed image according to the perspective transformation matrix to obtain the backward mapping relationship.

4. The image processing method according to claim 3, wherein The processing the deformation field corresponding to the preprocessed image according to the perspective transformation matrix to obtain the backward mapping relationship includes: Converting the coordinates of each point in the deformation field corresponding to the preprocessed image in the coordinate system of the preprocessed image into homogeneous coordinates, and updating the deformation field according to the results of multiplying each homogeneous coordinate by the inverse matrix of the perspective transformation matrix to obtain the deformation field corresponding to the input image.

5. The image processing method according to claim 1, characterized in that The selecting the target coordinate points corresponding to the original coordinate points in the quasi-mapping region according to the distances between each source pixel point and the original coordinate points includes: Sorting the source pixel points in the input image in ascending order of the distances between each source pixel point and the original coordinate points to obtain a source pixel point sequence; Selecting the Nth source pixel point from the source pixel point sequence, and determining the pixel point corresponding to the Nth source pixel point in the obtained corrected image as the target coordinate point corresponding to the original coordinate points.

6. The image processing method according to claim 1, wherein The obtaining of the original coordinate points in the input image includes: Determining the imaging position of the first target object in the input image based on a target detection algorithm; Determining at least one original coordinate point in the input image according to the imaging position of each first target object in the input image.

7. The image processing method according to claim 1, characterized in that The method further includes: Obtaining the image within a preset range around the target coordinate point in the corrected image, performing preset processing on the image within the preset range, and presenting the processing result to the user; And / or The determining of the quasi-mapping area of the original coordinate point in the corrected image according to the feature information of the input image and the feature information of the corrected image obtained after the input image is corrected includes: If the forward mapping relationship corresponding to the original coordinate point has not been obtained yet, determining the quasi-mapping area of the original coordinate point in the corrected image according to the feature information of the input image and the feature information of the corrected image obtained after the input image is corrected.

8. An image processing apparatus, characterized in that, The device includes: A coordinate acquisition module, configured to obtain the original coordinate points in the input image, where the original coordinate points represent the imaging position of the first target object in the input image; A quasi-mapping module, configured to determine the quasi-mapping area of the original coordinate point in the corrected image according to the feature information of the input image and the feature information of the corrected image obtained after the input image is corrected; A pixel traversal module, configured to determine the source pixel points in the input image of each pixel point in the quasi-mapping area according to the obtained backward mapping relationship, where the backward mapping relationship is the mapping relationship from the pixel points in the corrected image to the pixel points in the input image; A forward mapping module, configured to select the target coordinate points corresponding to the original coordinate points in the quasi-mapping area according to the distance between each source pixel point and the original coordinate points, where the target coordinate points are used to determine the forward mapping relationship corresponding to the original coordinate points, and the forward mapping relationship is the mapping relationship from the pixel points in the input image to the pixel points in the corrected image.

9. An electronic device, characterized in that, Including a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor implements the method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program, when executed by the processor, implements the method according to any one of claims 1 to 7.