Image processing method, dictionary pen, and storage medium

CN115660952BActive Publication Date: 2026-09-11ZHEJIANG MAOJING ARTIFICIAL INTELLIGENCE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211222183.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-08
Publication Date
2026-09-11
Estimated Expiration
2042-10-08

AI Technical Summary

Benefits of technology

[0008] According to a third aspect of the embodiments of this application, a computer storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the image processing method as described above.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115660952B_ABST
    Figure CN115660952B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide an image processing method, a dictionary pen and a storage medium. The image processing method comprises: acquiring a plurality of continuous to-be-recognized text images; determining, as a target region, a region close to a starting side of image acquisition and having an image quality greater than a first image threshold in a last image; determining, as a to-be-stitched region, a region close to the starting side of image acquisition and having an image quality greater than a second image threshold in an adjacent previous image; determining, as a stitching region, a region in the to-be-stitched region having a same pixel distribution as the target region; stitching based on the target region and the stitching region to obtain a stitched image; obtaining a plurality of first to-be-processed characters corresponding to the stitched image and position information of each first to-be-processed character in the stitched image, and second to-be-processed characters corresponding to the last image; and performing deletion processing on the plurality of first to-be-processed characters according to the position information and a relationship between the second to-be-processed characters and the plurality of first to-be-processed characters.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to an image processing method, a dictionary pen, and a storage medium. Background Technology

[0002] Currently, electronic devices with text recognition capabilities are constantly emerging in the field of educational hardware, such as dictionary pens.

[0003] Typically, multiple frames of images are continuously captured from the contents of a book using electronic devices. Each image may contain text fragments from the book. The captured multiple frames can then be recognized by electronic devices to obtain the corresponding text content. Based on the text content, the corresponding original text, translation, interpretation, etc., can be determined.

[0004] However, this requires electronic devices to accurately identify the text included in the captured multi-frame images. Summary of the Invention

[0005] In view of this, embodiments of this application provide an image processing solution to at least partially solve the above-mentioned problems.

[0006] According to a first aspect of the embodiments of this application, an image processing method is provided, comprising: acquiring a series of consecutive frames of text images to be recognized; performing image quality analysis on the last frame of the series of text images to be recognized, and determining a region in the last frame of the last frame that is close to the acquisition start side and has an image quality greater than a first image threshold as a target region; performing image quality analysis on the previous frame of the last frame of the last frame, and determining a region in the previous frame of the next adjacent frame that is close to the acquisition start side and has an image quality greater than a second image threshold as a region to be stitched together, wherein the first image quality threshold is greater than the second image quality threshold; and processing the pixels in the target region... The pixels of the target region are compared with those of the region to be stitched, and a stitching region with the same pixel distribution as the target region is determined from the region to be stitched. Based on the target region and the stitching region, the last frame image and the adjacent previous frame image are stitched together to obtain a stitched image. Multiple first characters to be processed corresponding to the stitched image and the position information of each first character to be processed in the stitched image are obtained, as well as a second character to be processed corresponding to the last frame image. According to the position information and the relationship between the second character to be processed and the multiple first characters to be processed, the multiple first characters to be processed are deleted.

[0007] According to a second aspect of the embodiments of this application, a dictionary pen is provided, comprising: an image acquisition device, a processor, and an output device. The image acquisition device is configured to acquire multiple consecutive frames of text images to be recognized. The processor is configured to perform image quality analysis on the last frame of the multiple frames of text images to be recognized, and determine a region in the last frame of the image that is close to the acquisition start side and has an image quality greater than a first image threshold as a target region; perform image quality analysis on the previous frame of the last frame of the image, and determine a region in the previous frame of the image that is close to the acquisition start side and has an image quality greater than a second image threshold as a region to be stitched, wherein the first image quality threshold is greater than the second image quality threshold; and stitch the pixels in the target region with the region to be stitched. The pixels of the regions are compared to determine a splicing region with the same pixel distribution as the target region. Based on the target region and the splicing region, the last frame image and the adjacent previous frame image are spliced ​​to obtain a spliced ​​image. Multiple first characters to be processed corresponding to the spliced ​​image and the position information of each first character to be processed in the spliced ​​image are obtained, as well as a second character to be processed corresponding to the last frame image. According to the position information and the relationship between the second character to be processed and the multiple first characters to be processed, the multiple first characters to be processed are deleted to obtain processed characters, and the corresponding output content is determined according to the processed characters. The output device is used to output the output content.

[0008] According to a third aspect of the embodiments of this application, a computer storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the image processing method as described above.

[0009] According to the solution provided in this application, multiple consecutive frames of text images to be recognized are acquired; image quality analysis is performed on the last frame of the multiple frames of text images to be recognized, and the region in the last frame that is close to the acquisition start side and has an image quality greater than a first image threshold is determined as the target region; image quality analysis is performed on the previous frame adjacent to the last frame, and the region in the previous frame that is close to the acquisition start side and has an image quality greater than a second image threshold is determined as the region to be stitched, where the first image quality threshold is greater than the second image quality threshold. Therefore, a target region with higher image quality can be determined from the side of the last frame close to the image acquisition start side, and a region to be stitched with higher image quality can be determined from the previous frame adjacent to the last frame. By ensuring that the first image threshold is greater than the second image threshold, the area of ​​the region to be stitched is greater than the area of ​​the target region. Thus, when comparing the pixels in the target region with the pixels in the region to be stitched, a stitching region with the same pixel distribution as the target region can be determined from the region to be stitched. Based on the target region and the stitching region... By stitching the last frame image with the adjacent previous frame image to obtain a stitched image, the last frame image can be preserved relatively completely at the end of the stitched image. Then, multiple first characters to be processed corresponding to the stitched image and their position information within the stitched image, as well as the second characters to be processed corresponding to the last frame image, are obtained. Since the last frame image is preserved relatively completely at the end of the stitched image, the positional matching degree between the first characters to be processed corresponding to the stitched image and the second characters to be processed corresponding to the last frame image is high. This minimizes positional errors caused by character deformation in the image, especially improving the positional error of the first characters to be processed corresponding to the last frame image in the stitched image. This results in smaller errors when deleting the multiple first characters to be processed based on the positional information and the relationship between the second characters to be processed and the multiple first characters to be processed, thus improving accuracy. Furthermore, the solution provided in this application deletes the multiple first characters to be processed based on positional information, so errors will not occur even when the first characters to be processed exist in multiple lines. Attached Figure Description

[0010] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings.

[0011] Figure 1AA schematic diagram of the structure of a dictionary pen provided in an embodiment of this application;

[0012] Figure 1B for Figure 1A A schematic diagram of an image recognition method in the illustrated embodiment;

[0013] Figure 2A A schematic flowchart of an image processing method provided in an embodiment of this application;

[0014] Figure 2B for Figure 2A A schematic diagram of a scenario example in the illustrated embodiment;

[0015] Figure 3A A flowchart illustrating the steps of an image processing method provided in this application embodiment;

[0016] Figure 3B for Figure 3A A schematic diagram of an image stitching method in the illustrated embodiment;

[0017] Figure 3C for Figure 3A A schematic diagram of a scenario example in the illustrated embodiment;

[0018] Figure 3D for Figure 3A A schematic diagram of another image stitching method in the illustrated embodiment;

[0019] Figure 4 A schematic diagram of the structure of a dictionary pen provided in an embodiment of this application;

[0020] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0021] To enable those skilled in the art to better understand the technical solutions in the embodiments of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art should fall within the protection scope of the embodiments of this application.

[0022] The following describes the application scenario of this application. Referring to Figure 1, a schematic diagram of a dictionary pen collecting data from a book is shown. The dictionary pen includes a pen body and a pen tip 10. A cover plate 11 is provided on the pen tip 10, and a camera (not shown in the figure) is provided on the inner side of the cover plate 11 for collecting the content on the book (such as "whale" shown in Figure 1).

[0023] In practice, users can use the dictionary pen to trace the text in the book according to the order of the words. The pen's camera can quickly record or capture multiple frames of the text. The multiple frames are then stitched together, and text recognition is performed based on the stitched image. The text is then converted to speech and played back, with the corresponding translation (such as English to Chinese) output simultaneously.

[0024] Specifically, when using the dictionary pen to stroke text, taking a left-to-right stroke as an example, the content seen by the user is on the left side of the obscuring panel, while the image captured by the camera is on the right side. When the user lifts the pen, the camera will still capture the content on the right side of the obscuring panel, which is the last frame captured by the camera. This content is after the pen is lifted and needs to be discarded during processing.

[0025] Generally, the last frame of the image can be identified, and the text corresponding to the identification result of the last frame can be discarded from the text identified from multiple frames. However, if the last frame contains multiple lines of text, especially if these lines contain words or phrases that are repeated from the results identified from previous images, erroneous discarding can easily occur. For example, see... Figure 1B When there are multiple lines of text in an image, the dictionary pen generally only recognizes the middle line of text. For example, in the image "I love China.", the first line "ive in China.China" is not in the middle, so it is not recognized. Figure 1B The part after the pen lift shown is the last frame of the image, which only contains the first line "China". That is, the recognition result corresponding to the last frame of the image is "China". If the "China" in the previous recognition results is discarded, the content that is retained is "I love", which means that there is a case of erroneous discarding.

[0026] In addition, multiple frames can be stitched together directly, and the last frame can be extracted from the stitched image. However, due to the limited stitching accuracy, there may be cases of deleting too many or too few frames.

[0027] Therefore, embodiments of this application provide an image processing method to solve or alleviate the above-mentioned problems as much as possible.

[0028] The specific implementation of the embodiments of this application will be further described below with reference to the accompanying drawings.

[0029] Figure 2A A flowchart illustrating an image processing method provided in an embodiment of this application is shown in the figure, which includes:

[0030] S201. Obtain multiple consecutive frames of text images to be recognized;

[0031] In this embodiment, the electronic device for acquiring the image of the text to be recognized can be a dictionary pen or other devices capable of acquiring images; this embodiment does not limit this.

[0032] When collecting images of the text to be recognized, the user can hold the electronic device and swipe it across the book. The device's camera will then capture multiple frames of the text to be recognized. Of course, the above example using a book is merely illustrative; any object with text can be used as the object to be collected.

[0033] S202. Perform image quality analysis on the last frame of the multi-frame text images to be identified, and determine the region in the last frame that is close to the acquisition start side and has an image quality greater than the first image threshold as the target region.

[0034] By performing image quality analysis on the last frame of the multi-frame text images to be identified, a region with high image quality that is close to the starting side of the acquisition can be selected as the target region, which improves the accuracy of pixel comparison in the subsequent step S204.

[0035] For example, image quality analysis may include, but is not limited to, brightness analysis, contrast analysis, text sharpness analysis, etc., and this embodiment does not limit it.

[0036] S203. Perform image quality analysis on the previous frame image adjacent to the last frame image, and determine the region in the previous frame image that is close to the acquisition start side and has an image quality greater than the second image threshold as the region to be stitched, wherein the first image quality threshold is greater than the second image quality threshold.

[0037] By performing image quality analysis on the previous frame adjacent to the last frame image, a region with high image quality that is close to the acquisition start side can be selected as the region to be stitched, which can also improve the accuracy of pixel comparison in the subsequent step S204.

[0038] In addition, in this embodiment of the application, since the first image quality threshold is greater than the second image quality threshold, the area of ​​the determined region to be stitched is greater than the area of ​​the target region, so as to ensure that when comparing the pixels in the target region with the pixels in the region to be stitched, a stitching region with the same pixel distribution as the target region can be determined from the region to be stitched.

[0039] The first image threshold and the second image threshold can be determined by those skilled in the art according to their needs, as long as it can be ensured that there is a splicing area in the area to be spliced ​​that has the same pixel distribution as the target area.

[0040] S204. Compare the pixels in the target area with the pixels in the area to be stitched, and determine a stitching area in the area to be stitched that has the same pixel distribution as the target area.

[0041] Furthermore, since both the target area and the stitching area are close to the starting side of the acquisition, the last frame can occupy a relatively complete portion of the stitched image after being stitched together with the adjacent previous frame. For specific methods of pixel comparison, please refer to relevant technologies; they will not be elaborated upon here.

[0042] S205. Based on the target region and the stitching region, the last frame image and the adjacent previous frame image are stitched together to obtain a stitched image.

[0043] For the last frame of a multi-frame image of text to be recognized, the pixels of the target area on the starting side of the last frame can be matched with the stitching area of ​​the adjacent previous frame to determine the stitching area that matches the pixels of the target area in the adjacent previous frame, and the image can be stitched accordingly.

[0044] For example, taking the acquisition of multiple frames of text images to be recognized from left to right by a dictionary pen, for the last frame of the multiple frames of text images to be recognized, the splicing area can be determined based on the correspondence between the pixels of the target area on the left side of the last frame and the pixels of the splicing area of ​​the adjacent previous frame. The target area and at least half of the right side of the last frame are spliced ​​with the adjacent previous frame, so that the last frame can be relatively completely preserved on the right side of the spliced ​​image.

[0045] Additionally, it should be noted that in this embodiment, multiple frames of text images to be recognized can be stitched together sequentially, and step S202 can be executed when the last frame is stitched together; alternatively, step S202 can be executed directly without stitching in sequence. In this case, for any two adjacent frames other than the last frame, the pixels of any one frame can be compared with the adjacent previous frame to determine the area with the same pixel distribution in the two frames. After determining the area with the same pixel distribution in all any two adjacent frames, multiple consecutive frames can be stitched together to obtain a stitched image.

[0046] The method for stitching images can be found in related technologies, and this embodiment does not limit it.

[0047] S206. Obtain a plurality of first characters to be processed corresponding to the spliced ​​image and the position information of each first character to be processed in the spliced ​​image, as well as a second character to be processed corresponding to the last frame image;

[0048] In this embodiment, by performing text recognition on the spliced ​​image, multiple first characters to be processed corresponding to the spliced ​​image can be obtained, as well as the position information of each first character to be processed in the spliced ​​image. For example, the offset distance of each first character to be processed relative to the left edge of the spliced ​​image can be obtained.

[0049] In this embodiment, by performing text recognition on the last frame image, the second character to be processed corresponding to the last frame image can be obtained.

[0050] For specific methods of text recognition, please refer to relevant technologies, which will not be elaborated here.

[0051] S207. Based on the location information and the relationship between the second character to be processed and the plurality of first characters to be processed, delete the plurality of first characters to be processed.

[0052] Since in steps S202-S205 above, when stitching the last frame image, it can be ensured that the last frame image can be retained relatively completely at the end of the stitched image, the position matching degree between the first character to be processed corresponding to the stitched image and the second character to be processed corresponding to the last frame image is relatively high. Therefore, the position error caused by the deformation of characters in the image can be minimized, and the accuracy is improved. Furthermore, the solution provided in this application performs deletion processing on the multiple first characters to be processed based on position information, so no error will occur when the first characters to be processed exist in multiple lines.

[0053] For example, in this embodiment, if the last frame image corresponds to a second character to be processed, the first character to be processed corresponding to the last frame image can be deleted based on the position information of each first character to be processed in the concatenated characters, thus avoiding errors introduced by other matching and improving accuracy.

[0054] The following example illustrates the solution of this application through a specific implementation scenario.

[0055] See Figure 2B It shows a series of consecutive frames of text images to be recognized, which can be acquired in order from left to right;

[0056] By stitching together multiple frames of text images to be recognized, a stitched image can be obtained. Specifically, when stitching the last frame, the target area on the left side of the last frame can be stitched together with the stitching area near the left side of the adjacent previous frame. This ensures that the last frame occupies a relatively complete portion of the stitched image, and the area corresponding to the last frame in the stitched image is... Figure 2B The part inside the wireframe.

[0057] Then, text recognition can be performed on the stitched image and the last frame image to obtain multiple first characters to be processed corresponding to the stitched image, as well as the position information of each first character to be processed in the stitched image, and the second character to be processed corresponding to the last frame image. Specifically, the position information is the offset of the first character to be processed from the left edge of the stitched image. Figure 2B The example shows the position information of some of the first characters to be processed in the spliced ​​image. For example, the position information of the character "ina.ch" is 285, 315, 351, 381, 424, and 463 respectively.

[0058] By performing text recognition on the last frame of the image, it can be determined whether there is a corresponding character to be processed in the last frame. If there is a corresponding character, the first character to be processed corresponding to the position information of the last frame can be deleted based on the position information. The deleted first character to be processed may include "I love china." After obtaining the deleted first character to be processed, the translation, interpretation, etc. corresponding to the deleted first character to be processed can be determined and displayed.

[0059] The solution provided in this embodiment acquires multiple consecutive frames of text images to be recognized; performs image quality analysis on the last frame of the multiple frames of text images to be recognized, and determines the region in the last frame that is close to the acquisition start side and has an image quality greater than a first image threshold as the target region; performs image quality analysis on the previous frame adjacent to the last frame, and determines the region in the previous frame that is close to the acquisition start side and has an image quality greater than a second image threshold as the region to be stitched, where the first image quality threshold is greater than the second image quality threshold. Therefore, a target region with higher image quality can be determined from the side of the last frame close to the image acquisition start side, and a region to be stitched with higher image quality can be determined from the previous frame adjacent to the last frame. By ensuring that the first image threshold is greater than the second image threshold, the area of ​​the region to be stitched is greater than the area of ​​the target region. Thus, when comparing the pixels in the target region with the pixels in the region to be stitched, a stitching region with the same pixel distribution as the target region can be determined from the region to be stitched. Based on the target region and the stitching region, the text is stitched... The last frame image is stitched together with the adjacent previous frame image to obtain a stitched image. This ensures that the last frame image is relatively intact at the end of the stitched image. Then, multiple first characters to be processed corresponding to the stitched image and their position information within the stitched image are obtained, along with the second characters to be processed corresponding to the last frame image. Since the last frame image is relatively intact at the end of the stitched image, the positional matching degree between the first characters to be processed in the stitched image and the second characters to be processed in the last frame image is high. This minimizes positional errors caused by character deformation in the image, especially improving the positional error of the first characters to be processed corresponding to the last frame image in the stitched image. This results in smaller errors when deleting the multiple first characters to be processed based on the positional information and the relationship between the second characters to be processed and the multiple first characters to be processed, thus improving accuracy. Furthermore, the solution provided in this application deletes the multiple first characters to be processed based on positional information, so errors will not occur even when the first characters to be processed exist in multiple lines.

[0060] The image processing method of this embodiment can be executed by any suitable electronic device with data processing capabilities, including but not limited to: servers, mobile terminals (such as mobile phones, PADs, etc.) and PCs.

[0061] Figure 3A A flowchart of an image processing method provided in this application embodiment is shown in the figure, which includes:

[0062] S301. Acquire multiple consecutive frames of text images to be recognized.

[0063] S302. The second frame image and the penultimate frame image are sequentially determined as images to be stitched together. The target region in the images to be stitched together and the stitching region in the previous frame image adjacent to the images to be stitched together, which has the same pixel distribution as the target region, are determined.

[0064] Optionally, in this embodiment of the application, for any image to be stitched, image quality analysis can be performed on the image to be stitched, and the region near the acquisition start side and whose image quality is greater than a first image threshold can be determined as the target region. Image quality analysis can also be performed on the adjacent previous frame image of the image to be stitched, and the region near the acquisition start side and whose image quality is greater than a second image threshold can be determined as the region to be stitched. The pixels in the target region are compared with the pixels in the region to be stitched, and a stitching region with the same pixel distribution as the target region is determined from the region to be stitched.

[0065] The specific methods for determining the target area and the splicing area can be found in the above embodiments, and will not be repeated here.

[0066] See Figure 3B , Figure 3B The right-hand area in the image is the image to be stitched together, and the left-hand area is the previous frame image adjacent to it. The box in the right-hand area shows a schematic diagram of a target area.

[0067] S303. Perform image quality analysis on the last frame of the multi-frame text images to be identified, and determine the region in the last frame that is close to the acquisition start side and has an image quality greater than the first image threshold as the target region.

[0068] It should be noted that the method for determining the target region in step S302 is the same as the method for determining the target region in the last frame image.

[0069] Optionally, in this embodiment, step S303 may include: performing image quality analysis on at least half of the image region near the acquisition start side in the last frame image; and, based on the image quality analysis results, determining a region with image quality greater than the first image threshold from the at least half of the image region near the acquisition start side in the last frame image as the target region. This allows for the selection of target regions with higher image quality for matching, improving matching accuracy. Furthermore, by selecting at least half of the image region near the acquisition start side for analysis, the resource consumption of image quality analysis is reduced compared to performing image quality analysis on the entire frame image.

[0070] Optionally, in this embodiment, the image quality analysis includes at least one of the following: brightness analysis, deformation analysis, brightness uniformity analysis, contrast analysis, signal-to-noise ratio analysis, and sharpness analysis. Correspondingly, the first image threshold includes at least one of the following: a first brightness threshold, a first deformation threshold, a first brightness uniformity threshold, a first contrast threshold, a first signal-to-noise ratio threshold, and a first resolution threshold; correspondingly, the second image threshold includes at least one of the following: a second brightness threshold, a second deformation threshold, a second brightness uniformity threshold, a second contrast threshold, a second signal-to-noise ratio threshold, and a second resolution threshold. Since the camera is at a certain angle (e.g., 45 degrees) to the surface being captured during the acquisition process using the dictionary pen, the content in the captured image will be deformed. Therefore, in this embodiment, it is preferable to use at least brightness analysis and deformation analysis as image quality analysis, thereby improving the quality of the determined target area.

[0071] Brightness analysis

[0072] Generally, images captured by a dictionary pen are typically white backgrounds with black text. Therefore, when performing brightness analysis, the brightness of the white background in a specific image region is mainly analyzed as the brightness of that image region.

[0073] Deformation analysis

[0074] When performing deformation analysis, the degree of deformation perpendicular to the acquisition direction can be analyzed. Of course, the degree of deformation parallel to the acquisition direction can also be analyzed, which is also within the scope of protection of this application.

[0075] Brightness uniformity analysis

[0076] For a specific image region, the brightness of each pixel in the image region can be collected, the variance of the brightness can be calculated, and the variance result can be used as the brightness uniformity analysis result; alternatively, the deviation between the brightness of each pixel and the highest or lowest brightness in the image region can be calculated, and the deviation result can be used as the brightness uniformity analysis result.

[0077] Contrast Analysis

[0078] For a specific image region, the brightness of each pixel within that region can be collected, and the contrast of the image region can be calculated based on the brightness. For example, the contrast of an image region can be determined by calculating the difference between the average brightness values ​​of different colors.

[0079] Signal-to-noise ratio analysis

[0080] When performing signal-to-noise ratio (SNR) analysis, the dictionary pen's camera can be used to capture images with a single color fill, and the SNR can be calculated for a specific image region based on the captured image. Furthermore, images with different single color fills can be scanned, and the SNR for the same image region can be calculated based on each scanned image to obtain the final SNR analysis result.

[0081] Sharpness analysis

[0082] During sharpness analysis, the dictionary pen's camera can be used to capture images of lines in a single direction, and based on the captured images, it can be determined whether the lines can be distinguished. Furthermore, images of lines in a single direction with different thicknesses or spacing can be scanned, and the distinguishability of each line can be determined based on the scanned images, yielding the final sharpness analysis result.

[0083] In addition, grayscale testing and convergence testing can be performed on dictionary pens.

[0084] Grayscale test

[0085] When conducting grayscale testing, the dictionary pen's camera can be used to capture an image with multiple rows of color boxes. Based on the captured image, it can be determined whether the color boxes can be distinguished to obtain the grayscale test results.

[0086] Convergence test

[0087] During convergence testing, the dictionary pen's camera can be used to capture images with the upper half being black background and white text, and the lower half being white background and black text. The number of stable exposure frames can be determined based on the captured images, and the convergence test results can be obtained based on the number of frames.

[0088] S304. Perform image quality analysis on the previous frame image adjacent to the last frame image, and determine the region in the previous frame image that is close to the acquisition start side and has an image quality greater than the second image threshold as the region to be stitched.

[0089] It should be noted that the method for determining the region to be stitched in step S302 above is the same as the method for determining the region to be stitched in the last frame image.

[0090] S305. Compare the pixels in the target area with the pixels in the area to be stitched, and determine a stitching area in the area to be stitched that has the same pixel distribution as the target area.

[0091] Specifically, in this embodiment, the region to be stitched can be determined from the previous frame image. The area of ​​the region to be stitched is larger than the target region. The specific region to be stitched is as follows: Figure 3BAs shown in the box on the left. Then, a sliding window identical to the target area can be identified within the area to be stitched, and the pixels of the target area and the sliding window can be compared to determine the area matching the target area from the area to be stitched, which will then be used as the stitching area.

[0092] S306. Based on the target area and the stitching area, perform image stitching to obtain the stitched image.

[0093] Optionally, in this embodiment, step S305 may include: stitching the last frame image and the adjacent previous frame image together, using the center line of the target area and the center line of the stitching area as the stitching position. Specifically, the center line can be referenced... Figure 3B The vertical center lines of the stitching area and the target area are shown in the diagram. Since the deformation is generally minimal at the center line, using the center line as the stitching location can improve the image quality of the stitched image.

[0094] Of course, it should be noted that the above method can also be used when splicing two adjacent frames other than the last frame. Of course, other methods can also be used, and this embodiment does not limit this.

[0095] Optionally, before step S306, the method may further include: performing image brightness preprocessing and / or image jump preprocessing on the target region and the stitching region. This ensures that the stitched image does not experience brightness jumps or image content jumps, thus improving the quality of the stitched image. Specific methods for performing image brightness preprocessing and / or image jump preprocessing can be found in related technologies and will not be elaborated upon here.

[0096] S307. Perform text recognition on the spliced ​​image to obtain multiple first characters to be processed corresponding to the spliced ​​image and the position information of each first character to be processed in the spliced ​​image.

[0097] S308. Perform text recognition on the last frame image to obtain the second character to be processed corresponding to the last frame image.

[0098] S309. Determine whether the second character to be processed is empty.

[0099] S310. If the second character to be processed is not an empty character, then according to the position information, delete the character in the first character to be processed that corresponds to the position of the second character to be processed.

[0100] Optionally, in this embodiment of the application, the position information of the first character to be processed in the spliced ​​image includes: the position offset of each first character to be processed relative to the edge of the spliced ​​image at the start of acquisition. Correspondingly, step S309 may include: determining the first image length of the spliced ​​image in the acquisition direction and the second image length of the last frame image in the acquisition direction, calculating the difference between the first image length and the second image length to obtain the distance threshold corresponding to the last frame image; and deleting the first characters to be processed whose position offset is greater than the distance threshold.

[0101] For example, see Figure 3C In this embodiment, when the acquisition direction is from left to right, the distance threshold is calculated as: the total length of the stitched images in the horizontal direction (first image length) - the length of the last frame image in the horizontal direction (second image length). The length can be a pixel value. Of course, if the acquisition direction is from top to bottom or diagonally, the first image length and second image length in the corresponding direction are used for calculation, which is also within the scope of protection of this application.

[0102] S311. If the second character to be processed is an empty character, then all of the plurality of first characters to be processed are retained.

[0103] See Figure 3D The second character to be processed is an empty character, which means that the last frame image may be a blank book. A blank book has less effective information, and its position is very easy to be deviated during splicing, resulting in the inaccurate position of the last frame image. Figure 3D The image above shows multiple consecutive frames of the text to be recognized before it was stitched together. Figure 3D The image in the middle is the desired stitched image. Figure 3D Below is the actual stitched image. Figure 3D The dashed lines in the middle and lower parts of the stitched image correspond to the last frame, as shown below. Figure 3D As shown, the position of the last frame in the actual stitched image is inaccurate. Therefore, in this embodiment, the first character to be processed is deleted only when the second character to be processed is not an empty character; when the second character to be processed is an empty character, the first character to be processed can be completely retained.

[0104] Figure 4 A schematic diagram of the structure of a dictionary pen provided in this application embodiment is shown in the figure. It includes: an image acquisition device 401, a processor 402, and an output device 403. The image acquisition device 401 can be any device with image acquisition function, such as a camera; the output device can include, but is not limited to, at least one of the following: a display and a speaker.

[0105] The image acquisition device 401 is used to acquire multiple consecutive frames of text images to be recognized.

[0106] The processor 402 is configured to perform image quality analysis on the last frame of the multi-frame text images to be recognized, and determine the region in the last frame that is close to the acquisition start side and has an image quality greater than a first image threshold as the target region; perform image quality analysis on the previous frame adjacent to the last frame, and determine the region in the previous frame that is close to the acquisition start side and has an image quality greater than a second image threshold as the region to be stitched, wherein the first image quality threshold is greater than the second image quality threshold; compare the pixels in the target region with the pixels in the region to be stitched, and determine the region from the region to be stitched that matches the target region. The target region is a splicing region with the same pixel distribution; based on the target region and the splicing region, the last frame image and the adjacent previous frame image are spliced ​​together to obtain a spliced ​​image; multiple first characters to be processed corresponding to the spliced ​​image and the position information of each first character to be processed in the spliced ​​image are obtained, as well as a second character to be processed corresponding to the last frame image; according to the position information and the relationship between the second character to be processed and the multiple first characters to be processed, the multiple first characters to be processed are deleted to obtain processed characters, and the corresponding output content is determined according to the processed characters.

[0107] The output device 403 is used to output the output content.

[0108] Optionally, in this embodiment, the dictionary pen further includes a memory for storing a knowledge base. The processor is specifically used to query the knowledge base based on the processed characters and determine the corresponding output content based on the query results. The knowledge base can be any knowledge base, such as a bilingual dictionary or the Three Hundred Tang Poems. This embodiment does not limit this.

[0109] Reference Figure 5 The diagram shows a structural schematic of an electronic device provided in an embodiment of this application. The specific embodiments of this application do not limit the specific implementation of the electronic device.

[0110] like Figure 5 As shown, the electronic device may include: a processor 502, a communications interface 504, a memory 506, a communications bus 508, an image acquisition device 510, and an output device 512.

[0111] in:

[0112] The processor 502, communication interface 504, memory 506, image acquisition device 510, and output device 512 communicate with each other through communication bus 508.

[0113] Image acquisition device 510 is used to acquire multiple consecutive frames of text images to be recognized. Image acquisition device 510 can be a camera, etc.

[0114] Communication interface 504 is used to communicate with other electronic devices or servers.

[0115] The processor 502 is used to execute program 514, which can specifically perform the relevant steps in the above image processing method embodiment for multiple frames of text images to be recognized, and determine the output content.

[0116] The output device 512 is used to output the output content. The output device 512 can be a display or a speaker, etc.

[0117] Specifically, program 514 may include program code that includes computer operation instructions.

[0118] The processor 502 may be a CPU (central processing unit), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application. The one or more processors included in the smart device may be processors of the same type, such as one or more CPUs; or they may be processors of different types, such as one or more CPUs and one or more ASICs.

[0119] Memory 506 is used to store program 514. Memory 506 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.

[0120] The specific implementation of each step in program 514 can be found in the corresponding steps and units described in the above-described image processing method embodiments, and will not be repeated here. Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the devices and modules described above can be referred to the corresponding process descriptions in the aforementioned method embodiments, and will not be repeated here.

[0121] This application also provides a computer storage medium storing a computer program that, when executed by a processor, implements any of the image processing methods described in the above-described method embodiments.

[0122] This application also provides a computer program product, including computer instructions that instruct a computing device to perform an operation corresponding to any of the image processing methods in the above-described multiple method embodiments.

[0123] It should be noted that, depending on the implementation needs, the various components / steps described in the embodiments of this application can be broken down into more components / steps, or two or more components / steps or parts of the operation of components / steps can be combined into new components / steps to achieve the purpose of the embodiments of this application.

[0124] The methods described in the embodiments of this application can be implemented in hardware, firmware, or as software or computer code that can be stored in a recording medium (such as a CD-ROM, RAM, floppy disk, hard disk, or magneto-optical disk), or as computer code downloaded over a network that is originally stored in a remote recording medium or a non-transitory machine-readable medium and will be stored in a local recording medium. Thus, the methods described herein can be stored as software processing on a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an ASIC or FPGA). It is understood that the computer, processor, microprocessor controller, or programmable hardware includes storage components (e.g., RAM, ROM, flash memory, etc.) capable of storing or receiving software or computer code that, when accessed and executed by the computer, processor, or hardware, implements the image processing methods described herein. Furthermore, when a general-purpose computer accesses code used to implement the image processing methods shown herein, the execution of the code transforms the general-purpose computer into a dedicated computer for executing the image processing methods shown herein.

[0125] Those skilled in the art will recognize that the units and method steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the embodiments of this application.

[0126] The above embodiments are only used to illustrate the embodiments of this application, and are not intended to limit the embodiments of this application. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the embodiments of this application. Therefore, all equivalent technical solutions also fall within the scope of the embodiments of this application, and the patent protection scope of the embodiments of this application should be defined by the claims.

Claims

1. An image processing method, comprising: Acquire consecutive multi-frame images of the text to be recognized; Image quality analysis is performed on the last frame of the multi-frame text images to be identified, and the region in the last frame that is close to the acquisition start side and has an image quality greater than the first image threshold is determined as the target region. Image quality analysis is performed on the previous frame image adjacent to the last frame image. The region in the previous frame image that is close to the acquisition start side and has an image quality greater than the second image threshold is determined as the region to be stitched. The first image threshold is greater than the second image threshold. The pixels in the target region are compared with the pixels in the region to be stitched, and a stitching region with the same pixel distribution as the target region is determined from the region to be stitched. Based on the target region and the stitching region, the last frame image and the adjacent previous frame image are stitched together to obtain a stitched image; Obtain multiple first characters to be processed corresponding to the stitched image and the position information of each first character to be processed in the stitched image, as well as the second character to be processed corresponding to the last frame image; Based on the location information and the relationship between the second character to be processed and the plurality of first characters to be processed, the plurality of first characters to be processed are deleted. The position information of the first character to be processed in the spliced ​​image includes: the position offset of each first character to be processed relative to the edge of the spliced ​​image at the acquisition start side. Correspondingly, the deletion process of the plurality of first characters to be processed based on the position information and the relationship between the second character to be processed and the plurality of first characters to be processed includes: If the second character to be processed is not an empty character, then determine the first image length of the spliced ​​image in the acquisition direction and the second image length of the last frame image in the acquisition direction, calculate the difference between the first image length and the second image length, and obtain the distance threshold corresponding to the last frame image; Delete the first character to be processed whose position offset is greater than the distance threshold.

2. The method according to claim 1, wherein, The step of stitching the last frame image and the adjacent previous frame image together based on the target region and the stitching region to obtain the stitched image includes: Using the center line of the target area and the center line of the stitching area as the stitching positions, the last frame image and the adjacent previous frame image are stitched together to obtain the stitched image.

3. The method according to claim 1, wherein, The step of performing image quality analysis on the last frame of the multi-frame text images to be recognized, and determining the region in the last frame image that is close to the acquisition start side and has an image quality greater than a first image threshold as the target region, includes: Image quality analysis is performed on at least half of the image region closest to the acquisition start side in the last frame image; Based on the image quality analysis results, from at least half of the image region near the acquisition start side in the last frame image, the region with image quality greater than the first image threshold is determined as the target region.

4. The method according to claim 1, wherein, The image quality analysis includes at least one of the following: brightness analysis, deformation analysis, brightness uniformity analysis, contrast analysis, signal-to-noise ratio analysis, and sharpness analysis; Correspondingly, the first image threshold includes at least one of the following: a first brightness threshold, a first deformation threshold, a first brightness uniformity threshold, a first contrast threshold, a first signal-to-noise ratio threshold, and a first resolution threshold; Correspondingly, the second image threshold includes at least one of the following: a second brightness threshold, a second deformation threshold, a second brightness uniformity threshold, a second contrast threshold, a second signal-to-noise ratio threshold, and a second resolution threshold.

5. The method according to claim 1, wherein, The step of deleting the plurality of first characters to be processed based on the location information and the relationship between the second character to be processed and the plurality of first characters to be processed further includes: If the second character to be processed is an empty character, then all of the multiple first characters to be processed will be retained.

6. The method according to any one of claims 1-5, wherein, Before stitching the last frame image and the adjacent previous frame image together to obtain the stitched image based on the target region and the stitching region, the method further includes: performing image brightness preprocessing and / or image transition preprocessing on the target region and the stitching region.

7. A dictionary pen, comprising: Image acquisition device, processor, output device, The image acquisition device is used to acquire multiple consecutive frames of text images to be recognized; The processor is configured to: perform image quality analysis on the last frame of the multi-frame text image to be recognized; determine the region in the last frame that is close to the acquisition start side and has an image quality greater than a first image threshold as the target region; perform image quality analysis on the previous frame adjacent to the last frame; determine the region in the previous frame that is close to the acquisition start side and has an image quality greater than a second image threshold as the region to be stitched, wherein the first image threshold is greater than the second image threshold; compare the pixels in the target region with the pixels in the region to be stitched, and determine the stitching region in the region to be stitched that has the same pixel distribution as the target region; stitch the last frame and the previous frame based on the target region and the stitching region to obtain a stitched image; obtain multiple first characters to be processed corresponding to the stitched image and the position information of each first character to be processed in the stitched image, and a second character to be processed corresponding to the last frame; delete the multiple first characters to be processed according to the position information and the relationship between the second character to be processed and the multiple first characters to be processed to obtain processed characters, and determine the corresponding output content according to the processed characters; The position information of the first character to be processed in the spliced ​​image includes: the position offset of each first character to be processed relative to the edge of the spliced ​​image at the start of acquisition. Correspondingly, the deletion process of the plurality of first characters to be processed based on the position information and the relationship between the second character to be processed and the plurality of first characters to be processed includes: if the second character to be processed is not an empty character, determining the first image length of the spliced ​​image in the acquisition direction and the second image length of the last frame image in the acquisition direction, calculating the difference between the first image length and the second image length to obtain the distance threshold corresponding to the last frame image; deleting the first characters to be processed whose position offset is greater than the distance threshold. The output device is used to output the output content.

8. The dictionary pen according to claim 7 further includes a memory for storing a knowledge base, wherein the processor is specifically used to query the knowledge base according to the processed characters and determine the corresponding output content according to the query result.

9. The dictionary pen according to claim 7, wherein, The processor is specifically used to stitch the last frame image and the adjacent previous frame image together, using the center line of the target area and the center line of the stitching area as the stitching positions, to obtain the stitched image.

10. The dictionary pen according to claim 7, wherein, The processor is specifically used to retain all of the plurality of first characters to be processed if the second character to be processed is an empty character.

11. A computer storage medium having a computer program stored thereon, which, when executed by a processor, implements the method as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Character recognition method and device, equipment, storage medium and intelligent dictionary pen

    CN113642584A

  • Image processing method and image processing device

    CN114078245A