Scanning image recognition method and device, electronic equipment and storage medium
By identifying valid boundary image frames in the scanning device and segmenting them according to the validity of the boundary characters, the problem of overscanning or underscanning of text in the scanning device is solved, thus improving the accuracy of the scanned images and recognition results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- IFLYTEK CO LTD
- Filing Date
- 2022-10-31
- Publication Date
- 2026-04-21
AI Technical Summary
Existing scanning equipment, in the absence of reference and with deviations in viewing angle, may result in overscanning or underscanning of text, affecting the accuracy of the scanned image and the accuracy of the recognition results.
By identifying valid boundary image frames, determining the validity of boundary characters, and determining the segmentation position based on validity, the scanning image is segmented to extract image frames containing valid characters and remove invalid character frames, thereby improving the accuracy of the scanning image.
This effectively avoids missing valid characters and scanning invalid characters multiple times, improving the accuracy of the scanned image and thus improving the accuracy of the recognition results.
Smart Images

Figure CN115565189B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular to a scanning image recognition method, apparatus, electronic device, and storage medium. Background Technology
[0002] With the continuous upgrading and development of image character recognition technology in recent years, information from books, documents and other materials can be recognized by scanning the images and then identifying the characters in the images. As a result, scanning devices such as scanning pens with character scanning and recognition functions have gained widespread recognition.
[0003] When a user uses a scanning pen to scan, due to a lack of reference, unclear reference positioning, or deviation between the viewing angle and the viewing angle of the image acquisition component in the scanning pen, inconsistencies may occur in human-computer perception. This can lead to overscanning or underscanning of text, resulting in a final scanned image that differs from the user's needs, affecting the accuracy of the scanned image and thus causing a low recognition accuracy rate. Summary of the Invention
[0004] Based on the deficiencies and shortcomings of the prior art, this application proposes a scanning image recognition method, apparatus, electronic device, and storage medium, which can improve the accuracy of scanning images and thus improve the accuracy of recognition results.
[0005] The first aspect of this application provides a method for scanning image recognition, including:
[0006] From the acquired scanned image frames, valid boundary image frames are determined, including the pen-dropping image frame at the moment of pen drop and / or the pen-lifting image frame at the moment of pen lift.
[0007] Determine whether the boundary character in the valid boundary image frame is a valid character. If it is a valid character, determine the first character boundary of the boundary character from the target image stitched together from the acquired scanned image frames, as the segmentation position. The first character boundary ensures that the image to be recognized obtained by segmenting the target image according to the segmentation position contains the boundary character.
[0008] If the character is invalid, the second character boundary of the boundary character is determined from the target image stitched together from the acquired scanned image frames, and used as the segmentation position; the second character boundary ensures that the image to be recognized obtained by segmenting the target image according to the segmentation position does not contain the boundary character;
[0009] Character recognition is performed on the image to be recognized obtained by segmenting the target image according to the segmentation position to obtain the scanning result.
[0010] Optionally, determining the valid boundary image frames from the acquired scanned image frames includes:
[0011] Illumination intensity is detected on the acquired scanned image frames. The first image frame in a continuous image frame sequence with illumination intensity within a preset stable range is taken as the pen-starting image frame, and / or the last image frame in a continuous image frame sequence with illumination intensity within a preset stable range is taken as the pen-lifting image frame.
[0012] Optionally, after determining the valid boundary image frames from the acquired scanned image frames, the method further includes:
[0013] Based on the pen-starting image frame and the pen-removing image frame, extract the image sequence to be stitched from the acquired scanned image frames; wherein, the image sequence to be stitched includes at least all scanned image frames from the pen-starting image frame to the pen-removing image frame;
[0014] Determining the first character boundary of the boundary character from the target image stitched together from the acquired scanned image frames, as the segmentation position, includes:
[0015] The first character boundary of the boundary character is determined from the target image formed by stitching together scanned image frames from the image sequence to be stitched, and used as the segmentation position;
[0016] Determining the second character boundary of the boundary character from the target image stitched together from the acquired scanned image frames, as the segmentation position, includes:
[0017] The second character boundary of the boundary character is determined from the target image formed by stitching together scanned image frames from the image sequence to be stitched, and used as the cutting position.
[0018] Optionally, based on the pen-starting image frame and the pen-removing image frame, the image sequence to be stitched is extracted from the acquired scanned image frames, including:
[0019] Extract all scanned image frames from the pen-starting image frame to the pen-lifting image frame from the acquired scanned image frames as the first image sequence;
[0020] According to a preset number of expansions, extract the image sequence adjacent to the first image sequence from the acquired scanned image frames as the expanded image sequence;
[0021] According to the image acquisition order of the extended image sequence and the first image sequence, the extended image sequence and the first image sequence are combined to obtain the image sequence to be stitched together.
[0022] Optionally, determining whether a boundary character in the valid boundary image frame is a valid character includes:
[0023] The integrity of the boundary characters in the valid boundary image frame is detected, and the rectangular enclosing boundary of the boundary characters in the valid boundary image frame is predicted.
[0024] Based on the integrity of the boundary characters in the valid boundary image frame and the rectangular enclosing boundary of the boundary characters in the valid boundary image frame, it is determined whether the boundary characters in the valid boundary image frame are valid characters.
[0025] In this context, the boundary character in the pen-starting image frame is the first character, and the boundary character in the pen-lifting image frame is the last character.
[0026] Optionally, detecting the integrity of boundary characters in the valid boundary image frame and predicting the rectangular enclosing boundary of the boundary characters in the valid boundary image frame includes:
[0027] Extract a predetermined number of image frames before the pen-starting image frame from the acquired scanned image frames as auxiliary image frames for the pen-starting image frame, and / or extract a predetermined number of image frames after the pen-lifting image frame from the acquired scanned image frames as auxiliary image frames for the pen-lifting image frame.
[0028] Based on the attention mechanism, the image coding features corresponding to the effective boundary image frame are determined by using the image features of the auxiliary image frame of the effective boundary image frame and the image features of the effective boundary image frame.
[0029] Based on the image encoding features corresponding to the valid boundary image frame, predict the integrity of the boundary characters and the rectangular enclosing boundary in the valid boundary image frame.
[0030] Optionally, determining whether a boundary character in a valid boundary image frame is a valid character based on the integrity of the boundary character in the valid boundary image frame and the rectangular enclosing boundary of the boundary character in the valid boundary image frame includes:
[0031] If the integrity of the boundary characters in the valid boundary image frame indicates that the characters are complete, then the boundary characters in the valid boundary image frame are determined to be valid characters.
[0032] If the integrity of the boundary character in the valid boundary image frame indicates that the character is incomplete, then based on the rectangular bounding boundary of the boundary character in the valid boundary image frame, the target proportion of the display area of the boundary character in the valid boundary image frame to the overall area of the boundary character is determined, and based on the target proportion corresponding to the boundary character in the valid boundary image frame, it is determined whether the boundary character in the valid boundary image frame is a valid character.
[0033] Optionally, determining whether a boundary character in the valid boundary image frame is a valid character based on the target proportion corresponding to the boundary character in the valid boundary image frame includes:
[0034] If the target proportion corresponding to the boundary character in the effective boundary image frame is within a preset effective proportion range, then the boundary character in the effective boundary image frame is determined to be a valid character.
[0035] If the target proportion corresponding to the boundary character in the valid boundary image frame is within a preset invalid proportion range, then the boundary character in the valid boundary image frame is determined to be an invalid character.
[0036] If the target proportion corresponding to the boundary character in the valid boundary image frame is not within the valid proportion range, and the target proportion corresponding to the boundary character in the valid boundary image frame is not within the invalid proportion range, then the boundary character in the valid boundary image frame is determined to be a valid character based on the influence of the boundary character in the valid boundary image frame on the integrity of the scanning statement corresponding to the target image.
[0037] Optionally, determining whether a boundary character in the valid boundary image frame is a valid character based on its impact on the integrity of the scanning statement corresponding to the target image includes:
[0038] Calculate the first statement integrity probability when the boundary characters in the effective boundary image frame are added to the scanning statement corresponding to the target image, and the second statement integrity probability when the boundary characters in the effective boundary image frame are not added to the scanning statement corresponding to the target image;
[0039] If the ratio between the first statement integrity probability and the second statement integrity probability is greater than a preset ratio, then the boundary character in the valid boundary image frame is determined to be a valid character;
[0040] If the ratio between the second statement integrity probability and the first statement integrity probability is greater than a preset ratio, then the boundary character in the valid boundary image frame is determined to be an invalid character.
[0041] Optionally, determining the first character boundary of the boundary character from the target image stitched together from the acquired scanned image frames, as the segmentation position, includes:
[0042] If the effective boundary image frame includes the pen stroke image frame, then the left boundary of the rectangular enclosing boundary of the boundary character is determined from the target image stitched together from the acquired scanned image frames, and used as the cutting position;
[0043] If the effective boundary image frame includes the pen-lifting image frame, then the right boundary of the rectangular enclosing boundary of the boundary character is determined from the target image stitched together from the acquired scanned image frames, and used as the cutting position.
[0044] Optionally, determining the second character boundary of the boundary character from the target image stitched together from the acquired scanned image frames, as the segmentation position, includes:
[0045] If the effective boundary image frame includes the pen stroke image frame, then the right boundary of the rectangular enclosing boundary of the boundary character is determined from the target image stitched together from the acquired scanned image frames, and used as the cutting position;
[0046] If the effective boundary image frame includes the pen-lifting image frame, then the left boundary of the rectangular enclosing boundary of the boundary character is determined from the target image stitched together from the acquired scanned image frames, and used as the cutting position.
[0047] Optionally, the method further includes:
[0048] Extract sample scan image frame sequences from the historical scan image frame sequences corresponding to historical scans within a preset time period;
[0049] The vertical offset of the sample image is determined based on the text bounding box and the image center position of the sample image spliced from the sample scan image frame sequence.
[0050] The lateral offset of the sample image is determined based on the segmentation position corresponding to the sample image spliced from the sample scan image frame sequence.
[0051] The scanning parameters of the scanning device are adjusted based on the horizontal and vertical offsets of the sample image.
[0052] A second aspect of this application provides a scanning image recognition device, comprising:
[0053] The image frame determination module is used to determine the effective boundary image frames from the acquired scanned images. The effective boundary image frames include the pen-dropping image frame at the moment of pen drop and / or the pen-lifting image frame at the moment of pen lift.
[0054] The segmentation determination module is used to determine whether the boundary character in the valid boundary image frame is a valid character. If it is a valid character, the first character boundary of the boundary character is determined from the target image stitched together from the acquired scanned image frames as the segmentation position. The first character boundary ensures that the image to be recognized obtained by segmenting the target image according to the segmentation position contains the boundary character.
[0055] The segmentation determination module is further configured to, if an invalid character is found, determine a second character boundary of the boundary character from the target image assembled from the acquired scanned image frames, as the segmentation position; the second character boundary ensures that the image to be recognized obtained by segmenting from the target image according to the segmentation position does not contain the boundary character;
[0056] The recognition module is used to perform character recognition on the image to be recognized obtained by segmenting the target image according to the segmentation position, and obtain the scanning result.
[0057] A third aspect of this application provides an electronic device, including: a memory and a processor;
[0058] The memory is connected to the processor and is used to store programs;
[0059] The processor is used to implement the above-described scanned image recognition method by running the program in the memory.
[0060] A fourth aspect of this application provides a storage medium storing a computer program, which, when executed by a processor, implements the above-described scanned image recognition method.
[0061] The scanning image recognition method proposed in this application determines valid boundary image frames from the acquired scanning image; determines whether the boundary characters in the valid boundary image frames are valid characters; if they are valid characters, the first character boundary of the boundary character is determined from the target image stitched from the acquired scanning image frames as the segmentation position; if they are invalid characters, the second character boundary of the boundary character is determined from the target image stitched from the acquired scanning image frames as the segmentation position; and performs character recognition on the image to be recognized obtained by segmenting from the target image according to the segmentation position to obtain the scanning result. The first character boundary ensures that the image to be recognized obtained by segmenting from the target image according to the segmentation position contains the boundary character, thus avoiding missed scanning of valid characters; the second character boundary ensures that the image to be recognized obtained by segmenting from the target image according to the segmentation position does not contain the boundary character, thus avoiding over-scanning of invalid characters, thereby improving the accuracy of the scanning image and consequently improving the accuracy of the recognition result. Attached Figure Description
[0062] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0063] Figure 1This is a flowchart illustrating a scanning image recognition method provided in an embodiment of this application;
[0064] Figure 2 This is a flowchart illustrating another scanning image recognition method provided in an embodiment of this application;
[0065] Figure 3 This is a schematic diagram of the process for determining whether a boundary character is a valid character, provided in an embodiment of this application.
[0066] Figure 4 This is a schematic diagram of the integrity of boundary characters and the rectangular bounding boundary in the predicted pen stroke image frame provided in the embodiments of this application;
[0067] Figure 5 This is a schematic diagram of the processing flow for determining whether a boundary character is a valid character based on the target proportion corresponding to the boundary character, provided in an embodiment of this application.
[0068] Figure 6 This is a schematic diagram of the process for calibrating a scanning device provided in an embodiment of this application;
[0069] Figure 7 This is a schematic diagram of the structure of a scanning image recognition device provided in an embodiment of this application;
[0070] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0071] The technical solutions of this application are applicable to image processing scenarios, especially to image processing of scanned images from scanning devices. By employing the technical solutions of this application, the accuracy of scanned images can be improved, thereby increasing the accuracy of recognition results.
[0072] In recent years, with the rapid development of character recognition technology, scanning devices such as scanning pens with image scanning and recognition functions have gained widespread acceptance. For example, a scanning pen can scan books, documents, and other materials, and then use character recognition technology to identify the characters in the scanned image, thereby recognizing the information. The general structure of a scanning pen is similar to a common fountain pen. Specifically, the scanning method involves adding a front baffle to guide the user's scan, and incorporating optical structures and image acquisition devices to assist in information acquisition. When the user scans, pressing the front baffle with the pen triggers the optical scanning head to continuously capture images of the scanned content. Releasing the pen stops the capturing of images, obtaining continuous video or image frames, and then using artificial intelligence algorithms such as character recognition and translation to complete subsequent processing.
[0073] However, existing scanning devices such as scanning pens have limitations in use. Due to a lack of reference, unclear reference positioning, and discrepancies between the user's line of sight and the image acquisition direction of the scanning pen, the content actually scanned by the user may differ from the content expected by the scanning device. This results in a discrepancy between the scanned content and the actual content received, making it impossible to achieve "what you see is what you get." This leads to inconsistencies between human and machine perception, causing the scanning device to overscan or miss text during scanning, thus affecting the accuracy of the scanned image and consequently the accuracy of the image recognition results.
[0074] In view of the shortcomings of the prior art and the problems of low accuracy of scanned images and low accuracy of recognition results, the inventors of this application have proposed a scanned image recognition method after research and experimentation. This method can avoid missing effective characters and scanning invalid characters, thereby improving the accuracy of scanned images and thus improving the accuracy of image recognition results.
[0075] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0076] This application provides a method for scanning image recognition; see [link to relevant documentation]. Figure 1 As shown, the method includes:
[0077] S101. Determine the valid boundary image frames from the acquired scanned image frames.
[0078] A barcode scanner uses its optical imaging unit to scan, which consists of a light source and an image acquisition component (such as a camera). Typically, the light source's response time is faster than the image acquisition component. For example, when a user presses the scanner, the light source turns on first, then the image acquisition component starts working; when the user lifts the scanner, the light source turns off first, then the image acquisition component stops working.
[0079] Because the supplementary light may turn on and automatically adjust exposure when the user puts down the pen during scanning, the first few scanned image frames may have unstable lighting. When the user lifts the pen, the supplementary light turns off first, and the image acquisition component stops working only after the supplementary light turns off. Therefore, the scanned image frames acquired after the pen is lifted, before the supplementary light turns off and the image acquisition component stops working, may have low lighting intensity or capture image frames unrelated to the target content. Therefore, this embodiment first needs to determine the pen-putting image frame and / or pen-lifting image frame from all the scanned image frames acquired in this scan, in order to exclude images unrelated to the target content to be identified in the target image stitched from the acquired scanned image frames. This embodiment uses the pen-putting image frame and / or pen-lifting image frame as the valid boundary image frames. The pen-putting image frame is the first image frame after the lighting stabilizes, and the pen-lifting image frame is the image frame before the sudden change in lighting, that is, the last image frame after the lighting stabilizes.
[0080] Regarding the pen-starting image frame, since the fill light works before the image acquisition component, the scanned image frame acquired by the image acquisition component is the initial image frame where the user has pressed the button and is about to start scanning. When the fill light does not have an automatic exposure adjustment function, the initial image frame acquired at the start of scanning is an image frame with stable lighting. In this case, the first image frame acquired by the image acquisition component can be used as the pen-starting image frame. When the fill light has an automatic exposure adjustment function, the light intensity of the initial image frame acquired at the start of scanning is not stable, and there may be image frames with strong light intensity and / or weak light intensity. In this case, the first image frame acquired after the light intensity stabilizes should be used as the pen-starting image frame.
[0081] Regarding the pen-lifting image frame, since the fill light stops working earlier than the image acquisition component, there is a clear process from illumination equalization to illumination abrupt change in the image frames acquired by the image acquisition component before and after the user lifts the pen. The fill light turns off when the user lifts the pen, and the illumination of the image frames acquired afterward changes abruptly. Therefore, it is possible to determine whether the user has lifted the pen based on the change in illumination intensity of the scanned image frame, and the image frame before the illumination abrupt change in the scanned image frame is taken as the pen-lifting image frame.
[0082] Specifically, this embodiment first performs illumination intensity detection on the scanned image frames acquired by the image acquisition component. Then, it determines a continuous sequence of image frames from all scanned image frames whose illumination intensity is within a preset stable range. All scanned image frames in this sequence are continuously acquired by the image acquisition component, and the illumination intensity is stable. The preset stable range is the illumination intensity range after the supplementary light has undergone automatic exposure adjustment and the illumination has stabilized, or a pre-set fixed illumination intensity range for the supplementary light. Since the effective boundary image frames include the image frame at the moment of pen placement and / or the image frame at the moment of pen lifting, determining the effective boundary image frames requires taking the first image frame in the determined image frame sequence as the pen placement image frame, and / or taking the last image frame in the determined image frame sequence as the pen lifting image frame. In this embodiment, existing algorithms for detecting illumination intensity in images can be used to detect the illumination intensity of the scanned image frames, such as histogram statistical algorithms.
[0083] S102. Determine whether the boundary character in the valid boundary image frame is a valid character. If yes, proceed to step S103; otherwise, proceed to step S104.
[0084] Because users lack reference points or have unclear reference positioning when scanning with a scanning pen, and because there is a discrepancy between the user's viewing angle and the image acquisition angle of the scanning pen's image acquisition component, the position where the pen touches the ground may deviate from the position the user actually wants to start scanning, and the position where the pen is lifted may deviate from the position the user actually wants to stop scanning. This results in the truncation of boundary characters in the valid boundary image frame (i.e., the first character in the pen-touching image frame or the last character in the pen-lifting image frame), so that only a portion of the boundary character is captured in the valid boundary image frame. When a boundary character in the valid boundary image frame is truncated, it may be because the boundary character was missed or scanned multiple times. If the boundary character was missed, it indicates that the boundary character is valid; if the boundary character was scanned multiple times, it indicates that the boundary character is invalid.
[0085] Therefore, this embodiment first needs to determine the boundary characters in the valid boundary image frames, and then detect whether the boundary characters are valid characters, thereby determining whether the start and stop positions of this scan have been over-scanned or under-scanned. Specifically, if the valid boundary image frames include a pen-starting image frame, then the boundary character of the pen-starting image frame is the first character in the pen-starting image frame; if the valid boundary image frames include a pen-lifting image frame, then the boundary character of the pen-lifting image frame is the last character in the pen-lifting image frame.
[0086] In this embodiment, the completeness of the valid character is very high, or the valid character has a significant impact on the completeness of the scanning statement corresponding to this scan. Therefore, whether the boundary character is a valid character can be determined based on the completeness of the boundary character in the valid boundary image frame or the impact of the boundary character on the completeness of the statement corresponding to this scan. If the boundary character in the valid boundary image frame is determined to be a valid character, then step S103 is executed; if the boundary character in the valid boundary image frame is determined to be a invalid character, then step S104 is executed.
[0087] S103. Determine the first character boundary of the boundary character from the target image stitched together from the acquired scanned image frames, and use it as the segmentation position.
[0088] In this embodiment, when a user scans with a scanning pen, the image acquisition component in the scanning pen needs to stitch the scanned image frames acquired for the current scan into a target image. Then, the image to be recognized is segmented from the target image for character recognition. The stitching of scanned image frames can be done while scanning; that is, the image acquisition component acquires a scanned image frame and stitches it with the previously stitched image until the current scan is complete. Alternatively, the image acquisition component can acquire all scanned image frames for the current scan and then stitch them together. The image frame stitching can employ existing image stitching algorithms, and this embodiment will not elaborate on the stitching process.
[0089] Specifically, if the boundary character in the valid boundary image frame is determined to be a valid character, then in this embodiment, the first character boundary of the boundary character in the target image obtained by this scan is used as the segmentation position. This ensures that the image to be recognized after segmenting the target image according to the first character boundary contains the boundary character in the valid boundary image frame, thereby preventing the boundary character, which is considered a valid character, from being missed. If the valid boundary image frame includes a pen stroke image frame, then the first character boundary of the boundary character in the pen stroke image frame is the left boundary of the first character in the pen stroke image frame. Segmenting from the left boundary of the first character in the pen stroke image frame of the target image, the image to be recognized is located to the right of the left boundary of the first character, and the first character is also in the image to be recognized. If the valid boundary image frame includes a pen lift image frame, then the first character boundary of the boundary character in the pen lift image frame is the right boundary of the last character in the pen lift image frame. Segmenting from the right boundary of the last character in the pen lift image frame of the target image, the image to be recognized is located to the left of the right boundary of the last character, and the last character is also in the image to be recognized.
[0090] S104. Determine the second character boundary of the boundary character from the target image stitched together from the acquired scanned image frames, and use it as the segmentation position.
[0091] Specifically, if the boundary character in the valid boundary image frame is determined to be an invalid character, this embodiment needs to use the second character boundary of the boundary character in the target image obtained from this scan as the segmentation position. This ensures that the image to be recognized after segmenting the target image according to the second character boundary does not contain the boundary character from the valid boundary image frame, thereby preventing the invalid boundary character from being scanned multiple times. If the valid boundary image frame includes a pen stroke image frame, then the second character boundary of the boundary character in the pen stroke image frame is the right boundary of the first character in the pen stroke image frame. Segmentation is performed from the right boundary of the first character in the pen stroke image frame in the target image. The image to be recognized is located to the right of the right boundary of the first character, and the first character is not in the image to be recognized, thus achieving the removal of the first character as an invalid character. If the valid boundary image frame includes the pen lifting image frame, then the second character boundary of the boundary character in the pen lifting image frame is the left boundary of the last character in the pen lifting image frame. The image is segmented from the left boundary of the last character in the pen lifting image frame in the target image. The image to be recognized is to the left of the left boundary of the last character. The last character is not in the image to be recognized, thus realizing the removal of the last character as an invalid character.
[0092] S105. Perform character recognition on the image to be recognized obtained by segmenting the target image according to the segmentation position to obtain the scanning result.
[0093] Specifically, in this embodiment, after determining the segmentation position of the target image based on whether the boundary characters in the valid boundary image frame are valid characters, the target image is segmented according to the segmentation position to obtain the image to be recognized. Then, character recognition technology is used to recognize the characters in the image to be recognized, and the recognition result is used as the scanning result of this scan. The character recognition technology used in this embodiment can be an existing character recognition technology, such as OCR technology.
[0094] As described above, the scanning image recognition method proposed in this application determines valid boundary image frames from the acquired scanning images; determines whether the boundary characters in the valid boundary image frames are valid characters; if they are valid characters, the first character boundary of the boundary character is determined from the target image stitched from the acquired scanning image frames as a segmentation position; if they are invalid characters, the second character boundary of the boundary character is determined from the target image stitched from the acquired scanning image frames as a segmentation position; and character recognition is performed on the image to be recognized obtained by segmenting from the target image according to the segmentation position to obtain the scanning result. The first character boundary ensures that the image to be recognized obtained by segmenting from the target image according to the segmentation position contains the boundary character, thus avoiding missed scanning of valid characters; the second character boundary ensures that the image to be recognized obtained by segmenting from the target image according to the segmentation position does not contain the boundary character, thus avoiding over-scanning of invalid characters, thereby improving the accuracy of the scanning image and consequently improving the accuracy of the recognition result.
[0095] Furthermore, in this embodiment, before scanning with the scanning pen, the user can perform initial calibration to preliminarily align the baffle area of the scanning pen with the imaging area from a specific viewing angle. This embodiment can use existing scanning pen initial calibration methods to perform initial calibration, and the process of initial calibration will not be described in detail here.
[0096] As an optional implementation method, see [link to implementation details]. Figure 2 As shown, another embodiment of this application discloses that after performing step S101, the method further includes:
[0097] S202. Based on the pen-starting image frame and the pen-lifting image frame, extract the image sequence to be stitched from the acquired scanned image frames.
[0098] Specifically, after determining the pen-starting image frame and pen-removing image frame from the acquired scanned image frames, to reduce image frame stitching work and improve scanned image recognition efficiency, it is not necessary to stitch all acquired scanned image frames; only the necessary scanned image frames need to be extracted and stitched. Therefore, this embodiment can extract the image sequence to be stitched from the acquired scanned image frames based on the pen-starting and pen-removing image frames, and then stitch all scanned image frames in the image sequence to be stitched to obtain the corresponding target image. The image sequence to be stitched includes at least all scanned image frames from the pen-starting image frame to the pen-removing image frame.
[0099] When the boundary character is a valid character, and the valid boundary image frame only captures a portion of the boundary character, the first character boundary of the boundary character is in an image frame other than all scanned image frames from the pen-starting image frame to the pen-removing image frame. In this case, if all scanned image frames from the pen-starting image frame to the pen-removing image frame are stitched together to obtain the target image, the segmentation position is not in the target image, and segmentation cannot be achieved. Therefore, the image sequence to be stitched can also include the image frame sequence adjacent to the pen-starting image frame and the image frame sequence adjacent to the pen-removing image frame.
[0100] Furthermore, this step specifically includes:
[0101] First, extract all scanned image frames from the pen-starting image frame to the pen-lifting image frame from the acquired scanned image frames as the first image sequence.
[0102] In this embodiment, according to the acquisition order of the scanned image frames, all scanned image frames from the pen-starting image frame to the pen-lifting image frame are extracted from the acquired scanned image frames to form the first image sequence. For example, if a4 is taken as the pen-starting image frame and b is taken as the pen-lifting image frame, then the corresponding first image sequence is (a, a+1, a+2, ..., b-2, b-1, b).
[0103] Second, according to a preset number of extensions, image sequences adjacent to the first image sequence are extracted from the acquired scanned image frames as extended image sequences.
[0104] In this embodiment, an extended image sequence is extracted from the acquired scanned image frames, adjacent to the first image sequence, according to a preset extension number. Specifically, a preset number of preceding image sequences adjacent to the preceding pen-writing image frame are extracted from the acquired scanned image frames, and / or a preset number of following image sequences adjacent to the following pen-lifting image frame are extracted from the acquired scanned image frames. The extended image sequence includes the preceding image sequence and / or the following image sequence. In this embodiment, the user can set the preset extension number based on experience; the preset extension number is preferably 5 to 10 frames, but this embodiment does not impose any limitation.
[0105] If the number of scanned image frames preceding the pen-starting image frame does not reach the preset expansion number, then only all scanned image frames preceding the pen-starting image frame need to be used as the preceding image sequence. Similarly, if the number of scanned image frames following the pen-lifting image frame does not reach the preset expansion number, then only all scanned image frames following the pen-lifting image frame need to be used as the following image sequence. For example, if the fill light in the scanning pen is not set to automatic exposure adjustment, and the first image frame among all captured scanned image frames is used as the pen-starting image frame, then since there are no other image frames preceding the pen-starting image frame, only the following image sequence needs to be extracted as the extended image sequence.
[0106] Third, according to the image acquisition order of the extended image sequence and the first image sequence, the extended image sequence and the first image sequence are combined to obtain the image sequence to be stitched together.
[0107] According to the image acquisition order of the scanned image frames, the extended image sequence and the first image sequence are combined into an image sequence to be stitched. If the extended image sequence includes a previous image sequence and a subsequent image sequence, the previous image sequence, the first image sequence, and the subsequent image sequence are combined into an image sequence to be stitched. For example, if k is used as the preset extension number, and the first image sequence is (a, a+1, a+2, ..., b-2, b-1, b), then the image sequence to be stitched is (ak, a-k+1, ..., a, a+1, a+2, ..., b-2, b-1, b, b+1, ..., b+k).
[0108] S204. Determine the first character boundary of the boundary character from the target image formed by stitching together scanned image frames from the image sequence to be stitched, and use it as the cutting position.
[0109] In this embodiment, the target image is formed by stitching together all scanned image frames in the image sequence to be stitched as determined in the above steps. The method for determining the first character boundary of the boundary character is the same as step S103 in the above embodiment, and will not be described in detail in this embodiment.
[0110] S205. Determine the second character boundary of the boundary character from the target image formed by stitching together scanned image frames from the image sequence to be stitched, and use it as the cutting position.
[0111] In this embodiment, the target image is formed by stitching together all scanned image frames in the image sequence to be stitched as determined in the above steps. The method for determining the second character boundary of the boundary character is the same as step S104 in the above embodiment, and will not be described in detail in this embodiment.
[0112] Figure 2 Step S201 and Figure 1 The steps in step S101 are the same. Figure 2 Step S203 and Figure 1 The steps in step S102 are the same. Figure 2 Step S206 and Figure 1 The steps S105 are the same as those in the previous embodiment, and steps S201, S203 and S206 will not be described in detail in this embodiment.
[0113] As an optional implementation method, see [link to implementation details]. Figure 3 As shown, another embodiment of this application discloses that step S102 includes:
[0114] S301. Detect the integrity of the boundary characters in the valid boundary image frame and predict the rectangular enclosing boundary of the boundary characters in the valid boundary image frame.
[0115] Specifically, this embodiment can detect the integrity of boundary characters in the valid boundary image frame to determine whether the boundary characters captured in the valid boundary image frame are incomplete. If the boundary characters in the valid boundary image frame are incomplete, it indicates that the user scanned too many or too few characters when using the scanning pen, resulting in the truncation of the scanned boundary characters. This embodiment can also predict the rectangular enclosing boundary of the boundary characters in the valid boundary image frame. This rectangular enclosing boundary is the smallest complete rectangular enclosing boundary of the boundary character. If the boundary characters in the valid boundary image frame are incomplete, then the rectangular enclosing boundary when the boundary character is complete needs to be predicted as the rectangular enclosing boundary of the boundary character.
[0116] Furthermore, this step specifically includes:
[0117] First, extract a preset number of image frames before the pen-down image frame from the collected scanned image frames as the auxiliary image frames of the pen-down image frame, and / or extract a preset number of image frames after the pen-up image frame from the collected scanned image frames as the auxiliary image frames of the pen-up image frame.
[0118] In this embodiment, there may be a situation where the boundary characters in the valid boundary image frame are incomplete. If the boundary characters are incomplete, then only using the image features of the valid boundary image frame to predict the integrity of the boundary characters and the rectangular bounding boundary has a low accuracy. Therefore, in this embodiment, the auxiliary image frames of the valid boundary image frame can be extracted from the collected scanned image frames, and the image features of the auxiliary image frames and the image features of the valid boundary image frame are used to predict the integrity of the boundary characters and the rectangular bounding boundary, which can improve the prediction accuracy.
[0119] Since the valid boundary image frame includes the pen-down image frame and / or the pen-up image frame, therefore, in this embodiment, it is necessary to extract the auxiliary image frames of the pen-down image frame from the collected scanned image frames, and / or extract the auxiliary image frames of the pen-up image frame from the collected scanned image frames. When the valid boundary image frame includes the pen-down image frame, extract a preset number of image frames before the pen-down image frame from the collected scanned image frames as the auxiliary image frames of the pen-down image frame, and the image frame sequence composed of the auxiliary image frames of the pen-down image frame is adjacent to the pen-down image frame. When the valid boundary image frame includes the pen-up image frame, extract a preset number of image frames after the pen-up image frame from the collected scanned image frames as the auxiliary image frames of the pen-up image frame, and the image frame sequence composed of the auxiliary image frames of the pen-up image frame is adjacent to the pen-up image frame.
[0120] As Figure 4 shown, the first character in the pen-down image frame is a残缺的字符 (incomplete character). If only using the image features of the pen-down image frame to predict the rectangular bounding boundary, then the first character in the pen-down image frame may be "聆", "冷", "玲", etc. Different characters result in different predicted rectangular bounding boundaries. Therefore, the auxiliary image frames of the pen-down image frame can be used to predict the integrity of the first character in the pen-down image frame and the rectangular bounding boundary, thereby improving the accuracy of the prediction result. As Figure 4 shown, the preset number is set to n, and the pen-down image frame is used as the 0th image frame in the collected scanned image frames. Then, the -nth frame to the -1st frame in the collected scanned image frames are the auxiliary image frames of the pen-down image frame. If the pen-up image frame is the xth image frame in the collected scanned image frames, then the x + 1st frame to the x + nth frame in the collected scanned image frames are used as the auxiliary image frames of the pen-up image frame.
[0121] Second, based on the attention mechanism, use the image features of the auxiliary image frames of the valid boundary image frame and the image features of the valid boundary image frame to determine the image coding feature corresponding to the valid boundary image frame.
[0122] This embodiment requires image feature extraction from the effective boundary image frame and each auxiliary image frame of the effective boundary image frame to obtain the image features of the auxiliary image frames of the effective boundary image frame and the image features of the effective boundary image frame. This embodiment can employ existing image feature extraction algorithms to extract image features from each image frame, such as LBP (Local Binary Patterns), HOG (Histogram of Oriented Gradients), and SIFT (Scale-Invariant Feature Transform).
[0123] This embodiment utilizes an attention mechanism to perform weighted fusion of image features from auxiliary image frames and image features from the effective boundary image frame to obtain the image coding features corresponding to the effective boundary image frame. This embodiment can first perform feature fusion on the image features of each auxiliary image frame of the effective boundary image frame to obtain auxiliary image features, and then perform a second weighted fusion on the auxiliary image features and the image features of the effective boundary image frame to obtain the image coding features corresponding to the effective boundary image frame. Since the effective boundary image frame has greater predictive significance, the weight of the image features of the effective boundary image frame should be set larger during the second weighted fusion. The formula for calculating the image coding features corresponding to the effective boundary image frame is: f fuse =λ1att(f -n f -n+1 , ..., f -1 )+λ2f T .
[0124] Among them, f fuse f represents the image coding features corresponding to the valid boundary image frames. T Att(f) represents the image features of the valid boundary image frame. -n f -n+1 , ..., f -1 ) represents the auxiliary graphic features obtained after fusing the image features of auxiliary image frames (i.e., all auxiliary image frames of the effective boundary image frames) from frame -n to frame -1. λ1 represents the weight of the auxiliary image features, and λ2 represents the weight of the image features of the effective boundary image frames. In this embodiment, the weights of the auxiliary image features and the weights of the image features of the effective boundary image frames can be set empirically. In this embodiment, it is preferred to set λ1 = 0.2 and λ2 = 0.8.
[0125] In this embodiment, a sequence coding network based on an attention mechanism (such as...) can be used. Figure 4The attention-based sequence coding module shown uses the image features of the auxiliary image frame and the image features of the effective boundary image frame to determine the image code features corresponding to the effective boundary image frame (e.g., ...). Figure 4 (Hidden layer features shown).
[0126] Third, based on the image coding features corresponding to the valid boundary image frames, predict the integrity of the boundary characters and the rectangular enclosing boundary in the valid boundary image frames.
[0127] This embodiment determines the image coding features corresponding to the valid boundary image frame, and then predicts the integrity of the boundary characters and the rectangular enclosing boundary in the valid boundary image frame based on these image coding features. This embodiment can employ a character integrity prediction network (such as...). Figure 4 The pre-truncation judgment module in the network predicts the integrity of the boundary character, that is, whether the boundary character is truncated. The character integrity prediction network can use a binary classification model. If the boundary character is predicted to be truncated, the output is 1, which means that the integrity of the boundary character is not complete. If the boundary character is predicted not to be truncated, the output is 0, which means that the integrity of the boundary character is complete.
[0128] This embodiment can employ a complete boundary prediction network (such as...). Figure 4 The complete boundary prediction module in the code predicts the rectangular bounding boundary of the boundary character. This rectangular bounding boundary can include the coordinates of the four vertices of the bounding box, or it can include the coordinates of the top-left vertex of the bounding box and the length and width of the bounding box. For example... Figure 4 In the formula [x, y, w, h] = [-30, 10, 60, 70], x is the x-coordinate of the top-left vertex of the bounding box, y is the y-coordinate of the top-left vertex of the bounding box, w is the width of the bounding box, and h is the length of the bounding box.
[0129] Furthermore, this embodiment can combine an attention-based sequence encoding network, a character integrity prediction network, and a complete boundary prediction network into a prediction model. Then, the prediction model is trained using a pre-collected set of pen-starting image frames and pen-lifting image frames. The pen-starting image frame sample set includes pre-collected sample pen-starting frames and corresponding sample pen-starting auxiliary frames. Each sample pen-starting frame carries the integrity identifier of the boundary character and rectangular bounding boundary data. The image features of the sample pen-starting frames and the image features of each sample pen-starting auxiliary frame are then input into the prediction model. The prediction model outputs the predicted integrity and predicted rectangular bounding boundary of the sample pen-starting frame. A loss value is calculated using a pre-set optimization objective function, the predicted integrity, the predicted rectangular bounding boundary, and the integrity identifier of the boundary character and the rectangular bounding boundary data carried by the sample pen-starting frame. The parameters of each network in the prediction model are then adjusted using the loss value. The optimization objective function is:
[0130] L=-λ1L cls +λ2L reg
[0131]
[0132] L reg =∑ i∈{x,y,w,h} smooth L1 (t i -u i )
[0133] Among them, L cls The objective function is classification, calculated using cross-entropy loss, where T represents the number of samples, and S... j L represents the confidence level of the j-th sample. reg For the regression objective function, t i u represents the bounding rectangle of the predicted rectangle for the i-th sample. i This represents the rectangular bounding boundary data carried by the sample stroke frame corresponding to the i-th sample, smooth. L1 This represents a smoothing process for the L1 function, where the L1 function is t. i -u i λ1 is the weighting coefficient of the classification objective function, and λ2 is the weighting coefficient of the regression objective function. λ1 and λ2 can be set empirically; in this embodiment, they are preferably set to λ1 = 0.3 and λ2 = 0.7.
[0134] This embodiment can train the prediction model as a whole, or it can train the attention-based sequence coding network, character integrity prediction network, and complete boundary prediction network in the prediction model separately. The method of training each network separately is similar to the method of training the whole model, and will not be described in detail in this embodiment.
[0135] S302. Determine whether the boundary characters in the valid boundary image frame are valid characters based on the integrity of the boundary characters in the valid boundary image frame and the rectangular enclosing boundary of the boundary characters in the valid boundary image frame.
[0136] Specifically, this embodiment first determines whether the boundary characters in the valid boundary image frame are complete based on their integrity. If the integrity of the boundary characters in the valid boundary image frame indicates that the characters are complete, it is determined that no extra scans or omissions occurred during this scan, and the boundary characters in the valid boundary image frame are valid characters. If the integrity of the boundary characters in the valid boundary image frame indicates that the characters are incomplete, it means that a deviation occurred during this scan, resulting in extra scans or omissions. In this case, it is necessary to further determine whether the incomplete boundary characters are valid or invalid based on the rectangular enclosing boundary of the boundary characters in the valid boundary image frame. If the boundary character is valid, it means that the deviation of this scan caused the characters to be missed; if the boundary character is invalid, it means that the deviation of this scan caused the characters to be scanned multiple times.
[0137] When the integrity of a boundary character in a valid boundary image frame indicates that the character is incomplete, the determination of whether the boundary character is a valid character is based on the rectangular enclosing boundary of the boundary character in the valid boundary image frame. First, the target proportion of the display area of the boundary character in the valid boundary image frame within the overall area of the boundary character needs to be determined based on the rectangular enclosing boundary of the boundary character in the valid boundary image frame. Then, the determination of whether the boundary character in the valid boundary image frame is a valid character is based on the target proportion corresponding to the boundary character in the valid boundary image frame.
[0138] The specific steps for determining the target proportion of the display area of the boundary characters in the effective boundary image frame relative to the overall area of the boundary characters, based on the rectangular enclosing boundary of the boundary characters in the effective boundary image frame, are as follows:
[0139] When the valid boundary image frame includes the pen stroke image frame, to calculate the target proportion corresponding to the first character in the pen stroke image frame, it is first necessary to calculate the width of the display area of the first character in the pen stroke image frame based on the coordinates of the upper left vertex of the rectangular enclosing boundary of the first character. That is, the overall width of the first character minus the horizontal distance between the upper left vertex of the rectangular enclosing boundary and the cutting position of the first character. The cutting position of the first character is the left boundary of the pen stroke image frame. For example, if the horizontal coordinate of the left boundary of the pen stroke image frame is set to 0, and the rectangular enclosing boundary data is [x, y, w, h] = [-30, 10, 60, 70], then the coordinates of the upper left vertex of the rectangular enclosing boundary are (x, y) = (-30, 10), the overall width of the first character is w = 60, and the horizontal distance w1 between the upper left vertex of the rectangular enclosing boundary and the cutting position of the first character is |-30-0| = 30. The width of the display area of the first character in the pen stroke image frame is w-w1 = 60-30 = 30. Finally, the ratio between the width of the display area of the first character in the image frame and the overall width of the first character is taken as the target proportion corresponding to the first character.
[0140] When the valid boundary image frame includes the pen-lifting image frame, to calculate the target proportion corresponding to the last character in the pen-lifting image frame, it is first necessary to calculate the width of the display area of the last character in the pen-lifting image frame based on the coordinates of the top-left vertex of the rectangular bounding boundary of the last character. That is, the horizontal distance between the top-left vertex of the rectangular bounding boundary and the cutting position of the last character, where the cutting position of the last character is the right boundary of the pen-lifting image frame. Finally, the ratio between the width of the display area of the last character in the pen-lifting image frame and the overall width of the last character is taken as the target proportion corresponding to the last character.
[0141] As an optional implementation method, see [link to implementation details]. Figure 5 As shown, another embodiment of this application discloses that, in step S302 of the above embodiment, determining whether a boundary character in a valid boundary image frame is a valid character based on the target proportion corresponding to the boundary character in the valid boundary image frame includes:
[0142] S501. If the target proportion corresponding to the boundary character in the valid boundary image frame is within the preset valid proportion range, then the boundary character in the valid boundary image frame is determined to be a valid character.
[0143] Specifically, in this embodiment, an effective proportion range and an invalid proportion range are preset. If most of the content of the boundary character is captured in the effective boundary image frame, it is considered a valid boundary character. If only a small part of the boundary character is captured in the effective boundary image frame, it is considered an invalid boundary character. Therefore, the effective proportion range set in this embodiment is a large range, and the invalid proportion range is a small range. The effective proportion range is preferably 0.9 to 1, and the invalid proportion range is preferably 0 to 0.1.
[0144] If the target proportion corresponding to the boundary character in the calculated valid boundary image frame is within the preset valid proportion range, it means that most of the content of the boundary character in the valid boundary image frame is displayed in the valid boundary image frame, and the boundary character in the valid boundary image frame is missed. Therefore, the boundary character in the valid boundary image frame is determined to be a valid character.
[0145] S502. If the target proportion corresponding to the boundary character in the valid boundary image frame is within the preset invalid proportion range, then the boundary character in the valid boundary image frame is determined to be an invalid character.
[0146] Specifically, if the target proportion corresponding to the boundary character in the calculated valid boundary image frame is within the pre-set invalid proportion range, it means that only a small part of the boundary character in the valid boundary image frame is displayed in the valid boundary image frame, and the boundary character in the valid boundary image frame is over-scanned. Therefore, the boundary character in the valid boundary image frame is determined to be an invalid character.
[0147] S503. If the target proportion corresponding to the boundary character in the valid boundary image frame is not within the valid proportion range, and the target proportion corresponding to the boundary character in the valid boundary image frame is not within the invalid proportion range, then the boundary character in the valid boundary image frame is determined to be a valid character based on the influence of the boundary character in the valid boundary image frame on the integrity of the scanning statement corresponding to the target image.
[0148] Specifically, if the target proportion corresponding to the boundary character in the calculated valid boundary image frame is neither within a pre-set valid proportion range nor within a pre-set invalid proportion range (if the valid proportion range is set to 0.9–1 and the invalid proportion range is set to 0–0.1), then when the target proportion corresponding to the boundary character in the valid boundary image frame is between 0.1 and 0.9, the boundary character in the valid boundary image frame is determined to be a valid character based on its impact on the integrity of the scanning statement corresponding to the spliced target image. Here, the scanning statement corresponding to the spliced target image mentioned in this step refers to the statement obtained after recognizing the target image, which does not contain the boundary character from the valid boundary image frame. This embodiment needs to determine the impact of the boundary character on the integrity of the scanning statement corresponding to the spliced target image based on the statement integrity when the boundary character from the valid boundary image frame is added to the scanning statement corresponding to the spliced target image and the statement integrity when the boundary character from the valid boundary image frame is not added to the scanning statement corresponding to the spliced target image.
[0149] Furthermore, in this step, the determination of whether a boundary character in a valid boundary image frame is a valid character is based on the impact of the boundary character in the valid boundary image frame on the integrity of the scanning statement corresponding to the target image. Specifically, this includes:
[0150] First, calculate the first statement integrity probability when the boundary characters in the valid boundary image frame are added to the scanning statement corresponding to the target image, and the second statement integrity probability when the boundary characters in the valid boundary image frame are not added to the scanning statement corresponding to the target image.
[0151] In this embodiment, to determine the impact of boundary characters in the valid boundary image frame on the integrity of the scanning statement corresponding to the target image, it is necessary to calculate the first statement integrity probability when the boundary characters in the valid boundary image frame are added to the scanning statement corresponding to the target image, and the second statement integrity probability when the boundary characters in the valid boundary image frame are not added to the scanning statement corresponding to the target image. This embodiment can utilize existing n-gram models or deep learning methods to calculate the statement integrity probability; the calculation process will not be described in detail here.
[0152] Second, if the ratio between the integrity probability of the first statement and the integrity probability of the second statement is greater than a preset ratio, then the boundary characters in the valid boundary image frame are determined to be valid characters.
[0153] This embodiment calculates the first statement integrity probability P when boundary characters in the effective boundary image frame are added to the scanning statement corresponding to the target image. wi The second statement integrity probability P when the boundary characters in the valid boundary image frame are not added to the scan statement corresponding to the target image wo Then, it is necessary to determine the probability P of the integrity of the first statement. wi With the probability P of the second statement integrity wo The ratio between them (i.e., P) wi / P wo If the probability of the first statement integrity is greater than the preset ratio β, then the probability of the first statement integrity is greater than the preset ratio β. wi With the probability P of the second statement integrity wo The ratio between them is greater than the preset ratio (i.e., P). wi / P wo If the ratio is greater than β), then the boundary character in the valid boundary image frame is determined to be a valid character. Here, the preset ratio β is a scaling factor, and in this embodiment, the preset ratio β is preferably set to 0.5.
[0154] Third, if the ratio between the integrity probability of the second statement and the integrity probability of the first statement is greater than a preset ratio, then the boundary characters in the valid boundary image frame are determined to be invalid characters.
[0155] This embodiment also requires determining the probability P of the integrity of the second statement. woWith the probability P of the first statement integrity wi The ratio between them (i.e., P) wo / P wi If the probability of the second statement's integrity is greater than the preset ratio β, then... wo With the probability P of the first statement integrity wi The ratio between them is greater than the preset ratio β (i.e., P). wo / P wi >β, which is P wi / P wo If <1 / β), then the boundary character in the valid boundary image frame is determined to be an invalid character.
[0156] If the probability of the first statement being complete is P wi With the probability P of the second statement integrity wo The ratio between them is not greater than the preset ratio β, and the probability P of the second statement integrity is... wo With the probability P of the first statement integrity wi The ratio between them is no greater than the preset ratio β, that is, 1 / β≤P wi / P wo If ≤β, it is considered that the scan can form complete content whether or not boundary characters are added. In this case, whether or not the boundary characters in the valid boundary image frame are retained has little impact on the scan result. At this time, the boundary characters in the valid boundary image frame can be regarded as valid characters or invalid characters. In this embodiment, it is preferred to regard the boundary characters in the valid boundary image frame as valid characters.
[0157] In this embodiment, the three steps S501, S502, and S503 are not distinguished by the order of execution.
[0158] As an optional implementation method, see [link to implementation details]. Figure 6 As shown in another embodiment of this application, the scanned image recognition method further includes:
[0159] S601. Extract sample scan image frame sequence from the historical scan image frame sequence corresponding to the historical scan within a preset time period.
[0160] Specifically, during the scanning process using the scanning device, this embodiment can calibrate the scanning device to improve the accuracy of scanned image recognition. Therefore, this embodiment needs to extract sample data for calibrating the scanning device, that is, extract sample scanned image frame sequences from the historical scanned image frame sequences corresponding to historical scans within a preset time period. Specifically, this embodiment can extract historical scanned image frame sequences corresponding to certain scans as sample scanned image frame sequences. This embodiment can also determine the historical stitched images corresponding to each scan in the historical scan based on the historical scanned image frames, and identify the scanned content of each historical image. If the user scans the same content multiple times consecutively, it indicates that there may be inconsistencies in human perception, such as missed scans or multiple scans, which leads to the user repeatedly scanning the same content. In this case, the historical scanned image frame sequences corresponding to these repeated scans are used as sample scanned image frame sequences to calibrate the scanning device. Furthermore, each extracted sample scanned image frame sequence carries the segmentation position corresponding to the sample image determined during scanning recognition.
[0161] S602. Determine the vertical offset of the sample image based on the text bounding box and the center position of the image of the sample image spliced from the sample scan image frame sequence.
[0162] Specifically, after obtaining several sample scan image frame sequences, the sample image formed by stitching together each sample scan image frame sequence is determined. Then, the text bounding box of each sample image is predicted. Based on the position data of each text bounding box in its corresponding sample image and the image center position of its corresponding sample image, the vertical offset of each sample image is determined, thereby enabling the determination of the vertical offset when scanning each sample scan image frame sequence.
[0163] S603. Determine the lateral offset of the sample image based on the segmentation position corresponding to the sample image spliced from the sample scan image frame sequence.
[0164] Specifically, based on the segmentation position corresponding to the sample image carried by each sample scan image frame sequence, the distance between the segmentation position and the boundary position of the effective boundary image frame in the sample scan image frame sequence is determined, and the lateral offset of each sample image is determined.
[0165] S604. Adjust the scanning parameters of the scanning device according to the horizontal offset and vertical offset of the sample image.
[0166] Based on the lateral and longitudinal offsets of each sample image, the scanning offset data is analyzed and statistically analyzed using methods such as Gaussian fitting. The resulting scanning offset data is then used as calibration parameters to adjust the scanning parameters of the scanning device. These scanning parameters include the acquisition parameters of the image acquisition component or the image processing parameters of the image acquisition component. The acquisition parameters of the image acquisition component can be its shooting angle, while the image processing parameters can be its effective imaging area, etc.
[0167] Corresponding to the above-described scanning image recognition method, this application also proposes a scanning image recognition device, see [link to relevant documentation]. Figure 7 As shown, the device includes:
[0168] The image frame determination module 100 is used to determine the effective boundary image frames from the acquired scanned images. The effective boundary image frames include the pen-dropping image frame at the moment of pen drop and / or the pen-lifting image frame at the moment of pen lift.
[0169] The segmentation determination module 110 is used to determine whether the boundary character in the valid boundary image frame is a valid character. If it is a valid character, the first character boundary of the boundary character is determined from the target image spliced from the acquired scanned image frames as the segmentation position. The first character boundary ensures that the image to be recognized obtained by segmenting from the target image according to the segmentation position contains the boundary character.
[0170] The segmentation determination module 110 is further configured to, if an invalid character is found, determine a second character boundary of the boundary character from the target image composed of the acquired scanned image frames, as the segmentation position; the second character boundary ensures that the image to be recognized obtained by segmenting from the target image according to the segmentation position does not contain the boundary character.
[0171] The recognition module 120 is used to perform character recognition on the image to be recognized obtained by segmenting the target image according to the segmentation position, and obtain the scanning result.
[0172] The scanning image recognition device proposed in this application uses an image frame determination module 100 to determine valid boundary image frames from the acquired scanning image; a segmentation determination module 110 determines whether the boundary characters in the valid boundary image frames are valid characters. If they are valid characters, the first character boundary of the boundary character is determined from the target image stitched from the acquired scanning image frames as the segmentation position; if they are invalid characters, the second character boundary of the boundary character is determined from the target image stitched from the acquired scanning image frames as the segmentation position; and a recognition module 120 performs character recognition on the image to be recognized obtained by segmenting from the target image according to the segmentation position to obtain the scanning result. The first character boundary ensures that the image to be recognized obtained by segmenting from the target image according to the segmentation position contains the boundary character, thus avoiding missed scanning of valid characters. The second character boundary ensures that the image to be recognized obtained by segmenting from the target image according to the segmentation position does not contain the boundary character, thus avoiding over-scanning of invalid characters, thereby improving the accuracy of the scanning image and consequently improving the accuracy of the recognition result.
[0173] As an optional implementation, another embodiment of this application also discloses that the image frame determination module 100 is specifically used for:
[0174] Illumination intensity is detected on the acquired scanned image frames. The first image frame in a continuous image frame sequence with illumination intensity within a preset stable range is taken as the pen-starting image frame, and / or the last image frame in a continuous image frame sequence with illumination intensity within a preset stable range is taken as the pen-lifting image frame.
[0175] As an optional implementation, another embodiment of this application also discloses that the scanned image recognition device further includes: a first extraction module.
[0176] The first extraction module is used to extract the image sequence to be stitched from the acquired scanned image frames based on the pen-starting image frame and the pen-lifting image frame; wherein the image sequence to be stitched includes at least all scanned image frames from the pen-starting image frame to the pen-lifting image frame.
[0177] The segmentation determination module 110 is also used to determine whether the boundary character in the valid boundary image frame is a valid character. If it is a valid character, the first character boundary of the boundary character is determined from the target image formed by splicing the scanned image frames in the image sequence to be spliced, and used as the segmentation position. If it is an invalid character, the second character boundary of the boundary character is determined from the target image formed by splicing the scanned image frames in the image sequence to be spliced, and used as the segmentation position.
[0178] As an optional implementation, another embodiment of this application also discloses that the first extraction module is specifically used for:
[0179] Extract all scanned image frames from the pen-starting image frame to the pen-lifting image frame from the acquired scanned image frames as the first image sequence;
[0180] According to a preset number of expansions, extract the image sequence adjacent to the first image sequence from the acquired scanned image frames as the expanded image sequence;
[0181] Following the image acquisition order of the extended image sequence and the first image sequence, the extended image sequence and the first image sequence are combined to obtain the image sequence to be stitched together.
[0182] As an optional implementation, another embodiment of this application also discloses a segmentation determination module 110, including a prediction unit and a character determination unit.
[0183] The prediction unit is used to detect the integrity of boundary characters in the valid boundary image frame and predict the rectangular enclosing boundary of the boundary characters in the valid boundary image frame.
[0184] The character determination unit is used to determine whether a boundary character in a valid boundary image frame is a valid character based on the integrity of the boundary characters in the valid boundary image frame and the rectangular enclosing boundary of the boundary characters in the valid boundary image frame.
[0185] In this context, the boundary character in the pen-starting image frame is the first character, and the boundary character in the pen-lifting image frame is the last character.
[0186] As an optional implementation, another embodiment of this application also discloses a prediction unit, specifically used for:
[0187] Extract a predetermined number of image frames before the pen-starting image frame from the acquired scanned image frames as auxiliary image frames for the pen-starting image frame, and / or extract a predetermined number of image frames after the pen-lifting image frame from the acquired scanned image frames as auxiliary image frames for the pen-lifting image frame.
[0188] Based on the attention mechanism, the image coding features corresponding to the effective boundary image frame are determined by using the image features of the auxiliary image frame of the effective boundary image frame and the image features of the effective boundary image frame.
[0189] Based on the image coding features corresponding to the valid boundary image frames, predict the integrity of the boundary characters and the rectangular enclosing boundary in the valid boundary image frames.
[0190] As an optional implementation, another embodiment of this application also discloses a character determination unit, specifically used for:
[0191] If the integrity of the boundary characters in a valid boundary image frame indicates that the characters are complete, then the boundary characters in the valid boundary image frame are determined to be valid characters.
[0192] If the integrity of the boundary character in the valid boundary image frame indicates that the character is incomplete, then the target proportion of the display area of the boundary character in the valid boundary image frame within the overall area of the boundary character is determined based on the rectangular enclosing boundary of the boundary character in the valid boundary image frame. Based on the target proportion corresponding to the boundary character in the valid boundary image frame, it is determined whether the boundary character in the valid boundary image frame is a valid character.
[0193] As an optional implementation, another embodiment of this application also discloses that the character determination unit determines whether a boundary character in a valid boundary image frame is a valid character based on the target proportion corresponding to the boundary character in the valid boundary image frame, including:
[0194] If the target proportion corresponding to the boundary character in the valid boundary image frame is within the preset valid proportion range, then the boundary character in the valid boundary image frame is determined to be a valid character.
[0195] If the target proportion corresponding to the boundary character in the valid boundary image frame is within the pre-set invalid proportion range, then the boundary character in the valid boundary image frame is determined to be an invalid character.
[0196] If the target proportion corresponding to the boundary character in the valid boundary image frame is not within the valid proportion range, and the target proportion corresponding to the boundary character in the valid boundary image frame is not within the invalid proportion range, then the boundary character in the valid boundary image frame is determined to be a valid character based on the impact of the boundary character in the valid boundary image frame on the integrity of the scanning statement corresponding to the target image.
[0197] As an optional implementation, another embodiment of this application also discloses that the character determination unit determines whether a boundary character in the valid boundary image frame is a valid character based on the influence of the boundary character in the valid boundary image frame on the integrity of the scanning statement corresponding to the target image, including:
[0198] Calculate the first statement integrity probability when the boundary characters in the valid boundary image frame are added to the scanning statement corresponding to the target image, and the second statement integrity probability when the boundary characters in the valid boundary image frame are not added to the scanning statement corresponding to the target image;
[0199] If the ratio between the integrity probability of the first statement and the integrity probability of the second statement is greater than a preset ratio, then the boundary character in the valid boundary image frame is determined to be a valid character.
[0200] If the ratio between the integrity probability of the second statement and the integrity probability of the first statement is greater than a preset ratio, then the boundary character in the valid boundary image frame is determined to be an invalid character.
[0201] As an optional implementation, another embodiment of this application also discloses that the segmentation determination module 110 is further used for:
[0202] If the valid boundary image frame includes the pen stroke image frame, then the left boundary of the rectangle surrounding the boundary character is determined from the target image stitched together from the acquired scanned image frames, and used as the cutting position.
[0203] If the valid boundary image frame includes the pen lifting image frame, then the right boundary of the rectangle surrounding the boundary character is determined from the target image stitched together from the acquired scanned image frames, and used as the cutting position.
[0204] As an optional implementation, another embodiment of this application also discloses that the segmentation determination module 110 is further used for:
[0205] If the valid boundary image frame includes the pen stroke image frame, then the right boundary of the rectangle surrounding the boundary character is determined from the target image stitched together from the acquired scanned image frames, and used as the cutting position.
[0206] If the valid boundary image frame includes the pen lifting image frame, then the left boundary of the rectangle surrounding the boundary character is determined from the target image stitched together from the acquired scanned image frames, and used as the cutting position.
[0207] As an optional implementation, another embodiment of this application also discloses that the scanned image recognition device further includes: a second extraction module, a first offset determination module, a second offset determination module, and an adjustment module.
[0208] The second extraction module is used to extract sample scan image frame sequences from the historical scan image frame sequences corresponding to historical scans within a preset time period.
[0209] The first offset determination module is used to determine the vertical offset of the sample image based on the text bounding box and the image center position of the sample image spliced from the sample scan image frame sequence.
[0210] The second offset determination module is used to determine the lateral offset of the sample image based on the segmentation position corresponding to the sample image spliced from the sample scan image frame sequence.
[0211] The adjustment module is used to adjust the scanning parameters of the scanning device based on the horizontal and vertical offsets of the sample image.
[0212] The scanning image recognition device provided in this embodiment belongs to the same concept as the scanning image recognition method provided in the above embodiments of this application. It can execute the scanning image recognition method provided in any of the above embodiments of this application and has the corresponding functional modules and beneficial effects for executing the scanning image recognition method. Technical details not described in detail in this embodiment can be found in the specific processing content of the scanning image recognition method provided in the above embodiments of this application, and will not be repeated here.
[0213] Another embodiment of this application discloses an electronic device, see [link to relevant documentation] Figure 8 As shown, the device includes:
[0214] Memory 200 and processor 210;
[0215] The memory 200 is connected to the processor 210 and is used to store programs;
[0216] The processor 210 is configured to implement the scanned image recognition method disclosed in any of the above embodiments by running the program stored in the memory 200.
[0217] Specifically, the aforementioned electronic device may also include: a bus, a communication interface 220, an input device 230, and an output device 240.
[0218] The processor 210, memory 200, communication interface 220, input device 230, and output device 240 are interconnected via a bus. Among them:
[0219] A bus can include a pathway for transmitting information between various components of a computer system.
[0220] The processor 210 can be a general-purpose processor, such as a general-purpose central processing unit (CPU), a microprocessor, etc., or an application-specific integrated circuit (ASIC), or one or more integrated circuits used to control the execution of the program of the present invention. It can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), an off-the-shelf programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0221] Processor 210 may include a main processor, as well as a baseband chip, modem, etc.
[0222] The memory 200 stores a program that executes the technical solution of this invention, and may also store an operating system and other key business functions. Specifically, the program may include program code, which includes computer operation instructions. More specifically, the memory 200 may include read-only memory (ROM), other types of static storage devices capable of storing static information and instructions, random access memory (RAM), other types of dynamic storage devices capable of storing information and instructions, disk storage, flash memory, etc.
[0223] Input device 230 may include a device for receiving user input data and information, such as a keyboard, mouse, camera, scanner, light pen, voice input device, touch screen, pedometer, or gravity sensor.
[0224] Output device 240 may include devices that allow information to be output to a user, such as a display screen, printer, speaker, etc.
[0225] The communication interface 220 may include a device that uses any transceiver to communicate with other devices or communication networks, such as Ethernet, Radio Access Network (RAN), Wireless Local Area Network (WLAN), etc.
[0226] The processor 210 executes the program stored in the memory 200 and calls other devices, which can be used to implement the various steps of the scanning image recognition method provided in the embodiments of this application.
[0227] Another embodiment of this application provides a storage medium storing a computer program, which, when executed by a processor, implements the various steps of the scanning image recognition method provided in any of the above embodiments.
[0228] For the foregoing method embodiments, in order to simplify the description, they are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, because according to this application, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0229] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For apparatus embodiments, since they are basically similar to method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.
[0230] The steps in the methods of the various embodiments of this application can be adjusted, combined, or deleted according to actual needs.
[0231] The modules and sub-modules in the various embodiments of the present application's devices and terminals can be merged, divided, and deleted according to actual needs.
[0232] It should be understood that the disclosed terminals, devices, and methods can be implemented in other ways, given the several embodiments provided in this application. For example, the terminal embodiments described above are merely illustrative. For instance, the division of modules or sub-modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple sub-modules or modules may be combined or integrated into another module, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or modules, and may be electrical, mechanical, or other forms.
[0233] The modules or submodules described as separate components may or may not be physically separate. The components that constitute a module or submodule may or may not be physical modules or submodules; that is, they may be located in one place or distributed across multiple network modules or submodules. Some or all of the modules or submodules can be selected to achieve the purpose of this embodiment's solution, depending on actual needs.
[0234] Furthermore, the functional modules or sub-modules in the various embodiments of this application can be integrated into one processing module, or each module or sub-module can exist physically separately, or two or more modules or sub-modules can be integrated into one module. The integrated modules or sub-modules described above can be implemented in hardware or in the form of software functional modules or sub-modules.
[0235] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0236] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software unit executed by a processor, or a combination of both. The software unit can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0237] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0238] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for recognizing scanned images, characterized in that, include: Illumination intensity is detected on the acquired scanned image frames. The first image frame in a continuous image frame sequence with illumination intensity within a preset stable range is taken as the pen-starting image frame, and / or the last image frame in a continuous image frame sequence with illumination intensity within a preset stable range is taken as the pen-lifting image frame. The pen-starting image frame and / or the pen-lifting image frame are used as valid boundary image frames; Determine whether the boundary character in the valid boundary image frame is a valid character. If it is a valid character, determine the first character boundary of the boundary character from the target image stitched together from the acquired scanned image frames, as the segmentation position. The first character boundary ensures that the image to be recognized obtained by segmenting the target image according to the segmentation position contains the boundary character. If the character is invalid, the second character boundary of the boundary character is determined from the target image stitched together from the acquired scanned image frames, and used as the segmentation position; the second character boundary ensures that the image to be recognized obtained by segmenting the target image according to the segmentation position does not contain the boundary character; Character recognition is performed on the image to be recognized obtained by segmenting the target image according to the segmentation position to obtain the scanning result.
2. The method according to claim 1, characterized in that, After performing illumination intensity detection on the acquired scanned image frames, and taking the first image frame in a continuous image frame sequence where the illumination intensity is within a preset stable range as the pen-starting image frame, and / or taking the last image frame in a continuous image frame sequence where the illumination intensity is within a preset stable range as the pen-lifting image frame, the method further includes: Based on the pen-starting image frame and the pen-removing image frame, extract the image sequence to be stitched from the acquired scanned image frames; wherein, the image sequence to be stitched includes at least all scanned image frames from the pen-starting image frame to the pen-removing image frame; Determining the first character boundary of the boundary character from the target image stitched together from the acquired scanned image frames, as the segmentation position, includes: The first character boundary of the boundary character is determined from the target image formed by stitching together scanned image frames from the image sequence to be stitched, and used as the segmentation position; Determining the second character boundary of the boundary character from the target image stitched together from the acquired scanned image frames, as the segmentation position, includes: The second character boundary of the boundary character is determined from the target image formed by stitching together scanned image frames from the image sequence to be stitched, and used as the cutting position.
3. The method according to claim 2, characterized in that, Based on the pen-starting image frame and the pen-removing image frame, extract the image sequence to be stitched from the acquired scanned image frames, including: Extract all scanned image frames from the pen-starting image frame to the pen-lifting image frame from the acquired scanned image frames as the first image sequence; According to a preset number of expansions, extract the image sequence adjacent to the first image sequence from the acquired scanned image frames as the expanded image sequence; According to the image acquisition order of the extended image sequence and the first image sequence, the extended image sequence and the first image sequence are combined to obtain the image sequence to be stitched together.
4. The method according to claim 1, characterized in that, Determining whether a boundary character in the valid boundary image frame is a valid character includes: The integrity of the boundary characters in the valid boundary image frame is detected, and the rectangular enclosing boundary of the boundary characters in the valid boundary image frame is predicted. Based on the integrity of the boundary characters in the valid boundary image frame and the rectangular enclosing boundary of the boundary characters in the valid boundary image frame, it is determined whether the boundary characters in the valid boundary image frame are valid characters. In this context, the boundary character in the pen-starting image frame is the first character, and the boundary character in the pen-lifting image frame is the last character.
5. The method according to claim 4, characterized in that, Detecting the integrity of boundary characters in the valid boundary image frame and predicting the rectangular enclosing boundary of the boundary characters in the valid boundary image frame includes: Extract a predetermined number of image frames before the pen-starting image frame from the acquired scanned image frames as auxiliary image frames for the pen-starting image frame, and / or extract a predetermined number of image frames after the pen-lifting image frame from the acquired scanned image frames as auxiliary image frames for the pen-lifting image frame. Based on the attention mechanism, the image coding features corresponding to the effective boundary image frame are determined by using the image features of the auxiliary image frame of the effective boundary image frame and the image features of the effective boundary image frame. Based on the image encoding features corresponding to the valid boundary image frame, predict the integrity of the boundary characters and the rectangular enclosing boundary in the valid boundary image frame.
6. The method according to claim 4, characterized in that, Determining whether a boundary character in a valid boundary image frame is a valid character based on the integrity of the boundary characters in the valid boundary image frame and the rectangular enclosing boundary of the boundary characters in the valid boundary image frame includes: If the integrity of the boundary characters in the valid boundary image frame indicates that the characters are complete, then the boundary characters in the valid boundary image frame are determined to be valid characters. If the integrity of the boundary character in the valid boundary image frame indicates that the character is incomplete, then based on the rectangular bounding boundary of the boundary character in the valid boundary image frame, the target proportion of the display area of the boundary character in the valid boundary image frame to the overall area of the boundary character is determined, and based on the target proportion corresponding to the boundary character in the valid boundary image frame, it is determined whether the boundary character in the valid boundary image frame is a valid character.
7. The method according to claim 6, characterized in that, Determining whether a boundary character in a valid boundary image frame is a valid character based on the target proportion corresponding to the boundary character in the valid boundary image frame includes: If the target proportion corresponding to the boundary character in the effective boundary image frame is within a preset effective proportion range, then the boundary character in the effective boundary image frame is determined to be a valid character. If the target proportion corresponding to the boundary character in the valid boundary image frame is within a preset invalid proportion range, then the boundary character in the valid boundary image frame is determined to be an invalid character. If the target proportion corresponding to the boundary character in the valid boundary image frame is not within the valid proportion range, and the target proportion corresponding to the boundary character in the valid boundary image frame is not within the invalid proportion range, then the boundary character in the valid boundary image frame is determined to be a valid character based on the influence of the boundary character in the valid boundary image frame on the integrity of the scanning statement corresponding to the target image.
8. The method according to claim 7, characterized in that, Determining whether a boundary character in the valid boundary image frame is a valid character based on its impact on the integrity of the scanning statement corresponding to the target image includes: Calculate the first statement integrity probability when the boundary characters in the effective boundary image frame are added to the scanning statement corresponding to the target image, and the second statement integrity probability when the boundary characters in the effective boundary image frame are not added to the scanning statement corresponding to the target image; If the ratio between the first statement integrity probability and the second statement integrity probability is greater than a preset ratio, then the boundary character in the valid boundary image frame is determined to be a valid character; If the ratio between the second statement integrity probability and the first statement integrity probability is greater than a preset ratio, then the boundary character in the valid boundary image frame is determined to be an invalid character.
9. The method according to claim 4, characterized in that, Determining the first character boundary of the boundary character from the target image stitched together from the acquired scanned image frames, as the segmentation position, includes: If the effective boundary image frame includes the pen stroke image frame, then the left boundary of the rectangular enclosing boundary of the boundary character is determined from the target image stitched together from the acquired scanned image frames, and used as the cutting position; If the effective boundary image frame includes the pen-lifting image frame, then the right boundary of the rectangular enclosing boundary of the boundary character is determined from the target image stitched together from the acquired scanned image frames, and used as the cutting position.
10. The method according to claim 4, characterized in that, Determining the second character boundary of the boundary character from the target image stitched together from the acquired scanned image frames, as the segmentation position, includes: If the effective boundary image frame includes the pen stroke image frame, then the right boundary of the rectangular enclosing boundary of the boundary character is determined from the target image stitched together from the acquired scanned image frames, and used as the cutting position; If the effective boundary image frame includes the pen-lifting image frame, then the left boundary of the rectangular enclosing boundary of the boundary character is determined from the target image stitched together from the acquired scanned image frames, and used as the cutting position.
11. The method according to claim 1, characterized in that, The method further includes: Extract sample scan image frame sequences from the historical scan image frame sequences corresponding to historical scans within a preset time period; The vertical offset of the sample image is determined based on the text bounding box and the image center position of the sample image spliced from the sample scan image frame sequence. The lateral offset of the sample image is determined based on the segmentation position corresponding to the sample image spliced from the sample scan image frame sequence. The scanning parameters of the scanning device are adjusted based on the horizontal and vertical offsets of the sample image.
12. A scanning image recognition device, characterized in that, include: The image frame determination module is used to detect the light intensity of the acquired scanned image frames, and to take the first image frame in the continuous image frame sequence where the light intensity is within a preset stable range as the pen-starting image frame, and / or to take the last image frame in the continuous image frame sequence where the light intensity is within a preset stable range as the pen-lifting image frame. The pen-starting image frame and / or the pen-lifting image frame are used as valid boundary image frames; The segmentation determination module is used to determine whether the boundary character in the valid boundary image frame is a valid character. If it is a valid character, the first character boundary of the boundary character is determined from the target image stitched together from the acquired scanned image frames as the segmentation position. The first character boundary ensures that the image to be recognized obtained by segmenting the target image according to the segmentation position contains the boundary character. The segmentation determination module is further configured to, if an invalid character is found, determine a second character boundary of the boundary character from the target image assembled from the acquired scanned image frames, as the segmentation position; the second character boundary ensures that the image to be recognized obtained by segmenting from the target image according to the segmentation position does not contain the boundary character; The recognition module is used to perform character recognition on the image to be recognized obtained by segmenting the target image according to the segmentation position, and obtain the scanning result.
13. An electronic device, characterized in that, include: Memory and processor; The memory is connected to the processor and is used to store programs; The processor is configured to implement the scanned image recognition method as described in any one of claims 1 to 11 by running a program in the memory.
14. A storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, implements the scanned image recognition method as described in any one of claims 1 to 11.
Citation Information
Patent Citations
Scanning method and related equipment thereof
CN113723420A
Auxiliary reading method and device
CN114550174A
Single-point character recognition method and device
CN114973255A