Optical character recognition method and device, electronic equipment and storage medium
By performing target segmentation and image block processing on the target image in optical character recognition technology, the problem of insufficient accuracy in the prior art when processing complex images is solved, and higher optical character recognition accuracy is achieved.
Patent Information
- Application Number
- CN202510005716.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-02
- Publication Date
- 2025-05-06
AI Technical Summary
Existing optical character recognition technologies are difficult to achieve ideal accuracy when processing complex and diverse images, especially when fuzzy fonts, special fonts, multiple language mixing, background interference or irregular text layout.
By acquiring the target image of the character area, target segmentation is performed to determine multiple targets, each target is further divided into image blocks, and optical character recognition is performed on these image blocks to improve the accuracy of recognition.
By segmenting and processing image blocks, this method can extract character features more accurately, reduce interference from adjacent characters, and improve the accuracy of optical character recognition.
Smart Images

Figure CN119942561A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of text recognition, and in particular to an optical character recognition method, device, electronic equipment and storage medium. Background Art
[0002] The core goal of optical character recognition is to accurately and efficiently convert the text contained in the image into a coded text form that can be recognized and processed by the machine. However, in terms of the current actual situation, optical character recognition technology still has obvious shortcomings in accuracy. When faced with complex and diverse images, such as those containing blurred handwriting, special fonts, mixed languages, background interference or irregular text layout, the existing optical character recognition system often finds it difficult to achieve the ideal accuracy. Therefore, how to improve the accuracy of optical character recognition is a problem that needs to be solved urgently. Summary of the invention
[0003] The embodiments of the present application provide an optical character recognition method, device, electronic device and storage medium, which improve the accuracy of optical character recognition.
[0004] In a first aspect, an embodiment of the present application provides an optical character recognition method, the method comprising:
[0005] Get the target image corresponding to the character area;
[0006] Segment the target image and determine n targets;
[0007] Determine the image block corresponding to each of the n targets to obtain n image blocks;
[0008] Perform optical character recognition on each of the n image blocks to obtain m characters; m is a positive integer less than or equal to n;
[0009] Determine the target text content based on m characters.
[0010] In a second aspect, an embodiment of the present application provides an optical character recognition device, the device comprising: an acquisition unit and a processing unit;
[0011] A processing unit, used for acquiring a target image corresponding to the character area;
[0012] An acquisition unit, used for segmenting the target image and determining n targets;
[0013] Determine the image block corresponding to each of the n targets to obtain n image blocks;
[0014] Perform optical character recognition on each of the n image blocks to obtain m characters; m is a positive integer less than or equal to n;
[0015] Determine the target text content based on m characters.
[0016] In a third aspect, an embodiment of the present invention provides an electronic device, comprising: a processor, a memory, a communication interface, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the processor so that the electronic device executes the method of the first aspect.
[0017] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and the computer program is executed by a processor to implement the method of the first aspect.
[0018] In a fifth aspect, an embodiment of the present invention provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, so that a computer executes the method of the first aspect.
[0019] The implementation of the present invention has the following beneficial effects:
[0020] It can be seen that the optical character recognition method described in the embodiment of the present invention includes: acquiring a target image corresponding to a character area, performing target segmentation on the target image, determining n targets, determining an image block corresponding to each of the n targets, obtaining n image blocks, performing optical character recognition on each of the n image blocks, and obtaining m characters; m is a positive integer less than or equal to n, and the target text content is determined based on the m characters, thereby improving the accuracy of optical character recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the implementation methods of the present application or the background technology, the drawings required for use in the implementation methods of the present application or the background technology will be described below.
[0022] Figure 1 is a flow chart of an optical character recognition method provided by an embodiment of the present application;
[0023] Figure 2 is a flow chart for determining m characters provided by an embodiment of the present application;
[0024] Figure 3 is a schematic structural diagram of an optical character recognition device provided in an embodiment of the present application;
[0025] Figure 4 It is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0026] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the implementation mode of the present application will be clearly and completely described below in conjunction with the drawings in the implementation mode of the present application. Obviously, the described implementation mode is only a part of the implementation mode of the present application, not all the implementation modes. Based on the implementation mode in the present application, all other implementation modes obtained by ordinary technicians in the field without creative work are within the scope of protection of the present application.
[0027] The terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally includes steps or units that are not listed, or optionally includes other steps or units inherent to these processes, methods, products or devices.
[0028] Reference to "embodiment" herein means that a particular feature, structure, or characteristic described in conjunction with the embodiment may be included in at least one embodiment of the present application. The appearance of the phrase in various locations in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0029] The following is an explanation of some professional terms involved in this application:
[0030] Optical Character Recognition (OCR) is a technology that converts text information in images into a computer-editable text format. It uses optical technology and computer technology to convert text information in images into a machine-editable text format. These images can be paper document images scanned by scanners, photos containing text taken by digital cameras, screenshots, etc. Its core is to use a series of complex algorithms and models to identify the contours, structures, strokes and other features of characters in the image, and then convert them into character codes that can be understood and processed by computers, and finally generate text files.
[0031] See also Figure 1 , Figure 1 is a flow chart of an optical character recognition method provided by an embodiment of the present application, including but not limited to the following steps:
[0032] S101: Acquire a target image corresponding to a character area.
[0033] In this embodiment, the character area refers to a specific part of the entire image that contains text content that requires optical character recognition. This area can be a regular shape, such as a rectangular text paragraph area on a document page, or an irregular shape, such as the part where the text on a sign in a photo is located. Its shape may change with the shape of the sign and the text layout. The boundary of the character area is usually determined by the boundary between the text and the surrounding non-text parts (such as blank areas, patterns, etc.).
[0034] The target image is the image portion corresponding to the character area extracted from the entire image. It is the image that contains the character area and is the direct object of subsequent optical character recognition processing. The size and shape of the target image depend on the size and shape of the character area, and the target image only contains the part of the image where the text information we care about is located. This can reduce the amount of data for subsequent processing and improve processing efficiency and accuracy.
[0035] In this embodiment, a pre-trained target detection model can be used to locate the character area. The target detection model is trained with a large amount of image data and can learn the characteristics and position of the text area. The semantic segmentation model can also be used to classify each pixel in the image into different categories, such as text category and non-text category. By inputting the image into the semantic segmentation model, the category label of each pixel is obtained, and then the pixels labeled as text category are extracted, the character area can be determined, and then the target image can be obtained.
[0036] S102: Segment the target image to determine n targets.
[0037] In this embodiment, the target image can be segmented using a connected region-based segmentation method. Specifically, first, scan each pixel of the target image and mark the pixels that have been visited. For unvisited pixels, if its grayscale value (or color value) meets the character condition (for example, in a binary image, a white pixel represents a character), a depth-first search or a breadth-first search is performed starting from the pixel, and all pixels connected to it are marked as the same group. This process is repeated until all pixels are visited and marked, thereby obtaining multiple different connected regions, that is, multiple targets are determined. The target image can also be segmented based on a watershed algorithm. Specifically, first, the gradient image of the target image is calculated, which indicates the severity of the change in the grayscale value of the image. The places with large gradients are usually the edges of the characters. Then, the minimum value point is found in the gradient image as the initial water injection point. As the water injection process (implemented by simulation in the algorithm), the image is segmented into different regions. In order to avoid over-segmentation, the gradient image usually needs to be preprocessed, such as marking some known target areas or performing gradient processing. Line correction; The target image can also be segmented by a method based on an adaptive threshold. Specifically, taking the adaptive threshold segmentation based on the local mean as an example, a suitable neighborhood size is selected. For each pixel in the target image, the mean of the pixels in its neighborhood is calculated, and then an offset is set (which can be a fixed value or dynamically adjusted according to the image characteristics). If the pixel is larger than the sum of the mean and the offset, it is divided into a character pixel, otherwise it is divided into a background pixel. In this way, the local features of different regions of the image are segmented to determine multiple targets; The target image can also be segmented by a convolutional neural network segmentation method. Specifically, first, a large amount of training data is required, including the target image and the corresponding pixel-level annotation (that is, the label of which target each pixel belongs to). Then, the convolutional neural network model is trained with these data, and the parameters of the model are adjusted by the back propagation algorithm so that the output of the model is as close to the real annotation as possible. During segmentation, the target image is input into the trained model. After obtaining the probability map, the pixels are classified according to the probability threshold to determine multiple targets. There are many choices for segmentation methods, which are not limited here. After the target image is segmented, n targets can be obtained.
[0038] It can be seen that in the target image, characters may be adhered to each other, overlapped or arranged closely. Through target segmentation, these characters or character groups can be separated to form relatively independent targets. For example, in handwritten text or some artistic font typesetting, there may be connected strokes or decorative connecting parts between characters. After the target image is segmented, each target contains only one or a group of relatively independent characters. In this way, when performing optical character recognition, the features of each character can be extracted more accurately, avoiding the interference of adjacent characters, thereby improving the recognition accuracy. Through segmentation, the target area where the characters are located can be clearly identified, and only these targets can be subjected to feature extraction and recognition operations, avoiding invalid calculations on irrelevant areas, thereby saving computing resources and improving processing efficiency.
[0039] S103: Determine an image block corresponding to each of the n targets to obtain n image blocks.
[0040] In this embodiment, illustratively, the contour of each of the n targets is first determined to obtain n contour sets. Specifically, first, an edge detection algorithm can be used to find the edge pixels of the target. For example, the Canny edge detection algorithm is a commonly used method, which determines the edge points by calculating the gradient of the image. For each pixel in the target image, the Canny algorithm calculates its grayscale change rate in the horizontal and vertical directions (or other directions, depending on the gradient operator used). When this change rate exceeds a certain threshold, the pixel is marked as an edge pixel. After obtaining the edge pixels, these edge pixels need to be connected to form a complete contour. A contour tracking algorithm can be used to implement this process. Adjacent edge pixels are searched in a certain direction (such as clockwise or counterclockwise) until returning to the starting pixel, thus forming a closed contour. For complex targets, there may be multiple unconnected edge parts, and these parts need to be contour tracked separately. For example, if the target is composed of several independent characters, the edge of each character needs to be tracked separately, and finally multiple contours are formed. The collection of these contours is the contour set for this target. For n targets, this process is repeated to obtain n contour sets.
[0041] Exemplarily, the image block corresponding to each of n contour sets is determined to obtain n image blocks. Specifically, for each contour set, the contours it contains surround the area where the target is located. The area can be determined by calculating the pixel range inside the contour. One method is to use a filling algorithm, starting from a point on the contour, and filling pixels inward until the contour boundary or the filled pixels are encountered. The filling method can be based on scan line filling or seed filling. Through the filling operation, all pixels inside the contour can be clearly defined. These pixels constitute the approximate area of the target in the image. The boundary of this area is the boundary defined by the contour, which determines the range of the image block corresponding to the target. After determining the area range corresponding to each contour set, the corresponding image block can be extracted from the target image. The specific operation is to copy these pixels from the target image according to the previously determined area range (such as the coordinate range of the pixel) to form a new image block. This image block is the target image block corresponding to the contour set. For all n contour sets, this extraction process is repeated to obtain n image blocks, which will serve as basic units for subsequent optical character recognition.
[0042] It can be seen that the image block corresponding to each target contains the character information related to the target. By determining these image blocks, it is possible to accurately focus on the characters in each target so as to better extract the features of the characters and improve the effectiveness of feature extraction. The processing methods and parameters can be flexibly adjusted or different recognition models can be used to process these problematic image blocks without reprocessing the entire target image, thereby improving the processing flexibility.
[0043] S104: Perform optical character recognition on each of the n image blocks to obtain m characters.
[0044] In this embodiment, m is a positive integer less than or equal to n. Figure 2 , Figure 2 A flowchart for determining m characters provided by an embodiment of the present application includes but is not limited to the following steps:
[0045] S201: Perform optical character recognition on each of the n image blocks to obtain n characters.
[0046] In this embodiment, first, preprocess the n image patches. If an image patch is in color, first convert it to a grayscale image because color information is usually not necessary for character recognition, and grayscaling can reduce the amount of data and simplify subsequent processing. The image patches may contain various noises, such as salt-and-pepper noise (randomly appearing black and white pixel points) or Gaussian noise (causing normal distribution changes in pixel values). To improve the recognition accuracy, noise reduction is required. For example, median filtering is a commonly used method that replaces the current pixel value with the median of the pixel values in the neighborhood around the pixel point, which is effective for removing salt-and-pepper noise. Gaussian filtering, on the other hand, performs weighted averaging on the pixel neighborhood according to the Gaussian distribution, and while removing Gaussian noise, can better preserve the shape and edge information of the characters. Then, convert the pixel values of the image patches into black and white according to a certain threshold, making the characters stand out clearly from the background. Then, find the edge pixels of the characters through an edge detection algorithm, and then analyze the connection of the edge pixels to determine the stroke endpoints and intersection points. For example, for the Chinese character "木", the endpoints and intersection points of the horizontal, vertical, left-falling and right-falling strokes can be detected. The positions and quantities of these feature points can help identify the characters. Then, extract the contour information of the characters, including the shape, length, curvature, etc. of the contour, analyze the direction and distribution of the character strokes. For Chinese characters, the direction vector of each stroke can be calculated, and the quantity and distribution of strokes in different directions can be counted. Then, calculate the pixel density of the character region, that is, the ratio of the character pixels to the total pixels of the entire image patch. Different characters usually have different pixel densities. For example, characters with simple strokes have lower pixel densities, while characters with complex strokes have higher pixel densities. Comparing the pixel densities can help distinguish different characters. For handwritten text, the stroke width may vary, and information such as the average width, maximum width, and minimum width of the character strokes in the image patch can be counted. Finally, match the features of the extracted image patches with a pre-established template library to obtain n characters. It should be noted that the pre-established template library is a collection of standard templates of various characters (such as all English letters, numbers, common Chinese characters, etc.) in advance. These templates can be ideal character shapes made manually or extracted from high-quality character samples, and can include forms such as grayscale images, binary images, or extracted feature vectors of the characters.
[0047] S202: Determine the credibility corresponding to each of the n characters to obtain n credibilities.
[0048] In this embodiment, the strength of the logical relationship between the first character and the adjacent characters is first determined according to the preset rules. The first character is any one of the n characters. Specifically, the strength of the logical relationship refers to the quantitative expression of the closeness of the relationship between the two characters in terms of semantics, grammar or context. It is used to measure the degree of mutual dependence and reasonable collocation of the two characters in a text unit (such as a word, phrase or sentence). The preset rules include part-of-speech collocation rules, sentence component relationship rules, fixed collocation rules and word meaning combination rules. Among them, the part-of-speech collocation rules include determining the part of speech of the words to which the first character and the adjacent characters belong (if they are Chinese characters, they may form a word). For example, in a sentence, the word where the first character belongs is an adjective, and the word where the adjacent character belongs is a noun. According to grammatical rules, adjectives can modify nouns. This collocation is grammatically reasonable and will be assigned a certain logical relationship strength score. For example, "beautiful (adjective) flowers (noun)", the strength of the logical relationship between "beautiful" and "flowers" based on part-of-speech collocation is high. The sentence component relationship rules include analyzing the roles of the two characters in the sentence structure. If the first character belongs to the subject part and the adjacent character belongs to the predicate part, then the relationship between them is There is a logical relationship between the performer of the action and the action. This relationship is important in the sentence structure and will have a high logical relationship strength. The fixed collocation rule includes checking whether the first character and the adjacent characters belong to a known fixed phrase or vocabulary combination. If the two characters are part of these fixed collocations, the logical relationship strength between them will be very high. The word meaning combination rule includes determining whether the semantics of the first character and the adjacent characters are accurate after collocation. For example, in "red car", "red" modifies "car", which semantically represents the color attribute of the car. This modification relationship has a certain logical relationship strength according to the semantic combination rule. The semantic roles of the two characters can be determined by semantic role labeling, such as agent, patient, tool, time, place, etc., and then the logical relationship strength can be judged. The word vector model can also be used to convert the vocabulary composed of characters into semantic vectors, and then the cosine similarity between vectors is calculated to measure the logical relationship strength. More complex language models can also be used to determine the strength of logical relationships. These models learn the deep semantic and grammatical relationships between characters during training. By calculating the joint probability or conditional probability of two characters in the language model, a score representing the strength of the logical relationship is obtained. For example, a sequence of two characters is input into the language model, and the probability value output by the model or the probability value after a certain transformation can be used as a measure of the strength of the logical relationship.
[0049] Exemplarily, if the logical relationship strength is less than the preset logical relationship strength, the credibility of the first character is determined to be 0. Specifically, first, it is necessary to pre-set a suitable threshold of the logical relationship strength, that is, the preset logical relationship strength, based on factors such as the specific application scenario, the expectation of character recognition accuracy, and the characteristics of the processed text. The logical relationship strength between the first character and the adjacent characters is compared with the preset logical relationship strength. If the calculated logical relationship strength is less than this preset value, it means that from the perspective of logical association, the combination of the first character and the adjacent characters is not reasonable, there may be recognition errors or the character does not conform to normal language logic here, so the credibility of the first character is directly determined to be 0, indicating that the recognition result of this character is basically unreliable.
[0050] Exemplarily, if the logical relationship strength is greater than or equal to the preset logical relationship strength, the credibility corresponding to the first character is determined based on the logical relationship strength. Specifically, the mapping relationship between the logical relationship strength and the credibility is first obtained. The mapping relationship between the logical relationship strength and the credibility can be constructed through prior experiments, statistical analysis, or based on existing linguistic knowledge. For example, after a large number of text sample tests, it is found that when the logical relationship strength is in a certain interval, the corresponding character credibility is usually in another corresponding interval, thereby forming a specific corresponding rule. This mapping relationship may be linear (such as the credibility increases by a fixed proportion for each increase in the logical relationship strength), or it may be nonlinear (such as in the stage of low logical relationship strength, the credibility increases slowly, and after reaching a certain strength, the credibility increases faster), or the general mapping relationship summarized by predecessors in similar text processing and character recognition fields is used for operation. Exemplarily, the reference credibility corresponding to the logical relationship strength is determined based on the mapping relationship. Specifically, according to the mapping relationship obtained previously, the credibility value corresponding to the logical relationship strength between the current first character and the adjacent character is found. This value is the reference credibility.
[0051] When the clarity of the image block is high, the structural features of the character such as the strokes and contours can be extracted more accurately. For example, in the optical character recognition process, for a clear image block, the stroke endpoints and intersections of the Chinese characters and the line direction and curve details of the English letters can be clearly distinguished. These accurately extracted features can better match the standard features in the character template or machine learning model, thereby improving the accuracy of character recognition and increasing the credibility of the characters. High clarity helps to avoid character misjudgment due to image blur or missing details. Low-definition image blocks may cause some features of the characters to be blurred or lost. Therefore, the clarity of the first image block corresponding to the first character is determined, and the first image block is the image block corresponding to the first character among the n image blocks, and then the target adjustment parameter corresponding to the clarity is determined. Specifically, it can be a mapping relationship between the preset clarity and the adjustment parameter, and the target adjustment parameter corresponding to the clarity can be determined based on the mapping relationship. Exemplarily, the reference credibility is adjusted based on the target adjustment parameter to obtain the credibility corresponding to the first character. Specifically, the specific calculation formula is as follows: credibility corresponding to the first character = reference credibility × (1 + target adjustment parameter). The credibility corresponding to the first character can be obtained according to the above formula.
[0052] It can be seen that by first evaluating the strength of the logical relationship between characters, unreasonable character combinations can be screened out from the semantic, grammatical and other language levels. Clear images can provide more reliable character features and have a positive impact on credibility, while blurred images will increase recognition uncertainty. Appropriately lowering credibility and considering the clarity of the image will make the credibility assessment more accurate. The final credibility is determined based on the adjustment parameters corresponding to the strength of the logical relationship and clarity, achieving fine adjustment of the credibility. For texts of poor quality, the credibility can be reasonably lowered to remind users of possible recognition problems. For texts of good quality, the credibility can be determined more accurately, improving the availability of recognition results.
[0053] S203: Determine m credibility levels among the n credibility levels that are greater than a preset credibility level.
[0054] In this embodiment, first, a suitable preset credibility number is determined in advance based on factors such as expectations for character recognition accuracy and past experience, and then the credibility corresponding to each of the n characters obtained is numerically compared with the preset credibility one by one, and all credibility greater than the preset credibility is selected. The number of these selected credibility is m, thereby obtaining m credibility.
[0055] S204: Determine m characters corresponding to m credibility levels.
[0056] In this embodiment, after m credibility levels greater than a preset credibility level are determined, characters corresponding to the m credibility levels are searched one by one according to the previously recorded corresponding relationships to obtain m characters.
[0057] S105: Determine the target text content based on the m characters.
[0058] In this embodiment, illustratively, the language type corresponding to each of m characters is determined to obtain m language types. Specifically, the corresponding language type can be determined based on factors such as the character's own encoding range, the character set to which it belongs, and the language environment in which it commonly appears. The language type can also be assisted in determining from aspects such as the character's appearance, structural characteristics, and the context in which it appears in the text. For example, some characters with writing style characteristics of a specific language, such as the unique letter form of Arabic and kana contained in Japanese, can easily be identified by observing their appearance. If in a text, a character is often surrounded by other characters that conform to the grammatical rules and common vocabulary of a certain language, then it can also be inferred that the character belongs to this language.
[0059] Exemplarily, the language type that appears most frequently among the m language types is determined to obtain the target language type. Specifically, a statistical analysis is performed on the m language types, the frequency of occurrence of each language type is calculated, the number of occurrences of various language types is compared, and the language type that appears most frequently is selected as the target language type.
[0060] Exemplarily, the language types of m characters are converted into the target language type to obtain m target characters. Specifically, if the language type to which the character originally belongs is the target language type, then the character remains unchanged. If the language type to which the character originally belongs is not the target language type, it can be translated into characters corresponding to the target language type with the help of translation tools or language models. For example, online translation tools (such as Google Translate, Baidu Translate, etc.) or machine translation models based on deep learning (these models have been trained with a large amount of bilingual or multilingual parallel corpora and can achieve conversion between languages) can be used.
[0061] Exemplarily, the target text content is obtained by combining m target characters. Specifically, the target characters are connected in sequence according to their position order in the original text, or based on a certain language logic (such as conforming to normal grammatical and semantic rules, etc.) to form a complete and semantically coherent text content. Sometimes, the text simply combined in sequence may not be perfect in terms of grammar and semantics. At this time, the combined text can be appropriately adjusted and optimized according to the grammatical rules and common expressions of the target language to make it more in line with normal language reading habits and semantic logic, and finally determine accurate and fluent target text content.
[0062] It should be noted that after determining the target text content, it is also necessary to determine whether there may be errors in the target text content. Therefore, k sentences in the target text content are first determined, where k is a positive integer less than n. Then, the k sentences are scored based on a preset language model to obtain k scores. Specifically, different types of language models can be used to score sentences. Common ones include statistical-based language models, neural network language models, etc. For example, the n-gram model counts the probabilities of character sequences of different lengths in the language, and measures the rationality score of the sentence by calculating the probability product of these character sequences in the sentence. The pre-trained language model based on the Transformer architecture (Bid i rect iona l Encoder Representations from Transformers, BERT) gives a score representing the quality of the sentence after the sentence is input into the model, based on the hidden layer representation output by the model or the prediction result after specific task training.
[0063] Exemplarily, the scores less than the preset score among the k scores are determined to obtain p scores, where p is an integer less than or equal to k. Specifically, the preset score can be set according to factors such as the specific application scenario and the expectation of sentence quality. For example, if you want to filter out sentences with relatively poor quality and possible errors for special attention, you may set the preset score to 0.5. This value can be set to a more appropriate threshold through multiple experiments and combined with manual judgment of sentence quality. After determining the preset score, the score less than the preset score among the k scores is determined to obtain p scores.
[0064] Exemplarily, a marking signal is generated based on p sentences corresponding to p scores, and the target user is prompted that the p sentences may have errors based on the marking signal. Specifically, the marking signal is an identifier used to indicate that a specific sentence may have problems, which can be a specific code, a Boolean value, or a text label. The target user can be prompted in a variety of ways. In a graphical user interface application, the sentence with the marking signal can be displayed in a special color, or a prompt icon can be displayed next to the sentence. When the user hovers the mouse over the icon, a prompt box pops up to indicate that "the sentence may have errors". In a command line interface, the sentences with the marking signal can be listed, and the corresponding prompt text can be added in front, and then the sentences can be listed line by line. If it is in a document processing software, annotations can also be added in the annotation function of the document to indicate that the corresponding sentence may have errors, etc., so that the target user can intuitively know which sentences may need further inspection and modification, thereby improving the accuracy of the text content.
[0065] It can be seen that by scoring the sentences separately, the quality of each sentence can be accurately located. Scoring based on the preset language model utilizes the language knowledge and rules obtained by training the language model on a large amount of text data. By comparing the sentence scores with the preset scores, it can effectively focus on those sentences with lower quality, which may not meet the language norms or have unclear semantic logic, generate marking signals and prompt users to sentences that may have errors, so that users can intuitively notice these problems.
[0066] In summary, the implementation of the present invention has the following beneficial effects:
[0067] It can be seen that the optical character recognition method described in the embodiment of the present invention includes: acquiring a target image corresponding to a character area, performing target segmentation on the target image, determining n targets, determining an image block corresponding to each of the n targets, obtaining n image blocks, performing optical character recognition on each of the n image blocks, and obtaining m characters; m is a positive integer less than or equal to n, and the target text content is determined based on the m characters, thereby improving the accuracy of optical character recognition.
[0068] See also Figure 3 , Figure 3 It is a structural schematic diagram of an optical character recognition device provided in an embodiment of the present application. The optical character recognition device 300 includes: an acquisition unit 301 and a processing unit 302;
[0069] An acquisition unit 301 is used to acquire a target image corresponding to a character area;
[0070] The processing unit 302 is used to segment the target image and determine n targets;
[0071] Determine the image block corresponding to each of the n targets to obtain n image blocks;
[0072] Perform optical character recognition on each of the n image blocks to obtain m characters; m is a positive integer less than or equal to n;
[0073] Determine the target text content based on m characters.
[0074] In some possible implementations, in determining an image block corresponding to each of the n targets to obtain the n image blocks, the processing unit 302 is specifically configured to:
[0075] Determine the contour of each of the n targets and obtain n contour sets;
[0076] An image block corresponding to each contour set in the n contour sets is determined to obtain n image blocks.
[0077] In some possible implementations, in performing optical character recognition on each of the n image blocks to obtain m characters, the processing unit 302 is specifically configured to:
[0078] Performing optical character recognition on each of the n image blocks to obtain n characters;
[0079] Determine the credibility of each character in the n characters to obtain n credibility;
[0080] Determine m credibility levels among n credibility levels that are greater than a preset credibility level;
[0081] Determine the m characters corresponding to the m credibility levels.
[0082] In some possible implementations, in determining the credibility corresponding to each of the n characters to obtain the n credibility, the processing unit 302 is specifically configured to:
[0083] Determine the strength of the logical relationship between the first character and the adjacent characters according to a preset rule; the first character is any one of the n characters;
[0084] If the logical relationship strength is less than the preset logical relationship strength, the credibility of the first character is determined to be 0;
[0085] If the logical relationship strength is greater than or equal to the preset logical relationship strength, the credibility corresponding to the first character is determined based on the logical relationship strength.
[0086] In some possible implementations, in determining the credibility corresponding to the first character based on the strength of the logical relationship, the processing unit 302 is specifically configured to:
[0087] Obtaining the mapping relationship between the strength of the logical relationship and the credibility;
[0088] Determine the reference credibility corresponding to the strength of the logical relationship based on the mapping relationship;
[0089] Determine the clarity of a first image block corresponding to a first character; the first image block is an image block corresponding to the first character among the n image blocks;
[0090] Determine target adjustment parameters corresponding to clarity;
[0091] The reference credibility is adjusted based on the target adjustment parameter to obtain the credibility corresponding to the first character.
[0092] In some possible implementations, in determining the target text content based on m characters, the processing unit 302 is specifically configured to:
[0093] Determine the language type corresponding to each of the m characters to obtain m language types;
[0094] Determine the language type that appears most frequently among the m language types and obtain the target language type;
[0095] Convert the language types of the m characters into the target language type to obtain m target characters;
[0096] Based on the m target characters, the target text content is obtained.
[0097] In some possible implementations, the processing unit 302 is further specifically configured to:
[0098] Determine k sentences in the target text content; k is a positive integer less than n;
[0099] Score k sentences based on the preset language model to obtain k scores;
[0100] Determine the score that is less than the preset score among the k scores, and obtain p scores; p is an integer less than or equal to k;
[0101] Generate a tag signal based on p sentences corresponding to p scores;
[0102] Based on the marked signal, the target user is prompted that p sentences may have errors.
[0103] See also Figure 4 , Figure 4 Schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. Figure 4 As shown, the electronic device 400 includes a transceiver 401, a processor 402 and a memory 403. They are connected via a bus 404. The memory 403 is used to store computer programs and data, and the transceiver 401 can transmit the data stored in the memory 403 to the processor 402. The above program includes instructions for executing the following steps:
[0104] Get the target image corresponding to the character area;
[0105] Segment the target image and determine n targets;
[0106] Determine the image block corresponding to each of the n targets to obtain n image blocks;
[0107] Perform optical character recognition on each of the n image blocks to obtain m characters; m is a positive integer less than or equal to n;
[0108] Determine the target text content based on m characters.
[0109] In some possible implementations, in terms of determining an image block corresponding to each of the n targets to obtain the n image blocks, the program includes instructions for executing the following steps:
[0110] Determine the contour of each of the n targets and obtain n contour sets;
[0111] An image block corresponding to each contour set in the n contour sets is determined to obtain n image blocks.
[0112] In some possible implementations, in terms of performing optical character recognition on each of the n image blocks to obtain m characters, the program includes instructions for executing the following steps:
[0113] Performing optical character recognition on each of the n image blocks to obtain n characters;
[0114] Determine the credibility of each character in the n characters to obtain n credibility;
[0115] Determine m credibility levels among n credibility levels that are greater than a preset credibility level;
[0116] Determine the m characters corresponding to the m credibility levels.
[0117] In some possible implementations, in determining the credibility corresponding to each of the n characters to obtain n credibility, the above program includes instructions for performing the following steps:
[0118] Determine the strength of the logical relationship between the first character and the adjacent characters according to a preset rule; the first character is any one of the n characters;
[0119] If the logical relationship strength is less than the preset logical relationship strength, the credibility of the first character is determined to be 0;
[0120] If the logical relationship strength is greater than or equal to the preset logical relationship strength, the credibility corresponding to the first character is determined based on the logical relationship strength.
[0121] In some possible implementations, in terms of determining the credibility of the first character based on the strength of the logical relationship, the program includes instructions for performing the following steps:
[0122] Obtaining the mapping relationship between the strength of the logical relationship and the credibility;
[0123] Determine the reference credibility corresponding to the strength of the logical relationship based on the mapping relationship;
[0124] Determine the clarity of a first image block corresponding to a first character; the first image block is an image block corresponding to the first character among the n image blocks;
[0125] Determine target adjustment parameters corresponding to clarity;
[0126] The reference credibility is adjusted based on the target adjustment parameter to obtain the credibility corresponding to the first character.
[0127] In some possible implementations, in determining the target text content based on m characters, the program includes instructions for performing the following steps:
[0128] Determine the language type corresponding to each of the m characters to obtain m language types;
[0129] Determine the language type that appears most frequently among the m language types and obtain the target language type;
[0130] Convert the language types of the m characters into the target language type to obtain m target characters;
[0131] Based on the m target characters, the target text content is obtained.
[0132] In some possible implementations, the above program includes instructions for performing the following steps:
[0133] Determine k sentences in the target text content; k is a positive integer less than n;
[0134] Score k sentences based on the preset language model to obtain k scores;
[0135] Determine the score that is less than the preset score among the k scores, and obtain p scores; p is an integer less than or equal to k;
[0136] Generate a tag signal based on p sentences corresponding to p scores;
[0137] Based on the marked signal, the target user is prompted that p sentences may have errors.
[0138] It should be understood that the electronic devices in this application may include optical character recognition devices, smart phones (such as Android phones, iOS phones, Windows Phone phones, etc.), tablet computers, PDAs, laptop computers, mobile Internet devices MID (Mobile Internet Devices, referred to as: MID) or wearable devices or servers, edge computing nodes, etc. The above electronic devices are only examples, not exhaustive, and include but are not limited to the above electronic devices.
[0139] The present application also provides a computer-readable storage medium, which stores a computer program. The computer program is executed by a processor to implement part or all of the steps of any optical character recognition method described in the above method implementation.
[0140] The present application also provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to cause a computer to execute part or all of the steps of any one of the optical character recognition methods described in the above method embodiments.
[0141] It should be noted that, for the above-mentioned various method implementations, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the implementations described in the specification are all optional implementations, and the actions and modules involved are not necessarily required by this application.
[0142] In the above-mentioned embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0143] In the several embodiments provided in this application, it should be understood that the disclosed device can be implemented in other ways. For example, the device implementation described above is only schematic, such as the division of units, which is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical or other forms.
[0144] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0145] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above integrated unit may be implemented in the form of hardware or in the form of a software program module.
[0146] If the integrated unit is implemented in the form of a software program module and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art or all or part of the technical solution, can be embodied in the form of a software product, which is stored in a memory and includes several instructions for a computer device (which can be a personal computer, server or network device, etc.) to execute all or part of the steps of each implementation method of the present application. The aforementioned memory includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, disk or optical disk and other media that can store program codes.
[0147] A person skilled in the art can understand that all or part of the steps in the various methods of the above-mentioned embodiments can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable memory, and the memory can include: a flash drive, a read-only memory (English: Read-Only Memory, abbreviated as: ROM), a random access memory (English: Random Access Memory, abbreviated as: RAM), a magnetic disk or an optical disk, etc.
[0148] The above is a detailed introduction to the implementation methods of the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above implementation methods is only used to help understand the method and core idea of the present application. At the same time, for general technical personnel in this field, according to the idea of the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. An optical character recognition method, characterized in that: The method comprises: Get the target image corresponding to the character area; Performing target segmentation on the target image to determine n targets; Determine an image block corresponding to each of the n targets to obtain n image blocks; Performing optical character recognition on each of the n image blocks to obtain m characters, where m is a positive integer less than or equal to n; The target text content is determined based on the m characters.
2. The method according to claim 1, characterized in that The step of determining an image block corresponding to each of the n targets to obtain n image blocks includes: Determine the contour of each of the n targets to obtain n contour sets; An image block corresponding to each contour set in the n contour sets is determined to obtain the n image blocks.
3. The method according to claim 2, characterized in that The performing optical character recognition on each of the n image blocks to obtain m characters comprises: Performing optical character recognition on each of the n image blocks to obtain n characters; Determine the credibility corresponding to each of the n characters to obtain n credibility; Determining m credibility levels among the n credibility levels that are greater than a preset credibility level; The m characters corresponding to the m credibility levels are determined.
4. The method according to claim 3, characterized in that The step of determining the credibility corresponding to each of the n characters to obtain n credibility comprises: Determine the strength of the logical relationship between the first character and adjacent characters according to a preset rule; the first character is any one of the n characters; If the logical relationship strength is less than the preset logical relationship strength, determining the credibility of the first character to be 0; If the logical relationship strength is greater than or equal to the preset logical relationship strength, the credibility corresponding to the first character is determined based on the logical relationship strength.
5. The method according to claim 4, characterized in that The determining the credibility corresponding to the first character based on the strength of the logical relationship includes: Acquire a mapping relationship between the logical relationship strength and the credibility; Determining a reference credibility corresponding to the strength of the logical relationship based on the mapping relationship; Determining the clarity of a first image block corresponding to the first character; the first image block is an image block corresponding to the first character among the n image blocks; determining a target adjustment parameter corresponding to the definition; The reference credibility is adjusted based on the target adjustment parameter to obtain the credibility corresponding to the first character.
6. The method according to any one of claims 2 to 5, characterized in that: The determining the target text content based on the m characters comprises: Determine the language type corresponding to each of the m characters to obtain m language types; Determine the language type that appears most frequently among the m language types to obtain the target language type; Convert the language types of the m characters into the target language type to obtain m target characters; The target text content is obtained by combining the m target characters.
7. The method according to claim 6, characterized in that The method further comprises: Determine k sentences in the target text content; k is a positive integer less than n; Scoring the k sentences based on a preset language model to obtain k scores; Determine the score that is less than the preset score among the k scores to obtain p scores; p is an integer less than or equal to k; Generate a label signal based on the p sentences corresponding to the p scores; The target user is prompted based on the marking signal that the p sentences may have errors.
8. An optical character recognition device, characterized in that: The device comprises: an acquisition unit and a processing unit; The processing unit is used to obtain a target image corresponding to the character area; The acquisition unit is used to segment the target image to determine n targets; Determine an image block corresponding to each of the n targets to obtain n image blocks; Performing optical character recognition on each of the n image blocks to obtain m characters, where m is a positive integer less than or equal to n; The target text content is determined based on the m characters.
9. An electronic device, characterized in that: The method comprises a processor, a memory, a communication interface and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the processor, and the one or more programs include instructions for executing the steps in the method described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program is executed by a processor to implement the method according to any one of claims 1 to 7.