Character recognition method and system based on language technology
Through the text recognition method based on language technology, the text areas in bank notes are filtered and segmented using fractal dimensions and topological invariants, the problem of low accuracy of handwritten amount numerals in complex scenarios is solved, and the high-precision text recognition effect is achieved.
Patent Information
- Application Number
- CN202510604376.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-05-12
AI Technical Summary
In bank notes containing anti-counterfeiting watermarks, miniature texts, and crease interference, how to accurately segment and identify overlapping, deformation, and low contrast handwritten amount numbers, while eliminating interference from anti-counterfeiting patterns, the existing OCR method has low recognition accuracy in complex scenarios.
Using a text recognition method based on language technology, text candidate areas are filtered through fractal dimensions, graph structure is constructed and topological invariants are calculated, topographic maps are generated for dynamic segmentation path planning, adhesion detection and elastic separation are performed, and character recognition results are finally converted into standardized text format.
It realizes accurate identification of handwritten amounts in complex scenarios, avoids interference from anti-counterfeiting watermarks, improves the accuracy and stability of text recognition, and is suitable for multiple types of bills.
Smart Images

Figure CN120126153A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of financial bill recognition, and in particular to a text recognition method and system based on language technology. Background Art
[0002] Bank bills are the core vouchers in financial transactions, including but not limited to checks, drafts, promissory notes, etc., and are mainly used for payment, settlement or transfer of rights. Bills themselves have the following characteristics: High security: Anti-counterfeiting technologies such as anti-counterfeiting watermarks, microtext, and security threads are used to prevent forgery.
[0003] Information density: It contains key fields such as amount, date, payee, etc., and some fields use microprinting (such as the tiny characters of the check amount).
[0004] Physical complexity: It is easily interfered by creases, stains, fading, etc., resulting in blurred or deformed handwritten or printed content.
[0005] Currently, for the text recognition of bank bills, text recognition technology (Optical Character Recognition, OCR) is usually used to convert the text information in the bill image into editable and retrievable text data. However, in scenarios such as non-uniform illumination, complex background interference, image blur or low resolution, the accuracy of character segmentation and feature extraction of traditional OCR methods significantly decreases. The anti-counterfeiting patterns of bills (such as background patterns, microtext) are easily misjudged as valid characters, resulting in background noise polluting the recognition result. Moreover, the handwritten content such as signatures and annotations has significant differences in font and strokes from printed numbers / letters, and traditional OCR methods are prone to failure.
[0006] Therefore, in bank bills containing anti-counterfeiting watermarks, microtext, and crease interference, how to accurately segment and recognize overlapping, deformed, and low-contrast handwritten amount numbers while excluding the interference of anti-counterfeiting patterns is an urgent problem to be solved and optimized in financial bill recognition. Summary of the Invention
[0007] The technical problem to be solved by the present invention is to provide a text recognition method and system based on language technology to solve the problem of accurately segmenting and recognizing handwritten amount numbers in bank bills containing anti-counterfeiting watermarks, microtext, and crease interference while excluding the interference of anti-counterfeiting patterns.
[0008] To solve the above technical problem, the technical solution of the present invention is as follows: In a first aspect, a text recognition method based on language technology, the method includes: Obtain a bank bill image, divide the bank bill image into grids, and screen according to the fractal dimension of each grid area to obtain a first text candidate area; Take the pixels of the first text candidate region as nodes, construct a graph structure with the edge weight between adjacent pixels being the gradient magnitude, and calculate topological invariants to eliminate abnormal regions based on the topological invariants, obtaining the second text candidate region; Generate a topographic map by performing brightness mapping based on the fractal dimension of the second text candidate region, and perform dynamic segmentation path planning on the topographic map to obtain the third text candidate region; Perform adhesion detection on the third text candidate region, and elastically separate the adhesion points to obtain the target text region; Perform character recognition on the target text region according to language technology, and convert the character recognition result into a standardized text format.
[0009] Further, obtain a bill image, divide the bill image into grids, and screen according to the fractal dimension of each grid region to obtain the first text candidate region, including: Convert the bank bill image into a grayscale image using the weighted average method, and divide the grayscale image into grids; Gradually reduce the sub-region size of each grid, count the number of the smallest boxes covering the stroke pixels, and calculate the fractal dimension of each grid; Calculate the mean and standard deviation of the fractal dimensions of all grids to determine the fractal dimension threshold; Screen the fractal dimension of each grid according to the fractal dimension threshold, mark the grids that meet the screening conditions as 1, and the rest as 0 to obtain the first text candidate region.
[0010] Further, take the pixels of the first text candidate region as nodes, construct a graph structure with the edge weight between adjacent pixels being the gradient magnitude, and calculate topological invariants to eliminate abnormal regions based on the topological invariants, obtaining the second text candidate region, including: Regard each pixel of the first text candidate region as a graph node, only connect the 4-neighborhood pixels of up, down, left, and right, take the gradient magnitude as the edge weight between adjacent pixels to obtain a graph structure, and use an adjacency matrix to store the weight relationship between nodes and edges; Count the number of connected regions in the graph structure through breadth-first search to obtain the number of connected components; Detect the closed loops in the graph structure through the matrix tree theorem to obtain the number of circular structures; Retain the regions with the number of connected components being 1 and the number of circular structures being 0 to obtain the second text candidate region.
[0011] Further, map the fractal dimension of each second text candidate region to a brightness value and generate a grayscale matrix, where each pixel value corresponds to a brightness value; Construct a pheromone concentration matrix according to the grayscale matrix, set a brightness threshold, assign a brightness level to each region, and initialize the pheromone concentration matrix according to the brightness level to obtain an initialized pheromone matrix; Update the pheromone of the initialized pheromone matrix according to the brightness gradient, move along the pheromone gradient, and generate a preliminary segmentation path; Calculate the path energy, and update the preliminary segmentation path according to the path energy to obtain the target segmentation path; Convert to a binary mask according to the target segmentation path to obtain the third text candidate region.
[0012] Furthermore, updating the pheromone of the initialized pheromone matrix according to the brightness gradient, moving along the pheromone gradient, and generating a preliminary segmentation path includes: Calculate the dynamic evaporation factor and the dynamic diffusion factor, and obtain the updated pheromone matrix according to the dynamic evaporation factor and the dynamic diffusion factor; Calculate the brightness gradient direction of each pixel, move along the gradient direction, and if moving into a low-brightness area, perform backtracking and re-plan the path to an adjacent high-brightness area; among them, when moving along the gradient direction, the step size is positively correlated with the pheromone concentration.
[0013] Furthermore, perform adhesion detection on the third text candidate region, and elastically separate the adhesion points to obtain the target text region, including: Use 8-neighborhood connectivity analysis on the third text candidate region, mark all connected regions, if the similarity of the boundary pixels of two adjacent connected regions is less than the similarity threshold, it is determined as adhesion, and an adhesion region marking matrix is obtained, and the adhesion pixel positions are marked as 1; Calculate the repulsive force magnitude of the adhesion pixel pairs, and calculate the attractive force for non-adhesion pixel pairs within the same grid, where the repulsive force magnitude is inversely proportional to the pixel distance, and the attractive force magnitude is inversely proportional to the pixel distance; For each adhesion pixel, calculate the vector sum of the repulsive force and the attractive force, and adjust the pixel position according to the vector sum, where if the pixel moves beyond the image boundary, it is pulled back within the boundary; If after 5 consecutive iterations, the energy change of the adhesion region is less than the threshold, or the maximum number of iterations is reached, stop the iteration to obtain the separated mask; Perform an opening operation on the separated mask, and perform connectivity verification. If the regional connectivity is insufficient, re-combine the adjacent high-brightness regions to obtain the target text region mask.
[0014] Furthermore, perform character recognition on the target text region according to language technology, and convert the character recognition result into a standardized text format, including: Apply morphological closing operation to the mask of the target text area, extract the brightness gradient features of the text area, and divide the text area into independent characters; Input the segmented characters into the OCR engine, binarize the character images, adjust the character sizes to the standard size, and obtain the initially recognized character sequence; Load a predefined format template according to the bill type, replace similar characters in the OCR recognition with the standard form, add thousand - separator commas according to the format template, and unify the letter case; Verify whether the text conforms to the format rules through regular expression matching, classify and store the text by fields, and generate a standardized text format.
[0015] In a second aspect, a text recognition system based on language technology includes: An acquisition module, configured to acquire a bank bill image, divide the bank bill image into grids, and filter according to the fractal dimension of each grid area to obtain a first text candidate area; A generation module, configured to use the pixels of the first text candidate area as nodes, construct a graph structure with the edge weight between adjacent pixels being the gradient amplitude, and calculate topological invariants to eliminate abnormal areas according to the topological invariants to obtain a second text candidate area; A calculation module, configured to generate a topographic map through brightness mapping according to the fractal dimension of the second text candidate area, and perform dynamic segmentation path planning on the topographic map to obtain a third text candidate area; An encryption module, configured to perform adhesion detection on the third text candidate area and elastically separate the adhesion points to obtain a target text area; A transmission module, configured to perform character recognition on the target text area according to language technology and convert the character recognition result into a standardized text format.
[0016] In a third aspect, a computing device includes: One or more processors; A storage system, configured to store one or more programs, which when executed by the one or more processors, cause the one or more processors to implement the above - mentioned method.
[0017] In a fourth aspect, a computer - readable storage medium stores a program, which when executed by a processor, implements the above - mentioned method.
[0018] The above - mentioned solutions of the present invention have at least the following beneficial effects: In the above solution of the present invention, by screening the candidate text regions through the fractal dimension and constructing a graph structure in combination with the brightness gradient, the adhered characters can be accurately distinguished, and the problem of adhered microtext can be solved; and while separating the adhered regions, the stroke continuity can be retained, avoiding over-segmentation caused by the morphological watershed algorithm; the high-brightness interference regions of the anti-counterfeiting watermark can be avoided, realizing the semantic-level separation of the anti-counterfeiting pattern and the text; the brightness threshold can be dynamically adjusted according to the type of bill to adapt to multiple types of bills, and the problem of insufficient text recognition accuracy in complex scenarios such as anti-counterfeiting watermarks, creases, and low contrast can be solved. Description of the Drawings
[0019] Figure 1 is a schematic flowchart of a text recognition method based on language technology provided by an embodiment of the present invention.
[0020] Figure 2 is a schematic diagram of a text recognition system based on language technology provided by an embodiment of the present invention. Detailed Embodiments
[0021] Hereinafter, exemplary embodiments of the present disclosure will be described in more detail with reference to the drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be completely conveyed to those skilled in the art.
[0022] As Figure 1 shown, an embodiment of the present invention proposes a text recognition method based on language technology, and the method includes: Step 1, obtaining a bank bill image, dividing the bank bill image into grids, and screening according to the fractal dimension of each grid region to obtain a first candidate text region; Step 2, taking the pixels of the first candidate text region as nodes, constructing a graph structure with the edge weight between adjacent pixels being the gradient amplitude, and calculating topological invariants to eliminate abnormal regions according to the topological invariants to obtain a second candidate text region; Step 3, generating a topographic map by brightness mapping according to the fractal dimension of the second candidate text region, and performing dynamic segmentation path planning on the topographic map to obtain a third candidate text region; Step 4, performing adhesion detection on the third candidate text region, and elastically separating the adhered parts to obtain a target text region; Step 5, performing character recognition on the target text region according to language technology, and converting the character recognition result into a standardized text format.
[0023] In the text recognition method based on language technology according to the embodiments of the present invention, the image is decomposed into local regions through grid division. Combining with fractal dimension screening, candidate text regions can be quickly located, background interference can be reduced, and the fractal dimension is sensitive to complex textures, which can effectively distinguish text from anti-counterfeiting watermarks, wrinkles and other low-complexity noise regions. Through topological invariants, abnormal structures such as isolated points and broken regions can be identified, non-text interferences (such as stains and creases) can be removed, and highly connected text regions can be retained. The topographic map generated by the fractal dimension reflects the regional complexity, and then dynamic path planning can preferentially cover high-complexity regions (such as dense text areas), improving the segmentation accuracy. Moreover, the fractal topographic map has tolerance for uneven illumination and perspective distortion and is applicable to complex bank bills. Through adhesion detection and elastic separation, the adhesion regions of strokes (such as the connected strokes of "1" and "2") can be separated, and independent characters can be retained, and complex adhesions can be processed without manual intervention. Furthermore, through language technology, the character recognition results can be converted into a standardized text format for storage.
[0024] In another optional embodiment of the present invention, in step 1 above, obtaining a bill image, performing grid division on the bill image, and screening according to the fractal dimension of each grid region to obtain a first text candidate region, including: Step 11, converting the bank bill image into a grayscale image using the weighted average method and dividing the grayscale image into grids; Step 12, gradually reducing the size of the sub-regions for each grid, counting the minimum number of boxes covering the stroke pixels, and calculating the fractal dimension of each grid; Step 13, calculating the mean and standard deviation of the fractal dimensions of all grids to determine the fractal dimension threshold; Step 14, screening the fractal dimension of each grid according to the fractal dimension threshold, marking the grids that meet the screening conditions as 1, and the rest as 0 to obtain the first text candidate region.
[0025] In this embodiment, in step 11, after the bank bill image is converted into a grayscale image, Gaussian filtering (kernel size 3×3) needs to be applied to eliminate noise, and then the number of grids can be adjusted according to the complexity of the bill. For example, for an image of 1024×768, 128×96 grids are obtained after division.
[0026] In step 12, through the formula , calculate the fractal dimension D of each grid, where is the reduced pixel, is the minimum number of boxes covering the stroke pixels. Preferably, pixels.
[0027] In step 14, the fractal dimension threshold can be set to a fixed value empirically. Due to the different complexities of different business bills of banks, in this embodiment, it is preferably determined according to the mean and standard deviation of the fractal dimensions of all current grids. Thus, the grids with fractal dimensions less than the fractal dimension threshold are marked as text candidates, denoted as 1, to obtain the first text candidate region.
[0028] In the text recognition method based on language technology described in the embodiment of the present invention, since the text regions in bank bills usually have dense and regular strokes and a low fractal dimension; the anti-counterfeiting patterns (such as microtext) are randomly distributed and have a high fractal dimension; therefore, retaining the regions with a low fractal dimension as text candidates and removing the high-dimensional regions can preliminarily screen the text regions in the bill. And in step 11, the Sobel operator can be used to detect the edges of the image, calculate the gradient intensity of each pixel point, and refine the edges through non-maximum suppression to remove the cluttered texture noise and suppress the refined edges.
[0029] In another alternative embodiment of the present invention, in step 2 above, the pixels of the first text candidate region are used as nodes, and the edge weights between adjacent pixels are the gradient amplitudes to construct a graph structure, and the topological invariant is calculated to remove the abnormal regions according to the topological invariant to obtain the second text candidate region, including: Step 21, regarding each pixel of the first text candidate region as a graph node, only connecting the 4-neighborhood pixels of up, down, left, and right, taking the gradient amplitude as the edge weight between adjacent pixels to obtain a graph structure, and using an adjacency matrix to store the weight relationship between nodes and edges; Step 22, counting the number of connected regions in the graph structure through breadth-first search to obtain the number of connected components; Step 23, detecting the closed loops in the graph structure through the matrix tree theorem to obtain the number of circular structures; Step 24, retaining the regions with the number of connected components being 1 and the number of circular structures being 0 to obtain the second text candidate region.
[0030] In this embodiment, the normal handwritten amount should be a single connected region. If it is split into multiple isolated parts, it may be noise or an anti-counterfeiting pattern; and the normal text region should have no closed loops. If there are closed loops, it means there is adhesion or redundant links, and it is marked as an abnormal region for removal. Further, when removing the abnormal regions, morphological closing operation (first dilating and then eroding, kernel size 3×3) can be used to merge slightly adhered strokes, and then morphological opening operation (first eroding and then dilating, kernel size 5×5) can be used to remove isolated noises (such as paper creases), so as to retain the regions that are not marked as abnormal and have been morphologically optimized, which are the second text candidate regions.
[0031] In the text recognition method based on language technology according to the embodiments of the present invention, by constructing the pixels of the text candidate region into a 4-neighborhood connected graph structure, using the gradient magnitude as the edge weight, and storing the node relationship with an adjacency matrix, an efficient modeling of the text region features is achieved; by quickly counting the number of connected components through breadth-first search, combining with the matrix tree theorem to accurately detect the closed-loop structure, an abnormal region elimination mechanism based on topological invariants is formed, and the graph structure simplifies the adjacency relationship of complex textures (only retaining the 4-neighborhood of up, down, left, and right), the gradient weight strengthens the edge information representation ability, and the breadth-first search ensures the linear time complexity of connectivity analysis, which can screen out pure text regions without adhesion and isolated loops. While retaining the continuity of the text strokes, it can effectively eliminate the circular interference of anti-counterfeiting watermarks (such as the closed structure of "VOID") and isolated points of background noise, and significantly improve the extraction accuracy of text candidate regions of high-density printed bills (such as invoice details) compared with traditional morphological methods.
[0032] In another alternative embodiment of the present invention, in step 3 above, generating a topographic map by performing brightness mapping according to the fractal dimension of the second text candidate region, and performing dynamic segmentation path planning on the topographic map to obtain a third text candidate region, including: Step 31, mapping the fractal dimension of each second text candidate region to a brightness value, and generating a grayscale matrix, where each pixel value corresponds to a brightness value; Step 32, constructing a pheromone concentration matrix according to the grayscale matrix, setting a brightness threshold, assigning a brightness level to each region, and initializing the pheromone concentration matrix according to the brightness level to obtain an initialized pheromone matrix; Step 33, updating the pheromone of the initialized pheromone matrix according to the brightness gradient, moving along the pheromone gradient, and generating a preliminary segmentation path; Step 34, calculating the path energy, and updating the preliminary segmentation path according to the path energy to obtain a target segmentation path; Step 35, converting to a binary mask according to the target segmentation path to obtain a third text candidate region.
[0033] In this embodiment, in step 31 above, when mapping the brightness value according to the fractal dimension, the higher the fractal dimension, the lower the brightness value. Furthermore, in step 32, the brightness threshold can be preset according to experience, so as to divide different brightness values into three levels: high brightness, medium brightness, and low brightness. Among them, the high-brightness region is a low-fractal-dimension region (such as the main body of the text), the medium-brightness region is a transition region where the fractal dimension is close to the brightness threshold (such as the edge of microtext), and the low-brightness region is a high-fractal-dimension region (such as anti-counterfeiting patterns or noise); further, different brightness levels correspond to fixed pheromone concentrations, so as to initialize the pheromone concentration matrix to obtain an initialized pheromone matrix.
[0034] In the text recognition method based on language technology according to the embodiments of the present invention, the regional complexity is quantified by the fractal dimension, the luminance gradient guides the path to fit the text edge, and the pheromone concentration is dynamically adjusted to enhance the weight of the highlighted area, which can significantly improve the accuracy of text area extraction for high-density printed bills (such as invoices and contracts); the gradient update mechanism can avoid stroke breakage caused by morphological operations, and the anti-noise ability of energy optimization is enhanced, which can effectively suppress the interference of anti-counterfeiting watermarks (such as the closed-loop structure of "VOID") and background noise adhesion. Furthermore, the generated binary mask can directly adapt to the OCR engine to realize the full-process automation processing from the candidate area to the structured text.
[0035] In another optional embodiment of the present invention, step 33, updating the pheromone of the initialized pheromone matrix according to the luminance gradient and moving along the pheromone gradient to generate a preliminary segmentation path, includes: Step 331, calculating the dynamic evaporation factor and the dynamic diffusion factor, and obtaining the updated pheromone matrix according to the dynamic evaporation factor and the dynamic diffusion factor; Step 332, calculating the luminance gradient direction of each pixel, moving along the gradient direction, and if moving into a low-luminance area, performing backtracking and re-planning the path to an adjacent high-luminance area; wherein, when moving along the gradient direction, the step size is positively correlated with the pheromone concentration.
[0036] In another optional embodiment of the present invention, step 331 includes: Step 3311, through the formula: , calculating the updated pheromone matrix ; wherein, is the balance coefficient of the evaporation term, is the balance coefficient of the diffusion term, and ; is the pixel value at the position in the image, is the luminance gradient amplitude vector, and ; is the dynamic evaporation factor; is the dynamic diffusion factor; Step 3312, through the formula: , calculating the dynamic diffusion factor , wherein, is the medium luminance threshold, and K is the steepness of the Sigmoid function; Step 3313, through the formula: , calculating the dynamic evaporation factor , wherein, is the standard evaporation rate of the high-luminance area, is the standard evaporation rate of the low-luminance area.
[0037] In this embodiment, the above parameters are preferably 0.9, 0.3, 150, K is 5, 0.6, 0.4. Through the above formula calculation, for the high-brightness area, is approximately equal to , the volatilization is extremely slow, the pheromone accumulates, is approximately equal to 0, the pheromone hardly diffuses, thus forming a high-concentration pheromone area to attract path coverage; for the medium-brightness area, smoothly transitions to a medium rate, the pheromone volatilizes moderately, is approximately equal to 1, the pheromone diffuses to the surroundings, connecting the high / low brightness areas, thus serving as a transition area to balance the path extension between the text and the anti-counterfeiting area; for the low-brightness area, is approximately equal to , the volatilization is extremely fast, the pheromone decays rapidly, is approximately equal to 0, the pheromone hardly diffuses, thus inhibiting the pheromone concentration and preventing the path from entering the anti-counterfeiting interference area.
[0038] In this embodiment, step 332 first calculates the brightness gradient direction angle through the brightness grayscale matrix to obtain the area where the brightness rises fastest, and then starting from the current pixel position, according to the gradient direction, through the formula: , calculates the step size, where is the basic step size coefficient, is the gradient amplitude; thus, the higher the pheromone concentration, the longer the step size, covering wider strokes; furthermore, it is necessary to check whether the pixel position is within the image boundary after moving; if the moved pixel is in the low-brightness area, it will return to the original position along the reverse direction of the original gradient, and select the adjacent pixel with the largest gradient amplitude (i.e., the area closest to the text edge) among the adjacent pixels, and recalculate the gradient direction; in path planning, the termination conditions include three consecutive backtracks, indicating that the path cannot be extended and the maximum number of iterations is reached.
[0039] In another alternative embodiment of the present invention, in the above step 34, the path energy is calculated, and the preliminary segmentation path is updated according to the path energy to obtain the target segmentation path, including: Step 341, through the formula: , calculates the path energy H(P), where P is the preliminary segmentation path, is the topological penalty weight, is the brightness penalty weight, is the pixel 's annular structure number, is the pixel 's brightness value, is the fracture penalty coefficient, and S(P) is the number of break points on path P; Step 342: Aggregate each preliminary segmentation path to generate a set of preliminary segmentation paths, calculate the path energy of each path, and fine-tune the step size of the current path to generate a new path; Step 343: Calculate the energy of the new path. If the energy of the new path is less than the energy of the current path, directly select the new path and update the current path; if the energy of the new path is not less than the energy of the current path, then use the formula: , calculate the selection probability, and decide whether to select the new path based on the selection probability. If the new path is selected, update the current path; where is the new path, is the current path, is the Boltzmann constant, is the temperature decay coefficient; Step 344: Repeat the process to obtain the globally optimal path, whose path energy is the minimum value.
[0040] In this embodiment, after calculating the selection probability, generate a uniformly distributed random number , if the random number is greater than the selection probability, reject the new path and retain the current path; otherwise, select the new path and update the current path.
[0041] In another alternative embodiment of the present invention, in the above step 4, perform adhesion detection on the third text candidate region and elastically separate the adhesion points to obtain the target text region, including: Step 41: Use 8-neighborhood connectivity analysis on the third text candidate region to mark all connected regions. If the similarity of the boundary pixels of two adjacent connected regions is less than the similarity threshold, it is determined as adhesion, and an adhesion region marking matrix is obtained, and the adhesion pixel positions are marked as 1; Step 42: Calculate the repulsive force magnitude of the adhesion pixel pairs, and calculate the attractive force for non-adhesion pixel pairs within the same grid. Among them, the repulsive force magnitude is inversely proportional to the pixel distance, and the attractive force magnitude is inversely proportional to the pixel distance; Step 43: For each adhesion pixel, calculate the vector sum of the repulsive force and the attractive force, and adjust the pixel position according to the vector sum. Among them, if the pixel moves beyond the image boundary, pull it back inside the boundary; Step 44: If after 5 consecutive iterations, the energy change of the adhesion region is less than the threshold, or the maximum number of iterations is reached, stop the iteration to obtain the separated mask; Step 45: Perform an opening operation on the separated mask and perform connectivity verification. If the regional connectivity is insufficient, re-merge the adjacent high-brightness regions to obtain the target text region mask.
[0042] In the text recognition method based on language technology according to the embodiments of the present invention, the adhesion area is accurately recognized through 8-neighborhood connectivity analysis, the adhesion pixels are dynamically separated by combining the repulsive force and attractive force models, efficient segmentation can be achieved by using boundary constraints and iterative optimization, and the integrity of the text area is ensured through morphological post-processing; it can prevent stroke breakage caused by pixel overflow, and can adapt to the adhesion scenarios of anti-counterfeiting watermark interference (such as the closed-loop structure of "VOID") and background noise.
[0043] In another alternative embodiment of the present invention, in step 5 above, character recognition is performed on the target text area according to language technology, and the character recognition result is converted into a standardized text format, including: Step 51, applying a morphological closing operation to the mask of the target text area, extracting the brightness gradient features of the text area, and dividing the text area into independent characters; Step 52, inputting the segmented characters into an OCR engine, binarizing the character images, and adjusting the character sizes to the standard size to obtain a preliminarily recognized character sequence; Step 53, loading a predefined format template according to the bill type, replacing similar characters in the OCR recognition with standard forms, adding thousand-separator symbols according to the format template, and unifying the case of letters; Step 54, verifying whether the text conforms to the format rules through regular expression matching, classifying the text by fields and storing it, and generating a standardized text format.
[0044] In the text recognition method based on language technology according to the embodiments of the present invention, independent characters are accurately segmented through morphological closing operation and gradient feature extraction, and standardized output is achieved by combining the OCR engine and the format template, so as to realize efficient conversion from the text area to the structured text.
[0045] As Figure 2 shown, the present application also provides a text recognition system based on language technology, and the system includes: An acquisition module 10, configured to acquire a bank bill image, divide the bank bill image into grids, and screen according to the fractal dimension of each grid area to obtain a first text candidate area; A generation module 20, configured to use the pixels of the first text candidate area as nodes, construct a graph structure with the edge weight between adjacent pixels being the gradient amplitude, and calculate topological invariants to eliminate abnormal areas according to the topological invariants to obtain a second text candidate area; A calculation module 30, configured to generate a topographic map by performing brightness mapping according to the fractal dimension of the second text candidate area, and perform dynamic segmentation path planning on the topographic map to obtain a third text candidate area; An encryption module 40, configured to perform adhesion detection on the third text candidate region, and perform elastic separation on the adhesion part to obtain a target text region; A transmission module 50, configured to perform character recognition on the target text region according to language technology, and convert the character recognition result into a standardized text format.
[0046] It should be noted that this system corresponds to the above method. All implementation manners in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.
[0047] An embodiment of the present invention further provides a computing device, including: a processor and a memory storing a computer program. When the computer program is run by the processor, the above-mentioned method is executed. All implementation manners in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.
[0048] An embodiment of the present invention further provides a computer-readable storage medium storing instructions. When the instructions are run on a computer, the computer is made to execute the above-mentioned method. All implementation manners in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.
[0049] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0050] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described system, system, and unit can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.
[0051] In the embodiments provided by the present invention, it should be understood that the disclosed system and method can be implemented in other ways. For example, the system embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the system or unit can be in an electrical, mechanical, or other form.
[0052] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0053] In addition, in each embodiment of the present invention, each functional unit may be integrated in a processing unit, may exist separately as individual physical units, or two or more units may be integrated in one unit.
[0054] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, ROM, RAM, magnetic disks, or optical discs that can store program codes.
[0055] In addition, it should be noted that in the systems and methods of the present invention, obviously, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of the present invention. And, the steps of performing the above series of processes can naturally be executed in chronological order according to the described order, but it is not necessary to execute them in chronological order. Some steps can be executed in parallel or independently of each other. For those of ordinary skill in the art, it is understandable that all or any steps or components of the methods and systems of the present invention can be implemented in any computing system (including processors, storage media, etc.) or in a network of computing systems in the form of hardware, firmware, software, or a combination thereof, which can be achieved by those of ordinary skill in the art using their basic programming skills after reading the description of the present invention.
[0056] Accordingly, the object of the present invention can also be achieved by running a program or a set of programs on any computing system. The computing system can be a well-known general-purpose system. Therefore, the object of the present invention can also be achieved merely by providing a program product containing program code for implementing the method or system. That is to say, such a program product also constitutes the present invention, and a storage medium storing such a program product also constitutes the present invention. Obviously, the storage medium can be any well-known storage medium or any storage medium developed in the future. It should also be noted that in the system and method of the present invention, obviously, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of the present invention. Moreover, the steps of performing the above series of processes can naturally be executed in chronological order according to the described order, but it is not necessary to execute them in chronological order. Certain steps can be executed in parallel or independently of each other.
[0057] The above is the preferred embodiment of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.
Claims
1. A text recognition method based on language technology, characterized in that: The method comprises: Obtain a bank note image, divide the bank note image into grids, and screen according to the fractal dimension of each grid area to obtain a first text candidate area; wherein screening according to the fractal dimension of each grid area includes: gradually reducing the sub-area size of each grid, counting the minimum number of boxes covering the stroke pixels, and calculating the fractal dimension of each grid; calculating the mean and standard deviation of the fractal dimensions of all grids, and determining the fractal dimension threshold; screening the fractal dimension of each grid according to the fractal dimension threshold; The pixels of the first text candidate area are used as nodes, and the edge weights between adjacent pixels are used as gradient amplitudes to construct a graph structure, and the topological invariant is calculated to eliminate abnormal areas according to the topological invariant to obtain the second text candidate area; Performing brightness mapping according to the fractal dimension of the second text candidate region to generate a topographic map, and performing dynamic segmentation path planning on the topographic map to obtain a third text candidate region; Performing adhesion detection on the third text candidate area and elastically separating the adhesion to obtain the target text area; Perform character recognition on the target text area based on language technology and convert the character recognition results into a standardized text format.
2. The method for character recognition based on language technology according to claim 1, characterized in that: Obtain a bill image, divide the bill image into grids, and screen according to the fractal dimension of each grid area to obtain the first text candidate area, including: The bank note image is converted into a grayscale image using the weighted average method, and the grayscale image is divided into grids.
3. The method for character recognition based on language technology according to claim 2, characterized in that: The pixels of the first text candidate area are used as nodes, and the edge weights between adjacent pixels are used as gradient amplitudes to construct a graph structure, and the topological invariant is calculated to eliminate abnormal areas according to the topological invariant to obtain the second text candidate area, including: Each pixel in the first text candidate area is regarded as a graph node, and only the four neighboring pixels above, below, left and right are connected. The gradient amplitude is used as the edge weight between adjacent pixels to obtain a graph structure, and an adjacency matrix is used to store the weight relationship between nodes and edges; The number of connected regions in the statistical graph structure is counted by breadth-first search to obtain the number of connected branches; Detect closed loops in the graph structure through the matrix tree theorem and obtain the number of ring structures; The region where the number of connected branches is 1 and the number of ring structures is 0 is retained to obtain the second text candidate region.
4. The method for character recognition based on language technology according to claim 3, characterized in that: Mapping the fractal dimension of each second text candidate region to a brightness value and generating a grayscale matrix in which each pixel value corresponds to a brightness value; Constructing a pheromone concentration matrix according to the grayscale matrix, setting a brightness threshold, assigning a brightness level to each region, and initializing the pheromone concentration matrix according to the brightness level to obtain an initialized pheromone matrix; The initialized pheromone matrix is updated according to the brightness gradient, and the pheromone is moved along the pheromone gradient to generate a preliminary segmentation path; Calculate the path energy, and update the preliminary segmentation path according to the path energy to obtain the target segmentation path; According to the target segmentation path, it is converted into a binary mask to obtain the third text candidate area.
5. The method for character recognition based on language technology according to claim 4, characterized in that: According to the brightness gradient, the initialized pheromone matrix is updated, and the pheromone is moved along the pheromone gradient to generate a preliminary segmentation path, including: Calculating the dynamic volatility factor and the dynamic diffusion factor, and obtaining an updated pheromone matrix according to the dynamic volatility factor and the dynamic diffusion factor; Calculate the brightness gradient direction of each pixel and move along the gradient direction. If the movement enters a low-brightness area, backtrack and re-plan the path to the adjacent high-brightness area. When moving along the gradient direction, the step size is positively correlated with the pheromone concentration.
6. The method for character recognition based on language technology according to claim 5, characterized in that: The third text candidate area is subjected to adhesion detection, and the adhesion is elastically separated to obtain the target text area, including: The 8-domain connectivity analysis is used for the third text candidate area to mark all connected areas. If the similarity of the boundary pixels of two adjacent connected areas is less than the similarity threshold, they are judged as adhesion, and the adhesion area marking matrix is obtained, and the adhesion pixel position is marked as 1; Calculate the repulsive force of the adhering pixel pairs, and calculate the attractive force for the non-adhering pixel pairs in the same grid, where the repulsive force is inversely proportional to the distance between pixels, and the attractive force is inversely proportional to the distance between pixels; For each sticky pixel, the vector sum of the repulsive force and the attractive force is calculated, and the pixel position is adjusted according to the vector sum. If the pixel moves beyond the image boundary, it is pulled back into the boundary. If the energy change of the adhesion area is less than the threshold after 5 consecutive iterations, or the maximum number of iterations is reached, the iteration is stopped and the separated mask is obtained; The separated masks are opened and their connectivity is verified. If the regional connectivity is insufficient, the adjacent high-brightness regions are re-merged to obtain the target text region mask.
7. The method for character recognition based on language technology according to claim 6, characterized in that: Perform character recognition on the target text area based on language technology and convert the character recognition results into a standardized text format, including: Apply morphological closing operation to the target text area mask, extract the brightness gradient features of the text area, and divide the text area into independent characters; The segmented characters are input into the OCR engine, and the character images are binarized and the character sizes are adjusted to the standard size to obtain the initially recognized character sequence; Load predefined format templates according to the bill type, replace similar characters in OCR recognition with standard forms, add thousand separators according to the format template, and unify the upper and lower case letters; Verify whether the text conforms to the format rules through regular expression matching, store the text by field classification, and generate a standardized text format.
8. A text recognition system based on language technology, characterized in that: include: The acquisition module is used to acquire a bank note image, divide the bank note image into grids, and screen according to the fractal dimension of each grid area to obtain a first text candidate area; wherein the screening according to the fractal dimension of each grid area includes: gradually reducing the sub-area size of each grid, counting the minimum number of boxes covering the stroke pixels, and calculating the fractal dimension of each grid; calculating the mean and standard deviation of the fractal dimensions of all grids, and determining the fractal dimension threshold; and screening the fractal dimension of each grid according to the fractal dimension threshold; A generation module is used to construct a graph structure by taking pixels of the first text candidate area as nodes and edge weights between adjacent pixels as gradient amplitudes, and to calculate topological invariants, so as to eliminate abnormal areas according to the topological invariants and obtain a second text candidate area; A calculation module, used for performing brightness mapping according to the fractal dimension of the second text candidate area to generate a topographic map, and performing dynamic segmentation path planning on the topographic map to obtain a third text candidate area; An encryption module is used to perform adhesion detection on the third text candidate area and elastically separate the adhesion to obtain a target text area; The transmission module is used to perform character recognition on the target text area according to language technology and convert the character recognition results into a standardized text format.
9. A computing device, characterized in that include: one or more processors; A storage system for storing one or more programs, when the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a program, which, when executed by a processor, implements the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Method for detecting surface defects of mechanical parts based on image texture and fractal dimension
CN101710081A
Robot visual image segmentation method based on statistics and local fractal dimension
CN102509274A
Hand segmentation method based on multi-clue fusion
CN114862894A
Road crack change detection method and system based on graph structure
CN117788455A