Seal identification method and device for medical bills

Through the multi-level feature fusion network and global multi-dimensional branch module, combined with seal position detection, boundary recognition and OCR recognition, the problem of difficult to identify curved texts and complex seals by traditional verification methods is solved, and efficient identification and verification of medical bill seals is achieved.

CN120126149APending Publication Date: 2025-06-10ZHEJIANG UNIV OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510071652.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

Traditional medical bill seal verification methods are difficult to effectively identify curved texts and complex seal styles, resulting in insufficient audit efficiency and accuracy.

Method used

Multi-level feature fusion network (DGMN) is used to combine global multi-dimensional branch modules to perform seal position detection, boundary recognition, text expansion and polar coordinate expansion, and finally obtain text information on the seal through OCR recognition.

Benefits of technology

It realizes accurate identification of the seal content on medical bills, improves the efficiency and accuracy of review, and can effectively verify the effectiveness of medical bills.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126149A_ABST
    Figure CN120126149A_ABST
Patent Text Reader

Abstract

The invention discloses a medical bill seal identification method and device, and the method comprises the steps: carrying out the seal position detection of an original medical image, cutting out a seal image, and carrying out the curved text identification of the cut seal image, mainly employing a multi-level feature fusion network model, the multi-level feature fusion network model fuses a global multi-dimensional branch module, so that a better seal text boundary result can be obtained; and finally, carrying out post-processing, including edge expansion and polar coordinate transformation, on the seal image based on the text boundary result, obtaining a final character recognition result through an OCR recognition engine, comparing the final character recognition result with a registered name through an editing distance, and verifying the validity of the bill. The method is of great significance in seal identification on the medical bill, and seal content on the medical bill can be identified more accurately and effectively.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of image processing, and particularly relates to a method and device for identifying seals on medical bills. Background Art

[0002] In today's rapidly developing medical industry, medical bills, as important vouchers for patients' medical treatment and financial settlement of medical institutions, their validity and authenticity have received extensive attention. Medical bills not only record patients' medical information, treatment processes, and expense details, but also often require various seals of medical institutions for verification. These seals are not only identifiers of institutional legitimacy but also the key to preventing bill forgery and ensuring patients' rights and interests. However, with the continuous progress of bill forgery technology, traditional bill verification methods are gradually becoming inadequate, and innovative technical means are urgently needed to improve the efficiency and accuracy of auditing.

[0003] Seals are usually formed by stamping, and in the step of making seals, the shapes and styles of seals are various, and they are very likely to present a curved form. The existence of this curved text poses challenges to traditional text extraction and verification. In addition, there may be other patterns or texts printed on medical bills, and these contents may overlap or be close to the seals, making the recognition of curved text even more difficult. In order to effectively identify this information and adapt to complex text forms and diverse seal styles, there is an urgent need for a method to accurately and effectively identify the seal images on medical bills. Summary of the Invention

[0004] The purpose of this application is to propose a new method for identifying seals on medical bills to more accurately and effectively identify the seal contents on medical bills in view of the above problems.

[0005] To achieve the above purpose, the technical solutions adopted by this application are as follows: A method for identifying seals on medical bills, comprising: Detect the seal position on the original medical bill image to locate the seal area; According to the located seal area, crop the original medical bill image, and then perform boundary recognition on the cropped seal image to obtain the text boundary; Perform dilation and polar coordinate expansion on the text boundary area on the seal image, and then perform text recognition to obtain the text recognition result; Among them, performing boundary recognition on the cropped seal image to obtain the text boundary includes: Extract features from the seal image through a multi-level feature fusion network, and a global multi-dimensional branch module is added after each dense block of the DenseNet in the multi-level feature fusion network, and the features output by each global multi-dimensional branch module are fused to obtain a shared feature map; Generate a classification map and a distance field map using a shared feature map; Binarize the distance field map to generate candidate boundaries, and remove candidate boundaries with low confidence according to the classification map to obtain rough text boundaries; Refine the rough text boundaries to obtain the final text boundaries.

[0006] Furthermore, the global multi-dimensional branch module performs the following operations: Obtain weighted weights by performing convolution operations and normalization operations on the input feature map, then multiply the weighted weights by the input feature map, and then add the result to the input feature map after convolution operations to obtain the first feature map of global context aggregation; Pass the first feature map through three branches respectively to capture cross-dimensional features, and the output feature maps of the three branches are averaged to obtain the output feature map of the global multi-dimensional branch module.

[0007] Furthermore, the three branches are respectively: The first branch first performs channel pooling, then passes through convolution and activation functions, and then multiplies with the elements of the input feature map; For the second and third branches, first perform transposition, then perform Z pooling, convolution and activation functions, then multiply with the transposed input feature map element by element, and finally perform transposition operations.

[0008] Furthermore, dilating and polar coordinate expanding the text boundary area on the seal image, and then performing text recognition to obtain the text recognition result, including: Extract the text boundary area from the seal image through masking and perform edge dilation to obtain a dilated image; Divide the dilated image into four quadrants, and then perform polar coordinate transformation to obtain a flattened seal image; Perform OCR recognition on the flattened seal image to obtain the text information on the seal.

[0009] Furthermore, the refinement of the rough text boundaries to obtain the final text boundaries is implemented using a boundary conversion module, and the boundary conversion module includes an encoder and a decoder.

[0010] This application also proposes a seal recognition device for medical bills, including a processor and a memory storing a number of computer instructions, and when the computer instructions are executed by the processor, the steps of the above-mentioned seal recognition method for medical bills are implemented.

[0011] This application proposes a method and device for seal recognition of medical bills, which fully considers the problems and difficulties existing in existing seal recognition. First, perform seal target detection on the original medical image; then perform curved text recognition on the cropped seal image. The main method used is the multi-level feature fusion network model DGMN, which is a better model adopted after experimental comparison with ResNet and DenseNet. It integrates the global multi-dimensional branch module (GM block) and can obtain better seal text boundary results. Finally, post-process the seal image based on the above text boundary results, including edge dilation, polar coordinate transformation, and obtain the final text recognition result through the OCR recognition engine, and compare it with the registered name through the edit distance to verify the validity of the bill. This application has great significance for seal recognition on medical bills, and the validity of medical bills can be fully verified through this seal recognition method. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 It is a flowchart of the seal recognition method for medical bills of this application.

[0013] Figure 2 It is a structural diagram of the multi-level feature fusion network of this application.

[0014] Figure 3 It is a schematic diagram of the structure of the global multi-dimensional branch module of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0015] In order to make the purpose, technical solutions and advantages of this application clearer, the following further describes this application in detail with reference to the drawings and embodiments. It should be understood that the specific embodiments described here are only used to explain this application and are not used to limit this application.

[0016] Embodiment 1. As Figure 1 shown, a method for seal recognition of medical bills is proposed, including: Step S1: Detect the seal position on the original medical bill image and locate the seal area.

[0017] Regarding target detection, it is already a relatively mature technology in this field. For example, the YOLOv8 network model is a commonly used target detection network model. The YOLOv8 network model trained with a medical bill dataset containing seals can be directly used to predict the original medical bill image to locate the seal area on the image.

[0018] The medical bill dataset is a private dataset consisting of 500 medical bills with seals, and the positions of the seals in each image are marked. The YOLOv8 network model adopted is a relatively new version of the yolo algorithm in the field of object detection. It uses an improved neural network architecture, combines deeper feature extraction and more efficient model design, enabling richer features to be extracted during object detection. At the same time, while maintaining high detection accuracy, it can still achieve real-time inference, making it suitable for applications in scenarios that require quick response.

[0019] Step S2: According to the located seal area, crop the original medical bill image, and then perform boundary recognition on the cropped seal image to obtain the text boundary.

[0020] Refer to the detected seal area above, crop the original medical bill image, and retain the content of the seal area as the seal image to prepare for further recognition later.

[0021] After obtaining the seal image, there are currently many methods to obtain the boundary of the acquired image. This application proposes an improved boundary recognition method, which will be described below through specific embodiments.

[0022] In a specific embodiment, the performing boundary recognition on the cropped seal image to obtain the text boundary includes: Step 2.1: Extract features from the seal image through a multi-level feature fusion network. The multi-level feature fusion network adds a global multi-dimensional branch module after each dense block of the DenseNet, and the features output by each global multi-dimensional branch module are fused to obtain a shared feature map.

[0023] This embodiment proposes a multi-level feature fusion network, abbreviated as DGMN, for realizing boundary recognition based on edge detection. As Figure 2 shown, the multi-level feature fusion network DGMN is an improvement based on the densely connected convolutional network (DenseNet). Specifically, a global multi-dimensional branch module (GM Block) is added after each dense block.

[0024] In the multi-level feature fusion network DGMN, the input image first passes through a 7×7 convolutional layer conv to extract input features, then through a batch normalization layer BatchNorm for feature standardization, followed by the introduction of a non-linear transformation through the ReLu function and a 3×3 max pooling layer for feature map dimensionality reduction. After that, it contains four dense blocks DenseBlock, which can pass the feature map to subsequent layers through dense connections, four global multi-dimensional branch modules GM Block (following each dense block), and there is a transition layer Transiton after each of the first three global multi-dimensional branch modules to control the number of channels and spatial dimensions of the feature map. Finally, there is another BatchNorm layer for feature standardization.

[0025] In this embodiment, the global multi-dimensional branch module, as Figure 3 shown, introduces global context modeling for the feature map, forms global context features by weighted averaging the features at all positions, and then captures the dependencies between channels, aggregates the global context features into the features at each position, and improves the network performance with almost no increase in computational and parameter amounts. The following operations are performed:

[0026] Step 211: Obtain the weighted weights by performing convolutional and normalization operations on the input feature map, then multiply the weighted weights by the input feature map, and after performing convolutional operations, add them to the input feature map to obtain the first feature map of global context aggregation.

[0027] This step is expressed by the following formula:

[0028] where x and y represent the input feature map and the first feature map of this part respectively, represents the number of positions in the feature map, represents the weighted weights obtained through 1×1 convolution (conv) and normalization softmax, represents matrix multiplication, represents using the weight to combine the features at all positions by weighted averaging, represents capturing dependencies through feature transformation, that is, the 1×1 conv part, represents element-wise addition, aggregates the global context features into the features at each position, and finally obtains the first feature map.

[0029] Step 212: Pass the first feature map through three branches respectively to capture cross-dimensional features, and the output feature maps of the three branches are averaged to obtain the output feature map of the global multi-dimensional branch module.

[0030] For the obtained first feature map, which includes the channel dimension C and the spatial dimension H or W, it is divided into three branches. Each branch is used to capture the cross-dimensional interaction features between the channel dimension and the spatial dimension. Finally, a Sigmoid activation layer is used to generate attention weights, which are applied to the input tensor, and its dimension is transformed back to the original shape. Through this method, cross-dimensional features are effectively captured, providing rich feature representations. The operations performed by the global multi-dimensional branch module in this embodiment are expressed by the following formula:

[0031]

[0032] Among them, y represents the first feature map, z represents the output of the global multi-dimensional branch module, y1 and y2 respectively represent the results after the first feature map is rotated 90° counterclockwise along the H axis and the W axis, ′ represents dimensionality reduction by the pooling layer, represents convolution of k×k, represents the Sigmoid operation, represents element-wise multiplication, and respectively represent the results after rotating 90° clockwise along the H axis and the W axis, that is, restoring the input dimension.

[0033] As Figure 3 shown in the three branches, the first branch first performs channel pooling, then goes through convolution (Conv+BatchNorm) and activation function (Sigmoid), and then multiplies element-wise with the input feature map; for the second and third branches, they first perform permutation, then Z-pooling, convolution (Conv+BatchNorm) and activation function (Sigmoid), then multiply element-wise with the transpose of the input feature map, and finally perform a transpose operation. Finally, an average (Avg) operation is performed on the three branches to obtain the final output.

[0034] It should be noted that ResNet and DenseNet are both commonly used models in the field. To select a better network model, two models, ResNet50 and DenseNet121, are compared: ResNet introduces the concept of residual learning and uses "skip connections" to directly add the input to the output, while DenseNet adopts the "dense connection" method, where the input of each layer comes not only from all the previous layers but also directly connects to the outputs of all the previous layers; after comprehensive consideration and experimental verification, DenseNet121 has better performance than ResNet50, and the improved model DGMN used in this application is further improved based on DenseNet121 and obtains better performance compared to the previous two.

[0035] The specific experimental results are shown in Table 1. Recall represents the recall rate, Precision represents the precision rate, and F-measure represents the harmonic mean of the precision rate and the recall rate.

[0036] Table 1

[0037] In this embodiment, there are four global multi-dimensional branch modules GM Block in the multi-level feature fusion network DGMN. The feature maps output by each global multi-dimensional branch module are fused by adding elements after upsampling in the order from bottom to top to obtain a shared feature map.

[0038] Step 2.2: Generate a classification map and a distance field map using the shared feature map.

[0039] The classification map contains the classification confidence of each pixel (text / non-text) in the shared feature map, and is a single-channel classification map obtained by passing the shared feature map through a 1×1 conv.

[0040] The distance field map is a normalized distance map, which represents the normalized distance from a text pixel to the nearest pixel on the text boundary. The formula is as follows:

[0041] Where T represents the text region obtained by passing the classification map through a classification threshold, represents each pixel in the text region T, represents to the nearest pixel on the text boundary. If is in the non-text region, the distance is defined as 0. For example, for the classification map, a classification threshold of 0.8 is set, and then those greater than the threshold are used as the text region, and those less than it are used as the non-text region to obtain the text region T.

[0042] And the specific definition of L is as follows:

[0043] It represents the maximum distance from the text pixel in the text region T to the nearest boundary pixel .

[0044] It should be noted that the generation of the classification map and the distance field map belongs to a relatively mature technology in this field and will not be elaborated here.

[0045] Step 2.3: Binarize the distance field map to generate candidate boundaries, and remove the candidate boundaries with low confidence according to the classification map to obtain the rough text boundaries.

[0046] In this step, candidate boundaries are generated by binarizing the distance field map, and then the confidence scores of each candidate boundary are calculated according to the classification map. The candidate boundaries with lower confidence are removed, and the one with the highest confidence is the obtained rough text boundary.

[0047] Step 2.4: Refine the rough text boundary to obtain the final text boundary.

[0048] In this embodiment, a boundary transformation module is used to effectively perform feature learning and predict the accurate offset of each vertex pointing to the text boundary, so as to obtain the final text boundary.

[0049] Uniformly sample the rough text boundary to obtain N control points, such as 20. Concatenate the shared feature map and the distance field map, and then calculate and extract the features of the control points as feature vectors through bilinear interpolation, and input them into the boundary transformation module.

[0050] The boundary transformation module adopts an encoder-decoder structure. The encoder consists of three layers, and each layer contains a multi-head self-attention block and a multi-layer perceptron network MLP, which can encode the feature map (B×N×36) of the rough text boundary into an embedded feature map (B×N×128). The decoder is a simple multi-layer perceptron network, consisting of three layers of perceptron networks and 1×1 convolution activated by ReLU. The decoder can learn to predict the offset between the control point and the target point, and the target point represents the nearest point on the text boundary annotated on the training image.

[0051] It should be noted that using the boundary transformation module to refine the rough text boundary to obtain the final text boundary belongs to a relatively mature technology in this field and will not be elaborated here.

[0052] Step S3: Dilate and perform polar coordinate expansion on the text boundary area on the seal image, and then perform text recognition to obtain the text recognition result.

[0053] Specifically, this step performs the following operations: Step 3.1: Extract the text boundary area from the seal image through a mask and perform edge dilation to obtain a dilated image.

[0054] First, create a completely black mask with the same size as the seal image, fill the text boundary area, then perform edge dilation on the mask, and then apply the mask on the original seal image, only keep the new polygon area, and set other areas to black to obtain the dilated image; Step 3.2: Divide the dilated image into four quadrants, and then perform polar coordinate transformation to obtain a flattened seal image.

[0055] For the inflated image, with the center point as the origin, the image is divided into four quadrants. Calculate the density of the boundary control points, and select the quadrant with the fewest control points as the starting quadrant. Each quadrant corresponds to a starting splitting angle (the first quadrant: -45°, the second quadrant: -135°, the third quadrant: -225°, the fourth quadrant: -315°), and perform polar coordinate expansion in the clockwise direction. Convert the picture from polar coordinates to rectangular coordinates to obtain the flattened seal image.

[0056] Step 3.3: For the flattened seal image, use OCR for recognition to obtain the text information on the seal.

[0057] Finally, use the OCR optical recognition engine to recognize the flattened seal image obtained to obtain the seal text content.

[0058] Subsequently, according to known information such as the source of the medical bill, etc., compare with the text content to verify the validity of the medical bill. When verifying, compare the seal text recognition result with the unit name at the time of uploading the medical bill, calculate the similarity, and verify the validity of the bill, including:

[0059] Calculate the similarity between the seal text recognition result and the uploaded unit name by using the method of edit distance, and judge whether the bill is valid through the set similarity threshold.

[0060] First, for the seal text recognition result and the uploaded unit name, use dynamic programming to calculate the seal text recognition result and the uploaded unit name of the edit distance, and perform initialization. The formula is as follows: dp[i][j]= &0 , i = j = 0 &i , j = 0 &j , i = 0 Among them, dp[i][j] represents the minimum distance required to convert the first i characters of to the first j characters of .

[0061] Secondly, after initializing the edit distance, perform state transition. The formula is as follows:

[0062] Then, for the result of the state transition, the final edit distance obtained is , and the similarity sim calculation formula between the two is as follows: Among them, m is the length of , and n is the length of .

[0063] Finally, for the similarity sim, if sim ≥ 0.8, the bill is considered valid; if sim < 0.8, the bill is considered invalid. The set similarity threshold can be changed according to actual requirements.

[0064] This application fully considers the problems and difficulties existing in the existing seal recognition. First, it performs seal target detection on the original medical image; then it performs curved text recognition on the cropped seal image, mainly using the multi-level feature fusion network model DGMN, which is a better model adopted after experimental comparison with ResNet and DenseNet. It integrates the global-multi-dimensional branch module (GM block) and can obtain better seal text boundary results; finally, based on the above text boundary results, post-processing is performed on the seal image, including edge dilation, polar coordinate transformation, and the final text recognition result is obtained through the OCR recognition engine and compared with the registered name through the edit distance to verify the validity of the bill; therefore, this application has great significance for the seal recognition on medical bills, and the validity of medical bills can be fully verified through this seal recognition method.

[0065] In another embodiment, this application also provides a seal recognition device for medical bills, including a processor and a memory storing a number of computer instructions. When the computer instructions are executed by the processor, the steps of the above seal recognition method for medical bills are implemented.

[0066] For the specific limitations of the seal recognition device for medical bills, reference can be made to the limitations of the seal recognition method for medical bills in the above text, which will not be elaborated here. The above seal recognition device for medical bills can be implemented in whole or in part through software, hardware, and their combinations. It can be embedded in the processor of a computer device in hardware form or be independent of it, or stored in the memory of a computer device in software form for the processor to call and execute the corresponding operations above.

[0067] The memory and the processor are directly or indirectly electrically connected to achieve data transmission or interaction. For example, these components can be electrically connected to each other through one or more communication buses or signal lines. The memory stores a computer program that can run on the processor. The processor realizes the seal recognition method for medical bills in the embodiments of the present invention by running the computer program stored in the memory.

[0068] Among them, the memory can be, but is not limited to, random access memory (RAM), read only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc. Among them, the memory is used to store a program, and after receiving an execution instruction, the processor executes the program.

[0069] The processor may be an integrated circuit chip with data processing capabilities. The above-mentioned processor may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc. It can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0070] The above-mentioned embodiments only represent several implementation manners of the present application, and the description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several deformations and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.

Claims

1. A medical bill seal recognition method, characterized in that: The seal recognition method of the medical bill comprises: Perform seal position detection on the original medical bill image to locate the seal area; The original medical bill image is cropped according to the located seal area, and then the boundary of the cropped seal image is recognized to obtain the text boundary; Dilate and expand the text boundary area on the seal image using polar coordinates, then perform text recognition to obtain the text recognition result; Among them, the boundary recognition is performed on the cropped seal image to obtain the text boundary, including: The seal image is subjected to feature extraction through a multi-level feature fusion network, wherein a global multi-dimensional branch module is added after each dense block of the dense connection network DenseNet, and the features output by each global multi-dimensional branch module are fused to obtain a shared feature map; Generate classification map and distance field map using shared feature map; Binarize the distance field map to generate candidate boundaries, remove candidate boundaries with low confidence according to the classification map, and obtain coarse text boundaries; The coarse text boundary is refined to obtain the final text boundary.

2. The medical bill seal recognition method according to claim 1, characterized in that: The global multi-dimensional branch module performs the following operations: The input feature map is convolved and normalized to obtain a weighted weight, and then the weighted weight is multiplied by the input feature map, and then added to the input feature map after the convolution operation to obtain the first feature map of global context aggregation; The first feature map is passed through three branches respectively to capture cross-dimensional features, and the output feature maps of the three branches are averaged to obtain the output feature map of the global multi-dimensional branch module.

3. The medical bill seal recognition method according to claim 2, characterized in that: The three branches are: The first branch first performs channel pooling, then passes through convolution and activation functions, and then multiplies the input feature map elements; The second and third branches first perform transposition, then perform Z pooling, convolution and activation function, then perform element-wise multiplication with the transposition of the input feature map, and finally perform transposition operation.

4. The medical bill seal recognition method according to claim 1, characterized in that: The text boundary area is expanded and polar coordinates are expanded on the seal image, and then text recognition is performed to obtain a text recognition result, including: Extract the text boundary area of ​​the seal image by masking, and perform edge dilation to obtain a dilated image; The expanded image is divided into four quadrants, and then polar coordinate transformation is performed to obtain the seal flattened image; Use OCR to recognize the flattened image of the seal and obtain the text information on the seal.

5. The medical bill seal recognition method according to claim 1, characterized in that: The rough text boundary is refined to obtain the final text boundary, which is achieved by using a boundary conversion module, and the boundary conversion module includes an encoder and a decoder.

6. A medical bill seal recognition device, comprising a processor and a memory storing a plurality of computer instructions, characterized in that: When the computer instructions are executed by a processor, the steps of the method described in any one of claims 1 to 5 are implemented.

Citation Information

Cited By

  • A seal character recognition method

    CN122531031A