OCR (Optical Character Recognition) system for medical instrument registration certificate
Through image acquisition, intelligent preprocessing and multimodal recognition technology, combined with layout analysis and database verification, the image acquisition accuracy problem of the OCR recognition system of the medical device registration certificate is solved, and efficient and accurate extraction and verification of the registration certificate information is achieved.
Patent Information
- Application Number
- CN202510309318.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-07-11
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, the OCR recognition system with a medical device registration certificate is affected by the image acquisition accuracy of the viewing angle and light angle, resulting in low recognition accuracy and even unreadable. The traditional manual review method is time-consuming and labor-intensive, and the efficiency and accuracy are difficult to guarantee.
The image acquisition module, intelligent preprocessing module, multimodal recognition engine, dynamic verification module and online learning mechanism are adopted, combined with set correction, image enhancement, layout analysis, text recognition and table analysis algorithms, and image correction and content analysis are realized through mixed edge detection, non-local mean filtering, ResNeXt-101 feature extraction, OpenCV line segment detection and database verification.
It improves the accuracy and recognition accuracy of image analysis, realizes a closed-loop recognition, analysis and verification process, and enhances the recognition efficiency and accuracy of medical device registration certificates.
Smart Images

Figure CN120299046A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of OCR recognition, and particularly to an OCR recognition system for medical device registration certificates. Background Art
[0002] Medical device registration refers to the process of systematically evaluating the safety and effectiveness of medical devices intended for marketing and use in accordance with legal procedures to determine whether to approve their marketing and use.
[0003] OCR (Optical Character Recognition) is a technology that can convert text in images, scanned documents, or photos into an editable and searchable digital text format. OCR technology is widely used in various scenarios, such as document digitization, data entry, license plate recognition, ID card recognition, etc.
[0004] In the registration process of medical device products, the registration certificate is a key document proving the legal marketability of the product, which contains key information such as product name, model specifications, manufacturer, expiration date, registration number, etc. The traditional manual review method is not only time-consuming and laborious but also prone to errors. Especially when dealing with a large number of registration certificates, it is difficult to ensure efficiency and accuracy.
[0005] Therefore, the existing technology uses OCR to recognize the registration certificates of medical devices, which can effectively solve the problems of efficiency and accuracy.
[0006] However, in the existing technology, the accuracy of image acquisition of the registration certificate is affected by factors such as the shooting angle and light angle. When performing content analysis and reading subsequently, the accuracy is relatively low, and there may even be a situation where the content cannot be read.
[0007] Therefore, the present invention proposes an OCR recognition system for medical device registration certificates. Summary of the Invention
[0008] The purpose of the present invention is to solve the disadvantages existing in the prior art, and to propose an OCR recognition system for medical device registration certificates.
[0009] To achieve the above purpose, the present invention adopts the following technical solutions:
[0010] An OCR recognition system for medical device registration certificates, comprising:
[0011] An image acquisition module that scans the medical device registration certificate to obtain its image data;
[0012] An intelligent preprocessing module that is built-in with a set correction unit and an image enhancement unit, which respectively perform set correction and image enhancement on the acquired image;
[0013] A multi-modal recognition engine that recognizes the pre-processed images. The multi-modal recognition engine is built-in with a layout analysis unit, a text recognition model, and a table parsing algorithm;
[0014] A dynamic verification module that dynamically verifies the recognized data. The dynamic verification module is built-in with a logical verification layer and a database verification layer;
[0015] An online learning mechanism that collects the error types fed back by users through the error sample annotation section in the data stream and performs incremental training;
[0016] A template library update module that updates the template library.
[0017] Preferably, the processing logic of the set correction unit includes the following steps:
[0018] A1: Edge detection. Using the Hybrid Edge Detection network HED, fusing the VGG16 feature extraction and the large-scale edge detection algorithm to output the document body edge coordinate set: E = {(x1, y1),......(x1, y1)};
[0019] A2: Projection matrix calculation. Fitting the optimal homography matrix H through the RANSAC algorithm 3×3 , H 3×3 is a 3x3 projection transformation matrix, where p i is the coordinate of the i-th feature point of the source image, in the form of (x i , y i , 1) T ; p' i is the coordinate of the corresponding point of the target image, in the form of (x' i , y' i , 1) T ; |||| is the Euclidean distance;
[0020] A3: Perspective transformation execution. Applying bilinear interpolation to achieve pixel remapping. The model of bilinear interpolation is: I dst = ∑ i,j I src (x'+i, y'+j)·(1-α)(1-β), I dst is the pixel value of the target image at the (x, y) position; I src is the pixel value of the source image at the (x'+i, y'+j) position; α, β are the fractional parts of the sub-pixel offsets, i, j are positive offsets, taking values of 0 or 1.
[0021] Preferably, the processing logic of the image enhancement unit includes the following steps:
[0022] B1: Noise suppression, using an improved Non-local Means algorithm, whose model is h is the smoothing parameter, σ is the variance, is the intensity value of pixel x after denoising; C(x) is the normalization coefficient, N(x) is the local neighborhood pixel block centered on x, and Ω is the search window range;
[0023] B2: Contrast enhancement. First, the image is segmented into 8×8 sub-blocks, then the histogram of each sub-block is clipped, and finally the results are combined using bilinear interpolation;
[0024] B3: Stamp separation, based on threshold segmentation in the HSV color space, whose model is: Mask = (H ∈ [0, 15] ∪ [160, 180]) ∩ (S > 0.4) ∩ (V > 0.3), where H, S, and V are hue, saturation, and value respectively; then the Telea image inpainting algorithm is used to fill the stamp area.
[0025] Preferably: in the B2 step, clip limit = 2.0.
[0026] Preferably: the analysis logic of the layout analysis unit includes the following steps:
[0027] C1: Feature extraction, using the feature extraction layer to extract features. The feature extraction layer uses ResNeXt-101 instead of the original ResNet50 and adds a deformable convolutional layer;
[0028] C2: Set the anchor box ratio;
[0029] C3: Establish a multi-task loss function. The loss function is L = γ1L1 + γ2L2 + γ3L3 + γ4L4; L1, L2, L3, and L4 are the classification loss, bounding box regression loss, mask prediction loss, and newly added affinity loss respectively, and γ1, γ2, γ3, and γ4 are the weight coefficients;
[0030] C4: Result output, output the analysis results.
[0031] Preferably: in the C2 step, the anchor box ratio is set to [0.5, 1, 2, 4].
[0032] Preferably: in the C3 step, where is the topological feature vector of the i-th predicted region, is the topological feature vector of the i-th corresponding real region, and N is the number of valid regions,
[0033] Preferably: the logic of the table parsing algorithm is:
[0034] D1: Cell detection, extract the table area based on the layout analysis result, and use OpenCV line detection;
[0035] D2: Topological relationship construction, establish a two-dimensional relationship matrix M m×n , calculate the cell merging relationship where M is the two-dimensional table structure matrix, and Merge i,j is the flag indicating whether the cell in the i-th row and j-th column is merged with the cell below;
[0036] D3: Semantic alignment, based on the semantic matching of the table title, Hr is the set of recognized table title words, DB is the standard field library of the drug administration, and Ct is the count of intersection elements of the set.
[0037] Preferably: The database verification layer performs field missing matching, and its model is: where V oer (f i ) is the value of the i-th field recognized by ocr; V db (f i ) is the corresponding field value in the database; [] is the Iverson bracket, which is 1 when the condition is true and 0 otherwise; ω i is the field weight.
[0038] Preferably: In the database verification layer, ω i corresponding
[0039] The beneficial effects of the present invention are:
[0040] 1. The present invention performs geometric correction and image enhancement on the image, which can effectively improve the accuracy of subsequent analysis and processing. At the same time, it can perform layout analysis on the image, and then perform verification and missing matching, realizing a "recognition - parsing - verification" closed loop, and improving the recognition accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 is the architecture function diagram of a medical device registration certificate OCR recognition system proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0042] The technical solutions of the present invention will be further described in detail below in conjunction with the specific embodiments.
[0043] Embodiment 1:
[0044] A medical device registration certificate OCR recognition system, which includes:
[0045] An image acquisition module, which scans the medical device registration certificate to obtain its image data;
[0046] An intelligent preprocessing module, which is built-in with a set calibration unit and an image enhancement unit, and performs set calibration and image enhancement on the acquired images respectively;
[0047] A multi-modal recognition engine, which recognizes the preprocessed images. The multi-modal recognition engine is built-in with a layout analysis unit, a text recognition model, and a table parsing algorithm;
[0048] A dynamic verification module, which dynamically verifies the recognized data. The dynamic verification module is built-in with a logical verification layer and a database verification layer;
[0049] An online learning mechanism, which collects the error types feedback by users through the error sample annotation section in the data stream and performs incremental training;
[0050] A template library update module, which updates the template library.
[0051] The processing logic of the set calibration unit includes the following steps:
[0052] A1: Edge detection, using the Hybrid Edge Detection network HED, fusing VGG16 feature extraction and large-scale edge detection algorithm to output the document body edge coordinate set: E = {(x1, y1),......(x1, y1)};
[0053] A2: Projection matrix calculation, fitting the optimal homography matrix H through the RANSAC algorithm 3×3 , H 3×3 is a 3-row and 3-column projection transformation matrix, where p i is the coordinate of the i-th feature point of the source image, in the form of (x i , y i , 1) T ; p' i is the coordinate of the corresponding point of the target image, in the form of (x' i , y' i , 1) T ; |||| is the Euclidean distance;
[0054] A3: Perspective transformation execution, applying bilinear interpolation to achieve pixel remapping. The model of bilinear interpolation is: I dst = ∑ i,j I src (x'+i, y'+i)·(1-α)(1-β), I dst is the pixel value of the target image at the (x, y) position; I src is the pixel value of the source image at the (x'+i, y'+j) position; α, β are the decimal parts of the sub-pixel offsets, i and j are positive offset values, taking values of 0 or 1.
[0055] The processing logic of the image enhancement unit includes the following steps:
[0056] B1: Noise suppression, using an improved Non-local Means algorithm, whose model is h is the smoothing parameter, σ is the variance, is the intensity value of pixel point x after denoising; C(x) is the normalization coefficient, N(x) is the local neighborhood pixel block centered on x, and Ω is the search window range;
[0057] B2: Contrast enhancement. First, the image is segmented into 8×8 sub-blocks, then histogram clipping is performed on each sub-block, and finally the processing results are merged using bilinear interpolation;
[0058] B3: Stamp separation, based on threshold segmentation in the HSV color space, whose model is: Mask = (H ∈ [0, 15] ∪ [160, 180]) ∩ (S > 0.4) ∩ (V > 0.3), where H, S, and V are hue, saturation, and value respectively; then the Telea image inpainting algorithm is used to fill the stamp area.
[0059] In the B2 step, clip limit = 2.0.
[0060] The analysis logic of the layout analysis unit includes the following steps:
[0061] C1: Feature extraction, using the feature extraction layer for feature extraction. The feature extraction layer uses ResNeXt-101 instead of the original ResNet50, and a deformable convolutional layer is added;
[0062] C2: Set the anchor box ratios;
[0063] C3: Establish a multi-task loss function. The loss function is L = γ1L1 + γ2L2 + γ3L3 + γ4L4; L1, L2, L3, and L4 are the classification loss, bounding box regression loss, mask prediction loss, and newly added affinity loss respectively, and γ1, γ2, γ3, and γ4 are the weight coefficients;
[0064] C4: Result output, outputting the analysis results.
[0065] In the C2 step, the anchor box ratios are set to [0.5, 1, 2, 4].
[0066] In the C3 step, where is the topological feature vector of the i-th prediction region, is the topological feature vector corresponding to the i-th true region, and N is the number of valid regions.
[0067] In the step C3,
[0068] Embodiment 2:
[0069] A medical device registration certificate OCR recognition system, which includes:
[0070] An image acquisition module that scans the medical device registration certificate to obtain its image data;
[0071] An intelligent preprocessing module with a set correction unit and an image enhancement unit built in, which respectively perform set correction and image enhancement on the acquired images;
[0072] A multi-modal recognition engine that recognizes the preprocessed image. The multi-modal recognition engine has a layout analysis unit, a text recognition model, and a table parsing algorithm built in;
[0073] A dynamic verification module that dynamically verifies the recognized data. The dynamic verification module has a logical verification layer and a database verification layer built in;
[0074] An online learning mechanism that collects the error types feedback by users through the error sample annotation section in the data stream and performs incremental training;
[0075] A template library update module that updates the template library.
[0076] The processing logic of the set correction unit includes the following steps:
[0077] A1: Edge detection. Use the hybrid edge detection network HED to fuse the VGG16 feature extraction and the large-scale edge detection algorithm to output the document body edge coordinate set: E = {(x1, y1),......(x1, y1)};
[0078] A2: Projection matrix calculation. Fit the optimal homography matrix H through the RANSAC algorithm 3×3 , H 3×3 is a 3x3 projection transformation matrix, where p i is the coordinate of the i-th feature point of the source image, in the form of (x i , y i , 1) T ; p' i is the coordinate of the corresponding point in the target image, in the form of (x' i , y' i , 1) T ; |||| is the Euclidean distance;
[0079] A3: Perspective transformation is performed, and bilinear interpolation is applied to achieve pixel remapping. The model of bilinear interpolation is: I dst = ∑ i,j I src (x′ + i, y′ + j)·(1 - α)(1 - β), where I dst is the pixel value of the target image at the position (x, y); I src is the pixel value of the source image at the position (x' + i, y' + j); α and β are the fractional parts of the sub-pixel offsets, and i, j are positive offsets, taking values of 0 or 1.
[0080] The processing logic of the image enhancement unit includes the following steps:
[0081] B1: Noise suppression is performed using an improved Non-local Means algorithm, and its model is where h is the smoothing parameter, σ is the variance, is the intensity value of pixel point x after denoising; C(x) is the normalization coefficient, N(x) is the local neighborhood pixel block centered on x, and Ω is the search window range;
[0082] B2: Contrast enhancement. First, the image is segmented into 8×8 sub-blocks, then histogram clipping is performed on each sub-block, and finally the processing results are merged using bilinear interpolation;
[0083] B3: Stamp separation is based on threshold segmentation in the HSV color space, and its model is: Mask = (H ∈ [0, 15] ∪ [160, 180]) ∩ (S > 0.4) ∩ (V > 0.3), where H, S, and V are hue, saturation, and value respectively; then the Telea image inpainting algorithm is used to fill the stamp area.
[0084] In the B2 step, clip limit = 2.0.
[0085] The analysis logic of the layout analysis unit includes the following steps:
[0086] C1: Feature extraction is performed using a feature extraction layer. The feature extraction layer uses ResNeXt-101 instead of the original ResNet50 and adds a deformable convolutional layer;
[0087] C2: Set the anchor box ratio;
[0088] C3: Establish a multi-task loss function. The loss function is L = γ1L1 + γ2L2 + γ3L3 + γ4L4; L1, L2, L3, and L4 are the classification loss, bounding box regression loss, mask prediction loss, and newly added affinity loss respectively, and γ1, γ2, γ3, and γ4 are the weight coefficients;
[0089] C4: Result output, output the analysis result.
[0090] In the C2 step, set the anchor box ratios as [0.5, 1, 2, 4].
[0091] In the C3 step, where is the topological feature vector of the i-th predicted region, is the topological feature vector of the i-th corresponding ground-truth region, and N is the number of valid regions.
[0092] In the C3 step,
[0093] Example 3:
[0094] A medical device registration certificate OCR recognition system, which includes:
[0095] An image acquisition module, which scans the medical device registration certificate to obtain its image data;
[0096] An intelligent preprocessing module, which is built-in with a set correction unit and an image enhancement unit to perform set correction and image enhancement on the acquired image respectively;
[0097] A multi-modal recognition engine, which recognizes the preprocessed image. The multi-modal recognition engine is built-in with a layout analysis unit, a text recognition model, and a table parsing algorithm;
[0098] A dynamic verification module, which dynamically verifies the recognized data. The dynamic verification module is built-in with a logical verification layer and a database verification layer;
[0099] An online learning mechanism, which collects the error types of user feedback through the error sample annotation section in the data stream and performs incremental training;
[0100] A template library update module, which updates the template library.
[0101] The processing logic of the set correction unit includes the following steps:
[0102] A1: Edge detection, using the hybrid edge detection network HED, fusing the VGG16 feature extraction and the large-scale edge detection algorithm to output the document body edge coordinate set: E = {(x1, y1),......(x1, y1)};
[0103] A2: Projection matrix calculation, fitting the optimal homography matrix H through the RANSAC algorithm 3×3 H 3×3 is a 3x3 projection transformation matrix, where p iis the coordinate of the i-th feature point of the source image, in the form of (x i , y i , 1) T ; p′ i is the coordinate of the corresponding point in the target image, in the form of (x′ i , y′ i , 1) T ; |||| is the Euclidean distance;
[0104] A3: Perspective transformation is performed, and bilinear interpolation is applied to achieve pixel remapping. The model of bilinear interpolation is: I dst = ∑ i,j I src (x′ + i, y′ + j)·(1 - α)(1 - β), I dst is the pixel value of the target image at the position (x, y); I src is the pixel value of the source image at the position (x' + i, y' + j); α and β are the fractional parts of the sub-pixel offsets, i and j are positive offsets, taking values of 0 or 1.
[0105] The processing logic of the image enhancement unit includes the following steps:
[0106] B1: Noise suppression, using an improved Non-local Means algorithm, whose model is h is the smoothing parameter, σ is the variance, is the intensity value of pixel point x after denoising; C(x) is the normalization coefficient, N(x) is the local neighborhood pixel block centered on x, and Ω is the search window range;
[0107] B2: Contrast enhancement. First, the image is divided into 8×8 sub-blocks, then the histogram of each sub-block is clipped, and finally the processing results are combined using bilinear interpolation;
[0108] B3: Stamp separation, based on threshold segmentation in the HSV color space, whose model is: Mask = (H ∈ [0, 15] ∪ [160, 180]) ∩ (S > 0.4) ∩ (V > 0.3), where H, S, and V are hue, saturation, and value respectively; then the Telea image inpainting algorithm is used to fill the stamp area.
[0109] In the B2 step, clip limit = 2.0.
[0110] The analysis logic of the layout analysis unit includes the following steps:
[0111] C1: Feature extraction. Feature extraction is performed using a feature extraction layer. The feature extraction layer uses ResNeXt-101 instead of the original ResNet50 and a deformable convolutional layer is added.
[0112] C2: Set the anchor box ratios.
[0113] C3: Establish a multi-task loss function. The loss function is L = γ1L1 + γ2L2 + γ3L3 + γ4L4; L1, L2, L3, and L4 are the classification loss, bounding box regression loss, mask prediction loss, and newly added affinity loss respectively, and γ1, γ2, γ3, and γ4 are the weight coefficients.
[0114] C4: Result output. The analysis results are output.
[0115] In the C2 step, the anchor box ratios are set to [0.5, 1, 2, 4].
[0116] In the C3 step, where is the topological feature vector of the i-th predicted region, is the topological feature vector of the i-th corresponding true region, and N is the number of valid regions.
[0117] In the C3 step,
[0118] Example 4:
[0119] A medical device registration certificate OCR recognition system, which includes:
[0120] An image acquisition module, which scans the medical device registration certificate to obtain its image data;
[0121] An intelligent preprocessing module, which is built-in with a set correction unit and an image enhancement unit to perform set correction and image enhancement on the acquired images respectively;
[0122] A multi-modal recognition engine, which recognizes the preprocessed image. The multi-modal recognition engine is built-in with a layout analysis unit, a text recognition model, and a table parsing algorithm;
[0123] A dynamic verification module, which dynamically verifies the recognized data. The dynamic verification module is built-in with a logical verification layer and a database verification layer;
[0124] An online learning mechanism, which collects the error types of user feedback through the error sample annotation section in the data stream and performs incremental training;
[0125] A template library update module, which updates the template library.
[0126] The processing logic of the set calibration unit includes the following steps:
[0127] A1: Edge detection. The hybrid edge detection network HED is used to fuse the VGG16 feature extraction and the large-scale edge detection algorithm to output the document body edge coordinate set: E = {(x1, y1),......(x1, y1)};
[0128] A2: Projection matrix calculation. The optimal homography matrix H is fitted through the RANSAC algorithm 3×3 , H 3×3 is a 3x3 projection transformation matrix, where p i is the coordinate of the i-th feature point of the source image, in the form of (x i , y i , 1) T ; p i ′ is the coordinate of the corresponding point of the target image, in the form of (x′ i , y′ i , 1) T ; |||| is the Euclidean distance;
[0129] A3: Perspective transformation execution. Bilinear interpolation is applied to implement pixel remapping. The model of bilinear interpolation is: I dst = ∑ i,j I src (x′ + i, y′ + i)·(1 - α)(1 - β), I dst is the pixel value of the target image at the (x, y) position; I src is the pixel value of the source image at the (x'+i, y'+j) position; α, β are the fractional parts of the sub-pixel offsets, i, j are positive offsets, taking values of 0 or 1.
[0130] The processing logic of the image enhancement unit includes the following steps:
[0131] B1: Noise suppression. The improved Non-local Means algorithm is used, and its model is h is the smoothing parameter, σ is the variance, is the intensity value of the pixel point x after denoising; C(x) is the normalization coefficient, N(x) is the local neighborhood pixel block centered on x, and Ω is the search window range;
[0132] B2: Contrast enhancement. First, the image is segmented into 8x8 sub-blocks, then the histogram of each sub-block is clipped, and finally the processing results are merged using bilinear interpolation;
[0133] B3: Separation of the seal. Threshold segmentation based on the HSV color space, with the model: Mask = (H ∈ [0, 15] ∪ [160, 180]) ∩ (S > 0.4) ∩ (V > 0.3), where H, S, and V are hue, saturation, and value respectively; then use the Telea image inpainting algorithm to fill the seal area.
[0134] In the B2 step, clip limit = 2.0.
[0135] The analysis logic of the layout analysis unit includes the following steps:
[0136] C1: Feature extraction. Use the feature extraction layer to extract features. The feature extraction layer uses ResNeXt-101 instead of the original ResNet50 and adds a deformable convolutional layer;
[0137] C2: Set the anchor box ratio;
[0138] C3: Establish a multi-task loss function. The loss function is L = γ1L1 + γ2L2 + γ3L3 + γ4L4; L1, L2, L3, and L4 are the classification loss, bounding box regression loss, mask prediction loss, and newly added affinity loss respectively, and γ1, γ2, γ3, and γ4 are the weight coefficients;
[0139] C4: Output the result. Output the analysis result.
[0140] In the C2 step, the anchor box ratio is set to [0.5, 1, 2, 4].
[0141] In the C3 step, where is the topological feature vector of the i-th predicted region, is the topological feature vector of the i-th corresponding real region, and N is the number of valid regions.
[0142] In the C3 step,
[0143] The logic of the table parsing algorithm is as follows:
[0144] D1: Cell detection. Extract the table area based on the layout analysis result and use OpenCV line segment detection;
[0145] D2: Construct the topological relationship. Establish a two-dimensional relationship matrix M m×n , and calculate the cell merging relationship where M is the two-dimensional table structure matrix, and Merge i,j is the flag indicating whether the cell in the i-th row and j-th column is merged with the cell below;
[0146] D3: Semantic alignment, based on semantic matching of table titles Hr is the set of recognized table title words, DB is the standard field library of the drug administration, and Ct is the count of intersection elements of the set.
[0147] The database verification layer performs missing field matching, and its model is as follows: where V oer (f i ) is the i-th field value recognized by ocr; V db (f i ) is the corresponding field value in the database; [] is the Iverson bracket, which is 1 when the condition holds, otherwise 0; ω i is the field weight.
[0148] In the database verification layer, ω i corresponding
[0149] Example 5:
[0150] A medical device registration certificate OCR recognition system, which includes:
[0151] An image acquisition module that scans the medical device registration certificate to obtain its image data;
[0152] An intelligent preprocessing module with a set correction unit and an image enhancement unit built in, which respectively perform set correction and image enhancement on the acquired images;
[0153] A multi-modal recognition engine that recognizes the preprocessed image. The multi-modal recognition engine has a layout analysis unit, a text recognition model, and a table parsing algorithm built in;
[0154] A dynamic verification module that dynamically verifies the recognized data. The dynamic verification module has a logical verification layer and a database verification layer built in;
[0155] An online learning mechanism that collects the error types feedback by users through the error sample annotation section in the data stream and performs incremental training;
[0156] A template library update module that updates the template library.
[0157] The processing logic of the set correction unit includes the following steps:
[0158] A1: Edge detection, using the hybrid edge detection network HED, fusing the VGG16 feature extraction and the large-scale edge detection algorithm to output the document body edge coordinate set: E = {(x1, y1),......(x1, y1)};
[0159] A2: Projection matrix calculation, fitting the optimal homography matrix H through the RANSAC algorithm 3×3 , H 3×3 is a 3x3 projection transformation matrix, where p i is the coordinate of the i-th feature point in the source image, in the form of (x i , y i , 1) T ; p i ′ is the coordinate of the corresponding point in the target image, in the form of (x′ i , y′ i , 1) T ; |||| is the Euclidean distance;
[0160] A3: Perspective transformation execution, applying bilinear interpolation to achieve pixel remapping. The model of bilinear interpolation is: I dst = ∑ i,j I src (x′ + i, y′ + j)·(1 - α)(1 - β), where I dst is the pixel value of the target image at the position (x, y); I src is the pixel value of the source image at the position (x'+i, y'+j); α and β are the fractional parts of the sub-pixel offsets, i and j are positive offsets, taking values of 0 or 1.
[0161] The processing logic of the image enhancement unit includes the following steps:
[0162] B1: Noise suppression, using an improved Non-local Means algorithm, whose model is h is the smoothing parameter, σ is the variance, is the intensity value of pixel point x after denoising; C(x) is the normalization coefficient, N(x) is the local neighborhood pixel block centered on x, and Ω is the search window range;
[0163] B2: Contrast enhancement. First, the image is segmented into 8x8 sub-blocks, then the histogram of each sub-block is clipped, and finally the processing results are merged using bilinear interpolation;
[0164] B3: Stamp separation, based on threshold segmentation in the HSV color space, whose model is: Mask = (H ∈ [0, 15] ∪ [160, 180]) ∩ (S > 0.4) ∩ (V > 0.3), where H, S, and V are hue, saturation, and value respectively; then the Telea image inpainting algorithm is used to fill the stamp area.
[0165] In the B2 step, clip limit = 2.0.
[0166] The analysis logic of the layout analysis unit includes the following steps:
[0167] C1: Feature extraction. Use the feature extraction layer to extract features. The feature extraction layer uses ResNeXt-101 instead of the original ResNet50 and adds a deformable convolutional layer;
[0168] C2: Set the anchor box ratio;
[0169] C3: Establish a multi-task loss function. The loss function is L = γ1L1 + γ2L2 + γ3L3 + γ4L4; L1, L2, L3, and L4 are the classification loss, bounding box regression loss, mask prediction loss, and newly added affinity loss respectively, and γ1, γ2, γ3, and γ4 are the weight coefficients;
[0170] C4: Result output. Output the analysis results.
[0171] In the C2 step, the set anchor box ratio is [0.5, 1, 2, 4].
[0172] In the C3 step, where is the topological feature vector of the i-th predicted region, is the topological feature vector of the i-th corresponding real region, and N is the number of valid regions.
[0173] In the C3 step,
[0174] The logic of the table parsing algorithm is as follows:
[0175] D1: Cell detection. Extract the table area based on the layout analysis result and use OpenCV line segment detection;
[0176] D2: Topological relationship construction. Establish a two-dimensional relationship matrix M m×n , and calculate the cell merging relationship where M is the two-dimensional table structure matrix, and Merge i,j is the flag indicating whether the cell in the i-th row and j-th column is merged with the cell below;
[0177] D3: Semantic alignment. Based on the semantic matching of the table title, Hr is the set of recognized table title words, DB is the drug administration standard field library, and Ct is the count of the intersection elements.
[0178] The database verification layer performs field missing matching, and its model is: where V oer (f i ) is the value of the i-th field recognized by ocr; V db (fi ) is the corresponding field value in the database; [] is the Iverson bracket, which is 1 when the condition is true and 0 otherwise; ω i is the field weight.
[0179] In the database verification layer, ω i corresponding
[0180] Embodiment 6:
[0181] A medical device registration certificate OCR recognition system, which includes:
[0182] An image acquisition module that scans the medical device registration certificate to obtain its image data;
[0183] An intelligent preprocessing module with a set correction unit and an image enhancement unit built in, which respectively perform set correction and image enhancement on the acquired images;
[0184] A multi-modal recognition engine that recognizes the preprocessed image. The multi-modal recognition engine has a layout analysis unit, a text recognition model, and a table parsing algorithm built in;
[0185] A dynamic verification module that dynamically verifies the recognized data. The dynamic verification module has a logical verification layer and a database verification layer built in;
[0186] An online learning mechanism that collects the error types feedback by users through the error sample annotation section in the data stream and performs incremental training;
[0187] A template library update module that updates the template library.
[0188] The processing logic of the set correction unit includes the following steps:
[0189] A1: Edge detection. The hybrid edge detection network HED is used to fuse the VGG16 feature extraction and the large-scale edge detection algorithm to output the document body edge coordinate set: E = {(x1, y1),......(x1, y1)};
[0190] A2: Projection matrix calculation. The optimal homography matrix H is fitted through the RANSAC algorithm 3×3 H 3×3 is a 3x3 projection transformation matrix, where p i is the coordinate of the i-th feature point of the source image, in the form of (x i , y i , 1) T ; p' i is the coordinate of the corresponding point in the target image, in the form of (x' i , y'i , 1) T ; |||| is the Euclidean distance;
[0191] A3: Perspective transformation is performed, and bilinear interpolation is applied to achieve pixel remapping. The model of bilinear interpolation is: I dst = ∑ i,j I src (x′ + i, y′ + i)·(1 - α)(1 - β), I dst is the pixel value of the target image at the position (x, y); I src is the pixel value of the source image at the position (x'+i, y'+j); α and β are the fractional parts of the sub-pixel offsets, i and j are positive offsets, taking values of 0 or 1.
[0192] The processing logic of the image enhancement unit includes the following steps:
[0193] B1: Noise suppression, using an improved Non-local Means algorithm, whose model is h is the smoothing parameter, σ is the variance, is the intensity value of the pixel point x after denoising; C(x) is the normalization coefficient, N(x) is the local neighborhood pixel block centered on x, and Ω is the search window range;
[0194] B2: Contrast enhancement. First, the image is segmented into 8×8 sub-blocks, then the histogram of each sub-block is clipped, and finally the processing results are merged using bilinear interpolation;
[0195] B3: Stamp separation, based on threshold segmentation in the HSV color space, whose model is: Mask = (H ∈ [0, 15] ∪ [160, 180]) ∩ (S > 0.4) ∩ (V > 0.3), where H, S, and V are hue, saturation, and value respectively; then the Telea image inpainting algorithm is used to fill the stamp area.
[0196] In the B2 step, clip limit = 2.0.
[0197] The analysis logic of the layout analysis unit includes the following steps:
[0198] C1: Feature extraction, using the feature extraction layer to perform feature extraction. The feature extraction layer uses ResNeXt-101 instead of the original ResNet50 and adds a deformable convolutional layer;
[0199] C2: Set the anchor box ratio;
[0200] C3: Establishment of multi-task loss function. The loss function is \(L = \gamma_1L_1+\gamma_2L_2+\gamma_3L_3+\gamma_4L_4\); \(L_1\), \(L_2\), \(L_3\), and \(L_4\) are classification loss, bounding box regression loss, mask prediction loss, and newly added affinity loss respectively, and \(\gamma_1\), \(\gamma_2\), \(\gamma_3\), \(\gamma_4\) are weight coefficients respectively.
[0201] C4: Result output. Output the analysis results.
[0202] In the C2 step, the anchor box ratios are set to \([0.5, 1, 2, 4]\).
[0203] In the C3 step, Among them is the topological feature vector of the \(i\)-th predicted region, is the topological feature vector of the \(i\)-th corresponding real region, and \(N\) is the number of valid regions.
[0204] In the C3 step,
[0205] The logic of the table parsing algorithm is as follows:
[0206] D1: Cell detection. Extract the table area based on the layout analysis result and use OpenCV line segment detection.
[0207] D2: Topological relationship construction. Establish a two-dimensional relationship matrix \(M\) m×n , and calculate the cell merging relationship where \(M\) is a two-dimensional table structure matrix, and \(Merge\) i,j is the flag indicating whether the cell in the \(i\)-th row and \(j\)-th column is merged with the cell below.
[0208] D3: Semantic alignment. Based on the semantic matching of the table title, \(H_r\) is the set of recognized table title words, \(DB\) is the drug administration standard field library, and \(Ct\) is the count of intersection elements of the set.
[0209] The database verification layer performs field missing matching, and its model is: Among them, \(V\) oer (\(f\) i ) is the value of the \(i\)-th field recognized by ocr; \(V\) db (\(f\) i ) is the corresponding field value in the database; \([ ]\) is the Iverson bracket, which is 1 when the condition is true and 0 otherwise; \(\omega\) i is the field weight.
[0210] In the database verification layer, \(\omega\) i corresponding
[0211] The above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent substitutions or changes, shall be covered by the protection scope of the present invention.
Claims
1. A medical device registration certificate OCR recognition system, characterized in that, Including: An image acquisition module that scans the medical device registration certificate to obtain its image data; An intelligent preprocessing module with a set correction unit and an image enhancement unit built in, which perform set correction and image enhancement on the acquired images respectively; A multi-modal recognition engine that recognizes the preprocessed images. The multi-modal recognition engine has a layout analysis unit, a text recognition model, and a table parsing algorithm built in; A dynamic verification module that dynamically verifies the recognized data. The dynamic verification module has a logical verification layer and a database verification layer built in; An online learning mechanism that collects the error types feedback by users through the error sample annotation section in the data stream and performs incremental training; A template library update module that updates the template library.
2. The OCR recognition system for a medical device registration certificate according to claim 1, wherein The processing logic of the set correction unit includes the following steps: A1: Edge detection. The hybrid edge detection network HED is used to fuse the VGG16 feature extraction and the large-scale edge detection algorithm to output the document main body edge coordinate set: E = {(x1, y1),......(x1, y1)}; A2: Projection matrix calculation, fitting the optimal homography matrix H through the RANSAC algorithm 3×3 , H 3×3 is a 3x3 projection transformation matrix, where p i is the coordinate of the i-th feature point in the source image, in the form of (x i , y i , 1) T ; p' i is the coordinate of the corresponding point in the target image, in the form of (x' i , y' i , 1) T ; |||| is the Euclidean distance; A3: Perspective transformation is performed, and bilinear interpolation is applied to achieve pixel remapping. The model of bilinear interpolation is: I dst = ∑ i, j I src (x′ + i, y′ + j)·(1 - α)(1 - β), I dst is the pixel value of the target image at the position (x, y); I src is the pixel value of the source image at the position (x' + i, y' + j); α and β are the fractional parts of the sub-pixel offsets, i and j are positive offsets, taking values of 0 or 1.
3. The OCR recognition system for a medical device registration certificate according to claim 1, characterized in that, The processing logic of the image enhancement unit includes the following steps: B1: Noise suppression, using an improved Non-local Means algorithm, whose model is h is the smoothing parameter, σ is the variance, is the intensity value of pixel point x after denoising; C(x) is the normalization coefficient, N(x) is the local neighborhood pixel block centered on x, and Ω is the search window range; B2: Contrast enhancement. First, the image is divided into 8×8 sub-blocks, then the histogram of each sub-block is cropped, and finally the bilinear interpolation is used to merge the processing results; B3: Seal separation. Based on the threshold segmentation in the HSV color space, the model is: Mask = (H∈[0, 15]∪[160, 180])∩(S>0.4)∩(V>0.3), where H, S, and V are hue, saturation, and value respectively; then the Telea image repair algorithm is used to fill the seal area.
4. The OCR recognition system for a medical device registration certificate according to claim 3, characterized in that, In the step B2, clip limit = 2.
0.
5. The OCR recognition system for a medical device registration certificate according to claim 1, characterized in that, The analysis logic of the layout analysis unit includes the following steps: C1: Feature extraction. The feature extraction layer is used for feature extraction. The ResNeXt-101 is used instead of the original ResNet50 in the feature extraction layer, and a deformable convolutional layer is added; C2: Set the anchor box ratio; C3: Establish a multi-task loss function. The loss function is L = γ1L1 + γ2L2 + γ3L3 + γ4L4; L1, L2, L3, and L4 are the classification loss, the bounding box regression loss, the mask prediction loss, and the newly added affinity loss respectively, and γ1, γ2, γ3, and γ4 are the weight coefficients; C4: Output the result. Output the analysis result.
6. The OCR recognition system for a medical device registration certificate according to claim 5, characterized in that, In the step C2, the set anchor box ratio is [0.5, 1, 2, 4].
7. The OCR recognition system for a medical device registration certificate according to claim 5, wherein, In the step C3, where P1 i is the topological feature vector of the i-th predicted region, and P2 i is the topological feature vector of the i-th corresponding true region, and N is the number of valid regions, 8. The OCR recognition system for a medical device registration certificate according to claim 1, characterized in that, The logic of the table parsing algorithm is: D1: Cell detection. Extract the table area based on the layout analysis result and use OpenCV line segment detection; D2: Topological relationship construction, establishing a two-dimensional relationship matrix M m×n , calculating the cell merging relationship where M is a two-dimensional table structure matrix, and Merge i,j is a flag indicating whether the cell in the i-th row and j-th column is merged with the cell below; D3: Semantic alignment, based on semantic matching of table titles Hr is the set of recognized table title words, DB is the standard field library of the drug administration, and Ct is the count of intersection elements of the set.
9. The OCR recognition system for a medical device registration certificate according to claim 1, wherein The database verification layer performs field missing matching, and its model is as follows: Among them, V oer (f i ) is the value of the i-th field recognized by ocr; V db (f i ) is the corresponding field value in the database; [] is the Iverson bracket, which is 1 when the condition is true and 0 otherwise; ω i is the field weight.
10. The OCR recognition system for a medical device registration certificate according to claim 9, characterized in that, In the database verification layer, ω i corresponding
Citation Information
Cited By
Extensible OCR intelligent identification method and system based on template priority and adaptive model optimization
CN121354141A