Clinical examination report character recognition method and system and related equipment
By training clinically dedicated OCR models and table processing models, the clinical test report is preprocessed and structured identification, which solves the problems of insufficient identification accuracy and limited structural analysis capabilities in the existing technology, and realizes high-precision clinical test report identification and analysis.
Patent Information
- Application Number
- CN202510212228.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-13
AI Technical Summary
When processing clinical test reports, the prior art lacks recognition accuracy and limited structural analysis capabilities, resulting in the limitation of the digitalization level of the medical system.
By training the text recognition OCR model based on the historical clinical test report image set and related knowledge base set, a clinically dedicated OCR model was obtained; receiving the clinical test report image uploaded by the user and pre-processing; table recognition is performed through the table processing model, structured text/numerical recognition is performed in combination with the clinically dedicated OCR model, and finally arranged and displayed according to the report format of the display page.
It significantly improves the text recognition accuracy and table structure analysis capabilities of clinical test reports, meeting the medical system's needs for high-precision recognition.
Smart Images

Figure CN120148055A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of electronic processing of reports, and particularly to a method, a system and related devices for recognizing text in clinical test reports. Background Art
[0002] With the rapid development of medical informatization, the demand for digital processing of clinical test reports is increasing day by day. At present, in scenarios such as health screening programs, chronic disease management, drug efficacy research, and cross-institutional collaboration, we need to manually enter the paper-based clinical test reports provided by patients into an electronic system, which is time-consuming, laborious and error-prone. Traditional text recognition technology can automatically recognize the text in the paper report and convert it into an electronic format for easy storage and retrieval. However, since clinical test reports contain important diagnostic data of patients and have complex formats and diverse contents, including various data forms such as text, tables, and pathological images, there are many deficiencies in the existing technology when processing such reports, specifically manifested as:
[0003] Image quality problems: The report may be affected by blurry, tilted or blocked shooting, which affects the recognition accuracy;
[0004] Complex table structures: Diverse table formats and item data relationships result in incorrect parsing of the text recognition results; Insufficient recognition of medical terms: Clinical reports contain a large number of medical terms (such as "ALT", "WBC"), and traditional text recognition models lack targeted optimization, resulting in low recognition rates.
[0005] In summary, when processing clinical test reports, the existing technology has insufficient recognition accuracy and limited structural analysis ability, which restricts the digital level of the medical system. There is an urgent need for an optimized text recognition solution specifically for medical scenarios.
[0006] Therefore, the existing technology still needs to be improved and developed. Summary of the Invention
[0007] The present invention provides a method, a system and related devices for recognizing text in clinical test reports. The main purpose of the present invention is to solve the technical problems mentioned in the background art of the existing technology.
[0008] The first aspect of the present invention provides a method for recognizing text in clinical test reports, including:
[0009] Training an optical character recognition (OCR) model for text recognition based on a historical clinical test report image set and a related knowledge base set to obtain a clinical-specific OCR model;
[0010] Receive the clinical test report form image uploaded by the user, and preprocess the clinical test report form image to obtain an enhanced clinical test report form image;
[0011] Perform table recognition on the enhanced clinical test report form image through a table processing model to obtain a structured cell image set;
[0012] Input the structured cell image set into the clinical-specific OCR model to obtain a structured text / value recognition data set;
[0013] Arrange and display the structured text / value recognition data set according to the report format of the display page.
[0014] In an alternative embodiment of the first aspect of the present invention, the receiving the clinical test report form image uploaded by the user and preprocessing the clinical test report form image to obtain an enhanced clinical test report form image includes:
[0015] Receive the clinical test report form image uploaded by the user;
[0016] Denoise the clinical test report form image using a Gaussian filter;
[0017] Enhance the contrast of the denoised clinical test report form image using a histogram equalization algorithm;
[0018] Perform skew correction on the contrast-enhanced clinical test report form image using the Hough transform algorithm;
[0019] Perform Fourier transform and band-stop filtering on the skew-corrected clinical test report form image to obtain an enhanced clinical test report form image.
[0020] In an alternative embodiment of the first aspect of the present invention, the performing Fourier transform and band-stop filtering on the skew-corrected clinical test report form image to obtain an enhanced clinical test report form image includes:
[0021] Perform Fourier transform on the skew-corrected clinical test report form image to convert the clinical test report form image from the spatial domain to the frequency domain;
[0022] Determine the moiré frequency region by analyzing the clinical test report form image in the frequency domain;
[0023] Perform band-stop filtering on the clinical test report form image in the moiré frequency region through a band-stop filter;
[0024] Perform inverse Fourier transform on the band-stop filtered clinical test report form image to obtain an enhanced clinical test report form image.
[0025] In an alternative embodiment of the first aspect of the present invention, the table recognition of the enhanced clinical test report image by the table processing model to obtain a structured cell image set includes:
[0026] Based on the Tesseract engine, perform table area detection on the enhanced clinical test report image to determine the target table area;
[0027] Perform row-column structure analysis on the target table area to obtain the grid structure of the target table area;
[0028] Establish the logical relationship between cells in the grid structure;
[0029] Extract the cell images in the grid structure and store them according to the logical relationship to obtain a structured cell image set.
[0030] In an alternative embodiment of the first aspect of the present invention, the input of the structured cell image set into the clinical-specific OCR model to obtain a structured text / numerical recognition data set includes:
[0031] Perform data type recognition on the structured cell image set through the clinical-specific OCR model to obtain a text cell image set and a numerical cell image set;
[0032] Perform font type recognition on the text cell image set to obtain a standard font cell image set and a non-standard font cell image set;
[0033] Match the standard font cell image set with the font type character library to obtain a first standard term text set;
[0034] Use the Jaccard similarity algorithm to perform text matching on the non-standard font cell image set with the thesaurus to obtain a second standard term text set of the non-standard font cells;
[0035] Perform digital recognition on the numerical cell image set to obtain a cell numerical set;
[0036] Restructure and bind the first standard term text set, the second standard term text set, and the cell numerical set to obtain a structured text / numerical recognition data set.
[0037] In an alternative embodiment of the first aspect of the present invention, the arrangement and display of the structured text / numerical recognition data set according to the report format of the display page includes:
[0038] Based on the report format of the display page, obtain the display position information of the structured text / numerical recognition data set on the display page;
[0039] Display the structured text / numerical recognition data set into the display page in sequence according to the order within the data set and the display position information.
[0040] In an optional implementation manner of the first aspect of the present invention, after arranging and displaying the structured text / numerical recognition data set according to the report format of the display page, it includes:
[0041] Receive the click of the user on the specified content on the display page;
[0042] Generate an edit box for the specified content in response to the click;
[0043] Receive the editing confirmation of the user for the content input in the edit box;
[0044] Update the specified content with the input content.
[0045] The second aspect of the present invention provides a clinical test report form text recognition system, and the clinical test report form text recognition system includes:
[0046] A model training module, configured to train an optical character recognition (OCR) model based on a historical clinical test report form image set and a related knowledge base set to obtain a clinical-specific OCR model;
[0047] An image preprocessing module, configured to receive a clinical test report form image uploaded by the user and preprocess the clinical test report form image to obtain an enhanced clinical test report form image;
[0048] A table recognition module, configured to perform table recognition on the enhanced clinical test report form image through a table processing model to obtain a structured cell image set;
[0049] An OCR recognition module, configured to input the structured cell image set into the clinical-specific OCR model to obtain a structured text / numerical recognition data set;
[0050] A report display module, configured to arrange and display the structured text / numerical recognition data set according to the report format of the display page.
[0051] The third aspect of the present invention provides a clinical test report form text recognition device, and the clinical test report form text recognition device includes: a memory and at least one processor, instructions are stored in the memory, and the memory and the at least one processor are interconnected through a line;
[0052] The at least one processor invokes the instructions in the memory to cause the clinical test report form text recognition device to execute the clinical test report form text recognition method according to any one of the first aspects of the present invention as described above.
[0053] The fourth aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the clinical test report form text recognition method according to any one of the first aspects of the present invention as described above.
[0054] Beneficial effects: The present invention provides a clinical test report form text recognition method, system and related devices. The method includes training an optical character recognition (OCR) model based on a historical clinical test report form image set and a related knowledge base set to obtain a clinical-specific OCR model; receiving a clinical test report form image uploaded by a user and preprocessing the clinical test report form image to obtain an enhanced clinical test report form image; performing table recognition on the enhanced clinical test report form image through a table processing model to obtain a structured cell image set; inputting the structured cell image set into the clinical-specific OCR model to obtain a structured text / numerical value recognition data set; and arranging and displaying the structured text / numerical value recognition data set in the report format of the display page. The solution of the present invention solves the problems of low text recognition accuracy caused by the quality of the report image, complex tables and insufficient terms in the prior art, which are difficult to meet the requirements of the medical system. Description of the Drawings
[0055] Figure 1 It is a schematic diagram of an embodiment of the main method steps of a clinical test report form text recognition method of the present invention;
[0056] Figure 2 It is a schematic diagram of an embodiment of a clinical test report form text recognition system of the present invention;
[0057] Figure 3 It is a schematic diagram of an embodiment of a clinical test report form text recognition device of the present invention. Detailed Embodiments
[0058] In the description, claims and the above-mentioned drawings of the present invention, the terms "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments described herein can be implemented in an order different from that shown or described herein. In addition, the term "comprising" or "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0059] For ease of understanding, the specific process of the embodiments of the present invention will be described below. Please refer to Figure 1 , the first aspect of the present invention provides a method for recognizing text in a clinical test report. The method can run in the form of software and programs, and the execution subject of the method can be a mobile terminal and a computer. The method for recognizing text in a clinical test report includes:
[0060] S100. Based on the historical clinical test report image set and the relevant knowledge base set, train an optical character recognition (OCR) model for text recognition to obtain a clinical-specific OCR model. In this step of the present invention, in order to better train the model, the present invention can establish a medical term database and a synonym database, collect a large number of medical terms, including common disease names, test item names, index units, etc. At the same time, sort out the synonyms of each medical term. For example, the synonyms of "Quantitative HBcAb" may include "Hepatitis B core antibody", etc. Organize these terms and synonyms in the form of a database for convenient subsequent matching use.
[0061] After that, based on a large amount of clinical test report data and the relevant knowledge base, train an OCR model specifically for the medical field (such as the OCR model based on Tencent Cloud). The knowledge base contains information such as common medical terms and the structural characteristics of test reports. During the training process, the model not only learns the characteristics of the text, but also combines medical knowledge and report structure to better understand and recognize the text in the test report. For example, by learning the common item arrangement order in the test report, the model can more accurately identify the item names and corresponding data. Compared with the general OCR model, the trained medical-specific OCR model has significantly enhanced efficiency and accuracy in recognizing the text in clinical test reports.
[0062] S200: Receive the clinical test report form image uploaded by the user, and preprocess the clinical test report form image to obtain an enhanced clinical test report form image. In the preprocessing of this solution, it includes, but is not limited to, denoising, enhancing contrast, correcting skew, and detecting and eliminating moiré patterns.
[0063] More specifically, in an alternative implementation manner of step S200 of the present invention, the receiving the clinical test report form image uploaded by the user and preprocessing the clinical test report form image to obtain an enhanced clinical test report form image includes:
[0064] Receive the clinical test report form image uploaded by the user. In this step, an entry for the user to upload the clinical test report form image is provided on the software interface of the method.
[0065] Use a Gaussian filter to denoise the clinical test report form image. In the present invention, the clinical test report form image input by the user may be interfered by various noises, such as Gaussian noise and salt-and-pepper noise during the shooting process. Use a Gaussian filter to perform denoising processing on the image. The principle of the Gaussian filter is to perform weighted averaging on each pixel point and its neighborhood in the image according to the Gaussian function. For a two-dimensional Gaussian function, determines the width of the Gaussian function, that is, the smoothing degree of the filter. By adjusting the value, the edges and details of the image can be retained as much as possible while removing the noise.
[0066] Adopt the histogram equalization algorithm to enhance the contrast of the denoised clinical test report form image. In the present invention, the method of histogram equalization is used to enhance the contrast of the image. In histogram equalization, the gray histogram of the image is adjusted so that the gray value distribution of the image is more uniform, thereby enhancing the contrast of the image. The specific approach is to calculate the gray histogram of the image, representing the number of pixel points with gray value, and then map the gray value of each pixel according to the cumulative distribution function.
[0067] Use the Hough transform algorithm to correct the skew of the clinical test report form image with enhanced contrast. In the present invention, the Hough transform is used to detect straight lines in the image. By detecting the horizontal and vertical straight lines in the image, the skew angle of the image is calculated. The Hough transform maps the straight lines in the image space to the parameter space. For the straight line equation, the straight lines in the image are determined by finding the peaks in the parameter space.
[0068] Perform Fourier transform and band-stop filtering on the clinically examined report form image after skew correction to obtain an enhanced clinically examined report form image. In the present invention, this step may specifically include: performing Fourier transform on the clinically examined report form image after skew correction to convert the clinically examined report form image from the spatial domain to the frequency domain. In the frequency domain, moiré patterns are manifested as components of a specific frequency; by analyzing the clinically examined report form image in the frequency domain, determine the moiré frequency region; then design a suitable filter, such as a band-stop filter, to suppress these frequency components, and perform band-stop filtering on the clinically examined report form image in the moiré frequency region through the band-stop filter; finally, perform inverse Fourier transform on the clinically examined report form image after band-stop filtering to convert the clinically examined report form image from the frequency domain back to the spatial domain to obtain an image with moiré patterns removed (i.e., obtain an enhanced clinically examined report form image).
[0069] S300. Perform table recognition on the enhanced clinically examined report form image through a table processing model to obtain a structured cell image set. Specifically, when traditional character recognition technologies process tables, they can usually only recognize data in a single row or column, and have a low processing accuracy for complex tables (such as nested tables and irregular tables), and cannot accurately capture the logical relationships between cells. The technical solution of the present invention introduces table area detection and row-column structure analysis technologies of a table processing model (based on the Tesseract engine), which can accurately recognize various table styles (such as fixed tables, nested tables, and irregular tables), and establish logical relationships between cells, such as the correspondence between titles and data, and the association between the main table and the sub-table, improving the parsing accuracy and complexity of the table structure.
[0070] That is, in an optional implementation manner of step S300 of the present invention, the performing table recognition on the enhanced clinically examined report form image through a table processing model to obtain a structured cell image set includes:
[0071] Perform table area detection on the enhanced clinically examined report form image based on the Tesseract engine to determine the target table area. Based on the Tesseract engine, the present invention first performs table area detection to accurately identify the part containing table content, avoiding interference from other noises (such as text or images), and supporting various table styles (fixed tables, nested tables, and irregular tables).
[0072] Perform row-column structure analysis on the target table area to obtain the grid structure of the target table area; establish the logical relationships between cells in the grid structure. After the Tesseract engine identifies the table area, it will perform table row-column structure analysis to construct a complete grid structure; and establish logical relationships between cells, such as the correspondence between titles and data, and the association between the main table and the sub-table.
[0073] Extract the cell images in the grid structure and store them according to the logical relationship to obtain a structured cell image set. In the present invention, the Tesseract engine initially identifies the content types (text, numbers, dates, special symbols, etc.) of the cells, supports inferring hidden structure information from the context, such as the correspondence between multi-level table headers and data columns, and finally completes the parsing of the table structure. Through the parsed table structure, the picture data of each cell is extracted in the order of rows and columns. During the extraction process, combined with the results of image preprocessing, the text in the cell is more accurately located and recognized. Since the structure of the table has been clarified, misrecognition of other areas during the recognition process is avoided, greatly improving the accuracy of table data recognition.
[0074] S400. Input the structured cell image set into the clinical-specific OCR model to obtain a structured text / value recognition data set; Traditional OCR technologies are difficult to effectively recognize medical-specific terms (such as disease names, test item names, etc.) and their variants when processing medical text. The coverage of the thesaurus and medical terms is limited, and the recognition rate is low. In the present invention, a dedicated medical term library and thesaurus are established, and a medical-specific OCR model is trained using a training set of historical clinical test report images and related knowledge base sets. Then, combined with a similarity algorithm (such as the Jaccard similarity algorithm), the OCR recognition results are matched, thereby improving the matching accuracy of medical terms. In addition, an OCR model for the medical field is trained to optimize the understanding and recognition of specific terms and structures in clinical test reports.
[0075] Specifically, in an optional implementation manner of step S400 of the present invention, the inputting the structured cell image set into the clinical-specific OCR model to obtain a structured text / value recognition data set includes:
[0076] Perform data type recognition on the structured cell image set through the clinical-specific OCR model to obtain a text cell image set and a numerical cell image set; In the technical solution of the present invention, the text image part and the numerical image part in the structured cell image set are processed separately. In the present invention, the clinical OCR model architecture adopts a multi-task learning framework, such as integrating the YOLOv8 table detection module and the PP-OCRv4 recognition module. Domain adaptation training: Use 3 million medical reports for adversarial training to enhance the recognition ability of scribbled medical orders. Dynamic threshold segmentation: Adaptively adjust the binarization parameters based on local contrast to handle problems such as faded text.
[0077] Perform font type recognition on the set of text cell images to obtain a set of standard font cell images and a set of non-standard font cell images. In the present invention, for the set of text cell images, further distinction is made based on whether the text cell images are in standard fonts. The comparison of standard fonts mainly uses a pre-constructed standard font library during use. Different matching methods for the set of standard font cell images and the set of non-standard font cell images can improve the matching efficiency of the text while ensuring the accuracy of text recognition.
[0078] Perform matching of the font type's character library on the set of standard font cell images to obtain a first set of standard term texts. In the present invention, the standard font library (i.e., the character library corresponding to the font type) is constructed by establishing a feature matrix covering 12 types of medical standard fonts (FangSong_GB2312, KaiTi_GB2312, etc.), and using StyleGAN2 to generate font mutation samples to enhance the generalization ability of the model. The Fourier descriptor is introduced for font contour feature extraction to construct a 128-dimensional feature vector.
[0079] Perform text matching on the set of non-standard font cell images with a thesaurus using the Jaccard similarity algorithm to obtain a second set of standard term texts for the non-standard font cell images. In the present invention, after obtaining the data of non-standard font cell images through OCR recognition, use the Jaccard similarity algorithm, etc. to match with the data in the thesaurus. Jaccard is a measure for measuring the similarity between two sets. It is defined by calculating the ratio of the size of the intersection of two sets to the size of their union. The value range of the Jaccard similarity coefficient is [0, 1], where 1 indicates that the two sets are exactly the same, and 0 indicates that the two sets have no common elements. For example, assume the text recognized by OCR is "aspartate aminotransferase", and search for terms with high similarity in the thesaurus, and it may match "AST", etc. In this way, the success rate of data matching is improved, ensuring that the recognized medical terms can accurately correspond to the standard terms. During the Jaccard similarity comparison process, the n-gram weighting algorithm can also be introduced: tri-gram weight 0.6, bi-gram 0.3, uni-gram 0.1, and the Levenshtein distance calculation is fused to calculate the mixed similarity: Similarity = αJaccard+(1-α)(1-Levenshtein / max_len); set a dynamic threshold mechanism: automatically adjust the matching threshold according to the word length (4-character word threshold 0.85, 8-character word 0.7).
[0080] Perform digital recognition on the numerical cell image set to obtain a cell numerical set; in the present invention, the digital recognition can use a numerical sequence recognition model developed based on CRF, which can support the construction of a medical unit conversion rule library for complex expressions such as "12.3×10^9 / L" (e.g., "ten thousand units" → "10^4U / L"), as well as an outlier detection module for dynamic range verification based on Tukey's fences algorithm to further improve the accuracy of digital recognition.
[0081] Restructure and bind the first standard term text set, the second standard term text set, and the cell numerical set again to obtain a structured text / numerical recognition data set. In the invention, the recognized text and numerical values will be bound and stored according to the mapping relationship in the table for convenient subsequent data loading on the display interface.
[0082] S500. Arrange and display the structured text / numerical recognition data set according to the report format of the display page. In the present invention, step S500 may specifically include: obtaining the display position information of the structured text / numerical recognition data set on the display page based on the report format of the display page; sequentially displaying the structured text / numerical recognition data set into the display page according to the order within the data set and the display position information. In the present invention, after the recognition result comes out, all the report data information is displayed in a clear format on the system page, and information such as test items, results, and reference ranges is arranged according to a certain layout for the convenience of users to view.
[0083] In an optional implementation manner of the first aspect of the present invention, after arranging and displaying the structured text / numerical recognition data set according to the report format of the display page, it includes: receiving a click by the user on the specified content on the display page; generating an edit box for the specified content in response to the click; receiving the user's editing confirmation of the content input in the edit box; and updating the specified content with the input content. In the present invention, the method also provides functions of manual correction and outlier marking. The technical solution of the present invention allows the user to manually correct possible errors in the recognition result. While displaying the data, an editing function is provided, and the user can directly modify the recognized incorrect text or data on the page. At the same time, according to the reference range of the reference index, outliers are marked. A confirmation function is provided. After the user checks and corrects the data, clicks the confirmation button, and the system will save the final recognition and correction results to complete the text recognition process of the entire clinical test report form.
[0084] Generally speaking, the specific improvements of the present invention are reflected in the following aspects:
[0085] 1. Enhancement of table structure parsing module: Prior art: When traditional text recognition technology processes tables, it can usually only recognize data in a single row and column. The processing accuracy for complex tables (such as nested tables and irregular tables) is low, and the logical relationship between cells cannot be accurately captured. Improvement of the present invention: The table area detection and row and column structure parsing technology based on the Tesseract engine is introduced, which can accurately identify a variety of table styles (such as fixed tables, nested tables, irregular tables), and establish logical relationships between cells, such as the correspondence between titles and data, the association between main tables and sub-tables, etc., which improves the parsing accuracy and complexity of the table structure.
[0086] 2. Medical terminology recognition and synonym matching: Prior art: When processing medical texts, traditional OCR technology has difficulty in effectively identifying medical terminology (such as disease names, test item names, etc.) and their variants. The coverage of synonym libraries and medical terms is limited, and the recognition rate is low. Improvement of the present invention: This solution improves the matching accuracy of medical terms by establishing a special medical terminology library and synonym library, and matching the OCR recognition results with a similarity algorithm (such as the Jaccard similarity algorithm). In addition, an OCR model for the medical field has been trained to optimize the understanding and recognition of specific terms and structures in clinical test reports.
[0087] 3. Integrate multiple optimization methods: Existing technology: When processing medical documents, image preprocessing, table parsing and term matching are usually performed separately, lacking overall optimization and linkage. Improvement of the present invention: This solution achieves a close integration between image preprocessing, table parsing and medical term recognition, and greatly improves the accuracy of recognition results and processing efficiency through accurate recognition and matching of table structures and medical terms.
[0088] Based on the improvements of the present invention described above, in general, the technical solution of the present invention can achieve the following beneficial effects:
[0089] (1) Improve OCR recognition accuracy: By introducing customized thesaurus and synonyms for medical terms and combining them with similarity algorithms, the occurrence of misrecognition and missed recognition can be significantly reduced.
[0090] (2) Enhanced table structure parsing capabilities: Effectively handle complex table structures, including cross-row and cross-column cells, etc. Output standardized data format to facilitate subsequent analysis and processing.
[0091] (3) Improve image quality: Introduce a moiré removal algorithm (frequency domain-based algorithm) to optimize the quality of input images and provide a better basis for OCR recognition. Solve the problem of image blur caused by shooting conditions and other reasons.
[0092] (4) Outlier detection and annotation: Identify and mark outliers to improve the efficiency and accuracy of data analysis.
[0093] (5) Wide applicability: Support various formats and types of inspection reports, suitable for the actual needs of different medical institutions. It has strong scalability and can adjust and optimize the module design according to specific scenario requirements. The technical solution of the present invention effectively solves the problems of insufficient recognition accuracy and weak table parsing ability in the prior art, and provides efficient and reliable technical support for the digital transformation of the medical industry.
[0094] See Figure 2 , the second aspect of the present invention provides a clinical inspection report text recognition system, and the clinical inspection report text recognition system includes:
[0095] A model training module 10, configured to train an optical character recognition (OCR) model based on a historical clinical inspection report image set and a related knowledge base set to obtain a clinical-specific OCR model;
[0096] An image preprocessing module 20, configured to receive a clinical inspection report image uploaded by a user and preprocess the clinical inspection report image to obtain an enhanced clinical inspection report image;
[0097] A table recognition module 30, configured to perform table recognition on the enhanced clinical inspection report image through a table processing model to obtain a structured cell image set;
[0098] An OCR recognition module 40, configured to input the structured cell image set into the clinical-specific OCR model to obtain a structured text / numerical recognition data set;
[0099] A report display module 50, configured to arrange and display the structured text / numerical recognition data set according to the report format of the display page.
[0100] In an optional implementation manner of the second aspect of the present invention, the image preprocessing module includes:
[0101] An image receiving unit, configured to receive a clinical inspection report image uploaded by a user;
[0102] An image denoising unit, configured to denoise the clinical inspection report image using a Gaussian filter;
[0103] An image enhancement unit, configured to enhance the contrast of the denoised clinical inspection report image by using a histogram equalization algorithm;
[0104] An image correction unit, configured to perform skew correction on the contrast-enhanced clinical inspection report image by using a Hough transform algorithm;
[0105] The moiré pattern elimination unit is configured to perform Fourier transform and band-stop filtering on the clinically examined report form image after skew correction to obtain an enhanced clinically examined report form image.
[0106] In an optional implementation manner of the second aspect of the present invention, the moiré pattern elimination unit includes:
[0107] The Fourier transform subunit is configured to perform Fourier transform on the clinically examined report form image after skew correction, and convert the clinically examined report form image from the spatial domain to the frequency domain;
[0108] The region determination subunit is configured to determine the moiré pattern frequency region by analyzing the clinically examined report form image in the frequency domain;
[0109] The band-stop filtering processing subunit is configured to perform band-stop filtering on the clinically examined report form image in the moiré pattern frequency region through a band-stop filter;
[0110] The inverse Fourier transform subunit is configured to perform inverse Fourier transform on the clinically examined report form image after band-stop filtering to obtain an enhanced clinically examined report form image.
[0111] In an optional implementation manner of the second aspect of the present invention, the form recognition module includes:
[0112] The form detection unit is configured to detect the form area of the enhanced clinically examined report form image based on the Tesseract engine to determine the target form area;
[0113] The form relationship parsing unit is configured to perform row-column structure parsing on the target form area to obtain the grid structure of the target form area;
[0114] The cell relationship establishment unit is configured to establish the logical relationship between cells in the grid structure;
[0115] The cell image storage unit is configured to extract the cell images in the grid structure and store them according to the logical relationship to obtain a structured cell image set.
[0116] In an optional implementation manner of the second aspect of the present invention, the OCR recognition module includes:
[0117] The data type recognition unit is configured to perform data type recognition on the structured cell image set through the clinical dedicated OCR model to obtain a text cell image set and a numerical cell image set;
[0118] A font type recognition unit for recognizing the font type of the text cell image set to obtain a standard font cell image set and a non-standard font cell image set;
[0119] A first term text matching unit for matching the standard font cell image set with the font type character library to obtain a first standard term text set;
[0120] A second term text matching unit for performing text matching on the non-standard font cell image set with a thesaurus using the Jaccard similarity algorithm to obtain a second standard term text set of the non-standard font cell image set;
[0121] A numerical text matching unit for performing digital recognition on the numerical cell image set to obtain a cell numerical set;
[0122] A structuring processing unit for restructuring and binding the first standard term text set, the second standard term text set, and the cell numerical set to obtain a structured text / numerical recognition data set.
[0123] In an optional implementation manner of the second aspect of the present invention, the report display module includes:
[0124] A display position information acquisition unit for acquiring the display position information of the structured text / numerical recognition data set on the display page based on the report format of the display page;
[0125] A data display unit for sequentially displaying the structured text / numerical recognition data set into the display page according to the order within the data set and the display position information.
[0126] In an optional implementation manner of the second aspect of the present invention, the report display module includes:
[0127] A click reception unit for receiving a click by the user on a specified content on the display page;
[0128] An edit box generation unit for generating an edit box for the specified content in response to the click;
[0129] An edit confirmation unit for receiving an edit confirmation of the content input by the user in the edit box;
[0130] A display content update unit for updating the specified content with the input content.
[0131] Figure 3FIG. 0 is a schematic structural diagram of a clinical test report form text recognition device provided by an embodiment of the present invention. The clinical test report form text recognition device may vary greatly due to different configurations or performances, and may include one or more processors 60 (central processing units, CPUs) (for example, one or more processors) and a memory 70, and one or more storage media 80 (for example, one or more mass storage devices) for storing application programs or data. Among them, the memory and the storage medium may be transient storage or persistent storage. The program stored in the storage medium may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the clinical test report form text recognition device. Further, the processor may be configured to communicate with the storage medium and execute a series of instruction operations in the storage medium on the clinical test report form text recognition device.
[0132] The clinical test report form text recognition device of the present invention may further include one or more power supplies 90, one or more wired or wireless network interfaces 100, one or more input / output interfaces 110, and / or, one or more operating systems, such as Windows Serve, Mac OS X, Unix, Linux, FreeBSD, etc. Those skilled in the art can understand that Figure 3 the shown structural diagram of the clinical test report form text recognition device does not limit the clinical test report form text recognition device, and may include more or fewer components than shown, or combine some components, or have different component arrangements.
[0133] The present invention also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions run on a computer, the computer is caused to execute the steps of the clinical test report form text recognition method.
[0134] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described system or system and unit can refer to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0135] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0136] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A method for recognizing text in a clinical test report, characterized in that: include: Based on the historical clinical test report image set and related knowledge base set, the text recognition OCR model is trained to obtain a clinical-specific OCR model; Receiving a clinical test report image uploaded by a user, and preprocessing the clinical test report image to obtain an enhanced clinical test report image; Performing table recognition on the enhanced clinical test report image through a table processing model to obtain a structured cell image set; Inputting the structured cell image set into the clinical-specific OCR model to obtain a structured text / numeric recognition data set; The structured text / numeric recognition data set is arranged and displayed according to the report format of the display page.
2. The clinical test report text recognition method according to claim 1, characterized in that: The receiving of the clinical test report image uploaded by the user and preprocessing the clinical test report image to obtain an enhanced clinical test report image includes: Receive clinical test report images uploaded by users; De-noising the clinical test report image using a Gaussian filter; Using a histogram equalization algorithm to enhance the contrast of the denoised clinical test report image; Using the Hough transform algorithm to perform tilt correction on the contrast-enhanced clinical test report image; The clinical test report image after tilt correction is subjected to Fourier transformation and band-stop filtering to obtain an enhanced clinical test report image.
3. The clinical test report text recognition method according to claim 2, characterized in that: The performing Fourier transform and band-stop filtering on the tilt-corrected clinical test report image to obtain an enhanced clinical test report image comprises: Performing Fourier transformation on the tilt-corrected clinical test report image to convert the clinical test report image from the spatial domain to the frequency domain; Determine the moiré frequency region by analyzing the clinical test report image in the frequency domain; Performing band-stop filtering on the clinical test report single image in the moiré frequency region by means of a band-stop filter; The clinical test report image after band-stop filtering is subjected to inverse Fourier transformation to obtain an enhanced clinical test report image.
4. The clinical test report text recognition method according to claim 3, characterized in that: The step of performing table recognition on the enhanced clinical test report image by using a table processing model to obtain a structured cell image set comprises: Performing table area detection on the enhanced clinical test report image based on the Tesseract engine to determine the target table area; Performing row and column structure analysis on the target table area to obtain a grid structure of the target table area; Establishing a logical relationship between cells in the grid structure; The cell images in the grid structure are extracted and stored according to the logical relationship to obtain a structured cell image set.
5. The clinical test report text recognition method according to claim 4, characterized in that: The step of inputting the structured cell image set into the clinically dedicated OCR model to obtain a structured text / numerical value recognition data set comprises: Using the clinically dedicated OCR model, data type recognition is performed on the structured cell image set to obtain a text cell image set and a numerical cell image set; Performing font type recognition on the text cell image set to obtain a standard font cell image set and a non-standard font cell image set; Matching the standard font cell image set with a text library of the font type to obtain a first standard terminology text set; Using the Jaccard similarity algorithm to perform text matching on the non-standard font cell image set and a synonym library to obtain a second standard term text set for the non-standard font cell image; Performing digital recognition on the numerical cell image set to obtain a cell numerical value set; The first standard terminology text set, the second standard terminology text set and the cell value set are restructured and bound to obtain a structured text / value recognition data set.
6. The clinical test report text recognition method according to claim 1, characterized in that: Arranging and displaying the structured text / numeric recognition data set according to the report format of the display page includes: Based on the report format of the display page, obtaining display position information of the structured text / numeric recognition data set on the display page; The structured text / numeric recognition data set is sequentially displayed on the display page according to the order in the data set and the display position information.
7. The clinical test report text recognition method according to claim 1, characterized in that: Arranging and displaying the structured text / numeric recognition data set according to the report format of the display page includes: Receiving a user's click on designated content on the displayed page; generating an edit box of the specified content in response to the click; Receiving a user's edit confirmation of the content input in the edit box; The designated content is updated according to the input content.
8. A clinical test report text recognition system, characterized in that: The clinical test report text recognition system comprises: Model training module, used to train text recognition OCR model based on historical clinical test report image set and related knowledge base set to obtain clinical-specific OCR model; An image preprocessing module is used to receive a clinical test report image uploaded by a user, and preprocess the clinical test report image to obtain an enhanced clinical test report image; A table recognition module, used to perform table recognition on the enhanced clinical test report image through a table processing model to obtain a structured cell image set; An OCR recognition module, used for inputting the structured cell image set into the clinical-specific OCR model to obtain a structured text / numeric recognition data set; The report display module is used to arrange and display the structured text / numeric recognition data set in a report format of a display page.
9. A clinical test report text recognition device, characterized in that: The clinical test report text recognition device comprises: a memory and at least one processor, the memory stores instructions, and the memory and the at least one processor are interconnected via a line; The at least one processor calls the instructions in the memory to enable the clinical test report text recognition device to execute the clinical test report text recognition method as described in any one of claims 1-7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for text recognition in a clinical test report as described in any one of claims 1 to 7 is implemented.