Medical record picture OCR (Optical Character Recognition) method and device, computer equipment and medium

By using the semantic attention model to segment and rotate the medical record picture, the problems of complex background interference and rotation picture processing are solved, and high-accurate medical record picture OCR recognition is achieved.

CN119942512APending Publication Date: 2025-05-06北京大学长沙计算与数字经济研究院 +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510035572.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

When using OCR technology for medical picture recognition, the prior art faces the problems of complex image background interference and limited processing capabilities of rotating image, resulting in reduced recognition accuracy and confusion in text information.

Method used

The initial medical record picture is segmented using a semantic attention model (SAM) based on preset propt, the position of the text box is recognized, and the picture is rotated according to the position of the text box is finally OCR recognition of the rotated picture.

Benefits of technology

Through segmentation and rotation processing, complex background interference is effectively removed, the accuracy of OCR recognition and the clarity of text information are improved, and accurate OCR recognition of medical record pictures is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942512A_ABST
    Figure CN119942512A_ABST
Patent Text Reader

Abstract

The invention relates to a medical record picture OCR recognition method and device, computer equipment and a medium, and the method comprises the steps: obtaining an initial medical record picture; segmenting the initial medical record picture based on an SAM model of preset prompt to obtain segmented medical record pictures; identifying a textbox position in the segmented medical record picture; performing picture rotation on the segmented medical record picture according to the textbox position to obtain a rotated medical record picture; and performing OCR identification on the rotated medical record picture. In the whole process, the initial medical record picture is segmented based on the SAM model of the preset prompt, irrelevant noise images are removed, picture rotation is carried out according to the position of the textbox, content disorder is avoided, and accurate OCR recognition of the medical record picture can be achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of OCR recognition technology, and in particular to a method, device, computer equipment, storage medium and computer program product for OCR recognition of medical record images. Background Art

[0002] In the current field of medical image processing, OCR (Optical Character Recognition) technology plays a vital role, especially in the extraction of text information from medical images. Medical images, such as test reports, prescriptions, and medical records, usually contain a large amount of key text information, which is crucial for doctors' diagnosis, patient treatment records, and subsequent medical research. However, existing technologies face many challenges when applying OCR technology to medical image recognition.

[0003] A significant problem is the interference of complex image background on OCR recognition accuracy. In practical applications, medical images are often accompanied by complex and changeable backgrounds, such as noise, lines, graphic marks, etc. These background elements not only distract the attention of the OCR system, but may also cause the system to mistakenly recognize information in the background that is not related to the text as text, thereby mixing it into the medical record text information that should be clear and accurate. This misrecognition phenomenon not only reduces the quality of text information extraction, but also may cause confusion and misunderstanding in the interpretation of medical information, posing a potential threat to the accuracy and efficiency of medical work. In addition, medical images often rotate during the collection and transmission process due to different equipment, angles or shooting conditions. Existing OCR technology has limited processing capabilities for rotated images and often cannot accurately recognize and adapt to the rotation changes of image content. When the image is rotated, the originally orderly text information may become disorganized, resulting in a larger edit distance in the OCR system during the recognition process, that is, the degree of difference between the recognition result and the original text increases significantly. This not only increases the difficulty of subsequent text correction and processing, but also further reduces the accuracy of OCR technology in medical image recognition.

[0004] It can be seen that accurate OCR recognition of medical record images cannot be achieved using traditional technologies. Summary of the invention

[0005] Based on this, it is necessary to provide an accurate medical record image OCR recognition method, device, computer equipment, computer readable storage medium and computer program product to address the above technical problems.

[0006] In a first aspect, the present application provides a method for OCR recognition of medical record images. The method comprises:

[0007] Obtain initial medical record images;

[0008] Segmenting the initial medical record image based on a preset prompt SAM model (Semantic Attention Model) to obtain a segmented medical record image;

[0009] Identifying the location of the text box in the segmented medical record image;

[0010] Rotate the segmented medical record image according to the position of the text box to obtain a rotated medical record image;

[0011] Perform OCR recognition on the rotated medical record image.

[0012] In one embodiment, segmenting the initial medical record image based on the SAM model of the preset prompt to obtain the segmented medical record image includes:

[0013] Set the preset prompt in the SAM model to the center of the picture;

[0014] The initial medical record image is segmented by a SAM model based on the center position of the image to obtain a segmented medical record image.

[0015] In one embodiment, performing SAM model segmentation on the initial medical record image based on the center position of the image to obtain the segmented medical record image includes:

[0016] Based on the SAM model, identify and segment the pixel clusters of the same type of objects as the center of the image;

[0017] A rectangle is drawn with the outermost coordinates of the pixel cluster to obtain a segmented medical record image.

[0018] In one embodiment, rotating the segmented medical record image according to the position of the text box to obtain the rotated medical record image includes:

[0019] Calculating multiple text box slopes according to the positions of the multiple text boxes;

[0020] Calculate the rotation angle based on the slopes of the multiple text boxes;

[0021] Rotate the segmented medical record image according to the rotation angle to obtain a rectangular medical record image flush with the horizontal plane;

[0022] Rotating the rectangular medical record image in different directions, and performing content recognition on the rectangular medical record images rotated in different directions to obtain multiple content recognition results;

[0023] Performing a text content quality score based on each of the content recognition results, and selecting an optimal rotation direction corresponding to the best text content quality score;

[0024] The rectangular medical record image under the optimal rotation direction is extracted to obtain the rotated medical record image.

[0025] In one embodiment, calculating the rotation angle based on the slopes of the multiple text boxes includes:

[0026] For each of the text boxes, a triangle is formed according to the slope of the text box and the coordinates of the upper left corner and the upper right corner of the text box;

[0027] Based on the triangle, the inverse trigonometric value of the tilt angle of each text box is calculated to obtain the rotation degree corresponding to different text boxes;

[0028] The average value of the rotation degrees corresponding to the different text boxes is calculated to obtain the rotation angle.

[0029] In one embodiment, the rectangular medical record image is rotated in different directions, and content recognition is performed on the rectangular medical record image rotated in different directions, and multiple content recognition results are obtained, including:

[0030] Rotate the rectangular medical record image by 90°, 180° and 270° respectively;

[0031] Content recognition is performed when the rectangular medical record image is rotated by 0°, 90°, 180° and 270° to obtain multiple content recognition results.

[0032] In one embodiment, the performing of text content quality scoring based on each of the content recognition results and selecting the optimal rotation direction corresponding to the optimal text content quality score comprises:

[0033] Based on the content recognition results, extract the content recognition results within a preset area close to a preset edge in the rectangular medical record image when rotated by 0°, 90°, 180° and 270°, where the preset edge is any rectangular edge in the rectangular medical record image;

[0034] Recording the number of preset keywords in the extracted content recognition results;

[0035] Score the text content quality of each of the content recognition results according to the number of the preset keywords;

[0036] Select the optimal rotation direction corresponding to the best text content quality score.

[0037] In a second aspect, the present application also provides a medical record image OCR recognition device. The device comprises:

[0038] An initial image acquisition module is used to acquire initial medical record images;

[0039] A segmentation module, used for segmenting the initial medical record image based on a SAM model of a preset prompt to obtain a segmented medical record image;

[0040] A text box recognition module, used to recognize the position of the text box in the segmented medical record image;

[0041] A rotation module, used to rotate the segmented medical record image according to the position of the text box to obtain a rotated medical record image;

[0042] The OCR recognition module is used to perform OCR recognition on the rotated medical record image.

[0043] In a third aspect, the present application further provides a computer device. The computer device includes a memory and a processor, the memory stores a computer program, and the processor implements the following steps when executing the computer program:

[0044] Obtain initial medical record images;

[0045] Segmenting the initial medical record image based on the SAM model of the preset prompt to obtain a segmented medical record image;

[0046] Identifying the location of the text box in the segmented medical record image;

[0047] Rotate the segmented medical record image according to the position of the text box to obtain a rotated medical record image;

[0048] Perform OCR recognition on the rotated medical record image.

[0049] In a fourth aspect, the present application further provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the following steps are implemented:

[0050] Obtain initial medical record images;

[0051] Segmenting the initial medical record image based on the SAM model of the preset prompt to obtain a segmented medical record image;

[0052] Identifying the location of the text box in the segmented medical record image;

[0053] Rotate the segmented medical record image according to the position of the text box to obtain a rotated medical record image;

[0054] Perform OCR recognition on the rotated medical record image.

[0055] In a fifth aspect, the present application further provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the following steps are implemented:

[0056] Obtain initial medical record images;

[0057] Segmenting the initial medical record image based on the SAM model of the preset prompt to obtain a segmented medical record image;

[0058] Identifying the location of the text box in the segmented medical record image;

[0059] Rotate the segmented medical record image according to the position of the text box to obtain a rotated medical record image;

[0060] Perform OCR recognition on the rotated medical record image.

[0061] The above-mentioned medical record image OCR recognition method, device, computer equipment, storage medium and computer program product obtain an initial medical record image; segment the initial medical record image based on the SAM model of the preset prompt to obtain a segmented medical record image; identify the position of the text box in the segmented medical record image; rotate the segmented medical record image according to the position of the text box to obtain a rotated medical record image; and perform OCR recognition on the rotated medical record image. During the whole process, the initial medical record image is segmented based on the SAM model of the preset prompt, irrelevant noise images are removed, and the image is rotated according to the position of the text box to avoid disorder of the content, so that accurate OCR recognition of medical record images can be achieved. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] Figure 1 A diagram showing an application environment of a medical record image OCR recognition method in one embodiment;

[0063] Figure 2 A schematic diagram of a process of an OCR recognition method for medical record images in one embodiment;

[0064] Figure 3 A schematic diagram of a flow chart of a medical record image OCR recognition method in another embodiment;

[0065] Figure 4 is a schematic diagram of a sub-process of S400 in one embodiment;

[0066] Figure 5 is a structural block diagram of a medical record image OCR recognition device in one embodiment;

[0067] Figure 6 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0068] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0069] The medical record image OCR recognition method provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or it can be placed on the cloud or other network servers. The terminal 102 sends a medical record image OCR recognition request to the server 104, and the server 104 responds to the request and extracts the initial medical record image carried in the request; the initial medical record image is segmented based on the SAM model of the preset prompt to obtain a segmented medical record image; the text box position in the segmented medical record image is identified; the segmented medical record image is rotated according to the text box position to obtain a rotated medical record image; and the rotated medical record image is OCR recognized. Further, the server 104 can feed back the obtained medical record image OCR recognition result to the terminal 102. Among them, the terminal 102 can be, but is not limited to, various personal computers, laptops, smart phones, tablet computers, Internet of Things devices and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart car-mounted devices, etc. Portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server 104 may be implemented as an independent server or a server cluster consisting of multiple servers.

[0070] In one embodiment, Figure 2 As shown, a medical record image OCR recognition method is provided, and the method is applied to Figure 1 Taking the server 104 in the example as an example, the following steps are included:

[0071] S100: Obtaining an initial medical record image.

[0072] The initial medical record images can come from the hospital's paper medical records, electronic medical record system or medical record images provided by the patient. These images may contain text, charts, images and other information, and the text direction, size and background may vary.

[0073] S200: Segment the initial medical record image based on the SAM model of the preset prompt to obtain a segmented medical record image.

[0074] This embodiment uses the SAM model to segment medical record images. The SAM model is a semantic segmentation model based on deep learning, which can perform fine segmentation on images according to the input preset prompt. Before segmentation, it is necessary to set a preset prompt to guide the SAM model to perform correct segmentation. The preset prompt can be customized according to the characteristics of the medical record image. The initial medical record image is input into the SAM model and segmented according to the preset prompt. The segmented medical record image will contain multiple independent text areas, which will be used for subsequent text box position recognition and image rotation processing.

[0075] S300: Identify the position of the text box in the segmented medical record image.

[0076] In the segmented medical record image, the position of the area containing the text (i.e., the text box) is identified using image processing techniques (such as edge detection, contour extraction, etc.). This position information will be used for subsequent image rotation processing. The identified text box position information is recorded for use in subsequent processing. This position information can be expressed as coordinates, rectangular boxes, or other forms of geometric shapes. Preferably, it can be a rectangular box.

[0077] S400: Rotate the segmented medical record image according to the position of the text box to obtain a rotated medical record image.

[0078] Based on the identified text box position information, calculate the angle that the medical record image needs to be rotated. This angle can be the angle required to make the text box have a uniform direction (such as horizontal or vertical direction). Use image processing software or libraries (such as OpenCV) to rotate the medical record image. The rotated medical record image will have a uniform text direction, which is convenient for subsequent OCR recognition processing.

[0079] S500: Perform OCR recognition on the rotated medical record image.

[0080] Select a suitable OCR recognition engine (such as Tesseract, Baidu OCR, etc.) to recognize and process the rotated medical record image. Output the OCR recognition results as editable and searchable digital text. These texts can be further used for medical record review, analysis, statistics and other applications.

[0081] The above-mentioned medical record image OCR recognition method obtains an initial medical record image; segments the initial medical record image based on the SAM model of the preset prompt to obtain a segmented medical record image; identifies the position of the text box in the segmented medical record image; rotates the segmented medical record image according to the position of the text box to obtain a rotated medical record image; and performs OCR recognition on the rotated medical record image. During the entire process, the initial medical record image is segmented based on the SAM model of the preset prompt, irrelevant noise images are removed, and the image is rotated according to the position of the text box to avoid disordered content, thereby achieving accurate OCR recognition of medical record images.

[0082] In one embodiment, if Figure 3 As shown, S300 includes:

[0083] S320: Set the preset prompt in the SAM model to the center position of the image.

[0084] In this application, in order to solve the problem of complex background, a SAM model is added to segment the irrelevant content of the medical record text. Furthermore, as a segmentation model, SAM needs to set the prompt in advance so that the model can segment according to the selected objects. For medical record picture scene text, the medical record pictures are often centered and occupy more than 60% of the page. Based on this feature, in this application, it is assumed that the center pixel of the picture must be the medical record picture paper, and the prompt point is set to the center of the picture by default.

[0085] S340: Perform SAM model segmentation on the initial medical record image based on the center position of the image to obtain a segmented medical record image.

[0086] After determining that the center of the image is the prompt point, the objects of the same type as the center of the image are segmented to obtain all the pixels in the image that are considered by the model to be paper. Therefore, the initial medical record image can be segmented by the SAM model based on the center of the image to obtain a segmented medical record image.

[0087] In one embodiment, the initial medical record image is segmented by a SAM model based on the center position of the image, and the segmented medical record image includes:

[0088] Based on the SAM model, pixel clusters of the same type of objects as the center of the image are identified and segmented; a rectangle is drawn with the outermost coordinates of the pixel cluster to obtain a segmented medical record image.

[0089] After determining that the center of the image is the prompt point, the objects of the same type as the center of the image are segmented to obtain all the pixels in the image that the model considers to be paper. A rectangle is then drawn using the outermost coordinates of the pixel cluster to automatically segment the medical record paper image, effectively removing information interference caused by non-paper content.

[0090] In one embodiment, if Figure 4 As shown, S400 includes:

[0091] S410: Calculate slopes of multiple text boxes according to the positions of the multiple text boxes.

[0092] First, the position information of multiple text boxes in the medical record image is identified. Using this position information (for example, the coordinates of the four vertices of the text box), the slope of each text box is calculated. The slope can reflect the degree of inclination of the text box relative to the horizontal direction.

[0093] S420: Calculate a rotation angle based on the slopes of the multiple text boxes.

[0094] Based on the calculated slopes of multiple text boxes, an overall rotation angle is determined by statistical analysis or average calculation, etc. This angle is intended to adjust most or all text boxes to a state close to horizontal or vertical.

[0095] S430: Rotate the segmented medical record image according to the rotation angle to obtain a rectangular medical record image flush with the horizontal plane.

[0096] The segmented medical record image is rotated using the calculated rotation angle, with the goal of making the text area in the medical record image as flush with the horizontal plane as possible, forming a rectangular medical record image.

[0097] S440: Rotate the rectangular medical record image in different directions, and perform content recognition on the rectangular medical record images rotated in different directions to obtain multiple content recognition results.

[0098] In order to further ensure the best content quality of the rotated medical record image, the rectangular medical record image obtained above is rotated in different directions (for example, 0°, 90°, 180°, 270°). Content recognition, such as OCR (optical character recognition), is performed on the medical record image rotated in each direction to obtain multiple content recognition results.

[0099] S450: Score the text content quality based on each content recognition result, and select the optimal rotation direction corresponding to the text content quality score with the best quality.

[0100] Based on the content recognition results, the text content quality of the medical record images rotated in each direction is scored. The scoring can be based on multiple dimensions such as recognition accuracy, text completeness, format clarity, etc. According to the text content quality score, the rotation direction corresponding to the best score is selected as the optimal rotation direction. The medical record images in this direction have the best text content quality.

[0101] S460: Extract the rectangular medical record image in the optimal rotation direction to obtain the rotated medical record image.

[0102] Finally, the rectangular medical record image in the optimal rotation direction is extracted as the final rotated medical record image.

[0103] In one embodiment, calculating the rotation angle based on the slopes of multiple text boxes includes:

[0104] Step 1: For each text box, construct a triangle based on the slope of the text box and the coordinates of the upper left corner and the upper right corner of the text box.

[0105] For each identified text box, a right triangle is constructed using the coordinates of its upper left and upper right corners, and the third side (hypotenuse) implied by these two points and the slope of the text box. The hypotenuse here is not drawn directly, but exists conceptually through the slope, that is, a virtual line connecting the upper left and upper right corners.

[0106] Step 2: Based on the triangle, calculate the inverse trigonometric value of the tilt angle of each text box to obtain the rotation degree corresponding to different text boxes.

[0107] Using the above triangle, we can calculate the tilt angle of each text box. This is usually done by using the inverse tangent function (arctan), which accepts a slope value as input and outputs an angle value that represents the angle with the horizontal line. Therefore, for each text box, we calculate the inverse trigonometric value of the tilt angle based on its slope, that is, how many degrees the text box needs to be rotated to make it flush with the horizontal plane.

[0108] Step 3: Calculate the average rotation degree corresponding to different text boxes to get the rotation angle.

[0109] Now that we have the rotation of each text box, we need to calculate an overall rotation angle. This is usually done by taking the average of all the text box rotations. The average calculation ensures that the rotation angle is as close to the corrective needs of all text boxes as possible, even though some text boxes may be tilted differently than others. By doing this, we get a rotation angle that is required to rotate the medical record image to make most text boxes close to horizontal.

[0110] In one embodiment, the rectangular medical record image is rotated in different directions, and content recognition is performed on the rectangular medical record images rotated in different directions, and multiple content recognition results are obtained, including:

[0111] Step 1: Rotate the rectangular medical record image by 90°, 180°, and 270° respectively.

[0112] First, the rectangular medical record image is rotated by 0° (i.e., not rotated, as a reference), 90°, 180°, and 270°. These four directions are usually sufficient to cover all possible text directions, because text can only appear in one of these four basic directions on a plane (regardless of the tilt angle).

[0113] Step 2: Perform content recognition when the rectangular medical record image is rotated by 0°, 90°, 180°, and 270° to obtain multiple content recognition results.

[0114] In each rotation direction, the medical record image is subjected to content recognition. This is usually achieved through OCR (Optical Character Recognition) technology, which can recognize text in the image and convert it into an editable text format. The purpose of content recognition is to obtain text information in the medical record image for subsequent analysis and processing. By rotating the medical record image in different directions and performing content recognition, four content recognition results will be obtained. Each result corresponds to a specific rotation direction and contains the text information recognized in that direction. Next, these content recognition results can be further processed and analyzed. For example, the recognition results in different directions can be compared to determine which direction has the clearest and most accurate text information.

[0115] In one embodiment, the text content quality is scored based on each content recognition result, and the optimal rotation direction corresponding to the best text content quality score is selected, including:

[0116] Step 1: Based on each content recognition result, extract the content recognition results within a preset area close to a preset edge in the rectangular medical record image when rotated 0°, 90°, 180° and 270°, where the preset edge is any rectangular edge in the rectangular medical record image.

[0117] First, for the content recognition results of each rotation direction (0°, 90°, 180°, 270°), extract the text content within the preset area close to any rectangular edge in the rectangular medical record image. This preset area can be set according to the actual situation. For example, it can be an area with a certain pixel width away from the edge. The purpose is to focus on the text near the edge that may contain important information (such as title, date, patient information, etc.). Specifically, the preset area close to the preset edge can be 30% of the upper edge of the image.

[0118] Step 2: Record the number of preset keywords in the extracted content recognition results.

[0119] In the extracted content recognition results, search and record the number of preset keywords. These preset keywords can be words closely related to the medical record content, such as "name", "age", "diagnosis", "drug", etc. The number of keywords can be used as an important indicator to measure the quality of text content.

[0120] Step 3: Score the text content quality of each content recognition result based on the number of preset keywords.

[0121] The text content quality score is calculated for the content recognition results of each rotation direction according to the number of preset keywords. The score can be linearly or nonlinearly weighted based on the number of keywords, or combined with other factors (such as text clarity, recognition accuracy, etc.) for comprehensive scoring.

[0122] Step 4: Select the optimal rotation direction corresponding to the best text content quality score.

[0123] Finally, the rotation direction with the best text content quality score is selected as the optimal rotation direction. The medical record image in this direction should have the best text content quality, that is, it should contain the largest number of preset keywords, and the text should be clear and accurately recognized.

[0124] In practical applications, the improvement of the accuracy of rotated image recognition includes the following: For rotated images, most open source tools on the market can only classify images into four categories, namely 0°, 90°, 180°, and 270°. After adding the preset rotated image recognition algorithm, it can process images rotated at all angles of 0-360°, and can correct rotated images at all angles. Using the open source framework of paddleOCR, the positions of all detected text boxes can be obtained. Based on this, a new rotated image recognition algorithm is proposed in this application (specifically as Figure 4As shown). After obtaining the coordinate points of all text boxes, the coordinate points of the first five text boxes are extracted to obtain five rectangular boxes. The slope of the top edge of each rectangular box is calculated in turn. According to the obtained slope and the coordinate points of the upper left and upper right corners of the rectangular box, a triangle is constructed to calculate the inverse trigonometric value of the text box tilt angle, and then the specific degree of rotation is obtained. After calculating the rotation angle of the first five text boxes, the average rotation angle of these five boxes is calculated. Finally, the image is rotated according to this average rotation angle. The rotated medical record image will be a rectangle flush with the horizontal plane, but it may also need to be rotated 90°, 180° or 270° to straighten the text. The method adopted in this scheme is to perform content recognition on the medical record when it is rotated 0°, 90°, 180° and 270°. The final recognition result is scored according to the text quality. The basis for scoring the text quality is to judge whether there are some medical keywords in 30% of the image. If there is one, one point is recorded. These keywords include: name, gender, age, chief complaint, current medical history, family history, allergy history, diagnosis, examination, etc. Finally, according to the scoring, the best result with the highest score among the four text recognition results is selected.

[0125] It should be understood that, although the steps in the flowcharts involved in the above embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps is not strictly limited in order, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above embodiments may include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.

[0126] Based on the same inventive concept, the embodiment of the present application also provides a medical record image OCR recognition device for implementing the medical record image OCR recognition method involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in one or more medical record image OCR recognition device embodiments provided below can refer to the limitations of the medical record image OCR recognition method above, and will not be repeated here.

[0127] In one embodiment, Figure 5 As shown, a medical record image OCR recognition device is provided, comprising:

[0128] An initial picture acquisition module 100 is used to acquire an initial medical record picture;

[0129] A segmentation module 200 is used to segment the initial medical record image based on a SAM model of a preset prompt to obtain a segmented medical record image;

[0130] A text box identification module 300 is used to identify the location of the text box in the segmented medical record image;

[0131] A rotation module 400 is used to rotate the segmented medical record image according to the position of the text box to obtain a rotated medical record image;

[0132] The OCR recognition module 500 is used to perform OCR recognition on the rotated medical record image.

[0133] In one of the embodiments, the segmentation module 200 is further used to set the preset prompt in the SAM model to the center position of the picture; perform SAM model segmentation on the initial medical record picture based on the center position of the picture to obtain a segmented medical record picture.

[0134] In one embodiment, the segmentation module 200 is further used to identify and segment a pixel cluster of the same type of object as the center of the image based on the SAM model; draw a rectangle with the outermost coordinates of the pixel cluster to obtain a segmented medical record image.

[0135] In one embodiment, the rotation module 400 is also used to calculate the slopes of multiple text boxes according to the positions of multiple text boxes; obtain a rotation angle based on the calculation of the slopes of multiple text boxes; rotate the segmented medical record image according to the rotation angle to obtain a rectangular medical record image flush with the horizontal plane; rotate the rectangular medical record image in different directions, and perform content recognition on the rectangular medical record images rotated in different directions to obtain multiple content recognition results; perform text content quality scoring based on each content recognition result, and select the optimal rotation direction corresponding to the best text content quality score; extract the rectangular medical record image under the optimal rotation direction to obtain a rotated medical record image.

[0136] In one embodiment, the rotation module 400 is also used to form a triangle for each text box according to the slope of the text box and the coordinates of the upper left corner and the upper right corner of the text box; based on the triangle, calculate the inverse trigonometric value of the inclination angle of each text box to obtain the rotation degree corresponding to different text boxes; calculate the average value of the rotation degrees corresponding to different text boxes to obtain the rotation angle.

[0137] In one of the embodiments, the rotation module 400 is further used to rotate the rectangular medical record image by 90°, 180° and 270° respectively; and to perform content recognition when the rectangular medical record image is rotated by 0°, 90°, 180° and 270° to obtain multiple content recognition results.

[0138] In one of the embodiments, the rotation module 400 is also used to extract content recognition results within a preset area near a preset edge in the rectangular medical record image when rotated 0°, 90°, 180° and 270° based on each content recognition result, where the preset edge is any rectangular edge in the rectangular medical record image; record the number of preset keywords in the extracted content recognition results; perform a text content quality score on each content recognition result according to the number of preset keywords; and select the optimal rotation direction corresponding to the optimal text content quality score.

[0139] Each module in the above-mentioned medical record image OCR recognition device can be implemented in whole or in part by software, hardware and their combination. Each of the above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the corresponding operations of each of the above modules.

[0140] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 6 As shown. The computer device includes a processor, a memory and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store preset data. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a medical record image OCR recognition method is implemented.

[0141] Those skilled in the art will understand that Figure 6 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0142] In one embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the above-mentioned medical record image OCR recognition method when executing the computer program.

[0143] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the above-mentioned medical record image OCR recognition method is implemented.

[0144] In one embodiment, a computer program product is provided, including a computer program, which implements the above-mentioned medical record image OCR recognition method when executed by a processor.

[0145] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but are not limited to this.

[0146] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0147] The above embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.

Claims

1. A medical record image OCR recognition method, characterized in that: The method comprises: Obtain initial medical record images; Segmenting the initial medical record image based on the SAM model of the preset prompt to obtain a segmented medical record image; Identifying the location of the text box in the segmented medical record image; Rotate the segmented medical record image according to the position of the text box to obtain a rotated medical record image; Perform OCR recognition on the rotated medical record image.

2. The method according to claim 1, characterized in that The initial medical record image is segmented based on the SAM model of the preset prompt to obtain the segmented medical record image, including: Set the preset prompt in the SAM model to the center of the picture; The initial medical record image is segmented by a SAM model based on the center position of the image to obtain a segmented medical record image.

3. The method according to claim 2, characterized in that The performing SAM model segmentation on the initial medical record image based on the center position of the image to obtain the segmented medical record image includes: Based on the SAM model, identify and segment the pixel clusters of the same type of objects as the center of the image; A rectangle is drawn with the outermost coordinates of the pixel cluster to obtain a segmented medical record image.

4. The method according to claim 1, characterized in that: The step of rotating the segmented medical record image according to the position of the text box to obtain a rotated medical record image comprises: Calculating multiple text box slopes according to the positions of the multiple text boxes; Calculate the rotation angle based on the slopes of the multiple text boxes; Rotate the segmented medical record image according to the rotation angle to obtain a rectangular medical record image flush with the horizontal plane; Rotating the rectangular medical record image in different directions, and performing content recognition on the rectangular medical record images rotated in different directions to obtain multiple content recognition results; Performing a text content quality score based on each of the content recognition results, and selecting an optimal rotation direction corresponding to the best text content quality score; The rectangular medical record image under the optimal rotation direction is extracted to obtain the rotated medical record image.

5. The method according to claim 4, characterized in that The calculating the rotation angle based on the slopes of the multiple text boxes comprises: For each of the text boxes, a triangle is formed according to the slope of the text box and the coordinates of the upper left corner and the upper right corner of the text box; Based on the triangle, the inverse trigonometric value of the tilt angle of each text box is calculated respectively to obtain the rotation degree corresponding to different text boxes; The average value of the rotation degrees corresponding to the different text boxes is calculated to obtain the rotation angle.

6. The method according to claim 4, characterized in that The step of rotating the rectangular medical record image in different directions and performing content recognition on the rectangular medical record image rotated in different directions to obtain multiple content recognition results includes: Rotate the rectangular medical record image by 90°, 180° and 270° respectively; Content recognition is performed when the rectangular medical record image is rotated by 0°, 90°, 180° and 270° to obtain multiple content recognition results.

7. The method according to claim 4, characterized in that The text content quality scoring is performed based on each of the content recognition results, and the optimal rotation direction corresponding to the optimal text content quality score is selected, which includes: Based on the content recognition results, extract the content recognition results within a preset area close to a preset edge in the rectangular medical record image when rotated by 0°, 90°, 180° and 270°, where the preset edge is any rectangular edge in the rectangular medical record image; Recording the number of preset keywords in the extracted content recognition results; Score the text content quality of each of the content recognition results according to the number of the preset keywords; Select the optimal rotation direction corresponding to the best text content quality score.

8. A medical record image OCR recognition device, characterized in that: The device comprises: An initial image acquisition module is used to acquire initial medical record images; A segmentation module, used for segmenting the initial medical record image based on a SAM model of a preset prompt to obtain a segmented medical record image; A text box recognition module, used to recognize the position of the text box in the segmented medical record image; A rotation module, used to rotate the segmented medical record image according to the position of the text box to obtain a rotated medical record image; The OCR recognition module is used to perform OCR recognition on the rotated medical record image.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.