Method, device and equipment for recognizing image text in test sheet and storage medium

CN116453149BActive Publication Date: 2026-08-07PING AN TECH (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
PING AN TECH (SHENZHEN) CO LTD
Filing Date
2023-04-18
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0004]有鉴于此,本发明提供了一种检验单中图像文本的识别方法、装置及设备,主要目的在于解决现有技术在校验单中图像文本的识别步骤中每个环节都是串联的方式,容易引入级联错误预测,影响检验单中图像文本的识别结果的问题

Benefits of technology

[0050] By means of the above technical solution, the present invention provides a method, apparatus, device and storage medium for recognizing text in an image of a test form. By acquiring the positional information of different regions in the test form image, the text information in different positional regions of the test form image is converted into text feature vectors. A first network model trained using an object detection algorithm encodes the test form image into visual feature vectors. Each visual feature vector represents a visual feature in a preset image region in the test form image. After combining the text feature vectors and visual feature vectors into multimodal sequence features, the features representing entities in the test form image are obtained through linear network mapping. Using the features representing entities in the test form image, the text and text correspondence in the test form image are recognized, and the text array contained in the test form is obtained. Compared with the existing technology that uses a pipeline approach to extract structured text from test images, this application treats the structured extraction of test images as an extraction task of entities and relationships from multimodal information. On the one hand, it enriches the text features through multimodal features, avoids the introduction of cascaded error prediction, and improves the recognition results of text in test images. On the other hand, by extracting joint information, it combines the correspondence between text in the image on the basis of the test image structure, and can accurately obtain the array relationship of text in the test image without adapting to different test image templates, thus improving the recognition efficiency of text in test images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116453149B_ABST
    Figure CN116453149B_ABST
Patent Text Reader

Abstract

The present application relates to the field of digital medical treatment, and discloses a kind of identification method of image text in test sheet, comprising: obtaining the position information of different regions in test sheet image, convert the text information in different position regions in check sheet image into text feature vector, using the first network model trained by target detection algorithm, test sheet image is encoded into visual feature vector, each visual feature vector represents the visual feature in the preset image region in test sheet image, after text feature vector and visual feature vector are combined into multimodal sequence feature, linear network mapping is carried out to obtain the feature representing entity in test sheet image, using the feature representing entity in test sheet image, the text in test sheet image and the corresponding relationship of text are identified, and the text array contained in test sheet is obtained.The present application enriches text feature by the mode of multimodal feature, avoids introducing cascade error prediction, and improves the identification result of text in test sheet image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of digital healthcare, and in particular to methods, apparatus, devices, and storage media for recognizing image text in test reports. Background Technology

[0002] With the development of technology, images play a significant role in information dissemination. To better serve their promotional purposes, text is increasingly being incorporated into images. For example, in the medical field, medical institutions need to upload images of test results to their systems so that they can view the results based on the content of these images. Therefore, since text within images often contains rich information, structured extraction and recognition of text from test result images is of great significance for image content analysis, understanding, and information retrieval.

[0003] Typically, test reports are presented to users in tabular form on a single sheet of paper. Currently, there are two main methods for structured text extraction from test report images. One method involves a streamlined process: first, layout recognition is performed to identify the table and main text sections. Then, image segmentation is applied to the table section of the test report image. The purpose of segmentation is to annotate the table lines. Next, geometric analysis is performed on the segmented image to extract connected components, thereby reconstructing the main framework structure of the table. Further, text recognition technology is used to extract the main text and table text. Finally, based on keywords, the location and relationships of keywords are identified from the integrated information to achieve text extraction. One approach is structured text extraction. However, considering that each step in the workflow is sequential, cascading error prediction is easily introduced. If a problem occurs in any step, it will affect the recognition result of the text in the test report image. Another approach is to predefine a test report template and extract the text in the region of interest based on the text recognition result, thus achieving structured text extraction. However, since different diseases correspond to different test reports, and the same examination item corresponds to different test report templates, this text recognition method can only recognize predefined test report templates and cannot adapt to different test report templates, resulting in low recognition efficiency of text in test report images. Summary of the Invention

[0004] In view of this, the present invention provides a method, apparatus and device for recognizing image text in a verification form. The main purpose is to solve the problem that in the prior art, each step in the recognition process of image text in a verification form is serial, which easily introduces cascading error prediction and affects the recognition result of image text in the verification form.

[0005] According to one aspect of the present invention, a method for recognizing text in an image is provided, comprising:

[0006] Obtain the location information of different regions in the single image of the test, and convert the text information in the different location regions of the single image of the test into text feature vectors;

[0007] The first network model trained using the object detection algorithm encodes the test single image into a visual feature vector, where each visual feature vector represents a visual feature within a preset image region in the test single image.

[0008] After combining the text feature vector and the visual feature vector into a multimodal sequence feature, the feature representing the entity in the test single image is obtained through linear network mapping. Using the feature representing the entity in the test single image, the text and text correspondence in the test single image are identified to obtain the text array contained in the test single.

[0009] Further, the step of obtaining the location information of different regions in the verification single image and converting the text information in different location regions of the verification single image into text feature vectors includes:

[0010] Acquire an inspection form image, perform layout analysis on the inspection form image, and obtain the position information of different regions in the inspection form image;

[0011] Using the positional information of different regions in the verification single image, the text information in different positional regions of the verification single image is converted into text feature vectors;

[0012] Further, the step of acquiring the inspection form image and performing layout analysis on the inspection form image to obtain the positional information of different regions in the inspection form image includes:

[0013] A test form image is acquired, and a second network model trained using an object detection algorithm is used to perform layout analysis on the test form image to obtain detection boxes corresponding to different regions in the test form image.

[0014] The location information of different regions in the test form image is determined based on the detection boxes corresponding to different regions in the test form image.

[0015] Furthermore, the step of using the positional information of different regions in the single verification image to convert the text information within different positional regions of the single verification image into text feature vectors includes:

[0016] Based on the positional information of different regions in the test image, text recognition technology is used to extract text information in different regions of the test image and the coordinate layout information of the text information in the test image.

[0017] Using the coordinate layout information of the text information in the verification single image, the text information in different position areas of the verification single image is serialized according to a set method, and then the text and position coordinates in the text information are vectorized and embedded using an embedded vector model to obtain text feature vectors.

[0018] Further, the step of combining the text feature vector and the visual feature vector into a multimodal sequence feature, and then mapping it through a linear network to obtain the features representing entities in a single image, includes:

[0019] After combining the text feature vector and the visual feature vector into a multimodal sequence feature, they are respectively input into two linear networks for mapping to obtain the features representing entities in a single image.

[0020] The features representing entities include start features and end features. The process involves using these entity features in the inspection form image to identify the text and text correspondences within the inspection form image, resulting in a text array contained in the inspection form, including:

[0021] The start and end features representing entities are transformed using a dual affine attention mechanism to obtain the matrix vector representing the start-to-end pairs of entities in the verification single image.

[0022] Identify the index relationship between entities in the verification single image in the matrix vector representing the start-to-end pairs of entities;

[0023] By utilizing the index relationship between entities in the verification form image, the text and text correspondence in the verification form image are obtained, resulting in a text array contained in the verification form.

[0024] Further, identifying the index relationship between entities in the verification single image within the matrix vector representing the start-to-end pairs of entities includes:

[0025] Different entity category labels are predefined. In the matrix vector representing the start-to-end pairs of entities, the position categories representing different entity category labels are determined. The position categories include the start position relationship category and the end position relationship category.

[0026] Based on the start position relationship category and end position relationship category representing different entity category labels, the index relationship between entities in the verification sheet image is identified.

[0027] Further, by utilizing the index relationships between entities in the verification form image, the text and text correspondence relationships in the verification form image are obtained, resulting in a text array contained in the verification form, including:

[0028] By utilizing the index relationship between entities in the verification single image, the entity category labels and the start and end index positions of all entities in the verification single image are obtained;

[0029] For each entity, check the start and end index positions to determine if there is a corresponding entity category label at each of the entity's start and end index positions.

[0030] If so, then based on the start and end index positions of all entities in the inspection form image, obtain the text and text correspondence in the inspection form image, and obtain the text array contained in the inspection form.

[0031] According to another aspect of the present invention, an image text recognition device for verification forms is provided, comprising:

[0032] The acquisition unit is used to acquire the position information of different regions in the verification single image and convert the text information in different position regions in the verification single image into text feature vectors.

[0033] The encoding unit is used to encode the test single image into a visual feature vector using a first network model trained with an object detection algorithm, whereby each visual feature vector represents a visual feature within a preset image region in the test single image.

[0034] The mapping unit is used to combine the text feature vector and the visual feature vector into a multimodal sequence feature, and then map it through a linear network to obtain the features representing entities in a single image for verification.

[0035] The recognition unit is used to identify the text and text correspondence in the inspection form image by utilizing the features representing entities in the inspection form image, and to obtain the text array contained in the inspection form.

[0036] Furthermore, the acquisition unit includes:

[0037] The analysis module is used to acquire the inspection form image, perform layout analysis on the inspection form image, and obtain the position information of different regions in the inspection form image;

[0038] The conversion module is used to convert text information in different regions of the verification single image into text feature vectors by utilizing the positional information of different regions in the verification single image.

[0039] Furthermore, the analysis module is specifically used to acquire the inspection form image, perform layout analysis on the inspection form image using a second network model trained with an object detection algorithm, obtain detection boxes corresponding to different regions in the inspection form image, and determine the position information of different regions in the inspection form image based on the detection boxes corresponding to different regions in the inspection form image.

[0040] Furthermore, the conversion module is specifically used to extract text information and coordinate layout information of the text information in different regions of the verification single image based on the position information of different regions in the verification single image using text recognition technology; using the coordinate layout information of the text information in the verification single image, the text information in different regions of the verification single image is serialized according to a set method, and then the text and position coordinates in the text information are vectorized and embedded using an embedded vector model to obtain text feature vectors.

[0041] Furthermore, the mapping unit is specifically used to combine the text feature vector and the visual feature vector into a multimodal sequence feature, and then input them into two linear networks for mapping to obtain the features representing entities in a single image.

[0042] The features representing an entity include a start feature and an end feature, and the identification unit includes:

[0043] The radiating module is used to perform affine transformations on the start features and end features of the entity representation using a dual affine attention mechanism to obtain a matrix vector representing the start-to-end pairs of entities in the verification single image.

[0044] The recognition module is used to identify the index relationship between entities in the verification single image in the matrix vector representing the start-to-end pairs of entities;

[0045] The acquisition module is used to obtain the text and text correspondence in the verification form image by utilizing the index relationship between entities in the verification form image, and obtain the text array contained in the verification form.

[0046] Furthermore, the recognition module is specifically used to predefine different entity category labels, determine the position category representing different entity category labels in the matrix vector representing the start-to-end pairs of entities, the position category including the start position relationship category and the end position relationship category; and identify the index relationship between entities in the verification single image based on the start position relationship category and the end position relationship category representing different entity category labels.

[0047] Furthermore, the acquisition module is specifically used to obtain entity category labels and the start and end index positions of all entities in the verification form image by utilizing the index relationship between entities in the verification form image; to check the start and end index positions of all entities and determine whether there is a corresponding entity category label at each of the start and end index positions of the entities; if so, to obtain the text in the verification form image and the text correspondence based on the start and end index positions of all entities in the verification form image, thereby obtaining the text array contained in the verification form.

[0048] According to another aspect of the present invention, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of a method for recognizing image text in a verification form.

[0049] According to another aspect of the present invention, a computer storage medium is provided, on which a computer program is stored, wherein when the computer program is executed by a processor, the steps of a method for recognizing image text in a verification form are implemented.

[0050] By means of the above technical solution, the present invention provides a method, apparatus, device and storage medium for recognizing text in an image of a test form. By acquiring the positional information of different regions in the test form image, the text information in different positional regions of the test form image is converted into text feature vectors. A first network model trained using an object detection algorithm encodes the test form image into visual feature vectors. Each visual feature vector represents a visual feature in a preset image region in the test form image. After combining the text feature vectors and visual feature vectors into multimodal sequence features, the features representing entities in the test form image are obtained through linear network mapping. Using the features representing entities in the test form image, the text and text correspondence in the test form image are recognized, and the text array contained in the test form is obtained. Compared with the existing technology that uses a pipeline approach to extract structured text from test images, this application treats the structured extraction of test images as an extraction task of entities and relationships from multimodal information. On the one hand, it enriches the text features through multimodal features, avoids the introduction of cascaded error prediction, and improves the recognition results of text in test images. On the other hand, by extracting joint information, it combines the correspondence between text in the image on the basis of the test image structure, and can accurately obtain the array relationship of text in the test image without adapting to different test image templates, thus improving the recognition efficiency of text in test images. Attached Figure Description

[0051] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:

[0052] Figure 1 This is a schematic diagram of an application environment for a method for recognizing image text in a verification form according to an embodiment of the present invention;

[0053] Figure 2 This is a flowchart illustrating a method for recognizing image text in a verification form according to an embodiment of the present invention;

[0054] Figure 3 yes Figure 2 A schematic diagram of a specific implementation method for step S10;

[0055] Figure 4 yes Figure 2 A schematic diagram of a specific implementation of step S40;

[0056] Figure 5 This is another flowchart illustrating the method for recognizing image text in a verification form according to another embodiment of the present invention;

[0057] Figure 6 This is a schematic diagram of a device for recognizing image text in a verification form according to an embodiment of the present invention;

[0058] Figure 7 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention;

[0059] Figure 8 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation

[0060] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.

[0061] The image text recognition method in the inspection form provided in this embodiment of the invention can be applied to, for example... Figure 1In this application environment, the client communicates with the server via a network. The client receives a verification form image, and the server obtains the location information of different regions in the verification form image. The server converts the text information within these regions into text feature vectors. A first network model trained using an object detection algorithm encodes the verification form image into visual feature vectors. Each visual feature vector represents a visual feature within a preset image region in the verification form image. After combining the text feature vectors and visual feature vectors into a multimodal sequence feature, a linear network is used to map the features representing entities in the verification form image. Using these entity features, the server identifies the text and its correspondences within the verification form image, obtaining the text array contained in the verification form. In this invention, the structured extraction of a single image is viewed as an extraction task of entities and relationships from multimodal information. On one hand, multimodal features enrich text features, avoiding cascading error prediction and improving the recognition results of text in the single image. On the other hand, by extracting joint information, the corresponding relationships between text in the image are combined with the structure of the single image, accurately obtaining the array relationships of text in the single image without needing to adapt to different single image templates, thus improving the recognition efficiency of text in the single image. The client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The invention will be described in detail below through specific embodiments.

[0062] Please see Figure 2 As shown, Figure 2 A flowchart illustrating a method for recognizing image text in a test form according to an embodiment of the present invention includes the following steps:

[0063] S10. Obtain the location information of different regions in the verification single image, and convert the text information in different regions of the verification single image into text feature vectors.

[0064] A test report is a pre-prepared detailed list used to record specific behaviors or characteristics. In the medical field, test reports typically include examination information for different items. For example, a complete blood count test report includes information on hemoglobin count, white blood cell count, platelets, and red blood cells, while a urinalysis test report includes information on cytokines, red blood cells, and urine color. The location information of different areas in the test report image can include, but is not limited to, text areas and table areas. The text area mainly includes basic test information, such as patient name, patient age, patient information, and test items. The table area mainly includes the test results, such as the item name, test value, and normal value range.

[0065] It's important to understand that, considering the layout and structure of the test report, different regions within the test report image require different reading methods, necessitating the detection of these different regions. Specifically, for example... Figure 3 As shown, step S10, which involves obtaining the location information of different regions in the single verification image and converting the text information within different location regions of the single verification image into text feature vectors, includes the following steps:

[0066] S11. Obtain the inspection form image, perform layout analysis on the inspection form image, and obtain the position information of different areas in the inspection form image.

[0067] S12. Using the positional information of different regions in the verification single image, convert the text information in different positional regions of the verification single image into text feature vectors.

[0068] Specifically, in the process of layout analysis of inspection form images, the inspection form image can be acquired, and a second network model trained using an object detection algorithm can be used to perform layout analysis on the inspection form image. This yields bounding boxes corresponding to different regions in the inspection form image, and the positional information of these regions is determined based on these bounding boxes. Here, the second network model trained by the object detection algorithm can be a Faster R-CNN network. By inputting the inspection form image into the trained Faster R-CNN network, predictions can be made, identifying the positional coordinates of different regions such as text and tables.

[0069] Specifically, in the process of converting text information in different regions of a single image into text feature vectors, text recognition technology can be used to extract the text information and its coordinate layout within the image based on the positional information of different regions. Using this coordinate layout information, the text information in different regions is serialized according to a predetermined method. Then, an embedded vector model is used to vectorize and embed the text and its coordinates to obtain text feature vectors. Here, text recognition technology can use OCR to scan the positional information of different regions in the single image, and through image processing and pattern recognition techniques, the optical characters can be identified, thus retrieving the text information in different regions of the single image.

[0070] It is understandable that each character in the text information within different regions corresponds to a text region. The coordinate layout information of the text information in the verification image mainly includes the coordinate information of each text region and the coordinate information of the text position. In order to accurately extract the text features in the verification image, the text information needs to be sorted according to the usual reading method. Here, the text information can be serialized according to the usual reading method from left to right and from top to bottom. Then, combined with the BERT vector embedding method, the characters and position coordinates are functionally embedded separately, and finally the text feature vector is obtained.

[0071] For example, the text feature vector here can be represented as W = E1 + E2 + E3, where W represents the final text vector representation of word w, E1 represents the vector embedding of word w, E2 represents the position embedding of word w in the sequence, and E3 represents the coordinate position embedding of word w in the image.

[0072] S20. The first network model trained using the object detection algorithm encodes the single image of the test into a visual feature vector.

[0073] Each visual feature vector represents a visual feature within a preset image region in the test image. Here, the first network model trained by the object detection algorithm can be a Faster R-CNN network. By inputting the test image into the trained Faster R-CNN network to encode the test image, multiple visual feature vectors can be obtained.

[0074] S30. After combining the text feature vector and the visual feature vector into a multimodal sequence feature, the feature representing the entity in the single image is obtained by mapping through a linear network.

[0075] To enrich the feature representation in the single test image, the BERT network model can be used to combine text feature vectors and visual feature vectors into multimodal sequence features. The BERT network model consists of multiple multimodal coding modules. By concatenating the text feature vectors and visual feature vectors into the multimodal coding module in sequence, a rich feature representation of the single test image after the interaction of text, image and layout information can be obtained.

[0076] The multimodal coding module here includes two modules: a self-attention submodule, which learns to capture long-range dependent contextual information in text feature vectors and visual feature vectors, and integrates text, layout, and image information; and a forward propagation network module, which contains two fully connected layers and can extract features after the interaction of text, layout, and image. Each submodule is followed by a normalization network to improve the model's generalization ability.

[0077] Specifically, the entities in the inspection form image are equivalent to the inspection item information in the inspection form image. In the process of mapping the multi-modal sequence features to the features representing the entities in the inspection form image through a linear network, two linear networks can be set. After combining the text feature vector and the visual feature vector into a multi-modal sequence feature, they are respectively input into the two linear networks for mapping to obtain the features representing the entities in the inspection form image. Here, the network structures corresponding to the two linear networks are the same, but the network parameters used are different. One is used to map the start feature representing the entity, and the other is used to map the end feature representing the entity. For example, for the entity of normal numerical range, the start feature representing the entity is the text "正", and the end feature representing the entity is the text "围". After passing through the two linear networks, the inspection form image can be represented as the features of different inspection items.

[0078] S40. Utilize the features representing the entities in the inspection form image to identify the text in the inspection form image and the text correspondence relationship, and obtain the text array included in the inspection form.

[0079] Here, the text correspondence relationship in the inspection form image is the correspondence relationship between the inspection items and the item results, item units, and item values respectively. This correspondence relationship can be in the form of triples. For example, if the inspection item is white blood cells, the text correspondence relationship can at least include: (white blood cell count, value, 5.8), (white blood cell count, unit, 10^9 / L), (white blood cell count, normal range, 3.5 - 9.5).

[0080] It should be understood that the above features representing the entities include the start feature and the end feature representing the entities. Specifically, as Figure 4 shown, in step S40, that is, utilize the features representing the entities in the inspection form image to identify the text in the inspection form image and the text correspondence relationship, and obtain the text array included in the inspection form, which includes the following steps:

[0081] S41. Use the double affine attention mechanism to perform affine transformation on the start feature representing the entity and the end feature representing the entity, and obtain the matrix vector representing the pair from the start to the end of the entity in the verification form image.

[0082] S42. Identify the index relationship between the entities in the verification form image in the matrix vector representing the pair from the start to the end of the entity.

[0083] S43. Utilize the index relationship between the entities in the verification form image to obtain the text in the inspection form image and the text correspondence relationship, and obtain the text array included in the inspection form.

[0084] Specifically, in the process of identifying the index relationship between entities in a single image in the matrix vector representing the start-to-end pairs of entities, different entity category labels can be predefined. The position categories representing different entity category labels are determined in the matrix vector representing the start-to-end pairs of entities. Here, the position categories include the start position relationship category and the end position relationship category. Based on the start position relationship category and the end position relationship category representing different entity category labels, the index relationship between entities in the single image is identified.

[0085] Here, we can learn and infer the positional relationships between different entity labels and the start and end features of an entity in the matrix vector, such as check items, numerical values, units, and normal ranges. Furthermore, the positional correspondence can be used as an index relationship between entities in a single image. For example, the white blood cell count corresponds to the entity start-to-end pair (1,5), so the position (1,5) in the matrix vector represents the label of the check item.

[0086] In practical applications, different entity category labels can be predefined. For example, the white blood cell count corresponds to the entity start-to-end pair as (0,4). In this case, position (0,4) in the matrix vector represents the check item label. Furthermore, the position relationship category representing the start of the entity and the position relationship category representing the end of the entity can be determined in the matrix vector. For example, the start position of the white blood cell count is 0 and the corresponding value 5.8 starts at position 5. In this case, position (0,5) in the matrix vector represents the position relationship category of the entity start. Similarly, the end position of the white blood cell count is 4 and the corresponding value 5.8 ends at position 7. In this case, position (4,7) in the matrix vector represents the position relationship category of the entity end.

[0087] Specifically, in the process of obtaining the text and text correspondence in the inspection form image to obtain the text array contained in the inspection form, the index relationship between entities in the inspection form image can be used to obtain the entity category labels and the start and end index positions of all entities in the inspection form image. The start and end index positions of all entities are checked to determine whether there is a corresponding entity category label at the start and end index positions of the entities. If so, the text and text correspondence in the inspection form image are obtained based on the start and end index positions of all entities in the inspection form image to obtain the text array contained in the inspection form. Here, we can first check whether there is a corresponding entity category label at the beginning position of the entity, and then check whether there is a corresponding entity category label at the end position of the entity. If there are corresponding entity category labels, the text in the verification form is decoded into a triple form according to the text correspondence relationship in the verification form image. For example, white blood cell count (0,4), 5.8 (5,7), the relationship value corresponds to the beginning position (0,5) satisfying that 0 is the beginning position of white blood cell count, 5 is the beginning position of value 5.8, the end position of white blood cell count 4 and the end position of value 5.8 7 form the position (4,7), which corresponds to an entity category label. Further decoding yields the triple form (white blood cell count, value, 5.8).

[0088] In practical applications, the process of recognizing text in images on a specific inspection form can be as follows: Figure 5 As shown, firstly, the laboratory report is obtained. On one hand, the report's layout is analyzed to obtain text and table areas. OCR is then used to obtain text and layout information, further extracting text and layout features. On the other hand, visual features are extracted from the report to obtain image visual features. These text and layout features are combined with the image visual features to form multimodal text features. These features are then processed through a cross-modal model to obtain multimodal sequence features. Feature mapping of the multimodal feature sequence is performed through a feedforward network (start) and a feedforward network (end) to obtain start and end features representing entities. Finally, through B... The iaffine affine transformation yields a matrix vector of entity start-end pairs. In this matrix vector, indices between entities are learned and inferred. These indices are used to retrieve the text and its correspondences from the lab report. Based on these correspondences, the text is decoded into triplet forms, specifically: (white blood cell count, value, 5.8), (white blood cell count, unit, 10^9 / L), (white blood cell count, normal range, 3.5-9.5), (hemoglobin, value, 139), (hemoglobin, unit, g / L), (hemoglobin, normal range, 115-150).

[0089] This embodiment provides a method for recognizing text in an image of a verification form. By acquiring the positional information of different regions in the verification form image, the text information in different regions of the verification form image is converted into text feature vectors. A first network model trained using an object detection algorithm encodes the verification form image into visual feature vectors. Each visual feature vector represents a visual feature in a preset image region in the verification form image. After combining the text feature vectors and visual feature vectors into multimodal sequence features, the features representing entities in the verification form image are obtained through linear network mapping. Using the features representing entities in the verification form image, the text and text correspondence in the verification form image are recognized, resulting in a text array contained in the verification form. Compared with the existing technology that uses a pipeline approach to extract structured text from test images, this application treats the structured extraction of test images as an extraction task of entities and relationships from multimodal information. On the one hand, it enriches the text features through multimodal features, avoids the introduction of cascaded error prediction, and improves the recognition results of text in test images. On the other hand, by extracting joint information, it combines the correspondence between text in the image on the basis of the test image structure, and can accurately obtain the array relationship of text in the test image without adapting to different test image templates, thus improving the recognition efficiency of text in test images.

[0090] In one embodiment, a device for recognizing image text in a test form is provided, which corresponds one-to-one with the image text recognition method in the test form described in the above embodiments. For example... Figure 6 As shown, the image text recognition device in this inspection form includes: an acquisition module 101, an encoding unit 102, a mapping unit 103, and a recognition unit 104. Detailed descriptions of each functional module are as follows:

[0091] The acquisition unit 101 is used to acquire the position information of different regions in the verification single image and convert the text information in different position regions in the verification single image into text feature vectors.

[0092] Encoding unit 102 is used to encode the test single image into a visual feature vector using a first network model trained with an object detection algorithm, wherein each visual feature vector represents a visual feature within a preset image region in the test single image;

[0093] The mapping unit 103 is used to combine the text feature vector and the visual feature vector into a multimodal sequence feature, and then map it through a linear network to obtain the features representing entities in a single image for verification.

[0094] The recognition unit 104 is used to identify the text and text correspondence in the inspection form image by utilizing the features representing entities in the inspection form image, and to obtain the text array contained in the inspection form.

[0095] In one embodiment, the acquisition unit 101 includes:

[0096] The analysis module is used to acquire the inspection form image, perform layout analysis on the inspection form image, and obtain the position information of different regions in the inspection form image;

[0097] The conversion module is used to convert text information in different regions of the verification single image into text feature vectors by utilizing the positional information of different regions in the verification single image.

[0098] In one embodiment, the analysis module is specifically used to acquire a test form image, perform layout analysis on the test form image using a second network model trained with an object detection algorithm to obtain detection boxes corresponding to different regions in the test form image, and determine the position information of different regions in the test form image based on the detection boxes corresponding to different regions in the test form image.

[0099] In one embodiment, the conversion module is specifically used to extract text information and coordinate layout information of the text information in different regions of the verification single image based on the position information of different regions in the verification single image using text recognition technology; using the coordinate layout information of the text information in the verification single image, the text information in different regions of the verification single image is serialized according to a set method, and then the text and position coordinates in the text information are vectorized and embedded using an embedded vector model to obtain text feature vectors.

[0100] In one embodiment, the mapping unit 103 is specifically used to combine the text feature vector and the visual feature vector into a multimodal sequence feature, and then input them into two linear networks for mapping to obtain the features representing entities in a single image.

[0101] The features representing an entity include a start feature and an end feature, and the identification unit 104 includes:

[0102] The radiating module is used to perform affine transformations on the start features and end features of the entity representation using a dual affine attention mechanism to obtain a matrix vector representing the start-to-end pairs of entities in the verification single image.

[0103] The recognition module is used to identify the index relationship between entities in the verification single image in the matrix vector representing the start-to-end pairs of entities;

[0104] The acquisition module is used to obtain the text and text correspondence in the verification form image by utilizing the index relationship between entities in the verification form image, and obtain the text array contained in the verification form.

[0105] In one embodiment, the recognition module is specifically used to predefine different entity category labels, determine the position category representing the different entity category labels in the matrix vector representing the start-to-end pairs of entities, the position category including the start position relationship category and the end position relationship category; and identify the index relationship between entities in the verification single image according to the start position relationship category and the end position relationship category representing the different entity category labels.

[0106] In one embodiment, the acquisition module is specifically used to obtain entity category labels and the start and end index positions of all entities in the verification form image by utilizing the index relationship between entities in the verification form image; to check the start and end index positions of all entities and determine whether there is a corresponding entity category label at each of the start and end index positions of the entities; if so, to obtain the text in the verification form image and the text correspondence based on the start and end index positions of all entities in the verification form image, thereby obtaining the text array contained in the verification form.

[0107] This embodiment provides a device for recognizing text in an image of a verification form. By acquiring the positional information of different regions in the verification form image, the text information in different regions of the verification form image is converted into text feature vectors. A first network model trained using an object detection algorithm encodes the verification form image into visual feature vectors. Each visual feature vector represents a visual feature in a preset image region in the verification form image. After combining the text feature vectors and the visual feature vectors into multimodal sequence features, the features representing entities in the verification form image are obtained through linear network mapping. Using the features representing entities in the verification form image, the text and text correspondence in the verification form image are recognized, and the text array contained in the verification form is obtained. Compared with the existing technology that uses a pipeline approach to extract structured text from test images, this application treats the structured extraction of test images as an extraction task of entities and relationships from multimodal information. On the one hand, it enriches the text features through multimodal features, avoids the introduction of cascaded error prediction, and improves the recognition results of text in test images. On the other hand, by extracting joint information, it combines the correspondence between text in the image on the basis of the test image structure, and can accurately obtain the array relationship of text in the test image without adapting to different test image templates, thus improving the recognition efficiency of text in test images.

[0108] Specific limitations regarding the image-text recognition device on inspection forms can be found in the above-mentioned limitations on the image-text recognition method on inspection forms, and will not be repeated here. Each module in the aforementioned image-text recognition device on inspection forms can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the corresponding operations of each module.

[0109] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 7 As shown. The computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a server-side method for recognizing image text in a verification form.

[0110] In one embodiment, a computer device is provided, which may be a client, and its internal structure diagram may be as follows: Figure 8 As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface is used to communicate with an external server via a network connection. When executed by the processor, the computer program implements client-side functions or steps of a method for recognizing image text in a verification form.

[0111] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps:

[0112] Obtain the location information of different regions in the single image of the test, and convert the text information in the different location regions of the single image of the test into text feature vectors;

[0113] The first network model trained using the object detection algorithm encodes the test single image into a visual feature vector, where each visual feature vector represents a visual feature within a preset image region in the test single image.

[0114] After combining the text feature vector and the visual feature vector into a multimodal sequence feature, the feature representing the entity in the single image is obtained by mapping through a linear network.

[0115] By utilizing the features representing entities in the inspection form image, the text and text correspondence in the inspection form image are identified, and the text array contained in the inspection form is obtained.

[0116] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:

[0117] Obtain the location information of different regions in the single image of the test, and convert the text information in the different location regions of the single image of the test into text feature vectors;

[0118] The first network model trained using the object detection algorithm encodes the test single image into a visual feature vector, where each visual feature vector represents a visual feature within a preset image region in the test single image.

[0119] After combining the text feature vector and the visual feature vector into a multimodal sequence feature, the feature representing the entity in the single image is obtained by mapping through a linear network.

[0120] By utilizing the features representing entities in the inspection form image, the text and text correspondence in the inspection form image are identified, and the text array contained in the inspection form is obtained.

[0121] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0122] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0123] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0124] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A method for recognizing text in an image, characterized in that, The method includes: Obtain the location information of different regions in the test image, and convert the text information in the different location regions of the test image into text feature vectors; The first network model trained using the object detection algorithm encodes the test single image into a visual feature vector, where each visual feature vector represents a visual feature within a preset image region in the test single image. After combining the text feature vector and the visual feature vector into a multimodal sequence feature, the feature representing the entity in the single test image is obtained through linear network mapping. In the process of mapping the multimodal sequence feature to the feature representing the entity in the single test image through the linear network, two linear networks are set up. After combining the text feature vector and the visual feature vector into a multimodal sequence feature, the two linear networks are respectively input into the two linear networks for mapping to obtain the feature representing the entity in the single test image. The two linear networks have the same network structure, but use different network parameters. One is used to map the start feature representing the entity, and the other is used to map the end feature representing the entity. Using the features representing entities in the inspection form image, the text and text correspondences in the inspection form image are identified to obtain a text array contained in the inspection form. This includes: using a dual affine attention mechanism to perform affine transformations on the start features and end features representing entities to obtain a matrix vector representing the start-to-end pairs of entities in the inspection form image; identifying the index relationships between entities in the inspection form image in the matrix vector representing the start-to-end pairs of entities; using the index relationships between entities in the inspection form image to obtain the text and text correspondences in the inspection form image to obtain a text array contained in the inspection form; pre-defining different entity category labels; and determining the positional relationship categories representing the start and end positions of entities in the matrix vector.

2. The method according to claim 1, characterized in that, The step of obtaining the location information of different regions in the test image and converting the text information within the different location regions in the test image into text feature vectors includes: Acquire an inspection form image, perform layout analysis on the inspection form image, and obtain the position information of different regions in the inspection form image; Using the location information of different regions in the test image, the text information in different location regions of the test image is converted into text feature vectors.

3. The method according to claim 2, characterized in that, The process of acquiring the inspection form image, performing layout analysis on the inspection form image, and obtaining the positional information of different regions in the inspection form image includes: A test form image is acquired, and a second network model trained using an object detection algorithm is used to perform layout analysis on the test form image to obtain detection boxes corresponding to different regions in the test form image. The location information of different regions in the test form image is determined based on the detection boxes corresponding to different regions in the test form image.

4. The method according to claim 2, characterized in that, The step of converting text information within different regions of the test image into text feature vectors using positional information of different regions in the test image includes: Based on the positional information of different regions in the test image, text recognition technology is used to extract text information in different regions of the test image and the coordinate layout information of the text information in the test image. Using the coordinate layout information of the text information in the single image, the text information in different regions of the single image is serialized according to a set method, and then the text and position coordinates in the text information are vectorized and embedded using an embedded vector model to obtain text feature vectors.

5. The method according to claim 1, characterized in that, The step of identifying the index relationship between entities in the single image of the test image in the matrix vector representing the start-to-end pairs of entities includes: Different entity category labels are predefined. In the matrix vector representing the start-to-end pairs of entities, the position categories representing different entity category labels are determined. The position categories include the start position relationship category and the end position relationship category. Based on the start position relationship category and end position relationship category representing different entity category labels, identify the index relationship between entities in the single image being inspected.

6. The method according to claim 1, characterized in that, The method utilizes the index relationships between entities in the inspection form image to obtain the text and text correspondence relationships in the inspection form image, resulting in a text array contained in the inspection form, including: By utilizing the index relationship between entities in the single inspection image, the entity category labels and the start and end index positions of all entities in the single inspection image are obtained; For each entity, check the start and end index positions to determine if there is a corresponding entity category label at each of the entity's start and end index positions. If so, then based on the start and end index positions of all entities in the inspection form image, obtain the text and text correspondence in the inspection form image, and obtain the text array contained in the inspection form.

7. A device for recognizing image text in a verification form, characterized in that, The device includes: The acquisition unit is used to acquire the position information of different regions in the test image and convert the text information in the different position regions of the test image into text feature vectors. The encoding unit is used to encode the test single image into a visual feature vector using a first network model trained with an object detection algorithm, whereby each visual feature vector represents a visual feature within a preset image region in the test single image. The mapping unit is used to combine the text feature vector and the visual feature vector into a multimodal sequence feature, and then map it through a linear network to obtain the features representing entities in the test image. In the process of mapping the multimodal sequence feature to the features representing entities in the test image through the linear network, two linear networks are set up. After the text feature vector and the visual feature vector are combined into a multimodal sequence feature, they are respectively input into the two linear networks for mapping to obtain the features representing entities in the test image. The two linear networks have the same network structure, but use different network parameters. One is used to map the start feature representing the entity, and the other is used to map the end feature representing the entity. The recognition unit is used to identify the text and text correspondences in the inspection form image by utilizing the features representing entities in the inspection form image, and to obtain a text array contained in the inspection form. This includes: performing affine transformations on the start features and end features representing entities using a dual affine attention mechanism to obtain a matrix vector representing the start-to-end pairs of entities in the inspection form image; identifying the index relationships between entities in the inspection form image within the matrix vector representing the start-to-end pairs of entities; obtaining the text and text correspondences in the inspection form image using the index relationships between entities in the inspection form image, and obtaining a text array contained in the inspection form; pre-defining different entity category labels; and determining the positional relationship categories representing the start and end positions of entities in the matrix vector, respectively.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.

9. A computer storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Text structured extraction method, device and equipment and storage medium

    CN112001368A

  • Text extraction method, text extraction model training method, device and equipment

    CN114821622A

  • Text recognition method and device and computer readable storage medium

    CN115730043A