Method and device for extracting verification information in report and storage medium
By identifying and processing the page layout in the verification report, the verification information in the report is extracted, and the problems of low efficiency and poor accuracy of information extraction in the prior art are solved, and more efficient and accurate information extraction is achieved.
Patent Information
- Application Number
- CN202411993643.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art is inefficient and less accurate in extracting information in the verification report, especially when processing reports of complex page layouts, the accuracy of identification information is low.
By obtaining the image of the report, identify the layout of each page, and extract the calibration information from the report image according to the different page layouts. This method uses processor and memory to execute program instructions to realize the extraction of report information.
It improves the accuracy of extracting information and can effectively process reports of different page layouts, thereby improving the efficiency and accuracy of information extraction.
Smart Images

Figure CN119992573A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to the field of image processing technology. More specifically, the present disclosure relates to a method and apparatus and a computer storage medium for extracting verification information in a report. Background Art
[0002] A calibration report is a written record of the calibration results of a calibration object such as a measuring instrument, equipment, or meter, and is used to prove the accuracy and reliability of the measurement results of the calibration object.
[0003] In the prior art, the information in the verification report is usually manually entered into the information system to extract the information in the verification report. This method of extracting information is inefficient and prone to errors, resulting in low accuracy of the extracted information. In addition, those skilled in the art will also extract report information through an intelligent extraction system. This method has a low accuracy rate in identifying information in the verification report when extracting information from a verification report with a complex page layout.
[0004] In view of this, there is an urgent need to provide a method for extracting assay information in a report so as to improve the accuracy of the extracted information. Summary of the invention
[0005] In order to at least solve one or more technical problems mentioned above, the present disclosure proposes, in multiple aspects, a method, an apparatus, and a computer storage medium for extracting verification information from a report.
[0006] In a first aspect, the present disclosure provides a method for extracting verification information in a report, wherein the report comprises one or more pages, and the method comprises: acquiring a report image corresponding to one or more of the pages of the report; identifying a page layout corresponding to each of the report images; and extracting the verification information from the report image according to each of the page layouts.
[0007] In a second aspect, the present disclosure provides a device for extracting report information, comprising: a processor configured to execute program instructions; and a memory configured to store the program instructions, wherein when the program instructions are loaded and executed by the processor, the device executes the method for extracting verification information in a report as described in the first aspect.
[0008] In a third aspect, the present disclosure provides a computer storage medium having stored thereon a computer program for extracting report information, and when the computer program is executed by a processor, the method for extracting assay information in a report as described in the first aspect is implemented.
[0009] Through the method for extracting verification information in a report as provided above, the disclosed embodiment improves the accuracy of the extracted information by identifying the page layout corresponding to each report image and, based on different page layouts, being able to extract corresponding verification information from report images with different page layouts. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] By reading the following detailed description with reference to the accompanying drawings, the above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become readily understood. In the accompanying drawings, several embodiments of the present disclosure are shown in an exemplary and non-limiting manner, and the same or corresponding reference numerals represent the same or corresponding parts, wherein:
[0011] Figure 1 An exemplary flow chart showing a method for extracting assay information in a report according to some embodiments of the present disclosure;
[0012] Figure 2 An exemplary flow chart showing a method for obtaining a report image corresponding to a page according to some embodiments of the present disclosure is shown;
[0013] Figure 3 An exemplary flow chart showing a method for removing noise from an initial image according to some embodiments of the present disclosure;
[0014] Figure 4 An exemplary flow chart showing a method for extracting authentication information according to some embodiments of the present disclosure is shown;
[0015] Figure 5 An exemplary flow chart showing a method for extracting verification information according to other embodiments of the present disclosure is shown;
[0016] Figure 6 An exemplary structural block diagram of an apparatus for extracting verification information from a report according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0017] The following will be combined with the drawings in the embodiments of the present disclosure to clearly and completely describe the technical solutions in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present disclosure.
[0018] It should be understood that the terms "include" and "comprising" used in the specification and claims of the present disclosure indicate the presence of described features, integers, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or collections thereof.
[0019] It should also be understood that the terms used in this disclosure are only for the purpose of describing specific embodiments and are not intended to limit the disclosure. As used in this disclosure and claims, the singular forms of "a", "an", and "the" are intended to include the plural forms unless the context clearly indicates otherwise. It should also be further understood that the term "and / or" used in this disclosure and claims refers to any combination of one or more of the associated listed items and all possible combinations, including these combinations.
[0020] As used in this specification and claims, the term "if" may be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" may be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.
[0021] The specific implementation of the present disclosure is described in detail below with reference to the accompanying drawings.
[0022] Exemplary application scenarios.
[0023] In an existing method of extracting the verification information from the report, the verification information in the report is specifically identified manually and entered into the information system. This method of extracting information is time-consuming, inefficient, and may enter wrong information, resulting in errors in the extracted information, thereby causing low accuracy of the extracted information.
[0024] In another existing method of extracting calibration information from a report, a specific intelligent extraction system is used to extract information from the report. When this method extracts information from a calibration report with a complex page layout, the accuracy of identifying the information in the calibration report is low. For example, table headers, merged cells, data across rows or columns, etc. may lead to recognition errors.
[0025] Exemplary application scenarios.
[0026] In view of this, the presently disclosed embodiment provides a method for extracting verification information in a report, which identifies the page layout corresponding to each report image in the report and, according to different page layouts, can extract corresponding verification information from report images of different page layouts, thereby improving the accuracy of the extracted information.
[0027] In a first aspect, embodiments of the present disclosure provide a method for extracting verification information in a report.
[0028] Figure 1An exemplary flow chart of a method for extracting assay information in a report according to some embodiments of the present disclosure is shown.
[0029] As shown in the figure, in step S100, a report image corresponding to one or more pages is obtained; in step S200, a page layout corresponding to each report image is identified; in step S300, verification information is extracted from the report image according to each page layout.
[0030] In the embodiments of the present disclosure, a report refers to a written record of the results of the calibration of a calibration object such as a measuring instrument, equipment, or meter, and may also be referred to as a calibration report. For example, the report may be a report on the power usage for detection of a transposition. The report is usually in PDF (Portable Document Format), which has one or more pages, each of which is used to record and save corresponding calibration information.
[0031] First, in step S100, "report image" refers to the image corresponding to each page in the report, and the page in PDF format can be converted into a picture through PyMuPDF to obtain the report image. The above-mentioned "PyMuPDF" is a high-performance Python library for data extraction, analysis, conversion and operation of PDF (and other) documents. In some embodiments, each page can also be photographed with a camera after camera calibration to obtain the report image. It can be understood by those skilled in the art that obtaining the report image is not subject to any limitation.
[0032] In the embodiments of the present disclosure, page layout refers to various elements on a page, and the elements on the page include but are not limited to text, pictures, charts, etc.
[0033] Next, in step S200, the page layout corresponding to each report image is identified, that is, the arrangement of the text, pictures, charts and other elements in each report image, where the arrangement refers to the position, direction and size of each element in the page. For example, there is a report image, the upper half of which has a row of titles and the lower half has a chart. Through step S200, the elements constituting the report image and the arrangement of each element can be identified.
[0034] In some embodiments, elements in the report image can be extracted by PyMuPDF, and the positions and sizes of the elements can be obtained.
[0035] Then, step S300 is performed to extract the calibration information from the report images corresponding to different page layouts according to each page layout. The calibration information here refers to the relevant data and results recorded and reported during the calibration process of the test object. For example, there is a page, the upper half of which is a schematic diagram of a power supply, and the lower half is the rated power of the power supply, which is obtained by testing the power supply. Step S300 is executed, and according to the page layouts corresponding to different pages, elements in different pages, that is, calibration information, can be extracted for different page layouts, thereby improving the accuracy of extracting calibration information in reports with complex page layouts.
[0036] In summary, through steps S100-S300, the method for extracting verification information in a report provided by the disclosed embodiment can identify the page layout corresponding to each report image, and according to different page layouts, can extract corresponding verification information from report images with different page layouts, thereby improving the accuracy of the extracted information.
[0037] Figure 2 FIG. 1 is an exemplary flow chart of a method for obtaining a report image corresponding to a page in some embodiments of the present disclosure. It can be understood that the method for obtaining a report image corresponding to a page is a specific implementation of the aforementioned step S100, so the aforementioned method is combined with the method described in FIG. Figure 1 The features described can analogously apply here.
[0038] As shown in the figure, in step S110, an initial image corresponding to one or more pages is received; in step S120, a watermark of the initial image is removed to obtain a report image.
[0039] In some embodiments of the present disclosure, the "initial image" here refers to the picture converted from the PDF page, and the initial image may contain information unrelated to the calibration information, namely noise and watermarks. The noise can be other information unrelated to the calibration information that already exists in the calibration report, and it can also be a partial pixel block introduced when the page is converted into a picture. In step S110, the page in PDF format can be converted into a picture using PyMuPDF to obtain the above-mentioned initial image. Those skilled in the art will understand that PyMuPDF is an interface of the programming language Python, which will not be described in detail here.
[0040] Then, in step S120, the noise in the initial image is removed, and the initial image after the noise is removed is used as the above-mentioned report image, so as to avoid the reduction of the accuracy of identifying the page layout corresponding to the report image due to the influence of noise.
[0041] Through steps S110-S120, information irrelevant to the verification information, i.e., noise, can be removed, thereby avoiding the reduction in the accuracy of identifying the page layout corresponding to the report image due to the influence of noise, thereby improving the accuracy of identifying the page layout. In this way, the verification information can be extracted from the report image according to the page layout in the subsequent process, reducing the interference of noise on the extraction of verification information, thereby improving the accuracy of extracting verification information.
[0042] Figure 3 An exemplary flow chart of a method for removing noise from an initial image according to some embodiments of the present disclosure is shown. It can be understood that the method for removing noise from an initial image is a specific implementation of the aforementioned steps S100 and S120, so the aforementioned method is combined with Figure 1 and Figure 2 The features described can analogously apply here.
[0043] As shown in the figure, in step S121, the initial image is grayed to obtain a gray image; in step S122, the gray image is binarized to obtain a binary page image; in step S123, the noise of the binary page image is removed to obtain a report image.
[0044] First, the initial image is grayscaled through step S121. The "grayscale processing" here refers to the process of converting a color image into a grayscale image. The initial image that has been grayscaled is called a "grayscale image". It should be noted that the initial image is a color image. In a color image, each pixel is composed of three channels: red (R), green (G) and blue (B). The value of each channel represents the intensity of the color, and the value range is generally 0-255. The grayscale image is a single-channel image, and each pixel has only one grayscale value, indicating different grayscale levels from black to white. The grayscale image in step S121 refers to the initial image without color, only black and white.
[0045] Next, step S122 is executed, that is, binarization processing is performed on the grayscale image. Here, "binarization processing" refers to the process of further converting the grayscale image into an image containing only black and white, and the grayscale value corresponding to each pixel is 0 or 255. For the convenience of description, the grayscale image after binarization processing is referred to as a "binarized page image".
[0046] Next, the noise in the binary page image is removed by step S123, so that unnecessary small pixels in the binary page image can be removed, thereby removing information irrelevant to the verification information, such as pixel blocks and watermarks irrelevant to the verification information. For ease of description, the binary page image after the noise is removed is used as the report image.
[0047] In some embodiments, the noise in the binary page image can be removed by erosion operation and dilation operation. Those skilled in the art will appreciate that erosion operation and dilation operation are commonly used operations for removing noise from binary images. In the present disclosure, the specific method of removing noise in the binary page image is not subject to any limitation.
[0048] By removing the noise in the binary page image in step S123, some small pixels or isolated pixels irrelevant to the verification information can be removed, thereby making the obtained report image clearer.
[0049] From the above description, it can be seen that since the binary page image is a grayscale image after binarization processing, the binary page image is an image containing only black and white. Therefore, the pixel blocks that are not related to the calibration information can be clearly distinguished by color, so that the watermark and other information that are not related to the calibration information can be accurately extracted, and the watermark and other information that are not related to the calibration information can be removed to obtain a binary page image without noise, that is, a report image, so that the report image is clearer and it is easier to identify the page layout or other information of the report image.
[0050] In some other embodiments, after step S121, the grayscale image may be subjected to contrast enhancement processing, and the grayscale image subjected to contrast enhancement processing is referred to as a "grayscale enhanced image". Then, the noise of the grayscale enhanced image is removed to obtain the above-mentioned report image. The method of removing the noise of the grayscale enhanced image is the same as the method of removing the noise of the binarized page image, which will not be described in detail here.
[0051] The above-mentioned "contrast enhancement processing" refers to an image processing operation used to increase the difference between bright and dark parts of an image, making the details in the image more clearly visible. Contrast enhancement processing is achieved by adjusting the pixel value distribution of the grayscale image. Common methods include histogram equalization, linear contrast stretching, etc. The difference between different grayscales becomes more obvious, and the overall image looks clearer. It should be noted that the role of binarization processing is the same as that of contrast enhancement processing. Both are used to improve the recognizability of the test information and page layout in the report image, making the information in the report image clearer.
[0052] Through the above steps S121-S123, the initial image is grayed to obtain a grayscale image, and then the grayscale image is binarized to obtain a binary page image. Then, some small pixels or isolated pixels that are irrelevant to the inspection information are removed to make the binary page image clearer, thereby improving the recognizability of the report image, making the inspection information and page layout in the report image clearer, and making it easier to identify the page layout or other information of the report image.
[0053] In some embodiments of the present disclosure, the page layout includes a home page layout and a non-home page layout. The home page refers to the first page in the report, which is usually the cover of the report. The home page may include a title, pictures, and tables, etc. The home page layout here refers to the page layout corresponding to the home page, that is, the arrangement of each element or content in the home page. Non-home pages refer to all pages in the report except the first page, that is, all pages except the home page layout. The above-mentioned non-home page layout refers to the page layout corresponding to the non-home page, that is, the arrangement of each element or content in the non-home page.
[0054] It should be noted that the contents of the home page and the non-home page are different, and the arrangement of the contents is also different, that is, the layout of the home page is different from that of the non-home page. In this embodiment, different methods are used to extract the verification information in the report image for different page layouts.
[0055] In some embodiments, the verification information includes the title of the page, entities, and the association relationship between different entities. The entity here refers to a person's name, a place name, an organization name, a date and time, a proper noun, etc. In the embodiments of the present disclosure, the entity can be preset. For example, the entity can be preset as a power supply, a rated power of the power supply, and a numerical value corresponding to the rated power. The association relationship between different entities refers to the corresponding relationship between different entities. For example, there is a corresponding relationship between the rated power of the power supply and the numerical value corresponding to the rated power. The rated power of the power supply can be used as a key, and the numerical value corresponding to the rated power can be used as a value.
[0056] Figure 4 FIG. 1 is an exemplary flow chart of a method for extracting verification information according to some embodiments of the present disclosure. It can be understood that the method for extracting verification information is a specific implementation of the aforementioned step S300, so the aforementioned method is combined with the method described in FIG. Figure 1 The features described can analogously apply here.
[0057] As shown in the figure, in step S310, in response to the page layout corresponding to the report image being the home page layout, a text recognition model is used to recognize the report text in the report image, and the report text includes handwritten text; in step S320, an entity recognition model is used to recognize entities in the report text; in step S330, a relationship extraction model is used to extract the association relationships between different entities.
[0058] In step S310, if the page layout corresponding to the report image is the home page layout, then the report image is the home page. Therefore, the content in the report image is relatively small and the arrangement of the content is relatively simple. At this time, the report text in the report image is identified by the trained text recognition model. The "report text" here refers to the text in the report image. The text in the report image includes handwritten text and machine-printed text. The handwritten text is referred to as "handwritten text" here.
[0059] The above-mentioned text recognition model is a machine learning model, and the text recognition model is used to detect and recognize handwritten text in the report image. The report image sample is a report image used for training, and the handwritten text and the area where the handwritten text is located are marked in the report image sample. The text recognition model is trained by the report image sample to obtain a trained word recognition model. Those skilled in the art can understand that the method of training the text recognition model can be set by oneself and is not subject to any limitation.
[0060] In step S320, entities in the report text are extracted by using the trained entity recognition model. The above-mentioned entity recognition model is a machine learning model, and the entity recognition model is used to recognize entities in the report image.
[0061] The above-mentioned entity recognition model is a machine learning model, and the entity recognition model is used to extract entities in the report text. The report text sample is a text used for training, and various texts have been annotated in the report text sample, and the report text sample is also annotated with handwritten text. The entity recognition model is trained by the report text sample to obtain a trained entity recognition model. Those skilled in the art can understand that the method of training the entity recognition model can be set by oneself and is not subject to any limitation.
[0062] In step S330, the association relationship between different entities is extracted by the trained relationship extraction model. The above-mentioned relationship extraction model is a machine learning model, and the relationship extraction model is used to extract the association relationship between different entities. The association relationship sample is an association relationship between different entities used for training, and the association relationship sample has been marked with the association relationship between different entities. The relationship extraction model is trained by the association relationship sample to obtain a trained relationship extraction model. Those skilled in the art can understand that the method of training the relationship extraction model can be set by oneself and is not subject to any limitation.
[0063] Through the above steps S310-S320, when the page layout corresponding to the report image is the home page layout, since the layout of the home page is relatively simple, there is no need to perform layout analysis on the report image. The calibration information in the report image can be quickly extracted through the corresponding model, which can reduce the time for layout analysis and improve the efficiency of extracting calibration information.
[0064] In some embodiments, the verification information includes a verification table, which includes attributes to be identified and attribute values corresponding to the attributes to be identified.
[0065] The calibration table here refers to a table used to store calibration information, which can be an Excel table. The attribute to be identified refers to the attribute in the calibration information, for example, the rated voltage or rated current of the power supply. The attribute value corresponding to the attribute to be identified refers to the value corresponding to the attribute to be identified, for example, the value corresponding to the rated voltage of the power supply, or the value corresponding to the rated current of the power supply.
[0066] Figure 5 FIG. 1 is an exemplary flow chart of a method for extracting verification information according to another embodiment of the present disclosure. It can be understood that the method for extracting verification information is a specific implementation of the aforementioned step S300, so the aforementioned method is combined with the method described in FIG. Figure 1 The features described can analogously apply here.
[0067] As shown in the figure, in step S340, in response to the page layout corresponding to the report image being a non-home page layout, a layout analysis model is used to perform layout analysis on the report image to obtain a calibration table in the report image; in step S350, the calibration table is processed using a table information extraction model to obtain the attributes to be identified in the calibration table and the attribute values corresponding to the attributes to be identified.
[0068] In step S340, if the page layout corresponding to the report image is a non-home page layout, then the report image is a non-home page. Therefore, the report image contains more content and the arrangement of the content is more complex. At this time, it is necessary to perform layout analysis on the report image through a layout analysis model, that is, to extract the content or elements in the report image and save the extracted content or elements in the verification table.
[0069] The layout analysis here refers to the structural analysis of the page, classifying and extracting the content contained in different parts of the page to better extract and utilize information.
[0070] It should be noted that, in step S340, the content or elements extracted by layout analysis are stored in a verification table in the form of key-value pairs. Specifically, the verification table is a form.
[0071] In step S350, the verification table is processed using the table information extraction model to extract the attributes to be identified and the attribute values corresponding to the attributes to be identified in the verification table. The "attributes to be identified" here refer to the attributes of the detected object in the report. For example, when the detected object is a computer, the attributes to be identified of the computer include the power and weight of the computer, and the attribute values corresponding to the attributes to be identified are 500 watts and 5 kilograms.
[0072] In the embodiments disclosed herein, the layout analysis model is a machine learning model, and the layout analysis model is used to perform layout analysis on report images. The report page sample is an image corresponding to a page used for training, and the content and arrangement of the content in the page are marked in the report page sample. The layout analysis model is trained by the report page sample to obtain a trained layout analysis model. Those skilled in the art can understand that the method of training the layout analysis model can be set by oneself and is not subject to any limitation.
[0073] The above-mentioned table information extraction model is a machine learning model, which is used to extract the attributes to be identified and the attribute values corresponding to the attributes to be identified in the verification table. The table sample is a verification table used for training, and the table sample has been marked with the attributes to be identified and the attribute values corresponding to the attributes to be identified. The table information extraction model is trained with the table sample to obtain a trained table information extraction model. Those skilled in the art can understand that the method of training the table information extraction model can be set by oneself and is not subject to any limitation.
[0074] Through steps S340-S350, the layout analysis model can be used to perform complex layout analysis on non-home pages, and the extracted content or elements can be saved in the form of key-value pairs in the verification table. Moreover, the verification table can be processed through the table information extraction model to obtain the attributes to be identified in the verification table and the attribute values corresponding to the attributes to be identified, thereby completing information extraction.
[0075] It can be seen from the description of the above embodiments that the method for extracting verification information in a report provided in the present disclosure can adopt different models for different page layouts, thereby flexibly extracting the content of the pages in the report.
[0076] In summary, the disclosed embodiment can identify the page layout corresponding to each report image and extract corresponding verification information from report images with different page layouts according to different page layouts, thereby improving the accuracy of the extracted information.
[0077] In a second aspect, the disclosed embodiments also provide a device for extracting verification information from a report. Figure 6 An exemplary structural block diagram of an apparatus for extracting verification information from a report according to an embodiment of the present disclosure is shown.
[0078] As shown in the figure, the apparatus 600 for extracting verification information in a report includes a processor 610 and a memory 620. The processor 610 is configured to execute program instructions. The memory 620 is configured to store program instructions. When the program instructions are loaded and executed by the processor, the apparatus executes the method for extracting verification information in a report provided in the first aspect.
[0079] In a third aspect, the disclosed embodiments further provide a computer storage medium storing a computer program for extracting verification information from a report. When the computer program is executed by a processor, the method provided in the first aspect is implemented.
[0080] Although multiple embodiments of the present disclosure have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Those skilled in the art may think of many changes, modifications, and alternatives without departing from the thought and spirit of the present disclosure. It should be understood that in the process of practicing the present disclosure, various alternatives to the embodiments of the present disclosure described herein may be adopted. The attached claims are intended to define the scope of protection of the present disclosure, and therefore cover equivalents or alternatives within the scope of these claims.
Claims
1. A method for extracting verification information in a report, characterized in that The report includes one or more pages, and the method includes: Obtaining report images corresponding to one or more of the pages; identifying a page layout corresponding to each of the report images; The certification information is extracted from the report image according to each of the page layouts.
2. The method according to claim 1, characterized in that The obtaining of report images corresponding to one or more pages of the report comprises: Receiving one or more initial images corresponding to the pages; The watermark of the initial image is removed to obtain the report image.
3. The method according to claim 2, characterized in that The removing the watermark of the initial image to obtain the report image comprises: Performing grayscale processing on the initial image to obtain a grayscale image; Binarizing the grayscale image to obtain a binary page image; The noise of the binarized page image is removed to obtain the report image.
4. The method according to claim 1, characterized in that: The page layout includes a home page layout and a non-home page layout.
5. The method according to claim 4, characterized in that The verification information includes at least one of the title of the page, entities, and associations between different entities.
6. The method according to claim 5, characterized in that The extracting the verification information from the report image according to each of the page layouts comprises: In response to the page layout corresponding to the report image being the home page layout, recognizing report text in the report image using a text recognition model, wherein the report text includes handwritten text; Using an entity recognition model to identify entities in the report text; The relationship extraction model is used to extract the association relationships between different entities.
7. The method according to claim 4, characterized in that The verification information includes a verification table, and the verification table includes attributes to be identified and attribute values corresponding to the attributes to be identified.
8. The method according to claim 7, characterized in that The extracting the verification information from the report image according to each of the page layouts comprises: In response to the page layout corresponding to the report image being the non-home page layout, performing layout analysis on the report image using a layout analysis model to obtain the calibration table in the report image; The verification table is processed using a table information extraction model to obtain the attributes to be identified in the verification table and the attribute values corresponding to the attributes to be identified.
9. A device for extracting verification information in a report, comprising: a processor configured to execute program instructions; as well as A memory configured to store the program instructions, which, when loaded and executed by the processor, causes the apparatus to perform the method according to any one of claims 1 to 8.
10. A computer storage medium having stored thereon a computer program for extracting assay information from a report, wherein when the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.