An extraction method, device, apparatus and storage medium

By integrating an OCR module into smart devices, the problem of high maintenance costs for cloud servers has been solved, enabling efficient and accurate extraction of financial report data.

CN115311669BActive Publication Date: 2026-01-13BANK OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210967708.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-12
Publication Date
2026-01-13
Estimated Expiration
2042-08-12

AI Technical Summary

Technical Problem

In existing technologies, integrating OCR modules into cloud servers results in high maintenance costs.

Method used

The OCR module is integrated into a smart device, which recognizes the text content in the image and generates a second table to extract the target data.

Benefits of technology

This reduces the maintenance costs of cloud servers and improves the efficiency and accuracy of data extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115311669B_ABST
    Figure CN115311669B_ABST
Patent Text Reader

Abstract

The application discloses an extraction method, device, equipment and storage medium, which can be applied to the field of artificial intelligence or the field of finance. The method comprises the following steps: receiving target extraction data and an identification result of a first table sent by an intelligent device; the identification result comprises text content in the first table and coordinates of a cell where the text content is located; a second table about the text content is generated according to the coordinates of the cell where the text content is located; the position information of the target extraction data in the second table is determined according to the target extraction data; and target data is extracted from the second table according to the position information. The first table is identified by the intelligent device to obtain an identification result (the identification step is transferred to the intelligent device to be executed), so that the maintenance cost of a cloud server is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data extraction technology, and in particular to an extraction method, apparatus, device and storage medium. Background Technology

[0002] Financial statements from corporate clients play a crucial role in risk control. Banks use these statements to assess clients and determine whether to grant loans. Therefore, accurate and efficient data entry into the bank's system is essential.

[0003] In existing technologies, firstly, the paper version of the financial report needs to be converted into an electronic version by scanning; then, the text in the financial report is recognized by an optical character recognition (OCR) module integrated in a cloud server; finally, the financial report data required for rating is extracted. Since the OCR module is integrated into the cloud server, the cloud server needs to be maintained. Therefore, this incurs significant maintenance costs.

[0004] Therefore, how to reduce the maintenance cost of cloud servers is a technical problem that needs to be solved by those skilled in the art. Summary of the Invention

[0005] To address the aforementioned issues, this application provides an extraction method, apparatus, device, and storage medium to reduce maintenance costs for cloud servers.

[0006] The embodiments of this application disclose the following technical solutions:

[0007] In a first aspect, this application provides an extraction method, the method comprising:

[0008] The system receives target extraction data and recognition results of a first table sent by a smart device; the recognition results include the text content in the first table and the coordinates of the cell where the text content is located.

[0009] Generate a second table about the text content based on the coordinates of the cell where the text content is located;

[0010] Based on the target extracted data, determine the position information of the target extracted data in the second table;

[0011] Based on the location information, target data is extracted from the second table.

[0012] Optionally, before generating the second table about the text content, the method further includes:

[0013] The received text content is corrected using natural language processing technology to obtain suggested modifications;

[0014] Based on the suggested modifications, the text content was revised.

[0015] Optionally, determining the location information of the target extracted data in the second table based on the target extracted data specifically includes:

[0016] The corresponding row and column numbers are extracted from the second table to match the target data using field matching.

[0017] Based on the row number and the column number, determine the location information of the target extracted data in the second table.

[0018] Optionally, extracting the target data from the second table specifically includes:

[0019] A deep learning model is used to extract target data from the second table; the deep learning model is trained using the extracted target data and the row and column information corresponding to the extracted target data.

[0020] Optionally, the method further includes:

[0021] Obtain the coordinates of the corresponding cell in the second table for the extracted target data;

[0022] Based on the coordinates of the corresponding cell in the second table, obtain the text content corresponding to that coordinate in the first table;

[0023] If the obtained target data matches the corresponding text content in the first table under the coordinates of the target data unit, then the extracted target data is correct.

[0024] Secondly, this application provides an extraction device, the device comprising: a receiving module, a generating module, a determining module, and an extraction module;

[0025] The receiving module is used to receive target extraction data and recognition results of the first table sent by the smart device; the recognition results include the text content in the first table and the coordinates of the cell where the text content is located.

[0026] The generation module is used to generate a second table about the text content based on the coordinates of the cell where the text content is located;

[0027] The determining module is used to determine the position information of the target extracted data in the second table based on the target extracted data;

[0028] The extraction module is used to extract target data from the second table based on the location information.

[0029] Optionally, the determining module is specifically used for:

[0030] The corresponding row and column numbers are extracted from the second table to match the target data using field matching.

[0031] Based on the row number and the column number, determine the location information of the target extracted data in the second table.

[0032] Optionally, the extraction module is specifically used for:

[0033] A deep learning model is used to extract target data from the second table; the deep learning model is trained using the extracted target data and the row and column information corresponding to the extracted target data.

[0034] Optionally, the device further includes: an identification module;

[0035] The recognition module is used by the smart device to recognize the first table and obtain the recognition result.

[0036] Optionally, the device further includes: a correction module;

[0037] The correction module is used to correct any errors in the received text content.

[0038] Optionally, the device further includes: a determination module;

[0039] The judgment module is used to determine whether the extracted target data is correct.

[0040] Thirdly, this application provides an extraction device, the device comprising: a memory and a processor;

[0041] The memory is used to store computer programs;

[0042] The processor is configured to implement the steps of the extraction method as described in any of the first aspects when executing the computer program.

[0043] Fourthly, this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the extraction method as described in any of the first aspects.

[0044] This application provides an input method, which includes: first, receiving target extraction data and recognition results of a first table sent by a smart device; the recognition results include text content in the first table and the coordinates of the cell where the text content is located; second, generating a second table about the text content based on the coordinates of the cell where the text content is located; then, determining the position information of the target extraction data in the second table based on the target extraction data; and finally, extracting target data from the second table based on the position information.

[0045] Compared with the prior art, this application has the following beneficial effects:

[0046] The first form is identified by a smart device to obtain the identification result (the identification steps are transferred to the smart device for execution), which reduces the maintenance cost of the cloud server. Attached Figure Description

[0047] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0048] Figure 1 A flowchart illustrating an extraction method provided in an embodiment of this application;

[0049] Figure 2 A schematic diagram of the first table provided in the embodiments of this application;

[0050] Figure 3 A schematic diagram of the second table provided in an embodiment of this application;

[0051] Figure 4 A flowchart illustrating another extraction method provided in this application embodiment;

[0052] Figure 5 This is a schematic diagram of an extraction device provided in an embodiment of this application. Detailed Implementation

[0053] As described above, current OCR modules are integrated into cloud servers, which will require corresponding maintenance of the cloud in the long run, thus inevitably generating a large amount of maintenance costs.

[0054] Through research, the inventors discovered that the OCR module can be integrated into smart devices. In this way, text content in images can be recognized through edge recognition, eliminating the need to maintain cloud servers and reducing maintenance costs.

[0055] In view of this, this application provides an extraction method, the method comprising: receiving target extraction data sent by a smart device and recognition results of a first table; the recognition results including text content in the first table and coordinates of the cell where the text content is located; generating a second table about the text content based on the coordinates of the cell where the text content is located; determining the location information of the target extraction data in the second table based on the target extraction data; and extracting target data from the second table based on the location information.

[0056] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present application.

[0057] See Figure 1 The figure is a flowchart of an extraction method provided in an embodiment of this application.

[0058] like Figure 1 As shown, the extraction method includes:

[0059] S101: Receive target extraction data and recognition results of the first table sent by the smart device; the recognition results include the text content in the first table and the coordinates of the cell where the text content is located.

[0060] The cloud server receives the target extraction data sent by the smart device and the recognition results of the first table.

[0061] Smart devices can be devices with smart chips, including mobile phones, computers, smartwatches, tablets, and similar devices. Smart devices can recognize the first table through an inherited OCR module.

[0062] The first table can be a table in image format, such as a photo of a financial statement or an information sheet. No restrictions are placed on the information in the table in the image.

[0063] The target data to be extracted can be the information that ultimately needs to be obtained from the first table.

[0064] As an example, if the first table is a financial statement, and the data in the statement includes accounts receivable as of December 31, 2021, accounts receivable as of December 31, 2020, notes receivable as of December 31, 2021, and notes receivable as of December 31, 2020, then the target data to be extracted can be the accounts receivable data corresponding to December 31, 2021. There is no limitation on the target data to be extracted here.

[0065] The recognition results may include the text content in the first table and the coordinates of the cell where the text content is located.

[0066] The text content can be the text, numbers, symbols, and other content recorded in the first table. For example, if the first table records "Accounts receivable as of December 31, 2021 are 1,231,082,748.72", then the text content includes the corresponding text, numbers, and symbols.

[0067] The coordinates of a cell can be the coordinates of the four vertices of the cell.

[0068] Refer to the instruction manual. Figure 2 Describe the content to be identified in the first table:

[0069] Project, [[214.0,23.0],[260.0,23.0],[260.0,54.0],[214.0,54.0]]

[0070] December 31, 2021, [[546.0,26.0],[736.0,26.0],[736.0,49.0],[546.0,49.0]]

[0071] December 31, 2020, [[972.0,25.0],[1162.0,25.0],[1162.0,47.0],[972.0,47.0]]

[0072] Current assets: [[51.0,74.0],[152.0,76.0],[152.0,106.0],[51.0,103.0]]

[0073] Cash and cash equivalents: [[98.0,128.0],[190.0,128.0],[190.0,157.0],[98.0,157.0]]

[0074] 9,733,956,725.88,[[687.0,131.0],[849.0,130.0],[849.0,149.0],[687.0,151.0]].

[0075] Among them, [[546.0,26.0],[736.0,26.0],[736.0,49.0],[546.0,49.0]] are the coordinates of the four vertices of the cell containing "December 31, 2021"; [[972.0,25.0],[1162.0,25.0],[1162.0,47.0],[972.0,47.0]] are the coordinates of the four vertices of the cell containing "December 31, 2020". The remaining information will not be repeated.

[0076] S102: Generate a second table about the text content based on the coordinates of the cell where the text content is located.

[0077] In the cloud server, a second table about the text content is generated based on the coordinates of the cell where the text content is located.

[0078] The second table is editable, so you can directly extract relevant data, text, and other information from it.

[0079] As an example, the second table is equivalent to filling the corresponding positions in the second table with the text from the first table after it has been recognized by OCR.

[0080] As an example, the instruction manual includes... Figure 3 This is the second table generated on the server.

[0081] S103: Based on the target extracted data, determine the location information of the target extracted data in the second table.

[0082] The target data to be extracted can be the information that is ultimately needed to be obtained from the second table.

[0083] As an example, if the first table is a financial statement, and the data in the statement includes accounts receivable as of December 31, 2021, accounts receivable as of December 31, 2020, notes receivable as of December 31, 2021, and notes receivable as of December 31, 2020, then the target data to be extracted can be the accounts receivable data corresponding to December 31, 2021. There is no limitation on the target data to be extracted here.

[0084] As an example, refer to the instruction manual appendix. Figure 3 To illustrate, if the target data to be extracted is "accounts receivable as of December 31, 2021", we can find the column where the target data is located by using "December 31, 2021" and the row where the target data is located by using "accounts receivable". Therefore, we know that the target data is located in the 7th row and 2nd column of the second table. Similarly, if the target data is "cash and cash equivalents as of December 31, 2020", it is located in the 3rd row and 3rd column of the second table.

[0085] S104: Extract target data from the second table based on the location information.

[0086] Location information can be the rows and columns of the target extracted data in the second table, through which a unique cell can be identified in the second table.

[0087] This application provides an input method, which includes: first, receiving target extraction data and recognition results of a first table sent by a smart device; the recognition results include text content in the first table and the coordinates of the cell where the text content is located; second, generating a second table about the text content based on the coordinates of the cell where the text content is located; then, determining the position information of the target extraction data in the second table based on the target extraction data; and finally, extracting target data from the second table based on the position information.

[0088] The first form is identified by a smart device to obtain the identification result (the identification steps are transferred to the smart device for execution), which reduces the maintenance cost of the cloud server.

[0089] See Figure 4 The figure is a flowchart of another extraction method provided in an embodiment of this application.

[0090] like Figure 4 As shown, the extraction method includes:

[0091] S401: Receive target extraction data and recognition results of the first table sent by the smart device; the recognition results include the text content in the first table and the coordinates of the cell where the text content is located.

[0092] The cloud server receives the target extraction data sent by the smart device and the recognition results of the first table.

[0093] Smart devices can be devices with smart chips, including mobile phones, computers, smartwatches, tablets, and similar devices. Smart devices can recognize the first table through an inherited OCR module.

[0094] The first table can be a table in image format, such as a photo of a financial statement or an information sheet. No restrictions are placed on the information in the table in the image.

[0095] The target data to be extracted can be the information that ultimately needs to be obtained from the first table.

[0096] As an example, suppose the first table is a financial statement, and the data in the statement includes accounts receivable as of December 31, 2021, accounts receivable as of December 31, 2020, notes receivable as of December 31, 2021, and notes receivable as of December 31, 2020. Then the target extraction data can be the accounts receivable data corresponding to December 31, 2021. Here, the target extraction data is not limited.

[0097] The recognition result can include the text content in the first table and the coordinates of the cell where the text content is located.

[0098] The text content can be the corresponding content such as words, numbers, symbols, etc. recorded in the first table. For example, if the first table records that "the accounts receivable corresponding to December 31, 2021 is 1,231,082,748.72", then the text content includes the corresponding words, numbers, and symbols.

[0099] The coordinates of the cell can be the coordinates of the 4 vertices of the cell.

[0100] S402: Use natural language processing technology to correct the received text content to obtain modification opinions; according to the modification opinions, correct the text content.

[0101] Natural language processing (NLP, Natural Language Processing) enables a computer to "understand" natural language.

[0102] As an example, by using NLP to train a machine learning model, the trained model can identify errors in the input text, such as: homophones or words with similar pronunciations (eyes, glasses), similar words or words with similar shapes (sorghum, gaoliang), reversed word order (born, give birth), full or abbreviated pinyin (Shanghai, Shanghai), extra words (good weather, good weather and smooth customs), missing words (Harbin City, Heilongjiang Province, Harbin City, Heilongjiang Province), etc.

[0103] Whether it is the picture information obtained through OCR recognition or the text input by the user through the input method, errors may occur. These errors will affect the readability of the text and are not conducive to the understanding of humans and machines. If these errors are not processed, they will spread to subsequent links and affect the effect of subsequent tasks.

[0104] S403: Generate a second table for the text content according to the coordinates of the cell where the text content is located.

[0105] S404: Match the corresponding row number and column number for the target extraction data in the second table by means of field matching; according to the row number and column number, determine the position information of the target extraction data in the second table.

[0106] Field matching can match the target extracted data with the text information in the second table.

[0107] As an example, if the target data to be extracted is "accounts receivable as of December 31, 2021", we can match "2021" with "December 31, 2021" in the second table to get the column number corresponding to "December 31, 2021"; we can match "accounts receivable" with "accounts receivable" in the second table to get the row number corresponding to "accounts receivable". In this way, we can obtain the row and column information of "accounts receivable as of December 31, 2021" in the second table as row 7, column 2.

[0108] S405: Extract target data from the second table based on the location information.

[0109] Location information can be the rows and columns of the target extracted data in the second table, through which a unique cell can be identified in the second table.

[0110] S406: Obtain the coordinates of the cell corresponding to the extracted target data in the second table; based on the coordinates of the cell corresponding to the target data in the second table, obtain the text content corresponding to the coordinates in the first table; if the obtained target data is consistent with the text content corresponding to the target data cell coordinates in the first table, then the extracted target data is correct.

[0111] The first table is identified by a smart device, and the identification result is obtained (the identification step is transferred to the smart device for execution), which reduces the maintenance cost of the cloud server; the content of the second table is corrected by correcting the text content, making the content of the second table closer to the first table; after the target data is extracted, the extracted target data is verified to further improve the accuracy of the extracted target data.

[0112] See Figure 5 The figure is a schematic diagram of an extraction device provided in an embodiment of this application.

[0113] like Figure 5 As shown, the extraction device includes: a receiving module 501, a generating module 502, a determining module 503, and an extraction module 504;

[0114] The receiving module 501 is used to receive target extraction data and recognition results of the first table sent by the smart device; the recognition results include the text content in the first table and the coordinates of the cell where the text content is located.

[0115] The generation module 502 is used to generate a second table about the text content based on the coordinates of the cell where the text content is located;

[0116] The determining module 503 is used to determine the position information of the target extracted data in the second table based on the target extracted data;

[0117] The extraction module 504 is used to extract target data from the second table based on the location information.

[0118] Optionally, the determining module 503 is specifically used for:

[0119] The corresponding row and column numbers are extracted from the second table to match the target data using field matching.

[0120] Based on the row number and the column number, determine the location information of the target extracted data in the second table.

[0121] Optionally, the extraction module 504 is specifically used for:

[0122] A deep learning model is used to extract target data from the second table; the deep learning model is trained using the extracted target data and the row and column information corresponding to the extracted target data.

[0123] In practical applications, the computer-readable storage medium can be any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. For example, a computer-readable storage medium can be, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this embodiment, the computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0124] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, capable of sending, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device.

[0125] Program code contained on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.

[0126] Computer program code for performing the operations of this invention can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0127] The extraction method, apparatus, device, and storage medium provided by this invention can be used in the financial field or other fields. For example, it can be used in financial data verification applications. Other fields refer to any field other than finance, such as the field of artificial intelligence. The above are merely examples and do not limit the application areas of the invention.

[0128] It should be noted that the various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, for the embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiments. The embodiments described above are merely illustrative, and the units described as separate components may or may not be physically separate. The components indicated as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of the solution in this embodiment according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0129] The above description is merely one specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. An extraction method characterized by, The method comprises: receiving target extraction data sent by an intelligent device and an identification result of a first table; the identification result comprises text content in the first table and coordinates of a cell where the text content is located; generating a second table about the text content according to the coordinates of the cell where the text content is located; determining position information of the target extraction data in the second table according to the target extraction data; extracting target data from the second table according to the position information; obtaining coordinates of a cell corresponding to the extracted target data in the second table; obtaining corresponding text content of the coordinates in the first table according to the coordinates of the cell corresponding to the target data in the second table; and if the obtained target data is consistent with the corresponding text content of the target data cell coordinates in the first table, the extracted target data is correct.

2. The method of claim 1, wherein, Before the generating of the second table about the text content, the method further comprises: correcting the received text content by a natural language processing technology to obtain a modification opinion; correcting the text content according to the modification opinion.

3. The method of claim 1, wherein, The determination of the position information of the target extraction data in the second table according to the target extraction data specifically comprises: matching corresponding row numbers and column numbers for the target extraction data in the second table by field matching; and determining the position information of the target extraction data in the second table according to the row numbers and the column numbers.

4. The method of claim 1, wherein, The extraction of the target data from the second table specifically comprises: extracting the target data from the second table by using a deep learning model; the deep learning model is trained by sample target extraction data and row and column information corresponding to the sample target extraction data.

5. An extraction device, characterized in that The device comprises a receiving module, a generating module, a determining module, an extracting module and a judging module. The receiving module is configured to receive target extraction data sent by an intelligent device and an identification result of a first table; the identification result comprises text content in the first table and coordinates of a cell where the text content is located. The generating module is configured to generate a second table about the text content according to the coordinates of the cell where the text content is located. The determining module is configured to determine position information of the target extraction data in the second table according to the target extraction data. The extracting module is configured to extract target data from the second table according to the position information. The judging module is configured to obtain coordinates of a cell corresponding to the extracted target data in the second table; obtain corresponding text content of the coordinates in the first table according to the coordinates of the cell corresponding to the target data in the second table; and if the obtained target data is consistent with the corresponding text content of the target data cell coordinates in the first table, the extracted target data is correct.

6. The apparatus of claim 5, wherein, The determining module is specifically configured to: match corresponding row numbers and column numbers for the target extraction data in the second table by field matching; and According to the row number and the column number, position information of the target extraction data in the second table is determined.

7. The apparatus of claim 5, wherein, The extraction module is specifically configured to: extract target data from the second table by using a deep learning model, wherein the deep learning model is trained by sample target extraction data and row and column information corresponding to the sample target extraction data.

8. An extraction apparatus, characterized by The device comprises a memory and a processor. The memory is configured to store a computer program. The processor is configured to implement the steps of the extraction method according to any one of claims 1 to 4 when the computer program is executed.

9. A computer-readable storage medium, characterized in that, The computer program is stored on the computer readable storage medium and is configured to implement the steps of the extraction method according to any one of claims 1 to 4 when executed by the processor.

Citation Information

Patent Citations

  • Table detection and identification method and medium

    CN113705286A

  • Table information extraction method and device, storage medium and electronic equipment

    CN113987112A