Method for the semi-automatic extraction of sensitive data from documents subject to strict confidentiality of these data
Patent Information
- Application Number
- EP2023822421
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-11-15
- Filing Date
- 2023-11-14
- Publication Date
- 2025-09-24
AI Technical Summary
Current methods for extracting data from tabular documents, such as financial and accounting records, face challenges in ensuring accuracy and confidentiality, particularly due to errors from OCR tools and variability in document formats, which compromise the reliability of extracted data and violate confidentiality obligations.
A semi-automatic extraction method involving a secure kernel computer that disfigures tabular documents by removing amounts, allowing verifiers to correct identifiers without accessing original amounts, and using shape proximity algorithms to group and verify digit images, ensuring accurate extraction while maintaining confidentiality.
This method enhances data extraction reliability and confidentiality by reducing human verification time and preventing data leaks, meeting the demands of financial and accounting sectors for accurate and secure data processing.
Smart Images

Figure 1.1
Abstract
Description
DESCRIPTION TITLE: Semi-Automatic Extraction Process for Sensitive Data in Documents, Subject to Strict Confidentiality of This Data
[0001] The present invention relates to the field of extracting data from material (paper) or digital (computer image files) documents of the Referenced Tabular Document type, the boxes of which are filled with data in printed characters.
[0002] In many areas, it is important to be able to extract such data in an automated or semi-automated manner to feed data into subsequent processes.
[0003] For this purpose, character recognition tools, known as OCR (Optical Character Recognition), are traditionally used, which enable the generation of ASCII (American Standard Code for Information Interchange) character strings from image files of the documents to be processed.
[0004] The data extracted from these files resulting from OCR can then be subject to computer processing adapted to different needs.
[0005] These treatments may include, for example, verification, calculation or evaluation operations.
[0006] An example of the application of the above concerns the field of finance and accounting.
[0007] In this area, Referenced Tabular Documents must be edited by companies, for example to declare their annual results, or to edit pay slips. This filling can be done by typing positioned on the correct boxes of the blank Referenced Tabular Document or, more modernly, through Document-image editing software based on, in the background, the standardized model, and the relevant data extracted from the company's Databases, and entered in the relevant boxes.
[0008] These Referenced Tabular Documents are presented as listed successions of table pages made up of rows and columns where the numerical data (financial, salary or accounting) are entered by users in rectangular boxes materialized and whose nature is indicated by alphanumeric headings at the top of rows and columns. Each box is thus identified as the intersection of a row and a column.
[0009] Often, the Referenced Tabular Document associates with the lines, Identifiers appearing at the head of the line, to the right of the alphanumeric heading, but this is not always the case (example of pay slips, where, for the needs of the invention, such identifiers will have to be arbitrarily defined), whereas, for the columns, in practically all cases, arbitrary identifiers will have to be agreed upon. Sometimes, certain boxes are preceded by an individual identifier characterizing it, but this is too exceptional to base the process on these. We will therefore be led to assign to each box a box identifier composed of the line identifier concatenated with the column identifier of the box.
[0010] The subject of the invention is the analysis of these Referenced Tabular Documents to extract the values for the purposes of various calculations of a financial, accounting or administrative nature in order to meet specific needs (for example, calculations of Research tax credit from researchers' salary slips).
[0011] By processing the Referenced Tabular Documents thus completed with OCR tools, one can attempt to extract the identifier / value pairs from each completed box, these pairs being placed in a structured results file for later use, subject to possible OCR errors, OCR techniques never being able to guarantee the accuracy of the recognized characters, a reservation which must be lifted by the invented Process.
[0012] This exploitation can include, according to another example linked to financial results, the use of the results file by the risk assessment department of a bank: the identifier / amount pairs extracted from the accounts of a company in fact make it possible to immediately and automatically reconstitute the profile of financial risk associated with this company, by applying ratios and classic calculations in the field of finance.
[0013] Depending on the risk profile obtained, the banking establishment may, for example, decide whether or not to grant a bank loan to the company concerned.
[0014] There is a strong demand, particularly from banking establishments, for the data extracted from the Referenced Tabular Documents completed by companies to be reliable, i.e. to have been verified.
[0015] There are in fact a number of factors which can lead to the production of errors in the identifier / amount pairs extracted with the automated or semi-automated processing tools available in the state of the art.
[0016] Among these error factors are errors inherent in the use of OCR tools: for example, a figure damaged by printing in a box of the Referenced Tabular Document supplied by the company can give rise to a misinterpretation by the OCR tool (this is called “substitution”).
[0017] Another error factor is related to the variability of the Referenced Tabular Documents used by companies.
[0018] There are in fact Tabular Document Models designed by the French Administration - designated by "CERFA" forms which we have become accustomed to using as reference documents for the financial results of Companies ("Balance Sheets"), having the advantage of existing and being standardized and managed by the State.
[0019] However, CERFA forms are Tabular Document Models for tax purposes, not accounting purposes, and publishers of Accounting Balance Sheets are led to create "modified" versions to meet their "business" needs. They are thus led to create Referenced Tabular Documents which are mixed documents playing the dual role of Documents sent to the Tax Administration as a tax return, and of Accounting Document reflecting the financial development of the Company.
[0020] As a result, the Referenced Tabular Documents are inspired by the CERFA forms while adding additional financial information, in unregulated forms, introducing uncontrolled variability, the Tax Administration accepting these Referenced Tabular Documents as long as it finds what it requests, but the accounting additions complicate, by their uncontrolled variability, their automatic extraction of data for the purposes of financial analyses.
[0021] These Referenced Tabular Documents, modified from the Tabular Document Templates, may include additional comments, boxes, lines, and columns. For example, in Figure 3, a 4th column is often added to report the results of the previous year, which can be compared to those shown in the 3rd column, allowing the direction and value of the change from one year to the next to be measured. The headings of the last two columns can be given variable forms, such as "N", "Nl" or "|31 | 12 | 2021 |", "| 31 | 12 | 2020 |", closing dates represented in rakes and therefore subject to reading errors. Other disturbances of the same order may appear on other pages or in other places (inversion of columns N and Nl, etc.).In addition, there are terminological, abbreviative and synonymous variants in the alphanumeric headings of the lines and the column titles, and sometimes additional or inverted lines.
[0022] This variability can lead in particular to the association of an incorrect identifier with an amount when processing data resulting from OCR.
[0023] Thus, to meet the demand for reliability of structured results files transmitted to banking establishments, companies providing verification services must use different verification methods, both for the Identifiers and for the values of the amounts contained in these boxes.
[0024] Some of these methods, which are fully automated and rely in particular on artificial intelligence, make it possible to improve the reliability of data transmitted to banking establishments, without however reaching the 100% level of reliability required by these establishments.
[0025] Other, semi-automated methods allow a higher level of reliability to be achieved.
[0026] These methods involve verification personnel who ensure, on appropriate interface screens, that the amounts resulting from the OCR of the form completed by the company do not contain any errors, and that they are associated with the correct box identifiers.
[0027] This human verification thus ensures that the value of the amount and its allocation are correct.
[0028] These semi-automated methods assume by construction that the verifying persons have in view the amounts and the associated identifiers.
[0029] From this visibility, auditors can therefore have access to all of a company's accounting information.
[0030] However, there is a growing demand from banking authorities for verification operations on structured data such as identifier / amount to be carried out in a completely confidential manner, i.e. avoiding, as far as possible, any data leaks from a malicious verifier, based on the principle that in this verification they are led to see this data, whereas the accounting information of a company is considered a priori as belonging to it without any body being able to arrogate to itself the right to decide otherwise.
[0031] This growing demand stems in particular from directives issued by the EBA (European Bank Authority), which in particular leads to prohibiting EU banking establishments from subcontracting data entry to external companies likely, as mentioned above, to cause data leaks, requiring them to internalize this verification / correction, based on the principle that this internalization allows the Bank to better control the risks of sensitive data leaks. This confidentiality obligation also applies to initial Referenced Tabular Documents which may be Salary Slips containing data personal data which, since 2016, have been protected for European citizens by Confidentiality rules defined in a General Data Protection Regulation (GDPR, Regulation No. 2016 / 679 of the European Parliament and of the Council of April 27, 2016).
[0032] In the two previous examples, the numerical values to be extracted are monetary data (in € for the EU), subject to the dual imperative of rigorous accuracy and guaranteed confidentiality, incompatible with each other since accuracy can only be guaranteed through human verification, opening the way to possible data leaks.
[0033] The aim of the present invention is to provide a solution to this apparent contradiction through a novel input method which circumvents the problem by combining new concepts which contradict the prejudice of such incompatibility.
[0034] The services offered by accounting and financial data entry / verification companies currently available on the market are incompatible with these confidentiality obligations, particularly with regard to the intervention of verifiers who have clear access to the amounts and their associated identifiers appearing in the Documents completed by the companies.
[0035] The present invention relates to a method for semi-automatic extraction of data located in a referenced tabular document, whether paper or digital, comprising boxes, identifiers of these boxes and amounts filled in these boxes, comprising the steps of: a. sending an image file of the referenced tabular document filled in a secure core computer without a verifying person being able to have access to it, c. removing by means of this core computer all the amounts filled in the boxes of the referenced tabular document, leading to a so-called "defaced" referenced tabular document, e. filling in the correct identifiers inside the boxes concerned emptied of their amounts, f. having the core computer extract the images of each figure making up each of said amounts, g. associate with each of the said images the corresponding figure obtained by OCRing the completed referenced tabular document, h. have a verifying person check on a control interface whether the conversion of the figure image into an OCRed figure has been carried out without error, by reading and comparing isolated figures whose origin the verifying person cannot identify, i. if this is not the case, have the verifying person correct the erroneous value on the control interface, - j. updating by means of the core computer the values associated with said images, k. having the core computer scan through all the boxes of the referenced tabular document in which the verified identifiers were filled in at step e), and having the core computer associate with each of these identifiers the underlying amounts reconstituted by this core computer.
[0036] By "OCRizing the referenced tabular document in this kernel computer" is meant OCRizing the image file of the completed referenced tabular document.
[0037] The referenced tabular document can be the completed referenced tabular document.
[0038] The invention may also include the following features taken alone or in combination.
[0039] According to one characteristic, the method further comprises the following step: - b) OCRize this referenced tabular document in this kernel computer.
[0040] According to one characteristic, the method further comprises the following step: - d) store these amounts in memory in the kernel computer, so that they become underlying, i.e. not accessible to the verifying person.
[0041] According to one characteristic, the method comprises a step of generating the image file of the completed referenced tabular document.
[0042] Advantageously, this allows the image file to be OCRed later.
[0043] According to one characteristic, said digit images are grouped by shape proximity class.
[0044] By "shape proximity classes" we mean similarity classes in which numbers or characters with a similar shape are grouped.
[0045] According to one characteristic, the method comprises a step of comparing the number images with each other.
[0046] According to one feature, the comparison of the digit images is made according to the pixel composition of said digit images in order to identify similarities in the pixel compositions of the digit images.
[0047] The pixel composition of an image is the number of pixels and the arrangement of pixels within the image.
[0048] Advantageously, this makes it possible to identify substantially similar digit images in a simple manner in order to group images corresponding to the same digit into the same group.
[0049] According to one characteristic, the method comprises a step of distributing the number images into similarity classes based on a result of the comparison of their pixel composition, so that two images whose pixel compositions are substantially identical are included in the same similarity class.
[0050] A "similarity class" means a collection of objects, such as images, that share a common property. For example, a class that relates to the number "8" includes a plurality of images that represent the number "8".
[0051] By "substantially identical pixel compositions" is meant, for example, pixel compositions which differ in the number of pixels in the composition only by a number of pixels less than 10, and preferably less than 4, and preferably less than 2.
[0052] Advantageously, this reduces the number of digit images to be checked by a verifier.
[0053] According to one characteristic, shape proximity classes are determined by grouping images that differ from each other only by a few isolated pixels. By "a few isolated pixels" we mean, for example, at least 2 pixels, for example, 2 pixels, 3 pixels or 4 pixels.
[0054] According to one characteristic, the method comprises a step of modifying at least one parameter of a digit image before comparing said digit image with another digit image.
[0055] The "parameter of a figure image" can be chosen from: a dimension of the image such as a width or a length of the image, a number of pixels in the image, or an arrangement of the pixels in the image.
[0056] Advantageously, this makes the comparison between digit images more accurate.
[0057] According to one feature, the at least one parameter is a dimension of the digit image.
[0058] The dimension of an image in pixels is the size of the image, expressed for example in number of pixels. Pixels are the smallest elements that make up a digital image. They are arranged in rows and columns, and the dimension of an image is defined for example by the number of pixels across the width and the number of pixels across the height of the image.
[0059] Advantageously, this ensures that a digit image is properly classified into the class where it should be classified despite a difference in its pixel composition that would be due to a fortuitous variation in the size of the digit image.
[0060] According to one feature, the correct identifiers are filled in at step e) by a verifying person.
[0061] According to one feature, the correct identifiers (33, 35) are pre-filled at least in part in step e) by the core computer, and a verifying person verifies that the identifiers pre-filled by the core computer are correct.
[0062] According to one characteristic, step e) is carried out by transferring to the empty boxes of the disfigured referenced tabular document corresponding identifiers entered in the boxes of a so-called “augmented” tabular document model and freezing the choices decided for these identifiers for the model considered.
[0063] According to one feature, a so-called "revealing" file associated with the referenced tabular document, for example verified referenced, is stored in the core computer, in accordance with any one of the preceding claims, this revealing file comprising a file of physical location of the identifiers in said referenced tabular document and a file of the figures verified by the verifying person, this revealing file not containing any of the amounts appearing in the boxes of the completed referenced tabular document, and during a passage of the referenced tabular document, this passage is detected by means of the core computer (27) and the revealing file is loaded into the memory of the core computer, making it possible to regenerate identically the results verified by the verifying person.
[0064] According to one feature, the method comprises the following step: grouping the digit images extracted in d) into shape similarity classes by an algorithm having the property that two images grouped in the same class cannot represent two different digits.
[0065] According to one feature, the method comprises the following step: using the strong coincidence algorithm as a shape proximity algorithm having the same property.
[0066] According to one characteristic, step h) comprises having a verifying person check (49) on a control interface (31) whether the conversion of these images (51) in OCRed figures (53) was carried out without error, by visualizing these labeled figures presented on the screen in order of their recognized values on several lines of a table, not allowing the verifying person to reconstruct the original amounts.
[0067] The invention further relates to a method for semi-automatic extraction of data located in a paper or digital referenced tabular document, comprising boxes, identifiers of these boxes and amounts filled in these boxes, comprising the steps of: a. sending an image file of the referenced tabular document filled in a secure core computer without a verifying person being able to have access to it, c. removing by means of this core computer all the amounts filled in the boxes of the referenced tabular document, leading to a so-called "defaced" referenced tabular document, e. filling in the correct identifiers inside the boxes concerned emptied of their amounts, f. grouping the extracted digit images in d) in shape proximity class by an algorithm having the property that two images grouped in the same class cannot represent two different digits, h.have a verification person check on a control interface whether the conversion of these images into OCR figures has been carried out without error, by viewing these labeled figures presented on the screen in order of their recognized values on several lines of a table, not allowing the verification person to reconstruct the original amounts, i. if this is not the case, have the verification person correct the erroneous value on the control interface,. - j. updating by means of the core computer the values associated with said images, k. having the core computer scan through all the boxes of the referenced tabular document in which the verified identifiers were filled in at step e), and having the core computer associate with each of these identifiers the underlying amounts reconstituted by this core computer.
[0068] According to one feature, the method may comprise one or more of the features presented above, in any of their technically possible combinations.
[0069] According to one feature, the method may comprise: g. associating by means of the kernel computer with each of said images the corresponding figure obtained by OCRing the completed referenced tabular document, and / or using the strong coincidence algorithm as a shape proximity algorithm having the same property.
[0070] The invention further relates to a computer program comprising instructions for implementing the method described above when the program is executed by a processor.
[0071] The present invention therefore relates to a method for semi-automatic extraction of data located in a paper or digital referenced tabular document, comprising boxes, identifiers of these boxes and amounts filled in these boxes, comprising the steps of: a. sending an image file of the referenced tabular document filled in a secure core computer without a verifying person being able to have access to it, b. OCRing the referenced tabular document in this core computer, c. removing by means of this core computer all the amounts filled in the boxes of the referenced tabular document, leading to a so-called "defaced" referenced tabular document, d. storing these amounts in memory in the core computer, so that they become underlying, that is to say not accessible to the verifying person, e. filling in the correct identifiers inside the boxes concerned emptied of their amounts, f.having the core computer extract the images of each figure making up each of said amounts, g. associating, by means of the core computer, with each of said images the corresponding figure obtained by OCRing the completed referenced tabular document, h. having a verifying person check on a control interface (31) whether the conversion of the image of the figure into an OCRed figure a. was carried out without error, by reading and comparing isolated figures whose origin the verifying person cannot identify, i. if this is not the case, have the verifying person correct the erroneous value on the control interface, - j. updating by means of the core computer the values associated with said images, k. having the core computer scan through all the boxes of the referenced tabular document in which the verified identifiers were filled in at step e), and having the core computer associate with each of these identifiers the underlying amounts reconstituted by this core computer.
[0072] The present invention thus aims in particular to provide the means for verifying the structured data extracted from completed Referenced Tabular Documents, while respecting the confidentiality obligations described above, despite the intervention of a verifying person.
[0073] The present invention thus relates to a method for semi-automatic extraction of sensitive digital Data, such as financial or salary amounts in Referenced Tabular Documents, which are image files filled with printed characters, where the identification of such Data can be ensured automatically by a Core Computer, when the recognition of this data, by an OCR incorporated in the Core for example, requires verification by verifiers in order to ensure a "faultless" while guaranteeing Total Confidentiality, despite the fact that these verifiers are likely to transgress this requirement if they see them, this method comprising the steps of: a) creating a "defaced" image of the Referenced Tabular Document where all the sensitive amounts have been erased by the Core, this defaced Referenced Tabular Document being the only Document presented to the verifiers,making any data leakage impossible, but allowing them to carry out their work of identifying each Data item in view of their textual and structural environment, and to, b) have the Core automatically extract from the undisfigured image all the figures making up the said data and present their images “in separate pieces” and in scattered order to the verifying persons, who verify and enter their individual values without being able to reconstruct the amounts, the Core being responsible for reconstructing the said data upon receipt of these results.
[0074] According to other optional characteristics of the method according to the invention:
[0075] - the number images extracted from the amounts by the Core are grouped into Digit models through an equivalence relationship carried out by a shape proximity algorithm ensuring that two equivalent images cannot represent different digits, making it possible to reduce recognition by OCR to a representative of each equivalence class, the Core recomposing the amounts by re-extracting the images of the digits from each box of said completed Referenced Tabular Document, finding for each of them to which Digit Model it belongs, by comparison through the proximity algorithm, this process reducing the verification by a human operator of the results found by the OCR, the reduction being able to be a division by at least 100, the quantitative gain becoming qualitative at this level of reduction.
[0076] In order to be able to associate, by means of the core computer, with each of said images the corresponding figure obtained by OCRing the completed referenced tabular document, the core computer can implement a distribution of the figure images into similarity classes.
[0077] A kernel computer is, for example, a computer that only has an operating system kernel and in which there are no applications or drivers installed.
[0078] In order to generate digit models (or similarity classes) as indicated for example in step g), the generated digit images can be scanned by the kernel computer. The first image can be chosen as representative of a first similarity class.
[0079] Then a second image can be compared to the first image. By "second image compared to the first image" we mean that the pixel composition of the two images is compared, that is, the number of pixels that make up the image and the arrangement of the pixels (the pixel coordinates).
[0080] If the second image is similar to the first image, the second image can be classified into the same similarity class as the first image.
[0081] Otherwise, a second similarity class can be created, and the second image can be classified into the second similarity class and can become the representative image of this second similarity class.
[0082] To classify a new image, the new image can be compared to the first-class image and the second-class image to determine the class into which it can be classified. If the new image is not similar to any of the representative images in the existing classes, then a new similarity class can be created and the new image can be classified into this new similarity class. This similarity class classification process can continue until all the digit images are classified into similarity classes.
[0083] A particular example of an implementation of the shape proximity algorithm that can be executed in step g) for example and which can carry out this classification into similarity classes is the algorithm called "strong coincidence". Thus, the association by means of the kernel computer with each of said images of the corresponding figure obtained by OCRing the completed referenced tabular document amounts to classifying the figure images into similarity classes.
[0084] The strong coincidence algorithm is as follows:
[0085] Let A and B be the digit images tested by the strong coincidence algorithm.
[0086] A first step of the strong coincidence algorithm can be to check whether A and B have the same dimension, for example the same width or the same height, or if the difference in pixel composition of images A and B is less than a difference limit value. By "difference limit value" we mean that images A and B can be considered similar if the difference in number of pixels between images A and B is less than or equal to four pixels for example.
[0087] If this is not the case, then the output of the algorithm may be that images A and B are not in strong coincidence, that is, the images are not significantly similar.
[0088] A second step of the strong coincidence algorithm may consist of modifying a parameter of image A or image B, for example a dimension of image A or a dimension of image B. This is to ensure that a fortuitous difference due to a variation in the size of the image following OCR is taken into account in the comparison.
[0089] To this end, we can create an image A0 by placing image A in a white frame one pixel thick (image A0 has its width and height increased by 2 pixels compared to those of A), and we can create an image B0 by placing image B in a white frame one pixel thick (image B0 has its width and height increased by 2 pixels compared to those of B). Then, we can modify image A0 by blackening the 8 pixels surrounding any black pixel in the image. Image A0 becomes An and we can modify image B0 by blackening the 8 pixels surrounding any black pixel in the image. Image B0 becomes Bn.
[0090] The principle of verification can consist in comparing image A in its two representations A0 and An to image B in its two representations B0 and Bn, by superimposing image B above image A with a possible shift of 1 pixel along a first dimension x and / or along a second dimension y, and by finding if there is a positioning in which: - no black pixel of B0 is superimposed on a white pixel of An, - no black pixel of A0 is superimposed on a white pixel of Bn.
[0091] This can be as simple as checking whether there is at least one positioning of B relative to A where the outer and inner contours of images A and B do not differ by more than 1 pixel from each other. If this is not the case, the output of the algorithm can be that the two images A and B are not in strong coincidence, in other words that images A and B are not significantly similar.
[0092] The strong coincidence relation can be an equivalence relation in the mathematical sense of the term, in other words if A coincides with B, then B coincides with A (symmetry) and if A coincides with B and if B coincides with C then A coincides with C (transitivity).
[0093] Based on this strong coincidence algorithm, digit images can thus be grouped into similarity classes.
[0094] Advantageously, based on the strong coincidence algorithm, two digit images belonging to the same similarity class necessarily correspond to the same digit, even if we do not yet know which one.
[0095] Advantageously, the strong coincidence algorithm ensures that images whose pixel compositions are recognized as similar (and therefore the images are in strong coincidence) cannot represent two different numbers.
[0096] Advantageously, thanks to the strong coincidence algorithm, the described method can meet obligations of the financial sector, in particular an obligation to guarantee the accuracy of the financial amounts entered, and guarantee the confidentiality of companies' financial data, for example against any possibility of data leakage by human operators.
[0097] Advantageously, the method can be limited to the recognition of each similarity class found through an internal OCR of the core computer applied to one of the representative images of the class, instead of having to recognize for example 3000 or 4000 figures contained in a document.
[0098] Advantageously, the strong coincidence algorithm can limit verification by a human operator of the accuracy of recognitions to a reduced number of "models" (a digit model is the association of the image of a class and its digit recognized by an internal OCR). The human verification time can be divided by a factor of at least 100, for example, and the intervention of the verification / correction operators is limited to around fifty models to be verified. At this level of reduction, the improvement is no longer quantitative but qualitative.
[0099] The strong coincidence algorithm can also be used to resolve special cases such as the accidental presence of an isolated white pixel that needs to be eliminated, or a thin black line that breaks in the middle of a digitized image that needs to be repaired. The isolated white pixel is blackened on the image An by extending one of the black pixels that touch it; the broken line can be repaired by the strong coincidence algorithm when the break is 1 or 2 pixels (the two black pixels that face each other will blacken these two white pixels, each of which will touch one of these black pixels. Breaks not exceeding 2 pixels are thus repaired.)
[0100] If the strong coincidence algorithm does not detect the similarity of two images of similar digits, one more model will have to be checked by a human operator. In any case, the number of models to be checked by the human operator remains small compared to the thousands of digits that would otherwise have to be checked one by one.
[0101] When the process is applied to Referenced Tabular Documents falling under a Tabular Document Model: a) this Tabular Document Model is associated with a set of identifiers for each box that can be filled with numerical values of amounts to be extracted, these identifiers being able to be defined semi-arbitrarily by the application manager, drawing inspiration from those possibly appearing on the Tabular Document Model, for table rows for example, b) a version of the Tabular Document Model is created, called the Augmented Model, where each box concerned is filled with the identifier assigned to it, c) the process then consists of transferring the corresponding Identifiers of the Augmented Model to the blanked boxes of the defaced Referenced Tabular Document, either automatically by the Core if its Artificial Intelligence algorithms allow it to recognize the identifiers concerned by the structural and textual environment of the boxes, or by transfer action from the Augmented Model carried out by the human intelligence of the verifying person, the final results being generated by the Core in the form of a list of pairs<identifiant, valeur> .
[0102] - an intelligent interface is used allowing the verifying persons to transfer the identifiers of the Augmented Model to the defaced Referenced Tabular Document in order to complete / correct any omissions and errors in the Core, in which an “operator” station is used consisting of two adjoining screens controlled by the same mouse passing from one to the other, allowing “Pick and Drop” operations of the Identifiers of the Augmented Model to the defaced Referenced Tabular Document;
[0103] - the verifiers are given optimized means of action according to the following specifications: a) transfer of an identifier from a box of the Augmented Model by clicking on this box, then on the location of the defaced Referenced Tabular Document where it must be placed, b) transfer of a rectangular block of boxes from the Augmented Model to the defaced Referenced Tabular Document, by clicking on the upper left corner, then on the lower right corner of the rectangle to be transferred, then by clicking on the upper left corner of the defaced Referenced Tabular Document where the verifier wishes to place the identifiers of the rectangular block, the interface driver completing the transfer by updating the boxes appearing on the defaced Referenced Tabular Document outside the transferred rectangle, likely to be filled by a remainder of identifiers otherwise transferred by deleting them, in order to respect the rule of non-ubiquity of Identifiers;
[0104] - said Core (27) archives, under an indexing linked to the Tabular Document Referenced (for example, Siren of the Company and year, or registration number and date of a pay slip) a Revealing File containing two files: a) the file of the defaced Referenced Tabular Document filled in by the identifiers as stated in the steps described above, b) the file of the Models of figures labeled and verified by a verifying person as stated above, these two files can be archived without transgressing the obligation of confidentiality of the amounts since they cannot allow them to be reconstituted on their own;
[0105] - when the Kernel detects, through its indexing, that a Tabular Document Referenced already processed during a previous pass, it reactivates the results obtained during this first pass as archived in its Developer File and Analyses the results as it had done during this first pass;
[0106] - the said Revealing File of the completed and archived Referenced Tabular Documents can be read to make selections of completed Referenced Tabular Documents through conditional expressions expressed in algebraic terms designating the positions involved through their identifiers of the Augmented Model and by browsing the original completed Referenced Tabular Documents archived in an ultra-protected Database, the ephemeral meeting of the Revealing File and the undisfigured Referenced Tabular Document making it possible to regenerate the numerical amounts and to check whether the conditional expressions are satisfied, then to erase the regenerated amounts, making a malicious intrusion unlikely in this very short period of time;
[0107] - the invention also proposes the possibility for the verifying persons to verify whether arithmetic relationships are satisfied between certain boxes of the Referenced Tabular Documents, without seeing the amounts filled in these boxes, by means of a Virtual Calculator which allows the Core to verify, through the underlying amounts whether the arithmetic relationships are satisfied and to inform the verifying persons of the result.
[0108] The present invention also relates to a system for implementing the method in accordance with the above, comprising: a) at least one secure core (27) equipped with means making it possible in particular to disfigure said completed Referenced Tabular Document, to extract from filled boxes the images of figures making up the digital amounts, to produce elementary images (51) of said data, to OCR said data, to store these elementary images and these OCRed data, b) at least one control interface (31) communicating with said secure core equipped with means making it possible to enter the identifiers on the screen, to compare said OCRed data (53) with said elementary images (51), to make the necessary corrections, and to return said identifiers and said corrections (55) to said secure core (27).
[0109] The present invention also relates to a computer program comprising instructions for implementing the method according to the above when the program is executed by a processor.
[0110] Other characteristics and advantages of the invention will emerge from reading the description which follows, with reference to the appended figures, which illustrate: [Fig. 1]: in the form of a flowchart the main steps of the method according to the invention; [Fig. 2]: an example of a tabular document model, CERFA type; [Fig. 3]: a completed Referenced Tabular Document, which is a completed modified version of the form in Figure 2; [Fig. 4]: a disfigured version of the Referenced Tabular Document of Figure 3; [Fig. 5]: an identified version of the Referenced Tabular Document of Figure 4; [Fig. 6]: an augmented version of the Tabular Document Model of Figure 2; [Fig. 7]: a version pre-identified by the Referenced Tabular Document Core of Figure 4; [Fig. 8]: an example of image and OCR number lines to compare; [Fig. 9]: an example of a Virtual Calculator that can be implemented within the framework of the invention.
[0111] For clarity, identical or similar elements are identified by identical or similar reference signs throughout the figures.
[0112] Numerical references in parentheses refer to the flowchart in Figure 1, and numerical references not in parentheses refer to Figures 2 to 8.
[0113] For the sake of clarity, it is also important to define precisely the following terms used in the description and claims of the present invention:
[0114] - OCR: conversion of an image of a number, letter, string of numbers or letters, into ASCII-type coded characters;
[0115] - OCRed number, letter, document: number, letter, document resulting from OCRing;
[0116] - Tabular document model: Reference document of a Document Referenced Tabular Document in a material (paper) or digital (file) form, such as a CERFA form from the Administration, showing the Reference tables and their cellular structures empty of all values, indicating the cellular structure that the Referenced Tabular Document must follow in order to be filled in by action of the Core and / or human data entry Operators;
[0117] - Referenced Tabular Document: a form inspired by the Tabular Document Model, but which may contain variations, such as added text, lines or additional columns, and lexicographic, synonymous and / or abbreviated versions of the alphanumeric headings on the left, edited and completed by the Companies through their chosen management databases;
[0118] - box: materialization, often in the form of a rectangle, of the field to be filled in on the Tabular Document Model or on the Referenced Tabular Document;
[0119] - identifier: code, in the form of several letters or numbers associated with a box, and providing information on the nature of the information contained by the box;
[0120] - Augmented Model: Tabular document model whose boxes have been filled in with corresponding identifiers fixing the choice decided for these identifiers for the Model considered;
[0121] - Completed Referenced Tabular Document: Referenced Tabular Document whose boxes have been filled with amounts by the company's management software and are therefore readable by any human being;
[0122] - amount: monetary or salary value (number) entered in a box of the Completed Referenced Tabular Document;
[0123] - figure: each of the figures making up the amount;
[0124] - verification person: human being responsible for verifying the accuracy of the amounts extracted from the completed Referenced Tabular Document, and the correct assignment of identifiers to these amounts;
[0125] - core: secure software and / or hardware space, i.e. to which, in particular, the verifying persons do not have access;
[0126] - underlying: not accessible to the verifying person, but stored in the core in the same location as the completed Referenced Tabular Document;
[0127] - Disfigured Referenced Tabular Document: Referenced Tabular Document whose boxes have been cleaned of their amounts by the core, these being stored in underlying manner in a structured file associated with this Referenced Tabular Document defaced and not visible to a verifying person;
[0128] - Identified Referenced Tabular Document: Referenced Tabular Document containing identifiers added by a verifying person and / or by the core;
[0129] - control interface: software and / or hardware assembly, managing the interactions of the verifying person with the data verification screen and the feedback of information to the Core according to the invention;
[0130] - end customer: ordering party for whom the data verified using the method according to the invention are intended;
[0131] - result file: file containing the verified structured data of the couples type <identifiant montant>, usable by the end customer.
[0132] We now refer to Figure 2, which shows an example of a CERFA-type tabular document template for a company's balance sheet declaration, provided by the Tax Administration.
[0133] In this illustrative and non-limiting example, the Tabular Document Template is a form for reporting a company's balance sheet.
[0134] Of course, this Tabular Document Template could be used for any other topic, such as a payroll template.
[0135] Referring therefore to this Tabular Document Model provided as an example, we can see that it includes zones 1, 3... for identifying the company and the accounting year, as well as a certain number of textual labels 5, 7... relating in this case to the assets of the company in question.
[0136] This Tabular Document Model also includes boxes 9, 11... arranged opposite the text labels 5, 7..., organized in the form of rows and columns, and intended to be completed by the company's accounting departments.
[0137] At least some of these boxes are identified by codes with several letters or numbers 13, 15..., the meaning of which is indicated by the textual wording associated with these boxes, and making it possible to identify the nature of the amount filled in the box concerned.
[0138] In Figure 3, we can see a modified version of this Tabular Document Model. This modified version of the basic Tabular Document Model, designated as Referenced Tabular Document and provided to the company's accounting firm by the publishers of Accounting documents, differs essentially from the Tabular Document Model of Figure 2, in that it includes an additional column 17 allowing the amounts relating to year N1 to be displayed.
[0139] In the specific case of Figure 3, the Referenced Tabular Document was completed with the amounts 19..., 21... relating respectively to the financial years 2022 (column 3 - reference 23) and 2021 (column 4 - reference 17).
[0140] The problem is therefore the following: the completed Referenced Tabular Document in Figure 3 will be subject to automatic data extraction, for example using an OCR tool, and the verifying person would be responsible for verifying the accuracy of the amounts in the completed OCRed Referenced Tabular Document, and their correct correspondence with the identifiers - without this verifying person being able to access these amounts at any time, in order to comply with the confidentiality obligations described above.
[0141] For this, an image file of the completed Referenced Tabular Document is generated. This allows the image file to be OCRed later.
[0142] Then the image file of the completed Referenced Tabular Document is sent (25) into the kernel (27) of the system according to the invention without the verifying person being able to access it.
[0143] The image file of the completed Referenced Tabular Document is then OCRed (29), and all the amounts filled in the boxes of the image file of the completed Referenced Tabular Document are removed (30), and stored in memory in the kernel (27): they therefore become underlying.
[0144] The first task of the verifying person will be to check that the identifiers 13, 15... are located in the right place, that is to say each one next to a box 9, 11... whose textual label 5, 7... corresponds to this identifier.
[0145] According to a first variant, the core (27) returns to the control interface (31) of the verifying person the disfigured Referenced Tabular Document visible in figure 4, that is to say the Referenced Tabular Document with empty boxes and others filled by the Core.
[0146] The task of the verifying person is to fill in (32) the correct identifiers inside the relevant boxes: this filling can confirm the identifier forming part of the distorted Referenced Tabular Document filled in by the Core (in which case, it leaves things as they are), or correct it if there is an error.
[0147] To help with this verification task, the verifier can use the Augmented Model in Figure 6: the boxes in this model have been filled in once and for all with identifiers that allow the nature of these boxes to be uniquely identified (additional identifiers of the AC_1 type are created for boxes without official identifiers, as is the case for the boxes in the 3rd column of the Tabular Document Model in Figure 6).
[0148] The reviewer can then perform “click and drop” operations on boxes or groups of boxes in the Tabular Document Model to place them in the appropriate areas of the defaced Referenced Tabular Document.
[0149] At the end of this intervention by the verifying person, we obtain the Referenced Tabular Document identified in figure 5, the relevant boxes of which have been filled in with the identifiers 33, 35... chosen by the verifying person.
[0150] Note, as can be seen in this figure 5, that the verifying person left the last column 17 empty: this column 17 relating to the financial year Nl was in fact added to the Tabular Document Referenced in relation to the Model of tabular document and is not part of the requested data. This example assumes that the Kernel has not found anything in preliminary.
[0151] It is therefore important to ignore the amounts appearing in this additional column, which will be done automatically by the Core when editing the results since none of these boxes are filled with an identifier.
[0152] If necessary, the modifications to the identifiers made by the verifying person are returned to the kernel (27) which is updated, ensuring perfect consistency with the image on the screen.
[0153] According to another possible variant, the identifiers 33, 35... can be at least partly pre-filled by the core (27) in the boxes of the disfigured Referenced Tabular Document: for example, this Referenced Tabular Document thus pre-filled can appear on the control interface (31) of the verifying person in the form visible in figure 7.
[0154] The work of the verifying person will then consist of verifying that the identifiers pre-filled by the core (27) in the boxes are correct.
[0155] In the example shown in Figure 7, we can also see that identifiers 41, 43... were placed by the kernel in the last column 17 of the pre-filled Referenced Tabular Document, relating to the financial year Nl, whereas they should have been placed in the penultimate column, relating to the financial year N.
[0156] An additional task for the verifier will therefore consist of selecting the last column on the Augmented Model in Figure 6 by clicking on its upper left corner (click 1) then on its lower right corner (click 2) then clicking, on the Referenced Tabular Document in Figure 7, on the upper left corner of the 3rd column (click 3), this last click causing the transfer, by the interface driver, of the 3rd column of the Augmented Model in Figure 6 to the 3rd column of the Referenced Tabular Document in Figure 7.
[0157] The interface driver will then have to correct the identifiers transferred elsewhere from the Augmented Model which therefore remain in the 4th column and must therefore be deleted, applying the rule of non-ubiquity of identifiers.
[0158] More generally, during such an identifier transfer from the Augmented Model to the Referenced Tabular Model, the interface driver must ensure that the repositioned identifiers do not remain visible in their position before transfer.
[0159] Indeed, the transferred copy of the 3rd column will automatically erase any identifiers that appeared in this column, but a remainder of these same identifiers could remain elsewhere, for example in the 4th column. In this case, the identifiers that were incorrect in the 4th column remain in place and will have to be erased by the pilot.
[0160] Another example: two consecutive lines are reversed; in this case, for example, we will take the lower line on the Augmented Model and place it on the upper line; the lower line will then be erased. We will then have to take the upper line on the Augmented Model and transfer it to the lower line.
[0161] The second verification task by a verification person will consist of ensuring the recognition of the amounts in the sensitive boxes of the completed Referenced Tabular Document to which only the Core (25) has access, the verification person having to carry out this work without ever seeing them. Two methods are successively proposed, the first Basic, the second Optimal.
[0162] Basic Method: Recognition of numbers in anonymously dispersed Spare Parts. To carry out this verification, the Core (27) retrieves the images of each number making up each amount located in each box of the completed Referenced Tabular Document, and associates (45) with each of these images the corresponding number obtained by OCRing the completed Referenced Tabular Document.
[0163] All couples<image du chiffre / valeur océrisée> are then grouped in ascending order of recognized values and sent (47) to the control interface (31), where the verifying person can check (49) whether the conversion of the image of the digit to OCR number has been performed without error and if this is not the case, it corrects the erroneous value on the screen. In doing so, it cannot deduce any information about the original position of these digits in this or that box, preserving the required Confidentiality.
[0164] Optimal Method: Grouping into Digit Models. Advantageously, to limit the number of checks to be carried out, digit images can be grouped by shape proximity classes, all images in the same class being distinguished from each other by very small differences, for example a difference on a few isolated pixels. By "a few isolated pixels" we mean for example at least 2 pixels, for example 2 pixels, 3 pixels or 4 pixels.
[0165] Clustering into digit patterns (or similarity classes) can be done based on the strong coincidence algorithm described previously.
[0166] The data extraction method that uses the strong coincidence algorithm may therefore comprise a step of comparing the digit images to each other, possibly as a function of the pixel composition of said digit images in order to identify similarities in the pixel compositions of the digit images, as well as a step of distributing the digit images into similarity classes as a function of a result of the comparison of their pixel composition, so that two images whose pixel compositions are substantially identical are included in the same similarity class.
[0167] Advantageously, this makes it possible to identify substantially similar digit images in a simple manner in order to group images corresponding to the same digit into the same similarity class.
[0168] However, due to OCR, a generated digit image may have unintended pixel differences compared to another similar digit image. To overcome this problem, the extraction method may comprise a step of modifying at least one parameter of a digit image before comparing said digit image with another digit image. The at least one parameter may be one dimension of the digit image. Advantageously, this makes the comparison between digit images more accurate. Advantageously, this ensures that a digit image is correctly classified into the class in which it should be classified despite a difference in its pixel composition that would be due to a chance variation in the size of the digit image.
[0169] The inventor of the present Method has developed a proprietary algorithm called "strong coincidence": two images that strongly coincide are placed in the same class, and the property of the strong coincidence relation is that it is an equivalence relation in the mathematical sense of the term: two images linked by the Strong Coincidence Relation always correspond to the same number, even if the value of this number is not yet known. Each class then falls under the mathematical concept of Model, and each Model is then OCRed on any image representing it.
[0170] By grouping these images into Number Models, we can substantially reduce the number of verifications of the pairs<image du chiffre / chiffre océrisé> , to be carried out by the verifying person: we typically go from a verification of several thousand to only a few dozen couples to verify. At this level, the quantitative gain can be considered qualitative.
[0171] As an illustration, Figure 8 shows two lines of numbers representative of the Number Patterns defined by the Strong Coincidence algorithm to be verified by the verifier: the upper line 51 reproduces the Number Image Patterns, through their representatives, and the lower line 53 includes the numbers obtained by OCRing each of the images in the upper line: the work of the verifier therefore consists of visually verifying that the numbers in the lower line correspond to the images in the upper line.
[0172] If a number in the lower line 53 does not correspond to the image of the number in the upper line 51, as seen at 55 in Figure 8 (number 9 in the lower line and image of a 4 in the upper line), the verifying person corrects the number in the lower line with a 4 on their control interface.
[0173] This correction is transmitted to the kernel (27), which will then update the values associated with the corrected Digit Models.
[0174] The amounts of each box can then be reconstructed automatically and without error by the Kernel by extracting the real images of the numbers in the box, successively "recognizing" them by comparison with the images of each Number Pattern through the Strong Coincidence algorithm, retaining the numerical value associated with the Coinciding Number Pattern.
[0175] The verifying person then carried out his task ensuring the error-free extraction of the values from the sensitive boxes without ever having access to these amounts or being able to reconstitute them in any way.
[0176] Once the two verification steps described above have been carried out, the core (27) goes through all the boxes of the Referenced Tabular Document in which the verified identifiers have been filled in, and associates (57) with each of these identifiers the underlying amounts reconstituted by the Core (27) as mentioned above.
[0177] We thus obtain a series of couples <identifiant montant>verified, which can be structured within a verified results file (59), edited in the format requested by the end customer, i.e. that of the Financial Analysis Software that he uses, and which generally does not depend on him, but on the Financial Analysis professional who is the publisher / designer of this software. As such, this Publisher can choose two modes of results entry formats: a mode of type <identifiant montant>, the identifiers being specific to it, and in this case, the Service Provider must have a dictionary translating its own Identifiers into those requested by the Financial Analysis software, a “Mapping” mode giving the order of appearance of the requested numerical results on a single-value file, the translation being able to be carried out from a grid indicating the list of the Service Provider’s identifiers corresponding to each position of the Financial Analysis software Publisher.
[0178] We therefore understand that the invention makes it possible to verify the reliability of the couples <identifiant montant>without ever accessing information enabling these amounts to be reconstituted: the control interface (31) allows this verification to be carried out indirectly, by reading and comparing isolated figures whose origin the verifying person cannot identify.
[0179] At the end of the two verification steps described above, the core (27) keeps in the service provider's database, and for the processed Referenced Tabular Document, a Revealing File associated with this Document, composed of: a file of physical location of the identifiers in the boxes of the Referenced Tabular Document, and the file of the Models of figures verified by the verifying person, this Revealing File being indexed in a unique manner by any appropriate means such as the combination of a company registration number (SI REN / SI RET) with the date of publication of the verified results file, or, more anonymously, by a chronological number for which the Core (27) has the analysis key.
[0180] This Disclosure File does not contain any of the amounts listed in the boxes on the completed form, which reinforces the request for data confidentiality since none of these amounts are archived.
[0181] When the Service Provider receives a new document to process, it looks in its database to see if a previous passage of the same document has been recorded, which the Core (27) detects through the indexing of the Revealing Files already archived, this allows the aforementioned Revealing File to be immediately loaded into the memory of the core (27), and thus, by exhuming it, to find the results transmitted at the time by the operators and to reissue the final results identical to what it had done at the time in complete confidentiality, this: without the amounts located in the boxes of the completed Referenced Tabular Document having been stored anywhere in the system, and without requiring any new verification operation on the part of the verifying person.
[0182] By way of illustration, if, between the initial and subsequent passages of the Referenced Tabular Document completed in the system according to the invention, certain amounts had been modified, the newly generated Results file would take this change into account.
[0183] Naturally, the invention is described in the foregoing with the aid of examples. It is understood that the person skilled in the art is able to carry out different variant embodiments of the invention without departing from the scope of the invention.
[0184] This is how, for example, the invention could be applied to the verification of data other than pairs <identifiant montant>: this could involve the verification of triplets, quadruplets, n-uplets of encrypted or lettered data, or even graphic elements.
[0185] This is also how we can predict that the Kernel ignores alphanumeric characters other than the 10 digits in upright / italic / normal / bold characters, and in particular letters and special characters in upper / lower case upright / italic / normal / bold, thus reducing the number of characters to be checked from at least 400 to 40.
[0186] This is also how human intervention could be replaced for the second step described above of verifying the numbers associated with each Number Model, by sending images of these Number Models, ordered in ascending order of the values found for these numbers by the internal OCR to a specialized external service provider with very powerful OCR resources, such as , replacing the human operator checking these values with total automation.
[0187] This is also how we could provide that the core (27) carries out additional checks by carrying out operations on the amounts underlying the identifiers entered in the completed Referenced Tabular Document.
[0188] These operations can be for example of the type XX1_2 + XX2_2 + XX3_2 + ... XXn_2 = TOT_2 (example for column _2, where XXi_2 designates the identifier of row i), or YY2 1 -YY2 2 = YY2 3 (example for row 2, where YY2_i designates the identifier of column i).
[0189] This is also how we could plan to use the Revealing Files stored in the core to carry out Company Selections according to certain criteria expressed in the form of conditional algebraic expressions using the Identifiers of the positions involved, and a hyper-protected archive containing the history of the initial Referenced Tabular Documents processed, allowing the amounts of the boxes to be furtively re-generated as indicated in
[0095] for a second pass, the time to use the regenerated data for the programmed selection.
[0190] Example: search for companies whose turnover (identifier XX) is between such and such amount: the time span during which the amounts regenerated for each processed Referenced Tabular Document could be hacked from the outside is reduced to a few milliseconds, making this risk unlikely.
[0191] This is also how the level of security of the invention could be further increased, by providing for the temporary connection, in another secure core, of the completed Referenced Tabular Document and the associated Revealing File, only during the very short time during which the Revealing File is queried to carry out a particular search, such as a search by turnover.
[0192] This is also how peripheral tools can be added to the method of the invention, proceeding from the same general spirit of hiding data from verifying persons.
[0193] An example of such a tool, which proves to be very practical in use, is a Virtual Calculator (which could also be called a Disfigured Calculator), allowing to the verifying persons to check whether arithmetic relationships between certain boxes of the Referenced Tabular Documents are satisfied, without seeing the amounts filled in these boxes.
[0194] The Virtual Calculator fits in a line placed at the bottom of the page of the Referenced Tabular Documents verification screen, like a footnote and remains fixed, the display screen reserved for the processed Referenced Tabular Document being limited to the frame of the full page, which refreshes each time you move to the next page of the Document, and can also have a vertical shuttle if necessary.
[0195] The Virtual Calculator consists of three rectangular boxes that follow one another on the line, interspersed with two multiple-choice symbols: arithmetic signs and relationship signs between amounts. A square box containing the symbol "OK" ends the line. The three rectangular boxes are filled with a background color specific to each of them: for example, blue for the first, green for the second, pink for the third, as shown in Figure 9.
[0196] The principle of the Virtual Calculator is that the verifying person indicates the boxes on the displayed page of the Referenced Tabular Document whose amount must be transferred into each of the colored boxes. The differentiated coloring of the boxes is used to display on the Document the box chosen to transfer its value: it is colored in the color of the receiving box. Indeed, the boxes of the Referenced Tabular Document are not necessarily filled with an identifier, quite the contrary, otherwise, there would be no need to remove an ambiguity.
[0197] When the verifying person has designated the boxes corresponding to each color, and has chosen the arithmetic signs and those of the relationship between amounts, he clicks on the square box "OK", which causes the location of the colored boxes of the Referenced Tabular Document, and the chosen arithmetic and relational signs to be sent to the Core, Core which: extracts the amounts underlying the colored boxes, performs the requested arithmetic calculations, checks the numerical relationship between amounts, with a tolerance of two monetary units (€ for example), returns its result to the interface driver, upon receipt, the interface driver fills in the OK box: in green if the result is positive, in red if it is negative.
[0198] The verifying person only has to make their choices based on the result and continue their verification.
[0199] EMBODIMENT: referring to figure 9, reference D designates the Referenced Tabular Document, and C the Virtual Calculator.
[0200] Initially, the three boxes of the Virtual Calculator are set to blank.
[0201] Then : Click 1: the box fills with blue; Click 2: the box clicked on the Document fills with blue; Click 3: the box fills with green; Click 4: the box clicked on the Document fills with green; Click 5: the box fills with pink; Click 6: the box clicked on the Document fills with pink; Click 7: choice of arithmetic operator; Click 8: choice of logical relationship; Click 9: the “OK” box fills with green: sending the request to the Core.
[0202] Return of results from the Kernel to the Interface Driver: if OK, the "OK" box fills with yellow; if not OK, the "OK" box fills with red.< / identifiant> < / identifiant> < / identifiant> < / identifiant> < / identifiant>
Claims
CLAIMS 1. Method for semi-automatic extraction of data located in a paper or digital referenced tabular document comprising boxes (9, 11), identifiers (13, 15) of these boxes (9, 11) and amounts (19, 21) filled in these boxes (9, 11), comprising the steps of: a) sending (25) the image file of the filled referenced tabular document into a secure core computer (27) without a verifying person being able to have access to it, c) removing (30) by means of this core computer (27) all the amounts (19, 21) filled in the boxes (9, 11) of the referenced tabular document, leading to a so-called "defaced" referenced tabular document, e) entering (32) the correct identifiers (33, 35) inside the boxes concerned emptied of their amounts, f) having the core computer (27) extract the images of each figure making up each of the said amounts,g) associating (45) by means of the kernel computer with each of said images (51) the corresponding figure (53) obtained by OCRing the completed referenced tabular document, h) having a verifying person check (49) on a control interface (31) whether the conversion of the image of the figure (51) into an OCRed figure (53) has been carried out without error, by reading and comparing isolated figures whose origin the verifying person cannot identify, i) if this is not the case, having the verifying person correct the erroneous value (55) on the control interface (31), - j) updating by means of the core computer (27) the values associated with said images (51), k) having the core computer (27) scan through all the boxes of the referenced tabular document in which the verified identifiers (33, 35) were filled in step e), and having the core computer (27) associate (57) with each of these identifiers (33, 35) the underlying amounts reconstituted by this core computer (27).
2. The method of claim 1 further comprising the following step: - b) OCR (29) in this kernel computer (27) this referenced tabular document. Method according to one of the preceding claims further comprising the following step: - d) storing these amounts (19, 21) in memory in the kernel computer (27), so that they become underlying, i.e. not accessible to the verifying person. Method according to one of the preceding claims, comprising a step of generating the image file of the completed referenced tabular document. Method according to one of the preceding claims, comprising a step of comparing the digit images with each other. Method according to claim 5, wherein the comparison of the digit images is done according to the pixel composition of said digit images in order to identify similarities in the pixel compositions of the digit images.Method according to claim 6, comprising a step of distributing the digit images into similarity classes according to a result of the comparison of their pixel composition, so that two images whose pixel compositions are substantially identical are included in the same similarity class. Method according to one of claims 5 to 7 comprising a step of modifying at least one parameter of a digit image before comparing said digit image with another digit image. Method according to claim 8 in which the at least one parameter is a dimension of the digit image. Method according to any one of the preceding claims, in which said digit images are grouped by shape proximity classes, and in which the shape proximity classes are determined by grouping the images which are distinguished from each other only by a few isolated pixels. Method according to any one of the preceding claims, in which the correct identifiers (33, 35) are filled in at step e) by a verifying person. Method according to any one of the preceding claims, in which the correct identifiers (33, 35) are pre-filled at least in part at step e) by the core computer (27), and a verifying person checks that the identifiers pre-filled by the core computer (27) are correct. Method according to any one of the preceding claims, in which step e) is carried out by transferring to the empty boxes of the defaced referenced tabular document corresponding identifiers entered in the boxes of a so-called "augmented" tabular document model and freezing the choices decided for these identifiers for the model in question.Method according to any one of the preceding claims, in which a so-called "revealing" file associated with the referenced tabular document is stored in the core computer in accordance with any one of the preceding claims, this revealing file comprising a file of physical location of the identifiers in said referenced tabular document and a file of the figures verified by the verifying person, this revealing file not containing any of the amounts appearing in the boxes of the completed referenced tabular document, and during a passage of the referenced tabular document, this passage is detected by means of the core computer (27) and the revealing file is loaded into the memory of the core computer (27) making it possible to regenerate identically the results verified by the verifying person.Computer program comprising instructions for implementing the method according to any one of claims 1 to 14 when the program is executed by a processor.