Method for segmenting and classifying an image including a structured document and corresponding terminal
Neural networks facilitate rapid and accurate segmentation and classification of stacked and distorted structured documents by encoding images into key vectors and applying geometric transformations, addressing the limitations of existing methods.
Patent Information
- Application Number
- PCT/EP2025/077239
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-10-08
- Filing Date
- 2025-09-23
- Publication Date
- 2026-04-16
AI Technical Summary
Existing methods for automated reading of structured documents like game slips are not fast enough or robust enough to handle multiple documents stacked together, geometric deformations, and ambient light interference, making real-time processing challenging.
A method using neural networks to encode images into key vectors, apply geometric transformations, and identify predefined graphic patterns, enabling rapid and accurate segmentation and classification of documents, even when partially overlapped or distorted.
Enables predictable and real-time analysis of stacked and distorted documents, ensuring quick processing without user delay, even with ambient light interference and geometric deformations.
Smart Images

Figure EP2025077239_16042026_PF_FP_ABST
Abstract
Description
Description Title of the invention: Method for segmenting and classifying an image including a corresponding structured and terminal document
[0001] The invention relates to the field of identification of documents containing structured information, in particular for documents such as game slips.
[0002] A game slip is an example of a document containing structured information. Other well-known structured documents include official forms or multiple-choice questionnaires. A game slip is printed on paper and has a predefined format and design. The structured information is typically presented as fields to be filled in, such as checkboxes.
[0003] Automated reading of betting slips can be performed using a terminal equipped with a camera. A user may need to stack several betting slips in succession for automated reading. This reading operation must be quick, requiring only the placement of a new slip under a camera, without precise positioning. Other slips can thus be visible beyond the edges of the last slip placed. This reading operation must also be tolerant of geometric deformations of the slip, such as folding or creasing. Furthermore, this reading operation must be robust against ambient light interference. Finally, this reading operation must be compatible with standard imaging devices such as cameras.
[0004] According to known technical processes, for example in document FR3105529, the automatic reading of a game slip includes a succession of steps, including the detection, segmentation and classification of game slips. R012357 PCT Deposit Text.docx
[0005] The detection function consists of determining whether or not a bulletin matching a listed bulletin template is present.
[0006] The segmentation function consists of determining the pixels of the image corresponding to the bulletin.
[0007] The classification function consists of determining the report card template.
[0008] The known methods do not allow these different functions to be performed quickly enough or in a sufficiently predictable time to allow a user to chain a stack of bulletins with robust reading.
[0009] A major contribution of the invention is the step of verifying particular graphic patterns which makes it possible to resolve ambiguities in the case of stacking several documents partially visible simultaneously, in order to determine which one is located on top.
[0010] The invention aims to resolve one or more of these drawbacks. The invention thus relates to a method for segmenting and classifying an image to be analyzed, as defined in attached claim 1.
[0011] The invention also relates to variants of the dependent claims. Those skilled in the art will understand that each of the features of the dependent claims and the description can be combined independently with the features of an intermediate claim, without thereby constituting an intermediate generalization.
[0012] The invention also relates to a terminal, as defined in the attached claims.
[0013] Other features and advantages of the invention will become clear from the description given below, by way of example and not limitation, with reference to the accompanying drawings, in which:
[0014] [Fig.1] is a schematic representation of a device for implementing a process according to the invention;
[0015] [Fig.2] is a schematic illustration of an example of a reference image of a structured document; R012357 PCT Deposit Text.docx
[0016] [Fig.3] is an illustration of an example of a shot of a structured document corresponding to the reference in Figure 2, which has undergone deformations and alterations;
[0017] [Fig.4] is an illustration of another example of a structured document intended to be identified;
[0018] [Fig.5] is a top view of a shot of a stack of documents on top of which is the document in figure 4 after various deformations;
[0019] [Fig.6] is a schematic representation of an example of an implementation method according to the invention.
[0020] Figure 2 illustrates an example of a reference image for a structured document 9. This reference image represents the ideal image of the reference document that can be expected during photography. Structured document 9 contains, in a way that is known in itself, various graphic patterns 901 to 907. Pattern 901 is a positioning block, patterns 903 are decorative elements, patterns 902 and 904 are grids, patterns 905 and 906 are boxes, and pattern 907 is a game code. A box typically refers to a pattern with any closed outline delimiting a space intended to be filled in by the user. Such an outline is usually identified by a line or by a difference in color between the background of the box and the rest of the document. In the example shown, some boxes have a square closed outline, and others have a star-shaped closed outline. Some boxes may also contain content such as one or more numbers or characters.Graphic patterns can, for example, correspond to signs, strings of characters, logos, positioning discs, geometric patterns, or small-sized images.
[0021] Figure 3 illustrates an example of a photograph of a corresponding structured document 9, which has undergone deformations and alterations, for example due to handling: creasing, folding, annotations, tears, optical distortions, variations in lighting, etc. Furthermore, the photograph of the structured document in Figure 3 could be tilted relative to the image of R012357 PCT Deposit Text.docx Reference. Shooting in an uncontrolled environment can induce more distortions and alterations.
[0022] The invention aims to implement the segmentation and classification of an image to be analyzed, as illustrated in Figure 3, in order to associate it with a structured document from among a set of other reference structured documents. Once the segmentation and classification have been implemented, other analysis methods can be used to determine how the structured document has been filled in by a user.
[0023] Figure 1 schematically illustrates an example of a device 1 for implementing a method according to the invention. Such a device 1 may, for example, be in the form of a terminal for processing a series of structured documents in paper form. The device 1 here includes an optical scanning element 11, for example, a camera. The camera is positioned at the top of an enclosure 12. The enclosure 12 has a receiving volume 13, at the bottom of which a structured paper document 9 to be analyzed is placed. To implement the method according to the invention, the device 1 has a database 14 containing the document templates and a digital processing system 15.
[0024] The segmentation and classification process can be implemented at the device level 1, in real time with image capture. Alternatively, the segmentation and classification process can be implemented in a dedicated device, receiving a stream of digitized images for digital processing.
[0025] Database 14 includes a set of structured document templates, each with an associated reference image, such as the one illustrated in Figure 2. A geometric transformation model defined by parameters and key vectors is also associated with each structured document template.
[0026] The method for segmenting and classifying an image to be analyzed according to the invention includes a preliminary step of retrieving an image to be analyzed. This image potentially includes a view of a structured document. R012357 PCT Deposit Text.docx This image to be analyzed can therefore come from the scanning device 11 or from an image stream or from an image database to be processed.
[0027] Next, a classification step is implemented for the image to be analyzed. During this classification, the following steps are performed (shown schematically in the flowchart of Figure 6): -a) the encoding (step 102 in Figure 6) of the image to be analyzed into a set of key vectors by a neural network. More precisely, the entire image to be analyzed is encoded into a set of key vectors by the neural network; -b) for each reference image in the set of models, identify via a neural network (step 103 in figure 6) the values of the parameters of its geometric transformation model between the key vectors of the image to be analyzed and the key vectors of this reference image; -c) for each reference image, apply (step 104 in figure 6), to the image to be analyzed, the geometric transformation associated with the identified parameters; -d) for each reference image, determine (step 105 in figure 6) the presence or absence of one or more predefined graphic patterns at a specific location in the image to be analyzed having undergone the respective geometric transformation; -e) associate (step 106 in figure 6) the image to be analyzed with a structured document of the set, according to the step of determining the presence or absence of the predefined graphic pattern(s) at the specific location.
[0028] Thus, the invention implements a test of each structured document model to determine the one that most likely corresponds to the image to be analyzed. Thanks to this systematic analysis and the use of neural networks, the segmentation and classification time is perfectly predictable. The method is therefore particularly well-suited for real-time image analysis, especially if the segmentation and classification process for each image requires minimal processing time. Due to the generation of key vectors for the entire image to be analyzed, the classification process R012357 PCT Deposit Text.docx proves to be particularly robust in the context of a stack of documents in the image to be analyzed.
[0029] Advantageously, the process includes a conversion step (step 100 in Figure 6) of the image to be analyzed into a fixed-resolution image prior to the encoding step of this image into a set of key vectors. The process is thus highly calibrated, regardless of the image acquisition method.
[0030] The step of converting the image to be analyzed into a fixed-resolution image advantageously includes either downsampling (step 101 in Figure 6) or upsampling of the image to be analyzed. Downsampling of the image to be analyzed can typically be performed with a dimension of 128 by 128 pixels.
[0031] The encoding of the image to be processed into a fixed-dimension key vector grid can be achieved with a dimension of 7 * 7 * 256, for example using a CNN (convolutional neural network) such as VGG16. The neural network will be advantageously chosen to achieve a compromise between the desired accuracy and computation speed.
[0032] Advantageously, the key vectors have a lower spatial dimension than the converted image and a higher per-pixel dimension than the converted image.
[0033] Advantageously, the determination of the presence of a predefined graphic pattern is implemented by a neural network.
[0034] Such a determination proves particularly robust for identifying the presence of a structured document superimposed on other elements that could interfere with its processing. For example, in Figure 5, a structured document 9 is superimposed on other documents, appearing during image scanning. Such surrounding elements appearing in the image can potentially interfere with the identification of the structured document 9 in the foreground. The structured document 9 on top of the stack has undergone alterations, and its image also contains distortions. R012357 PCT Deposit Text.docx
[0035] Structured document 9, present on the stack, is illustrated more precisely in Figure 4. Structured document 9 contains various predefined graphic patterns, 911 to 919, which are expected to be present in specific locations. These graphic patterns correspond here either to alphanumeric characters or to portions of an image.
[0036] Advantageously, step 106 of the association process can include additional checks to determine whether the image to be analyzed can be correctly associated with one of the structured reference documents. For example, additional checks can be performed on specific graphic patterns expected from the reference image. Among these additional checks, one could, for instance, verify pattern 907 of a game code.
[0037] Advantageously, the structured document is a game slip. The invention thus proves particularly useful in this scenario, where a user may wish to process a large number of documents sequentially.
[0038] Advantageously, several geometric transformations of each reference image are associated with each of the structured documents.
[0039] The identification (step 104) of a respective geometric transformation between the key vectors of the image to be analyzed and key vectors of this reference image is advantageously performed by a regression-type neural network. The geometric transformation can, for example, be an affine transformation or a homographic transformation.
[0040] In a particular embodiment, several different geometric transformations can be combined, such as an affine transformation or a homographic transformation, in order to determine the best one for each image to be analyzed.
[0041] A preliminary step may involve obtaining a digital image of the document to be analyzed. This preliminary digitization step can, for example, be carried out using a camera. A camera can generate a sequence of images at a rate of 5 to 20 frames per second. R012357 PCT Deposit Text.docx
[0042] In one variation, the process further includes a preliminary step of selecting the image(s) to be analyzed, followed by the decoding of the document's cells within a sequence of images based on the analysis of the graphic patterns in the preceding and current images of the sequence. Such a selection reduces processing time by focusing on the images with the greatest potential for processing.
[0043] The invention proves particularly advantageous when the image classification step is performed in real time with the scanning step of the next image to be analyzed. Indeed, it is in this real-time context that the reduced duration of the classification process according to the invention is benefited, since this classification can be carried out without slowing down the document submission process by the user. R012357 PCT Deposit Text.docx
Claims
9 Demands
1. A method for segmenting and classifying an image to be analyzed, comprising: -a step of retrieving an image to be analyzed, potentially including a view of a structured document; -a classification step including: -a) the encoding of the image to be analyzed into a set of key vectors by a neural network; -b) for each reference image of a set of structured document models, each having an associated reference image and a geometric transformation model defined by parameters and key vectors, identify via a neural network the values of the parameters of the respective geometric transformation model between the key vectors of the image to be analyzed and the key vectors of this reference image; -c) for each reference image, apply the respective geometric transformation with the parameters identified to the image to be analyzed; -d) for each reference image, determine the presence or absence of one or more predefined graphic patterns at a specific location in the image to be analyzed having undergone the respective geometric transformation; -e) associate the image to be analyzed with a structured document of the set, depending on the step of determining the presence or absence of the predefined graphic motif(s) at the specific location.
2. A method for segmenting and classifying an image to be analyzed according to claim 1, comprising a step of converting the image to be analyzed into a fixed-definition image prior to the step of encoding this image to be analyzed into a set of key vectors.
3. A method for segmenting and classifying an image to be analyzed according to claim 2, wherein said step of converting the image to be analyzed into a fixed-definition image includes downsampling or upsampling the image to be analyzed.
4. A method for segmenting and classifying an image to be analyzed according to claim 2 or 3, wherein the key vectors R012357 PCT Deposit Text.docx exhibit a spatial dimension lower than that of the converted image and a dimension per pixel higher than that of the converted image.
5. A method for segmenting and classifying an image to be analyzed according to any one of the preceding claims, wherein the determination of the presence of a predefined graphic pattern is implemented by a neural network.
6. A method for segmenting and classifying an image to be analyzed according to any one of the preceding claims, wherein the encoding of the image to be analyzed into a set of key vectors is implemented by a convolutional neural network.
7. A method for segmenting and classifying an image to be analyzed according to any one of the preceding claims, wherein the structured document is a game slip.
8. A method for segmenting and classifying an image to be analyzed according to any one of the preceding claims, wherein said graphic pattern comprises positioning tiles or disks or a geometric graphic pattern.
9. A method for segmenting and classifying an image to be analyzed according to any one of the preceding claims, wherein several geometric transformations of each reference image are associated with each of the structured documents.
10. A method according to any one of the preceding claims, wherein the identification of a respective geometric transformation between the key vectors of the image to be analyzed and key vectors of this reference image is carried out by a regression-type neural network.
11. A method according to any one of the preceding claims, comprising a preliminary step of obtaining a digital image of the document to be analyzed.
12. A method according to claim 11, wherein the preliminary digitization step is carried out via a camera. R012357 PCT Deposit Text.docx
13. A method according to claim 11 or 12, wherein the image classification step to be analyzed is carried out in real time with the subsequent image scanning step by the camera.
14. A method according to any one of the preceding claims, further comprising a preliminary step of selecting the image(s) to be analyzed up to the decoding of the boxes of the document in a sequence of images according to the result of analyzing the graphic patterns of the preceding and current images of the sequence.
15. Terminal, characterized in that it comprises: -a digital image capture device; -a digital processing module configured to retrieve an image from the digital imaging device and configured to implement a method according to any one of the preceding claims, j R012357 PCT Deposit Text.docx
Citation Information
Patent Citations
Method for analysing a content of at least one image of a deformed structured document
EP3153991A1
Method for segmenting an input image representing a document containing structured information
FR3105529A1