A method for segmenting and classifying an image, including a corresponding structured and terminal document.

The method uses neural networks to encode and transform images for rapid and accurate segmentation and classification of structured documents, addressing inefficiencies in existing technologies by ensuring predictable processing times and robustness against deformations and stacking.

FR3167236A1Pending Publication Date: 2026-04-10CARRUS GAMING
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
FR · FR
Patent Type
Applications
Current Assignee / Owner
CARRUS GAMING
Filing Date
2024-10-08
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing methods for automated reading of structured documents like game slips are inefficient and unreliable due to geometric deformations, stacking issues, and light disturbances, leading to unpredictable processing times and reduced robustness.

Method used

A method utilizing neural networks for image segmentation and classification, involving encoding images into key vectors, applying geometric transformations, and identifying predefined graphic patterns to accurately associate images with structured document templates, even under deformations and stacking conditions.

Benefits of technology

Enables rapid and robust image segmentation and classification, allowing real-time processing of multiple documents with minimal time delays and improved accuracy, even under geometric deformations and light disturbances.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The invention relates to a method for classifying an image to be analyzed, comprising: - a step of retrieving an image to be analyzed with a view of a structured document; - a classification step including: - a) encoding the image to be analyzed into a set of key vectors; - b) for each reference image in a set of document templates, identifying the parameter values ​​of a geometric transformation model between key vectors of the image to be analyzed and key vectors of that reference image; - c) for each reference image, applying the geometric transformation with the identified parameters to the image to be analyzed; - d) for each reference image, determining the presence of a graphic pattern at a specific location in the transformed image to be analyzed; - e) associating the image to be analyzed with a structured document, based on the determination of the presence of the graphic pattern. Figure to be published with the abstract: Fig. 1
Need to check novelty before this filing date? Find Prior Art

Description

Title of the invention: Method for segmenting and classifying an image including a corresponding structured and terminal document

[0001] The invention relates to the field of identification of documents containing structured information, in particular for documents such as game slips.

[0002] A game slip is an example of a document containing structured information. Other known structured documents include official forms or multiple-choice questionnaires. A game slip is printed on paper and has a predefined shape and design. The structured information typically takes the form of fields to be filled in, for example, checkboxes.

[0003] Automated reading of betting slips can be performed using a terminal equipped with a camera. A user may need to successively stack different betting slips in order to obtain their automated reading. This reading operation must be able to be performed quickly, simply by placing a new betting slip under a camera, without requiring precise positioning. Other slips can thus be visible beyond the edges of the last slip placed. This reading operation must also be able to be performed with a certain tolerance to geometric deformations of the slip, for example, its folding or creasing. This reading operation must also be robust against light disturbances related to ambient lighting. This reading operation must also be able to be performed by standard imaging devices such as cameras.

[0004] According to known technical processes, for example in document FR3105529, the automatic reading of a game slip comprises a succession of steps, including the detection, segmentation and classification of game slips.

[0005] The detection function consists of determining the presence or absence of a bulletin corresponding to a listed bulletin model.

[0006] The segmentation function consists of determining the pixels of the image corresponding to the bulletin.

[0007] The classification function consists of determining the model of the report card.

[0008] The known methods do not allow these different functions to be performed quickly enough or in a sufficiently predictable time to allow a user to chain a stack of bulletins with a robust reading.

[0009] A major contribution of the invention is the step of verifying particular graphic patterns which makes it possible to resolve ambiguities in the case of stacking several documents partially visible simultaneously, in order to determine which one is located on top.

[0010] The invention aims to resolve one or more of these drawbacks. The invention thus relates to a method for segmenting and classifying an image to be analyzed, comprising: -a step of retrieving an image to be analyzed, potentially including a view of a structured document; -a classification step including:

[0011] -a) encoding the image to be analyzed into a set of key vectors by a network of neurons;

[0012] -b) for each reference image of a set of document templates structured, each having an associated reference image and a geometric transformation model defined by parameters and key vectors, identify via a neural network the values ​​of the parameters of the respective geometric transformation model between the key vectors of the image to be analyzed and the key vectors of this reference image;

[0013] -c) for each reference image, apply the geometric transformation respective with the parameters identified in the image to be analyzed; -d) for each reference image, determine the presence or absence of one or more predefined graphic patterns at a specific location in the image to be analyzed having undergone the respective geometric transformation; -e) associate the image to be analyzed with a structured document of the set, depending on the step of determining the presence or absence of the predefined graphic motif(s) at the specific location.

[0014] The invention also relates to the following variants. Those skilled in the art will understand that each of the features of the following variants can be combined independently with the above features, without thereby constituting an intermediate generalization.

[0015] According to one variant, the process includes a step of converting the image to be analyzed into a fixed definition image prior to the step of encoding this image to be analyzed into a set of key vectors.

[0016] According to yet another variant, said step of converting the image to be analyzed into a fixed definition image includes the downsampling or upsampling of the image to be analyzed.

[0017] According to another variant, the key vectors have a spatial dimension lower than that of the converted image and a dimension per pixel higher than that of the converted image.

[0018] According to yet another variant, the determination of the presence of a predefined graphic pattern is implemented by a neural network.

[0019] According to one variant, the encoding of the image to be analyzed into a set of key vectors is implemented by a convolutional neural network.

[0020] According to another variant, the structured document is a game slip.

[0021] According to yet another variant, said graphic motif comprises positioning blocks or discs or a geometric graphic motif.

[0022] According to yet another variant, several geometric transformations of each reference image are associated with each of the structured documents.

[0023] According to one variant, the identification of a respective geometric transformation between the key vectors of the image to be analyzed and key vectors of this reference image is carried out by a regression-type neural network.

[0024] According to yet another variant, the process includes a preliminary step of obtaining a digital image of the document to be analyzed.

[0025] According to another variant, the preliminary digitization step is implemented via a camera.

[0026] According to yet another variant, the image classification step to be analyzed is carried out in real time with the next image scanning step by the camera.

[0027] According to one variant, the method further includes a preliminary step of selecting the image(s) to be analyzed up to the decoding of the boxes of the document in a sequence of images according to the result of the analysis of the graphic patterns of the previous and current images of the sequence.

[0028] The invention also relates to a terminal, comprising: -a digital imaging device; -a digital processing module configured to retrieve an image from the digital shooting device and configured to implement a process as defined above.

[0029] Other features and advantages of the invention will become clear from the following description, which is by way of example and not limitation, with reference to the accompanying drawings, in which:

[0030] [Fig-1] is a schematic representation of a device for implementing a process according to the invention;

[0031] [Fig.2] is a schematic illustration of an example of a reference image of a structured document;

[0032] [Fig.3] is an illustration of an example of a shot of a structured document corresponding to the reference of [Fig.2], having undergone deformations and alterations;

[0033] [Fig.4] is an illustration of another example of a structured document intended to be identified;

[0034] [Fig.5] is a top view of a shot of a stack of documents on top of which is the document of [Fig.4] after various deformations;

[0035] [Fig.6] is a schematic representation of an example of an implementation method according to the invention.

[0036] Figure 2 illustrates an example of a reference image for a structured document 9. This reference image corresponds to the ideal image of the reference document that can be expected during photography. The structured document 9 contains, in a manner known per se, various graphic motifs 901 to 907. Motif 901 is a positioning block, motifs 903 are decorative elements, motifs 902 and 904 are grids, motifs 905 and 906 are boxes, and motif 907 is a game code. A box usually refers to a motif with any closed outline delimiting a space intended to be filled in by the user. Such an outline is usually identified by a line or by a difference in color between the background of the box and the rest of the document. In the illustrated example, some boxes have a square closed outline and other boxes have a star-shaped closed outline. Some boxes may also contain content such as one or more numbers or characters.Graphic patterns can, for example, correspond to symbols, strings of characters, logos, positioning discs, geometric patterns, or small images.

[0037] Figure 3 illustrates an example of a photograph of a corresponding structured document 9, which has undergone deformations and alterations, for example, due to handling: creasing, folding, annotations, tears, optical distortions, variations in lighting, etc. Furthermore, the photograph of the structured document in Figure 3 may be tilted relative to the reference image. Photographing in an uncontrolled environment may induce further deformations and alterations.

[0038] The invention aims to implement the segmentation and classification of an image to be analyzed, as illustrated in [Fig. 3], in order to associate it with a structured document from among a set of other reference structured documents. Once the segmentation and classification have been implemented, other analysis methods can be used to determine how the structured document has been filled in by a user.

[0039] Figure 1 schematically illustrates an example of a device 1 for implementing a method according to the invention. Such a device 1 may, for example, be in the form of a terminal for processing a series of structured documents in paper form. The device 1 here comprises an optical element of Digitization 11, for example a camera. The camera is positioned at the top of an enclosure 12. The enclosure 12 has a receiving volume 13, at the bottom of which a structured paper document 9 to be analyzed is placed. To implement the method according to the invention, the device 1 has a database 14 containing document templates and a digital processing system 15.

[0040] The segmentation and classification process can be implemented at the device 1 level, in real time with the image capture. Alternatively, the segmentation and classification process can be implemented in a dedicated device, which receives a stream of digitized images for digital processing.

[0041] Database 14 comprises a set of structured document models, each having an associated reference image, such as that illustrated in [Fig. 2]. A geometric transformation model defined by parameters and key vectors is also associated with each structured document model.

[0042] The method for segmenting and classifying an image to be analyzed according to the invention includes a preliminary step of retrieving an image to be analyzed. This image potentially includes a view of a structured document. This image to be analyzed can therefore come from the scanning device 11 or from an image stream or from an image database to be processed.

[0043] Next, a classification step is implemented for the image to be analyzed. During this classification, the following steps are performed (steps schematically illustrated in the flowchart of [Fig. 6]): -a) the encoding (step 102 in [Fig.6]) of the image to be analyzed into a set of key vectors by a neural network; -b) for each reference image in the set of models, identify via a neural network (step 103 in [Fig.6]) the values ​​of the parameters of its geometric transformation model between the key vectors of the image to be analyzed and the key vectors of this reference image; -c) for each reference image, apply (step 104 in [Fig.6]), to the image to be analyzed, the geometric transformation associated with the identified parameters; -d) for each reference image, determine (step 105 in [Fig.6]) the presence or absence of one or more predefined graphic patterns at a specific location in the image to be analyzed having undergone the respective geometric transformation; -e) associate (step 106 in [Fig.6]) the image to be analyzed with a structured document of the set, according to the step of determining the presence or absence of the predefined graphic pattern(s) at the specific location.

[0044] Thus, the invention implements a test of each of the structured document models in order to determine the one that most likely corresponds to the image to be analyzed. Due to this systematic analysis and the use of neural networks, The segmentation and classification time is perfectly predictable. The process is therefore particularly well-suited for implementation in real-time image analysis, especially if the segmentation and classification process for each image requires minimal processing time.

[0045] Advantageously, the method includes a conversion step (step 100 in [Fig. 6]) of the image to be analyzed into a fixed-resolution image prior to the encoding step of this image into a set of key vectors. The method is thus particularly well-calibrated, regardless of the image acquisition process.

[0046] The step of converting the image to be analyzed into a fixed-resolution image advantageously includes either downsampling (step 101 in [Fig. 6]) or upsampling of the image to be analyzed. Downsampling of the image to be analyzed can typically be carried out with a dimension of 128 by 128 pixels.

[0047] Encoding the image to be processed into a fixed-dimension key vector grid can be achieved with a dimension of 7 * 7 * 256, for example by means of a CNN (convolutional neural network) such as VGG16. The neural network will advantageously be chosen to achieve a compromise between the desired accuracy and computational speed.

[0048] Advantageously, the key vectors have a spatial dimension lower than that of the converted image and a dimension per pixel higher than that of the converted image.

[0049] Advantageously, the determination of the presence of a predefined graphic pattern is implemented by a neural network.

[0050] Such a determination proves particularly robust for identifying the presence of a structured document superimposed on other elements that could interfere with its processing. Thus, in the example of [Fig. 5], a structured document 9 is superimposed on other documents, appearing during image scanning. Such surrounding elements appearing in the image can potentially interfere with the identification of the structured document 9 in the foreground. The structured document 9 in the stack has undergone alterations, and its image also contains distortions.

[0051] The structured document 9 present on the stack is illustrated more precisely in [Fig. 4]. The structured document 9 comprises various predefined graphic patterns 911 to 919, the presence of which is expected at specific locations. These graphic patterns correspond here either to alphanumeric characters or to portions of an image.

[0052] Advantageously, the association step 106 may include additional checks to determine whether the image to be analyzed can be properly associated with one of the structured reference documents. For example, checks may be performed Additional checks can be performed on specific graphic patterns expected from the reference image. For example, one can verify the 907 pattern of a game code.

[0053] Advantageously, the structured document is a game slip. The invention thus proves particularly useful in this scenario, where a user may wish to process a large number of documents sequentially.

[0054] Advantageously, several geometric transformations of each reference image are associated with each of the structured documents.

[0055] The identification (step 104) of a respective geometric transformation between the key vectors of the image to be analyzed and key vectors of this reference image is advantageously performed by a regression-type neural network. The geometric transformation can, for example, be an affine transformation or a homographic transformation.

[0056] In a particular embodiment, several different geometric transformations can be associated, such as an affine transformation or a homographic transformation, in order to determine the best one for each image to be analyzed.

[0057] A preliminary step may involve obtaining a digital image of the document to be analyzed. This preliminary digitization step may, for example, be carried out using a camera. A camera may, for example, generate a sequence of images at a rate of between 5 and 20 frames per second.

[0058] According to one variant, the method further comprises a preliminary step of selecting the image(s) to be analyzed, up to and including the decoding of the document's cells in a sequence of images based on the results of the analysis of the graphic patterns of the preceding and current images in the sequence. Such a selection reduces processing time by focusing on the images with the best processing capabilities.

[0059] The invention proves particularly advantageous when the image classification step is performed in real time with the scanning step of the next image to be analyzed. Indeed, it is in this real-time context that the reduced duration of the classification process according to the invention is benefited from, since this classification can be carried out without slowing down the document submission process by the user.

Claims

Demands

1. A method for segmenting and classifying an image to be analyzed, comprising: - a step of retrieving an image to be analyzed, potentially including a view of a structured document; - a classification step including: - a) encoding the image to be analyzed into a set of key vectors by a neural network; - b) for each reference image of a set of structured document models, each having an associated reference image and a geometric transformation model defined by parameters and key vectors, identifying, via a neural network, the values ​​of the parameters of the respective geometric transformation model between the key vectors of the image to be analyzed and the key vectors of that reference image; - c) for each reference image, applying the respective geometric transformation with the identified parameters to the image to be analyzed;-d) for each reference image, determine the presence or absence of one or more predefined graphic patterns at a specific location in the image to be analyzed after undergoing the respective geometric transformation; -e) associate the image to be analyzed with a structured document of the set, according to the step of determining the presence or absence of the predefined graphic pattern(s) at the specific location.

2. A method for segmenting and classifying an image to be analyzed according to claim 1, comprising a step of converting the image to be analyzed into a fixed-definition image prior to the step of encoding this image to be analyzed into a set of key vectors.

3. A method for segmenting and classifying an image to be analyzed according to claim 2, wherein said step of converting the image to be analyzed into a fixed-definition image includes downsampling or upsampling the image to be analyzed.

4. A method for segmenting and classifying an image to be analyzed according to claim 2 or 3, wherein the key vectors present a spatial dimension smaller than that of the converted image and a dimension per pixel larger than that of the converted image.

5. A method for segmenting and classifying an image to be analyzed according to any one of the preceding claims, wherein the determination of the presence of a predefined graphic pattern is implemented by a neural network.

6. A method for segmenting and classifying an image to be analyzed according to any one of the preceding claims, wherein the encoding of the image to be analyzed into a set of key vectors is implemented by a convolutional neural network.

7. A method for segmenting and classifying an image to be analyzed according to any one of the preceding claims, wherein the structured document is a game slip.

8. A method for segmenting and classifying an image to be analyzed according to any one of the preceding claims, wherein said graphic pattern comprises positioning tiles or disks or a geometric graphic pattern.

9. A method for segmenting and classifying an image to be analyzed according to any one of the preceding claims, wherein several geometric transformations of each reference image are associated with each of the structured documents.

10. A method according to any one of the preceding claims, wherein the identification of a respective geometric transformation between the key vectors of the image to be analyzed and key vectors of this reference image is carried out by a regression-type neural network.

11. A method according to any one of the preceding claims, comprising a preliminary step of obtaining a digital image of the document to be analyzed.

12. A method according to claim 12, wherein the preliminary digitization step is carried out via a camera.

13. A method according to claim 12 or 13, wherein the image classification step to be analyzed is carried out in real time with the subsequent image scanning step by the camera.

14. A method according to any one of the preceding claims, further comprising a preliminary step of selecting the image(s) to be analyzed up to the decoding of the document's cells in a sequence of images based on the result of analyzing the graphic patterns of the previous and current images in the sequence.

15. Terminal, characterized in that it comprises: -a digital image capture device; -a digital processing module configured to retrieve an image from the digital image capture device and configured to implement a method according to any one of the preceding claims.

Citation Information

Patent Citations

  • Method for analysing a content of at least one image of a deformed structured document

    EP3153991A1

  • Method for segmenting an input image representing a document containing structured information

    FR3105529A1