Image classification method, and corresponding electronic device and computer program product
Patent Information
- Application Number
- EP2023735041
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-06-29
- Filing Date
- 2023-06-26
- Publication Date
- 2025-05-07
Smart Images

Figure 1.1
Abstract
Description
[0001] DESCRIPTION
[0002] Title of the invention: Image classification method, electronic device and corresponding computer program product
[0003] 1. Technical field
[0004] This application relates to the field of classification (or categorization) of digital elements comprising a representation of information viewable on a screen of an electronic device, such as digital documents. These elements will also be referred to more simply as “images” hereinafter.
[0005] The present application relates in particular to a method for classifying such images by an electronic device, as well as a corresponding electronic device, a computer program product and a recording (or information) medium.
[0006] 2. State of the art
[0007] Many technical fields implement image classification techniques. Some artificial intelligence techniques, such as so-called "machine learning" techniques, require a training dataset, provided as examples of classified data, to train a classification model. These techniques sometimes need a large training dataset to obtain a relevant model. This is often the case for image classification techniques based on neural networks. Such techniques can, for example, require several thousand or millions of training data. For example, a widely used dataset such as the one from the ImageNet Large Scale Visual Recognition Challenge (ILSVRC 2012-2017), includes more than a million training images.
[0008] Preparing (especially annotating) these training datasets can be a long and tedious task.
[0009] The present application aims to propose improvements to at least some of the disadvantages of the state of the art.
[0010] 3. Statement of the invention
[0011] The present application aims to improve the situation using an image classification method implemented in an electronic device comprising: obtaining images comprising a plurality of visual items; dividing the images into a plurality of classes, based on at least one occurrence, in said images, of visual items extracted from said images. For example, the present application relates to an image classification method implemented in an electronic device and comprising: obtaining images comprising a plurality of visual items; dividing the images into a plurality of classes, based at least on a number of occurrences, in said images, of visual items extracted from said images.
[0012] By image, we mean here, as explained above, a representation of information that can be viewed on a screen of an electronic device, such as digital documents
[0013] According to at least one embodiment, said classes are identified during said distribution.
[0014] According to at least one embodiment, said method comprises labeling the identified classes.
[0015] According to at least one embodiment, said labeling is carried out automatically.
[0016] According to at least one embodiment, said images divided into said plurality of classes are used to train a neural network to classify other images according to said plurality of classes.
[0017] According to at least one embodiment, said method is implemented locally on said electronic device.
[0018] According to at least one embodiment, said images are obtained via a software probe and / or a scanning device.
[0019] According to at least one embodiment, said images comprise at least one digital document accessible to said device.
[0020] According to at least one embodiment, said visual items are words, groups of words and / or graphic objects of said images.
[0021] According to at least one embodiment, the method comprises replacing at least one first word, among said visual items, with at least one second word.
[0022] According to at least one embodiment, said at least one second word is a generic word or group of words describing a type of said first word.
[0023] For example, the "first" words "Dupond" and "Durand" can both be replaced by the same "second" word "Name".
[0024] According to at least one embodiment, said method comprises obtaining the positions of said extracted visual items in said plurality of images.
[0025] According to at least one embodiment, said visual items are words and / or groups of words and said method comprises, for an analyzed image of said plurality of images, an association with said analyzed image of the occurrence numbers of the distinct visual items extracted from said analyzed image. According to at least one embodiment, said method comprises a deletion, in said distinct visual items associated with said analyzed image, of at least one stop word.
[0026] According to at least one embodiment, said method comprises a deletion, in said distinct visual items associated with the analyzed images of said plurality of images, of the visual items associated with a number of images greater than a desired number of classes.
[0027] According to at least one embodiment, said method comprises: grouping the distinct visual items of said analyzed image taking into account the different values of said occurrence numbers of distinct visual items in said analyzed image; obtaining at least one candidate distribution of said images for a candidate value of the different occurrence numbers, taking into account distinct visual items common to at least two analyzed images grouped for said candidate value.
[0028] According to at least one embodiment, said distribution takes into account patterns, relating to the positions of said visual items in said images of said plurality of images, and present in at least two analyzed images of said plurality of images.
[0029] According to at least one embodiment, said method comprises: detecting at least one pattern in analyzed images of said plurality of images; associating said detected pattern with the analyzed images of said plurality of images in which an occurrence of said pattern has been detected; obtaining at least one candidate distribution of said analyzed images taking into account at least one occurrence of at least one pattern associated with at least two of said analyzed images.
[0030] According to at least one embodiment, said distribution takes into account the positions of said patterns in said analyzed images.
[0031] According to at least one embodiment, said method comprises an assignment of a confidence score to a class, taking into account a presence of at least one visual item associated with at least one image of said class, in at least one other image of at least one other class.
[0032] According to at least one embodiment, said method comprises obtaining at least two candidate distributions and said distribution is chosen, from said candidate distributions, taking into account a desired number of classes and / or the confidence score assigned to at least one of the classes of said candidate distributions.
[0033] According to at least one embodiment, said method comprises a modification of at least one of said images obtained before an analysis of said images.
[0034] According to at least one embodiment, the method comprises an association of a textual label with at least one of said classes.
[0035] According to at least one embodiment, the method comprises a rendering of at least one of said images and said label associated with the class of said rendered image is obtained from a user interface of said device.
[0036] The features presented in isolation in the present application in connection with certain embodiments of the method of the present application may be combined with each other according to other embodiments of the present method. According to another aspect, the present application also relates to an electronic device suitable for implementing the method of the present application in any of its embodiments. For example, the present application thus relates to an electronic device comprising at least one processor configured for image classification comprising: obtaining images comprising a plurality of visual items; distributing the images into a plurality of classes, based on at least one occurrence, in said images, of visual items extracted from said images.
[0037] For example, the present application also relates to an electronic device comprising at least one processor configured for image classification comprising: obtaining images comprising a plurality of visual items; dividing the images into a plurality of classes, based at least on a number of occurrences, in said images, of visual items extracted from said images.
[0038] The present application also relates to a computer program comprising instructions for implementing the various embodiments of the above method, when the program is executed by a processor and a recording medium readable by an electronic device and on which the computer programs are recorded.
[0039] For example, the present application thus relates to a computer program comprising instructions for implementing, when the program is executed by a processor of an electronic device, a method for classifying images comprising: obtaining images comprising a plurality of visual items; dividing the images into a plurality of classes, based on at least one occurrence, in said images, of visual items extracted from said images.
[0040] For example, the present application also relates to a computer program comprising instructions for implementing, when the program is executed by a processor of an electronic device, a method for classifying images comprising: obtaining images comprising a plurality of visual items; dividing the images into a plurality of classes, based at least on a number of occurrences, in said images, of visual items extracted from said images.
[0041] For example, the present application also relates to a recording medium (or information medium) readable by a processor of an electronic device and on which is recorded a computer program comprising instructions for implementing, when the program is executed by the processor, a method for classifying images comprising: obtaining images comprising a plurality of visual items; dividing the images into a plurality of classes, based on at least one occurrence, in said images, of visual items extracted from said images.
[0042] For example, the present application also relates to a recording medium readable by a processor of an electronic device and on which is recorded a computer program comprising instructions for implementing, when the program is executed by the processor, a method for classifying images comprising: obtaining images comprising a plurality of visual items; dividing the images into a plurality of classes, based at least on a number of occurrences, in said images, of visual items extracted from said images.
[0043] The above-mentioned programs may use any programming language, and may be in the form of source code, object code, or intermediate code between source code and object code, such as in a partially compiled form, or in any other desirable form.
[0044] The information carriers (or recording media) mentioned above may be any entity or device capable of storing the program. For example, a carrier may include a storage medium, such as a ROM, for example a CD ROM or a microelectronic circuit ROM, or a magnetic recording medium. Such a storage medium may, for example, be a hard disk, a flash memory, etc.
[0045] On the other hand, an information carrier may be a transmissible carrier such as an electrical or optical signal, which may be conveyed via an electrical or optical cable, by radio or by other means. A program according to the invention may in particular be downloaded from a network such as the Internet.
[0046] Alternatively, an information carrier may be an integrated circuit in which a program is incorporated, the circuit being adapted to execute or to be used in the execution of any of the embodiments of the method which is the subject of the present patent application.
[0047] 4. Brief description of the drawings
[0048] Other characteristics and advantages of the invention will appear more clearly on reading the following description of particular embodiments, given as simple illustrative and non-limiting examples, and the appended drawings, among which: [Fig 1] presents a simplified view of a system, cited as an example, in which at least certain embodiments of the method of the present application can be implemented,
[0049] [Fig 2] shows a simplified view of a device suitable for implementing at least certain embodiments of the method of the present application;
[0050] [Fig 3] presents an overview of the method of the present application, in at least some of its embodiments;
[0051] [Fig 4] shows in more detail certain treatments of the method of the present application, in certain embodiments compatible with the embodiments illustrated in Figure 3;
[0052] [Fig 5] shows in more detail certain treatments of the method of the present application, in certain embodiments compatible with the embodiments illustrated in Figure 3;
[0053] [Fig 6] shows in more detail certain treatments of the method of the present application, in certain embodiments compatible with the embodiments of Figures 4 and 5;
[0054] [Fig 7] shows, schematically and summarily, flow exchanges between some of the devices of the system 100 for implementing the method of the present application, in at least some of its embodiments.
[0055] [Fig 8] shows an example of patterns used in the embodiments illustrated in Figure 5.
[0056] 5. Description of the embodiments The present application proposes an automatic (or at least partially automatic) classification of images. More specifically, the present application proposes in at least certain embodiments to use automatic image analysis techniques to obtain, from these images, data characterizing elements represented by these images. These elements are for example textual elements such as words or groups of words, or graphic objects occupying a portion of the image. This data will then be used to divide the images into different classes, or categories (also called "clusters" according to English terminology).
[0057] Thanks to this division into classes (or "clustering" according to English terminology), the method of the present application, in at least some of its embodiments, can help to easily label (in other words annotate) a large set of images, for example by assigning one or more of the same labels to all the images of a class (or category). These labeled images can for example then be used as training data, in the context of supervised learning (for example for training a neural network intended to classify other images during its inference).
[0058] This automatic classification (or categorization) facilitating the labeling of images, it can also help, in at least certain embodiments of the method of the present application, to evolve the classes of a neural network over time (by successive learning, on variable data sets).
[0059] An automatic image classification can also make it possible to automatically label images, without intervention from an operator (for example by automatically assigning labels to the classes (such as successive numbers) (and possibly subsequently offering an operator the possibility of modifying the class labels as he wishes). As a result, an automatic image classification can therefore offer, at least in certain embodiments, advantages in terms of confidentiality of the images (which may for example correspond to personal data of an individual or group of individuals), and / or processing speed.Such automatic classification can also help to avoid, or at least limit, input errors, and simplify the choices of assigning a class (or category) to an image, in cases of distribution according to fairly complex criteria, or when the number of classes (or clusters) is high (for example of the order of a few dozen classes).
[0060] Automatic classification of images, when these correspond to captures or scans of administrative documents, can also offer advantages in terms of reliability, speed and / or confidentiality, for the electronic archiving of administrative documents. For example, thanks to the method of the present application, in at least some of its embodiments, a user can scan a stack of documents and obtain, automatically, a distribution of these documents into classes (pay slips, bank statements, social security statements, etc.) which he then only has to archive separately accordingly.
[0061] According to yet another example, automatic classification of training data may also enable, in at least some embodiments, training (e.g., federated training) of neural networks using, for training a neural network for inference on a device, self-labeled data local to that device, for example to meet obligations related to the protection of personal data.
[0062] The method of the present application can be implemented to classify various types of images. For example, as highlighted above, in certain embodiments it may involve scanning various types of digital documents: administrative documents (birth certificates, death notices, etc.), commercial documents (invoices, delivery notes, purchase orders, etc.), receipts, etc. In certain embodiments, it may also involve classifying images of products or objects (represented in the images) (for example, for the purpose of producing an advertising catalog).
[0063] The present application is now described in more detail in connection with Figure 1.
[0064] Figure 1 shows a telecommunications system 100 in which at least some embodiments of the invention can be implemented. The system 100 comprises one or more electronic devices, at least some of which can communicate with each other via one or more communication networks, possibly interconnected, such as a local area network or LAN (for Local Area Network according to English terminology) and / or a wide area network, or WAN (for Wide Area Network according to English terminology). For example, the network may comprise a corporate or domestic LAN network and / or a WAN network of the internet type, or cellular, GSM - Global System for Mobile Communications, UMTS - Universal Mobile Telecommunications System, Wifi - Wireless, etc.).
[0065] As illustrated in Figure 1, the system 100 may also comprise several electronic devices, such as a terminal (such as a laptop 110, a smartphone 130, a tablet 120), a server 140, and a storage device 150, on which may for example be stored parameters of a neural network, such as parameters relating to its structure (for example a description of its different layers, the number and sizes of the matrices associated respectively with these layers, etc.) and the current values of the coefficients of the matrices associated with these layers. The server 140 may for example use training data, for example previously stored on the storage device or on another of the devices of the system, to refine the values of the coefficients of the neural network during an (optional) training phase (for example prior) of the neural network.
[0066] One of the terminals 110, 120, 130 can also obtain, from the storage device for example, the current parameters of the neural network (including the current values of the coefficients, possibly previously learned using the server or another terminal) and carry out “local” learning of the neural network to refine the values of the coefficients as a function, for example, of learning data specific to the terminal, such as data stored locally by the terminal or stored remotely but accessible to the terminal and relating to the terminal.
[0067] The training data used by the server and / or the terminal may have been obtained by at least some embodiments of the method of the present application.
[0068] The system may also include network management and / or interconnection elements (not shown). These electronic devices may be associated with at least one user 132 (for example, via a user account accessible by login), some of the electronic devices 110, 130 being able to be associated with the same user 132. The system may also include a database 160, such as a lexical base.
[0069] Figure 2 illustrates a simplified structure of an electronic device 200 of the system 100, for example the device 110, 130 or 140 of Figure 1, adapted to implement the principles of the present application. Depending on the embodiments, it may be a server and / or a terminal.
[0070] The device 200 comprises in particular at least one memory M 210. The device 200 may in particular comprise a buffer memory, a volatile memory, for example of the RAM type (for “Random Access Memory” according to English terminology), and / or a non-volatile memory (for example of the ROM type (for “Read Only Memory” according to English terminology). The device 200 may also comprise a processing unit UT 220, equipped for example with at least one processor P 222, and controlled by a computer program PG 212 stored in memory M 210. At initialization, the code instructions of the computer program PG are for example loaded into a RAM memory before being executed by the processor P.The at least one processor P 222 of the processing unit UT 220 can in particular implement, individually or collectively, any one of the embodiments of the method of the present application (described in particular in relation to FIG. 3), according to the instructions of the computer program PG.
[0071] The device may also comprise, or be coupled to, at least one I / O input / output module 230, such as a communication module, allowing for example the device 200 to communicate with other devices of the system 100, via wired or wireless communication interfaces, and / or such as a module for interfacing with a user of the device (also called more simply “user interface” in the present application).
[0072] By “user interface” of the device, we mean for example an interface integrated into the device 200, or a part of a third-party device coupled to this device by wired or wireless communication means. For example, it may be a secondary screen of the device or a set of speakers connected by wireless technology to the device
[0073] A user interface may in particular be a user interface, called an “output” user interface, adapted to rendering (or controlling rendering) of an output element of a computer application used by the device 200, for example an application running at least partially on the device 200 or an “online” application running at least partially remotely, for example on the server 140 of the system 100. Examples of output user interfaces of the device include one or more screens, in particular at least one graphic screen (touch screen for example), one or more speakers, and / or a connected headset.
[0074] By rendering, we mean here a restitution (or “output” according to English terminology) on at least one user interface, in any form, for example including textual, audio and / or video components, or a combination of such components.
[0075] Furthermore, a user interface can be a user interface, called "input
[0076] ", adapted to an acquisition of information from a user of the device 200. This may in particular be information intended for a computer application accessible via the device 200, for example an application running at least partially on the device 200 or an "online" application running at least partially remotely, for example on the server 140 of the system 100. Examples of input user interface of the device 200 include a sensor, an audio and / or video acquisition means (microphone, camera (webcam) for example), a keyboard, a mouse. The device may also comprise at least one software module (or software probe) adapted to the capture of data entered or returned on the user interface of the device.
[0077] Said at least one microprocessor of the device 200 may in particular be adapted for a classification of images comprising: obtaining images comprising a plurality of visual items; a distribution of the images into a plurality of classes, as a function of at least one occurrence, in said images, of visual items extracted from said images. Said at least one microprocessor of the device 200 may in particular be adapted for a classification of images comprising: obtaining images comprising a plurality of visual items; a distribution of the images into a plurality of classes, as a function of at least a number of occurrences, in said images, of visual items extracted from said images.
[0078] Some of the above input-output modules are optional and may therefore be absent from the device 200 in certain embodiments. In particular, if the present application is sometimes detailed in connection with a device communicating with at least one second device of the system 100, the method may also be implemented locally by a device, for example using as input elements elements acquired for example by software and / or hardware probes executing on the device, to produce output elements stored locally on the device, or rendered via an output interface of the device.
[0079] Rather, in some of its embodiments, the method may be implemented in a distributed manner between at least two devices 110, 120, 130, 140 and / or 150 of the system 100.
[0080] The term "module" or the term "component" or "element" of the device is understood here to mean a hardware element, in particular wired, or a software element, or a combination of at least one hardware element and at least one software element. The method according to the invention can therefore be implemented in various ways, in particular in wired form and / or in software form.
[0081] Figure 3 illustrates certain embodiments of the method 300 of the present application. The method 300 can for example be implemented by the electronic device 200 illustrated in Figure 2.
[0082] As illustrated in FIG. 3, the method 300 may comprise obtaining 310 a set (or batch) of images to be classified.
[0083] As illustrated in Figure 3, the method 300 may comprise obtaining 320 visual items extracted from the images obtained. In the present application, “visual item” means a word, a group of words or a graphic object. This obtaining may comprise a search 322 for visual items in the images obtained (or alternatively access to at least one file associated with at least one of the images obtained and comprising visual items previously extracted from this image). The search 322 may for example implement, in certain embodiments, techniques for analyzing images and / or detecting elements represented by these images, such as character recognition techniques (for example techniques called OCR (for “Optical Character Recognition” according to the English terminology) leading to an extraction of word(s) from the images). It may also involve, for example, techniques for recognizing shapes and / or objects or and / or image segmentation.
[0084] In the case where the search 322 makes it possible to identify words in one of the images obtained, the method can also comprise a search 323 for at least one entity named in the words extracted by comparison with the elements of a lexical base for example (such as for example the element 160 of the system 100).
[0085] Such a search may for example comprise an execution of a service accessible to the device 200) for example a service of a software library (such as the service “Allganize” ©), in charge of detecting, in a text (here in words extracted from an analyzed image) words considered by the service as of a particular type (such as a name, a first name, an address, a telephone number, an organization); The method may also comprise a replacement (for example via a service as introduced above), in this text, of words having one of these particular types by a generic word, or group of words, describing this particular type (and subsequently called “named entity”). Examples of named entities may include the terms “name”, “first name”, “address”, “telephone number”, “organization”). For example, the group of words “Jean Dupont lives in Clermont-Ferrand” could become: “First name Last name lives in City”.This search 323 for named entity(ies) may be optional in certain embodiments.
[0086] In certain embodiments, the method 300 may comprise, prior to a search (and extraction) of visual items from at least one image, a pre-processing 321 (or “preprocessing” according to the English terminology) of at least one of the images obtained, in order for example to facilitate the extraction of visual items from the image.
[0087] This may involve, for example, in at least one embodiment, the application of at least one image processing technique. Such a technique may be, according to a first example, a color transformation (for example a conversion) applied to at least one of the images obtained, in order to retain, in the transformed image, only colors corresponding to different gray levels. Another technique may be, according to another example, a transformation applied to at least one of the images obtained, in order to vary the contrasts within the image (for example to increase the contrasts) and / or to modify the brightness of the image.
[0088] This pre-processing may be optional in certain embodiments.
[0089] The detection and / or extraction of visual items from the obtained (and possibly pre-processed) images can also make it possible, in certain embodiments, not only to identify visual items present in an image but also to obtain their positions in this image.
[0090] Following the extraction of visual items from the analyzed images (and the possible search for named entity(ies)), we therefore obtain, per image on which the extraction was carried out, a list of visual items extracted from this image and possibly their positions in this image.
[0091] As illustrated in Figure 3, the method 300 may comprise an analysis 330 of the visual items obtained (for example the words and named entities extracted from the images), to distribute 340 the images into a plurality of classes (or “clusters” according to English terminology) taking into account these visual items.
[0092] The analysis 330 and the distribution 340 based on this analysis, which are carried out, may depend on the embodiments. Thus, according to certain embodiments illustrated in FIG. 4, the analysis 331 (and therefore the distribution 341) may be based on a number of occurrences of visual items of the word or group of words type extracted from the analyzed images. According to other embodiments illustrated in FIG. 5, the analysis 332 (and therefore the distribution 342) may be based on the presence of “patterns” in the analyzed images relating to the positions of certain visual items in the images.
[0093] Figure 4 thus illustrates in more detail an example of analysis 331 of the images and associated visual items and distribution 341 of these images based on a number of occurrences of the visual items in the analyzed images.
[0094] In the example illustrated in Figure 4, the analysis 331 of an image is based on the visual items obtained 320 (Fig. 3) from this image (and associated with this analyzed image). More precisely, the analysis 331 comprises, for at least one visual item associated with the analyzed image, a count 3311 of the number of occurrences, in the analyzed image, of this visual item, that is to say the number of times this visual item (word or named entity) appears in this image. The analysis can also comprise a classification (or grouping) 3312 of the visual items associated with the analyzed image according to their number of occurrences (so as to gather together, for example, in a first group all the visual items appearing only once in the analyzed image, then in a second group all the visual items appearing exactly twice in the analyzed image, etc.).We therefore obtain, per analyzed image, one or more lists of visual items, each list being dedicated to a distinct number of occurrences.
[0095] The pseudo code below represents, as an example, the groupings
[0096] ClaWordlmg[1] .. ClaWordlmg[n] obtained respectively for one image among the plurality of images h .. I n analyzed: }
[0097] As shown in Figure 4, the analysis may include filtering 3313 some of the visual items associated with an image. In the example of Figure 4, this filtering may for example include removing visual item(s) that are useless for classifying the images. For example, the method may include removing at least one stop word (also called a “stopword”), or sometimes a “transition word,” “stop word,” “linking word,” or “portmanteau word”) (such as an article, or a linking word), the method may also include, in some embodiments, removing items appearing in a high number of images.Since the classification of the images is subsequently carried out based on the visual items extracted from the images, a visual item present in a number of images greater than a desired number of classes in the distribution would indeed be of little use for distinguishing images for the purpose of their distribution. Similarly, in certain embodiments, visual items associated with a number of images close to the number of images to be classified (such as associated with more than 90% of the analyzed images) may be deleted in certain embodiments. In addition, the inventors noted that certain elements, such as a logo or certain words (for example “Pay” or “Invoice” in the case of a type of administrative document), were often present a small number of times in an image while being characteristic of a type of image.As a result, certain embodiments, such as detailed embodiments, may favor low numbers of occurrences and delete items associated, for example, with a number of occurrences greater than a first number of occurrences (used as a high threshold, for example).
[0098] Depending on the embodiments, for example depending on the filtering performed, this filtering can be performed at least partially before and / or after grouping by number of occurrences.
[0099] Thus, filtering of useless words can, for example, be carried out before grouping visual items according to their number of occurrences, to help gain processing efficiency, for example, during this grouping.
[0100] Filtering may be optional in some embodiments (e.g., it may be enabled or disabled via a configuration setting).
[0101] The pseudo-code below thus represents, for the plurality of images h .. I n analyzed, the previously introduced ClaWordlmg[1] .. ClaWordlmg[n] groupings once these have been filtered:
[0102] 30
[0103] We see that the items 'word-ggg' and 'word-ddd' have for example been removed from the ClaWordlmg[1] and ClaWordlmg[n] groupings. The groupings of visual items by number of occurrences for each analyzed image can be used to divide the images into classes. As shown in Figure 4, the method can include obtaining a candidate distribution for at least one value k of the number of occurrences of visual items. Depending on the embodiments, this can be a value of "k" particular to each analyzed image or common to all analyzed images. More precisely, in certain embodiments, the method can include a search 3411 for items common to several images corresponding to a given number k of occurrences, which makes it possible to group these images into "x" clusters (where x is the number of desired classes). This involves identifying a class from the candidate distribution via a list of items common to the images of this class.For example, we search for the maximum number of common words in the analyzed images allowing us to differentiate the images by grouping them into "x" clusters. The pseudo-code below thus represents examples of classes (or clusters) obtained for the plurality of images h .. I. n analyzed, in conjunction with the pseudo-code examples previously introduced.
[0104] The association and identification steps can be performed for several occurrence values k of some images, for example until a number of classes corresponding to the number x of desired classes is found.
[0105] For example, in some embodiments, a first candidate distribution may be obtained for a first value of k, common to the images, chosen to correspond to the smallest number of occurrences of items on all the images (i.e. k=1 for example), at least one second candidate distribution being obtained, for at least one second value of occurrences, greater than the first value chosen for the first distribution. For example, candidate distributions may be obtained for successive, increasingly larger values of the number of occurrences, until a desired number of classes is obtained or until a number of occurrences corresponding to the largest number of occurrences associated with all the images is reached (i.e. the minimum value, over all the analyzed images, of the largest number of occurrences associated with each of these images).In certain embodiments, the identification of classes may take into account, for example, a minimum or maximum number or percentage of images per class, so as to obtain relatively homogeneous classes, and / or a desired number x of classes, for example, as explained above.
[0106] In some embodiments, the desired number of classes may not be fixed, a minimum and / or maximum number of occurrences to be used for the different images may for example be defined. Different occurrence values may for example be tested to arrive at a classification of the set of images respecting this minimum and / or maximum number of occurrences.
[0107] In some embodiments, when the desired number of classes is not fixed, the choice of a distribution, from among the candidate distribution(s), may take into account a confidence level associated with the distribution (or at least one class of the distribution). For example, the chosen distribution may be the first candidate distribution obtained associated with a confidence level greater than a first value (this first value may be a configuration parameter or be deduced from such a parameter). Thus, in some embodiments, candidate distributions may be searched for all occurrence numbers less than or equal to the largest occurrence number common to all images, the method then comprising a selection of the candidate distribution having the best confidence score.
[0108] In some embodiments, the method may include assigning a confidence score to at least one class (e.g., each class) of at least one of the candidate distributions.
[0109] This attribution may be optional in certain embodiments.
[0110] For example, in embodiments consistent with the illustration of Figure 4, a confidence level may be assigned to each item associated with a class, this confidence level being for example calculated by taking into account the number of classes in which the item is present. More specifically, the confidence level of an item in a class may be calculated in certain embodiments as a ratio between the number of images belonging to this class where this item appears, and the total number of images (in all classes) where the visual item is present.
[0111] A confidence score, at the level of the class considered, can for example be calculated by taking into account the respective confidence levels of the items in the class. For example, the calculation of the cluster confidence score can be based on the average of the confidence levels of the items, on the standard deviation relative to these confidence levels, etc.
[0112] The score of a class may, in certain embodiments, take into account compliance with at least one criterion relating to the values of the confidence levels of the items associated with it. For example, a value of a confidence level, for one of the items of the class, lower than a first value (for example a “threshold” value) may degrade (for example decrease), or in other embodiments improve (for example increase), the confidence score of the class (via a multiplicative coefficient for example). Similarly, a value of a confidence level, for one of the items of the class, higher than a second value (for example a “threshold” value) may improve, or in other embodiments degrade, the confidence score of the class.
[0113] Similarly, an overall score can be assigned to a candidate distribution by taking into account the confidence scores of all the different classes in the distribution.
[0114] The method may further comprise a selection of a distribution (called chosen or selected), from among the candidate distributions. One of the selection criteria may, for example, take into account the overall score obtained for a candidate distribution. Another example of a selection criterion may, for example, be compliance with at least one configuration parameter as detailed below.
[0115] Figure 5 thus illustrates in more detail an example of analysis 332 and distribution 342 based on the presence of patterns in the analyzed images, and optionally on the positions of these patterns in the analyzed images.
[0116] In certain embodiments, following the search (and detection) 322, 323 of visual elements (step 320), the method may comprise a search 324 of at least one pattern, or pattern, in terms of occupation of blocks by particular visual items, within an image or between the analyzed images, i.e. a repetition of an occupation of one or more blocks by one or more first visual items within the same image (intra-image pattern) or between at least two analyzed images (inter-image pattern). This may for example be a fixed pattern or a floating pattern.A fixed pattern corresponds to a repetition (possibly with a scaling factor, such as a multiplicative coefficient, in certain embodiments), in at least m analyzed images (with m being an integer greater than or equal to 2, definable for example by configuration), of a set of occupied block(s) of the same relative position(s) (of the same “offsets” according to English terminology), relative to a reference position (for example an origin position (0;0)), fixed between these at least two images. Thus, for example, a fixed pattern may correspond to a repetition of an identical set of occupied blocks of the same indices in several images. In certain embodiments, the visual items of a pattern are identical between the repetitions of the pattern. For example, a pattern may correspond to a repetition of the sequence of visual items of 3 entities named as name, first name, time, distributed over 3 consecutive blocks.
[0117] Figure 8 thus illustrates three images 810, 820, 830, some of whose blocks (hatched) are occupied by visual items. The groups of occupied blocks 811, 812, 813 present on several images correspond to (fixed) patterns. These patterns can have various shapes, more or less complex, as illustrated by element 831 of image 830 (this element is not present in images 810 and 820 but is deemed to be present in the context of this example in at least one other image not illustrated). A sliding pattern has a repetition (possibly with a scaling factor (such as a multiplicative coefficient) in certain embodiments), in at least two analyzed images, of a set corresponding to a set of occupied block(s) of the same relative position(s) (for example of the same offset), relative to a reference position, likely to vary within an image, or between different images.
[0118] According to the embodiments, the method can search only for fixed patterns or search for fixed patterns and floating patterns, or be limited to fixed patterns and floating patterns, the position of which, although variable, is located in a certain portion of the image (e.g.: right side of the images, center or left side, etc.).
[0119] The visual items of a pattern can be of different types. For example, a pattern can include at least one word, at least one named entity and / or at least one graphic object (e.g. a logo). In the case of textual items, the search for at least one pattern can thus take into account a repeated proximity between at least one named entity and at least one word, and / or a proximity between at least two named entities, and / or a proximity between at least two words.
[0120] Once the fixed and / or floating patterns have been detected, the method may comprise an analysis 332 of at least some of the obtained images 310, from the patterns (fixed and / or floating) detected during the above search. In certain embodiments, the method may comprise a search 324 for the presence of at least one pattern in an image and (optionally) an association 3321 with the patterns present in this image of their position(s) in this image.
[0121] For example, for an analyzed image, a first pattern P1, present several times in the image, will be associated with its positions (pos 11, pos 12, pos 13) in this image, a second pattern P2, present only once in the analyzed image, will be associated with its unique position pos21 in the image.
[0122] As in the embodiments illustrated in connection with Figure 4, a filtering 3322 (optional) can be performed on the patterns associated with an image (for example, a removal of empty words can be implemented before the detection of patterns). According to the example of Figure 5, the method can comprise a search 3421 for patterns common to several images and possibly their positions in these images (to detect a possible recurrence of their positions over the images). The identifications 3422 of patterns by each analyzed image can be used to distribute 342 the images into classes. As shown in Figure 5, the method can comprise obtaining a candidate distribution into “x” image classes, each identified class then being associated with a particular combination of patterns.The pattern combination associated with a class (and therefore the associated candidate distribution) may take into account in certain embodiments a proximity relationship between the positions of several patterns in the different images. For example, in certain embodiments, a proximity between absolute positions of at least two patterns in each of the at least two images where they are present, or relative positions of at least two patterns, in the images where they are present respectively, with respect to another element also present in these images (for example another pattern present in these at least 2 images) may be taken into account. For example, a recurrence of such proximity between several images may be taken into account, as illustrated by the blocks 811, 812, 813 of the images 810, 820, 830 in FIG. 8, to group the images.
[0123] As in the embodiments illustrated in connection with figure 4, this identification of classes can take into account for example a number or a percentage, minimum or maximum of images per class, so as to obtain relatively homogeneous classes (in terms of number of elements), of a desired number x of classes for example.
[0124] In the embodiments illustrated in connection with Figure 5, similarly to what has been explained in connection with the embodiments of Figure 4, the method can comprise an assignment of a confidence level to a pattern (for example as a ratio between the number of images of the class where an occurrence of the pattern is present and the total number of images analyzed where an occurrence of this pattern is present), and an assignment of a confidence score to a class and / or to a distribution.
[0125] According to certain embodiments of the method of the present application, the two analyses 331, 332 detailed above can be carried out sequentially and / or in parallel (as illustrated in FIG. 6), the method then comprising a selection 342 of at least one of the distributions obtained. This selection can take into account, for example, at least one selection criterion based, for example, on compliance with at least one configuration parameter, on a processing time and / or a memory occupation. In certain embodiments, the selection can take into account a confidence score assigned to at least one class of at least one of the distributions (as detailed above).
[0126] As indicated above, in certain embodiments, the method may comprise an association of a text label with at least one of the classes. This association may be optional in certain embodiments. The association of a label with a class may comprise an association of this label with all the images distributed in this class.
[0127] This labeling can for example be carried out by a user via a human-machine interface of said device 200.
[0128] In some embodiments, the method may comprise, prior to the analysis, obtaining at least one configuration data item, used to define a value of at least one parameter useful to the method of the present application. This may be, for example, at least one configuration data item accessible via at least one configuration file, or at least one configuration data item obtained via a user interface (or received from a third-party device). Such parameters may, in some embodiments, have default values, accessible via a storage means of the device for example, or be calculated automatically by the method of the present application. This step may be optional in some embodiments.
[0129] Obviously, the configuration data may vary depending on the embodiments. For example, in certain embodiments, at least one configuration data item may be obtained from the following data: a minimum number of desired clusters (for example, of the order of 5 to around ten clusters); a number of desired clusters (for example, of the order of around ten to a few dozen clusters, such as 12, 20, etc.); a maximum number of desired dusters (for example, of the order of around a hundred clusters, such as 100); a fixed, maximum, minimum, and / or average number of blocks dividing an image (as explained in more detail below), for example, a number of blocks of the order of one or a few hundred blocks (such as 99),
[0130] • an indication relating to a treatment to be carried out. This could be, for example, a Boolean specifying whether filtering on “empty” words (for their deletion) should be carried out or not;
[0131] • A minimum and / or maximum number of occurrences of visual items on which to base candidate distributions;
[0132] • An indication relating to the analysis and distribution to be carried out (selection of an analysis / distribution based on a number of occurrences of items, selection of an analysis / distribution based on the presence of patterns, or selection of both analysis / distribution (as illustrated in figure 6))
[0133] • A minimum confidence score to be respected (for example a minimum coefficient of 0.9 when the scores go from 0 to 1)
[0134] As discussed further, some of this configuration data (e.g., the number of classes into which to distribute the images, or the maximum number of such classes when the exact number is automatically defined by the method of the present application) may be optional in certain embodiments. These configuration parameters may intervene, for example, as criteria to be respected by a candidate distribution when selecting a distribution.
[0135] We now present, in connection with Figure 7 and by way of example, flow exchanges between some of the devices of the system 100 for the implementation of the method of the present application, for an application to learning a neural network. In the example illustrated, the method can for example be executed on the device 200. The device 200 receives 720 images from another device 710 (for example one of the terminals of the system 100), which it processes 300 as described in connection above with Figures 3 to 6 to distribute these images into classes. Information representative of this distribution can be provided 721 to the device 710.Such representative information may for example comprise at least one of the following elements: a number of classes, identifiers and / or labels of the classes, a number or percentage of images for at least one class, lists of images (or image identifiers) per class, lists of data structures each associating an image identifier with a class, images associated with metadata indicating their class, etc. The device 710 may for example name the classes as it wishes 722 and associate each distributed image with its class (so as to thus constitute a base of images annotated by their class). The device 710 may provide 723 these annotated images to a device 711 of the system 100, as a training data set for an artificial intelligence model 712. The device 711 may then carry out 724 the training of the model 712 using the annotated images received.In some embodiments, the parameters of the learned model may be provided 725 to the device 710, which may then use (infer) 726 the learned model on other images, to classify them.
Claims
CLAIMS 1. Image classification method implemented in an electronic device (200) and comprising: obtaining (310) images comprising a plurality of visual items; distributing (340, 341, 342) the images into a plurality of classes, based at least on a number of occurrences, in said images, of visual items extracted from said images.
2. Method according to claim 1 wherein said classes are identified during said distribution.
3. Method according to claim 2 comprising a labeling of the identified classes. Method according to claim 3 wherein said labeling is carried out automatically.
5. The method of claim 1 to 4 wherein said images divided into said plurality of classes are used to train a neural network to classify other images according to said plurality of classes.
6. Method according to claim 1 to 4 wherein said method is implemented locally to said electronic device.
7. Method according to claim 1 to 6 wherein said visual items are words, groups of words and / or graphic objects of said images.
8. The method of claim 7 wherein the method comprises replacing at least one first word, among said visual items, with at least one second word.
9. Method according to one of claims 1 to 8 wherein said method comprises obtaining the positions of said extracted visual items in said plurality of images.
10. Method according to one of claims 7 to 9 where said visual items are words and / or groups of words and where said method comprises, for an analyzed image of said plurality of images, an association with said analyzed image of the occurrence numbers of the distinct visual items extracted from said analyzed image.
11. The method of claim 10 wherein said method comprises deleting, in said distinct visual items associated with said analyzed image, at least one stop word.
12. Method according to one of claims 10 or 11 wherein said method comprises a deletion, in said distinct visual items associated with the analyzed images of said plurality of images, of the visual items associated with a number of images greater than a desired number of classes.
13. Method according to one of claims 10 to 12 wherein said method comprises: a grouping of the distinct visual items of said analyzed image taking into account the different values of said occurrence numbers of distinct visual items in said analyzed image; obtaining at least one candidate distribution of said images for a candidate value of the different occurrence numbers, taking into account distinct visual items common to at least two analyzed images grouped for said candidate value.
14. Method according to claim 9 wherein said distribution takes into account patterns relating to the positions of said visual items in said images of said plurality of images and present in at least two analyzed images of said plurality of images.
15. Method according to claim 14 wherein said method comprises: detecting at least one pattern in analyzed images of said plurality of images; associating said detected pattern with the analyzed images of said plurality of images in which an occurrence of said pattern has been detected; obtaining at least one candidate distribution of said analyzed images taking into account at least one number of occurrences of at least one pattern associated with at least two of said analyzed images.
16. Method according to claim 14 or 15 wherein said distribution takes into account the positions of said patterns in said analyzed images.
17. Method according to at least one of claims 13 to 16 wherein said method comprises an assignment of a confidence score to a class, taking into account an occurrence of at least one visual item and / or at least one pattern associated with at least one image of said class, in at least one other image of at least one other class.
18. Method according to one of claims 13 to 17 wherein said method comprises obtaining at least two candidate distributions and wherein said distribution is chosen, from among said candidate distributions, taking into account a desired number of classes and / or the confidence score assigned to at least one of the classes of said candidate distributions.
19. Method according to one of claims 1 to 18 comprising a modification of at least one of said images obtained before an analysis of said images.
20. Electronic device comprising at least one processor configured for image classification comprising: obtaining images comprising a plurality of visual items; a distribution of the images into a plurality of classes, based at least on a number of occurrences, in said images, of visual items extracted from said images.
21. Computer program comprising instructions for implementing, when the program is executed by a processor of an electronic device, a method for classifying images comprising: obtaining images comprising a plurality of visual items; dividing the images into a plurality of classes, based at least on a number of occurrences, in said images, of visual items extracted from said images.
22. Recording medium readable by a processor of an electronic device and on which is recorded a computer program comprising instructions for implementing, when the program is executed by the processor, a method for classifying images comprising: obtaining images comprising a plurality of visual items; dividing the images into a plurality of classes, based at least on a number of occurrences, in said images, of visual items extracted from said images.