Method for qualifying a training data set
The method enhances machine learning model training by segmenting images, incorporating user interaction, and leveraging crowdsourcing to create high-quality datasets efficiently, addressing the challenge of expert-dependent annotation in fields like medicine and astronomy.
Patent Information
- Application Number
- PCT/EP2025/065770
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-05
- Filing Date
- 2025-06-05
- Publication Date
- 2025-12-11
AI Technical Summary
Existing machine learning models require high-quality annotated datasets, particularly in fields like medicine and astronomy, where images lack a common ontological framework, necessitating domain expert labeling, which is scarce and costly.
A method involving image segmentation, labeling, and a two-step access process using graphical components to generate and qualify a training dataset, incorporating user interaction with identifiers and codes to enhance annotation precision and efficiency.
Improves the quality and efficiency of dataset annotation by leveraging crowdsourcing and user engagement, reducing reliance on expert labor and enhancing model training through precise labeling and verification processes.
Smart Images

Figure EP2025065770_11122025_PF_FP_ABST
Abstract
Description
[0001] METHOD FOR QUALIFYING A TRAINING DATASET
[0002] Scope of the invention
[0003] The invention relates to a method for qualifying a training dataset for learning a machine learning model. The invention also relates to a method for generating a crowdsourced training dataset for learning a machine learning model.
[0004] State of the art
[0005] Annotated data plays a crucial role in training machine learning models across various domains. By providing labels or annotations that describe the characteristics and categories of the data, annotated data enables machine learning algorithms to understand and generalize models from the provided examples. Whether in image recognition, natural language understanding, anomaly detection, or other tasks, annotated data serves as a reference for teaching models to recognize patterns and make relevant decisions. A high-quality annotated dataset is essential to ensure the accuracy and robustness of machine learning models, and often, its creation requires human intervention to guarantee the correctness of the annotations.
[0006] This is the case in the medical field, where new analytical techniques using artificial intelligence are being developed. In particular, machine learning algorithms are increasingly used to solve problems related to structure detection or the characterization of medical images. These algorithms require a considerable volume of annotation and training data.
[0007] One problem with image annotation, for example in the medical or astronomical fields, is that the images we want to annotate don't correspond to any structure commonly known to an individual. Indeed, classifying headlights, dogs, or cars relies on a known ontological framework because an individual can associate an image with an object whose ontology they understand. However, when we want to classify an image with patterns that are difficult to isolate because we are looking for an overall similarity or a similarity based on several criteria, only a domain expert is capable of accurately labeling the image.
[0008] One object of the invention is to propose a solution that improves upon the disadvantages of prior art solutions.
[0009] Summary of the invention
[0010] According to a first aspect, the invention relates to a method for qualifying a training dataset for learning a machine learning model, said method comprising:
[0011] ■ generation of a first set of images, called thumbnails, by a graphics component, corresponding to a segmentation of at least one given input image,
[0012] ■ display of the first set of images within a graphic element inscribed within a predefined geometric shape;
[0013] ■ generation of a target image representing the target of interest and associated with a label named "target label";
[0014] ■ selection of at least one first image from the first set using a selector from a graphical interface accessible via said graphical component;
[0015] ■ generation of a first image label associated with at least one first selected image;
[0016] ■ qualification of at least one first selected image by assigning the first label, so as to obtain a first set of labeled images intended to train a machine learning model.
[0017] All the embodiments relating to the labeling of images described according to the second aspect of the invention also relate to the first aspect.
[0018] Image segmentation refers to the operation of dividing an input image into different image segments. These segments can also be called fragments or extracts of the input image. When an input image extraction operation is performed, the input image is segmented, and conversely, when an input image is segmented, selecting a segmented image corresponds to its extraction.
[0019] The term "generation of a first set of images corresponding to a segmentation of at least one given input image" means: the generation of a first set of images resulting from a segmentation operation of at least one given input image.
[0020] The term "graphics component" refers to the software component that generates images. The term "graphics element" refers to the output of the display executed by the graphics component. In the following description, the "graphics component" will sometimes also be referred to as the "graphics element," particularly when the latter is activated and therefore configured to execute, for example, an action based on user interaction with the graphic element, such as image selection. According to a second aspect, the invention relates to a method for accessing a computing resource in a two-step sequence. This method comprises a first step for qualifying a training dataset for learning a machine learning model and a second step for unlocking access to a computing resource. The first step comprises:
[0021] ■ generation of a first set of images, called imagelets, corresponding to a segmentation of at least one given input image by a graphic component and comprising a first panel of images including at least one image likely to correspond to a target of interest and a second panel of images not corresponding to the target of interest;
[0022] ■ display of the first set of images within a graphic element inscribed within a predefined geometric shape;
[0023] ■ generation of a target image representing the target of interest and associated with a label named "target label";
[0024] ■ selection of at least one first image from the first set using a selector from a graphical interface accessible via said graphical component;
[0025] ■ generation of a first image label associated with at least one first selected image; ■ qualification of at least one first selected image by assigning the first label, so as to obtain a first set of labeled images intended to train a machine learning model, said graphic component generating an element enabling to engage a second step of said sequence of unlocking access to said computer resource, in which a code check is performed.
[0026] According to one embodiment, the selection of said at least one first image is associated with a first step in a sequence of access to a computer resource.
[0027] Selecting at least one image allows you to automatically move to the second step of the sequence.
[0028] According to one embodiment, the generation of a first set of images includes a first panel comprising at least one image representing a target of interest and a second panel of images not representing the target of interest.
[0029] According to one embodiment, the generation of the first set of images results in a random display by a graphic component of the set of images arranged within a graphic element that fits within a predefined geometric shape.
[0030] In one embodiment, the segmentation of a given input image by a graphics component is performed by randomly fragmenting the input image into fragments of identical resolution. When different input images are used, the selection and generation of the random arrangement can be performed so as to display images from different sources in the same window.
[0031] According to one embodiment, each image of the first set is extracted from a given input image containing an identifier.
[0032] In one embodiment, the qualification step includes qualifying a set of images by associating said set of images with the first label, so as to obtain a first set of qualified images. In another embodiment, the selection of at least one first image is associated with a first step in a sequence for unlocking access to a computer resource.
[0033] According to one embodiment, the graphic component displaying a graphic element within a predefined geometric shape and containing the images allows a second step of the unlocking sequence of access to said computer resource to be initiated, in which a code check is performed.
[0034] Advantageously, the method according to the invention encourages image labeling and motivates users to perform such actions by combining them with the entry of identifiers or codes. Indeed, when entering identifiers or codes, a user pays closer attention, which makes the annotation they perform in combination with the data entry action equally focused and deliberate. The probability that the resulting annotation will be correct, or more precise, is thus increased.
[0035] In one embodiment, each image in the first set is extracted from a given input image containing an identifier. Thus, the set of images in the first set can originate from a plurality of source images. Advantageously, each image extracted from a source image is associated with an identifier of the source image so as to allow subsequent labeling of the source image by reassociating it with the label of the extracted image fragment. Reassociation is possible if the image fragment included in the first set of images is associated with its source image, for example, through metadata associated with the fragment, such as an identifier.
[0036] In some embodiments, the process includes sending the first set of qualified images to a remote server.
[0037] In some embodiments, the images in the first set of images represent biological organisms or fragments of biological organisms, objects or fragments of objects.
[0038] In one embodiment, the images in the first image set represent fragments of elements whose image capture is defined at the microscopic scale. In another embodiment, the images in the first image set represent fragments of elements whose image capture is defined at the macroscopic scale.
[0039] According to one embodiment, the images in the first set of images are satellite images of portions or areas of space defining contours of shapes of terrestrial elements, such as trees, forests, waterways, urban areas, etc.
[0040] In some embodiments, the images in the first set of images represent stellar bodies, stars, galaxies and any other bodies present in space.
[0041] Thus, medical images can, for example, be annotated during the process according to the invention. Advantageously, large medical databases can be annotated and compiled in this way.
[0042] In some embodiments, the images in a panel of displayed images, fitting within a geometric shape, have identical or substantially equivalent resolution. Equivalent resolution is defined as a difference in resolution between two images of less than 10%.
[0043] CAPTCHA codes typically display images of varying resolutions to deceive a bot that could more easily analyze images of the same resolution. In other words, changes in resolution are one way to trick a bot in automated analysis. In this context, displaying images of the same or identical resolution provides a means to improve the training quality of a machine learning model configured to discriminate between image fragments corresponding to a given target.
[0044] In some embodiments, the second step of the unlocking sequence is successive to the first step of the unlocking sequence.
[0045] Thus, access to and completion of the second step of the unlocking sequence is conditional upon completion of the first step of the unlocking sequence.
[0046] In some embodiments, the process further includes a step for verifying the first label. In some embodiments, the process further includes:
[0047] - selection of a second image from the first panel associated with the target label, the selection allowing verification of the value of the first label.
[0048] Such a check allows, for example, the validation of the initial unlocking sequence of the process. It can also be used to verify that the user is entering information that is, in principle, correct for the other selected images.
[0049] Thus, the verification step allows the labeling to be validated or invalidated.
[0050] In some embodiments:
[0051] ■ at least one first selected image has been previously associated with a predefined label;
[0052] ■ The first label verification step includes a comparison of the first label with the predefined label.
[0053] Thus, advantageously, when information on at least one first selected image is known a priori, the labeling verification can be carried out in real time.
[0054] In one embodiment, the second step of the sequence involves authenticating a user with a data server. This second step can correspond to any authentication method, such as authentication involving the entry of a username or email address and a password, two-factor authentication, or authentication involving the entry of a code received on a terminal, etc.
[0055] In some embodiments, the step of verifying the first label is carried out by a human being.
[0056] Thus, advantageously, the labeling can be verified on a human scale, for example by experts.
[0057] According to different embodiments, the verification of said first label is carried out by computer equipment in an automated, supervised or unsupervised manner.
[0058] In some embodiments, the first label verification step includes a comparison with a prediction from a second machine learning model trained to generate a label prediction for the image. According to this embodiment, a model is trained to generate a statistic indicating whether an image belongs to a class within the image. One advantage of this solution is that it improves the training of a model whose statistics are not yet consolidated.
[0059] For example, if the second model assigns a class of images with a 60% confidence level for class assignment, the classification of the class generated by a user's image selection can be validated. In other words, beyond a given prediction threshold, a target class assigned to a user-selected image can be validated because the confidence in the predicted class is sufficient. One benefit is strengthening the training of a pre-trained machine learning model.
[0060] Conversely, if the second model assigns a 25% confidence level to the classification of a trusted image, and this image has been previously selected by a user, then the target class validation is not performed. In this latter case, a notification can be issued so that an expert can assign the target class to the selected image or not. This process allows for the qualification of images for which doubt remains regarding the target labeling. Such a process improves the training quality of a machine learning model.
[0061] In this case, the first learning model can result in a combination of the qualified returns of the image labels assigned to the images when the second model is implemented. According to one embodiment, the second model can be the first machine learning model.
[0062] In some embodiments, the selection, first label generation and qualification steps are implemented by a first user terminal, the process comprising, following the qualification step of the first set of images:
[0063] ■ display of the first set of images, called thumbnails, within a window defining or fitting into a predefined geometric shape on a display of a second user terminal; the window is, for example, a graphic element superimposed on a page or a digital document;
[0064] ■ generation of a second target image representing the first target of interest and associated with the label named "target label"; ■ selection of at least one image from said first set using a selector from a graphical interface accessible via a graphical component of the second terminal;
[0065] ■ generation of the first image label associated with at least one selected image;
[0066] ■ qualification of at least one selected image by assigning the first label in order to obtain a second set of labeled images,
[0067] ■ the first label verification step including a comparison of the first labeled images with the second labeled images.
[0068] Advantageously, the verification step can be performed automatically by pooling the results of the process according to the invention when implemented through the interaction of several user terminals with a remote server. When a minimum number of identical images generated on different terminals and selected by users correspond to the target and therefore have a target label, a verification step validates the qualification of the target label for each selected image. Thus, the target label assigned by a single user or terminal is a provisional or temporary label. When different users have qualified a selected image with the same target label, this label can be definitively validated.
[0069] In some other embodiments of the invention, several checks are performed on the same image before validating the target label of the selected image.
[0070] In some other embodiments of the invention, no human verification is performed before validating the target label of the selected image. Learning in this latter case can be more easily achieved and requires fewer validation steps.
[0071] In some embodiments, the step of selecting at least one first image includes a multiple selection of several images.
[0072] In some embodiments, the process includes, after selecting at least one initial image, a processing step using the graphical interface to obtain at least one processed initial image. Advantageously, this allows for simultaneous processing of the entire image set, so that the annotation action can be used to perform operations on the image set.
[0073] In some embodiments, the processing step includes a modification of a parameter of at least a first image, said parameter being chosen from contrast, sharpness, saturation, intensity, color, etc.
[0074] A third aspect of the invention relates to a method for generating a training database by supplying it with images from a population of users labeling images from a plurality of initial image sets for training a machine learning model, comprising:
[0075] - implementation, by a plurality of remote graphical components, of the process to qualify a training dataset for learning a previously described machine learning model, so as to obtain a plurality of first sets of qualified images;
[0076] - generation of the training database by union of the first sets of qualified images from said plurality of first sets of qualified images.
[0077] Advantageously, when the method for qualifying a training dataset includes a processing step, so as to obtain a plurality of at least one first processed image, the method for generating a training database by crowdsourcing further includes:
[0078] - determination of a treatment called global treatment from the plurality of at least one first processed image.
[0079] Thus, advantageously, by pooling the implementations of the process for qualifying a training dataset, the process for generating a training database through crowdsourcing makes it possible to determine and define a global treatment based on all the implementations of the process for qualifying a training dataset. In some embodiments, the treatment includes averaging.
[0080] Thus, advantageously, it is possible to define a processing representative of a plurality of processing carried out by remote users, which can serve as a standard for input data of the machine learning model.
[0081] In some embodiments, the process includes:
[0082] ■ saving the global processing, so that when the machine learning model is run, the execution of the machine learning model includes a preliminary step of applying the processing to an input image of the machine learning model.
[0083] Thus, the process for generating a training database by crowdsourcing makes it possible to generate a processing from the set of implementations of the process to qualify a training dataset that can then be applied to an input image of the machine learning model when it is executed.
[0084] According to one embodiment, the first set of images is obtained by segmenting a plurality of input images. This embodiment is compatible with every aspect of the invention, in particular the first, second, and third aspects.
[0085] According to a fourth aspect, the invention relates to a preliminary method for extracting images from several input images in order to produce thumbnails or tiles which will be integrated into the first set of images ENS1, also called captcha image, so as to allow labeling of said thumbnail or tile images.
[0086] According to this fourth aspect, the invention relates to a method for extracting small images of interest from at least one input image in a corpus of images from a domain of expertise comprising:
[0087] ■ Receiving parameters defining an objective of a domain of expertise, at least one parameter of which allows defining a type of images;
[0088] ■ Selection of at least one input image from a given image corpus from an objective of an expertise domain, said objective of an expertise domain including the definition of an image type, called a type parameter;
[0089] ■ Receiving a resolution parameter, a dimension parameter, and characteristic data from at least one area of interest in the input image;
[0090] ■ extraction of at least one image corresponding to a portion of the at least one input image within the area of interest from the at least one characteristic data, the dimension parameter and the resolution parameter;
[0091] ■ generation of a graphic element, called an image captcha, from a plurality of small images, also called the first set of images, extracted and aggregated within the graphic element, said small images coming from at least one of these image sources:
[0092] ■ of different input images from a corpus of images from a given area of expertise and / or;
[0093] ■ of a set of synthetic images, that is to say artificially generated from other images and / or;
[0094] ■ of a set of images whose classification label is determined and / or;
[0095] ■ of a set of images that have already been validated by another user in another first set of generated images, i.e. another image captcha.
[0096] This fourth aspect of the invention can also be implemented to constitute a preliminary step of the first aspect of the invention, the second aspect of the invention, or the third aspect of the invention.
[0097] According to one embodiment, the process includes a preliminary step comprising a selection of at least one input image from a given image corpus from an objective of an area of expertise, said objective of an area of expertise comprising the definition of an image type, called a type parameter.
[0098] According to one embodiment, the generation of the first set of images includes an extraction of at least one image corresponding to a portion of this input image within a defined area of interest from at least one characteristic data of the area of interest, said extraction including the determination of at least one dimension and at least one resolution deduced from a dimension parameter and a resolution parameter.
[0099] According to one embodiment, at least one characteristic data point of the area of interest possibly includes:
[0100] ■ a position parameter and / or;
[0101] ■ data describing an element of the image so as to produce area of interest markers in the input image from a software component and / or;
[0102] ■ pixel coordinates in the input image and / or;
[0103] ■ the definition of a bounding box in the image.
[0104] According to one embodiment, a training software component is configured to query an image database or a training configuration in order to generate at least one first characteristic parameter of a set of training images to be produced, said first characteristic parameter comprising at least one parameter among which: {resolution parameter, dimension parameter, type parameter, class parameter, parameter characterizing an area of interest or an element of the image, a number of images}.
[0105] According to one embodiment, a technology software component is configured to query a technology configuration of a machine learning function in order to generate at least one second characteristic parameter of a learning function, said second characteristic parameter comprising at least one parameter among the following: {resolution parameter, dimension parameter, type parameter, class parameter, parameter characterizing an area of interest or an element of the image, a number of images}.
[0106] According to another aspect, the invention relates to a system comprising at least one server and one terminal communicating via a data network and comprising computers to implement the method of the invention.
[0107] Brief description of the figures
[0108] Other features and advantages of the invention will become apparent from reading the detailed description that follows, with reference to the attached figures, which illustrate: Fig. 1: an example of a system configured to implement a method for qualifying a set of training data according to the invention;
[0109] Fig. 2: an example of a device configured to implement the method for qualifying a set of training data according to the invention;
[0110] Fig. 3: an example of a set of steps that can be carried out to implement the method for qualifying a set of training data according to the invention;
[0111] Fig. 4: a representation of an example of a set of images generated and displayed by a graphics component of the device of figure 2 and used in the process to qualify a training data set according to the invention;
[0112] Fig. 5: a second representation of the example image set illustrated in figure 4;
[0113] Fig. 6: a representation of an example of a set of images from different sources, generated and displayed by a graphics component randomly within the window;
[0114] Fig. 7: an example of a system configured to implement a method for generating a crowd-supplied training database according to the invention;
[0115] Fig. 8: an example of a set of steps that can be carried out to implement the method for generating a crowd-supplied training database according to the invention;
[0116] Fig. 9: A schematic representation of a training database generated by the method for generating a crowd-supplied training database according to the invention, at different times.
[0117] Fig. 10: An example of the steps in a process for extracting small images of interest from at least one input image in a corpus of images from a domain of expertise. These steps can be performed prior to the process in Figure 3 or independently.
[0118] Description of the invention
[0119] In many fields, new analytical techniques using artificial intelligence are being developed. As a result, training machine learning models often requires a significant volume of annotation and training data. This is the case in certain medical disciplines, such as histology, where the analysis of histological sections is increasingly automated. However, in this field, collecting training data is difficult because it is currently provided by trained users—students or qualified pathologists—who are scarce and have limited availability. These individuals are responsible for creating precise annotations, a large-scale, time-consuming, and costly undertaking.
[0120] In the problem of obtaining a larger volume of annotated data for the purpose of training machine learning models, one of the objects of the invention aims to promote data annotation actions, by associating such actions with frequent and daily actions, by conditioning the latter on one or more annotation actions.
[0121] One aspect of the invention relates to a method 100 for qualifying a training dataset for learning a machine learning model, illustrated in Figure 3. Another aspect of the invention relates to a method 200 for generating a training database by crowdsourcing for learning a machine learning model, illustrated in Figure 8.
[0122] Process 100 can be implemented using a system 10. Figure 1 shows an example of a set of elements of system 10. Figure 1 more specifically represents a data network NET1, which can be the internet. A user terminal T1 provides access to a remote server SERVI.
[0123] The user terminal T1 can send and receive information to and from the remote server SERVI using electronic communication systems.
[0124] The SERVI remote server stores a set of ENS0 images, also known as input images. This set of input images is also called an image corpus when associated with a domain of expertise. The images in the ENS0 image set are intended to train a machine learning model and require annotation, i.e., the assignment of a label. Images are defined as any element comprised of pixels, contained within a closed area, and capable of being displayed on a medium or device such as a screen. Images can be 2D, 3D, animated, or static. Images can be extracted from other images. Images can be ultrasound, MRI, CT scans, telescope or microscope images, and more generally, images from any optical, electromagnetic, or acoustic sensor.The images can be derived from a reconstruction step of a set of data acquired from one or more sensor(s).
[0125] As an example, the images in the ENSO image set are medical images, and the machine learning model is configured to classify these medical images and predict the presence of specific structures.
[0126] In the latter case, ENSO images can be grouped according to different image corpora associated with different areas of expertise, such as histology images, MRI images, CT images, ultrasound images, confocal microscopy images, etc.
[0127] In another example, the images in the ENSO image set are astronomical images representing astronomical objects, and the machine learning model is configured to classify these astronomical images and predict the presence of specific structures.
[0128] Specific structures can correspond to shapes, contrasts, colors or intensities, thicknesses of structures, patterns, dimensions of structures or a combination of these criteria.
[0129] Figure 2 shows a schematic diagram of the components of an example of the T1 user terminal. The T1 user terminal can be implemented as a single hardware device, for example, as a desktop PC, laptop, PDA, smartphone, smartwatch, server, or console, or it can be implemented across separate, interconnected hardware devices linked by one or more communication links, with wired and / or wireless segments. The T1 user terminal can, for example, communicate with one or more cloud computing systems, servers, or remote devices to implement the functions described herein for the device in question. The T1 user terminal can also be implemented itself as a cloud computing system.
[0130] As shown in Figure 2, the user terminal T1 comprises a computer, this computer including a memory 11 for storing program instructions that can be loaded into a circuit 12 and adapted to cause the circuit to execute steps of the process 100 illustrated in Figure 3, described below, when the program information is executed by the circuit. The memory can also store data and information useful for executing the steps of the present invention as described below.
[0131] Circuit 12 could be, for example:
[0132] - a processor or processing unit adapted to interpret instructions in a computer language, the processor or processing unit being able to understand, be associated with, or be attached to a memory containing the instructions, or
[0133] - the combination of a processor / processing unit and a memory, the processor or processing unit being adapted to interpret instructions in a computer language, the memory containing said instructions, or
[0134] - an electronic circuit board in which the steps of the invention are described in silicon, or
[0135] - a programmable electronic chip such as an FPGA chip (for "Field-Programmable Gate Array").
[0136] Memory 11 may include random access memory (RAM), cache memory, non-volatile memory, backup memory (e.g., programmable or flash memory), read-only memory (ROM), a hard disk drive (HDD), a solid-state drive (SSD), or any combination thereof. The ROM of memory 11 may be configured to store, among other things, an operating system and / or one or more computer program codes for one or more software applications. The RAM of memory 11 may be used by circuit 12 for temporary data storage.
[0137] The computer may also include an input interface 13 for receiving input data and an output interface 14 for providing output data. Examples of input and output data will be provided later.
[0138] To facilitate interaction with the computer, a 15-inch screen and a 16-inch keyboard can be provided and connected to the computer circuit.
[0139] Figure 3 is a flowchart representing an example of a set of steps that can be performed to implement process 100 to qualify a training dataset for learning a machine learning model.
[0140] Suppose a user of user terminal T1 wishes to access a computing resource via user terminal T1. In one example, the computing resource is user terminal T1 itself when access is locked, for example, due to prolonged inactivity. In another example, the computing resource is a secure application installed on user terminal T1 that can be launched from user terminal T1 by entering credentials. In yet another example, the computing resource is a secure application installed on a server and accessible from a data network and user terminal T1. The secure application can then be launched from user terminal T1 by entering credentials.
[0141] Advantageously, the method 100 according to the invention is part of a sequence for unlocking access to the computer resource via the user terminal T1. The unlocking sequence comprises a first step and a second step, which can be carried out successively in a predefined order or in reverse order, or simultaneously, in order to unlock access to the user terminal T1.
[0142] According to one embodiment, the method of the invention comprises a first step of generating GEN1 a set of images ENS1 from at least one input image. The set of images ENS1 are advantageously thumbnails, patches, or tiles intended to be integrated into a graphic element comprising several display areas for such thumbnails, also called an "image captcha".
[0143] Advantageously, the ENS1 image set groups images that relate to or are associated with an initial semantic structure. This initial semantic structure may itself be linked to the domain of expertise and / or the image corpus from which the input data is extracted. For example, the images in the ENS1 image set have metadata relating to their provenance, their source, a shared folder, or a label or tag that allows them to be categorized.
[0144] According to one example, the images in the first set ENS1 are associated with metadata labeling a batch of images relating to a cohort of patients with a type 1 and / or type 2 disease.
[0145] As an example, each input image and each image extracted from that input image can be associated with a semantic structure, such as metadata, that links the input image or extracted image to its original semantic structure. For example, the original structure could be a pathology, an elementary lesion, a transcriptomic signature, a genetic alteration, or any other semantic structure from a different domain.
[0146] Within the scope of the invention, an image captcha is considered without restriction to the implementation of the Turing test. It is a graphical element such as a graphical window or a portion intended to display a graphical element within which a plurality of images are displayed to be annotated by a user by clicking with a selection tool.
[0147] According to one embodiment, the input image is cut or segmented into a plurality of images defining images corresponding to thumbnails representing only a part of the input image.
[0148] According to another embodiment, the set of ENS images corresponds to as many input images
[0149] According to another embodiment, the image set ENS1 corresponds to subsets of images from different input images.
[0150] According to an example, a first set ENS1 of images generated in a display window includes a first panel PA1 likely to include at least one image representing a target of interest CIB1 and a second panel PA2 of images not representing the target of interest.
[0151] Figure 4 represents an example of the first ENS1 set of images.
[0152] In Figure 4, the first panel, PA1, is represented with a white fill, and the second panel, PA2, is represented with a fill of horizontal hatched lines. In this case, it is assumed that the images in the first panel, PA1, are those that the user is expected to select and annotate based on that selection. The user selects these images because they are representative of the target image, IMC.
[0153] The target image can also be called the reference image.
[0154] However, the first PA1 image panel is not known in advance except when certain labeled images are used to validate that the user is actively seeking to identify the actual representative images of the target BMI image. The second PAI panel is therefore defined after the user's selection.
[0155] For example, input images that are likely to include images corresponding to the target image, and therefore belong to the first PA1 panel, can be pre-identified using a confidence statistic. The aim is then to validate whether this pre-identification is correct or not.
[0156] The first ENS1 image set, for example, corresponds to the result of subdividing an IMO image, configured to be annotated, from the ENSO image set stored by the remote SERVi server. The first ENS1 image set can result from segmenting a given input image into thumbnails, each representing a portion of the given input image. The first ENS1 image set can also result from extracting thumbnails from several ENSO input images, each representing a portion of one or more given input images.
[0157] The thumbnails are also called images, tiles, or patches in the technical literature associated with image captchas. Therefore, these terms will be used interchangeably to refer to the images in an image captcha. In the request, an "image captcha" and a "graphic element" containing activatable or selectable images are referred to as the same object.
[0158] In one example, the resolution and / or size of the input image is determined so that an image matching the target image fits within a thumbnail. In another example, the resolution and / or size of the input image is determined so that an image from the first set matching the target image fits within a maximum of four thumbnails. Other configurations allow the input image to be adapted so that the elements to be selected—that is, a target image—fit within a maximum number of adjacent thumbnails. This ensures that the display window dimensions are optimized for displaying a target image.
[0159] According to one embodiment, the dimension of a target IMC image is estimated so as to ensure that the display of an element corresponding to a target is contained within a thumbnail.
[0160] Figure 6 illustrates a scenario where the target image IMC represents a cell, tissue, or microorganism with a unique structural, deformed, colored, or regular characteristic. This target image IMC is displayed at the top of the image in a dedicated area that can be supplemented with a description in the form of instructions to help a user identify images within the ENS1 set that contain an element resembling the target IMC. In this example, the elements being sought are smaller than the dimensions of an image generated within the ENS1 set in a window.
[0161] According to one embodiment, a target has dimensions less than 50% of the size of a thumbnail. One advantage is to minimize cases where such a target would be positioned on a border of the thumbnail, that is to say on a border of an image of the ENS1 set.
[0162] When the images of the ENS1 set form a coherent image, that is to say it is a single image decomposed into 9 pieces for example, or into 16 pieces, the targets contained in the images and positioned at the limit of a thumbnail can spill over into an adjacent thumbnail.
[0163] When the images in the ENS1 set form an incoherent image—that is, a plurality of images assembled randomly and comprising, for example, 9 pieces, or 16 seemingly independent pieces, possibly originating from different source images—the overflow of an element resembling a target at the boundary has no reason to appear on an adjacent image in the ENS1 set. In this latter case, it is understood that the dimensions of a target are a priori smaller than those of an image in the ENS1 set so that the user can recognize a target in the images presented.
[0164] The invention relates to an embodiment in which images are randomly generated in a window and whose dimensions allow visibility of target-like elements on several images of the set ENS1.
[0165] In one embodiment, each image generated in the ENS1 set comes from a different source image. More generally, this case relates to an embodiment in which the images of each thumbnail of an image displayed in the graphic element come from different source images.
[0166] According to one embodiment, the images of the ENS1 set are extracted from a larger original image that has been segmented.
[0167] The target of interest C IB 1 is, for example, a biological structure or a fragment of a biological organism within a biological tissue. Non-limiting examples of biological structures include: a gland, cross-section, a lymphocyte, a surface epithelium, apoptosis, mitosis, a lymphoid mass, a villus, a Helicobacter-type germ, a red blood cell, a mucus cell, an enterocyte nucleus, a plasma cell, a neutrophil (PMN), an eosinophil (EPN), a lymphocytic cryptitis, an abscess, an ulceration, a cryptic abscess, and a granuloma.
[0168] In another example, the target of interest CIB1 can be an astronomical object. Non-exhaustive examples of astronomical objects include: a star, a planet, a nebula, a galaxy, a quasar, a black hole.
[0169] According to another example, the target of interest CIB1 can be a microorganism-type object, an object, a landscape, etc.
[0170] The GENi generation step is advantageously implemented by components of the remote server SERVi. For example, the GENi generation step is advantageously implemented by a processing unit of the remote server SERVi. By processing unit, we mean an electronic component or a plurality of electronic components including, for example, a computer, comprising a set of at least one processor, and optionally a memory operationally coupled to the computer. According to a first example, the first set ENS1 is transmitted to the user terminal T1 in the form of a random display of thumbnails.
[0171] According to a second example, the first set ENS1 is transmitted to the user terminal T1, in the form of an ordered display of thumbnails arranged so as to maintain overall consistency of the input image which has been segmented.
[0172] The display of all generated images is inscribed within a geometric shape on screen 15 of user terminal T1.
[0173] In a second generation step (GEN2), a target image (IMC) representing the target of interest (CIB1) is generated and displayed on screen 15. The target image (IMC) is associated with a label called the "target label". The target label can be: "presence of the target of interest (CIB1)" within an image or a group of images.
[0174] According to one embodiment, during the second generation step GEN2, an INSTR instruction can also be generated and displayed near the target image IMC. For example, the INSTR instruction can prompt the selection of images showing the target of interest Cl B 1.
[0175] Non-exhaustive examples of instructions include:
[0176] - when the target of interest CIB1 is a cross-section gland: "Can you find one or more daisies?",
[0177] - when the target of interest CIB1 is a lymphocyte: "Can you find this little bead?",
[0178] - when the target of interest CIB 1 is a surface epithelium: "Can you identify the boundary of the great white?",
[0179] - when the target of interest CIB1 is apoptosis: "Can you find those very dense little grains?",
[0180] - when the target of interest Cl B 1 is a mitosis: "will you be able to find this badly wound ball of yarn?",
[0181] - when the target of interest CIB1 is a lymphoid cluster: "can you identify this pile of marbles?",
[0182] - when the target of interest CIB 1 is a villus: "can you spot one or more glove fingers?",
[0183] - when the target of interest CIB 1 is a Helicobacter-type germ: "can you spot this nasty little bug?",
[0184] - when the target of interest CIB1 is a red blood cell: "Look, it's a red blood cell! Can you identify them here?"
[0185] - when the target of interest CIB1 is a mucus-secreting cell: "can you identify this cell with the large belly?",
[0186] - when the target of interest CIB1 is an enterocyte nucleus: "a bit like a lychee pit? Can you spot them?",
[0187] - when the target of interest C IB 1 is a plasma cell: "It too has a big belly but it is not transparent, find it!"
[0188] - when the target of interest CIB1 is a PNN: "can you spot this king of contortion?",
[0189] - when the target of interest CIB1 is a PNE: "he always has two plump cheekbones on his pinkish-orange cheeks, will you be able to spot him?",
[0190] - when the target of interest CIB1 is a lymphocytic cryptitis: "these little beads don't belong in this daisy! Find the image that resembles it",
[0191] - when the target of interest CIB1 is an abscess, an ulceration, a cryptic abscess: "These contortionists must not play together! Not in the yard nor on a daisy! Find these scoundrels!"
[0192] - when the target of interest CIB1 is a granuloma: "can you find this cluster of pink cells?".
[0193] According to this latter embodiment, the instruction may correspond to a non-limiting popularization of structure to be identified in order to facilitate annotations and the choice of thumbnails.
[0194] Depending on the implementation, the choice may be limited to selecting an image, or it may involve drawing, outlining, delimiting, masking, coloring, or covering the structure to be visualized, in order to refine the annotation. In the latter case, an automatic analysis of the pixels of the selected image allows for the recording of metadata relating to the position, size, or shape of a structure within the image.
[0195] In one embodiment, instructions are not generated. For example, when images are randomly generated within the window containing all ENS1 images, an instruction is not necessarily generated. It should be noted that, within the scope of the present invention, instructions potentially implemented in the solution of the invention are generated optionally depending on the use case. In one embodiment, the images are 3D images, that is, images containing depth information.
[0196] In another example, the images are animated, for instance, images containing a short animated sequence such as a video. This can be useful for characterizing movement, flow, or a process, such as mutations of proteins, organoids, or other biological structures, or the aging of cells or other biological elements. For this purpose, animated images containing a chronological sequence of images can be used.
[0197] As an example, ultrasound images of the echographic type can be used to form the images of the ENS1 set displayed in the window.
[0198] According to an embodiment depending on the specific case, the resolution and size of the images of the first set ENS1 can be adapted so that a target likely to be present in the input image is included in one or more images of the first set ENS1.
[0199] In a SEL selection step, the user selects at least one initial image IM1 from the first set ENS1 using a selector on the user terminal T1. The first selected IM1 image is shown as a dotted fill in Figure 5. The selector is accessible via screen 15. For example, the selector is a pointer that can be moved on screen 15 using a mouse. The user may be prompted to select multiple images. This can occur when the target is present in several images of the set ENS1 or when the target spans multiple images.
[0200] The target(s) being sought are not necessarily identical to the target represented in the IMC target image. Indeed, variations in shape or color may be tolerated because the aim is precisely to label images of the same class that share common characteristics. One advantage of the invention is precisely to help an individual, through the display of the target image, to recognize similar images with common characteristics.
[0201] The result of the SEL selection step is sent to the remote server SERVI. In a third generation step, GEN3, a label called the first label LB1 is generated and associated with at least one first image IM1. In some embodiments, the third generation step, GEN3, is advantageously implemented by components of the remote server SERVI. For example, the third generation step, GEN3, is advantageously implemented by a processing unit of the remote server SERV1. In other embodiments, the third generation step, GEN3, is advantageously implemented by components of the user terminal T1, such as circuit 12.
[0202] In a QUAL qualification step, the first set of ENS1 images is associated with the first LB1 label, so as to form a first set of qualified ENS1q images.
[0203] In some embodiments, in which the QUAL qualification step is implemented by one or more components of the user terminal T1, the first set of qualified images ENS1q is sent to a remote server such as the remote server SERVI.
[0204] In other cases, the first label LB1 is sent to the remote server SERVI, so the qualification step is performed by one or more components of the remote server SERVI. In other words, the first label LB1 is assigned to the first set ENS1. For example, the qualification step QUAL is advantageously implemented by a processing unit of the remote server SERVI. The first qualified set ENS1q is stored in memory of the remote server SERVI. Alternatively, data encoding the association of the first label LB1 with the first set ENS1, called annotation data D1, is stored in memory of the remote server SERVI.
[0205] Once the QUAL qualification step has been completed, the first step of the sequence for unlocking access to the user terminal T1 is validated, this step is noted as VAL in Figure 3, and the second step of the sequence for unlocking access to the IT resource is accessible.
[0206] Different implementation variations allow for the validation of the first step of the sequence. As stated previously, in one implementation, the selection and validation (VAL) of the selected images allows access to the second step of the sequence. Access to the second step can be achieved without any selection validation.
[0207] In a second embodiment, a check is performed on at least one selected image to validate the first step. This first step could include generating an image of the ENS1 set whose label is known—in this case, the target label. The check can be performed automatically; for example, if the result is contained in metadata associated with the image, no exchange with a server is necessary to validate the step. Alternatively, an exchange with a remote server is used to verify that the image selection is correct.
[0208] If the set of images in the ENS1 set is coherent, that is to say that each image in the ENS1 set forms a portion of a larger image formed by different pieces representing the thumbnails, a priori, no image in the ENS1 set will be able to have a known label because by definition we seek to label unknown images.
[0209] It is understood that the embodiment in which a check of an image of the set can be carried out because its label is known is more suited to the mode in which the images entered in the window within the set ENS1 are generated randomly and come from at least one given source.
[0210] One advantage is the ability to introduce an image whose label is known in order to perform a verification.
[0211] Thus, this validation ensures that the user has selected a coherent set of images. In this case, at least one image from the set is generated at a known position, allowing the user's selection to validate that the image was indeed selected. Consequently, the other images selected by the user were chosen with the intention of providing an element of truth from the user's perspective. Conversely, when the image whose label is known has not been selected, this generates an indicator that the selection is not qualified. Validation for proceeding to the second step of the sequence can be performed only if the image whose label is known has been selected, or conversely, validation can be initiated regardless of the selection outcome.According to a third example, an input image containing images previously labeled by another user can be used by the user of terminal T1 to corroborate the selection results. One advantage is that the selection can be qualified directly in the first step, validating this first step before proceeding to the second. Similarly, proceeding to the second step can be done independently of checking the labels of the selected images, or conversely, it can only occur if the check results in a correct verification.
[0212] The user must complete the second step of the IT resource access unlock sequence to unlock access to the IT resource. In some embodiments, this second step includes the user entering and verifying a code.
[0213] In some embodiments, process 100 comprises, successively to the selection of the first image from the first PA1 panel associated with the target label, a step in which the user selects a second image. The second image is associated with the target label.
[0214] In some embodiments, process 100 includes a step of verifying the first LB1 label.
[0215] Advantageously, information associated with at least one first selected image (IM1) is known a priori. In other words, at least one first selected image (IM1) can be associated with a predefined label. Thus, the step of verifying the first label (LB1) includes a comparison of the first label (LB1) with the predefined label, which can be performed automatically and in real time.
[0216] According to one variant, the verification step of the first LB1 label is performed by a human. Thus, when the ENS0 set is a set of medical images, such as histological sections, the verification step can be carried out by a medical expert such as a pathologist.
[0217] According to a second variant, the verification step of the first label LB1 includes a comparison of the first label LB1 with a prediction from a second machine learning model. For example, the second machine learning model may be partially trained. The second model may correspond to the first model whose prediction quality we wish to improve. In this case, a prediction threshold can be used to verify that an image selected by the user in the first step is consistent with a minimum prediction threshold of the second machine learning model.
[0218] According to a third variant, the verification step of the first LB1 label can be followed by the reception of multiple sets of qualified images from several different user terminals communicating with the remote SERVI server. More specifically, process 100 can include, following the qualification step of the first set of ENS1 images, the implementation of the following steps:
[0219] ■ generation of another set of ENS2 images, at least one of which is likely to represent a target.
[0220] ■ said generation of said other set of images resulting in a random or ordered display by a graphic component of another user terminal of the set of images arranged within a geometric shape and defining a graphic element;
[0221] ■ generation of a second target image representing the second target of interest which may be a different image from the first image or the same but which includes the same target as the first image;
[0222] ■ selection of at least one image from said other set by means of a selector in a graphical interface accessible via said graphical component;
[0223] ■ generation of an image label associated with at least one other selected image corresponding to the target label;
[0224] ■ qualification of at least one selected image by assigning the first label LB1 or another label in order to obtain a second set of labeled images:
[0225] ■ verification of the first LB1 label including a comparison of the first labeled images with the second labeled images or a comparison of the labels with each other.
[0226] The steps above can be repeated for a plurality of other user terminals, so as to obtain a plurality of other sets of qualified images. The first label LB1 verification step then includes a comparison of the first label LB1 with each of the second labels or a comparison of the images with each other. As an example, the first label LB1 is validated when it matches a minimum number of second labels.
[0227] In some embodiments, process 100 includes, after the step of selecting at least one first image IM1, a TRAIT processing step using a graphical user terminal interface T1, so as to obtain at least one first treated image IM1 trait.
[0228] During the TRAIT processing step, an instruction is advantageously presented to the user, enabling them to perform an operation on at least one first image IM1 or on the first set ENS1. The operation is, for example, a change to a parameter of at least one first image IM1 or the first set ENS1. The parameter is, for example, the contrast, sharpness, saturation, intensity, or a color of at least one first image IM1 or the first set ENS1. Advantageously, the user can modify the parameter with a movable cursor displayed on screen 15.
[0229] According to one embodiment, the manipulation can correspond to a selection of a group of pixels, a selection of at least one image, an association of an image with a keyword, called a "tag".
[0230] Another aspect of the invention relates to a method 200 for generating a crowdsourcing training database for training a machine learning model, as shown in Figure 8. The output of method 200 is a training database B(Tf), where Tf is the completion date of the implementation of method 200. In some embodiments, Tf is a fixed date. In other embodiments, Tf represents a rolling time, corresponding to an evaluation point in the training database generated by method 200.
[0231] Advantageously, process 200 takes place over a time window [T, Tf], where T is a fixed date corresponding to the start date of the process 200 implementation. The training database B(Tf) is built progressively over the time window [Tj, Tf], so that B(t) denotes the training database at a time t within the time window [T, Tf]. Process 200 can be implemented using a system 20. Figure 7 shows an example of a set of elements of system 20. More specifically, Figure 7 represents a data network NET1, which could be the internet. Several user terminals, including user terminals Ti, T2, T3, ..., TN, are connected to the data network NET1, so that they can communicate with the remote server SERVI.
[0232] Information can be exchanged between the remote server SERVI and each of the user terminals T1, T2, T3...TN using electronic communication systems. Each user terminal T1, T2, T3...TN includes components similar to those of user terminal T1 described and illustrated in Figure 2. In particular, each user terminal T1, T2, T3...TN includes a graphical component such as a screen.
[0233] As before, the remote SERVI server stores a set of ENSO images. The images in the ENSO image set are intended to train a machine learning model and require annotation, i.e., the assignment of a label.
[0234] Process 200 comprises, during the time window [Ti,Tf], in a first step E1, one or more implementations of Process 100 described previously by a plurality of users, each user being associated with one of the set of user terminals T1, T2, T3...TN, such that at a current time t of the time window [Tj,Tf], to a user associated with the user terminal Ti, corresponds a plurality of qualified sets Pqi(t) stored in a memory of the remote server SERVI. In a second step E2, the training database B(t) at time t comprises the union of the plurality of qualified sets received and stored in a memory of the remote server SERVI.
[0235] For example, for a user associated with user terminal Ti, where i is an integer between 1 and N, at a current time t within the time window [Tj,Tf], a plurality of qualified sets Pqi(t) has been stored in the memory of the remote server SERVI. Thus, at time t, the training database B(t) comprises the union of the plurality of qualified sets Pq1(t), Pq2(t)...PqN(t). Figure 9 schematically illustrates the training database generated at time t, B(t), and the final input database obtained at time Tf, B(Tf). In some embodiments, when a processing step has been implemented by one or more user terminals from among the set of user terminals Ti, T2, T3...TN, process 200 includes a step for determining a processing function, referred to as global processing.According to an example, the overall processing is an average of the processing steps implemented by the user terminal(s) across all user terminals T1, T2, T3...TN. Thus, when the processing steps are image colorings by different users each associated with a user terminal, the overall processing can consist of a normalization of all the colorings.
[0236] Advantageously, the global processing is saved, for example in a memory of the remote SERVI server. Thus, when the machine learning model is executed, receiving an input image, its execution includes a preliminary step of applying the global processing to the input image.
[0237] Figure 10 illustrates another aspect of the invention, corresponding to a method for extracting relevant image tags as a preliminary step to generating the image captcha. Figure 10 describes the various steps or substeps aimed at extracting images from a panel of images according to a specific objective within a domain of expertise, a learning strategy based on the existing training data, and the learning technology that will be used to train a given model.
[0238] Defining the objective of an area of expertise
[0239] The process of extracting images of interest includes a first step aimed at receiving parameters defining an objective of an area of expertise, at least one parameter of which allows a type of image to be defined.
[0240] For example, in the medical field, image types can include, for instance, histology images for observing tissue sections or cytology images for cellular observations, on a micrometer to sub-micrometer scale; MRI images for visualizing soft tissue, brain, or organs on a millimeter scale; CT images to produce 3D cross-sectional images of the body, bones, or organs on a millimeter scale; ultrasound images on a millimeter to submillimeter scale; confocal microscopy images on a sub-micrometer scale; and electron microscopy images on a nanometer scale.
[0241] For example, in the field of materials characterization, the types of images can be, for example, electron microscopy images for observing crystalline structures at the nanometer or sub-nanometer scale, atomic force microscopy images for observing surface relief at the atomic scale, X-ray diffraction images for observing crystalline structure and grain orientation, X-ray tomography images called micro-CT for observing the internal 3D structure of a sample, analyzing porosity, cracks, or images produced by near-field optical imaging or Raman spectroscopy images to produce chemical maps of materials.
[0242] As an example, in the field of astronomy, image types can include, for example, telescopic images of galaxies, nebulae, stars with arcsecond resolution, radio astronomy images of nebulae, protoplanetary disks with a resolution ranging from minutes to milliarcseconds, images obtained by infrared imaging, X-ray imaging, spectro-imaging or coronagraphic imaging which can cover a resolution from arcseconds down to sub-arcseconds.
[0243] The objective of the area of expertise can therefore allow for the automatic determination of image resolution and size, based on a parameter specific to that objective. It is worth recalling that resolution corresponds to the pixel size, that is, the number of microns in the pixel, and dimension corresponds to the number of pixels in the image.
[0244] When the area of expertise is defined by a parameter designating the type of image, for example by designating an optical device, a resolution and the nature of the objects likely to be present in the acquired image, the method of the invention makes it possible to select an appropriate corpus of images.
[0245] When only one image corpus is present, the parameter can be used to select a subset of images having an attribute specific to the desired objective.
[0246] Furthermore, according to one embodiment, a parameter can correspond to a phenotype, gender, age, date, pathology, diagnosis, etc. This parameter can be used to select or control the characterization of a set of data from a data corpus.
[0247] For example, a specific medical objective might be to characterize cells from a particular organ in the human body exhibiting atrophy, hypertrophy, hyperplasia, or metaplasia. In this case, a selected corpus of images might be derived, for instance, from samples of a population that underwent a biopsy, which was then analyzed to generate images of a slide from the sample. The diagnosis of these images is likely to correspond to the presence of a specific or non-specific elementary lesion associated with a given pathology.
[0248] In one embodiment, selecting a corpus of images allows for the direct deduction of the lens's properties and thus its characterizing parameters. However, using a parameter characteristic of the lens within the domain of expertise allows for the selection or suggestion of suitable image corpora. Consequently, directly selecting a suitable image corpus does not necessarily require defining a parameter that will not necessarily be used.
[0249] Selected images from the image corpus
[0250] In one embodiment, the selected images from the image corpus are high-content images, that is, images of high resolution and possibly large dimensions. These images are generally dense and have a size of several hundred MB, or even GB.
[0251] According to an example in the field of pathology, input images are virtual slide images originating, for example, from a biopsy slide to a sectioning slide.
[0252] For example, an input image corresponds to an image extracted from a patient's file.
[0253] Taking into account existing training data
[0254] Within the image corpus, an initial filter can be applied based on at least one parameter defining the existing training datasets. This parameter can be generated by a training software component labeled DATA_ENT in Figure 10. This training software component allows for the analysis of a training image corpus used to train a learning function, also known as a machine learning model, and the deduction of:
[0255] - Characterizations of missing data in a training data corpus,
[0256] - characterizations of the representativeness of a training data corpus,
[0257] - characterizations of the homogeneity or heterogeneity of a training dataset,
[0258] - characterizations of a domain within a training dataset that is overrepresented in training, or more generally, of parameters describing the equilibrium of a training dataset,
[0259] - Characterizations of underrepresented categories of training data and parameters aimed at developing a strategy for increasing training data in certain areas,
[0260] - characterizations of proportions of real training data, with respect to synthetic training data, the real data being from real photographic samples and the synthetic data being from an algorithm configured to produce variations of real images.
[0261] The DATA_ENT training software component allows the generation of a data filtering and / or selection strategy within an existing image corpus that can be used to produce new IM1 images for a generated image captcha. These filters can be applied to data describing these images, such as dates, phenotype characterizations, acquisition region characterizations, pathology characterizations, etc. This data can include the name of the image or set of images, the location of their recording, a sender's email address, or any descriptive field of an image or set of images.
[0262] According to another, complementary embodiment, these filters can be applied to the images themselves, for example following image processing.
[0263] Consideration of the Technology Type: According to one embodiment, a software component for technologies, denoted ML_TEC in Figure 10, provides a parameter that characterizes the training technology, such as the nature of the network used and the type of training employed. For example, the parameter can define a learning type, such as machine learning or deep learning. The parameter can also describe the type of training, such as supervised training, semi-supervised training, self-supervised training, or unsupervised training.
[0264] According to various examples, the parameter specifying the supervised machine learning technology can be a linear regression algorithm, a logistic regression algorithm, an algorithm called in the Anglo-Saxon literature "Support Vector Machines" whose acronym is SVM, a K-nearest neighbors algorithm, called in the Anglo-Saxon literature "K-Nearest Neighbors" whose acronym is KNN, a decision tree algorithm, an algorithm called "Random Forests" in the Anglo-Saxon literature, a gradient-boosting algorithm such as XGBoost, LightGBM, CatBoost, a Bayesian network.
[0265] According to various examples, the parameter specifying the unsupervised machine learning technology can be an algorithm called "K-Means Clustering" in the Anglo-Saxon technical literature, or in Anglo-Saxon terminology the following algorithms: DBSCAN designating "Density-Based Spatial Clustering", Hierarchical Clustering, Gaussian Mixture Models called GMM algorithms, Principal Component Analysis called PCA algorithms, t-SNE / UMAP especially for dimensionality reduction techniques, Autoencoders also used in deep learning.
[0266] According to various examples, the parameter specifying the machine learning technology using deep learning can characterize, for example, certain neural networks such as the Multilayer Perceptron (MLP), a Convolutional Neural Network (CNN), a Recurrent Neural Network (RNN), an algorithm known in Anglo-Saxon terminology as "Long Short-Term Memory" (LSTM), a "Gated Recurrent Unit" (GRU), or transformers such as BERT, GPT, ViT, etc. According to various examples, the parameter specifying the reinforcement learning technology can characterize certain algorithms such as a "Q-Learning" algorithm, a Deep Q-Network (DQN), or other algorithms.
[0267] According to various examples, the parameter specifying hybrid or advanced machine learning technology can characterize certain algorithms such as a Bayesian network, a Markov decision process type algorithm, or other algorithms.
[0268] One advantage is to allow the selection of an appropriate number of images from a corpus of images according to the specific case of the technology under consideration, to divide these images into subsets to carry out different training phases such as labeling, testing and validation phases.
[0269] The data characterizing the objective of an area of expertise and the two software components ML_TEC and DATA_ENT allow action on the selection of images from a corpus of images in order to select IM1 imagelets to be produced in an image captcha.
[0270] When a set of images is selected from one or more image corpora, this set defines an input image set.
[0271] Detection of areas of interest in selected input images
[0272] At this stage of the process, it is optional to implement an algorithm for detecting the portions of interest in each input image. This algorithm is applied to extract IM1 images, which will be used to generate an image captcha containing a plurality of IM1 images.
[0273] If this algorithm is not applied, the IM1 images from the input images will be generated either by random segmentation of each input image or by a predefined structured segmentation strategy. In the latter case, such a segmentation strategy might involve segmenting the input image based on a given margin considered from one or more edges of the image, and according to the desired dimensions of the IM1 images to be produced. For example, an overlap ratio between images could be configured to extract the entire input image.
[0274] Alternatively, according to one embodiment, the IM1 images are extracted from certain portions of the input image using a region of interest selection strategy.
[0275] In one example, input images can be annotated with markers corresponding to the positions of regions of interest within the image. These regions of interest can contain patterns, singularities, shapes, or colors that may or may not correspond to entities whose labels are to be determined using the annotation performed when the image is selected in the image captcha. The extracted image IM1 from the segmentation can be defined relative to the region of interest.
[0276] As an example, a region of interest (ROI) can itself be segmented into several image patches or tiles, covering different portions of the ROI. ROIs can be defined according to the objective of the domain of expertise. For example, a ROI can be defined to contain multiple objects in the input image or to contain only one object in the input image.
[0277] As an example, if three areas of interest are identified in an input image, some extracted images may include all or part of the shape defining a pattern of an area of interest.
[0278] Patch generation (thumbnails)
[0279] In one embodiment, a GEN_PATCH extraction software component can be used to generate patches from the input image and markers defining areas of interest. The patches correspond to the tiles or thumbnails that will be produced to be incorporated into an image captcha.
[0280] These IM1 images can have predetermined dimensions, for example, a square format that allows for easy integration into an image captcha. The area of the input image extracted to produce an image can be configured according to the subject matter. In other words, the size of the extracted area determines the scale at which objects are represented in the image. The resolution and dimensions of the IM1 image are configured so that the extracted area is represented at a specific resolution or within a certain acceptable resolution range.
[0281] The IM1 images can, for example, correspond to the entire input image or a portion thereof. A cropping step can be applied to address the issue of the input image edges, which may be altered, have undergone optical edge effects, or define less interesting areas, for example, due to the centering of the sample on a slide, as is the case in the field of expertise related to anatomopathology or cytopathology.
[0282] Different strategies can be used to define the patches. In one embodiment, some patches are generated with the detected pattern so that it is fully represented in the image. Other patches can be generated with portions of patterns so that a pattern is not completely represented in the image. Another strategy consists of detecting similar patterns and generating an image containing two pattern portions.
[0283] These strategies of input image segmentation, detection of areas of interest and extraction of imagettes make it possible to strengthen training with a wide variety of annotated images in order to make a machine learning algorithm more robust.
[0284] According to one embodiment, the dimensions and resolution of the images are parameterized according to an objective of the domain of expertise or a specification from the software function DATA_ENT or the software function ML_TEC.
[0285] In one embodiment, a default resolution parameter for image IM1s can be used. Alternatively, this parameter can be configured according to the domain of expertise objective as mentioned above and / or according to a given learning strategy. Typically, if images have already been used to train a machine learning algorithm with a given resolution, the resolution of the imagelets produced to increase the number of training images can be configured to match a specific resolution corresponding to that of the images already used. In one example embodiment, a resolution error margin is configured to produce imagelets whose IM1 resolution falls within a specified range of tolerance values.
[0286] The resolution constraint can be a dual constraint that must be compatible both with the objective of the area of expertise and with a training-specific requirement that uses data of a certain resolution. In this case, the chosen image corpus takes into account the intersection of the two given resolution ranges.
[0287] In one example, the target image or reference image is selected from among the images produced by the invention's process. It forms a prototype that will serve as the reference for an individual to select the other images included in the image captcha that one wishes to label. In another case, the target image or reference image is selected from an image in a different set. In one embodiment, several target images can be used to generate the captcha.
[0288] The where target images are preferentially selected by a user so that they represent a pattern or singularity specific to the objective of the area of expertise.
[0289] Image processing
[0290] In one embodiment, image processing of the imagelets can be performed optionally. For example, processing could involve modifying the sharpness, contrast, or brightness of the images. In another example, noise is added to some of the imagelets. This also makes the learning process more robust, enabling the detection of elements in the image even when the images to be classified are noisy.
[0291] Separation of the thumbnails
[0292] When a set of images is generated, a separation step can be performed to address different images within different image CAPTCHAs that can be created. This step can involve addressing the images based on a predefined labeling strategy. Each image is referenced in a computer system with an identifier so that it can be associated with and / or aggregated with data that enriches the image to be used.
[0293] Indeed, image CAPTCHAs are created to be displayed on specific user interfaces, to be used by specific machines or applications, at a given frequency and for a predefined duration. This entire setup is called a "campaign." A campaign can be defined and associated with an objective within a specific area of expertise.
[0294] Consequently, considering an example, if 300 images are required to train a machine learning model from which a label is to be obtained using the invention's method, these 300 images could come, for example, from 100 images in the image corpus. In this example, we assume that, on average, 3 portions of interest were used to generate 3 images for each input image. If we consider that in an image captcha containing 12 images, 6 come from the 300 images and 6 are generated randomly or specifically from another corpus, then in this campaign, a display frequency of one image captcha per day for a user's machine interface can be determined. This campaign can then last for as many days as needed to collect the required number of processed image captchas per user.
[0295] Of the six CAPTCHA images selected from the input images, some will be chosen by the user and others will not. The goal is to identify and differentiate between these images in order to enrich our understanding of them and subsequently train a model. Therefore, 50 image CAPTCHAs are needed to process the 300 images for which we want information (a label) generated by a user action that labels or not these images. A 5-day campaign targeting 10 users will then allow us to process all the images.
[0296] Other examples of separation, distribution, or control strategies can be defined according to different time periods, a different number of users, or a different strategy for distributing the images in the image captcha. In one example, the set of images in the image captcha comes from the input images.
[0297] A strategy for separating images can also result from indicators reflecting the quality and number of images that one wishes to annotate to train a model.
[0298] In one example, an image analysis software component can be implemented to automatically analyze the proximity of content in thumbnails and remove thumbnails with content that is too similar. In this case, the separation step includes selecting thumbnails and / or rejecting thumbnails based on a proximity criterion.
[0299] In another example, an image analysis software component can be implemented to automatically analyze the differences or disparities in the content of thumbnails and remove those with excessively disparate content, or conversely, retain those with the most disparate content. In this case, the separation step includes selecting and / or rejecting thumbnails based on a disparity criterion.
[0300] Estimating the proximity or distance between two image sets can be done using various methods. One method might involve analyzing the spectrum and thus the spectral content of the image, for example, using a Fourier transform. Other methods can also be used. One advantage is generating optimized training image sets for a given training strategy, which limits the number of image sets needed to train a machine learning model.
[0301] The separation step can therefore implement a software component defining a selector, a controller and / or a distributor aimed at selecting certain produced imagettes to distribute them within different image captchas from a given training strategy.
[0302] Image captcha generation
[0303] In order to produce an image captcha, the invention includes a GEN1 step which aims to produce an image captcha from different image resources.
[0304] According to a first example, all the thumbnails of an image captcha come from the same and unique input image of the image corpus.
[0305] According to a second example, part of the thumbnails of an image captcha comes from several selected input images in the image corpus.
[0306] According to a third example, part of the thumbnails in an image captcha come from one or more selected input images in the image corpus and other images from other image resources.
[0307] In one example, an image resource is a memory storing synthetic images, that is, images artificially generated from other images. In another example, an image resource is a memory storing a set of ENS13 images whose label is known. The images in this ENS13 set are therefore already labeled and are preferably already processed to correspond in resolution and size substantially to the imagelets processed according to the method of the invention. Preferably, the images in the ENS13 set are imagelets from the same corpus of images as the imagelets that we seek to label using the method of the invention. However, the imagelets in the ENS13 set may have been manually labeled, for example, by an expert user or specialist.
[0308] These pre-labeled thumbnails can be used as a test to validate that the user selects thumbnails with the correct label. By simultaneously annotating some thumbnails from the input images and those with known labels, it's possible to verify that the user is performing a thorough annotation, meaning one with a high level of confidence. Indeed, if the user fails to annotate the thumbnail whose label matches the target image, the entire selection can be invalidated.
[0309] In another example, an image resource is a memory containing images that have already been validated by another user in a different image CAPTCHA. This set of images is labeled ENS12 in Figure 10. An arrow is shown from the output of the process for qualifying an image from a CAPTCHA via the QUAL step and reintroducing this image into another image CAPTCHA by saving it in the ENS12 set. These images are then reintroduced into another CAPTCHA for double validation.
[0310] This ensures that the label assigned to an image has been validated twice by two different users. If an image has not been validated identically by two different users, it is either left unlabeled, sent to a third user via a new image captcha, or forwarded to an administrator or subject matter expert for manual annotation.
[0311] In one embodiment, the use of the same image to obtain two labels from two different users is taken into account to calculate a confidence score associated with a user. To this end, when two labels annotated by two users differ, a comparison and additional labeling identify the user who annotated the image incorrectly. This allows a confidence score to be assigned to that user.
[0312] In another scenario, one or more images with a known class or label are included in an initial set of images, i.e., an image captcha, to verify whether a user or users are correctly annotating the images. When an error is made during annotation, which can be easily verified by the fact that the image's class or label is known, a confidence score is assigned to the user.
[0313] This makes it possible, in particular, to exclude certain user profiles during the execution of the process of the invention.
[0314] In another example, the thumbnails embedded in the image captcha come from a resource storing synthetic images. Synthetic images are produced from real images acquired and processed from a physical object or a living organism. Synthetic images can be produced by applying an image processing algorithm that generates variations of a starting image.
[0315] Figure 10 represents this ENS11 set. A link is shown from the ML TEC and / or DATA_ENT software component to this set because the configuration of a process aimed at producing synthetic images can be directly derived from a machine learning technology and training strategy or from existing data to train the learning function.
[0316] In other words, the method of the invention makes it possible to configure the amount of synthetic data to be produced and to generate data in the domain that meets the need, for example if a subdomain is poorly represented in learning.
[0317] The process of the invention therefore makes it possible to produce images in a subdomain that is poorly represented in images and to validate that users correctly assign the expected label in this subdomain to the images displayed in the image captcha.
Claims
DEMANDS 1. A method for accessing a computing resource in a two-step sequence, said method comprising a first step for qualifying a training dataset for learning a machine learning model and a second step for unlocking access to a computing resource, the first step comprising: ■ generation (GEN1) of a first set (ENS1) of images, called imagelets, by a graphics component, resulting from a segmentation of at least one given input image, ■ display of the first set of images within a graphic element inscribed within a predefined geometric shape; ■ generation (GEN2) of a target image (IMC) representing the target of interest (CIB1) and associated with a label named "target label"; ■ selection (SEL) of at least one first image (iM1) of the first set (ENS1) by means of a selector of a graphical interface accessible via said graphical component; ■ generation (GEN3) of a first image label (LB1) associated with at least one first selected image (IM1); ■ qualification (QUAL) of at least one first selected image (IM1) by assignment of the first label (LB1), so as to obtain a first set of labeled images (ENS1q) intended to train a machine learning model, said selection (SEL) of said at least one first image (IM1) being associated with a first step of a sequence of access to a computer resource, and said graphic component generating an element enabling to engage a second step of said sequence of access to said computer resource, in which an entry of a code and a control of said code is carried out.
2. A method according to the preceding claim, characterized in that the method comprises a preliminary step including: ■ Selection (SEL_CORP) of at least one input image from a given image corpus from an objective of an expertise domain, said objective of an expertise domain including the definition of an image type, called type parameter (PT).
3. A method according to the preceding claim, characterized in that the generation of the first set of images (ENS1) comprises: ■ extraction (EXT_POI, GEN_PATCH) of at least one image corresponding to a portion of this input image within an area of interest (AOI) defined from at least one characteristic data of the area of interest (AOI), said extraction including the determination of at least one dimension and at least one resolution deduced from a dimension parameter (DP) and a resolution parameter (RP).
4. A method according to the preceding claim, characterized in that at least one characteristic data point (PP) of the area of interest (AOI) possibly comprises: ■ a position parameter (PP) and / or; ■ data describing an element of the image so as to produce area of interest markers in the input image from a software component and / or; ■ pixel coordinates in the input image and / or; ■ the definition of a bounding box in the image.
5. A method according to the preceding claim, characterized in that a training software component (DATA_ENT) is configured to query an image database or a training configuration in order to generate at least one first characteristic parameter (PC1) of a set of training images to be produced, said first characteristic parameter (PC1) comprising at least one parameter among which: {resolution parameter, dimension parameter, type parameter, class parameter, parameter characterizing an area of interest or an element present in the image, a number of images}.
6. A method according to the preceding claim, characterized in that a software technology component (ML_TEC) is configured to query a technology configuration of a machine learning function in order to generate at least one second characteristic parameter (PC2) of a learning function, said second characteristic parameter (PC2) comprising at least one parameter among which: {resolution parameter, dimension parameter, type parameter, class parameter, parameter characterizing an area of interest or an element present in the image, a number of images}.
7. A method according to the preceding claim, characterized in that the generation of the first set of images (ENS1), referred to as imagelets, comprises an aggregation of imagelets from at least one of its image sources: ■ Different input images from a corpus of images in a given area of expertise and / or; ■ From a set of synthetic images (ENS11), that is to say artificially generated from other images and / or; ■ From a set of images (ENS13) whose classification label is determined and / or; ■ A set of thumbnails (ENS12) which have already been validated by another user in another first set of images (ENS1') generated, the different thumbnails integrated into the graphic element, called first set of images (ENS1), being positioned randomly, each thumbnail position being associated with its identification so that a selection of said thumbnail allows the exploitation of the data associated with said selected thumbnail.
8. A method according to claim 1 comprising a step of extracting small images of interest from at least one input image from a corpus of images in a domain of expertise comprising: ■ Reception (REC_OBJ) of parameters defining an objective of a domain of expertise of which at least one parameter allows to define a type of images; ■ Selection (SEL_CORP) of at least one input image from a given image corpus from an objective of an expertise domain, said objective of an expertise domain including the definition of an image type, called type parameter (PT); ■ Receiving a resolution parameter (PR) and a dimension parameter (PD) and a characteristic data of at least one area of interest (ZI) of the input image; ■ extraction (EXT_POI, GEN_PATCH) of at least one image corresponding to a portion of the at least one input image within the area of interest (AOI) from at least of the characteristic data, the dimension parameter (PD) and the resolution parameter (PR); ■ generation of a graphic element, called the first set of images (ENS1), from a plurality of imagelets, also called the first set of images (ENS1), extracted and aggregated within the graphic element, said imagelets coming from at least one of its image sources: ■ of different input images from a corpus of images from a given area of expertise and / or; ■ of a set of synthetic images (ENS11), that is to say artificially generated from other images and / or; ■ of a set of images (ENS13) whose classification label is determined and / or; ■ of a set of thumbnails (ENS12) which have already been validated by another user in another first set of images (ENST) generated.
9. A method according to any one of the preceding claims, characterized in that the first set (ENS1) comprises a first panel (PA1) of images including at least one image capable of corresponding to a target of interest and a second panel (PA2) of images not corresponding to the target of interest.
10. A method according to any one of the preceding claims, characterized in that the images of the first set have an identical or substantially identical resolution within 10%.
11. A method according to any one of the preceding claims, characterized in that the segmentation of a given input image by a graphic component is carried out by random fragmentation of the input image into fragments of identical resolution.
12. A method according to any one of the preceding claims, characterized in that each image of the first set (ENS1) is extracted from a given input image having an identifier.
13. A method according to any one of the preceding claims, characterized in that the images of the first set of images (ENS1) represent: ■ fragments of biological organisms and / or; ■ fragments of elements whose image capture is defined at the microscopic scale.
14. A method according to any one of the preceding claims, wherein the second step of the computer resource access sequence is successive to the first step for qualifying a training dataset for the computer resource access sequence.
15. A method according to any one of the preceding claims, comprising sending the first set of labeled images (ENS1 q) to a remote server and a step of verifying said first label (LB1).
16. A method according to any one of the preceding claims, wherein: ■ at least one first selected image (IM1) has been previously associated with a predefined label; ■ A first label verification step (LB1) includes a comparison of the first label (LB1) with the predefined label.
17. A method according to any one of the preceding claims, further comprising: ■ selection of a second image from the first panel (PA1) associated with said target label, said selection allowing verification of the value of the first label (LB1).
18. A method according to claim 15 or 16, characterized in that when the verification is correct, the method comprises: ■ validation of the first step of the two-step sequence and generation of a graphical window to acquire a code to validate a second step and / or; ■ validation of the second step concurrently with the validation of the first step, said validation of the second step corresponding to the expected code.
19. A method according to claim 15 or 16, characterized in that when the verification is incorrect, the method comprises: ■ generation of a new window within a new geometric shape containing a new set of images and / or; ■ validation of the first step of the two-step sequence and generation of a graphical window to acquire a code to validate a second step.
20. A method according to any one of the preceding claims, characterized in that the second step of the sequence involves authentication of a user with a data server or access to the individual computer workstation.
21. Method according to claim 11, wherein the first label verification step (LB1) includes a comparison of the first label (LB1) with a prediction from a second machine learning model trained to generate a prediction of the image label.
22. A method according to any one of the preceding claims, wherein the selection (SEL), generation (GEN2) of said first label (LB1) and qualification (QUAL) steps are implemented by a first user terminal, the method comprising following the qualification (QUAL) step of said first set (ENS1) of images: ■ display of the first set of images (ENS1) within a window that fits within a predefined geometric shape on a display of a second user terminal; ■ generation of a second target image representing the first target of interest and associated with the label named "target label"; ■ selection of at least one image from said first set (ENS1) using a selector from a graphical interface accessible via a graphical component of the second terminal; ■ generation of the first image label (LB1) associated with at least one selected image; ■ qualification of at least one selected image by assigning the first label (LB1) in order to obtain a second set of labeled images, ■ the first label verification step (LB1) including a comparison of the first labeled images with the second labeled images.
23. A method according to any one of the preceding claims, further comprising, after the selection of at least one first image (IM1), a processing step using said graphical interface, so as to obtain at least one first processed image (IM1 trait).
24. Method for generating a training database by supplying it with a population of users labeling images from a plurality of first image sets for training a machine learning model, comprising: ■ implementation (E1) of the method 100 according to any one of claims 1 to 23 by a plurality of remote graphic components, so as to obtain a plurality of first sets of qualified images; ■ generation (E2) of said training database by union of the first qualified image sets of said plurality of first qualified image sets.
25. System comprising at least one server and one terminal communicating via a data network and comprising computers for implementing the method of any one of claims 1 to 23.
Citation Information
Patent Citations
Captcha techniques utilizing traceable images
US20160055329A1
Image based captcha challenges
WO2017096022A1