System and methods for generation of ground truth imagery data of living creatures

EP4659209A1Pending Publication Date: 2025-12-10DIPTERA AI LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
EP2024749830
Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-01-30
Filing Date
2024-01-30
Publication Date
2025-12-10

AI Technical Summary

Technical Problem

Current methods for accurately sorting pre-adult insects and fish are labor-intensive and lack the necessary precision, especially under varying environmental conditions, due to the limited external features available for classification, hindering the efficient use of rearing space and increasing costs.

Method used

A system and method for generating annotated ground truth imagery data by capturing images of pre-adult larvae under controlled conditions, growing them to adulthood for feature identification, and annotating the larval images based on distinguishable features, enabling automated machine learning-based sorting and reducing manual labeling reliance.

Benefits of technology

This approach allows for precise and accurate automated sorting of larvae at the pre-adult stage, reducing operator costs and optimizing rearing space use by minimizing waste, while enabling the use of machine learning models for efficient classification and continuous learning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IL2024050117_08082024_PF_FP
    Figure IL2024050117_08082024_PF_FP
Patent Text Reader

Abstract

A method for generation of ground truth dataset of larval images of living creatures, the method comprising: a. creating a plurality of image batches, wherein each batch corresponds to a single living creature in the larval stage and comprises one or more images; b. growing the single living creature from the larval stage to a stage where visual features of interest are distinguishable; wherein the growing is done while the living creature is placed, separately, in a container; c. after growing the living creature to the stage where visual features of interest are distinguishable, identifying one or more distinguishable features and classifying the living creature according to one or more distinguishable features; and d. annotating the image batches that captured the living creature in the larval stage with the one or more distinguishable features and / or a classification resulting from the one or more distinguishable features.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] SYSTEM AND METHODS FOR GENERATION OF GROUND TRUTH IMAGERY DATA OF LIVING CREATURES FIELD OF INVENTION

[0002] The invention relates to a system and methods for generating large amounts of annotated data regarding pre-adult living creatures.

[0003] BACKGROUND

[0004] Mass rearing of insects and fish is implemented in numerous applications, ranging from food production and waste disposal (e.g., black soldier fly) to pest reduction (e.g., using the sterile insect technique (SIT)). The ability to cost-effectively improve the rearing yield and / or its quality, naturally has great financial implications. One main challenge facing growers is the need to accurately sort the reared population. For example, to ensure that only a specific species is grown or to control the sex of the reared population (e.g., ensuring it is males-only, females-only or at a pre-determined sex ratio). The earlier this sorting is performed, the greater the economic benefit, enabling to maximize use of the rearing space for just the population of interest while minimizing waste (e.g., labor, feed). Therefore, the ability to sort insects and fish at the pre-adult stage is commercially attractive.

[0005] Machine learning (ML)-based sorting holds the promise of automating and speeding-up such processes. In many instances, however, a high level of accuracy and robustness in the sorting is required, especially under varying environmental conditions. This in turn demands the generation of large amounts of data, to train the ML models and evaluate their performance.

[0006] SUMMARY OF THE INVENTION

[0007] The subject matter discloses systems and methods for annotating images of pre-adult insects and fish for ML applications. The subject matter is designed primarily for applications which sort insect or fish larvae according to their sex and / or species and / or genus. This is especially relevant in the context of SIT and for optimization of mass-rearing conditions. The invention is particularly useful as it provides a means to perform automated, ML-based sorting already at the larval stage, which currently is either impossible or extremely labor-intensive. This in turn, dramatically reduces the operator’s costs. Namely, by enabling optimal use of the insect / fish rearing space while minimizing waste by growing to adulthood only the population of interest (e.g., in the case of sex-sorting for SIT, only males are required).

[0008] The subject matter comprises a process of gathering a plurality of image batches of larvae under controlled illumination conditions and background. Each batch corresponds to a single larva and contains one or more images, possibly taken from different directions. The next steps include growing the photographed larva to adulthood or to a stage where the one or more features of interest are easily distinguishable in a separate and supervised manner. This manner is defined by the ability to identify or access a single living creature at a certain time. The separate and supervised manner may be implemented by growing each larva in a labeled container, wherein each label is uniquely associated with the image batch of the living creature grown inside, the batch captured in the pre-adult stage of the living creature; analyzing the living creature in an adult stage or at a stage where the one or more features of interest are easily distinguishable separately and classifying them according to one or more distinguishing features; and annotating the image batches captured at the larval stage with the classification(s), said classifications are provided in the adult stage or at a stage where the one or more features of interest are easily distinguishable. The entire process may be completely automated without any user intervention, partially automated or completely manual.

[0009] The subject matter takes advantage of the fact that for many insects and fish, a reliable and quick classification could be performed at the adult stage or at a stage where the one or more features of interest are easily distinguishable to generate precise and accurate larvae ground truth data, without requiring the usual stage of painstaking and less reliable manual labeling of raw data. The term “ground truth data” is defined as a sufficiently accurate labeled dataset used to train a software-based model such as machine learning models or deep learning models.

[0010] Thus, according to one broad aspect of the invention, there is provided a method and / or system for automatic generation of a ground truth dataset of larval images of insects or fish, the method comprising: creating a plurality of image batches, wherein each batch corresponds to a single living creature and comprises one or more images captured in the pre-adult, larval stage; growing each living creature, separately in a container, to adulthood or a stage where one or more features of interest of the living creature can be identified or distinguished. The one or more features of interest may include organs, shapes, sizes, voids, structures, color, and other recognizable features that represent conditions or identities of interest, for example, the living creature’s sex, species, health, age, and the like. Each container may have a label attached thereto or another unique identifier enabling to uniquely associate a container with the image batch of the living creature grown inside; analyzing the living creatures after they grow to the desired development stage that enables identifying the visual features of interest; classifying the living creatures according to one or more distinguishing features of interest and annotating the image batches with the corresponding classification(s).

[0011] According to another broad aspect of the invention, there is provided a method and / or system for automatic ground truth generation of larval images of insects or fish, the method comprising: creating a plurality of image batches of larvae, wherein each batch comprises one or more images and corresponds to a single larva known to have one or more sorting-linked phenotypes, said phenotypes are identifiable using a specific imaging technique. The larvae may be fluorescently labeled and / or fluorescently overexpressing and / or from a genetic strain. Then, the method comprises classifying the images of the larvae based on the one or more sorting-linked phenotypes; and annotating the image batches of each imaged larva with the sorting-linked classification.

[0012] In some embodiments, said annotated larval images are used for training and / or validating a machine learning (ML) algorithm.

[0013] In some embodiments, said larval images are used for continual learning of an ML algorithm.

[0014] In some embodiments, said ML algorithm is a classifier.

[0015] In some embodiments, said classification is done using one or more images where no sorting-linked phenotype is present or visible.

[0016] In some embodiments, said classification is according to sex.

[0017] In some embodiments, said classification is according to species or genus.

[0018] In some embodiments, said classification is of mosquitoes.

[0019] In some embodiments, said classification is based on imaging.

[0020] In some embodiments, said larval images are taken from different directions.

[0021] In some embodiments, said larval images are taken as the larvae are flowing.

[0022] In some embodiments, said larval images are taken under controlled background and illumination conditions. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to better understand the subject matter that is disclosed herein and to exemplify how it may be carried out in practice, embodiments will now be described, by way of non-limiting examples only, with reference to the accompanying drawings, in which:

[0024] Fig. 1 illustrates an exemplary flowchart summarizing an exemplary embodiment of the image annotation process used for training ML applications according to the present invention;

[0025] Fig. 2 illustrates another exemplary flowchart summarizing an exemplary embodiment of the image annotation process used for training ML applications according to the present invention; and

[0026] Fig. 3 illustrates a schematic block diagram summarizing operation of a classification system according to one embodiment of the present invention.

[0027] DETAILED DESCRIPTION OF SPECIFIC EMBODIMENTS

[0028] In the present invention, a rapid, accurate and easy to use technique of annotating images of pre-adult insects and fish for ML applications, is described.

[0029] The technical challenge addressed in the subject matter is classifying images of larvae as the larvae present only a small number of external features. This is especially true for determining a larva’s sex since the reproductive organs are internal and not fully developed. Thus, annotation of larvae images for generation of ground truth data is currently handicapped. This is highly problematic in the case of supervised machine learning and / or deep learning where large amounts of annotated data are required to train the model. In contrast to larvae, adult insects and fish are typically rich with differentiating external features (e.g., size, structure, color, shape of different organs) that can be identified from imagery data taken in the adult stage. This makes annotation for purposes of classification a much more straightforward, reliable, and non-invasive task. Thus, for example, image recognition algorithms have been able to sex-sort adult Aedes mosquitoes with an accuracy of over 99.9996%.

[0030] Some prior art solutions propose a deep-learning framework to classify mosquito larvae from larvae of other insects using smartphone images. The process of image labeling in these solutions is done manually and without any information regarding the adult classifications.

[0031] The techniques described herein are described with reference to insect and fish larvae but are applicable to other living creatures and other pre-adult stages, such as pupal stages of insects.

[0032] Reference is made to Fig. 1 illustrating, by way of a flowchart 100, a non-limiting example of the image annotation process used for training ML applications. In step 102, one or more images of a living creature in a larval stage are acquired. In some cases, the plurality of image batches is taken while the larvae are flowing. The images may be captured using a camera, either a still camera, a video camera, a camera in the visible wavelength, an Infrared camera, a radar sensor or any other sensor that can be used to produce a visual representation of the living creature. The plurality of image batches may be taken under a controlled background and illumination condition, for example, illumination configured to assist in identifying the features of interest. The images may be taken from one or more directions. This can facilitate the detection of structures and / or organs of interest, which may be hidden or obscured when viewed from a single direction. After capturing the image of the living creature, the living creature that appeared in the captured images is then stored in a separate, container in which the living creature grows to a stage where visual features of interest of the living creature are distinguishable 104. The container may be defined as a volume big enough to grow the living creature, having a lid, a door or an opening configured to enable insertion and removal of the living creature into and from the container. The container may contain water, feed and any other condition necessary for the living creature to grow to a stage where visual features of interest of the living creature are distinguishable. After growing the living creature, the method comprises a step of identifying one or more distinguishable features. Then, the acquired images are associated with the container (e.g., using a barcode or another unique identifier) in a one-to-one relationship 106. In certain embodiments, the pre-adult imaging system is automated and designed to pass one living creature in the larval stage at a time, capture one or more images of the living creature and transfer the living creature to an independent container. Once the unique identifier is assigned to the container of the specific living creature, a software program associates the images of the larva inside with the label’s unique identifier and the system moves on to the next larva.

[0033] After reaching adulthood or a stage where the one or more features of interest are easily distinguishable, the living creature in the container is analyzed 108 (e.g., by imaging). Based on this analysis, in step 110, a high-confidence classification for each living creature is provided (e.g., species, sex). The classification may be provided by a computer algorithm or a human operator. The classification may be stored in a computerized memory, associated with the images of the same living creature that were taken in the larval stage. The classification may then be associated with the appropriate container in which the living creature grew. Thus, by assigning a unique identifier or a label to the relevant container, the larval images of the living creature that grew inside the container may be tagged with the classification(s) provided in the later developmental stage of the living creature 112. Finally, an ML prediction model may be trained to classify larvae in real time using these labeled images 114. In certain embodiments, the living creature may be mosquitoes, and sex classification will be performed based on visual differences such as the shape of the adult head and body size.

[0034] In some exemplary embodiments, the ML prediction model may be trained using a training dataset that includes classifiers and a set of images that match to the classifier. The prediction model may include, directly or indirectly, a set of rules over specific classifiers to distinguish each classifier in the larval stage using the training dataset. Additionally, or alternatively, the prediction model may be generated using federated learning performed on a server. Each edge device may be configured to provide a model update to the predictive model based on pairs of fingerprints and labels available to the edge device, without exposing training data generated by the respective edge device. One technical effect of the disclosed subject may be reducing the time required to identify a classifier of a living creature in a larval stage, or otherwise in a stage prior to the development of visual features that enable to identify the classifier. The classifier may be the creature’s sex, species or genus or another classifier that can be used to sort living creatures.

[0035] In some embodiments, the ML prediction model is continuously trained with new data. Such continual learning reduces the risks of concept drift and less accurate predictions over time, by adapting to possible changes in the system (e.g., changes in the population due to seasonality).

[0036] Fig. 2 illustrates, by way of a flowchart 200, another non-limiting example of the image annotation process used for training ML applications. The flowchart 200 describes a process of imaging a living creature in the larval stage. A plurality of image batches may be taken under a controlled background and illumination condition, for example, illumination configured to assist in identifying the features of interest. The living creature is fluorescently labeled and / or fluorescently overexpressing and / or from a genetic strain. The labeling, overexpression, or genetic strain makes it easy to classify (e.g., species, sex) the living creature in the larval stage due to one or more sorting-linked phenotypes. This way, the images of the living creature in the larval stage may be used to train an ML prediction model classifier for wild-type larvae, which do not present this sorting-linked phenotype. The process of Fig. 2 may be extremely beneficial, for example, in use cases where larvae need to be sorted, but cannot be fluorescently labeled and / or fluorescently overexpressing and / or from a genetic strain (e.g., in the case of sex-sorting for SIT where regulations limit releases to wild-type insects).

[0037] In step 202, one or more images are acquired, from one or more regions of a living creature in a larval stage. In some cases, the plurality of image batches is taken while the larvae are flowing. The living creature in the larval stage is fluorescently labeled and / or fluorescently overexpressing and / or from a genetic strain, such that it has one or more sorting-linked phenotypes. The images may be captured using a camera, either a still camera, a video camera, a camera in the visible wavelength, an Infrared camera, a radar sensor or any other sensor that can be used to produce a visual representation of the living creature. The images may be taken from one or more directions. This can facilitate the detection of structures and / or organs of interest, which may be hidden or obscured when viewed from a single direction. Based on the sorting-linked phenotype(s), a computer algorithm, or a human operator, provides a high-confidence classification for the living creature in the larval stage 204 (e.g., species, sex). This classification is then associated with the images of the living creature in the larva stage 206, including images that do not present the sorting-linked phenotype. Then, an ML prediction model may be trained to classify living creatures in the larval stage in real time using the labeled images and where the sorting-linked phenotype is not present or not visible or has been removed from the image 208. In certain embodiments, the images of the living creature in the larval stage are acquired by two distinct imaging systems. One imaging system is designed to detect the sorting-linked phenotype(s) and the other imaging system is designed such that the sorting-linked phenotype(s) is not present or not visible or may be removed from the image. In such embodiments, the images acquired by the imaging system designed to detect the sorting-linked phenotype(s) are used to classify the images taken by the other imaging system, which will serve to generate a ground truth dataset. Thus, images taken of larvae which are fluorescently labeled and / or fluorescently overexpressing and / or from a genetic strain may be used to train an ML prediction model classifier for wild-type larvae, which do not present such sorting-linked phenotypes. Additionally, in certain embodiments, the insects may be mosquitoes, and sex classification will be performed based on eye color.

[0038] FIG. 3 is a schematic block diagram illustrating operation of a classification system 300, according to one embodiment. In one embodiment, a network or ML algorithm 302, may be trained and used for identifying and classifying or detecting larvae in an image. The network or ML algorithm 302 may include a neural network, such as a deep convolution neural network, or other ML model or algorithm for classifying or identifying larva sex and / or species and / or genus.

[0039] In one embodiment, the network or ML algorithm 302 is trained using a training algorithm 304 based on training data 306. The training data 306 may include images of larvae and their designated classifications. For example, the training data may include images classified as including larvae of a first type and images classified as including larvae of a second type. The types of the larvae may vary significantly based on the type of examination or report that is needed (classification by sex, species and / or genus). The training data is composed of image batches, wherein each batch corresponds to a single larva and comprises one or more images. Furthermore, each batch is labeled according to classification(s) performed on the respective insect or fish after it reached the stage where visual features of interest are distinguishable 308. Using the training data, the training algorithm 304 may train the ML algorithm 302. For example, the training algorithm 304 may use any type or combination of supervised or unsupervised ML algorithms.

[0040] In one embodiment, the training data 306 may include a plurality of images for the same larva. The plurality of images of the same specimen may be from different directions, yet all corresponding to the same identical label (e.g., ground truth classification of the specimen).

[0041] Once the network or ML algorithm 302 is trained, the network or ML algorithm 302 may be used to identify or predict the type of larva within an image. For example, an unclassified image 310 (or previously classified image with the classification information removed) is provided to the network or ML algorithm 302 and the network or ML algorithm 302 outputs a classification 312. The classification 312 may indicate a yes or no for the presence of a specific type of larva. For example, the network or ML algorithm 302 may be targeted to detecting whether a specific type of larva is present in the unclassified image 310. Alternatively, the classification 312 may indicate one of many types that may be detected by the network or ML algorithm 302. For example, the network or ML algorithm 302 may provide a classification that indicates which type of larva is present in the unclassified image 310. During training, the classification 312 may be compared to a human classification or an out-of-channel classification to determine how accurate the network or ML algorithm 302 is. If the classification 312 is incorrect, the unclassified image 310 may be assigned a classification from a human and used as training data 306 to further improve the network or ML algorithm 302.

[0042] In one embodiment, both offline and online training of the network or ML algorithm 302 may be performed. For example, after an initial number of rounds of training, an initial accuracy level may be achieved. The network or ML algorithm 302 may then be used to assist in classification with close review by human workers. As additional data comes in the data may be classified by the network or ML algorithm 302, reviewed by a human, and then added to a body of training data for use in further refining training of the network or ML algorithm 302. Thus, the more the network or ML algorithm 302 is used, the better accuracy it may achieve. As the accuracy is improved, less and less oversight of human workers may be needed.

[0043] The processes disclosed above are performed by a system having a processor and a memory. In some exemplary embodiments, the system may comprise one or more processors. The processor(s) may be a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), a microprocessor, an electronic circuit, an Integrated Circuit (IC) or the like. The processor may be utilized to perform computations required by the system or any of its subcomponents. It is noted that the processor may be a traditional processor, and not necessarily a quantum processor.

[0044] In some exemplary embodiments, the memory may be a hard disk drive, a Flash disk, a Random Access Memory (RAM), a memory chip, or the like. In some exemplary embodiments, the memory may retain program code operative to cause a processor to perform acts associated with any of the subcomponents of the system. The memory may comprise one or more components as detailed below, implemented as executables, libraries, static libraries, functions, or any other executable components.

[0045] The present disclosed subject matter may be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium (or media) having computer-readable program instructions thereon for causing a processor to carry out aspects of the present disclosed subject matter. The computer-readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer-readable storage medium may be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of the computer-readable storage medium includes the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD- ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or raised structures in a groove having instructions recorded thereon, and any suitable combination of the foregoing. A computer readable storage medium, as used herein, is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission media (e.g., light pulses passing through a fiber-optic cable), electrical signals transmitted through a wire, Quantum Random Access Memory (QRAM), photons, trapped ions, lasers, cold atoms, or the like.

[0046] It is appreciated that certain features of the invention, which are, for clarity, described in the context of separate embodiments, may also be provided in combination in a single embodiment. Conversely, various features of the invention, which are, for brevity, described in the context of a single embodiment, may also be provided separately or in any suitable subcombination or as suitable in any other described embodiment of the invention. Certain features described in the context of various embodiments are not to be considered essential features of those embodiments, unless the embodiment is inoperative without those elements.

[0047] Although the invention has been described in conjunction with specific embodiments thereof, it is evident that many alternatives, modifications and variations will be apparent to those skilled in the art. Accordingly, it is intended to embrace all such alternatives, modifications and variations that fall within the spirit and broad scope of the appended claims.

Claims

CLAIMS1. A method for generation of ground truth dataset of larval images of living creatures, the method comprising: a. creating a plurality of image batches, wherein each batch corresponds to a single living creature in the larval stage and comprises one or more images; b. growing the single living creature from the larval stage to a stage where visual features of interest are distinguishable; wherein the growing is done while the living creature is placed in a separate container; c. after growing the living creature to the stage where visual features of interest are distinguishable, identifying one or more distinguishable features and classifying the living creature according to one or more distinguishable features; and d. annotating the image batches that captured the living creature in the larval stage with the one or more distinguishable features and / or a classification resulting from the one or more distinguishable features.

2. The method according to claim 1, where the annotated larval images are used for training and / or validating a machine learning (ML) algorithm.

3. The method according to claim 2, where the larval images are used for continual learning of an ML algorithm.

4. The method according to claims 2 or 3, wherein the ML algorithm is a classifier.

5. The method according to claim 4, wherein the classification is done according to sex.

6. The method according to claim 4, wherein the classification is done according to species or genus.

7. The method according to claim 1, further comprises assigning a first unique identifier to the labeled container and a second unique identifier to the image batches; wherein the first unique identifier is associated with the second unique identifier.

8. The method according to claim 1, further comprises identifying additional distinguishable features from the living creature at the stage where visual features of interest are distinguishable, wherein the additional distinguishable features are not related to the classification and annotating the image batches to include the additional distinguishable features.

9. The method according to claim 1, wherein at least one batch of the plurality of image batches comprises images taken from different directions.

10. The method according to claim 1, wherein the plurality of image batches are taken under a controlled background and illumination conditions.

11. The method according to claim 1, wherein the plurality of image batches are taken while the larvae are flowing.10SUBSTITUTE SHEET (RULE 26)12. A system for automatic generation of ground truth datasetof larval images of insects and fish, the system comprising: one or more cameras configured to create a plurality of image batches, wherein each batch corresponds to a single living creature and comprises one or more images from the larval stage; a container configured to grow each living creature separately, to a stage where one or more features of interest are distinguishable; a processor configured to receive a classification of the living creatures according to one or more distinguishing features and to annotate the image batches with the classification of the living creature.

13. The system according to claim 12, wherein the annotated image batches are used for training and / or validating an ML algorithm.

14. The system according to claim 13, wherein the image batches are used for continual learning of an ML algorithm.

15. The system according to claims 14 or 13, wherein the ML algorithm is a classifier.

16. The system according to claim 14, wherein the classification is done according to sex.

17. The system according to claim 14, wherein the classification is done according to species or genus.

18. The system according to claim 11, wherein the image batches are taken from different directions.

19. The system according to claim 11, wherein the larval images are taken under a controlled background and illumination conditions.

20. A method for automatic generation of ground truth dataset of images of living creatures, the method comprising: creating a plurality of image batches, wherein each batch comprises one or more images and corresponds to a single living creature in a larval stage, said living creature is fluorescently labeled and / or fluorescently overexpressing and / or from a genetic strain, such that said living creature has one or more sorting-linked phenotypes; classifying each living creature based said on one or more sorting-linked phenotypes; and annotating the image batches that captured the living creature with the sorting-linked classification.

21. The method according to claim 20, wherein the annotated larval images are used for training and / or validating a machine learning (ML) algorithm.

22. The method according to claim 21, wherein the larval images are used for continual learning of an ML algorithm.

23. The method according to claims 21 or 22, wherein the ML algorithm is a classifier.

24. The method according to claim 23, wherein the classification is done according to sex.

25. The method according to claim 24, wherein the classification is done according to species or genus.11SUBSTITUTE SHEET (RULE 26)26. The method according to claim 24, wherein the classification is done using one or more images where no sorting-linked phenotype is present or visible.

27. The method according to claim 20, wherein the larval images are taken from different directions.

28. The method according to claim 20, wherein the larval images are taken under a controlled background and illumination conditions.

29. The method according to claim 22, wherein acquiring the plurality of image batches using two distinct imaging systems; wherein a first imaging system is designed to detect the sorting-linked phenotype(s) and a second imaging system is designed such that the sorting-linked phenotype(s) is not present or not visible or may be removed from the image; wherein the images acquired by the first imaging system are used to classify the images taken by the second imaging system.

30. A system for automatic generation of ground truth dataset of images of living creatures in the larval stage, the system comprising: a camera configured to create a plurality of image batches, wherein each batch comprises one or more images and corresponds to a single living creature in the larval stage, said living creature is fluorescently labeled and / or fluorescently overexpressing and / or from a genetic strain, such that said living creature has one or more sorting-linked phenotypes; a processor configured to classify each living creature based on one or more sorting-linked phenotypes and annotate the image batches of each living creature with the sorting-linked classification.

31. The system according to claim 30, wherein the annotated larval images are used for training and / or validating a ML algorithm.

32. The system according to claim 31, wherein the larval images are used for continual learning of an ML algorithm.

33. The system according to claims 31 or 32, wherein the ML algorithm is a classifier.

34. The system according to claim 33, wherein the classification is done according to sex.

35. The system according to claim 33, wherein the classification is done according to species or genus.

36. The system according to claim 33, wherein the classification is done using one or more images where no sorting-linked phenotype is present or visible.

37. The system according to claim 30, wherein the images are taken from different directions.

38. The system according to claim 30, wherein the images are taken under a controlled background and illumination conditions.

39. The system according to claim 30, further comprises two distinct imaging systems; wherein a first imaging system is designed to detect the sorting-linked phenotype(s) and a second imaging system is designed such that the sorting-linked phenotype(s) is not present or not visible or may be removed from12SUBSTITUTE SHEET (RULE 26)the image; wherein the images acquired by the first imaging system are used to classify the images taken by the second imaging system.13SUBSTITUTE SHEET (RULE 26)