Provisioning of annotated training images for use in training of trainable image content recognition algorithms for machine vision systems
By pre-acquiring annotation information in the machine vision system and automatically capturing and annotating images, the labor-intensive problem of training image selection and annotation processes is solved, achieving efficient and flexible training image supply and automated training, thus improving training efficiency and flexibility.
Patent Information
- Application Number
- CN202510645620.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-05-23
- Filing Date
- 2025-05-20
- Publication Date
- 2025-11-25
AI Technical Summary
In existing technologies, the process of selecting and annotating training images in machine vision systems is labor-intensive and difficult to automate, resulting in low training efficiency and poor flexibility. In particular, the delay is significant when retraining is required due to changes in applications and image recognition tasks.
By pre-acquiring annotation information in the machine vision system, images are automatically captured and annotated to generate annotated training images that can be directly used for algorithm training. Useful images are determined by checking the recognition content, reducing unnecessary image storage and repeated training.
It enables efficient and automated provisioning of training images, improves training flexibility and efficiency, reduces storage requirements, supports real-time feedback and user-friendly operation, and simplifies the retraining process.
Smart Images

Figure CN121010975A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments herein relate to an arrangement and a method for providing annotated training images for use in training of trainable image content recognition algorithms of a machine vision system. BACKGROUND
[0002] Machine vision systems, sometimes also referred to as computer vision systems, are for example used in industry, then often referred to as industrial vision systems. Such systems comprise one or more cameras. Sometimes the system is a camera or corresponds to a camera, i.e. the system is mainly or only a camera or is implemented in or in the same unit as the camera.
[0003] Furthermore, machine vision systems, or generally image content recognition solutions, are increasingly based on machine learning, more specifically on trainable image content recognition algorithms. There are many such algorithms, most commonly in the form of neural networks like ResNet or MobileNet, which can be trained, or in other words optimized, using numerical methods such as backpropagation.
[0004] Training of an algorithm like this for a machine vision system typically requires many images to train on, to learn the capability of image content recognition and / or to improve it to a sufficient degree that the algorithm after training provides a desired response regarding the images to be used. In other words, the training is performed to enable the algorithm after training to perform its task, for example when used in a machine vision system which can be part of a production line. What the task exactly is, what content to recognize, what the desired response is to provide and what is considered to be sufficient varies significantly between different applications. The same machine vision system and algorithm but trained differently can be used for different applications. An algorithm and machine vision system being used for a first application can be (re)configured and / or (re)trained for use in another, second, application. This can happen when the kind of objects to be recognized changes and / or if the image recognition task(s) to be performed changes. A retraining can also be performed if the result is that the algorithm that has been trained does not perform as it should.
[0005] Of course, what images are suitable for training depends on many different factors in practice, such as the intended application of the trained algorithm, the algorithm, the task of the algorithm, what content to recognize and the desired response to provide. The images for use in training are often and herein referred to as training images. In practice, the training images are often based on what the algorithm is known or has been identified to encounter and should be able to handle after training. The training images can include clear, simple and problem-free images, as well as more difficult to handle and perform the image recognition task on such images.
[0006] During training, the algorithm needs to be told what is in the images being used, i.e. in the training images, e.g. what is shown or not shown in the training images, or more generally some properties of what is in the images and relevant and useful for the training. For example, if the algorithm should be able to classify fruit images as apples or pears, it is recognized that it is beneficial to train on images where it is known that all shown fruit is either an apple or a pear, because the algorithm can then learn what the properties of apples and pears are, respectively.
[0007] The annotation information is information associated with the training image and the information indicates or identifies properties of the content in the image for use in the training using the training image. A training image with associated annotation information can be referred to as an annotated training image. Thus, the annotation information about the training image on which the content recognition algorithm is to be trained can be said to correspond to information about properties of some content in the training image that are useful for the algorithm when training on the training image. Another way of saying it is that the annotation information corresponds to information that a content recognition algorithm can exploit to improve the ability of the algorithm to perform its image-based task(s) by training.
[0008] Another example: if the algorithm should be able to detect damaged objects, it is relevant to train on images of differently damaged objects and information identifying these images as images of damaged objects, but it is also of interest to train on images of undamaged objects and information identifying these images as images of undamaged objects.
[0009] As is recognized thereby, the image content or its properties to which the annotation information is directly related need not be the same image content that the algorithm is to recognize according to its task after training.
[0010] Annotating a training image, i.e. performing the annotation, is about determining the annotation information and associating it with the training image.
[0011] Generating and / or collecting images for training and annotating them is often a labor-intensive task. Fig. 1A illustrates a basic prior art method with actions. The user first acquires images, e.g. generates them and / or obtains already generated images from elsewhere. Then, the user selects which images to use as training images, or can simply use all or as many of the acquired images as possible. Of course, these acquisition and selection actions can be combined. After that, the user annotates the images and the resulting annotated training images are used to train the algorithm.
[0012] It is to be realized that selecting training images requires effort from the user, takes time, and that the effort and time increases with the number of images, of course also depending on how the selection and annotation is made.
[0013] Fig. 1 B is a flow chart schematically illustrating another prior art approach, where the selection is automated. That is, there is an automatic selection of suitable training images for training an image content recognition algorithm for use by a machine vision system.
[0014] One can make use of so-called active learning, where the system in some way determines which of the annotated data points, such as images, will be most useful to annotate, and only presents those examples to a human user for annotation. An example of such automatic selection as illustrated in Fig. 1 B can be found in patent application US20220189185 A1, which describes an active learning system for a machine vision system.
[0015] There are also concepts such as semi-supervised learning or self-training. In these cases, the learning system does not present images for annotation, but rather determines images that the algorithm is confident about, then annotates these images autonomously and trains on them, hoping that it will become more confident about other images as well. It is to be realized that the result can not always be like this.
[0016] Generally, it is difficult to fully automate the selection and annotation with good results without some human involvement, which can even need to be substantial.
[0017] Fig. 1 C is a flow chart related to Figs. 1 A-1 B, schematically illustrating that the operation of the machine vision system using the trained algorithm is separate from such training. The images generated by the machine vision system during normal operation, i.e. post-training, and on which the trained algorithm is used, can be referred to as production images.
[0018] The training is typically spatially and temporally separate from the place and time where the algorithm is eventually used post-training. For example, training of an algorithm for a machine vision system is typically performed remotely with respect to the machine vision system, and can even not involve any physical machine vision system at all.
[0019] A typical procedure is to deliver and / or set up the machine vision system with a trained algorithm and / or a (re)trained algorithm that can be received, e.g. from a provider of the machine vision system, and install it on the machine vision system, and then operate the machine vision system using the installed (re)trained algorithm. SUMMARY
[0020] In view of the above, an object is to provide one or more improvements or alternatives to the prior art, such as to provide improvements regarding the provision of annotated training images for use in the training of trainable image content recognition algorithms of a machine vision system.
[0021] According to a first aspect of embodiments herein, the object is achieved by a method for providing one or more annotated training images for use in the training of trainable image content recognition algorithms of a machine vision system. The machine vision system is operable to recognize content in images captured by the machine vision system by means of the trainable image content recognition algorithms. Annotation information is acquired for one or more intermediate images to be captured by the machine vision system. The annotation information is indicative of an attribute of content in the intermediate images and is content for recognition by the trainable image content recognition algorithms. The attribute is such that, if the algorithms were to be trained on the intermediate images, the training would benefit from the knowledge that the intermediate images contain content having the attribute. The machine vision system is operated so that the one or more intermediate images are captured by the machine vision system in accordance with the acquired annotation information. One or more annotated intermediate images are provided corresponding to the one or more intermediate images captured by the machine vision system in accordance with the acquired annotation information and the annotation. The one or more annotated training images are provided based on at least one of the annotated intermediate images.
[0022] According to a second aspect of embodiments herein, the object is achieved by one or more devices for providing one or more annotated training images for use in the training of trainable image content recognition algorithms of a machine vision system. The machine vision system is operable to recognize content in images captured by the machine vision system by means of the trainable image content recognition algorithms. The one or more devices are configured to acquire annotation information for one or more intermediate images to be captured by the machine vision system. The annotation information is indicative of an attribute of content in the intermediate images and is content for recognition by the trainable image content recognition algorithms. The attribute is such that, if the algorithms were to be trained on the intermediate images, the training would benefit from the knowledge that the intermediate images contain content having the attribute. The one or more devices are further configured to operate the machine vision system so that the one or more intermediate images are captured by the machine vision system in accordance with the acquired annotation information. Furthermore, the one or more devices are configured to provide one or more annotated intermediate images corresponding to the one or more intermediate images captured by the machine vision system in accordance with the acquired annotation information and the annotation. Moreover, the one or more devices are configured to provide the one or more annotated training images based on at least one of the annotated intermediate images.
[0023] According to a third aspect of embodiments herein, the object is achieved by a computer program comprising instructions, which when executed by one or more processors, cause one or more devices to perform the method according to the second aspect.
[0024] According to a fourth aspect of embodiments herein, the object is achieved by a carrier comprising the computer program according to the third aspect.
[0025] Thanks to embodiments herein, which can be described as being of a "pre- annotation" type, annotated training images for use in the training of trainable image content recognition algorithms of a machine vision system can be provided without any need for inspection of images for annotation. Instead, the annotation of captured images can be performed automatically by the system, for example. The combination of "pre-annotation" and the use of a machine vision system like this for ingesting images has several further advantages. More efficient and useful annotated training images can be provided, and used directly on the system for training immediately after or right after they have been provided. Furthermore, a reduced need for storage and use of large amounts of annotated training images is achieved. Moreover, the training can be made more efficient, more flexible and / or of a higher degree of automation compared to what is normally the case.
[0026] The annotation, type and / or format of the annotated images like this can be as in the prior art. The difference compared to the prior art is that the annotation information is acquired in advance of the image to which the annotation information relates being captured by the machine vision system. Thus, the annotation information is already available when the image is captured, and can therefore be annotated directly and / or automatically by the system without any need for inspection of images like this. As identified from the discussion in the background art, this is different from the conventional approach to image annotation for training of trainable image content recognition algorithms.
[0027] Furthermore, embodiments herein enable the system to apply the algorithm to the respective training image being captured, preferably automatically, and then perform a check to see if the algorithm is able to recognize the content from the annotation information. In this way, the system can know which images will be most useful to train on, i.e. images for which the algorithm fails to recognize the content from the annotation information. Other images can be discarded, enabling fewer but more useful annotated training images to be available for training later.
[0028] The embodiments herein also support and / or enable user-friendly operations to efficiently complete informative and useful training images. For example, the embodiments herein in principle make it possible to give feedback in real-time to a control system or a user (e.g. an operator of a machine vision system) that controls what is being imaged (such as an object). This can more or less directly know when the imaged thing (e.g. a certain pose of an object) cannot be identified by the algorithm from the annotation information. The embodiments herein thus make information directly available after the system has ingested an image, which information indicates whether there is an interest in further training on that image. The control system or the user can learn from this feedback how to produce and be encouraged to next try to produce another image on which the algorithm cannot identify from the annotation information, i.e. another training image on which to train would be useful. BRIEF DESCRIPTION OF DRAWINGS
[0029] Examples of embodiments herein will be described in more detail with reference to the attached schematic drawings, in which:
[0030] Figs. 1A-1C are schematic block diagrams for illustrating prior art methods and actions.
[0031] Figure 2 schematically illustrates a simplified example of a machine vision system that can be used and / or configured to perform embodiments herein.
[0032] Figure 3A is a flow chart for schematically illustrating a method according to embodiments herein.
[0033] Figure 3B is a block diagram schematically illustrating and exemplifying some of the actions as part of the method in Figure 3A
[0034] Figure 4 is a block diagram schematically illustrating and exemplifying some embodiments of the method in Figure 3A
[0035] Figure 5 is a schematic block diagram for illustrating embodiments of one or more devices configured to perform the method in Figure 3A
[0036] Figure 6 is a schematic block diagram illustrating some embodiments related to computer program(s) and carriers thereof. DETAILED DESCRIPTION
[0037] The embodiments herein are exemplary embodiments. It should be noted that the embodiments are not necessarily mutually exclusive. Components from one embodiment can be assumed to be present in another embodiment, and it will be apparent to those skilled in the art how those components can be used in other exemplary embodiments. In order to enable a better understanding, prior art situations and problems indicated in the above background will be further elaborated before describing the embodiments herein.
[0038] In order to train an image content recognition algorithm as discussed in the background, a very large number of training images are typically used, although only some of them can actually add something informative and useful to the training. That is, there is something new that is useful in the sense that there is some improvement in the ability to "learn" from, and bring about the ability to recognize image content and perform the algorithm's task. However, it is typically not known in advance which training images are optimal for use. Instead, the very large number is used to make it possible that a sufficient number of useful images will be included. Thus, a large number of images are used, and thus also a large number of images must be annotated, although they can not actually add something useful.
[0039] As already mentioned in the background, the training is typically performed as a separate action, which can be spatially and temporally separate from the place and time at which the machine vision system will use the trained algorithm.
[0040] When the training is performed and separate from the use of the trained algorithm in the machine vision system, it is not a big problem to use a large number of annotated training images and train on them like this. There is then typically enough time, processing power and available memory capacity.
[0041] However, if the machine vision system in operation after the training turns out to perform poorly in content recognition based on the trained algorithm it uses, and for this reason or another reason needs to be retrained, then this conventional approach is found to be more problematic. For the above reasons, retraining will typically result in a significant delay, as new annotated training images must be acquired from and / or by someone. The retraining is then performed on these images, which can typically be involved remotely and / or by another party than the one operating the system.
[0042] Many machine vision systems are at least to some extent flexible and reconfigurable, so they can be configured for different image content recognition tasks (e.g. if the objects to be recognized change). That is, the system is not just for fixed use according to the first application and task for which the system is set up and used. It would be beneficial if such changes and reconfigurations could be done as easily and efficiently as possible without, for example, involving the system provider and / or any other party retraining. It is also desirable if it is not necessary to acquire and annotate a large number of new training images, perform training of these images, then have to test the newly trained algorithm on the machine vision system, and repeat the process if the algorithm does not perform well enough, etc.
[0043] In view of the above observations, it is recognized that there is a need to be able to more efficiently (re)train the image content recognition algorithm part of a machine vision system, in particular already installed and functioning properly, when the application and / or image recognition task(s) change.
[0044] For example, the training images can be captured locally by the machine vision system itself. However, for the same reasons as above, there would still be a need to ingest, annotate many images, then perform training on them, etc. This still takes time, and if performed by the operator of the machine vision system, this is different type of work than during normal operation. Another problem that can arise in such a case is the memory limitations of the machine vision system, e.g. the camera is typically not suited for storing and processing large amounts of images, as this is not needed during normal operation.
[0045] The embodiments herein are based on the recognition that a solution to the problem indicated above and in the background art would be to "pre-annotate" the images before they are ingested by the machine vision system, i.e. to first determine the annotation information, for example, then generate the images by the machine vision system according to this annotation information. If the system knows the annotation in advance, it can automatically perform the annotation of the ingested images. The system, e.g. the camera, can be configured to associate the annotation information that is already present from this with the images that are being captured according to this annotation information.
[0046] When using the system to generate annotated training images, such an algorithm is preferably also used directly on each captured annotated image for content recognition, i.e. although the main purpose at this time is to generate annotated training images, the algorithm is correspondingly used as during normal operation. The system should then preferably automatically apply the algorithm to the respective training image being captured and perform a check to see if the algorithm is able to recognize the content from the annotation information. In other words, in principle, the algorithm is used and tested directly on the annotated training images. Note that this is different from training the algorithm on the images. In this way, the system can know which images will be most useful to train on, i.e. the images for which the algorithm fails to recognize the content from the annotation information. Other images can be discarded. This enables less but more useful annotated training images to be saved and used for later training, enabling the need for storing large amounts of images to be reduced.
[0047] When the check indicates that useful annotated training images have been generated, or after a certain number of such images have been generated, the (re)training mode can even be entered temporarily and even automatically. After such a training session and the thereby improved image content recognition capability, the process of capturing new images for training can continue, but this time using the newly trained and thereby improved algorithm. Thus, the useful annotated training images and the training on these images, in which the algorithm and its capabilities are also tested, can be done in an iterative and combined manner. Note that at the start, the algorithm can have no or only basic content recognition capabilities, e.g. from an earlier basic training and / or some default capabilities accordingly.
[0048] Further, as realized from above, by said inspection, more or less directly after an image has been captured by the system, one can have available information whether or not the image is an image on which (further) training is of interest. This in principle enables providing feedback to a control system or user (e.g. an operator of the machine vision system) controlling what is being imaged (such as an object) in real time, so that it can be more or less directly known when something (e.g. a certain pose of an object) is being imaged which cannot be identified by the algorithm from the annotation information, so that training on which would be useful. The control system or user can use and / or learn from this feedback how to produce and be encouraged to produce training images which are more useful than otherwise. I.e. there is assistance in the supply of images according to the annotation information, but which are too tricky for the algorithm to handle without further training, and thus of particular interest on which to train. For example, the control system or user can “compete” with the algorithm and try to outdo the algorithm by trying to produce images according to the annotation information, but which the algorithm fails to identify as being according to the annotation information. Further, if (re)training is included in the process, such “competition” with the algorithm will be a test of the algorithm at the same time, and when the algorithm has or seems to have trained well enough for use during normal operation, the training and generation of further annotated training images can be stopped. Such testing and time for this is in normal cases needed anyway.
[0049] Hence, the above embodiments herein enable efficient supply of annotated training images in a process which can be combined with (re)training and testing of a trained algorithm, providing further advantages.
[0050] Figure 2 An example of a machine vision system 205 which can be used and / or configured to perform embodiments herein is schematically illustrated. The machine vision system 205 comprises a camera 230 having an image sensor, which in the figure and herein is exemplified by being an image sensor 234.
[0051] The camera 230 is arranged and / or configured to capture images as part of the operation of the machine vision system 205, i.e. as part of its normal operation. These images are images of what is located in the field of view 232 of the camera 230 when the images are captured. For example, the respective images can be imaging one or more objects. What is being imaged typically changes between images in some way, e.g. the object(s) change and / or the object(s) change position and / or pose between images, such as due to movement of the object(s) in and / or across the field of view 232. For example, there can be a conveyor belt (not shown) or similar that changes what is in the field of view 232. Of course, other ways of replacing and / or placing what is to be imaged in the field of view 232 are possible, such as due to operation and movement of a robotic arm or similar or even by a human, or due to any other way of doing this known from the prior art. In the illustrated example, there are exemplary illustrative objects of different geometrical shapes being imaged; a first object 220-1 that is a cube and a second object 220-2 that is a sphere. Thus, the camera 230 can image the cube as a square in 2D and the sphere as a circle.
[0052] In addition to the image sensor 234, the camera 230 can further comprise one or more processors 235, memory, etc., and / or the image sensor 234 itself can have some integrated processing capabilities. Of course, there are typically also input and / or output interfaces, etc., such as for input and / or output of control and / or information for the camera, such as the captured image(s) and / or information extracted from the captured images. The one or more processors 235 can be used to integrate the same or similar capabilities as a computer in the camera, i.e. in the same unit, making it more autonomous and able to perform more of its own processing and / or control. I.e. to obtain the corresponding functionality without the need for further and / or separate units, such as a separate computer, and / or to enable a simple setup procedure and / or faster processing, e.g. due to reduced need for communication with separate units. In some embodiments, the system 205 corresponds to the camera 230, i.e. the system 205 can be in the form of or implemented as a single unit that is the camera, corresponds to the camera or at least resembles the camera and contains the camera functionality. Furthermore, network capabilities, light sources, etc. can also be integrated in such a unit, e.g. in the camera corresponding to the camera 232.
[0053] However, in some embodiments, the camera 232 is a camera unit with more specific imaging capabilities only, e.g. some image processing capabilities can be in the camera 232 in addition to the imaging capabilities. The camera 232 can be connected to one or more other parts of the system, such a control device 242 with further processing capabilities as exemplified in the figure. The control device 242 can be configured to control the camera 230 as well as other parts of the system 205, if any. The control device 242 can comprise one or more processors, memory, etc. and can correspond to a computer.
[0054] By means of a trainable image content recognition algorithm, the machine vision system 205 is operable to recognize content in images captured by the camera 230. Such a trainable image content recognition algorithm can be a state-of-the-art algorithm and can be stored in and / or executed by the camera 230 or the control device 242 or other parts of the system 205 with suitable computing power. The same unit configured to execute the algorithm is preferably configured such that it can also train the algorithm using annotated training images. The algorithm can alternatively be trained by another unit which can be part of the system 205 or external to the machine vision system 205 even remotely.
[0055] The camera can be a 2D or 3D camera, can have built-in processing capabilities and / or have communication (such as streaming) capabilities of images to another device, e.g. to the control device 242 or other parts of the system 205, e.g. with suitable computing power for executing the trainable image content recognition algorithm and / or for its training and / or for executing the actions of the embodiments herein.
[0056] The system 230 can comprise a user interface and / or user interface device, e.g. a user interface arrangement 245, such as a display or computer screen, which can correspond to or be part of a computing unit 244, such as a laptop or stationary computer. As used herein, a user interface arrangement means an arrangement, such as a display, computer screen and / or computer, suitable for acting as a user's, typically a human, e.g. an operator of the system 230, interface, as exemplified by the user 252 in the figure. The user interface can be unidirectional, but is typically bidirectional, i.e. both for providing information to the user 252, such as visually and / or aurally, and for receiving input from the user 252.
[0057] In some embodiments, the user interface arrangement 245 is the same as or partly combined with the control unit 242. The user interface arrangement 245, e.g. a display, such as a computer screen, can be connected to another device, such as the control device 242 and / or the camera, with processing capabilities with respect to input and / or output via the user interface arrangement 245.
[0058] As described above, including the background art, trainable image content recognition algorithms that can be used with the embodiments herein, as those skilled in the art should recognize, can be described as trainable algorithms for performing image recognition-based tasks (i.e., image recognition tasks). Image content recognition algorithms are configured and / or at least partially trained and / or trainable to perform such tasks. These tasks, other than image recognition like this, can vary between different algorithms and applications. Therefore, the image content (e.g., objects(one or more) in an image) and the attributes of interest for annotations used for recognition can differ between different algorithms and tasks. However, such trainable image content recognition algorithms and any tasks they are configured to perform are not special to the embodiments herein and can be conventional, such as those in the prior art. The training of such trainable image attribute recognition algorithms based on annotated training images can correspond to that discussed above and in the background art.
[0059] Figure 3A This is a flowchart used to schematically illustrate an embodiment of the method according to the embodiments herein. Figure 3B It is a schematic diagram and illustration as Figure 3A A flowchart of some of the actions in a method.
[0060] The following actions of the method can be used to provide one or more annotated training images for use in a trainable image content recognition algorithm for training a machine vision system (such as machine vision system 205, which will be used to illustrate the machine vision system below). Figure 3B The annotated training images 336a-b are used below to illustrate the annotated training images. With the aid of a trainable image content recognition algorithm, the machine vision system 205 is operable to recognize the content in the images captured by the machine vision system 205.
[0061] Below Figure 3A The methods and / or operations indicated herein may be performed by one or more devices (i.e., one or more devices, such as one or more devices of the machine vision system 205, for example, camera 230 and / or control device 242). In some embodiments, the one or more devices are or correspond to the machine vision system 205. Therefore, the methods and / or their operations may be at least partially computer-implemented. The one or more devices for performing the methods and their actions are further described below.
[0062] Note that the following operations can be performed in any suitable order, and / or may be performed in complete or partial overlap where possible and appropriate.
[0063] Action 301
[0064] Annotation information is acquired for one or more intermediate images to be captured by the machine vision system 205. The annotation information indicates (such as reveals, identifies or describes) properties of content in the intermediate images. The content is for being recognized by a trainable image content recognition algorithm. These properties are such that, if the algorithm were to be trained on the intermediate images, the training would benefit from the knowledge that the intermediate images contain content having the properties.
[0065] In Figure 3B the example, the annotation information 337 reveals a "single object with a square shape" or correspondingly "a single cuboid object imaged". To simplify and facilitate understanding of the example, the annotation information 337 is referred to as "single square object" or "SSO" in Figure 3B the following. In the following actions, the annotation information 337 will be used as a non-limiting example of annotation information. What the annotation information can be (including further examples) will be discussed separately below.
[0066] The annotation information 337 can be provided (such as inputted and / or sent) to the device(s) performing the present action (e.g. the machine vision system 205) which thereby acquires (e.g. receives) the annotation information 337. This provision to the device(s) can be done by a user (such as the user 252, e.g. an operator of the machine vision system 205). The provision by inputting can be via a user interface (UI), preferably such as via the user interface arrangement 245 and / or a graphical user interface (GUI) thereon.
[0067] In some embodiments, the acquired annotation information is fully or partially predetermined. The user can select which specific annotations to use for the intermediate images to be captured by the machine vision system 205, e.g. from a predetermined annotation information (e.g. from a list), as part of the following action 302, e.g. via the UI.
[0068] Action 302
[0069] The machine vision system 205 is operated or operated such that the one or more intermediate images are captured by the machine vision system 205 in accordance with the acquired annotation information 337. In other words, the intermediate images being captured are valid for the annotation information.
[0070] In practice, for ease of implementation of the method and / or with existing systems, the operational actions of the method can partly involve a user, typically the same user, such as an operator of the system, as already mentioned, e.g. the user 252. In this operational step, the user 252 can control what is being imaged during the operation, based on the annotation information 337 that the user himself can have inputted into the system as described above, so that the intermediate images will be in accordance with the annotation information 337. In the example, it is thus ensured that the single cuboid object is imaged in different poses, one pose per image, such as the intermediate images 338a-c that will be used as examples in the following. Additionally, or alternatively, the user 252 can be prompted, suggested and / or instructed via the UI or another UI what is to be imaged, such as certain object(s), and / or variations to be covered, such as different poses or arrangements of what is being imaged, e.g. the certain object(s). Of course, this should also be done so that the intermediate images 338a-c captured by the system 205 will be in accordance with the annotation information, and / or facilitate capturing variations in the intermediate images that are useful for training.
[0071] The operational actions can alternatively be more or fully automated, e.g. without intervention of the user, wherein the machine vision system 205 can be configured to control what is to be imaged based on the annotation information directly itself. The annotation information can be predetermined and / or associated with information about what is to be imaged, so that the intermediate images will be in accordance with the annotation information. It is even conceivable with a system that itself determines what would be beneficial for training, the system providing annotation information accordingly, and then controlling what is to be imaged based on the determined and in accordance with the annotation information. For example, a robot that holds and / or places various combinations and / or poses of object(s) to be imaged based on instructions from the machine vision system 205.
[0072] Action 303
[0073] One or more annotated intermediate images are provided, preferably by means of the machine vision system 205. The one or more annotated intermediate images correspond to the one or more intermediate images 338 captured by the machine vision system 205 in accordance with the obtained annotation information 337, as well as the annotations. Again with reference to the example of Figure 3B annotated intermediate images are exemplified by the annotated intermediate images 339a-c in the following. The annotation information is further and separately discussed below.
[0074] Action 304
[0075] In some embodiments, the trainable image content recognition algorithm with the first capability of content recognition is applied on the one or more annotated intermediate images 339a-c to recognize the content of the respective annotated intermediate image 339a-c. In these embodiments, also a check is made whether the recognition is according to the annotation information of the respective annotated intermediate image.
[0076] It will be appreciated that this action can easily be automated. In addition to what the algorithm recognizes, also a check is thus made whether what is being recognized is according to the annotation information. For example, as in the example of Figure 3B if the annotation information 337 is about only a single square object in each image, a check is made whether the single square object has been recognized, such as detected, in the image by the algorithm as expected, or not. As a further example: if the annotation information is instead about a certain number of apples and / or only apples being imaged, this check recognizes whether this is according to the annotation information, and for example, the correct number and / or only apples (and for example, no pears) are recognized. As realized, in the context of the embodiments herein, it is mainly of interest to know when there is a failure, i.e. when the recognition is not according to the annotation information. This indicates that an improvement would be beneficial, and thus for which corresponding intermediate image to train on to achieve a second capability, which is an improvement of the first capability. This will be further discussed below under action 306.
[0077] Action 305
[0078] In some embodiments, in response to the algorithm with the first capability failing to recognize the content of the respective annotated intermediate image, e.g. intermediate image 339b, according to the annotation information 337 of the respective annotated intermediate image, failure information is provided via a user interface of the machine vision system 205. The user interface can for example belong to the user interface arrangement 245. In this way, a user, such as the user 252, who controls what is being imaged by the machine vision system 205, can obtain feedback about the failure via the user interface. Thereby, the user can control the next image(s) based on the failure information to increase the likelihood of further failures, and thereby increase the likelihood of useful training images.
[0079] Preferably, the feedback is direct and / or associated with what was imaged and the respective annotated intermediate image 339a-c generated thereby. That is, in this way the user can get direct or as fast as possible feedback and response to what was imaged by the user control. For example, in the example of Figure 3BIn the example of Fig. 3, the user 252 is provided with feedback on the failure of the algorithm with the first capability to recognize the content of the annotated intermediate image 339b. This feedback is provided in the form of the annotated intermediate image 339b itself, which is provided as an annotated training image 336a. In this example, if the user 252 placed a cube object with a pose that resulted in the annotated intermediate image 339b, and the machine vision system 205 with the first capability failed to recognize that the image showed a single square object, then the user 252 should directly get feedback on the failure, and thus be told in this example that the object pose was problematic for the algorithm. Such a failure is good in this context, as it means that an image has been produced that is particularly valuable to train on. Thus, the user 252 can discover and learn what the algorithm is not good at recognizing correctly, and the user can base on this control what is subsequently imaged. This means that annotated training images are generated more efficiently, i.e. such that if the algorithm is trained on these images, an improved capability should result.
[0080] Note that the user interface need not be advanced, or even a GUI, such as via a computer screen. Alternatively or additionally, simple light and / or audio signals (e.g. color-coded and / or via light emitting diodes) can be used.
[0081] Action 306
[0082] Based on at least one of the annotated intermediate images 339a-c, one or more annotated training images 336a-b are provided. The method need not provide all annotated intermediate images as annotated training images (in other words, as output from the method), although this can be the case.
[0083] In embodiments where action 304 is performed, in response to the algorithm with the first capability failing to recognize the content of a respective annotated intermediate image 339b in accordance with the annotation information of the respective annotated intermediate image 339b, the respective annotated intermediate image (e.g. annotated intermediate image 339b) of the annotated intermediate images 339a-c can be provided as a respective annotated training image (e.g. annotated training image 336a) of the annotated training images. In this way, the user 252 can be provided with feedback on the failure of the algorithm with the first capability to recognize the content of the respective annotated intermediate image 339b. Figure 3B In the example of Fig. 3, this can be the case if the algorithm with the first capability fails to detect that there is a single square object in the annotated intermediate image 339b (as the algorithm does not adapt to the pose of the object when the object is imaged in the annotated intermediate image 339b). In this way, useful annotated intermediate images can be efficiently identified and provided as annotated training images, while other less useful intermediate images can be discarded. There is no benefit, or at least less benefit, in storing or using intermediate images for which the algorithm is already able to correctly recognize the image content in accordance with the associated annotation information.
[0084] Note that this action (as well as action 304) can also be performed in an automated manner.
[0085] As indicated in the figure, when the present action has been performed, action 302 can be performed again to capture yet another image or images according to the annotation information, e.g. with different poses and placements of the cube object, such that further image(s) forming a single square object are formed, and then actions 303-306 can be performed for said further image(s), etc.
[0086] Action 307
[0087] In some embodiments, the respective annotated intermediate images provided in action 306 as respective annotated training images, such as annotated training image 336a, are stored in memory for later use, e.g. annotated intermediate image 339b.
[0088] As indicated in the figure, in embodiments with the present action and when the action has been performed, action 302 can be performed again to capture yet another image or images according to the annotation information, e.g. with different poses and placements of the cube object, such that further image(s) forming a single square object are formed, and then actions 303-307 can be performed for said further image(s), etc.
[0089] Action 308
[0090] In embodiments where action 307 is performed, the remaining annotated intermediate image(s) not provided as the one or more annotated training images, e.g. annotated training image 336a, are discarded, e.g. annotated intermediate image 339a. This can save memory storage space.
[0091] The respective annotated training images, such as annotated training image 336a, can first be stored locally, e.g. in the camera, such as camera 230, or other device of the machine vision system 205. This can be the case when the images are to be used for training soon, i.e. later but in close time. Additionally, or alternatively, the respective annotated training images can be stored outside of the system 205 or even remotely, such as on a server or computer cloud, thus stored after transmission. Such storage can be for long-term storage and later use, such as later use separate from such a method, e.g. later to be used for training or training of the same or similar systems and / or algorithms in the machine vision system 205.
[0092] As indicated in the figure, in embodiments with the present action and when the action has been performed, action 302 can be performed again to capture yet another image or images according to the annotation information, e.g. with a different pose and placement of the cube object, such that a further image(s) of the single square object is formed, then actions 303-308 can be performed for said further image(s), etc.
[0093] Action 309
[0094] In some embodiments, the trainable image content recognition algorithm is trained to implement a second capability of content recognition using a respective annotated intermediate image provided as a respective annotated training image. That is, the training uses training images with their associated annotation information. For example, if annotated intermediate image 339b is provided as annotated training image 336a, this image is used in the training with its annotation information 337, which reveals that this image belongs to the single square object, whereby the algorithm implements a second capability that it should better recognize square objects with an imaging as in annotated intermediate image 339b.
[0095] The second capability should be an improvement over the first capability. After having been used for the training, the respective annotated training image, such as annotated training image 336a, can be discarded at least locally to save memory storage space. Additionally or alternatively, it can be transmitted elsewhere and stored, and / or kept stored for a long time. It can be reused later, e.g. at a later point in time for training the same or similar system and / or for the same application.
[0096] In some embodiments, the training is performed in response to the algorithm with the first capability failing to recognize the content of a respective annotated intermediate image, e.g. annotated intermediate image 339b, according to the annotation information 337 of the respective annotated intermediate image. As already discussed above, the failure is indicative of a particular interest in training on the annotated intermediate image.
[0097] The training can be performed automatically and / or directly, e.g. as fast as possible, after a time period after a failure or after a certain (e.g. predetermined) time period and / or after a certain amount of annotated training images has been collected and / or in relation to the available storage of the local storage for annotated training images. The latter can be a trigger for performing the training, as the local storage can be a scarce resource. For example, after a certain amount of annotated sample images has been provided and locally stored, the training can be performed based on these images. Additionally, or alternatively, the training can be performed after a certain time period during operation and / or at a certain point in time dedicated for training and / or in response to a user input, e.g. the user triggers or allows the training to be performed, e.g. using the locally stored annotated sample image(s).
[0098] The training can be performed automatically, e.g. as fast as possible, when the system or camera has the possibility, capacity and ability to train and one or more provided annotated training images are available. In this way, the improved capabilities can be completed and made available as fast as possible without interrupting or disturbing other uses of the system.
[0099] As indicated in the figure, in embodiments with the present action, when the action has been performed, the action 302 can be performed again to capture yet another image or images according to the annotation information, e.g. with a different pose and placement of the cube object, such that further image(s) forming a single square object are formed, for which the actions 303-309 can be performed next, etc.
[0100] Thanks to the embodiments described herein and as outlined above with respect to Figure 3A-Figure 3B By virtue of the embodiments described herein and as outlined above with respect to
[0101] Annotating, types and / or formats of annotated images like this, how they are associated with images, etc. can be as in the prior art. The difference compared to the prior art is that the annotation information is pre-acquired before the image with which the annotation information is related is captured by the machine vision system. Thus, the annotation information is already available when the image is captured, and thus can be directly and / or automatically annotated by the system without the need for inspection of images like this. As identified from the discussion in the background and introductory sections above, this is different from the conventional approach to image annotation in training of trainable image content recognition algorithms.
[0102] Furthermore, as realized, the approach according to the embodiments here requires access to a machine vision system using the algorithm, and thus is not generally applicable to all trainable image content recognition algorithms and training in all cases. However, as should be realized from the examples herein, in the case of a machine vision system using such an algorithm, it is a benefit rather than a drawback to have to involve the system like this to provide annotated sample images for training.
[0103] Figure 4 is a block diagram schematically illustrating and exemplifying some embodiments of the approach discussed above in relation to Figure 3A-Figure 3B the discussion above. The example is basically an overview of how annotated training images can be provided in an iterative manner based on the embodiments herein. The actions involved in the example have been described above and correspond to actions 301-309. The annotation information is as in the previous example, thus annotation information 337 revealing that the image should be of a single square object (SSO). The machine vision system 205 and the trainable image content recognition algorithm start with a first capability. This is illustrated by:
[0104] first round a) the machine vision system 205 is operated such that the cube object 220-1 is imaged by the first camera 230, whereby a first intermediate image 338a is captured according to the acquired annotation information, i.e. as indicated in the figure, an SSO with a first pose and placement is imaged. Then, the machine vision system 205 involves, e.g. automatically, providing an annotated intermediate image 339a corresponding to the intermediate image 338a and the annotation information, e.g. according to the annotation information 337. The recognition is checked against the annotation information 337. I.e. it is checked whether the algorithm with the first capability is able to recognize that the image is imaging a single square object. In the present example, it is assumed that the algorithm succeeds in the present example, whereby the intermediate image 339a is discarded, e.g. from memory in the camera 230. The reason is as discussed above, if the algorithm and the machine vision system 205 already are able to recognize the imaged object in the image according to the annotation information, there is no use in saving the annotated intermediate image for use in later training of the algorithm and the machine vision system 205.
[0105] Second round b): This round is basically a repetition of the actions in round a), but now the machine vision system 205 is operated such that the cuboid object 220-1 is imaged by the camera 230 in a different placement and pose, and an intermediate image 338b is captured according to the obtained annotation information. Then, the machine vision system 205 involves providing an annotated intermediate image 339b corresponding to the intermediate image with the annotation according to the annotation information 337. The recognition is checked against the annotation information 337. In the present example, it is assumed that the algorithm fails in this check. Thus, the algorithm with the first capability fails to recognize that the image is imaging a single cuboid object. Therefore, the intermediate image 339b should be useful for training on, and is thus kept and provided as an annotated training image 336a. That is, the provided annotated training image 336a is saved for use in training. In the shown example, the training is performed directly using the annotated training image 336a, i.e. the machine vision system 205 can enter a training mode and perform training on the annotated training image 336a, whereby the training enables a new second capability. With the second capability, the algorithm should have learned to better recognize objects with a placement and pose as in the intermediate image 339b. Of course, the provided annotated training image 336a can also be saved and stored for later (re-)use.
[0106] Third round c): This round is just to show that as long as desired, the rounds as described above for round a) and round b) can continue, whereby further useful annotated training images can be provided and the algorithm trained to have further improved image content recognition capabilities.
[0107] It is to be realized that the check in combination with the provision of annotated training images and training implies that there will be a test of the trained algorithm, which test can be used to determine when the algorithm is considered trained sufficiently for use on production images. Anyway, such a test is often needed, but conventionally a separate action from both the provision of annotated training images and the training like this.
[0108] As already indicated above, the embodiments herein are based on very unconventional annotations regarding their origin, when and how they are used to some extent, but the annotation like this and how images are associated with annotation information can be conventional. Thus, the annotation like this, e.g. the type and / or format of the annotation and what the annotation is about, can be conventional, such as in the prior art. However, the annotation is related to properties of the image content to be recognized, which in turn are related to the algorithm and the task it is configured to perform. Thus, further examples of image content properties and annotation information can best be understood by considering different types of image recognition tasks and examples of image content that an image recognition algorithm as in the embodiments herein can be configured to perform.
[0109] Examples of image recognition tasks:
[0110] • Anomaly detection - does an object (e.g. a circuit board) look as usual, as it should or normal, or is there an abnormality with it, such as a scratch, misaligned solder spot, missing component?
[0111] • Classification - does an image look more like class A or class B, e.g. is there sunshine or rain in a scene?
[0112] • Counting - how many instances of a certain type of object are present in an image? For example, how many screws are present on this face of a machine part?
[0113] • Localization - how is an object in an image disposed? For example, what is the precise location at which a robot should place a screw in a hole?
[0114] Combinations are of course possible. For example: detect individual object(s) and classify them and find out if there is an anomaly.
[0115] To give another more specific example, content recognition can involve one or more of: detecting the localization(s) of individual fruit in an image, classifying the imaged fruit(s) based on type, such as apple(s) or pear(s), detecting if any fruit is damaged, and counting them.
[0116] As already indicated above, the relevant attributes identified by the annotation information should be those that can be utilized in training, or in other words, are attributes of the image content that it is useful to know for training using the training images having these attributes. Since the annotation information is for use when training is performed on the training image(s) with which the annotation information is associated, the attributes should be those that training would benefit from knowing that these attributes are present in the training image(s).
[0117] Examples of attributes of image content that can be the subject of annotation:
[0118] • One or more categories to which the image belongs.
[0119] • One or more categories to which the image does not belong.
[0120] • Presence or absence of a certain type of object in the image.
[0121] • Number of a certain type of object present in the image.
[0122] • Total number of objects present in the image.
[0123] • The object does not overlap or touch any other object in the image.
[0124] • The position of the center of the object in the image.
[0125] • The shape and size of the object in the image.
[0126] Here the combination is of course also possible.
[0127] If the image content contains only normal content or only abnormal content (such as a damaged fruit as in the previous example), the attributes can thus relate to the object type or class of the object(s), and / or the object material, and / or the object color, and / or the geometric form or shape of the object(s), and / or the number of object(s) and / or object overlap. For example, the annotation information can identify one or more of the following attributes:
[0128] • All objects in the image are apples and / or no object is a pear.
[0129] • All objects in the image are rectangular and / or no object is circular.
[0130] • All objects in the image overlap or no object overlaps.
[0131] • There is a certain number of objects in the image.
[0132] Here is a more specific example for further insight and understanding:
[0133] Suppose an Italian pasta manufacturer wants to check a station that should verify that the type of pasta present in the packages leaving its production line is the expected type. To train a machine vision system in the check station to recognize the type of pasta, the user can input the type of pasta as (pre-)annotation into the machine vision system as in the embodiments herein. The production line and the machine vision system can then be operated such that the type of pasta according to the annotation information is passed by the machine vision system. If the system is configured to automatically retrain when the system fails to correctly predict the type of pasta, the user can fully train the system by alternating between different types of pasta until the machine vision system no longer makes any mistakes. The advantage in this example will be that the user does not have to manage and label images at all, just make sure that the pasta production operates as expected during the image collection. It is enough to operate the production line, the machine vision system, and provide the correct (pre-)annotation information. The image collection process (i.e. the supply of annotated training images) as well as the training can be stopped as soon as the system reaches the expected performance corresponding to a sufficient image content recognition ability with respect to the pasta. In the case of using a regular regime, since it is not possible to know beforehand how many annotated training images are needed, it will either have to collect more annotated training images than are actually necessary for the sufficient training of the system, or iteratively go back and collect more images until satisfactory performance is reached.
[0134] Figure 5 is a schematic block diagram illustrating an embodiment of one or more devices 500 (i.e. the device(s) 500) that can correspond to the device(s) for performing the embodiments herein (such as for performing the methods and / or actions described above with respect to the embodiments herein) described above. Figure 3A
[0135] The device(s) 500 can for example correspond to any of the machine vision system 205, the camera 230, the control device 242, or other suitable device part(s) of or connected to the machine vision system 205. As will be realized by those skilled in the art, some embodiments of the method comprise actions that can be distributed for execution by a plurality of devices configured to perform the actions. Also, the control of the machine vision system can be at least partly executed remotely, for example by or via a remote computer, server or computer cloud, and there can be corresponding remote device(s) that can be configured to perform the method or some action(s) therein. However, it is typically preferred to use the device(s) configured to perform the method as part of or in close connection with the machine vision system 205 and imaging, as this enables faster and typically more stable execution and less information transfer.
[0136] The schematic block diagram is used to illustrate how the device(s) 500 are configured to perform embodiments of the method and actions discussed above in connection with Figure 3A-Figure 3B the machine vision system 205. Thus, the device(s) 500 are used to provide one or more annotated training images, such as the training images 336a-b, for use in training of a trainable image content recognition algorithm of a machine vision system, such as the machine vision system 205. The machine vision system is operable by means of the trainable image content recognition algorithm to recognize content in images captured by the machine vision system.
[0137] The device(s) 500 can comprise processing module(s) 501, such as processing components, one or more hardware modules, including for example one or more processing circuits, circuitry, such as processors, and / or one or more software modules for performing the described method and / or actions.
[0138] The device(s) 500 can further comprise one or more memories 502 that can comprise, such as contain or store, computer program(s) 503. The computer program(s) 503 comprise “instructions” or “code” that are directly or indirectly executable by the device(s) 500 respectively to perform the described method and / or actions. The one or more memories 502 can comprise one or more memory units, and can further be arranged to store data, such as configurations, data and / or values, relating to or used for performing the functions and actions of the embodiments herein.
[0139] Further, respective device(s) 500 can comprise processing circuitry 504 relating to processing (e.g. executing and training algorithms) as an exemplifying hardware module(s), and can comprise or correspond to one or more processors or processing circuitry. Processing module(s) 501 can comprise, e.g. be "embodied in the form of" or "implemented by" such processing circuitry 504. In these embodiments, memory 502 can comprise computer program(s) 503 executable by processing circuitry 504, respectively, whereby respective device(s) 500 are operative or configured to perform the described methods and / or actions thereof.
[0140] Generally, device(s) 500 (e.g. processing module(s) 501) comprise input / output (I / O) module(s) 505 configured to (e.g. by performing) any communication to and / or from other units and / or devices, such as sending and / or receiving information to and / or from other devices. When applicable, I / O module(s) 505 can be exemplified by taking (e.g. receiving) and / or providing (e.g. sending) modules.
[0141] Further, in some embodiments, device(s) 500 (e.g. processing module(s) 501) comprise one or more of acquisition module(s), operation module(s), providing module(s), application module(s), checking module(s), storage module(s), discarding module(s), training module(s), as exemplifying hardware and / or software modules for performing actions of embodiments herein. These modules can be implemented in whole or in part by processing circuitry 504.
[0142] Thus:
[0143] Device(s) 500, and / or processing module(s) 501, and / or processing circuitry 504, and / or I / O module(s) 505, and / or acquisition module(s) are operative or configured to acquire the annotation information for the one or more intermediate images.
[0144] Device(s) 500, and / or processing module(s) 501, and / or processing circuitry 504, and / or I / O module(s) 505, and / or operation module(s) are operative or configured to operate or enable operation of a machine vision system such that the machine vision system captures the one or more intermediate images in accordance with the acquired annotation information.
[0145] The device(s) 500, and / or the processing module(s) 501, and / or the processing circuitry 504, and / or the I / O module(s) 505, and / or the providing module(s) are operable or configured to provide the one or more annotated intermediate images, preferably by means of a machine vision system.
[0146] The device(s) 500, and / or the processing module(s) 501, and / or the processing circuitry 504, and / or the I / O module(s) 505, and / or the providing module(s) are operable or configured to provide the one or more annotated training images based on at least one of the annotated intermediate images.
[0147] In some embodiments, the device(s) 500, and / or the processing module(s) 501, and / or the processing circuitry 504, and / or the I / O module(s) 505, and / or the application module(s), and / or the checking module(s) are operable or configured to apply the trainable image content recognition algorithm having the first capability of content recognition to the one or more annotated intermediate images to recognize content of a respective annotated intermediate image, and to check whether the recognition is in accordance with annotation information of the respective annotated intermediate image.
[0148] In some embodiments, the device(s) 500, and / or the processing module(s) 501, and / or the processing circuitry 504, and / or the I / O module(s) 505, and / or the providing module(s) are operable or configured to provide the failure information via the user interface in response to the algorithm having the first capability failing to recognize content of a respective annotated intermediate image in accordance with annotation information of the respective annotated intermediate image.
[0149] In some embodiments, the device(s) 500, and / or the processing module(s) 501, and / or the processing circuitry 504, and / or the I / O module(s) 505, and / or the storage module(s) are operable or configured to store the respective annotated intermediate image provided as a respective annotated training image in a memory for later use.
[0150] In some embodiments, the device(s) 500, and / or the processing module(s) 501, and / or the processing circuitry 504, and / or the I / O module(s) 505, and / or the discarding module(s) are operable or configured to discard any remaining annotated intermediate images that are not provided as annotated training images.
[0151] In some embodiments, the device(s) 500, and / or the processing module(s) 501, and / or processing circuitry 504, and / or I / O module(s) 505, and / or training module(s) are operable or configured to train the trainable image content recognition algorithm using respective annotated intermediate images provided as respective annotated training images with associated annotation information to implement a second capability of the content recognition.
[0152] Figure 6 is a schematic diagram illustrating some embodiments of the device(s) 500 discussed above, causing the method and actions discussed above to be performed by the computer program 503 and its carrier.
[0153] The computer program(s) 503 comprise instructions which, when executed by the processing circuitry 504 and / or processing module(s) 501, cause the apparatus 500 to carry out any of the embodiments as described above. In some embodiments, one or more carriers, i.e. carrier(s), or more specifically data carrier(s) such as computer program product(s), comprising the computer program(s) are provided. The respective carrier can be one of an electronic signal, optical signal, radio signal, and computer readable storage medium (e.g. computer readable storage medium 601 as schematically depicted in the figure). The computer program(s) 503 can therefore be stored on a computer readable storage medium 501. A carrier can exclude a transitory signal per se, and a data carrier can correspondingly be named non-transitory data carrier. A non-limiting example of a data carrier as computer readable storage medium is a storage card or memory stick, a disk storage medium, or generally a mass storage device based on hard drive(s) or solid state drive(s) (SSD). The computer readable storage medium 601 can be used to store data accessible to a computer network 602, e.g. the Internet or a local area network (LAN). Furthermore, the computer program(s) 503 can be provided as pure computer programs, or include in one or more files. The file(s) can be stored on the computer readable storage medium 601 and available e.g. through download (e.g. through the computer network 602 as indicated in the figure, such as via a server). The server can be a Web or File Transfer Protocol (FTP) server or similar. The file(s) can be executable files for direct or indirect download to and execution on the apparatus to cause the file to be executed as described above, e.g. through execution by the processing circuitry 504. The file(s) can also or alternatively be used for intermediate download and compilation involving the same or another processor to make the file executable before further download and execution to cause the apparatus 500 to carry out any of the embodiments as described above.
[0154] As used herein, a training or sample image can be described as an image used in the training of a trainable image content recognition algorithm to recognize content, or in other words, with some properties about the image content known for the training image to solve the task of the algorithm and / or improve the ability of the algorithm to do so.
[0155] Further, as used herein, a production image can be described as an image for which image content is to be identified by a trainable image content recognition algorithm without prior knowledge about what the image shows or attributes of the image content as compared to a training image. Thus, a production image is an image on which a machine vision system is intended to operate in normal operation. In other words, a production image can be considered to be an image for which it is a task of the machine vision system to identify content therein by means of the algorithm (e.g., make predictions and / or assertions about the image content).
[0156] Further, in connection with the examples herein, examples of types of annotation information and some particular attributes. Note that these are just some examples of the many possible examples and embodiments herein are not limited to any particular trainable image content recognition algorithm, image content recognition task of such algorithms, attributes of image content, and / or types and / or formats of annotation information. As recognized, content recognition in connection with embodiments herein is generally about, but not necessarily about, object recognition. As already indicated and as should be recognized by those skilled in the art, embodiments herein are not about or limited to any particular trainable image content recognition algorithm, but still refer to some examples, which include deep learning based methods, e.g., convolutional neural networks such as ResNet or MobileNet, or any conventional machine vision algorithm, including any variants and / or combinations of such methods and algorithms. As already indicated before, annotation information like this (e.g., types and / or formats of this annotation information and what this annotation information is about) can be conventional, which includes e.g., that it can be graphical and / or alphanumeric form representations such as object mask(s) and / or shape(s) and / or count of object(s) correspond.
[0157] Note that any processing module(s) and circuit(s) mentioned in the foregoing can be implemented as software and / or hardware modules, e.g., in existing hardware, and / or as application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), etc. Also note that any hardware module(s) and circuit(s) mentioned in the foregoing can e.g., be included in a single ASIC or FPGA, or distributed among several separate hardware components, whether individually packaged or assembled into a system on a chip (SoC).
[0158] Those skilled in the art will also appreciate that modules and circuits discussed herein can refer to a combination of hardware modules, software modules, analog and digital circuits, and / or one or more processors configured with software and / or firmware (e.g., stored in memory), that when executed by the one or more processors, can cause the device(s), sensor(s), etc. to be configured to and / or perform the methods and acts described above.
[0159] The identification of any identifier in this document can be implicit or explicit. Identifiers can be unique within a certain context, such as for a specific computer program or program provider.
[0160] As used herein, the term "memory" can refer to a data storage device used to store digital information, typically a hard disk, magnetic storage device, medium, portable computer disk or optical disk, flash memory, random access memory (RAM), etc. Additionally, memory can refer to the processor's internal register memory.
[0161] Also note that any enumeration terms, such as first device, second device, first surface, second surface, etc., should be considered non-restrictive, and such names do not imply a defined hierarchical relationship. In the absence of any explicit information to the contrary, enumeration naming should only be considered as a way to implement different names.
[0162] As used herein, the expression “configured to” can mean that the processing circuitry is configured or adapted to perform one or more of the actions described herein by means of software or hardware configuration.
[0163] As used herein, the term "number" or "value" can refer to any type of number, such as binary, real, imaginary, or rational numbers. Furthermore, a "number" or "value" can be one or more characters, such as letters or a string of letters. Additionally, a "number" or "value" can also be represented by a bit string.
[0164] As used herein, the expressions “may” and “in some embodiments” are generally used to indicate that the described features can be combined with any other embodiments disclosed herein.
[0165] In the accompanying drawings, features that may only exist in some embodiments are typically drawn using dotted or dashed lines.
[0166] When the words “include” or “contain” are used, they should be interpreted as non-restrictive, meaning “consisting of at least…”.
[0167] The embodiments described herein are not limited to those described above. Various alternatives, modifications, and equivalents may be used. Therefore, the above embodiments should not be considered as limiting the scope of this disclosure as defined by the appended claims.
Claims
1. A method for providing one or more annotated training images (336) for use in training of a trainable image content recognition algorithm of a machine vision system (205) which is operable to recognize content in images captured by the machine vision system (205) by means of the trainable image content recognition algorithm, wherein the method comprises: - acquiring (301) annotation information (337) for one or more intermediate images (338) to be captured by the machine vision system (205), the annotation information (337) being indicative of properties of content in the intermediate images (338), the content being content for recognition by the trainable image content recognition algorithm, and wherein the properties are such that, if the algorithm were to be trained on the intermediate images, the training would benefit from knowledge that the intermediate images contain content having the properties; - operating (302) the machine vision system (205) so that the one or more intermediate images (338) are captured by the machine vision system (205) in accordance with the acquired annotation information; - providing (303) one or more annotated intermediate images (339) corresponding to the one or more intermediate images (338) captured by the machine vision system (205) in accordance with the acquired annotation information (337) and the annotations, and - providing (306) the one or more annotated training images (336) based on at least one of the annotated intermediate images (339).
2. The method of claim 1, wherein the method further comprises: - applying (304) the trainable image content recognition algorithm having a first capability of content recognition on the one or more annotated intermediate images (339) to recognize content of the respective annotated intermediate image (339) and to check whether the recognition is in accordance with the annotation information (337) of the respective annotated intermediate image (339).
3. The method of claim 2, wherein in response to the algorithm having the first capability failing to recognize content of a respective annotated intermediate image (339b) in accordance with the annotation information (337) of the respective annotated intermediate image (339b), the respective annotated intermediate image (239) is provided as a respective annotated training image (336a) of the annotated training images (336).
4. The method of any of claims 2-3, wherein the method further comprises: - in response to the algorithm having the first capability failing to recognize content of the respective annotated intermediate image (339b) from annotation information (337) of the respective annotated intermediate image (339b), providing (305) failure information via a user interface (245) of the machine vision system (205), whereby a user (252) of the machine vision system (205) who controls what is being imaged is able to obtain feedback about the failure via the user interface (245).
5. The method of any one of claims 1-4, wherein the method further comprises: - storing (307) the respective annotated intermediate image (339b) provided as the respective annotated training image (336a) in a memory for later use; and - discarding (308) any remaining annotated intermediate image (339a) that is not provided as an annotated training image (336), whereby memory storage space can be saved.
6. The method of any one of claims 2-5, wherein the method further comprises: - using the respective annotated intermediate image (339b) provided as the respective annotated training image (336a) and its associated annotation information to train (309) the trainable image content recognition algorithm to achieve a second capability of content recognition.
7. The method of claim 6, wherein the training is performed in response to the algorithm having the first capability failing to recognize content of the respective annotated intermediate image (339b) from annotation information (337) of the respective annotated intermediate image (339b).
8. One or more devices (205; 230; 500) for providing one or more annotated training images (336) for use in training of a trainable image content recognition algorithm of a machine vision system (205), the machine vision system (205) being operable to recognize content in images captured by the machine vision system (205) by means of the trainable image content recognition algorithm, wherein the one or more devices are configured to: - obtain (301) annotation information (337) for one or more intermediate images (338) to be captured by the machine vision system (205), the annotation information (337) being indicative of an attribute of content in the intermediate images (338), the content being content for recognition by the trainable image content recognition algorithm, and wherein the attribute is such that, if the algorithm were to train on the intermediate images, the training would benefit from knowledge that the intermediate images contain content having the attribute; - operate (302) the machine vision system (205) such that the one or more intermediate images (338) are captured by the machine vision system (205) in accordance with the obtained annotation information; providing (303) one or more annotated intermediate images (339) corresponding to the one or more intermediate images (338) captured by the machine vision system (205) according to the acquired annotation information (337), and providing (306) the one or more annotated training images (336) based on at least one of the annotated intermediate images (339).
9. One or more computer programs (503) comprising instructions which, when executed by one or more processors, cause an apparatus (205; 230; 242; 500) according to claim 8 to carry out the method according to any one of claims 1-7.
10. One or more carriers comprising the one or more computer programs (503) according to claim 9, wherein the one or more carriers are one or more of the following: an electrical signal, an optical signal, a radio signal or a computer readable storage medium (601).
Citation Information
Patent Citations
System and method for applying deep learning tools to machine vision and interface for the same
US20220189185A1