Physical markers for labeling
The use of physical marker devices simplifies and cost-reduces the generation of training data for machine learning models by automating the labeling process, enhancing the efficiency and quality of visual feature classification on objects with changing appearances.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-08-08
- Publication Date
- 2026-03-24
AI Technical Summary
Existing methods for generating labeled training data for machine learning, particularly deep learning, are time-consuming, costly, and inefficient, especially for classifying small visual features on reflective or transparent objects, and require significant manual labeling effort.
A method using physical marker devices, such as AR or QR markers, to automatically generate training data by attaching them to local features of objects, acquiring images, and using mask information to train models for classification, reducing manual effort and enhancing data quality.
This approach enables rapid generation of high-quality training data with minimal human intervention, improving the efficiency and cost-effectiveness of training models for visual detection and classification tasks, especially on objects with changing appearances due to lighting or pose.
Smart Images

Figure 2026052665000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to supervised machine learning, particularly deep learning using training data, and the field of computer vision. In particular, the present invention relates to techniques for generating labeled training data for learning to classify local features of objects in image data. Specifically, the present invention relates to a method and system for generating classification information for classifying local features of objects in an image.
Background Art
[0002] Detecting, classifying, identifying, and tracking objects, and particularly physical features that are part of an object, such as features that locally exist on the surface of an object in an image, is a task in the field of computer vision that seeks solutions for a wide range of technical applications.
[0003] For example, U.S. Patent No. 10,713,769B2 discloses a method for active learning to train a defect classifier. This method involves acquiring an image, selecting data points in that image, acquiring labels for one or more selected data points, generating a set of labeled data containing the selected data points and data of the acquired labels, and training a defect classifier using the set of labeled data. Deep learning, as a prior art solution to supervised learning, requires the use of a sufficient amount of labeled data, for example, containing hundreds or millions of data samples in the training data, depending on the complexity of the application. Thus, supervised machine learning, such as deep learning, requires a large set of labeled training data to train a classifier model. Established approaches to generating labeled training data include a human expert visually inspecting pre-recorded images of physical images and manually adding labels to the inspected images, for example, by drawing boxes around the relevant image regions in the inspected images using a labeling tool. The human expert stores the labeled images as part of the training data set.
[0004] However, visually inspecting each image and manually adding labels to the inspected images is time-consuming, requires a significant amount of training for the labeling software used, and is therefore costly.
[0005] Furthermore, known approaches fail to provide the large sets of training data essential for machine learning models, which often require thousands or millions of individual samples of training data to train models that can classify and track parts of objects with high certainty. Large sets of training data are particularly essential in applications that require the classification of small visual features on the surface of reflective or transparent objects, or in applications where background and lighting conditions significantly influence the visual appearance.
[0006] In different application areas, U.S. Patent No. 10,169,678B1 relates to generating training data for training models, such as perceptual models for identifying objects or structures in an image, for example, for identifying regions of interest in an image. In particular, perceptual models enable processing image data acquired from a perceptual component that processes data generated by sensors. The perceptual component identifies, classifies, and / or tracks objects in the environment, and the perceptual function includes (1) segmentation, (2) classification, and (3) tracking over time and image frames, which are performed by a perceptual model trained by machine learning, in particular using real-world image data and / or image data generated in a simulated environment. Rendering virtual objects using simulation, in particular, 3D representations of objects under changing conditions, including different spatial viewpoints and changing lighting conditions, benefits from the rendering framework's nearly perfect knowledge of its internal state. For example, the rendering framework has ideal knowledge of the transformation between the virtual object and the camera. Thus, a single label for the 3D object is sufficient to predict the corresponding label on a newly rendered image. Therefore, this approach offers the ability to provide large amounts of labeled data without requiring excessive manual labeling. However, simulated reproducibility generally differs from the real world, and such simulated data is often found to be suitable only for augmenting labeled real-world data. Furthermore, this approach to generating training data using simulation requires elaborate representations (modeling) of objects, computationally complex rendering tools, and labeling of regions of interest in images, which are also time-consuming. Generating training data using simulation incurs significant costs for complex software tools and requires thorough user training.
[0007] In different technical fields relating to localization applications, European Patent Application Publication 4053801A1 discloses generating training data for training a model, which is a perceptual model for identifying, for example, objects or structures in an image, for identifying regions of interest in an image. European Patent Application Publication 4053801A1 discloses a method for generating detection information for detecting objects representing landmarks for localization and navigation purposes in an image by training a neural network. The method uses an image that includes an image region containing a desired landmark and an image region showing a labeled object located in a certain spatial relationship to the desired landmark in order to generate training data. The method aims to identify unique and immovable landmarks that enable robust localization of an autonomous agent, and uses an image region that represents an object that is already available to the detector. European Patent Application Publication 4053801A1 provides an efficient process in terms of processing time for labeling regions of interest in an image, but requires the presence of a labeled object near the landmark, thereby limiting its applicability. [Overview of the project] [Problems that the invention aims to solve]
[0008] The objective of the present invention is to improve the process for generating training data for training a model for visual detection and classification of objects, in terms of process simplification, data generation speed, training data quality, and cost-effectiveness. [Means for solving the problem]
[0009] A method for generating image data for generating training information according to independent claim 1, and a method for generating classification information for automated image analysis relating to local features of objects in an image, a system for generating image data for generating training information, and a system for generating classification information for automated image analysis relating to local features of objects in an image, according to the corresponding independent claim, provide advantageous solutions to the aforementioned problems.
[0010] The dependent claims define further advantageous embodiments.
[0011] In a first aspect of the present invention, a method for generating image data for generating training information for automated image analysis relating to local features in an image includes the steps of: applying at least one physical marker device adjacent to a local feature of an object; acquiring a plurality of images of the object using at least one camera sensor; and storing the plurality of images.
[0012] A second embodiment of the method for generating classification information for automated image analysis relating to local features of an object in an image includes the steps of: acquiring a plurality of images of an object; detecting at least one physical marker device applied adjacent to local features of an object in at least one of the acquired images; for each detected at least one physical marker device, calculating a region of interest in at least one image based on predetermined relative position information associated with the at least one physical marker device; generating mask information based on the calculated region of interest; storing the generated mask information associated with at least one image as training information; and generating classification information for automated image analysis relating to local features by training a model using the stored training information.
[0013] A method according to the first embodiment enables economically advantageous marking of multiple images stored in a non-volatile memory device, for example, the method forms a basis for generating training information to generate classification information for use in automated image analysis relating to local features by training a model using the stored training information.
[0014] Alternatively, features of the methods according to the first and second embodiments may be combined within the scope of the appended claims.
[0015] Local features can include local physical features. Objects may be physical objects.
[0016] The generated training information enables the training of a model for generating classification information for automated image analysis related to local features in images. The classification information may, in particular, be classification information for detecting, segmenting, classifying, identifying, and / or determining the regression of local features in images using the stored classification information.
[0017] Mask information is information that allows for the isolation of specific regions (areas) or object representations within an image. Generally, image masking is an approach in image processing and computer vision that allows for the masking of unwanted parts of an image, concentrating processing resources on the area of interest, thereby contributing to precise and accurate processing results with acceptable processing resource usage. Mask information can define a binary image containing pixels with zero and non-zero pixel values. When a mask is applied to a corresponding image of the same size, all pixels in the image corresponding to pixels in the mask with zero pixel values are set to zero. All other pixels in the image corresponding to pixels in the mask with non-zero pixel values remain unchanged.
[0018] The model is a machine learning model (ML model). A trained model enables the classifier to identify and label local object features of objects in an image.
[0019] The first aspect of the method provides labeled image data for generating a large set of training data in a short time and with only limited involvement of human experts. Thus, the method is useful for generating large sets of training data for machine learning models using supervised learning and deep learning approaches.
[0020] This method makes it possible to build high-quality machine learning models in a short amount of time at an acceptable cost.
[0021] The training effort required to work with the method according to the first embodiment in a particular technical domain is minimal for a human expert in that domain, thanks to the use of a physical marker device, and the use of a physical marker device in this method is intuitive, in contrast to virtual labeling tools typically employed to label image data.
[0022] This method has proven particularly useful when applied to objects with reflective or at least partially transparent surfaces. Objects can change their appearance significantly due to changes in lighting conditions, such as changes in the direction of background lighting, changes in the camera's pose when taking an image, and changes in the object's pose. Under these conditions, a large amount of training data is essential to train models used in the process of reliably and with high quality detecting, classifying, identifying, segmenting, and tracking small features, such as small surface features on the surface of an object.
[0023] Surface features on the surface of a physical object also include features visible on the surface of the physical object, but these are features within the body of the physical object, or, in particular, features arising from defects within the body. For example, a defect may be visible on the surface of the physical object, but it may also be an inclusion or bubble contained in the material of the body of the physical object, located some distance below the surface plane of the body of the physical object. This specifically relates to an object made from a transparent material, for example, a light cover made of cast resin. The light may be a flashing light for a vehicle in the automotive, marine, or aerospace industry. This method provides a useful and efficient approach to transferring implicit domain knowledge across different domains from humans, for example, human experts in a specific field of application, to automated systems, in particular, autonomously operating systems. This method requires only limited training overhead for knowledge transfer.
[0024] The present invention is useful in detection scenarios where general-purpose models are not available for use in classifiers. Therefore, this method is particularly useful in industrial application areas or personal learning settings where the overhead of labeling to generate training data for training a model is high and therefore expensive for the number of times the trained model is applied.
[0025] Using a physical marker device in the method according to the first embodiment is advantageous compared to the conventional approach of drawing image labels in the images using software tools when acquiring numerous images of an object together with the attached physical marker device. Preferably, for a single placement of a physical marker device on the surface of an object, numerous images with different image capture parameters, from various viewing angles under various lighting conditions and with different backgrounds, are captured and stored.
[0026] According to one embodiment, the method includes applying at least one physical marker device adjacent to a local feature on the surface of an object.
[0027] The physical marker device can include an augmented reality (AR) marker or a QR marker disposed on a carrier material.
[0028] The AR marker is an image or a small object integrated into a system for aligning or positioning an augmented reality object using the position of the AR marker in the real world. A QR marker (QR-based marker) is an image including a QR code (registered trademark), which is an evolved version of a two-dimensional barcode. The QR code can be used to encode information in a plurality of pixels arranged in the shape of a square grid. There are established computationally efficient solutions for detecting, identifying, and evaluating AR markers and QR-based markers in an image.
[0029] In one embodiment, the method includes a carrier material made of a flexible material, particularly including a sheet of paper or a plastic plate.
[0030] Thus, the physical marker device has a flexible or bendable structure. The physical marker device conforms to the curved, e.g., concave or convex shape of the surface of the object, thereby facilitating attaching the physical marker device to the surface. Further, the flat layout of the attached physical marker device increases the angular range for capturing an image of the physical marker device attached on the object, thereby increasing the amount of training information for one attached physical marker device on the object. The quality of the training data including all the training information increases, and thus the classification result provided by the trained model applied to the classifier is also improved.
[0031] In one embodiment of the method, at least one physical marker device has an annular structure.
[0032] Annular or ring-shaped structures enclose a region of interest. The user can immediately understand the region of interest when applied to any structure of physical marker devices that enclose it. The process of labeling the region of interest and its boundaries does not require extensive training for the user. The disadvantage of annular structures lies in the potential visual impact of the physical marker devices on the visual appearance of the region of interest, and therefore on the outcome of training.
[0033] The surface of at least one physical marker device has a specific color or a specific visual pattern, in particular a specific dot pattern.
[0034] Therefore, physical marker devices are easily detectable by computer vision in images that represent objects.
[0035] According to one embodiment of the method, at least one physical marker device includes fastening means for attaching the at least one physical marker device to the surface of an object.
[0036] Thus, the physical marker device remains attached to the object when multiple images are captured by at least one camera sensor. The fastening means can also ensure that the physical marker device conforms to the shape of the object's surface and does not protrude significantly beyond the surface.
[0037] Preferably, the physical marker device is detachable and, in particular, is only temporarily attached to the surface of the object. The physical marker device can be removed after multiple images of the object have been taken together with the physical marker device attached to the surface.
[0038] In this way, the object can be reused to learn other local features of the same or different classes using the same or a different physical marker device.
[0039] The fastening means includes at least one of an adhesive layer, a removable adhesive, a magnet, a suction cup, and a clip device.
[0040] The adhesive layer ensures that the flat, flexible physical marker device conforms perfectly to the shape of the object. The clip device has the specific effect of not leaving any traces that affect the visual appearance after the physical marker device is removed from the object, and therefore does not affect a further series of images of the object for generating training information.
[0041] The adhesive layer may include a removable adhesive.
[0042] The removable adhesive combines the effects of holding the physical marker device flat against the object's surface, while simultaneously leaving nothing to interfere with the visual appearance of the object's surface after removing the physical marker device and any remaining adhesive, thereby ensuring that the object is ready to be used to generate further training information.
[0043] The clipping device may include a spring-loaded clipping device.
[0044] A fastening means including a magnet, and at least one physical device including a ferromagnetic material, enable the physical marker device to be attached to the surface of an object. The physical marker device can be later removed without leaving any residual visible traces, thereby avoiding any undesirable visual effect on subsequent images of the object.
[0045] One embodiment of a physical marker device includes a plurality of physical marker devices arranged in a pattern (spatial pattern) on the surface of an object that defines a region of interest.
[0046] Therefore, users can define local features of the regular or irregular shape of an object. Additionally, users can define regions of interest that have at least one of a shape or size that is not predetermined.
[0047] Defining the area of interest can include enclosing (enclosing) the area of interest on the surface of an object.
[0048] A method according to one embodiment includes a plurality of physical marker devices arranged in a closed-loop pattern connected by expandable connections, which defines a region of interest surrounded by the closed-loop pattern.
[0049] In one embodiment of the method, at least one physical marker device includes a pattern of invisible ink, the invisible ink includes a UV photofluorescent material, an NIR reflective material, a material that reflects light of a predetermined polarization, or a material that reflects electromagnetic waves in a predetermined frequency band.
[0050] NIR reflective materials are materials that reflect light in the near-infrared (NIR) spectrum of the electromagnetic spectrum, adjacent to the frequencies where the visible light spectrum decreases (wavelengths where it increases). The NIR spectrum includes wavelengths from 780 nm to 3000 nm.
[0051] This physical marker device is an invisible ink-based physical marker device that is visible only under certain conditions, for example, only when illuminated with ultraviolet (UV) or infrared (IR) light. In the visible spectrum of light, and therefore under normal lighting conditions, the invisible ink-based physical marker device is invisible. Since labeling essentially involves scribbling on an area of interest on the surface of an object with a pen for invisible ink, using an invisible ink-based physical marker device significantly simplifies the process of applying the physical marker device. Subsequently, in a step of acquiring multiple images, a (first) image is acquired using a first camera sensor that records image data in the visible light spectrum. Furthermore, a (second) image associated with the first image is acquired using a second camera sensor that records image data with a specific filter adapted to receive light in the spectrum of the invisible ink, in order to acquire an image of the physical marker device containing the physical ink. In this way, the unchanging appearance of both the area of interest and the physical marker device can be acquired simultaneously by each pair of the two camera sensors and associated images. Alternatively, a single camera sensor may be used to sequentially acquire the first and second images by utilizing different and changing illumination conditions between the first and second images, for example, by quickly switching the illumination of an object between visible light and IR / UV light.
[0052] Embodiments using invisible ink are advantageous because they minimize the impact of the applied label, in this case a physical marker device, on the visual appearance of the object in the image. Thus, the trained model will reliably function on new images where the physical marker device is absent, avoiding a scenario where the model actually learns to detect the physical marker device instead of the region of interest while training the model with the training information.
[0053] According to one embodiment of the method, the physical marker device includes at least one body part of the user's body, in particular at least one finger or hand positioned in a particular gesture.
[0054] In this way, users can simply point to local features, which is intuitive and reduces the training effort required for users to apply the method.
[0055] A method according to one embodiment includes at least one physical marker device which is one of several types of physical marker devices which differ by size, and the size of the at least one physical marker device defines the size of the region of interest in at least one of several images.
[0056] Using different sizes of the same general-purpose physical marker device allows for associating each size of the marker device with a specific relative size and offset of the region of interest relative to the physical marker device in terms of relative position information. Selecting different sizes of the same type of physical marker device has the effect of defining the region of interest with different associated sizes. Thus, rescaling the physical marker device using a simple user selection results in an automatic and intuitive rescaling of the region of interest. For example, physical marker devices may be available for square-shaped regions of interest with lateral lengths of 0.01m, 0.02m, 0.05m, and 0.10m.
[0057] According to one embodiment, each of several types of physical marker devices is associated with (encoded) one of the following: positive classification of the region of interest, negative classification of the region of interest, and classification confidence of the user to whom at least one physical marker device is applied.
[0058] In this way, users can intuitively increase the diversity of the training information generated, thereby improving the quality of the trained models that result from the training, and consequently improving the performance of the classifiers to which the trained models are applied.
[0059] Classification confidence can include one of the following for the associated area of interest: a certain positive classification, a likely positive classification, or an 80% negative classification.
[0060] A method according to one embodiment includes at least one physical marker device, which includes shielding means for visually shielding at least a portion of the physical marker device.
[0061] The shielding means may include a movable flap, which is movable between a first position in which the movable flap shields at least a portion of the physical marker device and a second position in which at least a detectable portion of the physical marker device is fully visible in the captured image. Since detecting the physical marker device requires that at least a detectable portion of the physical marker device be visible in multiple images, unintended labeling of areas in multiple images is avoided during the process of applying at least one physical marker device to an object, and the quality of the generated training information is maintained.
[0062] During the step of applying the physical marker device, a corresponding effect can be achieved if the user obscures at least part of the physical marker device, for example, by using their hand or finger. Since detecting the physical marker device using known algorithms usually fails if even a small portion of the detectable part of the physical marker device is obscured, unintended labeling of areas in multiple images is avoided during the process of applying at least one physical marker device to an object.
[0063] A method according to one embodiment includes the step of applying at least one physical marker device adjacent to a local feature of an object at a predetermined in-plane angle with respect to the orientation of the image plane of at least one camera sensor, wherein the predetermined in-plane angle encodes the degree of membership of the local feature into a specific class.
[0064] The degree of membership quantifies the degree of membership of a local feature into a specific class. In the case of regression, the degree of membership has a scalar value.
[0065] In this way, the same physical marker device encodes additional information about local features. Users performing the labeling process can perform encoding in a highly intuitive manner without requiring extensive training to transfer their expertise to the classifier during training.
[0066] A method according to one embodiment includes the step of detecting a specific pattern in a plurality of images based on a pre-trained computer vision model of the at least one physical marker device, in the step of detecting at least one physical marker device.
[0067] In this way, physical marker devices can be reliably detected with minimal computational effort.
[0068] According to one embodiment, the method includes a plurality of application modes. When operating in the first application mode, at least one physical marker device is associated with the region of interest with positive examples of local features, and other regions of the image represent regions with negative examples of local features.
[0069] When operating in the second application mode, the sensor view of the camera sensor for acquiring multiple images is restricted to a view that includes only the region of interest in the multiple images, or it covers areas other than the region of interest, preventing the acquisition of image information from those areas.
[0070] Thus, the method avoids training the model to detect physical marker devices in an image, instead of training it to detect regions of interest in the image.
[0071] A method according to one embodiment includes the step of displaying online at least one region of interest associated with at least one detected physical marker device on the screen of a handheld device or a wearable augmented reality / virtual reality device while training a model. The method further includes the step of obtaining user input from a user via a human-machine interface, which includes classifications to associate with the displayed region of interest or to terminate processing when a predetermined classification quality is reached.
[0072] Therefore, the method is well-suited for guiding inexperienced users in creating training data for machine learning classification models without requiring extensive training.
[0073] Further embodiments of the method include the steps of validating a trained model by applying the trained model to stored training information, and determining whether a positive classification of a region of interest occurred outside the area in a plurality of images associated with a detected physical marker device. If it is determined that a positive classification of a region of interest occurred outside the area in a plurality of images associated with a detected physical marker device, the method includes the steps of communicating the determined positive classification of the region of interest that occurred outside the area in the images associated with the detected physical marker device to a user via a human-machine interface, performing a method for generating new training information, and generating new classification information by further training the model using the stored training information.
[0074] A method according to one embodiment includes the steps of generating new mask information based on at least one image containing a visually altered representation of at least one detected physical marker device, and storing the generated new mask information associated with the at least one image as further training information.
[0075] Thus, the additional step minimizes the impact of the applied physical marker device on the visual appearance of the object. Therefore, the undesirable effect of the trained classifier learning to classify based on the representation of the physical marker in the image instead of the region of interest is avoided by explicitly learning the invariance of the marker. Changing the image content of the new mask information results in a model that learns to be invariant to the visual appearance of the physical marker device. Visual changes can include changes from data augmentation, or replacing the image content of the new mask with different pixels in each iteration of the model's training.
[0076] In a third aspect of the present invention, a system for generating image data for generating training information for automated image analysis relating to local features in an image comprises at least one physical marker device applied adjacent to the local features of an object. The system further comprises at least one camera sensor configured to acquire multiple images of an object, and a memory configured to store the multiple images.
[0077] A fourth aspect of the present invention relates to a system for generating classification information for automated image analysis relating to local features of an object in an image. The system comprises a processor configured to acquire a plurality of images of an object. The processor is further configured to detect at least one physical marker device in at least one of the acquired plurality of images, calculate a region of interest in at least one image based on predetermined relative position information associated with the physical marker device for each detected physical marker device, generate mask information based on the calculated region of interest, store the generated mask information associated with at least one image in memory as training information, and train a model using the stored training information to generate classification information for automated image analysis relating to local features.
[0078] The systems according to the third and fourth embodiments achieve the advantageous effects discussed with reference to the methods of the first and second embodiments.
[0079] The following description of the embodiments will be accompanied by reference to the figures. [Brief explanation of the drawing]
[0080] [Figure 1] This figure presents a flowchart illustrating the process of generating training data according to one embodiment. [Figure 2] This figure provides a flowchart illustrating the steps of a method for generating training data according to one embodiment. [Figure 3] This figure illustrates a first example of a physical marker device in an application scenario. [Figure 4] This figure illustrates a second example of a physical marker device in an application scenario. [Figure 5] This figure illustrates an application example where the in-plane angle of the applied physical marker device is used to encode information about the degree of membership. [Figure 6] This figure illustrates a third example of a physical marker device in an application scenario. [Figure 7A] This figure illustrates a fourth example of a physical marker device in an application scenario. [Figure 7B] This figure illustrates a further variation of the fourth example regarding physical marker devices in application scenarios. [Figure 8] This figure illustrates a set of physical marker devices used in one embodiment. [Figure 9] This figure provides a flowchart illustrating the steps for applying a physical marker device to an object according to one embodiment, utilizing a set of physical marker devices. [Figure 10] This figure illustrates a fifth example of a physical marker device in an application scenario. [Figure 11] This figure illustrates a sixth example of a physical marker device in an application scenario. [Figure 12] This figure illustrates a seventh example of a physical marker device in an application scenario. [Figure 13] This figure illustrates an eighth example of a physical marker device in an application scenario. [Figure 14] This figure provides a simplified block diagram illustrating the architecture of one embodiment of the system. [Figure 15] This diagram shows an example of a hardware structure for system implementation. [Modes for carrying out the invention]
[0081] In the diagrams, corresponding elements share the same reference numeral. In the discussion of the diagrams, we avoid discussing the same reference numerals in different diagrams as much as possible, for the sake of brevity and without negatively impacting comprehension.
[0082] Figure 1 presents a flowchart illustrating a process for generating training data according to one embodiment, which is applied in a process for detecting local features on the surface of an object.
[0083] The application field of the embodiment generally includes the visual learning of local features 26 that form part of an object 23. The local features may be features on the surface of the object 23.
[0084] In specific application areas of quality control, the method enables training a classifier to detect local defects on the surface of object 23, including local features such as scratches, bubbles, pinholes, and heterogeneous color defects on object 23, during a monitoring or testing process.
[0085] In specific application areas of quality control, the method can generate training data for training a classifier to classify local features, for example, to classify local bubbles in foamed materials by size in a manufacturing process.
[0086] In certain application areas of visual inspection during monitoring or testing, the method allows for training a classifier to detect local surface features, such as local defects on the surface of object 23 due to corrosion or wear during operation of object 23.
[0087] In specific application areas of operating autonomous devices that utilize computer vision for control, the method allows for training a classifier to classify plants in a garden in order to distinguish which plants should be cut or eradicated, which plants should be fertilized or watered, and which plants should generally be avoided by the autonomous device.
[0088] Object 23 includes, for example, the transport vehicle body of a land, sea, air, or space transport vehicle. Uses may include inspection of a ship's hull, propeller, or rudder assembly.
[0089] Object 23 may include elements of a garden or agricultural environment. Each garden may include a limited number of plant species as characteristics of the garden, which must be learned individually for each garden.
[0090] The flowchart in Figure 1 illustrates three stages. Generally, the embodiments discussed are based on a specific application that classifies local features that are part of object 23 and located on, within, or at least near the surface of object 23.
[0091] In the first stage, the method generates training data for training a classifier to classify local features on the surface of object 23. Steps S1 and S2 form part of the process of generating training data for training a classifier to detect local features on the surface of object 23.
[0092] In the second stage, the method uses the generated training data to train a classifier that detects local features on the surface of an object. Steps S3, S4, and S5 form part of the process of training a classifier to classify local features on the surface of an object using the generated training data.
[0093] In the third stage, the method uses a trained classifier to classify local features on the surface of object 23 in the new image data. Step S6 in the flowchart of Figure 1 represents the third stage, in which the trained classifier is used to detect local features on the surface of the object in the new image data.
[0094] In step S1, the method begins by applying at least one physical marker device onto the surface 24 of object 23 adjacent to a local feature 26 that forms part of object 23.
[0095] A physical marker device may be attached to an object near a local feature. Alternatively, the physical marker device may be attached to an object so that at least a portion of its body surrounds the local feature.
[0096] A physical marker device is a marker detectable in an image using a known, available detector in a low-complexity detection process. Alternatively, or additionally, a pre-trained physical marker device is placed on a portion of the surface 24 of object 23, which will be learned as a local feature 26 of object 23.
[0097] In step S2, following step S1, the method continues to acquire multiple images of the surface 24 of object 23 using at least one camera sensor, and to store the multiple images in memory module 6.
[0098] After recording that multiple images have been acquired and stored, the user can remove the physical marker device from the surface 24 of object 23.
[0099] The method then detects at least one physical marker device in at least one of the multiple images acquired of object 23.
[0100] The method can detect physical marker devices, in particular the presentation of physical marker devices, in multiple images using available or pre-trained detection software.
[0101] For each detected physical marker device, the method then calculates a region of interest 22, 44 containing a representation of a local feature in at least one image, based on predetermined relative position information associated with the physical marker device.
[0102] The region of interest is used to generate mask information that defines a labeled training mask that can be used directly to train the model. Alternatively, or additionally, the mask information is stored as an image in an image data format for later training using the stored mask information.
[0103] The mask information is then associated with at least one image and stored in memory module 7 as training information.
[0104] The process in step S2 will be discussed in more detail with reference to the flowchart in Figure 2.
[0105] In step S2.1, following step S2, the method determines whether to finish generating training information. If it determines that further training information should be generated and the process of generating training information should not be finished yet (NO), the method returns to step S1 and executes steps S1 and S2 again. The method collects further instances of training information before starting to train the model. If it determines in step S2.1 that the process of generating training information in steps S1 and S2 should be finished (YES), the method proceeds from step S2.1 to step S3.
[0106] In step S3, the model is trained using the training information generated in step S2. In particular, the method then generates classification information for classifying local features in an image by training the model using the stored training information.
[0107] Alternatively, in step S3, an already trained model (a pre-trained model) may be refined (retrained) using the stored training information generated in step S2.
[0108] In step S4, the method determines a predetermined quality metric based on the training process. The quality metric may be determined using a predetermined set of validation information. The quality metric may also be determined using statistics generated during the model training process, such as loss.
[0109] In step S5, the determined quality metric is compared to a predetermined threshold (quality threshold). If it is determined that the determined quality metric exceeds the predetermined threshold, the method proceeds to a third stage (application stage), which includes step S6.
[0110] Alternatively, or additionally, the method could output information to the user prompting them to decide whether to continue adding data or to use the trained model in subsequent application phases.
[0111] If the user decides to add further training information based on the outputted information, the user can repeat the method in a new iteration of the sequence of steps S1, S2, S3, S4, and S5 by placing the physical marker device, or, for example, a different type of physical marker device, on a new instance of object 23, i.e., on a different local part of the same object 23 containing the features of local interest.
[0112] In step S6, the method acquires a new image and continues by using the trained model to detect regions of interest 22, 44 in the new image.
[0113] In step S5, if the method determines that the determined quality metric is smaller than a predetermined threshold, the method returns to step S1. In subsequent iterations of the first and second stages, the method prompts the user to attach a physical marker device to another representation of a local feature on the surface of object 23.
[0114] Figure 2 provides a flowchart illustrating the steps of a method for generating training data according to one embodiment. Figure 2 illustrates, in particular, step S2 of the first stage, namely the process of generating training information for training a classifier to detect local features on the surface of object 23, in more detail than in Figure 1.
[0115] A method for generating classification information for classifying object representations of local features in an image of object 23 begins in step S1 by applying at least one physical marker device to the surface 24 of the physical object 23.
[0116] The generated detection information includes a machine learning model (ML model) for a classifier that enables the identification, segmentation, or tracking of local visual features on the surface 24 of object 23 across a sequence of images of object 23.
[0117] In step S21, the method acquires multiple images of the surface 24 of the object 23 using at least one camera sensor. The method stores the acquired multiple images in the memory module 6.
[0118] In step S22, the method performs a detection process to detect at least one physical marker device, in particular a visual representation of a physical marker device, in at least one of the multiple images acquired and stored in the memory module 6.
[0119] In step S23, the method determines (calculates) for each at least one physical marker device detected in at least one of the images a region of interest 22, 44 containing a representation of local object features in at least one of the images, based on predetermined relative position information associated with the detected at least one physical marker device.
[0120] The predetermined positional information may include offset information, shape information, and size information of the region of interest relative to the position of the physical marker device on the surface 24 of the object 23.
[0121] The predetermined location information may include information on whether the regions of interest 22 and 44 are positive regions of interest 22 and 44 or negative regions of interest 22 and 44. Positive regions of interest 22 and 44 are areas on the surface 24 of object 23 where instances of local features exist. Negative regions of interest 22 and 44 are areas on the surface 24 of object 23 where instances of local features 26 do not exist.
[0122] The predetermined position information is stored in memory and can be retrieved from memory for the calculation in step S23.
[0123] The predetermined location information may be stored in software code, and therefore may be static, and may be associated with each distinguishable type of detectable physical marker device.
[0124] Predetermined location information may be modified by the user using a user interface and stored in association with each type of detectable physical marker device.
[0125] In step S24, the method continues to generate mask information based on the calculated regions of interest 22, 44.
[0126] In step S25, the method stores the generated mask information associated with at least one image as training information. The method stores the generated training information in memory module 6.
[0127] After executing step S25, the method determines in step S2.1 whether further training information should be generated in a further processing cycle that executes steps S1, S21 to S2.1, or whether the first stage should be terminated. If it is determined that the generation of training information should be terminated (YES), the method proceeds to the second stage, and in step S3, generates detection information for detecting local features 26 (representations of local features) in an image by training a model using the training information stored in the memory module 6.
[0128] The model may be an image classification model or an image segmentation model.
[0129] Training a model in step S3 involves either training a new model directly using the generated mask and associated images, or retraining (refining) an existing, pre-trained model.
[0130] Figures 3 to 11 illustrate in more detail examples of physical marker devices in the methods and systems according to embodiments, and their use. In general, the methods are suitable for use with many different types and designs of physical marker devices. The figures illustrate several specific examples of physical marker devices. Before discussing specific embodiments of the examples illustrated with reference to the figures, some common embodiments relating to physical marker devices will be discussed.
[0131] The physical marker device may be a known AR marker or QR marker, as known in the field of computer vision. Detecting the AR marker or QR marker may be implemented using existing software solutions available in the field of computer vision.
[0132] Physical marker devices in the form of AR markers or QR markers can be printed on paper carrier materials, plastic carrier plates, or any other printable carrier material.
[0133] The physical marker device may include a single physical marker device corresponding to a single region of interest 22, 44, or an arrangement of multiple physical marker devices corresponding to a single region of interest 22, 44.
[0134] A physical marker device may have a specific visual appearance, such as a particular design or external appearance, such as a visual pattern detectable by a computer vision model. The pattern of the physical marker device is preferably learned before the implementation of a method for generating detection information begins. A simple example of such a pattern is designing an annular-shaped physical marker device with a specific color on its surface, facing away from the surface 24 of object 23.
[0135] Alternatively, or additionally, the surface of the physical marker device facing away from the surface 24 of object 23 may be designed with an easily detectable dot pattern, which is specific to the use scenario in which images for generating training information are acquired.
[0136] A physical marker device may be configured to be flexible and bendable, made of a flexible material (non-rigid material), or have a flexible layout that allows for the deformation of each physical marker device. In one embodiment, a paper or plastic carrier material can provide a certain degree of flexibility. Alternatively, or additionally, a physical marker device may include a polygonal ring of a base physical marker device connected via flexible means, for example, a cord, wire, or a retractable rod-shaped connector in a closed-loop arrangement. A flexible physical marker device offers the user the advantageous property of being able to conform to the shape of a particular surface on the surface of object 23. A flexible physical marker device is therefore a versatile tool for users tackling multiple specific labeling scenarios using a single type of physical marker device. Figure 11 shows a specific example of a flexible physical marker device.
[0137] A physical marker device may be applied to the surface 24 of object 23 using invisible ink and its respective ink pen.
[0138] The invisible ink may have UV light fluorescence or NIR detection capability, meaning that the invisible ink fluoresces or stands out against the contrast in the respective electromagnetic spectra of UV light or NIR light. In the visible light electromagnetic spectrum, the invisible ink is transparent or colored depending on the surface 24 of object 23.
[0139] A physical marker device may be applied to the surface 24 of object 23 using a specific ink and its respective ink pen. The specific ink is visible only in a narrow band of the electromagnetic spectrum when irradiated by a laser in each band of the electromagnetic spectrum, for example.
[0140] Physical marker devices can include, for example, carrier materials in the form of plates or polarized, distinguishable inks.
[0141] Figures 3, 4, 5, and 11 illustrate embodiments of mounting a physical marker device to the surface 24 of an object 23.
[0142] The user can temporarily hold the physical marker device in its position on the surface 24.
[0143] An alternative and advantageous solution includes a fastening means for temporarily securing a physical marker device on the surface 24 of object 23.
[0144] When a physical marker device is designed in the form of an adhesive sticker, the means of fastening can be provided as an integrated feature. A removable adhesive on the surface of the physical marker device facing the surface 24 of object 23 allows the physical marker device to be detachably attached to the surface 24. This is particularly advantageous in combination with a flexible carrier material or flexible carrier plate for the physical marker device, as the adhesive area conforms to the concave or convex shape of the surface 24 of object 23.
[0145] The marker device may provide fastening means in the form of a spring-loaded clip device, rivet, bracket, or clamp, for mechanically securing the physical marker device to the surface 24 of the object 23. The choice of fastening means may depend on the material of the object 23, so that, ideally, no visible trace is left on the surface 24 after the physical marker device is removed from the object 23. For example, fastening means in the form of rivets may be suitable for labeling specific species of plants as local features in a garden environment, but may be unsuitable for labeling local features on the surface 24 of a transport vehicle body.
[0146] In an alternative application scenario, the fastening means may include at least one magnet 48.11 for fixing the physical marker device onto the surface 24 of the object 23 having ferromagnetic properties.
[0147] Combined with the processing proposed in embodiments for generating training information, the discussed physical marker device provides optimal results for objects 23 having locally flat surfaces containing local features or object portions to be classified. When the surface containing local features 26 or object portions extends significantly into three dimensions from the flat surface 24 of the physical object, and therefore deviates from a locally flat structure, the accuracy of labeling by the physical marker device may decrease due to the effects of perspective drawing, particularly for images captured at low camera angles relative to the surface 24 of the object 23.
[0148] This problem can be overcome by utilizing 3D processing, for example, by using a 3D camera sensor such as an RGBD camera and a physical marker device such as an AR marker that defines a volume of interest in 3D relative to the AR marker. The volume of interest may include, for example, a cube with predetermined side lengths determined by a predetermined distance from the physical marker device. The cubic volume of interest may have its base on the flat plane of the physical marker device. The predetermined side lengths and predetermined distances of the cubic volume of interest may each be 2 cm in length.
[0149] The following diagram provides a specific example of a physical marker device.
[0150] Figure 3 illustrates a first example of a physical marker device 21.x in an application scenario.
[0151] The arrangement of multiple physical marker devices 21 corresponds to a single region of interest 22. Multiple physical marker devices 21.x (where x=1,...,8), each having a unique visual appearance, are positioned in a rectangular, and in particular, square-shaped, arrangement 21 that encloses the region of interest 22 on the surface 24 of the object 23.
[0152] The arrangement of multiple physical marker devices 21 includes single physical marker devices 21.1, 21.3, 21.5, and 21.7 positioned at each corner of a square-shaped arrangement.
[0153] The arrangement of multiple physical marker devices 21 further includes single physical marker devices 21.2, 21.4, 21.6, and 21.8 positioned at the center of each side of a square-shaped arrangement.
[0154] Each of the multiple physical marker devices 21.x (where x=1,...,8) has a unique optical appearance that can encode the relative position of each physical marker device 21.x (where x=1,...,8) with respect to the region of interest 22.
[0155] The arrangement of multiple physical marker devices 21 improves the probability of detecting a region of interest 22 due to the redundancy of eight physical marker devices 21.x (where x=1,...,8) associated with a single region of interest 22 in a predefined spatial manner, thereby also increasing robustness against visual occlusion in the acquired image.
[0156] In the example shown in Figure 3, the method generates classification information for classifying object representations corresponding to color defects or scratches on the surface 24 of the transport vehicle, which corresponds to object 23 in the image of the transport vehicle itself, as shown in the right portion of Figure 3.
[0157] Figure 4 illustrates a second example of a physical marker device 25 attached to the surface 24 of object 23 in the application scenario shown in the right portion of Figure 4.
[0158] The physical marker device 25 includes a rectangular carrier plate made of a flexible material such as a thin plastic sheet. The physical marker device 25 includes a window portion 25.1 and a marker portion 25.2.
[0159] The window portion 25.1 includes an opening (window) that allows an image of the region of interest 22 on the surface 24 of the object 23 to be captured when the physical marker device 25 is applied to the surface 24. The window portion 25.1 may include a frame surrounding the region of interest 22. Supported by the window portion 25.1 as a targeting aid for correctly applying the physical marker device 25 to the surface 24 for local features 26, the user can intuitively label the region of interest 22.
[0160] A physical marker device 25 attached to the surface 24 of the transport vehicle body as a physical object 23 is shown in the right portion of Figure 4. The region of interest 22 labeled by the physical marker device 25 includes local defects (coating defects 26) as local features 26 in the color coating of the transport vehicle body. The method is well suited for generating training information to train a classifier to detect coating defects such as orange peel, stringing, and other undesirable film characteristics on the bright, often convex surface 24 of the transport vehicle body.
[0161] The marker portion 25.2 of the physical marker device 25 includes a QR code and at least one optical indicator, for example, two arrows in the example of Figure 4, thereby providing further support to the user when applying the physical marker device 25 to the surface 24 for local features 26.
[0162] The marker portion 25.2 corresponds to the detectable portion of the physical marker device 25 that System 1 detects in one of several images. During the labeling process, in order to achieve sufficient quality of the training information generated and the model that will ultimately be trained for use in the classifier, it is necessary to avoid unintended labeling of areas in the image that are not intended to represent the region of interest 22. Unintended labeling of areas in the image may occur during the labeling process when the image has already been taken, but the user is still in the process of placing the physical marker device on the object 23.
[0163] For example, by using a trigger signal provided to the system when the user has finished the process of applying a physical marker device to object 23, such as by operating a button on a human-machine interface, an explicit and unambiguous separation between steps S1 and S21 of the method can be ensured. Thus, the method avoids unintended labeling of areas in the image.
[0164] Alternatively, or additionally, system 1 may require the user to explicitly start and stop step S21, which is to acquire multiple images or to record a video containing a sequence of images corresponding to multiple images.
[0165] Alternatively, or additionally, the physical marker device may include shielding means not explicitly shown in Figure 4. The shielding means visually shields the physical marker device, in particular, at least partially, the marker portion 25.2 of the physical marker device 25.
[0166] The shielding means includes, for example, a movable flap, which is movable between a first position in which the movable flap shields at least a portion of the physical marker device, in particular the marker portion 25.2 of the physical marker device 25, and a second position. In the second position, the detectable portion of the physical marker device 25 is fully visible in the captured image. Detecting the physical marker device requires that at least the detectable portion (marker portion 25.2) of the physical marker device 25 is fully visible in multiple images. Thus, unintended labeling of areas in multiple images is avoided during the process of applying at least one physical marker device 25 to the object 23, and the quality of the generated training information is maintained.
[0167] As an alternative to the movable flap, a slider covering at least a portion of the marker portion 25.2 may be used.
[0168] The user performing the labeling in step S1 can manually operate the slider or movable flap.
[0169] Alternatively, the marker device 25 may include a spring-loaded button configured to automatically release the cover when the physical marker device 25 is applied to the object 23, or in response to user interaction with the marker portion 25.2.
[0170] The QR code on the marker portion 25.2 of the physical marker device 25 may include encoded information about the marker, such as whether the physical marker device 25 is a positive marker or a negative marker.
[0171] Figure 5 illustrates an application that allows for the evaluation of the in-plane angle of the applied physical marker device 25 in order to obtain encoded class information as an additional input.
[0172] The embodiment in the scenario of Figure 5 includes applying at least one physical marker device 25 adjacent to a local feature 26 that forms part of an object 23 at a predetermined in-plane angle with respect to the orientation of the image plane of at least one camera sensor, sensor 1, and sensor 2. The predetermined in-plane angle encodes the degree of membership of the local feature 26 into a specific class, based on the user's expert knowledge. The degree of membership quantifies the degree of membership of the local feature 26 into a specific class.
[0173] In the case of regression, the degree of membership has a scalar value.
[0174] For example, in the example discussed earlier with reference to Figure 4, the physical marker device 25 indicated whether the region of interest 22 belonged to a particular class. The region of interest 22 is labeled as one positive member of the binary membership for classification. To label negative observations (cases) of local features 26, evaluation of the training information assumes that all unlabeled areas represent negative observations. Alternatively, an additional distinguishable physical marker device 25 is required for explicit labeling of negative observations of local features 26.
[0175] In the embodiment shown in Figure 5, the same physical marker device 25 encodes additional information about local features 26. Users performing the labeling process can perform encoding in a highly intuitive manner without requiring extensive training to transfer their expertise to the classifier during training.
[0176] In an example of coding for the degree of membership in a particular class, 100% (full) membership of a local feature 26 is coded into a particular class by applying a physical marker device 25 in an orientation corresponding to (equal to) the orientation of the image plane of at least one camera sensor, resulting in an in-plane angle of 0° (0 degrees). The left sub-diagram of Figure 5 illustrates this specific example.
[0177] A 45° interior angle encodes 75% of the membership. The central sub-diagram in Figure 5 illustrates a concrete example of this.
[0178] A 90° interior angle encodes 50% membership. The right-hand sub-diagram of Figure 5 illustrates a concrete example of this.
[0179] Although not illustrated in Figure 5, an interior angle of 135° encodes 25% membership, and an interior angle of 180° encodes 0% membership. A degree of 0% membership corresponds to a negative example (observation, case) of membership in a particular class.
[0180] Additionally, System 1 may include a human-machine interface that outputs at least one of the in-plane angle and degree of membership to the user during the ongoing labeling process in order to support the user when applying the physical marker device 25 to the object 23.
[0181] A human-machine interface can use a display device, such as a monitor or an augmented reality (AR) headset, to provide at least one of the in-plane angle and degree of membership in order to support the user in the form of easily understandable feedback without requiring a significant amount of prior training for the user.
[0182] Figure 6 illustrates a third example of a physical marker device 27 in an application scenario.
[0183] The physical marker device 27 has a rectangular layout formed by a frame structure 27.1 made of a flexible carrier material surrounding the opening (window), which defines a region of interest 22 of each rectangular shape. As shown in the example in Figure 4, the physical layout of the physical marker device 27, in particular the shape of the frame structure 27.1, encodes predetermined relative position information, which enables the determination (calculation) of the region of interest 22 in the image when the physical marker device 27 is detected. As in the example in Figure 4, the region of interest 22 in the right portion of Figure 6 includes a localized defect 26 in the color coating of the transport vehicle body, such as a scratch.
[0184] Figure 7A illustrates a fourth example of a physical marker device in an application scenario.
[0185] Figure 7A shows a rectangular region of interest 22 with a width corresponding to the diameter of the finger 28, where the region of interest 22 is located, for example, at a distance of approximately half the width of the finger 28, approximately the extension of the length of the finger 28.
[0186] A physical marker device in one embodiment includes at least one finger 28 of the user. In the example in Figure 7A, the physical marker device is the user's finger. Predetermined relative positional information associated with the at least one physical marker device implemented by the finger 28 can define a region of interest 22 as an area, with the size of the fingernail of the finger 28 in at least one image. The region of interest 22 may be positioned in an image that extends in the direction the finger is pointing, at a predetermined distance from the finger 28, with a predetermined shape of the region of interest 28.
[0187] Figure 7B illustrates a further variation of the fourth example regarding physical marker devices in application scenarios.
[0188] In the example in Figure 7B, the physical marker device corresponds to the user's hand 29, and in particular, the fingers 28 (index finger) and thumb 30 positioned in a specific gesture performed by the user to label local features 26. In Figure 7B, the specific gesture performed by the user surrounds the region of interest 22 by positioning the thumb 30 and fingers 28 so that their respective fingertips touch each other, thereby forming a ring-like structure surrounding the region of interest 22.
[0189] Both variations of the physical marker device in Figures 7A and 7B provide a specific, intuitive method for the labeling process that does not require specific, thorough training for the user to successfully apply the physical marker device to object 23.
[0190] Figure 8 illustrates a set of physical marker devices 31 used in one embodiment.
[0191] A set of physical marker devices 31 includes multiple physical marker devices 31.x (where x = 1, ..., 6).
[0192] The set of physical marker devices 31 includes one basic type of physical marker device in two subtypes. The first subtype includes physical marker devices 31.1, 31.3, and 31.5 for positive regions of interest 22 and 24. The second subtype includes physical marker devices 31.2, 31.4, and 31.6 for negative regions of interest 22 and 24.
[0193] The set of physical marker devices 31 includes three different sizes for each of the two subtypes of the physical marker device, for the positive regions of interest 22, 44 and the negative regions of interest 22, 44, respectively.
[0194] Different types of physical marker devices can encode different classes of regions of interest 22, 44 and local features 26 on the surface 24 of object 23.
[0195] For example, classes encoded by different types of marker devices may include a subset of certain positive rankings, possibly positive rankings, and 80% negative rankings associated with regions of interest 22, 44, indicated by the associated type of physical marker device. Each ranking depends on the assessment of the user performing the labeling, resulting in each user's classification confidence, although thorough knowledge of the labeling tool is not required when implementing and working with the method.
[0196] The set of physical marker devices 31 includes three further groups of physical marker devices, which differ depending on the size of the region of interest.
[0197] The first group of small-sized physical marker devices includes physical marker devices 31.1 and 31.2. The second group of medium-sized physical marker devices includes physical marker devices 31.3 and 31.4. The third group of large-sized physical marker devices includes physical marker devices 31.5 and 31.6.
[0198] The set of physical marker devices shown in Figure 8 provides the user with six physical marker devices for labeling regions of interest on the surface of object 23.
[0199] The set of physical marker devices 31 shown in Figure 8 allows the user to select the most suitable size and correct subgroup of physical marker devices 31.x (where x=1,··,6) from the set of physical marker devices 31 to label a specific region of interest 22 on the surface 24 of an object 23.
[0200] In one embodiment, a specific region of interest 22 is defined in terms of the associated relative size of a physical marker device compared to the sizes of other physical marker devices in a set of physical marker devices 31. This provides the effect of automatic scaling of the region of interest 22, 44 by selecting a specific physical marker device from the set of physical marker devices 31. For example, selecting a physical marker device that is of the same type as another physical marker device and has a relative size of 200% of the selected physical marker device compared to another physical marker device defines a region of interest 22, 44 associated with a size of 200% of the size of the other physical marker devices. In the step of applying physical marker devices, explicit programming by the user performing the labeling process is not required, resulting in an intuitive process that does not require extensive training for the user.
[0201] Figure 9 provides a flowchart illustrating the steps for applying a physical marker device to object 23 according to one embodiment using a set of physical marker devices.
[0202] Steps S11 and S12 represent substeps of step S1 to the physical marker device object 23 during the first stage of generating training information.
[0203] In step S11, the method includes selecting at least one physical marker device from a plurality of available physical marker devices based on the size and type of the region of interest 22, 44 on the surface 24 of the object 23.
[0204] Multiple physical marker devices may include, for example, sets of physical marker devices of different sizes for each of the positive regions of interest 22, 44 and negative regions of interest 22, 44 shown in Figure 8.
[0205] In step S12, the method proceeds by applying at least one selected physical marker device to a suitable position relative to the locations of interest 22, 44 on the surface 24 of the object 23.
[0206] Figure 10 illustrates a fifth example of a physical marker device in an application scenario.
[0207] The physical marker device 32 in Figure 10 is an example that uses invisible ink. To apply the physical marker device 32 to the surface 24 of the object 23, the user uses a pen to label local features 26 on the surface 24 of the object 23. The local features 26 may be local defects, such as scratches on the surface 24 of the object 23.
[0208] The left portion of Figure 10 illustrates a first image in the visible spectrum of light, using an image taken with a conventional camera sensor.
[0209] The right portion of Figure 10 illustrates a second image taken by an NIR camera sensor illustrating a second image in the NIR spectrum of light, the second image including a physical marker device 32 in the form of an anomalous ink trace 32 of invisible ink applied by a user on a surface 24 using an ink pen at the location of local features 26.
[0210] Step S23, which calculates the region of interest 22, may include a detected physical marker device, and the region of interest 22 includes local features 26 in at least one first image based on predetermined relative position information associated with at least one physical marker device. The predetermined position information in the example of Figure 10 may include calculating a frame 33, for example, a rectangular frame surrounding an anomalous ink trace 32 applied by the user as close as possible to the location of the local object feature 26. In step S24, the method generates mask information based on the calculated region of interest 22, as shown by frame 33 in the right image of Figure 10, and stores the generated mask information associated with at least one image corresponding to the left image of Figure 10 as training information.
[0211] Figure 11 illustrates a sixth example of a physical marker device in an application scenario. The application scenario in Figure 11 differs from the example in Figure 10 only in that, due to the transparent housing of the printed circuit board, the object 23 has a locally flat surface 24; however, the region of interest 44 shown in the right-hand panel of Figure 10 has a sharp change in the direction perpendicular to the image plane of the camera sensor 58.
[0212] Object 23 in Figure 11 is a printed circuit board equipped with multiple electrical circuit elements arranged inside a transparent enclosure.
[0213] The physical marker device in Figure 11 is an area of invisible ink applied by the user on the surface of a printed circuit board, as shown in Figure 10.
[0214] The left panel of Figure 11 includes an image 40 captured while illuminating a printed circuit board with only light in the visible portion of the electromagnetic spectrum. A physical marker device 43 applied to the printed circuit board using invisible ink is not visible in image 40 of the left panel of Figure 11.
[0215] The central image in Figure 11 includes images 41 captured while illuminating a printed circuit board with additional illumination of light in the visible portion of the electromagnetic spectrum (visible light) and light in the UV portion of the electromagnetic spectrum (UV light). A physical marker device 43 applied to the printed circuit board using invisible ink that reflects light in the UV portion of the electromagnetic spectrum is clearly visible in the central image of Figure 11.
[0216] The right-hand figure of Figure 11 illustrates the result of calculating a region of interest 44 based on at least one detected physical marker device 43, where the region of interest 44 includes a representation of local object features in the image 42 based on predetermined relative position information associated with at least one physical marker device 43.
[0217] The left and center figures of Figure 11 illustrate switching the light source illuminating object 23. In the left figure, the first light source emits visible light, in which case object 23 appears unchanged in the acquired image. In the center figure of Figure 11, the second light source emits UV light, enabling the acquisition of an image for detecting a physical marker device 43. For example, UV light makes the physical marker device 43, which includes UV fluorescent paint, visible in the image acquired by a second camera sensor adapted to capture UV images. Thus, in the embodiment of Figure 11, two images captured sequentially or simultaneously are required to be recorded.
[0218] To record continuously, the illumination of object 23 between acquiring two consecutive images requires synchronizing the exposure times of the two images from two camera sensors, including the first camera sensor and the second camera sensor. Synchronization for time multiplexing of image capture works perfectly when the camera sensors and object 23 are stationary, and therefore not moving in space.
[0219] When using a camera sensor with a high recording frame rate, the camera sensor and object 23 do not need to be completely still; even if the object 23 and camera sensor move slowly, sufficient results can be produced to generate training information. Therefore, cost-effective solutions are available to minimize the visual impact of labeling by the physical marker device 43 on object 23. Thus, the trained model is guaranteed to achieve good classification results on new images where the physical marker device 43 is not present, and avoids a scenario where the model learns to detect the physical marker device 43 instead of the region of interest 44 while being trained using the training information.
[0220] Figure 12 illustrates a seventh example of a physical marker device 45 having an annular body.
[0221] The physical marker device 45 has a ring-like structure in which the annular body 45.1 of the physical marker device 45 encloses the opening 46. The size of the opening 46 defines the location and size of the region of interest 22 when the physical marker device 45 is attached to the surface 24 of the object 23.
[0222] To fix the physical marker device 45 to the surface 24, the annular body 45.1 of the physical marker device 45 may have at least one level surface 47 that provides a flat surface for applying an adhesive to bond the physical marker device 45 to the surface 24 of the object 23.
[0223] The physical marker device 45 defines the region of interest 22, which has a circular shape surrounding the region of interest 22.
[0224] Alternatively, the body of the physical marker device may have a different shape, such as an elliptical, rectangular, or polygonal closed shape, instead of the annular shape of the annular body 45.1 in Figure 12.
[0225] The annular body 45.1 of the physical marker device 45 may be made of a flexible material, thereby allowing the user to adapt the physical marker device 45 to a surface 24 that deviates from a planar surface.
[0226] An example of an annular physical marker device 45 made of a flexible material, 45.1, allows the user to adapt the physical marker device 45 to define a region of interest that is not perfectly circular.
[0227] Figure 13 illustrates an eighth example of a physical marker device 48.
[0228] The physical marker device 48 includes a plurality of basic physical marker devices 48.1, which are connected between each pair of basic physical marker devices 48.1 by telescopic connectors 48.2. The six basic physical marker devices 48.1 and the telescopic connectors 48.2 form a closed structure surrounding area 49. Area 49 corresponds to a region of interest 22 when the physical marker device 48 is attached to the surface 24 of object 23.
[0229] The retractable connection 48.2 of the physical marker device 48 shown in Figure 13 includes multiple connection elements 48.21, 48.22, and 48.23, which allow the user to change the length of the retractable connection 48.2, as indicated by the arrows in partial view A.
[0230] By changing the length of the telescopic connection 48.2 between adjacent pairs of basic physical marker devices 48.1, the user can change the distance between the basic physical marker devices 48.1. This has the effect of changing the shape of the arrangement of multiple basic physical marker devices 48.1, as well as the shape and size of the area 49 that defines the region of interest 22 on the surface 24 of the object 23.
[0231] Each of the basic physical marker devices 48.1 in Figure 13 includes a magnetic layer 48.11 as an example of fastening means for fixing the basic physical marker device 48.1 to a metal surface.
[0232] Figure 14 provides a simplified block diagram illustrating the architecture of System 1 for generating detection information to detect object representations in an image according to one embodiment.
[0233] System 1 for generating detection information to detect object representations in an image comprises at least one physical marker device applied to the surface 24 of object 23.
[0234] System 1 includes at least one camera sensor configured to acquire multiple images of the surface 24 of object 23. The system in Figure 14 includes a first camera sensor (sensor 1) and a second camera sensor (sensor 2).
[0235] The first camera sensor captures an image of object 23 in the visual spectrum.
[0236] A second camera sensor captures an image of object 23 in a spectrum invisible to a human observer. The second camera sensor may be used in embodiments that utilize a physical marker device that is not visible in the electromagnetic spectrum of visible light.
[0237] System 1 further includes data storage (memory) configured to store information and data. For example, memory module 6 of System 1 stores multiple images acquired by perception module 2 of System 1. Image processing module 3 of System 1 retrieves the images acquired by perception module 2 and stored in memory module 6 and performs image preprocessing.
[0238] Next, the ROI determination module 4 of system 1 detects at least one physical marker device in at least one of the multiple images that have been acquired and preprocessed. The ROI determination module 4 then calculates a region of interest (ROI) for each detected physical marker device, which contains a representation of local object features in the image, based on predetermined relative position information associated with the detected physical marker device. The ROI determination module 4 then generates mask information based on the calculated region of interest.
[0239] Next, the training data generation module 5 of system 1 generates training information, which includes generated mask information associated with at least one image, and stores the generated training information in the memory module 6.
[0240] Next, the training module 7 generates classification information for classifying local features 26 in a new image by training a classification model using the training information stored in the memory module 6. The training module 7 then stores the trained classification model in the classification model memory 8.
[0241] The classification module 9 uses the trained classification model stored in the classification model memory 8 to detect, classify, segment, or track representations of local object features in a new image acquired by the perception module 2, and generates and outputs a classification signal 10 based on the classified, segmented, or tracked representations of local object features detected in the new image.
[0242] System 1 can implement modules such as a perception module 2, an image processing module 3, an ROI determination module 4, a training data generation module 5, a training module 7, and a classification module 9, which are included in the software modules that run on the processor (processing hardware).
[0243] The processing hardware may include multiple processors, microprocessors, signal processors, and microcontrollers. The memory module 6 and the classification model memory 8 may be implemented using the same or different data storage devices, or they may be at least partially distributed across data storage devices and servers located at different sites and connected via a communication network.
[0244] All steps performed by the various entities described in this disclosure, as well as the functionalities described as being performed by the various entities, are intended to mean that each entity is adapted or configured to perform its respective step and functionality.
[0245] The functions of the modules discussed in this description may be implemented using discrete electrical hardware circuits. Alternatively, or in addition, at least some of the functions may be implemented in software in combination with at least one programmed microprocessor, general-purpose computer, application-specific integrated circuit (ASIC), or digital signal processor.
[0246] Figure 15 displays an example of a computer hardware element architecture suitable for carrying out an embodiment of a computer implementation method at a high level of abstraction, illustrating, in particular, interfaces to further hardware elements that are useful for understanding the elements of the embodiment.
[0247] The system 50 in Figure 15 includes a processor 51, data storage 53 (memory 53), an input / output interface 55, and a network interface 54, which communicate via a data bus 52.
[0248] The input / output interface 55 provides the ability to output information to a human user via visual or audible signals. The input / output interface 55 also provides the ability to obtain information and commands from a human user.
[0249] The input / output interface 55 is an interface for connecting input / output devices, including, but not limited to, keyboards, mice, pointing devices, displays, microphones, loudspeakers, or combinations thereof.
[0250] The processor 51 may be any type of controller or processor, and may even be embodied as one or more processors 51 adapted to perform the functionality discussed herein. The term processor 51 may encompass a single integrated circuit (IC), or it may encompass multiple connected, arranged, or grouped integrated circuits or other components, such as controllers, microprocessors, digital signal processors (DSPs), parallel processors, multicore processors, custom ICs, application-specific integrated circuits (ASICs), or field-programmable gate arrays (FPGAs).
[0251] The processor 51 can, in particular, provide hardware on which software implementing the modules and submodules of System 1, as discussed with reference to Figure 14, runs.
[0252] The memory 53 may include a data repository or database and may be embodied in any number of forms, including within any computer or other machine-readable data storage medium, memory device, or other storage or communication device for the storage or communication of information, including, but not limited to, volatile, non-volatile, removable, or non-removable memory ICs or memory portions of integrated circuits, such as resident memory in the processor 51.
[0253] Memory 53 may be adapted to store various lookup tables, parameters, coefficients, other information and data, programs or instructions of the software of this disclosure, and other types of tables such as database tables. Memory 53 may, among other things, store the memory module 6 and classification model module 8 of System 1, which are discussed with reference to Figure 14.
[0254] The processor 51 is programmed, for example, to implement the methodology of the disclosed computer implementation using the software and data structures of the disclosed computer implementation. As a result, the system 1 and computer implementation of the present invention may be embodied as software that provides such programming or other instructions, such as a set of instructions and / or metadata that are embodied in the computer-readable medium discussed above.
[0255] The camera sensor 58 may form part of system 1. The camera sensor 58 may include a single camera sensor or multiple camera sensors, in particular a first camera sensor and a second camera sensor.
[0256] The camera sensor 58 may be connected to system 1 via the network interface 54 instead of being directly connected to the data bus 52.
[0257] The lighting module 60 may form part of system 1, or it may be connected to system 1 via a network interface 54 instead of being directly connected to the data bus 52. The lighting module 60 emits light in a specific part of the electromagnetic spectrum to illuminate a physical marker device, for example, in an embodiment using invisible ink, where the physical marker device is visible to the camera sensor 58 only when illuminated in a specific part of the electromagnetic spectrum.
[0258] The network interface 54 provides System 1 with the ability to connect to external data sources, such as at least one server 57, via a communication network 56. The network interface 54 enables the implementation of System 1 in a spatially distributed manner, among other things, by performing at least some of the individual method steps at least partially remotely from System 1, or by storing data remotely from System 1. All steps performed by the various entities described herein, as well as the functionalities described to be performed by the various entities, are intended to mean that each entity is adapted or configured to perform its respective steps and functionalities.
[0259] The functions of the modules discussed in this description may be implemented using discrete electrical hardware circuits. Alternatively, or additionally, at least some of the functions may be implemented in software using application-specific integrated circuits (ASICs) or one or more digital signal processors in combination with a programmed microprocessor, a general-purpose computer.
[0260] In the claims and in this description, the word “comprising” does not preclude the presence of other elements or steps.
[0261] The indefinite article "a" or "an" does not exclude plurality.
[0262] A single element or module can perform the functions of several entities or items enumerated in the claims.
[0263] The present invention, as defined in the attached claims, can be described in the discussion of specific embodiments and combined with the features shown in the figures.
Claims
1. A method for generating image data for generating training information for automated image analysis related to local features in an image, The steps include applying at least one physical marker device (21, 25, 27, 28, 29, 30, 31, 32, 43, 45, 48) adjacent to the local feature (26) of the object (23), The steps include acquiring multiple images of the object (23) using at least one camera sensor (sensor 1, sensor 2), and storing the multiple images. A method that includes this.
2. Steps include applying at least one of the physical marker devices (21, 25, 27, 28, 29, 30, 31, 32, 43, 45, 48) adjacent to the local feature (26) on the surface (24) of the object (23). The method according to claim 1, further comprising:
3. The physical marker devices (21, 25, 27, 31, 32, 43, 45, 48) include and / or AR markers or QR markers placed on a carrier material. The carrier material is made of a flexible material, in particular a sheet of paper or a plastic plate. The method according to claim 1 or 2.
4. The at least one physical marker device (45) has an annular structure, and in particular, the surface of the at least one physical marker device (45) has a specific color or a specific pattern, in particular a specific dot pattern. The method according to claim 1.
5. The at least one physical marker device (21, 25, 27, 31, 32, 43, 45, 48) includes fastening means for attaching the at least one physical marker device (21, 25, 27, 31, 32, 43, 45, 48) to the surface (24) of the object (23), The method according to claim 1.
6. The fastening means includes at least one of an adhesive layer, a removable adhesive, a magnet (48.11), a suction cup, and a clip device. The method according to claim 5.
7. The at least one physical marker device (48) includes a plurality of physical marker devices (48.1) arranged in a pattern on the surface (24) of the object (23) that define a region of interest (22, 44) in at least one of the plurality of images. The method according to claim 1.
8. The at least one physical marker device (43) includes a pattern of invisible ink, the invisible ink includes a UV photofluorescent material, an NIR reflective material, a material that reflects light of a predetermined polarization, or a material that reflects electromagnetic waves in a predetermined frequency band. The method according to claim 1.
9. The physical marker device includes at least one body part of the user's body, in particular at least one finger (28, 30) or hand (29) positioned in a particular gesture, The method according to claim 1.
10. The at least one physical marker device (21, 25, 27, 31, 32, 43, 45, 48) is one of a plurality of types of physical marker devices (21, 25, 27, 31, 32, 43, 45, 48) that differ in size, The size of the at least one physical marker device (21, 25, 27, 31, 32, 43, 45, 48) defines the size of the region of interest (22, 44) in at least one of the plurality of images. The method according to claim 1.
11. Each of the multiple types of physical marker devices (21, 25, 27, 31, 32, 43, 45, 48) is associated with one of the following: a positive classification of the region of interest (22, 44), a negative classification of the region of interest (22, 44), and the classification confidence of the user applying the at least one physical marker device (21, 25, 27, 31, 32, 43, 45, 48). The method according to claim 10.
12. The at least one physical marker device (21, 25, 27, 31, 32, 43, 45, 48) includes shielding means for visually shielding at least a portion of the physical marker device (21, 25, 27, 31, 32, 43, 45, 48), The method according to claim 1.
13. The method includes the step of applying the at least one physical marker device (21, 25, 27, 31, 32, 43, 45, 48) adjacent to the local feature (26) of the object (23) at a predetermined in-plane angle with respect to the orientation of the image plane of the at least one camera sensor (sensor 1, sensor 2), The predetermined in-plane angle encodes the degree of membership of the local feature (26) into a specific class. The method according to claim 1.
14. A method for generating classification information for automated image analysis relating to local features (26) of an object (23) in an image, The steps include obtaining multiple images of the object (23), The steps include detecting at least one physical marker device (21, 25, 27, 28, 29, 30, 31, 32, 43, 45, 48) applied adjacent to the local feature (26) of the object (23) in at least one of the acquired plurality of images, For each detected physical marker device (21, 25, 27, 28, 29, 30, 31, 32, 43, 45, 48), the steps of calculating a region of interest (22, 44) in the at least one image based on predetermined relative position information associated with the at least one physical marker device (21, 25, 27, 28, 29, 30, 31, 32, 43, 45, 48), The steps include generating mask information based on the calculated regions of interest (22, 44), and storing the generated mask information associated with the at least one image as training information. The steps include: training a model using the stored training information to generate classification information for automated image analysis related to the local features (26); and A method that includes this.
15. The step of detecting the at least one physical marker device (21, 25, 27, 31, 32, 43, 45, 48) includes the step of detecting a specific pattern in the plurality of images based on a pre-trained computer vision model of the at least one physical marker device (21, 25, 27, 31, 32, 43, 45, 48), The method according to claim 14.
16. The method includes a plurality of application modes, When operating in the first application mode, the at least one physical marker device (21, 25, 27, 31, 32, 43, 45, 48) is associated with the region of interest (22, 44) in a positive example of the local feature (26), and other regions of the image represent regions (21, 25, 27, 31, 32, 43, 45, 48) in a negative example of the local feature (26), or When operating in the second application mode, the sensor view of the camera sensors (sensor 1, sensor 2) for acquiring the multiple images is restricted to a view that includes only the regions of interest (22, 44) in the multiple images, or the regions other than the regions of interest (22, 44) are covered, preventing the acquisition of image information from there. The method according to one of claims 14 and 15.
17. The steps include: displaying online on the screen of a handheld device or a wearable augmented reality / virtual reality device at least one region of interest associated with the detected at least one physical marker device (21, 25, 27, 28, 29, 30, 31, 32, 43, 45, 48) while training the model; The steps include obtaining user input from the user via a human-machine interface, which includes classifications to associate with displayed regions of interest (22, 44) or to terminate processing when a predetermined classification quality is reached, and The method according to claim 14, including the method described in claim 14.
18. The steps include: verifying the trained model by applying the trained model to the stored training information; The steps include determining whether the positive classification of the region of interest (22, 44) occurred outside the area in the plurality of images associated with the detected physical marker device (21, 25, 27, 28, 29, 30, 31, 32, 43, 45, 48), When it is determined that the positive classification of the region of interest (22, 44) occurred outside the area in the plurality of images associated with the detected physical marker device (21, 25, 27, 28, 29, 30, 31, 32, 43, 45, 48), A step of communicating to the user via a human-machine interface the determined positive classification of the region of interest (22, 44) that occurred outside the area in the image associated with the detected physical marker devices (21, 25, 27, 28, 29, 30, 31, 32, 43, 45, 48), A step of performing the method described above to generate new training information, and The step of generating new classification information by further training the model using the stored training information. A step of performing at least one of the following: The method according to claim 14, including the method described in claim 14.
19. The steps include generating new mask information based on the at least one image which includes a visually modified representation of the at least one detected physical marker device (21, 25, 27, 28, 29, 30, 31, 32, 43, 45, 48), A step of storing the newly generated mask information associated with the at least one image as further training information. The method according to claim 14, including the method described in claim 14.
20. A system for generating image data for generating training information for automated image analysis related to local features in an image, At least one physical marker device (21, 25, 27, 28, 29, 30, 31, 32, 43, 45, 48) applied adjacent to the local feature (26) of the object (23), At least one camera sensor (sensor 1, sensor 2) configured to acquire multiple images of the object (23), A memory (6) configured to store the aforementioned multiple images and A system equipped with these features.
21. A system for generating classification information for automatic image analysis relating to local features (26) of an object (23) in an image, wherein the system The system includes a processor (51) configured to acquire multiple images of the object (23), The aforementioned processor, In at least one of the acquired images, at least one physical marker device (21, 25, 27, 28, 29, 30, 31, 32, 43, 45, 48) is detected. For each detected physical marker device (21, 25, 27, 28, 29, 30, 31, 32, 43, 45, 48), a region of interest (22, 44) in the at least one image is calculated based on predetermined relative position information associated with the physical marker device (21, 25, 27, 28, 29, 30, 31, 32, 43, 45, 48). Based on the calculated regions of interest (22, 44), mask information is generated. The generated mask information associated with at least one of the images is stored in memory (6, 8) as training information. By training the model using the stored training information, the classification information for the automated image analysis relating to the local features (26) is generated. Further components of the system.