Apparatus, system, and method for generating localization maps of synthetic images
The framework generates realistic 2D images from 3D CAD models using procedural generation and Grad-CAM localization maps to enhance object detection accuracy by evaluating synthetic image quality and identifying critical regions, addressing inefficiencies in existing synthetic image assessment methods.
Patent Information
- Application Number
- PCT/EP2024/052591
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-02-02
- Publication Date
- 2025-08-07
AI Technical Summary
Existing methods for assessing synthetic image quality, particularly for object detection and identification, are inefficient and lack spatial information, leading to inaccurate training and performance of AI modules.
A framework that generates realistic 2D images from 3D CAD models using procedural generation and image rendering, combined with Grad-CAM localization maps to evaluate synthetic image quality and improve object detection accuracy by identifying critical regions affecting classification results.
Enhances the robustness and accuracy of object detection modules by providing a visual representation of image quality, allowing for quick identification of areas of confusion and improving the training process.
Smart Images

Figure EP2024052591_07082025_PF_FP_ABST
Abstract
Description
APPARATUS, SYSTEM, AND METHOD FOR GENERATING LOCALIZATION MAPS OF SYNTHETIC IMAGESTECHNICAL FIELD
[0001] This disclosure relates to an apparatus, a system, and a method for generating localization maps associated with synthetic images, in particular synthetic images of one or more components of an object.BACKGROUND
[0002] Synthetic data generation, which is the process of generating artificial data that mimics the characteristics of real-world data, has been receiving attention due to its potential to alleviate issues such as data scarcity, privacy concerns, and class imbalance in the fields of machine learning and data science. Specifically, synthetic image data has been applied to the generation of 2D images in various applications including object recognition.
[0003] The generated synthetic image data may be assessed by various methodologies. One current methodology is using perceptual studies, which involve human raters, and another methodology is using quantitative metrics, which measure one or more mathematical properties of the generated synthetic images. Perceptual studies may be time-consuming and are not scalable for large datasets. Quantitative metrics such as Inception Score (IS) and Frechet Inception Distance (FID) may fail to capture all relevant aspects of image quality. As an example, an RGB histogram analysis may provide global information about color distribution in an image but lacks spatial information regarding the location of such colors in the image, and may thus be less effective in discerning the fine details of an image.
[0004] There exists a need to provide an improved solution for assessing synthetic images, in particular synthetic images comprising one or more components of an object.SUMMARY
[0005] This disclosure was conceptualized to address the complexities and constraints in assessing synthetic image quality, particularly in quantifying a necessary amount of synthetic data for the training of Al modules, such as Al modules used for automatic object detection and identification, to optimize performance.
[0006] The present disclosure seeks to provide a technical solution that provides an object detection module framework with two-dimensional (2D) image quality evaluation capabilities, generates realistic-looking 2D image representation of a three-dimensional (3D) object, given a 3D computer-aided design (CAD) model, and increase an object detection module‘s prediction accuracy whilst reducing false positive confidence levels via quality data augmentation (i.e., combination of real and high-quality synthetic data). In one aspect, there is provided an image-based disassembly solution that exploits the use of 3D CAD data with image rendering techniques and search algorithms to enable synthetic defect generation. Notably, the disclosure seeks to improve the robustness and accuracy of an object detection module used for image-based disassembly of electric motors.
[0007] The present disclosure describes and demonstrates the generation of at least one localization map, which may be in the form of at least one heatmap, overlaid on one or more synthetic images. In some embodiments, the one or more synthetic images may comprise 2D synthetic images, and the synthetic images may be configured as an input to an Al-based object detection module. The Al-based object detection module may be configured to detect or identify one or more components of an object. The localization map may be used to identify one or more regions or areas which the Al-based object detection module focuses on. In some embodiments, the localization map may then be configured as a visual representation of one or more heatmaps, the one or more heatmaps may subsequently be used to provide an indication of quality of one or more generated synthetic images which are used to train the Al-based object detection module. In particular, the one or more heatmaps can be used to visually assess the quality of synthetic data, quickly pinpointing areas or regions of confusion for the Al-based object detection module used to identify different components of an object from the 2D image. Consequently, this method streamlines the process of verifying module performance and enhances the quality of image-based synthetic datasets.
[0008] According to an aspect of the present disclosure, an apparatus as claimed in claim 1 is provided. According to another aspect of the present disclosure, a system as claimed in claim 9 is provided. According to another aspect of the present disclosure, a computer-assisted method according to the disclosure is defined in claim 13. A computer program comprising instructions to execute the computer-assisted method is defined in claim 15.
[0009] The dependent claims define some examples associated with the apparatus, system, and method, respectively.BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The disclosure will be better understood with reference to the detailed description when considered in conjunction with the non-limiting examples and the accompanying drawings, in which:- FIG. 1 is a schematic diagram of an embodiment of an apparatus 100 for generating a localization map of a synthetic image;- FIG. 2 is a data augmentation framework wherein the apparatus 100 may be applied or integrated thereto, the data augmentation framework suitable for detecting components of an object for an automated disassembly process;- FIG. 3 is an illustration of an overall procedural generation configuration, in the form of a pseudo-code flow chart, for generating photorealistic eAxle texture;- FIG. 4 illustrates an integration of the Grad-CAM module with an embodiment of the object detection module, the object detection module in the form of a YOLOv4;- FIG. 5 illustrates the output of the feature importance module highlighting with Grad- CAM via region size and gradient values;- FIG. 6 shows various generated Grad-Cam localization or heatmaps of the YOLOv4 convolutional layers of interest for small-sized components, medium-sized components, and large-sized components;- FIG. 7 depicts the observed variations in synthetic data quality, produced by two distinct rendering methods, as interpreted by the YOLOv4 module; and- FIG. 8 is a flow chart of a method for generating a localization map of a synthetic image.DETAILED DESCRIPTION
[0011] The following detailed description refers to the accompanying drawings that show, by way of illustration, specific details and embodiments in which the disclosure may be practiced. These embodiments are described in sufficient detail to enable those skilled in the art to practice the disclosure. Other embodiments may be utilized and structural, and logical changes may be made without departing from the scope of the disclosure. The various embodiments are not necessarily mutually exclusive, as some embodiments can be combined with one or more other embodiments to form new embodiments.
[0012] Embodiments described in the context of one of the systems or methods are analogously valid for the other systems or methods.
[0013] Features that are described in the context of an embodiment may correspondingly apply to the same or similar features in the other embodiments. Features that are described in the context of an embodiment may correspondingly apply to the other embodiments, even if not explicitly described in these other embodiments. Furthermore, additions and / or combinations and / or alternatives as described for a feature in the context of an embodiment may correspondingly be applicable to the same or similar feature in the other embodiments.
[0014] In the context of various embodiments, the articles “a”, “an” and “the” as used with regard to a feature or element include a reference to one or more of the features or elements.
[0015] As used herein, the term “and / or” includes any and all combinations of one or more of the associated listed items.
[0016] As used herein, the term “data” may be understood to include information in any suitable analog or digital form, for example, provided as a file, a portion of a file, a set of files, a signal or stream, a portion of a signal or stream, a set of signals or streams, and the like. The term data, however, is not limited to the aforementioned examples and may take various forms and represent any information as understood in the art.
[0017] As used herein, the term “module” refers to, or forms part of, or include an Application Specific Integrated Circuit (ASIC); an electronic circuit; a combinational logic circuit; a field programmable gate array (FPGA); a processor (shared, dedicated, or group) that executes code; other suitable hardware components that provide the described functionality; or a combination of some or all of the above, such as in a system-on-chip. The term module may include memory (shared, dedicated, or group) that stores code executed by the processor.
[0018] As used herein, the terms ‘first’, ‘second’, ‘third’, and so on, are used for purposes of clarity and do not imply order or precedence.
[0019] As used herein, the term “disassembly process” refers to a process whereby an object is separated into its components. In some embodiments, the disassembly process may be used to separate one or more components (e.g. bolts or other fasteners) from a product (e.g. electric motor) and / or subassemblies via non-destructive or semi-destructive processes which target the connectors / fasteners. Such disassembly process may have applications in various industries, in particular, but not limited to, a waste treatment facility, reverse engineering processes, etc. The disassembly process may broadly include a guided disassembly processusing augmented reality, virtual reality, and / or mixed reality; an automated disassembly (i.e. full automation).
[0020] As used herein, the term “three-dimensional (3D) scanner” may refer to one or more image capturing devices, and may include various 3D scanning techniques. In some embodiments, the 3D scanners include structured light 3D scanners, which can be configured to collect high quality data. In some embodiments, the 3D scanning techniques may be based on stereo vision methods, and / or time-of-flight methods.
[0021] As used herein, the term “localization map” refers to a visual representation of data points allowing users to identify the more salient or important data points. In some embodiments, the localization map may broadly include a two-dimensional or higherdimensional representation of data in which various values may be represented by different colors and / or color intensity, i.e. in the form of a heat map.
[0022] It is envisaged that structured light 3D scanners, may produce scanning having relatively higher precision and / or accuracy. A structured light 3D scanner may be configured to work using principles of triangulation. A light source of the 3D scanner may project a fringe pattern across the scan target object surface, and two cameras of the 3D scanner may be used to capture the surface geometry based on the pattern distortion, calculating 3D coordinate measurements. The 3D scanner may process coordinate data into a "point cloud," creating a digital image of the object. In some embodiments, the light source of the 3D scanners may be configured to emit white light. In some embodiments, the light source may be configured to emit other light such as blue light, which may be used to capture data on shinier and darker colored surfaces and filters out the ambient light.
[0023] As used herein, the term “procedural generation” broadly refers to any method of creating image data algorithmically as opposed to manually. Such methods may include a combination of human-generated content and algorithms coupled with computer-generated randomness and processing power. In some embodiments, procedural generation includes algorithmically generating textures, terrains, and objects, by utilizing mathematical algorithms to autonomously generate complex textures, landscapes, and objects by harnessing computer processing power. The term synthetic image broadly refers to one of more images created by such procedural generation methods.
[0024] As used herein, the term “processor” refers to a circuit, including analog circuits or components, digital circuits or components, or hybrid circuits or components. Any other kindof implementation of the respective functions which will be described in more detail below may also be understood as a "circuit" according to an alternative embodiment. A digital circuit may be understood as any kind of a logic implementing entity, which may be special purpose circuitry or a processor executing software stored in a memory, or a firmware.
[0025] As used herein, the term “obtain”, as used herein, refers to the processor which actively obtains the inputs, or passively receives inputs from a user interface and / or one or more sensors. The term obtain may also refer to the processor that receives or obtains inputs from a communication interface, e.g. the user interface. The processor or the dosing module may also receive or obtain the inputs via a memory, a register, and / or an analog-to-digital port.
[0026] As used herein, the term “heatmap-based methods” includes Gradient-weighted Class Activation Mapping (Grad-CAM) is a generalization of a class activation mapping (CAM) technique. Grad-CAM can also be applied to non-classification examples such as regression or semantic segmentation.
[0027] An embodiment of the disclosure is shown in FIG. 1, which illustrates a setup of an apparatus 100 for evaluating or assessing the quality of one or more generated synthetic images 104. Each synthetic image 104 may comprise an object 110, and at least one (one or more) regions 140. Each region may correspond to a component of the object 110. In some embodiments, apparatus 100 may be a server apparatus.
[0028] The apparatus 100 may be configured to provide a visual representation of the synthetic image 104, the visual representation may be in the form of a localization map indicating one or more regions 140 within the synthetic image 104 that correspond to areas that may affect the training and classification of an Al-based object detection and classification module. The identified one or more regions 140 may then be used to provide an indication of quality of the generated synthetic image 104. For example, the one or more regions 140 identified may be used for the subsequent determination of feature data within the synthetic image 104 that may affect the accuracy of the classification result of the Al-based object detection and classification module. In some embodiments, the indication of quality may be understood as a correlation between the accuracy of the classification result and the feature data, such as to enable further analysis carried out on the feature data so as to improve the accuracy. In some embodiments, the indication of quality may be an arrangement or configuration wherein each synthetic image 104 is assigned a quality score based on theaccuracy of the classification result. In some embodiments, such a quality score may be normalized.
[0029] The apparatus 100 may comprise at least one processor 150, the at least one processor 150 configured to: obtain a three-dimensional (3D) image 102 of the object 110; generate a two-dimensional (2D) realistic image 102A of the 3D image 102; generate, using a procedural generation algorithm, and the 2D realistic image 102A, one or more two- dimensional 2D synthetic images 104 of the 3D image data 102; augment the 2D realistic image 102 A and the one or more 2D synthetic images 104 to form a training dataset 106; train, using the training dataset 106, an artificial -intelligence (Al) based object detection module 107, to obtain a classification result 108; generate, using a gradient-based localization algorithm 170, a localization map 172 based on the one or more 2D synthetic images 104; wherein the localization map 172 identifies at least one region 140 of the object 110 that affects the classification result 108.
[0030] In some embodiments, the procedural generation algorithm is implemented as part of a computer graphics software toolset. In some embodiments, the computer graphics software toolset may be a Blender™ [1], configured for enabling procedural generation, and for crafting realistic scenes. The computer graphics tool may involve algorithmically generating textures, terrains, and objects, as opposed to manual sculpting or designing each facet, thus utilizing mathematical algorithms to autonomously generate complex textures, landscapes, and objects by harnessing computer processing power rather than depending on manual input. In some embodiments, the nodes, the foundational elements in Blender™ node-based editors like the shader editor or the geometry nodes editor, may be employed for implementing the procedural generation algorithm. The nodes may be instrumental in creating, modifying, and amalgamating various data types, including geometry, textures, and materials. For example, in the geometry nodes editor, a node network delineating the rules and algorithms for geometry generation may be created. In some embodiments, the complexity of the node networks may be tailored to the specific needs, allowing for optimal control and flexibility over the created content. Importantly, the generated data encompasses not only geometry but also textures, materials, and other scene components.
[0031] In some embodiments, the gradient-based localization algorithm 170 comprises a Gradient-weighted Class Activation Mapping (Grad-Cam) algorithm [2], and the localizationmap may comprise heatmap data, the heatmap data may be visualized as a heatmap overlay on the one or more synthetic images 104.
[0032] In some embodiments, the Al-based object detection module may comprise or is a real-time object detection module, which may include classification capabilities. For example, in a 2D image, the Al-based object detection module may be trained to detect one or more objects, one or more components, and classify the object and components. In some embodiments, the real-time object detection module may include a convolutional neural network. The convolutional neural network may comprise a plurality of convolutional layers and a series of fully connected layers. In some embodiments, the convolutional neural network may be a YOLOv4 [3] algorithm.
[0033] FIG. 2 shows a data augmentation framework 200 wherein the apparatus 100 may be applied or integrated thereto. The system 200 may include an image capturing device 202, and may incorporate the apparatus 100. The system 200 may be suited for the automated disassembly of an electric motor 204 using a robotic arm 208.
[0034] The image capturing device 202 may include a three-dimensional (3D) scanner and an RGB image capturing device (e.g. a RGB camera). The 3D scanner may be configured for point cloud data collection and the RGB camera may be configured to capture 2D images of the object 110. In the context of an automated disassembly of the electric motor 204, the components to be disassembled from the electric motor 204 may include one or more bolts 206.
[0035] In some embodiments, the 3D scanner may be mounted on the robotic arm 208. The robotic arm 208 may be configured to move in at least one rotational motion and at least one translational displacement. One or more platforms (not shown) may be made available for the object 110 to be positioned thereon so that 3D scans may be obtained. In some embodiments, the 3D scanner may be a structured light 3D scanner.
[0036] In some embodiments, the robotic arm 208 may be a robotic arm with six degree of freedom. In some embodiments, the robotic arm may be a six-axis articulated robotic arm, and may comprise one or more controllers for communication with a processor 250. It is appreciable that the robotic arm may have higher degrees of freedom, i.e. more than 6 degrees of freedom.
[0037] In some embodiments, the robotic arm 208 may be a collaborative robot (cobot).
[0038] In some embodiments, the robotic arm 208 may be equipped for the disassembly of the electric motor components, which include, but are not limited to, the removal of coil windings, differential gears, and unscrewing of exterior automotive bolts. It may be appreciable that the simplicity and homogeneity of the electric motor’s internal components may indicate a simpler way to remove the internal components than the exterior encasement, which is frequently exposed to unknown environmental conditions.
[0039] The system 200 depicted in FIG. 2 may be configured to receive 3D image data 102 of the electric motor 204 and coordinates information relating to one or more bolts 206. Based on the image data of the electric motor 204, the set of component data relating to the one or more components 140 may be derived or generated, and the relatively realistic looking (i.e., synthetic) deterioration of the components, i.e. the one or more bolts 206, may be overlaid on the components 140. In some embodiments, one or more 2D images of the object 110 may be captured under different light conditions, and the components may be manually labelled or annotated for subsequent training of the component localization algorithm and / or the feature extraction algorithm.
[0040] As depicted in FIG. 2, there is an edge computing device 250A arranged to acquire data from the image capturing device 202, the edge computing device 250A further comprises a robot control module 252 configured to send control signals to control the robotic arm; and an Al-based component detection module 254, the component detection module 254 configured to determine, using a trained object detection and classification algorithm, a bounding box around each of the one or more bolts 206 on the electric motor 204. The AI- based component detection module 254 may be configured as a computer vision-based bolt detection sensor hardware and algorithm to detect and account for all the bolts to be unscrewed whilst assessing the health status of individual bolts (i.e. degree of defect, such as rust).
[0041] The system 200 further comprises an edge computing server 250B, the edge computing server 250B configured for the training and deployment of the component detection module 254. The apparatus 100 may be used to train the edge server 250B, before deployment on the edge computer 250A, on a periodic or need-to-basis.
[0042] The Al-based module 254 may be trained to perform bounding box labeling using a pretrained Al module’s predictions as label data (hereinafter also referred to as labels). These labels may be prior reviewed and corrected by human and use to train the Al module. In some embodiments, the labelling may be automatically performed, i.e. directly use the module’spredictions as labels, no review or correction needed. In some embodiments, manual annotations or crowd-sourcing may be used, i.e. distributing the data labelling task to a large group of people via online platforms.
[0043] The system 200 includes an image synthesizer module 220, a heat-map feature importance module 230 and an augmented image database 240.
[0044] The image synthesizer module 220 is configured to generate, using the procedural generation algorithm and any 2D realistic image captured by the image capturing device 202, one or more two-dimensional 2D synthetic images of the 3D image data. The generated synthetic images may be automatically labelled using an annotation module 256. The annotation module 256 may form part of the Al-based object detection module or may be an independent annotation module 256. The annotation or labelling may be performed on the electric motor 204, as well as one or more bolts 206, for every image generated via the image synthesizer module 220. The generated synthetic images may then be stored in a synthetic image database 258 as synthetic dataset. The synthetic image database 258 may be arranged in data communication with the augmented image database 240.
[0045] The edge computing server 250B may be configured to obtain realistic images 102A of the electric motor 204 for manual annotation performed via a manual annotation module 262. The manually annotated realistic images 102A may then be stored in a real image database 264.
[0046] The augmented image database 240 may be configured to augment the realistic images 102 A and the 2D synthetic images 104 to form a training dataset for training or retraining the pre-deployed Al based object detection module 254 A before deployment or update on the edge computing device 250A.
[0047] When the object detection module 254 is triggered for module training or re-training, image samples are retrieved from the augmented image database 240. Within the augmented dataset, the training dataset comprises of both realistic image 102A and synthetic images 104, while a testing dataset (not shown) may only contain realistic images 102A.
[0048] The output of the trained object detection module 254, which may include one or more classification results or scores, may be sent to the heatmap feature module 230. The heatmap feature module 230 may be configured to generate, using the gradient-based localization algorithm 170, the localization map 172 based on the one or more 2D synthetic images 104, wherein the localization map 172 identifies at least one region 140 of the object110 which may have relevance and importance to the outcome of the classification results or scores.
[0049] The heatmap feature module 230 may provide a graphical user interface for a human evaluator to determine if the newly trained module 254 performs at least as good as the previously trained module and baseline module results. For interpretability purposes, the human evaluator can utilize the localization map to gain insights into the module 254 decisionmaking, and changes may be made to fine-tune or tweak the image synthesizer parameters for generating future synthetic images, i.e. the fine-tuning or tweaking may be in the form of a feedback parameter 266 and send the at least one feedback parameter 266 to the procedural generation algorithm for generation of the one or more two-dimensional (2D) synthetic images (104).
[0050] In some embodiments, the 3D image data 102 may be in the form of a 3D computer- aided design (CAD) data.
[0051] Regarding the effects of augmented data on the performance of the object detection module 254, empirical evidence illustrates a 1 :6 ratio of actual to synthetically generated images (i.e., 25 real images is to 150 synthetic images). Notably, producing further synthetic images 106 yields a reduced marginal increment based on a mean average precision (mAP) score. This may underscore the importance of focusing on the quality of the synthetic images generated, rather than simply increasing their quantity.
[0052] In some embodiments, comparative studies based on a known UV mapping, a 3D modelling process of projecting a 3D module's surface to a 2D image for texture mapping, is compared with the procedural generation algorithm of the present disclosure. The image resulting from UV mapping displays visual artifacts and is somewhat overexposed, whereas the image rendering using the procedural generation technique displays a markedly higher level of realism.
[0053] FIG. 3 refers to an embodiment of a module 300 for generating realistic images of an object and / or component thereof, based on procedural generation, based on input parameters for texture generation. The module is configured to generate photorealistic textual of images related to an electric drive system also referred to as eAxle. The module 300 comprises four sub-modules, an edge wear sub-module 301, a blemish sub-module 311, a roughness submodule 321, and a scratch sub-module 331. Each of the sub-modules 301, 311, 321, and 331are configured to receive respective, input parameters for textual generation, in order to replicate real-world textures on the eAxle module.
[0054] The edge wear sub-module 301 is configured to receive an edge wear input parameter and process the edge wear input parameter using an arithmetic function 302, ambient occlusion function 303, a threshold function 304, and a grayscale colour map function 305.
[0055] The blemish sub-module 311 is configured to receive a blemish input parameter and process the blemish input parameter through a arithmetic function 312, perlin noise (Musgrave texture) function 313, threshold function 314, and grayscale colour map function 315.
[0056] The output from the grayscale colour map functions 305, 315 may be combined by a colour map mix function 316 to provide a base colour output 341.
[0057] The roughness sub-module 321 is configured to receive an roughness input parameter and process the roughness input parameter using an arithmetic function 322, which in turns pass through a perlin noise (Musgrave texture) function 323, a metallic value function 324 (to output a metallic value 342) and a rough value function 325 (to output a roughness value 343). The output of the perlin noise function 323 is passed through a bump map generation function 326.
[0058] The scratch sub-module 331 is configured to receive a scratch input parameter, and then process the scratch input parameter using an arithmetic function 332, perlin noise function (to generate a wave texture) 333, perlin noise function (to generate a musgrave texture) 334, a threshold function 335, and a bump map generation function 336. The perlin noise function 333 may be configured to receive an input from a random number generator 337, and the perlin noise function 334 may be configured to receive an input from the arithmetic function 332 or another arithmetic function 339.
[0059] The output from the bump map generation functions 326, 336 may be merged by a bump map merge function 338. The output of bump map merge function 338 may provide a normal output 344. The various outputs 341, 342, 343, 344 may be applied as material attributes to the 3D model.
[0060] In summary, FIG. 3 illustrates a combination of nodes with specific parameters configurations to achieve the mentioned realism. For instance, a combinatorial sequence of nodes (i.e., multiply, subtract, ambient occlusion, color ramp), named “edge wear”, are used to simulate signs of physical wear at the edges of the 3D model. Due to the 2D nature of the CAD model’s surface, a combination of nodes are responsible for adding the perception of“roughness” type of visual realism to imitate the actual surface appearance of the eAxle. The same applies to the remaining node group combinations for “blemish” and “scratch”. Thereafter, the outputs of these combinatorial nodes are mixed and applied onto the 3D CAD model via the blender node “principled BSDF”, before visual output.
[0061] In some embodiments, the Grad-Cam algorithm may be configured to receive output from a dense prediction module (YOLO head) of the YOLOv4. The Grad-CAM algorithm may be configured to provide a visual representation of the heatmap to facilitate a user in the identification of one or more parts of the image 110 that are crucial for a particular classification decision made by a convolutional neural network (CNN). For example, the Grad-CAM algorithm may be configured to highlight the importance of different regions in an image for every prediction. The Grad-CAM algorithm is configured to use the gradients of the target class (for example, the classification “electric motor” and “bolts”) with respect to the feature maps of a last convolutional layer of the CNN. These gradients may be global-average-pooled to obtain the weights of each feature map for the target class. These weights may signify the importance of each feature map for the target class. Then, a weighted combination of the feature maps may be used to compute or generate the Grad-CAM heatmap, which highlights the important regions in the image for the target class. For example, the regions in the image where the gradient values are relatively larger may correspond to the object detection module focus regions as well, , which elucidates that a minor alteration in the electric drive's angle, combined with the repositioning of a single bolt, can affect the regions of feature importance within every image.
[0062] The Grad-CAM heatmap may be visualized as an overlay on the one or more synthetic images 104, to interpret the module's decision. While the generated Grad-CAM heatmap is primarily used for visual explanations, the use of the Grad-CAM heatmap may be used to estimate the quality of the synthetic data generated. In particular, Grad-CAM utilizes the concept of gradient- weighted class activation mapping to generate visual explanations from deep networks. The Grad-CAM may be configured to produce a coarse localization map, underscoring areas in the image which may be most relevant to the specific object class predictions under consideration. Such an approach is grounded in the notion that higher gradient values indicate more important regions for the model's decision-making process, thereby illuminating the 'why' behind the network's predictions. By mapping the highlighted gradient values with known bounding box positions of the object class to be detected, thisinformation can be used to infer the quality of the synthetic data image. For example, as the gradient values increases (i.e., becomes darker in color), this means that the Y0L0v4 model places more emphasis on the regions of interest within the image.
[0063] FIG. 4 illustrates an integration of the Grad-CAM module with an embodiment of the object detection module, the object detection module in the form of a YOLOv4 (You Only Look Once version 4) object-detection module 400. The module 400 comprises a YOLOv4 pipeline 401, the pipeline 401 comprises a backbone network 402, a YOLO head 403, and a sparse prediction sub-module 404. An input image data 420 may be fed into the backbone network 402 which processes the input image data 420 and feeds it to the YOLO head 403 for the classification of small objects (labelled Conv2d_93), medium objects (labelled conv2d_101), and large objects (labelled conv2d_109). The three outputs may be fed into a Grad-Cam module 405, which in turn generates localization or heatmaps 406, 407, 408 associated with the small, medium and large objects. The sparse prediction module provides a classification of the various objects 409 based on the trained YOLOv4. In some embodiments, due to framework incompatibility between Grad-CAM and YOLOv4, the YOLOv4 module framework may requires conversion from Darknet to Tensorflow. The YOLOv4 may be configured to utilize three convolutional layers to detect objects of different sizes, with each layer having a distinct feature map. In some embodiments, the three convolutional layers to determine the presence and relevance of features are labelled ‘conv2d_93’, ‘conv2d_101’, and ‘conv2d_109’ in FIG. 4. The three convolutional layers are for detecting objects ranging from small, medium and large sizes, respectively. When fed with images of real electrical drives, it was observed that the Grad-CAM heatmap responsible for detecting small objects, layer ‘conv2d_93’, successfully identifies regions of interest. However, the heatmaps of the other two layers lack any features, indicating that no potential medium or large objects could be identified, see FIG. 6, which shows the generated Grad-Cam localization or heatmaps of the YOLOv4 convolutional layers of interest for small-sized components 406, medium-sized components 407, and large-sized components 408. The finding is congruent with the proposed current framework of FIG. 2, as the YOLOv4 module was finetuned to detect only small bolts on the body of the electric drive. Hence, it may be appreciable that the use of ‘conv2d_93’ may be used for the detection of bolts and for gaining insights into the module interpretability. Likewise, the same method is used to validate the synthetic data quality that is generated (i.e.,training set comprise of real images and test set contains synthetic data). If necessary, module layers ‘conv2d_101’ and ‘conv2d_109’ can be also used for the detection of larger objects in the electric drive. For example, electric drive mechanical gears, ball bearings, and gear shafts.
[0064] In some embodiments, Grad-CAM may be applied to compute the gradients for layer ‘conv2d_109’. These layers may be before the bounding box prediction layers of the YOLOv4 module and may serve as an auxiliary method to interpret the YOLOv4 module’ s focus regions.
[0065] FIG. 5 illustrates two localisation maps (heat maps) 501, 502 generated by the Grad- CAM module of an electric motor with bolts, with the heat map 502 associated with an input image of a 10 degrees angle counter-clockwise and a bolt switch with respect to the input image associated with heat map 501. FIG. 5 elucidates that a minor alteration in the electric drive's angle, combined with the repositioning of a single bolt, can affect the regions of feature importance within every image.
[0066] FIG. 7 depicts the observed variations in synthetic data quality 520, 530, produced by two distinct rendering methods, as interpreted by the YOLOv4 module. The synthetic image labelled 520 is generated using a UV unwrapping (heat map labelled as 525), and the synthetic image labelled 530 is generated using the procedural generation algorithm (heat map labelled as 535) according to the present disclosure. In the absence of Grad-CAM, it may be difficult or impossible to evaluate the quality of the generated synthetic data and understand the subsequent effects on the prediction accuracy of object detection modules like YOLOv4.
[0067] In summary, the invention’s proposed approach serves dual purposes: facilitating YOLOv4 module interpretability and cross-validation of synthetic data quality with respect to the bounding box confidence scores by YOLOv4. Without these features, generating vast amounts of synthetic data will be pointless, as it would simply add noise to the training dataset and confuse the Al-based object detection module. Notably, this component’s contribution is key to enabling subsequent innovations presented in this invention report.
[0068] It is appreciable that the data labelling / annotation of the targeted objects may be automated with relatively higher accuracy whilst improving the human productivity levels for highly repetitive task of data annotation.
[0069] In some embodiments, the apparatus 100 may be suited for an automated disassembly process. In some embodiments, the automated disassembly process may be utilized for the disassembly of an electric motor, and may include the removal of one or more components 140, such as one or more bolts, from the electric motor. In some embodiments, theautomated disassembly process may form part of a waste or materials recovery facility, such as an e-waste recovery facility.
[0070] In some embodiments, the apparatus 100 further comprises a graphical user interface (GUI) 170 arranged in data communication with the processor 150, the GUI 170 configured to display a set of image parameters associated with the simulated defect image data for input by a user. The set of parameters may include a first image parameter 172 associated with a degree of defect, and a second image parameter 174 associated with an environment.
[0071] In some embodiments, the set of component data is obtained based on a comparison data between the set of 3D coordinates of the component 140 and the data of a trained component database.
[0072] According to another aspect and with reference to FIG. 8, there is provided a computer-aided method 600 for generating a localization map 172 of a generated synthetic image 104 of an object 110, the method 600 comprising: obtaining 602 a three-dimensional 3D image 102 of the object 110; generating or obtaining 604 a two-dimensional 2D realistic image 102 A of the 3D image 102; generating 606, using a procedural generation algorithm and the 2D realistic image 102A, one or more two-dimensional 2D synthetic images 104 of the 3D image data 102; augmenting 608, the 2D realistic image 102A and the one or more 2D synthetic images 104 to form a training dataset 106; training 610, using the training dataset 106, an artificial -intelligence (Al) based object detection module, to obtain a classification result 108; generating 612, using a gradient-based localization algorithm 170, a localization map 172 based on the one or more 2D synthetic images 104; wherein the localization map identifies at least one region 140 of the object 110 that affects the classification result 108.
[0073] In some embodiments, known methodology of explainable Al may be used to supplement the Grad-CAM heatmap in providing an indication of the quality of the generated synthetic images. Such methodology may include techniques such as feature importance analysis, partial dependence plots, and module-agnostic methods like LIME and SHAP. These methods aim to elucidate how Al modules make decisions while maintaining transparency and interpretability.
[0074] It may be contemplated that the present disclosure may be used to re-purpose Grad- CAM as an effective evaluator of synthetic image quality, thus providing more detailed, interpretable evaluation mechanisms from the Al module perspective.References[1] Blender Foundation. (2023). Blender (Version 3.1) htps: / / www.blender.org / [2] Selvaraju, R.R., Cogswell, M., Das, A., Vedantam, R., Parikh, D. and Batra, D., 2017. Grad-cam: Visual explanations from deep networks via gradient-based localization. In Proceedings of the IEEE international conference on computer vision (pp. 618-626).[3] Bochkovskiy, A., Wang, C.Y. and Liao, H.Y.M., 2020. Yolov4: Optimal speed and accuracy of object detection. arXiv preprint arXiv: 2004.10934.
Claims
CLAIMS1. An apparatus (100) for generating a localization map (172) associated with one or more synthetic images (104) of an object (110), the apparatus (100) comprising at least one processor (150), the at least one processor (150) configured to: obtain a three-dimensional (3D) image (102) of the object (110); obtain a two-dimensional (2D) realistic image (102A) of the 3D image (102); generate, using a procedural generation algorithm and the 2D realistic image (102A), the one or more synthetic images (104); augment the 2D realistic image (102A) and the one or more synthetic images (104) to form a training dataset (106); train, using the training dataset (106), an artificial -intelligence (Al) based object detection module (107), to obtain a classification result (108); generate, using a gradient-based localization algorithm (170), a localization map (172) based on the one or more synthetic images (104); wherein the localization map (172) identifies at least one region (140) of the object (110) on the one or more synthetic images (104) that affects the classification result (108).
2. The apparatus (100) of claim 1, wherein the procedural generation algorithm is implemented as a computer graphics software toolset or part thereof.
3. The apparatus (100) of claim 1 or 2, wherein the Al-based object detection module (107) comprises or is a real-time object detection module (107), wherein the real-time object detection module (107) comprises a plurality of convolutional layers and a series of fully connected layers, wherein the real-time object detection module (107) optionally includes a YOLOv4 pipeline (401).
4. The apparatus (100) of any one of claims 1 to 3, wherein the gradient-based localization algorithm (170) comprises a Gradient-weighted Class Activation Mapping (Grad-Cam) algorithm, and the localization map (172) comprises a heatmap (406, 407, 408), the heatmap (406, 407, 408) visualized as an overlay on the one or more synthetic images (104).
5. The apparatus (100) of claim 3 and claim 4, wherein the Grad-Cam algorithm is configured to receive output from a dense prediction module (403) of the convolutional neural network-based algorithm.
6. The apparatus (100) of any one of the preceding claims, wherein the apparatus (100) comprises a graphical user interface (GUI), the GUI configured to display the localization map (172) for input by a user.
7. The apparatus (100) of claim 6, wherein the GUI is configured to obtain at least one feedback parameter (226) and send the at least one feedback parameter (226) to the procedural generation algorithm for generation of the one or more synthetic images (104).
8. The apparatus (100) of any one of the preceding claims, wherein the at least one processor (150) is configured to provide a testing dataset, the testing dataset comprises only a plurality of the 2D realistic images (102A).
9. An automated disassembly system (200), the system (200) comprising the apparatus (100) of any one of the preceding claims, wherein the one or more synthetic images (104) output from the apparatus (100) is configured as training input, and the Al-based object detection module (107) comprises a computer vision-based component detection Al algorithm.
10. The system (200) of claim 9, further comprising a three-dimensional (3D) scanner and an RGB image capturing device (202), the 3D scanner and the RGB image capturing device (202) configured to obtain the 3D image of the object (110), wherein the object (110) is an electric motor (204), and the at least one region (140) comprises one or more bolts (206) of the electric motor (204), wherein the system (200) further comprises a robotic arm (208) operable to remove the one or more bolts (206), and wherein the three-dimensional scanner and the RGB image capturing device (202) are mounted on the robotic arm (208).
11. The system (200) of claim 10, wherein the system (200) further comprises an edge computing server (150B), the edge computing server (150B) comprises the Al-based object detection module (107).
12. The system (200) of claim 10 or 11, wherein the robotic arm (208) is arranged in data communication with an edge computing device (150A), the edge computing device (150A) comprises a robot control module (252) configured to send control signals to control the robotic arm; and a component detection module (254), the component detection module configured to determine, using the component localization algorithm, a bounding box around each of the one or more bolts on the electric motor, and the set of 3D coordinates of each of the one or more bolts.
13. A computer-aided method (600) for generating a localization map (172) associated with of one or more synthetic images (104) of an object (110), the method (600) comprising: obtaining a three-dimensional (3D) image (102) of the object (110); obtaining a two-dimensional (2D) realistic image (102 A) of the 3D image (102); generating, using a procedural generation algorithm and the 2D realistic image (102A), one or more synthetic images (104); augmenting the 2D realistic image (102A) and the one or more synthetic images (104) to form a training dataset (106); training, using the training dataset (106), an artificial -intelligence (Al) based object detection module (107), to obtain a classification result (108); generating, using a gradient-based localization algorithm (170), the localization map (172) based on the one or more synthetic images (104); wherein the localization map identifies at least one region (140) on the one or more synthetic images (104) of the object (110) that affects the classification result (108).
14. The computer-aided method of claim 13, wherein the generated synthetic image (104) comprises a two-dimensional (2D) image data and a label data associated with a component within the at least one region (140).
15. A computer program, the computer program comprising instructions to execute the computer-assisted method according to claim 13 or 14.
Citation Information
Cited By
Escalator step detection method fusing YOLO algorithm and digital twinning technology
CN122024174A