Experiential Learning in Virtual Worlds

By determining the third domain on the virtual scene dataset and learning the domain adaptive neural network, the problem of domain differences between virtual data and real data is solved, and object inference with high accuracy and robustness in real scenes is achieved, especially in manufacturing scenarios.

CN111898173BActive Publication Date: 2025-08-12DASSAULT SYSTEMES SA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010371530.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-05-06
Filing Date
2020-05-06
Publication Date
2025-08-12
Estimated Expiration
2040-05-06

AI Technical Summary

Technical Problem

When training deep models with virtual data, it is difficult to effectively narrow the domain differences between virtual data and real data, resulting in insufficient generalization capabilities of the model in real-world scenarios, especially when processing objects that can be reconstructed in space.

Method used

By providing a virtual scene dataset and determining the third domain to make it closer to the data distribution of the real scene, a domain adaptive neural network is learned and a spatially reconstructed neural network is used to infer objects in the real scene.

Benefits of technology

It improves the accuracy and robustness of the inference space of the object being reconstructed in real scenarios, especially in manufacturing scenarios, overcoming data scarcity and annotation difficulties, and achieving high-quality object recognition and detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111898173B_ABST
    Figure CN111898173B_ABST
Patent Text Reader

Abstract

The present invention particularly relates to a computer-implemented machine learning method. The method includes providing a dataset of a virtual scene. The dataset of the virtual scene belongs to a first domain. The method also includes providing a test dataset of a real scene. The test dataset belongs to a second domain. The method also includes determining a third domain. The third domain is closer to the second domain than the first domain in terms of data distribution. The method also includes learning a domain-adaptive neural network based on the third domain. The domain-adaptive neural network is a neural network configured to infer spatially reconstructable objects in a real scene. Such a method constitutes an improved machine learning method for a dataset including a scene with spatially reconstructable objects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer programs and systems, and more particularly to methods, systems, and programs for machine learning with a dataset of a scene including spatially reconstructable objects. Background Art

[0002] Numerous systems and programs are available on the market for the design, engineering, and manufacturing of objects. CAD stands for Computer-Aided Design, i.e., software solutions for designing objects. CAE stands for Computer-Aided Engineering, i.e., software solutions for simulating the physical behavior of future products. CAM stands for Computer-Aided Manufacturing, i.e., software solutions for defining manufacturing processes and operations. In such computer-aided design systems, graphical user interfaces play a significant role in technical efficiency. These technologies may be embedded within product lifecycle management (PLM) systems. PLM refers to a business strategy that helps companies share product data, apply common processes, and leverage corporate knowledge for product development across the extended enterprise, from concept to end-of-life. PLM solutions offered by Dassault Systèmes (under the trademarks CATIA, ENOVIA, and DELMIA) provide an engineering center for organizing product engineering knowledge, a manufacturing center for managing manufacturing engineering knowledge, and an enterprise center for enabling enterprise integration and connectivity between engineering and manufacturing centers. The entire system provides an open object model that connects products, processes, and resources to enable dynamic, knowledge-based product creation and decision support that drives optimized product definition, manufacturing preparation, production, and service.

[0003] In this and other contexts, machine learning is becoming increasingly important.

[0004] The following papers are relevant to this area and are cited below:

[0005] [1] A. Gaidon, Q. Wang, Y. Cabon, and E. Vig, "Virtual worlds as proxy for multi-object tracking analysis," in CVPR, 2016.

[0006] [2]SRRichter, V.Vineet, S.Roth, and V.Koltun, "Playing for data: Groundtruth from computer games," in ECCV, 2016.

[0007] [3]M.Johnson-Roberson,C.Barto,R.Mehta,S.N.Sridhar,K.Rosaen,andR.Vasudevan,“Driving in the matrix:Can virtual worlds replace human-generatedannotations for real world tasks?”in ICRA,2017.

[0008] [4]S.R.Richter,Z.Hayder,and V.Koltun,“Playing for benchmarks,”inICCV,2017.

[0009] [5]S.Hinterstoisser,V.Lepetit,P.Wohlhart,and K.Konolige,“Onpretrained image features and synthetic images for deep learning,”in arXiv:1710.10710,2017.

[0010] [6]D.Dwibedi,I.Misra,and M.Hebert,“Cut,paste and learn:Surprisinglyeasy synthesis for instance detection,”in ICCV,2017.

[0011] [7]J.Tobin,R.Fong,A.Ray,J.Schneider,W.Zaremba,and P.Abbeel,“Domainrandomization for transferring deep neural networks from simulation to thereal world,”in IEEE / RSJ International Conference on Intelligent Robots andSystems(IROS),2017.

[0012] [8]J.Tremblay,A.Prakash,D.Acuna,M.Brophy,V.Jampani,C.Anil,T.To,E.Cameracci,S.Boochoon,and S.Birchfield,“Training deep networks withsynthetic data:Bridging the reality gap by domain randomization,”in CVPRWorkshop on Autonomous Driving(WAD),2018.

[0013] [9]K.He,G.Gkioxari,P.Dollar,and R.Girshick.“Mask r-cnn“.arXiv:1703.06870,2017.

[0014]

[10] J.Deng,A.Berg,S.Satheesh,H.Su,A.Khosla,and L.Fei-Fei.ILSVRC-2017,2017.URL http: / / www.image-net.org / challenges / LSVRC / 2017 / .

[0015]

[11] Y.Chen,W.Li,and L.Van Gool.ROAD:Reality oriented adaptation forsemantic segmentation of urban scenes.In CVPR,2018.

[0016]

[12] Y.Chen,W.Li,C.Sakaridis,D.Dai,and L.Van Gool.Domain adaptiveFaster R-CNN for object detection in the wild.In CVPR,2018.

[0017]

[13] Y.Zou,Z.Yu,B.Vijaya Kumar,and J.Wang.Unsupervised domainadaptation for semantic segmentation via class-balanced self-training.InECCV,September 2018.

[0018]

[14] A.Geiger,P.Lenz,and R.Urtasun,“Are we ready for autonomousdriving?The KITTI vision benchmark suite,”in CVPR,2012.

[0019]

[15] A.Prakash,S.Boochoon,M.Brophy,D.Acuna,E.Cameracci,G.State,O.Shapira,and S.Birchfield.Structured domain randomization:Bridging thereality gap by contextaware synthetic data.arXiv preprint arXiv:1810.10093,2018.

[0020]

[16] J.Deng,W.Dong,R.Socher,L.-J.Li,K.Li,and L.Fei-Fei.ImageNet:ALarge-Scale Hierarchical Image Database.In CVPR09,2009.

[0021]

[17] T.-Y.Lin,M.Maire,S.Belongie,J.Hays,P.Perona,D.Ramanan,P.Dollar,and C.L.Zitnick.Microsoft COCO:Common objects in′context.In ECCV.2014.

[0022]

[18] Shaoqing Ren,Kaiming He,Ross B.Girshick,and Jian Sun.Faster R-CNN:towards real-time object detection with region proposal networks.CoRR,abs / 1506.01497,2015.

[0023]

[19] D.E.Rumelhart,G.E.Hinton,R.J.Williams,Learning internalrepresentations by error propagation,Parallel distributed processing:explorations in the microstructure of cognition,vol.1:foundations,MIT Press,Cambridge,MA,1986.

[0024]

[20] Jun-Yan Zhu,Taesung Park,Phillip Isola,and AlexeiA.Efros.Unpaired image-to-image translation using cycle-consistentadversarial networks.CoRR,abs / 1703.10593,2017.

[0025]

[21] M.Heusel,H.Ramsauer,T.Unterthiner,B.Nessler,and S.Hochreiter.Ganstrained by a two time-scale update rule converge to a local nashequilibrium.In Advances in Neural Information Processing Systems(NIPS),2017.

[0026]

[22] C.Szegedy,W.Liu,Y.Jia,P.Sermanet,S.Reed,D.Anguelov,D.Erhan,V.Vanhoucke,and A.Rabinovich.Going deeper with convolutions.In Proceedings ofthe IEEE Conference on Computer Vision and Pattern Recognition,pages 1–9,2015.

[0027] Deep learning techniques have shown excellent performance in several areas such as 2D / 3D object recognition and detection, semantic segmentation, pose estimation, or motion capture (see [9, 10]). However, such techniques typically (1) require a large amount of training data to fully realize their potential and (2) often require expensive manual labeling (also called annotation):

[0028] (1) In practice, due to data scarcity or confidentiality reasons, it is difficult to collect hundreds of thousands of labeled data in several fields (e.g., manufacturing, health, or indoor motion capture). Publicly available datasets usually consist of objects in daily life that are manually labeled via crowdsourcing platforms (see [16, 17]).

[0029] (2) In addition to the data collection problem, labeling is also very time-consuming. The annotation of the well-known ImageNet dataset (see

[16] ) took several years, where each image was labeled with “only” a single label. Pixel-by-pixel labeling per image takes more than an hour on average. In addition, manual annotation may have some errors and lack accuracy. Some annotations require domain-specific know-how (e.g., manufacturing equipment labeling).

[0030] For all these reasons, leveraging virtual data to learn deep models has attracted increasing attention in recent years (see [1, 2, 3, 4, 5, 6, 7, 8, 15]). Virtual (also known as synthetic) data is computer-generated data created using CAD or 3D modeling software, rather than data generated by actual events. In practice, virtual data can be generated to meet very specific needs or conditions that are unavailable in existing (real) data. This can be useful when privacy requirements limit the availability or use of data, or when the data required for the test environment simply does not exist. Note that in the case of virtual data, labeling is free and error-free.

[0031] However, the inherent domain difference between virtual and real data (also known as the reality gap) can render the learned models incapable of generalizing to real-world scenarios. This is primarily due to overfitting, where the network learns detailed information that only exists in the virtual data, failing to generalize well and extracting informative representations of the real data.

[0032] One approach that has become popular recently, especially for recognition problems related to self-driving cars, is to collect photo-realistic virtual data that can be automatically annotated at low cost, such as video game data.

[0033] For example, Richter et al. (see [2]) built a large-scale synthetic urban scene dataset for semantic segmentation from the game GTA V. The GTA-based data (see [2, 3, 4]) uses a large number of assets to generate a variety of realistic driving environments that can be automatically labeled at the pixel level (about 25K images). Such photorealistic virtual data generally does not exist in domains other than driving and urban contexts. For example, existing virtual manufacturing environments are non-photorealistic CAD-based environments (video games do not exist for the manufacturing domain). Manufacturing specifications do not require photorealism. In addition, having shadows and lightening in such virtual environments can be confusing. Therefore, all work focuses on producing semantically and geometrically accurate CAD models rather than photorealistic CAD models. In addition, creating photorealistic manufacturing environments for training purposes is not relevant due to the large variability between the layouts of different workshops and factories. Unlike urban contexts, where urban scenes have strong idiosyncrasies (e.g., the size and spatial relationships of buildings, streets, cars, etc.), there is no "inherent" spatial structure for manufacturing scenes. Also note that some previous works (e.g., VKITTI (see [1])) create copies of real-world driving scenarios and are therefore highly correlated with those original real data and lack variability.

[0034] Furthermore, even with a high degree of realism, it is not obvious how to effectively use photorealistic data to train neural networks to operate on real data. Typically, cross-domain adaptation components are used during training neural networks (see [11, 12, 13]) to narrow the gap between virtual and real representations.

[0035] In summary, photo-realistic virtual datasets require either carefully designed simulation environments or the existence of annotated real data as a starting point, which is not feasible in many fields (e.g., manufacturing).

[0036] To alleviate this difficulty, Tobin et al. (see [7]) introduced the concept of domain randomization (DR), in which realistic rendering is avoided in favor of random variations. Their method randomly varies the texture and color of foreground objects, the background image, the number of lights in the scene, the pose of the lights, the camera position, and the foreground objects. The goal is to close the reality gap by generating virtual data with sufficient variation that the network sees the real-world data as another kind of variation. They used DR to train a neural network to estimate the 3D world position of various shape-based objects relative to a robotic arm fixed to a table. Recent work (see [8]) has shown that DR achieves state-of-the-art results for 2D bounding box detection of cars on the real-world KITTI dataset (see

[14] ). However, with such a large amount of variation, DR requires very large amounts of data to train (about 100K per object category), the network often finds it difficult to learn the right features, and the lack of context causes DR to fail to detect smaller objects or occluded objects. In practice, the results of this work are limited to (1) large cars that are (2) fully visible (KITTI Easy dataset). When used for complex inference tasks such as object detection, this approach requires that there are enough pixels within the bounding box for the learned model to make decisions without surrounding context. All of these shortcomings were quantitatively demonstrated in (see

[15] ).

[0037] To handle more challenging criteria (e.g., smaller cars, partial occlusions), the authors of

[15] proposed a variant of DR where they exploited urban context information and urban structure by randomly placing objects according to a probability distribution generated by the specific problem at hand (i.e., car detection in an urban context) rather than the uniform probability distribution used in previous DR-based works.

[0038] Similarly, DR-based methods (see [5, 6, 7, 15]) have mainly focused on “easy” object classes with little intra-class variability (e.g., the “car” class). Significant intra-class variability can cause these methods to significantly degrade in performance, as the amount of variation to be considered during data creation can grow exponentially. This is the case for spatially reconstructable objects, such as articulated objects.

[0039] Spatially reconfigurable objects, or articulated objects, are objects that have components attached via joints and that can move relative to each other, which can give each object type an infinite range of possible states.

[0040] The same problem may arise while dealing with multiple classes of objects (inter-class and intra-class variability).

[0041] Within this context, there remains a need for improved machine learning methods using scene datasets that include spatially reconstructible objects. Summary of the Invention

[0042] Therefore, a computer-implemented machine learning method is provided. The method includes providing a dataset of a virtual scene. The dataset of the virtual scene belongs to a first domain. The method also includes providing a test dataset of a real scene. The test dataset belongs to a second domain. The method also includes determining a third domain. The third domain is closer to the second domain than the first domain in terms of data distribution. The method also includes learning a domain adaptive neural network based on the third domain. The domain adaptive neural network is a neural network configured to infer spatially reconstructable objects in a real scene.

[0043] The method may include one or more of the following:

[0044] -For each scene in the virtual scene dataset, determining the third domain includes:

[0045] Extracting one or more spatially reconstructable objects from the scene; and

[0046] For each extracted object, transform the extracted object into an object that is closer to the second domain than the extracted object in terms of data distribution;

[0047] - The determination of the third domain also includes:

[0048] for each transformed extracted object, placing the transformed extracted object in one or more scenes, each scene being closer to the second domain than the first domain in terms of data distribution, a third domain including each scene on which the transformed extracted object was placed;

[0049] -Placing the transformed extracted objects in one or more scenes is performed randomly;

[0050] - determining the third domain comprises placing one or more distractors in each respective scene of the one or more scenes comprised in the third domain, the distractors being objects that are not spatially reconstructable objects;

[0051] -Learning about domain adaptive neural networks includes:

[0052] Providing a teacher extractor, which is a machine-learned neural network configured to output an image representation of a real scene;

[0053] training a student extractor, which is a neural network configured to output an image representation of a scene belonging to the third domain, wherein training the student extractor comprises minimizing a loss,

[0054] The loss penalizes, for each of one or more true scenes, the difference between the result of applying the teacher extractor to the scene and the result of applying the student extractor to the scene;

[0055] - the result of applying the teacher extractor to the scene is a first Gramian matrix, and the result of applying the student extractor to the scene is a second Gramian matrix;

[0056] - calculating a first Gramian matrix over several neuron layers of the teacher extractor, said neuron layers including at least the last neuron layer, and calculating a second Gramian matrix over several neuron layers of the student extractor, said neuron layers including at least the last neuron layer;

[0057] - the discrepancy is the Euclidean distance between the result of applying the teacher extractor to the scene and the result of applying the student extractor to the scene;

[0058] - each virtual scene of the virtual scene dataset is a virtual manufacturing scene comprising one or more spatially reconfigurable manufacturing tools, the domain adaptive neural network being configured to infer the spatially reconfigurable manufacturing tools in a real manufacturing scene; and / or

[0059] - Each real scene of the test dataset is a real manufacturing scene including one or more spatially reconfigurable manufacturing tools, and the domain adaptation neural network is configured to infer the spatially reconfigurable manufacturing tools in the real manufacturing scene.

[0060] We further propose a domain-adaptive neural network that can be learned according to this method.

[0061] A computer program comprising instructions for performing the method is also provided.

[0062] Also provided is an apparatus comprising a data storage medium having recorded thereon a domain adaptive neural network and / or a computer program.

[0063] The device may form or act as a non-transitory computer-readable medium, such as on a SaaS (Software as a Service) or other server, or on a cloud-based platform, etc. The device may alternatively include a processor coupled to a data storage medium. Thus, the device may form a computer system in whole or in part (e.g., the device is a subsystem of the entire system). The system may also include a graphical user interface coupled to the processor. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] Embodiments of the invention will now be described by way of non-limiting examples and with reference to the accompanying drawings, in which:

[0065] - Figure 1A flowchart illustrating an example computer-implemented process incorporating the method;

[0066] - Figures 2 to 4 A flowchart illustrating an example of a computer-implemented process and / or an example of a method is shown;

[0067] - Figure 5 An example of a graphical user interface for the system is shown;

[0068] - Figure 6 An example of a system is shown; and

[0069] - Figures 7 to 32 A process and / or method is shown. DETAILED DESCRIPTION

[0070] A computer-implemented machine learning method is provided that utilizes a dataset of a scene including spatially reconstructable objects.

[0071] In particular, a first machine learning method is provided. The first method includes providing a dataset of a virtual scene. The dataset of the virtual scene belongs to a first domain. The first method also includes providing a test dataset of a real scene. The test dataset belongs to a second domain. The first method also includes determining a third domain. In terms of data distribution, the third domain is closer to the second domain than the first domain. The first method also includes learning a domain adaptive neural network based on the third domain. The domain adaptive neural network is a neural network configured to infer spatially reconstructable objects in a real scene. Hereinafter, this first method may be referred to as a "domain adaptive learning method."

[0072] Domain adaptation learning methods form an improved machine learning approach for exploiting scene datasets that include spatially reconfigurable objects.

[0073] Notably, the domain adaptation learning method provides a domain inference neural network that is capable of inferring spatially reconfigurable objects in real scenes. This allows the use of the domain adaptation learning method in the process of inferring spatially reconfigurable objects in real scenes, such as inferring spatially reconfigurable manufacturing tools (e.g., industrial articulated robots) in real manufacturing scenes.

[0074] Furthermore, it may not be possible to directly learn a domain adaptive neural network on a dataset of real scenes, but the domain adaptive neural network can still infer actual spatially reconstructable objects in the real scene. This improvement is achieved, in particular, by determining a third domain according to the domain adaptive learning method. In practice, because it is composed of virtual scenes, the first domain may be relatively far from the second domain (i.e., the domain of the test dataset of the domain adaptive neural network, which is composed of real scenes) in terms of data distribution. This reality gap may make learning a domain adaptive neural network based directly on a dataset of virtual scenes difficult and / or inaccurate. Such learning can indeed result in a domain adaptive neural network that may produce inaccurate results in inferring spatially reconstructable objects in real scenes. In contrast, the domain adaptive learning method determines a third domain that is closer to the second domain than the first domain in terms of data distribution, and bases the learning of the domain adaptive neural network on this third domain. This makes it possible for the learning of the domain adaptive neural network to accurately infer spatially reconstructable objects in the real scene, even when the only dataset provided (i.e., given as input to the domain adaptive learning method) is actually a dataset of virtual scenes. In other words, both the learning and learned domain adaptive neural networks are more robust.

[0075] This is equivalent to saying that the dataset of virtual scenes is theoretically intended to be the training set (or learning set) for the domain adaptive neural network, and that the domain adaptive learning method includes: before actual learning, a preprocessing stage of the training set (i.e., determining the third domain) in order to learn the domain adaptive neural network with the improvements discussed previously. In other words, the true training / learning domain of the domain adaptive neural network is the third domain, and it is derived from the preprocessing of the domain originally intended as the training / learning domain of the domain adaptive neural network (i.e., the first domain). However, relying on a dataset of virtual scenes as the theoretical training set (i.e., processed before use for training, as described above) allows the use of datasets with a large number of scenes. In practice, there may not be, or at least not be many, large datasets of real scenes that include spatially reconfigurable objects, such as manufacturing scenes that include spatially reconfigurable manufacturing tools. Such datasets can indeed be difficult to obtain due to privacy / confidentiality issues: for example, manufacturing scenes that include spatially reconfigurable manufacturing tools (e.g., robots) are often confidential and therefore not publicly available. Furthermore, even if such a large dataset of real scenes would be available, it would likely contain many annotation (e.g., labeling) errors: annotating spatially reconstructable objects (e.g., manufacturing tools, such as robots) may indeed be difficult for a person who does not possess the appropriate skills for it, because spatially reconstructable objects may have so many and / or such complex positions that their identification in a real scene and their manual annotation would require a certain amount of skill and knowledge. On the other hand, large datasets of virtual scenes can be easily obtained, for example by using simulation software to easily generate virtual scenes. Furthermore, annotation of objects in virtual scenes (e.g., spatially reconstructable objects) is relatively easy and, most importantly, can be performed automatically without errors. In fact, simulation software usually knows the specifications of the objects relevant to the simulation, so their annotation can be performed automatically by the software, and the risk of annotation errors is low.

[0076] In an example, each virtual scene in the dataset of virtual scenes is a virtual manufacturing scene that includes one or more spatially reconfigurable manufacturing tools. In these examples, the domain adaptation neural network is configured to infer spatially reconfigurable manufacturing tools in a real manufacturing scene. Additionally or alternatively, each real scene in the test dataset is a real manufacturing scene that includes one or more spatially reconfigurable manufacturing tools. In this case, the domain adaptation neural network is configured to infer spatially reconfigurable manufacturing tools in the real manufacturing scene.

[0077] In such examples, in any case, a domain adaptive neural network is configured to infer spatially reconfigurable manufacturing tools (e.g., spatially reconfigurable manufacturing robots) in real manufacturing scenarios. In these examples, the previously discussed improvements brought about by domain adaptive learning methods are particularly emphasized. In particular, it is difficult to rely on a large number of training sets of real manufacturing scenarios that include spatially reconfigurable manufacturing tools, because such sets are often confidential and / or publicly available but not adequately annotated. On the other hand, it is particularly convenient to provide a large dataset of virtual manufacturing scenarios that include spatially reconfigurable manufacturing tools as a training set, because such datasets can be easily obtained and automatically annotated, for example by using software that can generate virtual simulations in virtual manufacturing environments.

[0078] A domain adaptive neural network that can be learned according to a domain adaptive learning method is also provided.

[0079] A second computer-implemented machine learning method is also provided. The second method includes providing a test data set of a scene. The test data set belongs to a test domain. The second method includes providing a domain adaptive neural network. The domain adaptive neural network is a neural network for machine learning based on data learned from a training domain. The domain adaptive neural network is configured to infer spatially reconstructable objects in the test domain scene. The second method also includes determining an intermediate domain. The intermediate domain is closer to the training domain in terms of data distribution than the test domain. The second method also includes inferring spatially reconstructable objects from the scene of the test domain transmitted on the intermediate domain by applying the domain adaptive neural network. This second method can be referred to as a "domain adaptive inference method."

[0080] Domain adaptive inference methods form an improved machine learning approach that exploits scene datasets that include spatially reconstructable objects.

[0081] Notably, domain-adaptive inference methods allow for inferring spatially configurable objects from a scene. This allows for the use of domain-adaptive inference methods in inferring spatially reconfigurable objects, such as in inferring spatially reconfigurable manufacturing tools (e.g., industrial articulated robots) in manufacturing scenarios.

[0082] Furthermore, the method allows accurate inference of spatially reconstructable objects from the scenes of the test dataset, while the domain adaptive neural network used for such inference has already been machine-learned based on data obtained from the training domain, which may be far from the test domain (to which the test dataset belongs) in terms of data distribution. This improvement is achieved in particular by determining an intermediate domain. In fact, in order to better utilize the domain adaptive neural network, before applying the domain adaptive neural network for inference, the domain adaptive inference method transforms the scenes of the test domain into scenes belonging to the intermediate domain, in other words, scenes that are closer to the learning scenes of the domain adaptive neural network in terms of data distribution. In other words, the domain adaptive inference method allows the domain adaptive neural network to be fed not directly with scenes of the test domain as input, but with scenes that are closer to the scenes on which the domain adaptive neural network has been learned, so as to improve the accuracy and / or quality of the outputs of the domain adaptive neural network (which are each one or more spatially reconstructable objects inferred in the corresponding scenes of the test domain).

[0083] Therefore, the improvements brought by the domain adaptive inference method are particularly emphasized in examples where the test domain and the training domain are relatively far from each other. For example, this may be the case when the training domain consists of virtual scenes and the test domain consists of real scenes. For example, this may also be the case when both the training domain and the test domain consist of virtual scenes, but the scenes of the test domain are more photorealistic than those of the training domain. In all of these examples, the method includes a stage of preprocessing the scenes of the test domain (i.e., determining an intermediate domain), which are provided as input to the domain adaptive neural network (e.g., in the testing stage of the domain adaptive neural network) in order to make them closer to the scenes of the training domain in terms of photorealism (which can be quantified in terms of data distribution). As a result, the photorealism level of the scenes provided as input to the domain adaptive neural network is close to the photorealism of the scenes on which the domain adaptive neural network has been learned. This ensures better output quality of the domain adaptive neural network (i.e., the inferred spatially reconstructible objects inferred on the scenes of the test domain).

[0084] In an example, the training domain includes a training dataset of virtual scenes, each virtual scene including one or more spatially reconstructable objects. Additionally, each virtual scene of the dataset of virtual scenes may be a virtual manufacturing scene including one or more spatially reconstructable manufacturing tools, in which case the domain adaptation neural network is configured to infer the spatially reconstructable manufacturing tools in the real manufacturing scene. Additionally or alternatively, the test dataset includes real scenes. In an example of this case, each real scene of the test dataset is a real manufacturing scene including one or more spatially reconstructable manufacturing tools, and the domain adaptation neural network is configured to infer the spatially reconstructable manufacturing tools in the real manufacturing scene.

[0085] In these examples, domain adaptive neural networks can be configured to infer spatially reconfigurable manufacturing tools (e.g., spatially reconfigurable manufacturing robots) in real-world manufacturing scenarios. In these examples, the improvements previously achieved through domain adaptive inference methods are particularly emphasized. In particular, it is difficult to rely on training sets of large numbers of real-world manufacturing scenarios that include spatially reconfigurable manufacturing tools, as such sets are often confidential and / or publicly available but insufficiently annotated. On the other hand, using large datasets of virtual manufacturing scenarios that include spatially reconfigurable manufacturing tools as training sets is particularly convenient, as such datasets can be easily obtained and automatically annotated, for example, using software that can generate virtual simulations in virtual manufacturing environments. For at least these reasons, domain adaptive neural networks may have been learned on training sets of virtual manufacturing scenarios. Therefore, the preprocessing of the test dataset discussed above is performed by determining an intermediate domain, thereby considering the training set of the domain adaptive neural network by cleverly making the real-world manufacturing scenarios of the test dataset closer to those of the test scenarios in terms of data distribution, so as to ensure better output quality of the domain adaptive neural network.

[0086] The domain adaptive inference method and the domain adaptive learning method can be performed independently. In particular, the domain adaptive neural network provided according to the domain adaptive inference method may have been learned according to the domain adaptive learning method or by any other machine learning method. Providing the domain adaptive neural network according to the domain adaptive inference method may include, for example, remotely accessing a data storage medium over a network, on which the domain adaptive neural network is stored after it has been learned (e.g., according to the domain adaptive learning method) and retrieved from a database.

[0087] In an example, the data obtained from the training domain includes a scene of another intermediate domain. In these examples, a domain adaptive neural network has been learned on another intermediate domain. In terms of data distribution, the other intermediate domain is closer to the intermediate domain than the training domain. In such an example, the domain adaptive neural network can be a domain adaptive neural network that can be learned according to a domain adaptive learning method. The test data set according to the domain adaptive learning method can be equal to the test data set according to the domain adaptive inference method, and the test domain according to the domain adaptive inference method can be equal to the second domain according to the domain adaptive learning method. The training domain according to the domain adaptive inference method can be equal to the first domain according to the domain adaptive learning method, and the data obtained from the training domain can belong to (e.g., form) a third domain determined according to the domain adaptive learning method. The other intermediate domain can be a third domain determined according to the domain adaptive learning method.

[0088] These examples combine the improvements brought by the previously discussed domain adaptation learning methods and domain adaptation inference methods: using a large dataset of correctly annotated virtual scenes as a training set for the domain adaptation neural network, a preprocessing stage of the training set performed by determining a third domain (or another intermediate domain), which improves the quality and / or accuracy of the output of the domain adaptation neural network, and performing a test dataset of real scenes by determining an intermediate domain, which further improves the quality and / or accuracy of the output of the domain adaptation neural network.

[0089] The domain adaptive learning method and the domain adaptive inference method may alternatively be integrated into the same computer-implemented process. Figure 1 A flow chart illustrating this process is shown and will now be discussed.

[0090] The process includes an offline phase that integrates a domain adaptive learning method. The offline phase includes providing a dataset of S10 virtual scenes and a test dataset of real scenes according to the domain adaptive learning method. The dataset of the virtual scene belongs to a first domain, which forms (in theory, as described above) a training domain of a domain adaptive neural network learned according to the domain adaptive learning method. The test dataset of the real scene belongs to a second domain, which forms a test domain of a domain adaptive neural network learned according to the domain adaptive learning method. The offline phase also includes determining S20 a third domain according to the domain adaptive learning method. In terms of data distribution, the third domain is closer to the second domain than the first domain. The offline phase also includes learning 30 of the domain adaptive neural network according to the domain adaptive learning method.

[0091] After the offline phase, the domain adaptive neural network is processed according to the domain adaptive learning method. In other words, the domain adaptive neural network may form the output of the offline phase. The offline phase may be followed by a phase of storing the learned domain adaptive neural network, for example on a data storage medium of the device.

[0092] The process also includes an online phase that integrates a domain adaptive inference method. The online phase includes providing S40 the domain adaptive neural network learned during the offline phase according to the domain adaptive inference method. It is worth noting that providing S40 the domain adaptive neural network may include, for example, remotely accessing, for example, a data storage medium (on which the domain adaptive neural network was stored at the end of the offline phase) via a network and retrieving the domain adaptive neural network from a database. The online phase also includes determining S50 an intermediate domain according to the domain adaptive inference method. In terms of data distribution, the intermediate domain is closer to the first domain than the second domain. It will be understood that in the context of this process: the first domain of the domain adaptive learning method is the training domain of the domain adaptive inference method, the second domain of the domain adaptive learning method is the test domain of the domain adaptive inference method, the test dataset of the domain adaptive learning method is the test dataset of the domain adaptive inference method, and the data obtained from the training domain of the domain adaptive inference method belongs to (e.g., in the form of) a third domain. The online phase may also include inferring S60 a spatially reconstructable object according to the domain adaptive inference method.

[0093] This process combines the improvements brought by the domain adaptation learning method and the domain adaptation inference method discussed previously: using a large dataset of correctly annotated virtual scenes as the training set for the domain adaptation neural network, performing a preprocessing stage of the training set by determining a third domain, which improves the quality and / or accuracy of the output of the domain adaptation neural network, and performing a preprocessing stage of the test dataset of real scenes by determining an intermediate domain, which further improves the quality and / or accuracy of the output of the domain adaptation neural network.

[0094] This process can constitute a process for inferring spatially reconfigurable manufacturing tools (e.g., spatially reconfigurable manufacturing robots) in real-world manufacturing scenarios. In fact, in examples of this process, the virtual scene of the first domain can be each virtual manufacturing scenario that includes one or more spatially reconfigurable manufacturing tools, and the real scene of the second domain can be each real-world manufacturing scenario that includes one or more spatially reconfigurable manufacturing tools. Therefore, in these examples, the domain adaptive neural network learned during the offline phase is configured to infer spatially reconfigurable manufacturing tools in real-world manufacturing scenarios. The offline phase benefits from the previously discussed improvements provided by the domain adaptive learning method in the specific case of manufacturing tool inference: virtual manufacturing scenarios are readily available in large quantities and correctly annotated, while the previously discussed preprocessing of the training set can make the training set scenario closer to the real-world scenario of the test dataset, which improves the robustness and accuracy of the learned domain adaptive neural network. Combined with these improvements are improvements brought about by the domain adaptive inference method integrated into the online phase of the process. These improvements include the previously discussed preprocessing of the test dataset to make the real-world scenario of the test dataset closer to the scenario of the third domain, on which the domain adaptive neural network was learned during the offline phase, thereby improving the quality of the domain adaptive neural network output. As a result, inferring S60 a spatially reconfigurable manufacturing tool in a real manufacturing scenario in an online phase is particularly accurate and robust.Thus, the process may constitute a particularly robust and accurate process for inferring a spatially reconfigurable manufacturing tool in a real manufacturing scenario.

[0095] The domain adaptive learning method, the domain adaptive inference method and the process are computer-implemented. The concept of a computer-implemented method (or process) is now discussed.

[0096] "The method (or process) is computer-implemented" means that the steps (or substantially all steps) of the method (or process) are performed by at least one computer or any similar system. Thus, the steps of the method (or process) are performed by a computer, which may be fully or semi-automatically performed. In an example, the triggering of at least some steps of the method (or process) may be performed by user-computer interaction. The required level of user-computer interaction may depend on the level of automation envisioned and balanced with the need to implement the user's wishes. In an example, the level may be user-defined and / or predefined.

[0097] A typical example of a computer-implemented method (or process) is to perform the method (or process) using a system suitable for this purpose. The system may include a processor coupled to a memory and a graphical user interface (GUI), the memory having recorded thereon a computer program including instructions for performing the method (or process). The memory may also store a database. The memory is any hardware suitable for such storage and may include several physically distinct parts (e.g., one for the program and one for the database).

[0098] Figure 5 An example of a GUI for a system is shown, where the system is a CAD system.

[0099] GUI 2100 may be a typical CAD-like interface, having standard menu bars 2110 and 2120 and bottom and side toolbars 2140 and 2150. Such menus and toolbars contain a set of user-selectable icons, each associated with one or more operations or functions as known in the art. Some of these icons are associated with software tools suitable for editing and / or working on the 3D modeled object 2000 displayed in GUI 2100. Software tools may be grouped into workbenches. Each workbench contains a subset of software tools. In particular, one of the workbenches is an editing workbench suitable for editing the geometric features of the modeled product 2000. In operation, a designer may, for example, pre-select a portion of the object 2000 and then initiate an operation (e.g., changing the size, color, etc.) or edit geometric constraints by selecting the appropriate icon. For example, a typical CAD operation is modeling the punching or folding of a 3D modeled object displayed on the screen. The GUI may, for example, display data 2500 related to the displayed product 2000. In the example of the figure, data 2500 displayed as a "feature tree" and its 3D representation 2000 relate to a brake assembly including a brake caliper and a disc. The GUI may also display various types of graphical tools 2130, 2070, 2080, such as for facilitating 3D orientation of an object, for triggering simulation of an operation of the edited product, or for presenting various properties of the displayed product 2000. A cursor 2060 may be controlled by a haptic device to allow the user to interact with the graphical tools.

[0100] Figure 6 An example of the system is shown, where the system is a client computer system, such as a user's workstation.

[0101] The client computer of this example includes a central processing unit (CPU) 1010 connected to an internal communication bus 1000, and random access memory (RAM) 1070 also connected to the bus. The client computer is also provided with a graphics processing unit (GPU) 1110, which is associated with video random access memory 1100 connected to the bus. Video RAM 1100 is also known in the art as a frame buffer. A mass storage device controller 1020 manages access to mass storage devices (e.g., hard drive 1030). Mass storage devices suitable for tangibly embodying computer program instructions and data include all forms of non-volatile memory, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM disks 1040. Any of the above may be supplemented by or incorporated into specially designed ASICs (application-specific integrated circuits). A network adapter 1050 manages access to a network 1060. The client computer may also include haptic devices 1090, such as cursor control devices, keyboards, and the like. A cursor control device is used in the client computer to allow the user to selectively position the cursor at any desired location on the display 1080. Furthermore, the cursor control device allows the user to select various commands and input control signals. The cursor control device includes multiple signal generating devices for inputting control signals into the system. Typically, the cursor control device may be a mouse having buttons for generating signals. Alternatively or additionally, the client computer system may include a sensitive pad and / or a sensitive screen.

[0102] The computer program may include computer-executable instructions, including means for causing the aforementioned system to perform a domain-adaptive learning method, a domain-adaptive inference method, and / or a process. The program may be recordable on any data storage medium, including the system's memory. The program may be implemented, for example, in digital electronic circuitry or computer hardware, firmware, software, or a combination thereof. The program may be implemented as an apparatus, such as a product tangibly embodied in a machine-readable storage device for execution by a programmable processor. The processing / method steps (i.e., the steps of the domain-adaptive learning method, the domain-adaptive inference method, and / or the process) may be performed by programming a processor to execute a program of instructions to perform the functions of the process by operating on input data and generating output. Thus, the processor may be programmable and coupled to receive data and instructions from a data storage system, at least one input device, and at least one output device, and to send data and instructions to the data storage system, at least one input device, and at least one output device. If desired, the application may be implemented in a high-level procedural or object-oriented programming language or in assembly or machine language. In any case, the language may be a compiled or interpreted language. The program may be a complete installer or updater. In any case, application of the program to the system results in instructions for performing domain adaptive learning methods, domain adaptive inference methods, and / or processes.

[0103] The concept of providing a scene dataset will be discussed, involving domain adaptive learning methods, domain adaptive inference methods, and processes. Before discussing this concept, the data structures involved will now be discussed. As will be appreciated, in the context of providing domain adaptive learning methods, domain adaptive inference methods, and / or processes, the data structure definitions and examples provided herein can be applied to at least a portion (e.g., all) of any dataset.

[0104] In the context of the present disclosure, a scene specifies the arrangement (e.g., disposition) of one or more objects in a background at a point in time. The background typically represents the physical environment of the real world, the objects typically represent physical objects of the real world, and the arrangement of the objects typically represents the deployment of the real-world objects in the real-world environment at that point in time. This is equivalent to saying that a scene is a representation of a real-world object in its real-world environment at a point in time. The representation can be real, geometrically real, semantically real, photo-realistic, and / or virtual. The concepts of real, geometrically real, semantically real, photo-realistic, and virtual scenes will be discussed below. Scenes are typically (but not always) three-dimensional. In the following, for simplicity, no distinction will be made between scenes (e.g., virtual or real) and their representations (e.g., represented by virtual or real images).

[0105] In the context of the present disclosure, a dataset of scenes typically includes a large number of scenes, for example, more than 1,000, 10,000, 100,000, or 1,000,000 scenes. Any dataset of scenes disclosed herein may consist of or substantially consist of scenes from the same context (e.g., a manufacturing scene, a construction site scene, a port scene, or an apron scene). The scenes involved in the domain adaptive learning method, domain adaptive inference method, and / or process may actually be all scenes from the same context (e.g., a manufacturing scene, a construction site scene, a port scene, or an apron scene).

[0106] Any scene involved in the domain adaptive learning method, domain adaptive inference method and / or process may include one or more spatially reconfigurable objects. A spatially reconfigurable object is an object comprising several (e.g., mechanical) parts, which is characterized by the presence of one or more semantically realistic spatial relationships between the parts of the object. A spatially reconfigurable object can typically be an assembly of (e.g., mechanical) parts physically connected to each other, the assembly having at least one degree of freedom. This means that at least a first part of the object is movable relative to at least a second part of the object, while both the at least first part and the at least second part remain physically attached to the assembly. For example, at least the first part can translate and / or rotate relative to at least the second part. Therefore, a spatially reconfigurable object can have a very large number of spatial configurations and / or positions, which makes the object difficult to identify, for example, difficult to label, for people without appropriate skills. In the context of the present disclosure, a spatially reconfigurable object can be a spatially reconfigurable manufacturing tool (e.g., an industrial articulated robot) in a manufacturing scenario, a crane in a construction site scenario, an aerial walkway in a helipad scenario, or a port crane in a port scenario.

[0107] In the context of the present disclosure, a scene may be a manufacturing scene. A manufacturing scene is a scene that represents real-world objects in a real-world manufacturing environment. The manufacturing environment may be a factory or a part of a factory. Objects in the factory or a part of a factory may be manufactured products, i.e., products that have been manufactured by one or more manufacturing processes performed at the factory at a point in time that is relevant to the scene. Objects in the factory or a part of a factory may also be products being manufactured, i.e., products manufactured by one or more manufacturing processes performed at the factory at a point in time that is relevant to the scene. Objects in the factory or a part of a factory may also be (e.g., spatially reconstructable) manufacturing tools, which are tools involved in one or more manufacturing processes performed at the factory. Objects in the factory may also be objects that constitute the context of the factory.

[0108] In other words, a scene representing a factory or part of a factory may therefore include one or more products (each of which is being manufactured or has been manufactured by one or more manufacturing processes of the factory) and one or more manufacturing tools, each of which is involved in the manufacture of one or more of the products. Products may typically be (e.g., mechanical) parts or assemblies of parts (or equivalently, assemblies of parts, since assemblies of parts may be considered parts themselves from the perspective of this disclosure).

[0109] The mechanical part can be part of a land vehicle (including, for example, automobile and light truck equipment, racing cars, motorcycles, trucks and motor vehicle equipment, trucks and buses, trains), part of an aircraft (including fuselage equipment, aerospace equipment, propulsion equipment, defense products, aviation equipment, aerospace equipment), part of a naval vehicle (including naval equipment, commercial ships, offshore equipment, yachts and workboats, ship equipment), general mechanical parts (including, for example, industrial manufacturing machinery, heavy mobile machinery or equipment, mounted equipment, industrial equipment products, metal products, tire products, articulated and / or remanufacturing manufacturing equipment (e.g., robotic arms)), electromechanical or electronic parts (including, for example, consumer electronics, safety and / or control and / or instrumentation products, computing and communications equipment, semiconductors, medical devices and equipment), consumer products (including, for example, furniture, home and garden products, leisure products, fashion products, products of hard goods retailers, products of soft goods retailers), packaging (including food and beverage and tobacco, beauty and personal care, household product packaging).

[0110] The mechanical part may also be one of the following or any possible combination thereof: a molded part (i.e., a part manufactured by a molding manufacturing process), a machined part (i.e., a part manufactured by a machining manufacturing process), a drilled part (i.e., a part manufactured by a drilling manufacturing process), a turned part (i.e., a part manufactured by a turning manufacturing process), a forged part (i.e., a part manufactured by a forging manufacturing process), a stamped part (i.e., a part manufactured by a stamping manufacturing process) and / or a folded part (i.e., a part manufactured by a folding manufacturing process).

[0111] A manufacturing tool can be any of the following:

[0112] - machining tools (i.e., tools that perform at least part of a machining process), such as broaching machines, drilling machines, gear shaping machines, gear hobbing machines, grindstones, lathes, screw machines, milling machines, sheet metal shears), shaping machines, saws, planers, Stewart platform milling machines, grinders, multi-tasking machines (e.g., machines with multiple axes that combine turning, milling, grinding, and / or material handling into a highly automated machine tool);

[0113] - a compression molding machine (i.e., a machine that performs at least a portion of the compression molding process), typically comprising at least one mold or molding matrix, such as a rapid plunger mold, a straight plunger mold, or a floor plunger mold;

[0114] - an injection molding machine (i.e. a machine that performs at least a part of the injection molding process), such as a die casting machine, a metal injection molding machine, a plastic injection molding machine, a liquid silicon injection molding machine or a reaction injection molding machine;

[0115] - a rotating cutting tool (e.g., performing at least part of the drilling process), such as a drill bit, a countersink, a countersink, a tap, a die, a milling cutter, a reamer, or a cold saw blade;

[0116] - a non-rotating cutting tool (e.g., performing at least part of the turning process), such as a pointed tool or a forming tool;

[0117] a forging machine (i.e., performing at least a part of the forging process), such as a mechanical forging press (usually comprising at least one forging die), a hydraulic forging press (usually comprising at least one forging die), a forging hammer driven by compressed air, a forging hammer driven by electricity, a forging hammer driven by a hydraulic system, or a forging hammer driven by steam;

[0118] - a compression molding machine (i.e., performing at least a part of the compression molding process), such as a compression molding machine, a machine press, a punching press, a blanking machine, an embossing machine, a bending machine, a flanging machine, or a die-casting machine;

[0119] - bending or folding machines (i.e., machines that perform at least part of the bending or folding process), such as box and disc brakes, press brakes, folders, sheet metal bending machines or presses;

[0120] - a spot welding robot; or

[0121] - An electric painting tool (ie, performs at least a part of the product painting process), such as a (eg, electric) car painting robot.

[0122] Any of the manufacturing tools listed above may be a spatially reconfigurable manufacturing tool, i.e., a spatially reconfigurable manufacturing tool, or may form at least a portion of a spatially reconfigurable manufacturing tool. For example, a manufacturing tool held by a tool holder of a spatially reconfigurable manufacturing tool forms at least a portion of the spatially reconfigurable manufacturing tool. A spatially reconfigurable manufacturing tool (e.g., any of the tools listed above) may be a manufacturing robot, such as an (e.g., articulated) industrial robot.

[0123] Now let's discuss the concept of a real scene. As previously mentioned, a scene represents the arrangement of real-world objects in a real-world environment (i.e., the background of the scene). A scene is real when the representation of the arrangement is a physical arrangement derived from an acquisition of the same physical arrangement as it would be in the real world. This acquisition can be performed by a digital image acquisition process that creates a digitally encoded representation of the visual features of the physical arrangement of the objects in their environment. For simplicity, this digitally encoded representation may be referred to as an "image" or "digital image" below. The digital image acquisition process may optionally include processing, compression, and / or storage of this image. A digital image can be created directly from the physical environment using a camera or similar device, a scanner or similar device, and / or a video camera or similar device. Alternatively, a digital image can be obtained from another image in a similar medium, such as a photograph, photographic film (e.g., a digital image is a snapshot of the film at an instant), or printed paper, using an image scanner or similar device. A digital image can also be obtained by processing non-image data, such as non-image data collected using tomography equipment, scanners, side-scan sonar, radio telescopes, X-ray detectors, and / or photostimulable phosphor plates (PSPs). In the context of the present disclosure, a real manufacturing scene may be derived from a digital image obtained by a camera (or similar device) located in a factory, for example. Alternatively, the manufacturing scene may be derived from a snapshot of a film obtained by one or more cameras located in a factory.

[0124] Figures 7 to 12 Examples of real manufacturing scenarios are shown, each scenario including articulated industrial robots 70 and 72, 80, 90, 92, 94, 96 and 98, 100, 102 and 104, 110 and 112, 120, 122 and 124, respectively.

[0125] Let's now discuss the concept of a virtual scene. A scene is virtual when it originates from computer-generated data. Such data can be generated by (e.g., 3D) simulation software, (e.g., 3D) modeling software, CAD software, or a video game. This is equivalent to saying that the scene itself is computer-generated. Objects in a virtual scene can be placed in realistic locations within the virtual scene. Alternatively, objects in a virtual scene can be placed in unrealistic locations within the virtual scene, such as when the objects are randomly placed within an existing virtual scene, as discussed in more detail below. This is equivalent to saying that a virtual scene represents the arrangement of virtual objects (e.g., virtual representations of real-world objects) within a virtual background (e.g., a virtual representation of a real-world environment). For at least a portion of the virtual objects, the arrangement can be physically realistic, i.e., the arrangement of at least a portion of the virtual objects represents the real-world physical arrangement of real-world objects. Additionally or alternatively, at least another portion of the virtual objects can be arranged (e.g., relative to each other and / or relative to the at least portion of the objects) in an unrealistic arrangement (i.e., not corresponding to a real-world physical arrangement), such as when the at least portion of the virtual objects are randomly placed or arranged, as discussed in more detail below. Any object in the virtual scene can be automatically annotated or marked. The concept of marked objects will be discussed further below. The virtual manufacturing scene can be generated by 3D modeling software for the manufacturing environment, 3D simulation software for the manufacturing environment, or CAD software. Any object in the virtual manufacturing scene (e.g., a product or a manufacturing tool) can be virtually represented by a 3D modeling object representing the skin (e.g., outer surface) of an entity (e.g., a B-rep model). The 3D modeling object may have been designed by a user using a CAD system and / or may be derived from a virtual manufacturing simulation performed by 3D simulation software for a manufacturing environment.

[0126] A modeling object is any object defined by data stored, for example, in a database. By extension, the expression "modeling object" refers to the data itself. Depending on the type of system that performs the domain adaptive learning method, domain adaptive inference method and / or process, the modeling object can be defined by different kinds of data. The system can indeed be any combination of a CAD system, a CAE system, a CAM system, a PDM system and / or a PLM system. In those different systems, the modeling object is defined by the corresponding data. Thus, one can mention CAD objects, PLM objects, PDM objects, CAE objects, CAM objects, CAD data, PLM data, PDM data, CAM data, CAE data. However, these systems are not mutually exclusive, as a modeling object can be defined by data corresponding to any combination of these systems. It will be apparent from the definition of such a system provided below that the system can therefore also be a CAD and PLM system.

[0127] A CAD system additionally means any system, such as CATIA, that is at least suitable for designing a modeled object based on a graphical representation of the modeled object. In this case, the data defining the modeled object include data that allow the modeled object to be represented. A CAD system can, for example, use edges or lines (in some cases faces or surfaces) to provide a representation of a CAD modeled object. Lines, edges or surfaces can be represented in various ways, such as non-uniform rational B-splines (NURBS). In particular, a CAD file contains specifications from which geometric figures can be generated, which in turn allow a representation to be generated. The specifications for a modeled object can be stored in a single CAD file or in multiple CAD files. The typical size of a file representing a modeled object in a CAD system is in the range of 1 MB per part. And a modeled object can typically be an assembly of thousands of parts.

[0128] A CAM solution additionally means any solution, hardware or software suitable for managing the manufacturing data of a product. Manufacturing data generally includes data about the product to be manufactured, the manufacturing process and the required resources. A CAM solution is used to plan and optimize the entire manufacturing process of a product. For example, it can provide the CAM user with information about the feasibility, duration of the manufacturing process or the number of resources (e.g. a specific robot) that can be used at a specific step level of the manufacturing process; and thus, allow decisions on management or required investments. CAM is a subsequent process after the CAD process and potentially the CAE process. Such CAM solutions are sold by Dassault Systèmes under the trademark supply.

[0129] In the context of CAD, a modeled object may typically be a 3D modeled object, for example representing a product (e.g., a part or an assembly of parts), or possibly an assembly of products. A "3D modeled object" means any object that is modeled from data that allows its 3D representation. The 3D representation allows the part to be viewed from all angles. For example, when represented in 3D, a 3D modeled object can be manipulated and rotated about any of its axes or about any axis in the screen on which the representation is displayed. In particular, this does not include 2D icons that are not modeled in 3D. The display of a 3D representation facilitates design (i.e., increases the speed at which designers statistically complete their tasks). Since the design of a product is part of the manufacturing process, this can speed up the manufacturing process in industry.

[0130] Figures 13 to 17 Virtual manufacturing scenes are shown, each including virtual articulated industrial robots 130 , 140 , 150 and 152 , 160 and 162 , and 170 and 172 , respectively.

[0131] A virtual scene can be geometrically and semantically realistic, meaning that the geometry and functionality of virtual objects in the scene (e.g., products or manufacturing tools in a manufacturing scene) are realistic, e.g., realistically represent their corresponding geometry and functionality in the real world. A virtual scene can be geometrically and semantically realistic without being photorealistic. Alternatively, a virtual scene can be geometrically and semantically realistic and also photorealistic.

[0132] The concept of a photo-realistic image can be understood in terms of a number of visible features, such as: shading (i.e., how the color and brightness of a surface changes with lighting), texture mapping (i.e., the method of applying detail to a surface), bump mapping (i.e., the method of simulating small-scale bumps and bumps on a surface), fogging / participating media (i.e., how light darkens when passing through an opaque atmosphere or air), shadows (i.e., the effect of blocking light), soft shadows (i.e., the varying darkness caused by a partially blocked light source), reflections (i.e., specular or highly glossy reflections), transparency or opacity (i.e., sharp transmission of light through a solid object), translucency (i.e., highly diffuse transmission of light through a solid object), refraction (i.e., the bending of light associated with transparency), diffraction (i.e., the bending, scattering, and reflection of light passing through an object or hole that interferes with the light), and so on. Such properties exist so that such data (i.e., corresponding to or derived from a virtual photorealistic scene) has a similar distribution to real data (i.e., images representing real scenes). In this context, the closer the distribution of a virtual set is to the real-world data, the more photorealistic it appears.

[0133] Now let's discuss the concept of data distribution. In the context of building machine learning (ML) models, it is often desirable to train the model on data from the same distribution and test it on data from the same distribution. In probability theory and statistics, a probability distribution is a mathematical function that provides the probability of different possible outcomes occurring in an experiment. In other words, a probability distribution is a description of a random phenomenon in terms of the probability of an event. A probability distribution is defined with respect to an underlying sample space, which is the set of all possible outcomes of the observed random phenomenon. The sample space can be a set of real numbers or a higher-dimensional vector space, or it can be a non-numeric list. It is understood that when training on virtual data and testing on real data, one faces the problem of distribution shift. When modifying virtual images with real images to address the shift problem, domain adaptive learning methods, domain adaptive inference methods, and / or processes can transform the distribution of virtual data to better align with the distribution of real data. To quantify how well the two distributions align, several methods can be used. To illustrate, below are two examples of the methods.

[0134] The first method is a histogram method. Consider the case where N image pairs are created. Each image pair includes a first created virtual image and a second created virtual image, wherein the second created virtual image corresponds to a real image representing the same scene as the first virtual image, which real image is obtained by using a virtual-to-real generator. In order to evaluate the relevance of the modified training data, the domain adaptive learning method, the domain adaptive inference method and / or the process quantifies the degree of alignment of the virtual image distribution with the real image data before and after being synthesized into the real modification. To this end, the domain adaptive learning method, the domain adaptive inference method and / or the process may consider a set of real images and calculate the Euclidean distance between the average image of the real data and the virtual data before and after being synthesized into the real modification. Formally, consider D v ={v1, ..., v N} and D v′ ={v′1, ..., v′ N}, deep features are derived from a pre-trained network, e.g. trained on the ImageNet classification task (starting network, see

[22] ) on virtual images before and after modification. real Represents the deep features derived from the mean image of the real dataset, and the Euclidean difference is ED v ={v1-m, ..., v N -m} and ED v′ ={v′1-m, ..., v′ N -m}. To evaluate the alignment with the true data distribution, domain adaptation learning methods, domain adaptation inference methods and / or processes use histograms to compare ED v and ED v′The histogram is an accurate representation of the distribution of numerical data. It is an estimate of the probability distribution of a continuous variable. Figure 18 Shows ED v With ED v′ The histogram of the distribution of the Euclidean distance between . Figure 18 , compared to using the original virtual image ED v The distribution obtained when (blue) is compared to that when the representation comes from the modified virtual image ED v′ (red), the distribution is more concentrated around 0. This histogram validates the hypothesis being tested, which posits that the modified virtual data is transferred to a new domain that is closer to real-world data in terms of data distribution than the original virtual data. Regarding photorealism, consider a test that asks people to visually guess whether a given image is artificial or real. It's clear that the modified images are more visually likely to fool humans than the original virtual images. This directly correlates to the distribution alignment assessed by the aforementioned histogram.

[0135] The second method is the Frechet Inception Distance (FID) method (see

[21] ). FID compares the distributions of the inception embeddings of two sets of images (starting with the activations of the second-to-last layer, see

[22] ). Both distributions are modeled as multidimensional Gaussian models, parameterized by their respective means and covariances. Formally, let D A ={(x A )} is the first dataset of the scene (for example, the first dataset of the real scene), and let D B ={(x B )} is a second dataset (e.g., a second dataset of virtual scenes). The second method first uses a feature extractor network derived from a pre-trained network (initial network, see

[22] ) trained on, for example, the ImageNet classification task to extract features from x A and x B Extract deep features v A and v B Next, the second method collects v A This will produce multiple C-dimensional vectors, where C is the feature dimension. Then, the second method is to calculate the mean of these vectors and the covariance matrix Fit a multivariate Gaussian to these eigenvectors. Similarly, we have and The FID between these extracted features is:

[0136]

[0137] The smaller the distance, the greater the A and D BFrom the perspective of photo realism, when the first dataset is a dataset of a real scene and the second dataset is a dataset of a virtual scene, the smaller the distance is, the closer the distribution of D B The more photo-realistic the image is. This metric allows to assess the relevance of modifying our data from one domain to another in terms of distribution alignment.

[0138] Following the previous discussion on data distribution, the distance between two datasets of a scene can be quantified by distances such as the Euclidean distance or FID discussed previously. Therefore, in the context of the present disclosure, a "first dataset (e.g., family, e.g., domain, the concept of domain discussed below) of a scene is closer to a second dataset (e.g., family, e.g., domain) of the scene than a third dataset (e.g., family, e.g., domain) of the scene," which means that the distance between the first dataset (or family or domain) and the second dataset (or family or domain) is smaller than the distance between the third dataset (or family or domain) and the second dataset (or family or domain). In the case where the second dataset is a dataset of a real scene and both the third and first datasets are datasets of a virtual scene, it can be said that the first dataset is more photorealistic than the third dataset.

[0139] Now let's discuss the concept of a "domain." A domain specifies the distribution of a family of scenes. Therefore, a domain is also a family of distributions. A dataset of scenes belongs to a domain if the scenes in the dataset and the scenes in the family are close in terms of data distribution. In this example, this means that there is a predetermined threshold such that the distance (e.g., Euclidean distance or FID as described above) between the scenes in the dataset and the scenes in the family is less than the predetermined threshold. It should be understood that the dataset of scenes can also form a domain itself, in which case the domain can be referred to as the "domain of scenes" or the "domain formed by the scenes in the dataset." The first domain according to the domain adaptive learning method is the domain of virtual scenes, e.g., virtual scenes that are both geometrically and semantically realistic but not photorealistic. The second domain according to the domain adaptive learning method is the domain of real scenes. The test domain according to the domain adaptive learning method can be the domain of real scenes. Alternatively, the test domain can be the domain of virtual scenes, e.g., geometrically and semantically realistic and optionally photorealistic. The training domain according to the domain adaptive learning method can be the domain of virtual scenes, e.g., geometrically and semantically realistic and optionally photorealistic. In an example, the training domain is a domain of virtual scenes that is less photorealistic than scenes of the test domain, eg, the training domain is a domain of virtual scenes while the test domain is a domain of real scenes.

[0140] We now discuss the concept of providing scene datasets.

[0141] In the context of the present disclosure, providing a dataset of a scene can be performed automatically or by a user. For example, a user can retrieve a dataset from a memory storing a dataset. The user can further choose to complete the dataset by adding (e.g., one by one) one or more scenes to the retrieved dataset. Adding one or more scenes can include retrieving them from one or more memories and including them in the dataset. Alternatively, the user can create at least a portion (e.g., all) of the scene of the dataset. For example, as previously discussed, when providing a dataset of a virtual scene, the user can first obtain at least a portion (e.g., all) of the scene by using 3D simulation software and / or 3D modeling software and / or CAD software. More generally, in the context of domain adaptive learning methods, domain adaptive inference methods and / or processes, providing a dataset of a virtual scene can be preceded by a step of obtaining (e.g., calculating) the scene, such as by using 3D simulation software (e.g., for a manufacturing environment) and / or 3D modeling software and / or CAD software, as previously discussed. Similarly, in the context of domain adaptive learning methods, domain adaptive inference methods and / or processes, providing a dataset of real scenes may be preceded by a step of obtaining (e.g., acquiring) the scenes, for example, by a digital image acquisition process as described above, as described above.

[0142] In an example of a domain adaptive learning method and / or process, each virtual scene of the dataset of virtual scenes is a virtual manufacturing scene including one or more spatially reconstructable manufacturing tools. In these examples, the domain adaptive neural network configuration is used to infer spatially reconstructable manufacturing tools in a real manufacturing scene. In these examples, the virtual scenes of the dataset of virtual scenes can all be derived from simulations. Therefore, before step S10 of providing the dataset of virtual scenes, a step of performing simulations can be performed, for example, using 3D simulation software in a manufacturing environment, as described above. In an example, the simulation can be a 3D experience (i.e., a virtual scene with accurate modeling and faithful (e.g., compliance, e.g., consistent with reality) behavior), in which different configurations of a given scene are simulated, such as different arrangements of objects in a given scene and / or different configurations of certain objects (e.g., spatially reconstructable objects) in a given scene. The dataset can specifically include virtual scenes corresponding to the different configurations.

[0143] In examples of domain adaptive learning methods and / or processes, each real scene of the test data set is a real manufacturing scene including one or more spatially reconstructable manufacturing tools. In these examples, the domain adaptive neural network is configured to infer the spatially reconstructable manufacturing tools in the real manufacturing scene. In these examples, the real scenes of the test data set can all be derived from images collected during the digital acquisition process. Therefore, before providing the test data set of S10 virtual scenes, the following steps may be performed: acquiring images of the factory, and the scenes of the test data set are derived from images of the factory, for example by using one or more scanners and / or one or more cameras and / or one or more cameras.

[0144] In examples of domain adaptive inference methods and / or processes, the test data set includes real scenes. In these examples, each real scene of the test data set can be a real manufacturing scene including one or more spatially reconstructable manufacturing tools, and the domain adaptive neural network is optionally configured to infer the spatially reconstructable manufacturing tools in the real manufacturing scene. Therefore, in providing S10 a test data set of a virtual scene, the following steps can be performed: acquiring an image of a factory, the scene of the test data set being derived from the image of the factory, for example by using one or more scanners and / or one or more cameras and / or one or more cameras.

[0145] The determination of the third domain S20 is now discussed.

[0146] Determining the third domain S20 causes the third domain to be closer to the second domain than the first domain in terms of data distribution. Thus, determining the third domain S20 specifies an inference of another scene for each virtual scene in the dataset of virtual scenes belonging to the first domain, such that all inferred another scenes form (or at least belong to) a domain, namely, a third domain, which is closer to the second domain than the first domain in terms of data distribution. Inferring the other scene may include transforming (e.g., converting) the virtual scene into the other scene. The transformation may include, for example, transforming the virtual scene into the other scene using a generator capable of transferring images from one domain to another. It should be understood that such a transformation may involve only one or more portions of the virtual scene: the one or more portions are transformed, for example, using a generator, while the rest of the scene remains unchanged. The other scene may include a scene in which one or more portions are transformed while the rest of the scene remains unchanged. Alternatively, the transformation may include replacing the rest of the scene with a portion of some other scene (e.g., belonging to the second domain). In such a case, the other scene may include a scene in which one or more portions are transformed while the rest of the scene is replaced with the portion of the other scene. It should be understood that all scenarios of the first domain may be similarly inferred, eg, with the same paradigm being applicable to all scenarios.

[0147] Now refer to Figure 2 Discussing an example of S20 determining the third domain, Figure 2 A flowchart illustrating an example of S20 of determining the third domain is shown.

[0148] In an example, determining S20 the third domain includes, for each scene in the dataset of virtual scenes, extracting S200 one or more spatially reconstructable objects from the scene. In these examples, determining S20 also includes, for each extracted object, converting S210 the extracted object into an object that is closer to the second domain than the extracted object in terms of data distribution.

[0149] As previously discussed, prior to providing the virtual scene dataset ( S10 ), the virtual scene of the virtual scene dataset can be obtained, for example, by calculating a virtual simulation from which the scene is derived using 3D simulation software. The calculation of the simulation can include automatic positioning and annotation of each spatially reconstructable object in each virtual scene. For each virtual scene in the dataset, and for each spatially reconstructable object in the virtual scene, positioning can include calculating a bounding box around the spatially reconstructable object. A bounding box is a rectangular box whose axes are parallel to the sides of the image representing the scene. The bounding box is characterized by four coordinates and is centered around the spatially reconstructable object at an appropriate ratio and scale to enclose the entire spatially reconstructable object. In the context of the present disclosure, enclosing an object in the scene can mean encompassing all parts of the object that are fully visible in the scene, whether the part is the entire object or only a portion of the object. Annotation can include assigning a label to each calculated bounding box, indicating that the bounding box contains the spatially reconstructable object. It should be understood that the label can indicate other information about the bounding box, such as that the bounding box encloses a spatially reconstructable manufacturing tool, such as an industrial robot.

[0150] For each corresponding one of the one or more spatially reconstructable objects, extracting S200 the spatially reconstructable objects from the virtual scene can specify any action that results in separating the spatially reconstructable objects from the virtual scene, and can be performed by any method capable of performing such separation. In practice, extracting S200 the spatially reconstructable objects from the virtual scene can actually include: detecting a bounding box surrounding the spatially reconstructable object, the detection being based on a label of the bounding box and optionally including access to the label. Extracting S200 can also include separating the bounding box from the virtual scene, thereby also separating the spatially reconstructable object from the virtual scene.

[0151] The transformation S210 of the extracted object can specify any transformation that changes the extracted object into another object that is closer to the second domain in terms of data distribution. "Closer to the second domain" means that the object into which the extracted object is transformed is itself a scene that belongs to a domain that is closer to the second domain than the first domain. When the extraction S200 of the object includes separating the bounding box surrounding the object from the virtual scene as described above, the transformation S210 of the extracted object can include the entire bounding box (therefore including the extracted object) or just the modification of the extracted object, wherein the modification of the entire bounding box or the extracted object only makes it closer to the second domain in terms of data distribution. The modification can include applying a generator that can convert a virtual image into an image that is more photorealistic in terms of data distribution to the bounding box or an object that can be reconstructed in space with the bounding box.

[0152] In one embodiment, the transformation S210 can be performed by a CycleGAN virtual-to-real network (see

[20] ) trained on two sets of spatially reconstructable objects (virtual and real). It produces a virtual-to-real spatially reconstructable object (e.g., robot) generator that is used to modify the spatially reconstructable object (e.g., robot) in a virtual image. The CycleGAN model transfers images from one domain to another. It proposes a method that can learn to capture the special features of one set of images and figure out how to transfer these features to a second set of images without any paired training examples. In fact, this relaxation of having a one-to-one mapping makes this formulation very powerful. This is achieved by a generative model, specifically a generative adversarial network (GAN) called CycleGAN.

[0153] Figure 19 An example of extraction S200 and transformation S210 of two virtual industrial articulated robots 190 and 192 is shown. Each of these robots 190 and 192 has been extracted S200 from a corresponding virtual scene. The extracted robots 190 and 192 are then transformed S210 into more photorealistic robots 194 and 196, respectively.

[0154] Still refer to Figure 2 , determining the third domain may further include: for each transformed extracted object, placing the transformed extracted object in one or more scenes ( S220 ), each scene being closer to the second domain than the first domain in terms of data distribution. In this case, the third domain includes each scene in which the transformed extracted object is placed.

[0155] Each of the one or more scenes may be a scene that does not include spatially reconstructable objects, such as a background scene (i.e., a scene that includes only a background, although the background itself may include background objects). Each of the one or more scenes may also be a scene from the same context as the scene to which the transformed extracted object originally belonged. In fact, the same context may be the context of all scenes in the dataset of virtual scenes.

[0156] Each of the one or more scenes is closer to a second domain than the first domain. “Closer to the second domain” means that each of the one or more scenes belongs to a domain that is closer to the second domain than the first domain. In an example, each of the one or more scenes is a real scene directly obtained from reality, as previously discussed. Alternatively, each of the one or more scenes can be generated from a corresponding virtual scene of a dataset of virtual scenes, one of the corresponding virtual scenes being, for example, a virtual scene to which the extracted object belongs after transformation, such that the generated scene is more real and / or more photorealistic than the corresponding virtual scene. Generating a scene from a corresponding virtual scene of the dataset can include applying a generator capable of converting a virtual image into a real and / or more photorealistic image to the virtual scene. In one implementation, the generation of the scene can be performed by a CycleGAN virtual-to-real generator (see

[20] ) trained on two sets of backgrounds (e.g., factory backgrounds) (a virtual set and a real set). It produces a virtual-to-real background (e.g., factory background) generator that is used to modify the background (e.g., factory background) in the virtual image. The CycleGAN model (see

[20] ) transfers an image from one domain to another domain. It proposes a method that can learn to capture the special features of one set of images and figure out how to transform these features to a second set of images without any paired training examples. In fact, this relaxation of having a one-to-one mapping makes this formulation very powerful. This is achieved through generative models, specifically a generative adversarial network (GAN) called CycleGAN.

[0157] The placement S220 of the transformed extracted object can be performed (e.g., automatically) by any method capable of pasting a single and / or multiple images of the object into a background image. Placing S220 the transformed extracted object onto the one or more scenes may thus include applying such a method to paste the transformed extracted object onto each corresponding one of the one or more scenes. It should be understood that the scene on which the transformed extracted object is placed may include one or more other transformed extracted objects that have already been placed S220 on the scene.

[0158] As can be seen from the examples discussed above, in these examples, determining the third domain S20 can specifically transform the spatially reconstructable objects of the scene and the background of the scene to which they belong, respectively, and each transformed spatially reconstructable object can then be placed into one or more transformed backgrounds. First, this allows for a wider range of transformed object configurations within the transformed backgrounds compared to virtual scenes from a dataset that fully transforms the virtual scene (i.e., without separating the spatially reconstructable objects and their background). Because the learning S30 of the domain adaptation neural network is based on scenes from the third domain, the learning is more robust due to the third domain being closer to the test domain. Second, fully transforming the virtual scenes of the virtual scene dataset may lead to insufficient results because, due to practical background variability, scenes from the test domain may not have fundamental relationships with each other. For example, fully transforming a virtual scene including spatially reconstructable objects may result in a transformed scene in which the spatially reconstructable objects are indistinguishable from the background (e.g., blended with the background). For all of these reasons, fully transforming the virtual scenes of the virtual scene dataset may lead to insufficient results, e.g., by generating a third domain that may be no closer to the second domain than the first domain. Furthermore, scenes from the third domain may not be suitable for learning in this context, as the properties of the objects of interest (spatially reconstructible objects) in these scenes may be affected by the transformation.

[0159] Figure 20 An example of placement S220 of the transformed extracted object is shown. Figure 20 The S220 is placed in the factory background scenes 200 and 202 respectively. Figure 19 Examples of transformed extracted industrial articulated robots 194 and 196 .

[0160] Still refer to Figure 2 In the example, randomly placing the transformed extracted objects in one or more scenes S220 is performed. Such random placement S220 of the transformed extracted objects can be performed (e.g., automatically) by any method capable of randomly pasting a single and / or multiple images of an object into a background image. Therefore, randomly placing the transformed extracted objects in one or more scenes S220 can include applying such a method to randomly paste the transformed extracted objects onto each corresponding scene of the one or more scenes.

[0161] Randomly placing the transformed extracted objects S220 improves the robustness of the learning S30 of the domain adaptive neural network, which is based on a learning set consisting of scenes from the third domain. In fact, randomly placing the transformed extracted objects ensures that during learning S30, the domain adaptive neural network does not focus on the positions of the transformed extracted objects, because the sparse range of such positions is too large to make learning S30 rely on them and / or the correlation between the transformed extracted objects and their positions. Instead, learning S30 can truly focus on the inference of objects that can be spatially reconstructed, regardless of their positions. Incidentally, not relying on the positions of objects that can be spatially reconstructed allows for a significant reduction in the risk of the domain adaptive neural network missing spatially reconstructed objects that the domain adaptive neural network should infer simply because the objects do not occupy the appropriate positions.

[0162] In the above example, for each transformed extracted object, the third domain includes each scene in which the transformed extracted object is placed S220. For example, the third domain is composed of such scenes. In other words, the extraction S200, transformation S210, and placement S220 are performed similarly for all scenes in the dataset of virtual scenes. For example, they are performed so that all scenes in which one or more transformed extracted objects are placed have substantially the same level of photorealism. In other words, they belong to and / or form a domain, namely, the third domain.

[0163] Still refer to Figure 2 According to the flowchart of the embodiment of the present invention, determining S20 the third domain may include placing S230 one or more distractors in each corresponding one of the one or more scenes included in the third domain. A distractor is an object that is not a spatially reconstructable object.

[0164] The distractors may be objects having substantially the same level of photorealism as the transformed extracted objects. In practice, the distractors may be objects of a virtual scene of a dataset of virtual scenes, the objects of which are extracted from the virtual scene, transformed to be closer to objects of the second domain, and placed in a scene of the third domain in a manner similar to the spatially reconstructible objects with respect to the virtual scene. Alternatively, the distractors may be derived from another dataset of real or virtual distractors, in which case determining S20 may include the step of providing such a dataset. It should be understood that in any case, the distractors have a certain photorealism such that they can be placed S230 in scenes of the third domain, and that after placement S230, these scenes still belong to the third domain.

[0165] The distractors may be mechanical parts, e.g., extracted and transformed from the virtual manufacturing scene as described previously, or any other object (e.g., that is part of the virtual (e.g., manufacturing) scene background). Figure 21Three disruptors 210, 212, and 214 are shown to be placed in one or more scenes of the third domain in the context of a manufacturing scene. Disruptors 210, 212, and 214 are, respectively, a hard hat (which is a background object of the manufacturing scene), a metal plate (which is a machine part), and an iron horseshoe forging (which is a machine part).

[0166] Placing S230 distractors, in particular distractors obtained similar to the transformed extracted object (as described above), can reduce the sensitivity of the neural network to numerical (e.g., digital, such as imaging) artifacts. In fact, as a result of extraction S200 and / or transformation S210 and / or placement S220, a scene of a third domain that may include at least one transformed extracted object is characterized by one or more numerical artifacts at the location of at least one transformed extracted object at placement S220. Placing the distractor S230 in the scene may cause the same kind of artifacts to appear at the location where the distractor is placed. Therefore, when being learned S30, the domain adaptive neural network will not focus on the artifacts, or at least will not link the artifacts to objects that can be spatially reconstructed, precisely because the artifacts are also caused by the distractor. This makes learning S30 more robust.

[0167] Return Reference Figure 1 Now, we discuss the learning S30 of the domain adaptive neural network.

[0168] We begin with a brief discussion of the general concepts of "machine learning" and "learning neural networks."

[0169] As is known from the field of machine learning, the processing of an input by a neural network comprises applying an operation to the input, the operation being defined by data including weight values. Thus, learning of a neural network comprises determining the values of the weights based on a data set configured for such learning, such a data set being referred to as a learning data set or a training data set. To this end, the data set comprises data segments, each of which forms a respective training sample. The training samples represent the diversity of situations in which the neural network is to be used after learning. Any data set referred to herein may contain a number of training samples higher than 1,000, 10,000, 100,000 or 1,000,000. In the context of the present disclosure, "learning a neural network based on a data set" means that the data set is a learning / training set for the neural network. "Learning a neural network based on a domain" means that the training set for the neural network is a data set that belongs to and / or forms a domain.

[0170] Any neural network of the present disclosure, such as a domain adaptive neural network, or such as a detector, teacher extractor, or student extractor, will be discussed later, and they can be a deep neural network (DNN) learned by any deep learning technique. Deep learning techniques are a powerful set of learning techniques in neural networks (see

[19] ), which is a biologically inspired programming paradigm that enables computers to learn from observational data. In image recognition, the success of DNNs is attributed to their ability to learn rich mid-level media representations rather than hand-crafted low-level features (Zernike moments, HOG, Bag-of-Words, SIFT) used in other image classification methods (SVM, Boosting, Random Forest). More specifically, DNNs focus on end-to-end learning based on raw data. In other words, by completing end-to-end optimization starting from raw features and ending with labels, they can move away from feature engineering to the greatest extent possible.

[0171] The concept of "being configured to infer spatially reconstructable objects in a scene" is now discussed. This discussion is particularly applicable to domain adaptive neural networks according to any of domain adaptive learning methods, domain adaptive inference methods, or processes, regardless of the manner in which the domain adaptive neural network is learned. Inferring spatially reconstructable objects in a scene means inferring one or more spatially reconstructable objects in a scene that contains one or more such objects. This may mean inferring all spatially reconstructable objects in the scene. Thus, a neural network configured to infer spatially reconstructable objects in a scene takes as input the scene and output data related to the positions of one or more spatially reconstructable objects. For example, the neural network may output a (e.g., digital) image representation of the scene, wherein corresponding bounding boxes each enclose a corresponding one of the inferred spatially reconstructable objects. The inference may also include the computation of such bounding boxes and / or the labeling of bounding boxes (e.g., enclosing spatially reconstructable objects of the scene) and / or the labeling of the inferred spatially reconstructable objects (as spatially reconstructable objects). The output of the neural network may be stored in a database and / or displayed on a computer screen.

[0172] The learning S30 of the domain adaptive neural network may be performed by any machine learning algorithm capable of learning a neural network configured to infer spatially reconstructable objects in a real scene.

[0173] The learning S30 of the domain adaptive neural network is based on the third domain. Therefore, the third domain is or at least contains the training / learning set of the domain adaptive neural network. It is clear from the previous discussion about the determination S20 of the third domain that the training / learning set of the domain adaptive neural network is derived from (e.g., obtained from) the virtual scene dataset provided S10. In other words, the determination S20 can be considered as a preprocessing of the theoretical training / learning set, which is a dataset of virtual scenes, in order to transform it into an actual training / learning set, i.e., a scene of the third domain, which is closer to the set of inputs that the neural network (once learned) will accept. As previously discussed, this improves the robustness of the learning S30 and the output quality of the domain adaptive neural network.

[0174] The domain adaptive neural network may include (e.g., may be composed of) a detector, also called an object detector, which forms at least a part of the domain adaptive neural network. Thus, the detector itself is a neural network that is also configured to detect spatially reconstructable objects in a real scene. The detector may typically include millions of parameters and / or weights. Learning S30 of the domain adaptive method may include training the detector, which sets the values of these parameters and / or weights. Such training may typically include updating these values, which includes continuously correcting these values according to the output of the detector for each input obtained by the detector. During training, the input of the detector is the scene, and the output includes (at least after appropriate training) data related to the location of one or more (e.g., all) spatially reconstructable objects in the scene (e.g., bounding boxes and / or annotations, as described above).

[0175] The correction of the value can be based on the annotations associated with each input. An annotation is a set of data associated with a specific input that allows the output of the model to be evaluated as true or false. For example, and as previously discussed, each spatially reconstructable object belonging to a scene in the third domain can be surrounded by a bounding box. As previously mentioned, in the context of the present disclosure, surrounding an object in a scene may mean surrounding all parts of an object that are completely visible on the scene. The bounding box can be annotated to indicate that it surrounds the spatially reconstructable object. The correction of the value can therefore include evaluating such annotations of the bounding box of the third domain scene input to the detector, and based on the evaluated annotations, determining whether the spatially reconstructable object surrounded by the bounding box is indeed output by the detector. Determining whether the output of the detector includes the spatially reconstructable object surrounded by the bounding box can include evaluating the correspondence between the output of the detector and the spatially reconstructable object, for example, by evaluating whether the output is true (for example, whether it includes the spatially reconstructable object) or false (for example, whether it includes the spatially reconstructable object).

[0176] The above-described way of training a detector by using annotations of a virtual dataset can be called "supervised learning". The training of the detector can be performed by any supervised learning method. After the detector is trained, the correction of the values stops. At this point, the detector is able to process new inputs (i.e., inputs that were not seen during the training of the detector) and return detection results. In the context of the present disclosure, the new input to the detector is the scene of the test domain. After training, the detector will return two different outputs, because the "detection" task means jointly performing the recognition (or classification) task and the localization task of spatially reconstructable objects in the real scene of the test dataset.

[0177] The localization task involves computing bounding boxes, each of which encloses a spatially reconstructible object in the real-world scene fed to the detector. As previously mentioned, a bounding box is a rectangular box whose axes are parallel to the image edges and characterized by four coordinates. In this example, for each spatially reconstructible object in the scene fed to the detector, the detector returns a bounding box centered on the object at the appropriate ratio and scale.

[0178] The classification task consists of labeling each computed bounding box with a corresponding label (or annotation) and associating a confidence score with that label. The confidence score reflects the detector's confidence that the bounding box indeed encloses the spatially reconstructable object of the scene provided as input. The confidence score can be a real number between 0 and 1. In this case, the closer the confidence score is to 1, the greater the detector's confidence that the label associated with the corresponding bounding box truly annotates the spatially reconstructable object.

[0179] Now refer to Figure 3 Discuss the example of learning S30 of domain adaptive neural network, Figure 3 A flow chart illustrating an example of learning S30 of a domain adaptive neural network is shown.

[0180] refer to Figure 3 , the learning of the domain adaptive neural network S30 may include providing S300 a teacher extractor. The teacher extractor is a machine-learned neural network configured to output an image representation of a real scene. In this case, the learning of the domain adaptive neural network S30 also includes training S310 a student extractor. The student extractor is a neural network configured to output an image representation of a scene belonging to a third domain. The training of the student extractor S310 includes minimizing a loss. For each of the one or more real scenes, the loss penalizes the difference between the result of applying the teacher extractor to the scene and the result of applying the student extractor to the scene.

[0181] The student extractor is a neural network configured to output an image representation of a scene belonging to a third domain. The student extractor can be part of a detector, and thus, the student extractor can be further configured to output the image representation of the scene belonging to the third domain to other parts of the detector. As previously discussed, at least one of such other parts can be configured for a localization task, and at least one of such other parts can be configured for a classification task. Thus, the student extractor can output the image representation of the scene of the third domain to the other parts so that the other parts can perform the tasks of localization and classification. Alternatively, the student extractor can be part of a domain adaptive neural network, and in this case the image representation can be output to other parts of the domain adaptive neural network (e.g., a detector).

[0182] This is equivalent to saying that the student extractor supervises the output image representation of scenes belonging to the third domain, while at least one other part of the domain adaptation neural network (e.g., the detector discussed previously) supervises the detection (e.g., localization and classification) of spatially reconstructable objects in each scene of the third domain. The detection can be based on the image representation output by the student extractor. In other words, the image representation output by the student extractor is fed to the rest of the domain adaptation neural network, for example, if the student extractor forms part of the detector, then the rest of the detector.

[0183] The student extractor is trained with the help of a teacher extractor provided in S300. The teacher extractor is a machine-learned neural network configured to output an image representation of a real scene. In this example, this means that the teacher extractor has been learned on a real scene to output an image representation of the real scene, for example, by learning robust convolutional filters for the real data pattern. The neuron layers of the teacher extractor typically encode a representation of the real image to be output. The teacher extractor can be learned on any dataset of real scenes, for example, available from open source using any machine learning technology. Providing the teacher extractor in S300 may include accessing a database in which the teacher extractor has been stored after learning, and retrieving the teacher extractor from the database. It should be understood that the teacher extractor may already be available, that is, the domain adaptive learning method may not include learning the teacher extractor simply by providing the teacher extractor already available in S300. Alternatively, the domain adaptive learning method may include learning the teacher extractor using any machine learning technology before providing the teacher extractor in S300.

[0184] Still refer to Figure 2 , we now discuss the training of the student extractor S320.

[0185] The training S320 comprises minimizing a loss that penalizes, for each of one or more real scenes, the difference between the result of applying the teacher extractor to the scene and the result of applying the student extractor to the scene. At least a portion (e.g., all) of the one or more real scenes may be scenes of a test dataset, and / or at least a portion (e.g., all) of the one or more real scenes may be scenes of another dataset of real scenes belonging to a second domain, e.g., which are not annotated and not necessarily numerous. Minimizing the loss involves applying both the teacher extractor and the student extractor to the result of each corresponding real scene in the one or more real scenes, which means that the real scenes are fed into both the teacher extractor and the student extractor during the latter's training S310. Thus, while at least a portion of the domain adaptive neural network is trained on virtual data belonging to a third domain (e.g., the detector discussed previously), at least a portion of the domain adaptive neural network, i.e., the student extractor, is trained on real data.

[0186] The loss can be a quantity (e.g., a function) that measures the similarity and / or dissimilarity between the results of applying the teacher extractor to the scene and the results of applying the student extractor to the scene for each of the one or more real scenes. For example, the loss can be a function that takes each of the one or more real scenes as a parameter, the function taking as input the results of applying the teacher extractor to the scene and the results of applying the student extractor to the scene, and outputting a quantity (e.g., a positive real number) that represents the similarity and / or dissimilarity between the results of applying the teacher extractor to the scene and the results of applying the student extractor to the scene. For each of the one or more real scenes, the difference between the results of applying the teacher extractor to the scene and the results of applying the student extractor to the scene can be a quantification of the dissimilarity between the results of applying the teacher extractor to the scene and the results of applying the student extractor to the scene. Penalizing the difference can mean that the loss is an increasing function of the difference.

[0187] Minimizing the loss improves the robustness of the learning S30 of the domain adaptive neural network by penalizing the differences between the outputs of the teacher extractor and the student extractor. In particular, the student extractor is guided by a teacher extractor that has already been learned and whose weights and / or parameters are therefore not modified during the training S320 of the student extractor. Indeed, because the differences between the outputs of the student extractor and the teacher extractor are penalized, the parameters and / or weights of the student extractor are modified so that the student extractor learns to output realistic image representations that are close to the image representations output by the teacher extractor in terms of the data distribution. In other words, the student extractor is trained to mimic the teacher extractor in a way that outputs realistic image representations. Therefore, the student extractor can learn the same or substantially the same robust convolutional filters learned by the teacher extractor during training for realistic data patterns. As a result, the student extractor outputs realistic image representations of scenes that originally belonged to the third domain, even if these scenes may be relatively far away from the second domain (composed of real scenes) in terms of data distribution. This allows making the domain-adaptive neural network, or where appropriate, its detector invariant to the appearance of the data, since the student extractor is trained in a manner S320 that enables the domain-adaptive neural network to infer spatially reconstructable objects from real scenes while being trained on annotated virtual scenes in a third domain.

[0188] This loss may be referred to as a "distillation loss," and the training of the student extractor using the teacher extractor S320 may be referred to as a "distillation method." Minimization of the distillation loss may be performed by any minimization algorithm or any relaxation algorithm. The distillation loss may be part of the loss of the domain adaptive neural network, optionally also including the loss of the detector discussed previously. In one embodiment, the learning of the domain adaptive neural network S30 may include minimizing the loss of the domain adaptive neural network, which may be of the following type:

[0189] L training =L detector +L distillation

[0190] Among them L training is the loss corresponding to the learning S30 of the domain adaptive neural network, L detector is the loss corresponding to the training of the detector, L distillation is the distillation loss. Minimize the loss L of the domain adaptive neural network training It can therefore include minimizing the detector L by any relaxation algorithm simultaneously or independently detector Loss and distillation loss L distillation .

[0191] In the example, the result of applying the teacher extractor to the scene is a first Gramian matrix, and the result of applying the student extractor to the scene is a second Gramian matrix. The Gramian matrix of an extractor (e.g., student or teacher) is a form of representing the distribution of data on an image output by the extractor. More specifically, applying the extractor to the scene produces an image representation, and computing the Gramian matrix of the image representation produces an embedding of the distribution of the image representation, a so-called Gramian-based representation of the image. This is equivalent to saying that in the example, the result of applying the teacher extractor to the scene is a first Gramian-based representation of the image representation output by the teacher extractor, which is obtained by computing the first Gramian matrix of the image representation output by the teacher extractor, and the result of applying the student extractor to the scene is a second Gramian-based representation of the image representation output by the student extractor, which is obtained by computing the second Gramian matrix of the image representation output by the student extractor. For example, the Gramian matrix can contain non-localized information about the image output by the extractor, such as texture, shape, and weights.

[0192] In examples, the first Gramian matrix is calculated over several layers of neurons in the teacher extractor. The layers of neurons include at least the last layer of neurons. In these examples, the second Gramian matrix is calculated over several layers of neurons in the student extractor. The layers of neurons include at least the last layer of neurons.

[0193] The last layer of neurons in the extractor (teacher or student) encodes the image output by the extractor. Both the first and second Gram matrices are calculated over several layers of neurons, which are deeper layers of neurons. This means that they both contain more information about the output image than the image itself. This makes training the student extractor more robust because it increases the proximity of the true style between the output of the teacher extractor and the student extractor, which is the goal of minimizing the loss. Remarkably, the student extractor can thus infer a two-layer representation of the real image as accurately as its teacher.

[0194] In the example, the difference is the Euclidean distance between the result of applying the teacher extractor to the scene and the result of applying the student extractor to the scene.

[0195] Now let’s discuss an implementation that minimizes distillation loss. In this implementation, the extractor (teacher or student) The style representation of the image output at some layer of is encoded by the feature distribution of the image in that layer. In this embodiment, the feature correlation is encoded using the Gram matrix, which is a form of feature distribution embedding. Formally, for a layer with length The filter response Layer Feature correlation is given by the Gram matrix Given, where It is a layer The inner product between the vectorized feature maps i and j is:

[0196]

[0197] The intuition behind the proposed distillation method is to ensure that during the learning process, the weights of the student extractor are updated so that its output data style representation is close to the form generated by the teacher extractor based on the real data. To this end, learning S30 consists of minimizing the Gram-based distillation loss L at the feature extractor level distillation . Given a real image, let and In a layer of the student extractor and the teacher extractor respectively Therefore, the proposed layer The distillation losses are of the following types:

[0198]

[0199] in is the corresponding Gram matrix and Any distance between (e.g., Euclidean distance). Consider In the distillation knowledge, the type of distillation loss is:

[0200]

[0201] As previously described, the learning S30 of the domain adaptation neural network may then include minimizing a loss of the following type:

[0202] L training =L detector +L distillation .

[0203] Return Reference Figure 1 Referring to the flowchart of FIG. 1 , S50 of determining the intermediate domain is now discussed.

[0204] Determining the intermediate domain S50 causes the intermediate domain to be closer to the training domain than the test domain in terms of data distribution. Thus, determining the intermediate domain S50 specifies an inference of another scene for each scene in the test dataset of scenes, such that all inferred alternative scenes form a domain that is closer to the training domain in terms of data distribution than the test domain, i.e., an intermediate domain. Inferring the alternative scene may include transforming (e.g., converting) the scene into the alternative scene. The transformation may include, for example, transforming the scene into the alternative scene using a generator capable of transferring images from one domain to another. It should be understood that such a transformation may involve only one or more portions of the scene: the one or more portions are transformed, e.g., using a generator, while the rest of the scene remains unchanged. The alternative scene may include a scene in which one or more portions are transformed while the rest of the scene remains unchanged. Alternatively, the transformation may include replacing the rest of the scene with a portion of some other scene (e.g., belonging to the training domain). In such a case, the alternative scene may include a scene in which one or more portions are transformed while the rest of the scene is replaced with the portion of the other scene. It should be understood that similar inferences may be performed for all scenes in the test domain, for example, using the same paradigm applicable to all scenes.

[0205] The domain adaptive neural network learns not on the training domain, but rather on data obtained from the training domain. This data can be the result of data processing from the training domain. For example, the data obtained from the training domain can be a dataset of scenes, each of which is obtained by transforming the corresponding scene from the training domain. This is equivalent to saying that the data obtained from the training domain can form a dataset of scenes belonging to a domain, which is the true training / learning domain of the domain adaptive neural network. This domain can be precisely determined as a third domain by performing similar operations and steps. In any case, determining the intermediate domain ( S50 ) can be performed in such a way (e.g., by transforming the scenes as discussed above) that the intermediate domain is closer to the domain formed by the data obtained from the training domain than the test domain. In some examples, the data obtained from the training domain forms another intermediate domain that, in terms of data distribution, is closer to the intermediate domain than the training domain. This other intermediate domain can be the third domain discussed above. In all of these examples, determining the intermediate domain ( S50 ) is equivalent to preprocessing the test dataset to make it closer to the true training / learning set on which the domain adaptive neural network was learned. As discussed above, this improves the output quality of the domain adaptive neural network.

[0206] In an example, for each scene of the test dataset, determining S50 the intermediate domain includes transforming the scene of the test dataset into another scene that is closer to the training domain in terms of data distribution.

[0207] Transforming a scene into another scene may comprise using a generator that is capable of transferring images from one domain to another. It will be appreciated that such a transformation may involve only one or more parts of the scene: the one or more parts are transformed, e.g. using a generator, while the rest of the scene remains unchanged. The other scene may comprise a scene where one or more parts are transformed while the rest of the scene remains unchanged. Alternatively, the transformation may comprise replacing the rest of the scene with a part of some other scene (e.g. belonging to the training domain). In such a case, the other scene may comprise a scene where one or more parts are transformed while the rest of the scene is replaced by the part of some other scene. It will be appreciated that inferences may be made similarly for all scenes of the test domain, e.g. having the same exemplars for all scenes.

[0208] Now refer to Figure 4 Discussing an example of determining S50 of the middle domain, Figure 4 A flowchart illustrating an example of determination S50 of the intermediate domain is shown.

[0209] refer to Figure 4 , the transformation of the scene may include generating S500 a virtual scene from the scene of the test data set. In this case, another scene will be inferred based on the scene of the test data set and the generated virtual scene.

[0210] Generating S500 can be performed by applying any generator capable of generating a virtual scene from a scene of the test domain. Another scene can be inferred by mixing the scene of the test dataset and the generated virtual scene and / or by combining (e.g., mixing, seamless cloning, or merging (e.g., by merging image channels)) the scene of the test dataset and the generated virtual scene.

[0211] Still refer to Figure 4 Flowchart of generating S500 a virtual scene from a scene of a test data set may include applying a virtual scene generator to the scene of the test data set. In this case, the virtual scene generator is a machine learning neural network configured to infer the virtual scene from the scene of the test domain.

[0212] The virtual scene generator takes the scene of the test domain as input and outputs a virtual scene, for example, a virtual scene with lower photorealism than the scene of the test domain. It should be understood that the virtual generator may have already been trained / learned. Thus, generating the virtual scene S500 may include providing the virtual scene generator, for example, by accessing a memory storing the virtual scene generator after training / learning the virtual scene generator and retrieving the virtual scene generator from the memory. Additionally or alternatively, before providing the virtual scene generator, the virtual scene generator may be learned / trained by any machine learning technique suitable for such training.

[0213] In an example, a virtual scene generator has been learned on a dataset of scenes, each scene including one or more spatially reconstructable objects. In such an example, when the virtual scene generator is fed a scene including one or more spatially reconstructable objects, the virtual scene generator can output a virtualization of the scene, which is a virtual scene that also includes one or more spatially reconstructable objects, for example all located in the same position and / or in the same configuration. As previously mentioned, it should be understood that the domain adaptive inference method and / or process can include only providing a virtual scene generator (i.e., it is not learned), or can also include learning of the virtual scene generator.

[0214] In an embodiment of a domain adaptive inference method and / or process wherein the test dataset consists of a real manufacturing scene containing a manufacturing robot, the generator is a CycleGAN real-to-virtual network trained on two manufacturing scene datasets containing manufacturing robots (a virtual dataset and a real dataset). In this embodiment, the method and / or process comprises training a CycleGAN network on the two datasets to adapt the appearance of the robot in the test images of the test scene to more closely resemble the training set of the domain adaptive neural network, which in this embodiment is a third domain according to a domain adaptive learning method and consists of a virtual manufacturing scene containing a virtual manufacturing robot. The training results in a real-to-virtual robot generator that is capable of modifying the real test images to virtualize them.

[0215] Figure 22 An example of using a CycleGAN generator to generate S500 is shown. In this example, the test dataset consists of real manufacturing scenes, each of which contains one or more manufacturing robots. Figure 22 A first real scene 220 and a second real scene 222 of the test dataset are shown. Both scenes are manufacturing scenes that include an articulated robot. Figure 22 Also shown are a first virtual scene 224 and a second virtual scene 226 generated from the first scene 220 and the second scene 222, respectively, by applying the CycleGAN generator.

[0216] Still refer to Figure 4 In the flowchart of FIG. 5 , determining the intermediate domain S50 may include mixing the scene of the test dataset and the generated virtual scene S510 . In this case, another scene is obtained by mixing.

[0217] Formally, let I be the scene of the test dataset, let I' be the generated virtual scene, and let I" be another scene. Mixing S510 I and I' can refer to any method capable of mixing I and I' so as to obtain I". In an example, the mixing S510 of the scene of the test dataset and the generated virtual scene is a linear mixing. In one embodiment, the linear mixing includes calculating the other scene I" by the following formula:

[0218] I"=α*I+(1-α)*I'.

[0219] Figure 23 An example of linear blending is shown. Figure 23 A real manufacturing scene 230 including a real manufacturing robot and a virtual scene 232 generated from the real scene 230 is shown. Figure 23 Also shown is the result of the linear blending, which is a virtual manufacturing scene 234 .

[0220] Return Reference Figure 1 , we now discuss the objects that can be reconstructed in the scene inference S60 space from the test dataset of the test domain transferred on the intermediate domain.

[0221] First, inferring spatially reconstructable objects from a scene (S60) may include inferring spatially reconstructable objects from a scene, or inferring spatially reconstructable objects from each scene of the input dataset of the scene for each scene of the input dataset of the scene. Notably, inferring (S60) may constitute a testing phase of a domain adaptive neural network, wherein all scenes of the test dataset are provided (e.g., sequentially) as input to the domain adaptive neural network, which infers one or more spatially reconstructable objects in each input scene. Additionally or alternatively, inferring (S60) may constitute an application phase of a domain adaptive learning method and / or a domain adaptive inference method and / or process, wherein the input dataset is a dataset of a test domain scene that is not equal to the test dataset. In this case, the application phase may include providing the input dataset, inputting each scene of the input dataset into the domain adaptive neural network, and performing inferring (S60) for each scene. It should also be understood that inferring (S60) spatially reconstructable objects in a scene may include inferring (e.g., simultaneously or sequentially) one or more spatially reconstructable objects in a scene that includes one or more such objects, for example by inferring all spatially reconstructable objects in the scene.

[0222] Now discuss the inference S60 of a spatially reconstructable object in a scene, which may be part of the inference S60 of more spatially reconstructable objects in the scene. Inferring S60 a spatially reconstructable object in a scene by applying a domain adaptive neural network includes inputting the scene into the domain adaptive neural network. The domain adaptive neural network then outputs data related to the location of the spatially reconstructable object. For example, the neural network may output a (e.g., digital) image representation of the scene, where a corresponding bounding box surrounds the inferred spatially reconstructable object. The inference may also include the calculation of such a bounding box and / or the labeling of the bounding box (e.g., surrounding the spatially reconstructable object of the scene) and / or the labeling of the inferred spatially reconstructable object (as the spatially reconstructable object). The output of the neural network may be stored in a database and / or displayed on a screen of a computer.

[0223] Figures 24 to 29 Shown is the inference of a manufacturing robot in a real manufacturing scenario. Figures 24 to 29 Each shows the corresponding real scene of the test domain fed as input to the domain adaptation neural network, and the bounding boxes computed and annotated around all manufacturing robots included in each corresponding scene. Figures 26 to 30 As shown, multiple robots and / or at least partially occluded robots can be inferred in the real scene.

[0224] In practice, inference of spatially reconstructable objects can be performed by a detector, which forms at least part of a domain adaptation neural network that jointly performs the tasks of identifying (or classifying) and localizing spatially reconstructable objects in real-world scenes from a test dataset. The localization task involves computing bounding boxes, each of which encloses a spatially reconstructable object in the scene input to the detector. As previously described, a bounding box is a rectangular box whose axes are parallel to the image edges and characterized by four coordinates. In this example, for each spatially reconstructable object in the scene input to the detector, the detector returns a bounding box centered on the object at the appropriate scale and proportion. The classification task involves labeling each computed bounding box with a corresponding label (or annotation) and associating a confidence score with the label. The confidence score reflects the detector's confidence that the bounding box truly encloses the spatially reconstructable object in the scene provided as input. The confidence score can be a real number between 0 and 1. In this case, the closer the confidence score is to 1, the greater the detector's confidence that the label associated with the corresponding bounding box truly annotates the spatially reconstructable object.

[0225] In examples where a domain adaptive neural network is configured to infer objects in real scenes, such as when the test domain consists of real scenes and the intermediate domain consists of virtual scenes, the domain adaptive neural network can also include an extractor. The extractor forms part of the domain adaptive neural network, which is configured to output image representations of the real scene and feed them to other parts of the domain adaptive neural network, such as a detector or a part of a detector, for example, to supervise the detection of spatially reconstructable objects. In examples, this means that the extractor can process scenes from the intermediate domain to output image representations of the scene, for example by learning robust convolutional filters for the pattern of real data. The neuron layers of the extractor typically encode representations of real or photorealistic images to be output. This is equivalent to saying that the extractor supervises the output image representations of scenes belonging to the intermediate domain, while at least one other part of the domain adaptive neural network (e.g., the detector discussed previously) supervises the detection (e.g., localization and classification) of spatially reconstructable objects in each input scene of the third domain. The detection can be based on the image representations output by the student extractor. This allows for efficient processing and output of real data.

[0226] Each scene that is input to the domain adaptation neural network is a scene that belongs to the test domain, but is transferred on the determined intermediate domain. In other words, providing the scene as input to the domain adaptation neural network includes: transferring the scene on the intermediate domain and feeding the transferred scene as input to the domain adaptation neural network. The transfer of the scene can be performed by any method that can transfer a scene from one domain to another domain. In an example, transferring the scene can include transforming the scene into another scene and optionally mixing the scene with the transformed scene. In this case, it can be as previously referred to Figure 4 It is worth noting that the scene of the intermediate domain can actually be a scene of the test dataset transferred on the intermediate domain.

[0227] The implementation of this process is now discussed.

[0228] This embodiment proposes a learning framework in a virtual world to address the problem of insufficient manufacturing data. Specifically, this embodiment utilizes the efficiency of virtual simulation datasets and image processing methods for learning to learn domain adaptive neural networks, namely deep neural networks (DNNs).

[0229] Virtual simulation is a form of obtaining virtual data for training. It has emerged as a promising technique to address the difficulty of obtaining the training data required for machine learning tasks. Virtual data is generated algorithmically to mimic the characteristics of real data, while being automatically, error-free, and cost-free labeled.

[0230] This implementation proposes virtual simulation to better explore the working mechanisms of different manufacturing equipment. Formally, different manufacturing tools are considered during the learning process through the different states that these tools will assume when running.

[0231] Deep Neural Networks: are a powerful set of learning techniques in neural networks

[19] , a biologically inspired programming paradigm that enables computers to learn from observational data. In image recognition, the success of DNNs is attributed to their ability to learn rich mid-level representations, rather than the hand-crafted low-level features (Zernike moments, HOG, bag-of-words, SIFT, etc.) used in other image classification methods (SVM, boosting, random forest, etc.). More specifically, DNNs focus on end-to-end learning based on raw data. In other words, by completing an end-to-end optimization starting from raw features and ending with labels, they can move away from feature engineering as much as possible.

[0232] Domain adaptation is a field related to machine learning and transfer learning. This scenario arises when the goal is to learn a model from a source data distribution that performs well on a different (but related) target data distribution. The goal of domain adaptation is to find an effective mechanism to transfer or adapt knowledge from one domain to another. This is achieved by introducing a cross-domain bridging component to bridge the cross-domain gap between training and test data.

[0233] The proposed embodiment is particularly characterized by Figure 31 , which shows a diagram illustrating the offline and online stages of the process.

[0234] 1. Offline Phase: This phase aims to train a model using virtual data, which should be capable of inferring the real world. It consists of two main steps. Note that this phase is transparent to the user.

[0235] 1) Perform virtual simulation in a virtual manufacturing environment to generate representative virtual training data and automatically annotate it. The virtual manufacturing environment is geometrically realistic rather than photorealistic. The annotation type depends on the target inference task.

[0236] 2) Learning neural network models based on virtual data, including models based on domain-adaptive DNNs.

[0237] 2. Online phase: Given real-world media, the learned domain adaptation model is applied to infer manufacturing equipment.

[0238] This implementation of the process is particularly relevant to the field of fully supervised object detection. Specifically, in this particular implementation, we are interested in manufacturing environments, where the goal is to detect industrial robot arms in factory images. For this fully supervised domain, the training dataset should contain up to hundreds of thousands of annotated data points, which justifies the use of virtual data for training.

[0239] State-of-the-art object detectors are based on deep learning models. The values of the millions of parameters that characterize the model cannot be set manually. Therefore, these parameters must be set with the help of a learning algorithm. When the learning algorithm updates the model parameters, the model is said to be in "training mode." This involves continuously "correcting" the model based on its output for each input, thanks to annotations associated with each input. An annotation is a set of data associated with a specific input that allows the model's output to be evaluated as true or false. For example, an object classifier trained to distinguish between images of cats and dogs requires a dataset of annotated images of cats and dogs, each annotated as "cat" or "dog." Therefore, if the object classifier outputs "dog" for an input image of a cat in its training mode, the learning algorithm will correct the model by updating its parameters. This method of supervising model training with an annotated dataset is called "supervised learning." Once the model is trained, its parameters are no longer updated. The model is then used only to process new inputs (i.e., inputs not seen during training mode) and return detection results. This is called "testing mode."

[0240] In the context of this embodiment, the object detector returns two different outputs because the "detection" task means performing the recognition (or classification) task and the localization task together.

[0241] 1. Localization Output: Object localization is possible thanks to bounding boxes. A bounding box is a rectangular box whose axes are parallel to the image edges. It is characterized by four coordinates. Ideally, for each object, an object detector returns a bounding box centered on the object at the appropriate scale and proportions.

[0242] 2. Classification output: Object classification is performed by associating a class label with a confidence score for each bounding box. The confidence score is a real number between 0 and 1. The closer the value is to 1, the more confident the object detector is about the class label associated with the corresponding bounding box.

[0243] This embodiment is particularly characterized by Figure 32 FIGURE 2 shows a domain adaptation component used on both the offline and online stages.

[0244] It involves learning an object detection model using generated virtual data. Several domain adaptation methods are applied in both offline and online stages to address the gap between the virtual and real domains.

[0245] 1. Offline phase: This phase aims to train a model using virtual data, which should be capable of inferring the real world. This phase consists of two main steps.

[0246] 1) Perform virtual simulation in a virtual manufacturing environment to generate representative virtual training data that is automatically annotated with bounding boxes and “robot” labels.

[0247] 2) Domain Adaptive Transformation

[0248] i. At the training set level: transform the training set to a new training domain that is closer to the real world in terms of data distribution.

[0249] ii. At the model architecture level: A domain adaptation module is inserted into the model to address the problematic virtual-real domain shift.

[0250] 2. Online phase: Real-world media is converted to a new test domain, which is then better processed by the learned detection model and then inferred by the learned detection model. In other words, the new test domain is closer to the training domain than the original real domain.

[0251] It is important to note that a particularity of this implementation is that it does not suffer from the accuracy issue of virtual-to-real image transformations. The intuition behind the current work is to consider an intermediate domain between the virtual and real worlds, rather than attempting to migrate virtual data to the real world. Formally, virtual training data and real test data are pulled into the intermediate training and testing domains, respectively, where the domain shift is minimized.

[0252] Based on the above framework, it can be understood that the implementation currently discussed is based on two main steps:

[0253] Virtual data generation for learning.

[0254] Domain adaptation.

[0255] To address the scarcity of factory images, this implementation relies on a simulation-based virtual data generation method. The training set consists of 3D rendered factory scenes. These scenes were generated using the company's in-house design application and include masks that segment the objects of interest (robots) at the pixel level. Each scene is simulated into multiple frames to depict the various possible functional states of the robot.

[0256] This step generates a set of virtual images with unique and complex features for the task of detecting robots in a manufacturing environment. Although this appears to be a “simple” single-object detection task, baseline detectors and state-of-the-art domain adaptation networks fail at this problem. Formally, the generated set is not photorealistic but rather “cartoonish.” The simplicity of the generated virtual scenes reflects the ease of generating our training data, which is the main advantage of this implementation. At the same time, the complexity of the objects of interest makes it difficult to use state-of-the-art methods for this type of virtual data, namely domain randomization (DR) techniques (see [7, 15]).

[0257] In practice, domain randomization (DR) abandons photorealism by randomly perturbing the environment in non-photorealistic ways. The intuition behind this DR is to teach the algorithm to force the network to learn to focus on essential features of the image and ignore uninteresting objects in the scene. More formally, the main real features missing from the virtual data are low-level features of the image, such as texture, lighting, and shading. DR ensures that these features are perturbed in the training dataset so that the algorithm does not rely on them to recognize objects. By the same logic, DR will not affect the common features between real and virtual data. Within the range of virtual data generation used, one might learn that the learned model expects the correct features to learn to rely on are shape and contour. This is important in DR techniques.

[0258] Furthermore, the objects of interest in this embodiment are robots, which have many joint configurations (typically 6 degrees of freedom) that are calibrated over a continuous range of values. This gives each robot type an infinite range of possible states. Therefore, for such complex industrial objects, randomizing all unique features in addition to the learned shape and outline is not the best option. In fact, when the shape and outline of the object are complex, it is difficult to meet the variability required to represent the shape and outline of the object. Especially for the robot example, it is indeed unimaginable to learn a model that can extract and focus on the unique shapes and outlines of various robot types with infinite joint positions to recognize objects in images.

[0259] The proposed implementation sets up a learning problem that is competent when the data is complex (such as in the present case).

[0260] The proposed domain adaptation method can be considered a partial domain randomization technique. Formally, it proposes to enhance the correct features that the model is learning by intentionally randomizing or transforming the data features to a new domain that better represents the real world than the virtual environment. For example, it can better explore the specificity of texture features. To this end, three novel techniques are implemented in the implementation currently discussed:

[0261] 1. Conversion from virtual to real

[0262] To correct for cross-domain shift, a common approach in existing art papers is to alter the appearance of virtual data to make it appear as if it were extracted from the real domain, i.e., they obtain a distribution that better represents the real world. To do this, the virtual image is holistically modified using the real scene, without any correspondence between similar objects. However, to achieve the best image-to-image translation performance, it might be beneficial to transfer the real robot texture to the robot in the virtual scene and use it against the background.

[0263] The current implementation builds on this targeted appearance transfer by applying different image-to-image translation networks to the robot and factory contexts for each media in the training set, respectively.

[0264] To this end, we use the CycleGAN model

[20] to translate images from one domain to another. A method is proposed that can learn to capture the unique features of one set of images and figure out how to translate these features to a second set of images without any paired training examples. In fact, this relaxation of having a one-to-one mapping makes this representation very powerful. This is achieved through a generative model, specifically a generative adversarial network (GAN) called CycleGAN.

[0265] As mentioned above, two different models were trained

[0266] CycleGAN Virtual-to-Real Network 1: Trained on two sets of factory backgrounds: virtual and real. It produces a virtual-to-real factory generator that modifies the background in virtual images.

[0267] CycleGAN Virtual to Real 2: Trained on two sets of robots (virtual and real). This results in a virtual to real robot generator that is used to modify robots in virtual images.

[0268] 2. Gram-based distillation

[0269] To better emphasize the robustness of the model with respect to domain shift, we propose a novel knowledge distillation method implemented at the detector level. Note that most state-of-the-art detectors (see [18, 9]) contain a feature extractor, which is a deep network of convolutional layers. The feature extractor outputs an image representation to be fed to the rest of the detector. The intuition of distillation is to guide the detection model to learn convolutional filters that are robust to the real data style, which is consistent with the main goal of making the model invariant to the data appearance.

[0270] Formally, the implementation consists of training a detector on virtual media while updating the detector's feature extractor weights to approximate the style encoded by a two-layer representation of real images fed into both our detector and a second detector pre-trained on real data from open source.

[0271] Please note that a layer The style representation of the image at is encoded by the feature distribution in this layer. For this implementation, we use a Gram matrix that encodes feature correlations, which is a form of feature distribution embedding. Formally, for a layer with length The filter response Layer Feature correlation is given by the Gram matrix Given, where It is a layer The inner product between the vectorized feature maps i and j is:

[0272]

[0273] The intuition behind the proposed distillation is to ensure that during the learning process the detector’s feature extractor weights are updated so that the data-style representation it outputs is close to the data-style representation that another feature extractor pre-trained on real data would produce. To this end, the implementation currently discussed implements a Gram-based distillation loss L at the feature extractor level. distillation .

[0274] Given a real image, we use and Layers representing the detector's feature extractor and pre-trained feature extractors The corresponding feature map output of . Therefore, in the layer The proposed distillation loss at is the corresponding Gram matrix and The Euclidean distance between

[0275]

[0276] Consider the layer In the distillation knowledge, the distillation loss is defined as

[0277]

[0278] The total training loss is

[0279] L training =L detector +L distillation

[0280] 3. Conversion from real to virtual

[0281] One might expect that the task of translating virtual images to the real domain cannot be accomplished perfectly. This means that the modified data is closer to the real world than the virtual data, while not matching the real distribution exactly. Therefore, we assume that the modified virtual data is transferred to a new domain D1 that lies between the synthetic domain and the real domain. One contribution of this work is to extend the scope of domain adaptation to real test data by dragging it to another new domain D2. In order to use this approach, D2 should be closer to D1 than the real domain in terms of distribution; this is an assumption confirmed in this application case by comparing data representations. This form of adaptation is achieved by applying the real-to-virtual data transformation on the test data (i.e., the inverse operation previously applied to the virtual data during the training phase).

[0282] In the context of this embodiment, the transformation is ensured by training a CycleGAN real-to-virtual network trained on two robot sets (virtual and real). This results in a real-to-virtual robot generator that is used to modify the real test images. At this level, we choose to train the CycleGAN network on the robot set because we are primarily interested in adapting the appearance of the robots in the test images so that they can be recognized by the detector. Note that it is also possible to perform more complex image transformations on the test data by targeting regions individually, as was done on the training set. However, the improvement is limited, and we are content to apply one CycleGAN network to the entire test image.

Claims

1. A computer-implemented machine learning method comprising: -supply: - a dataset of a virtual scene, the dataset of the virtual scene belonging to a first domain; as well as - A test dataset of real scenes belonging to the second domain; - determining a third domain, said third domain being closer to said second domain than said first domain in terms of data distribution; as well as - learning a domain adaptive neural network based on the third domain, the domain adaptive neural network being a neural network configured to infer spatially reconstructable objects in a real scene, The learning of the domain adaptive neural network includes: - providing a teacher extractor, the teacher extractor being a machine-learned neural network configured to output an image representation of a real scene; - training a student extractor, the student extractor being a neural network configured to output an image representation of a scene belonging to the third domain, the training of the student extractor comprising minimizing a loss that penalizes, for each of one or more real scenes, a difference between a result of applying the teacher extractor to the scene and a result of applying the student extractor to the scene.

2. The method according to claim 1, wherein Determining the third domain includes performing the following operations for each scene in the dataset of the virtual scene: - extracting one or more spatially reconstructable objects from the scene; as well as - for each extracted object, transforming the extracted object into an object that is closer to the second domain than the extracted object in terms of data distribution.

3. The method according to claim 2, wherein: Determining the third domain further includes: - for each transformed extracted object, placing the transformed extracted object in one or more scenes, each scene being closer to the second domain than the first domain in terms of data distribution, the third domain including each scene on which the transformed extracted object is placed.

4. The method according to claim 3, wherein: Placing the transformed extracted objects in the one or more scenes is performed randomly.

5. The method according to claim 3, wherein The determining of the third domain includes placing one or more distractors in each respective one or more scenes included in the third domain, the distractors being objects that are not spatially reconstructable objects.

6. The method according to claim 1, wherein The result of applying the teacher extractor to the scene is a first Gramian matrix, and the result of applying the student extractor to the scene is a second Gramian matrix.

7. The method according to claim 6, wherein: - the first Gramian matrix is calculated on a plurality of neuron layers of the teacher extractor, the neuron layers including at least the last neuron layer; and - The second Gramian matrix is calculated over several neuron layers of the student extractor, including at least the last neuron layer.

8. The method according to claim 1, wherein The difference is the Euclidean distance between the result of applying the teacher extractor to the scene and the result of applying the student extractor to the scene.

9. The method according to any one of claims 1 to 8, wherein: Each virtual scene in the dataset of virtual scenes is a virtual manufacturing scene including one or more spatially reconfigurable manufacturing tools, and the domain adaptive neural network is configured to infer the spatially reconfigurable manufacturing tools in the real manufacturing scene.

10. The method according to any one of claims 1 to 8, wherein Each real scene in the test dataset is a real manufacturing scene including one or more spatially reconfigurable manufacturing tools, and the domain adaptation neural network is configured to infer the spatially reconfigurable manufacturing tools in the real manufacturing scene.

11. A computer program product comprising instructions for performing the method according to any one of claims 1 to 10.

12. A device comprising a data storage medium on which the computer program product according to claim 11 is recorded.

13. The apparatus of claim 12, further comprising a processor coupled to the data storage medium.

Citation Information

Patent Citations

  • Viewpoint invariant object recognition by synthesization and domain adaptation

    US20190066493A1

  • Refining Synthetic Data With A Generative Adversarial Network Using Auxiliary Inputs

    US20190080206A1