Experiential Learning in the Virtual World

Through domain adaptive learning methods and neural networks, virtual scene data sets are converted into intermediate domains, which solves the accuracy problem of virtual data sets being able to reconstruct objects in the real world in the identification space, and improves the model's inference ability in real scenes.

CN111898172BActive Publication Date: 2025-08-01DASSAULT SYSTEMES SA
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202010371444.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-05-06
Filing Date
2020-05-06
Publication Date
2025-08-01
Estimated Expiration
2040-05-06

AI Technical Summary

Technical Problem

The prior art is difficult to effectively use virtual data sets to train deep models to accurately identify and infer spatially reconstructed objects in real-world scenarios, especially in the field of manufacturing, due to domain differences between virtual and real data and data scarcity issues.

Method used

Through the domain adaptive learning method, virtual scene data sets are provided and converted into intermediate domains closer to real scenes. Domain adaptive neural networks are used for training and inference, and a hybrid data set of virtual and real scenes is combined to improve the robustness and accuracy of the model.

Benefits of technology

It realizes accurate inference of objects that can be reconstructed spatially in real scenes, solves the problem of generalization of virtual data sets in the real world, and improves the robustness and output quality of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111898172B_ABST
    Figure CN111898172B_ABST
Patent Text Reader

Abstract

The present invention particularly relates to a computer-implemented machine learning method. The method includes providing a test data set of a scenario. The test data set belongs to a test domain. The method includes providing a domain adaptation neural network. The domain adaptation neural network is a neural network of machine learning that learns data obtained from a training domain. The domain adaptation neural network is configured to infer spatially reconstructable objects in the scenario of the test domain. The method further includes determining an intermediate domain. In terms of data distribution, the intermediate domain is closer to the training domain than the test domain. The method further includes inferring spatially reconstructable objects according to the scenario of the test domain transferred on the intermediate domain by applying the domain adaptation neural network. Such a method constitutes an improved machine learning method using a data set of scenarios including spatially reconstructable objects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer programs and systems, and more particularly to methods, systems, and programs for machine learning with a dataset having a scene that includes spatially reconstructable objects. Background Art

[0002] Many systems and programs for the design, engineering, and manufacturing of objects are provided on the market. CAD is an acronym for computer-aided design, for example, which relates to software solutions for designing objects. CAE is an acronym for computer-aided engineering, for example, which relates to software solutions for simulating the physical behavior of future products. CAM is an acronym for computer-aided manufacturing, for example, which relates to software solutions for defining manufacturing processes and operations. In such computer-aided design systems, the graphical user interface plays an important role in terms of technical efficiency. These technologies may be embedded within a product lifecycle management (PLM) system. PLM refers to a business strategy that helps companies share product data, apply common processes, and leverage company knowledge for product development across extended enterprise concepts from concept to the end of product life. The PLM solution provided by Dassault Systèmes (trademarks CATIA, ENOVIA, and DELMIA) provides an engineering hub for organizing product engineering knowledge, a manufacturing hub for managing manufacturing engineering knowledge, and an enterprise hub for enabling enterprise integration and connectivity to the engineering and manufacturing hubs. The overall system provides an open object model that connects products, processes, and resources to enable dynamic, knowledge-based product creation and decision support, which drives optimized product definition, manufacturing readiness, production, and service.

[0003] In this context and others, machine learning is becoming increasingly important.

[0004] The following papers are relevant to this field and are cited below:

[0005] [1] A.Gaidon, Q.Wang, Y.Cabon, and E.Vig, “Virtual worlds as proxy for multi-object tracking analysis,” in CVPR, 2016.

[0006] [2] S.R.Richter, V.Vineet, S.Roth, and V.Koltun, “Playing for data: Groundtruth from computer games,” in ECCV, 2016.

[0007] [3] M. Johnson-Roberson, C. Barto, R. Mehta, S. N. Sridhar, K. Rosaen, and R. Vasudevan, “Driving in the matrix: Can virtual worlds replace human-generated annotations for real world tasks?” in ICRA, 2017.

[0008] [4] S. R. Richter, Z. Hayder, and V. Koltun, “Playing for benchmarks,” in ICCV, 2017.

[0009] [5] S. Hinterstoisser, V. Lepetit, P. Wohlhart, and K. Konolige, “On pretrained image features and synthetic images for deep learning,” in arXiv:1710.10710, 2017.

[0010] [6] D. Dwibedi, I. Misra, and M. Hebert, “Cut, paste and learn: Surprisingly easy synthesis for instance detection,” in ICCV, 2017.

[0011] [7] J. Tobin, R. Fong, A. Ray, J. Schneider, W. Zaremba, and P. Abbeel, “Domain randomization for transferring deep neural networks from simulation to the real world,” in IEEE / RSJ International Conference on Intelligent Robots and Systems (IROS), 2017.

[0012] [8] J. Tremblay, A. Prakash, D. Acuna, M. Brophy, V. Jampani, C. Anil, T. To, E. Cameracci, S. Boochoon, and S. Birchfield, “Training deep networks with synthetic data: Bridging the reality gap by domain randomization,” in CVPR Workshop on Autonomous Driving (WAD), 2018.

[0013] [9] K. He, G. Gkioxari, P. Dollar, and R. Girshick. “Mask r-cnn”. arXiv:1703.06870, 2017.

[0014]

[10] J. Deng, A. Berg, S. Satheesh, H. Su, A. Khosla, and L. Fei-Fei. ILSVRC-2017, 2017. URL http: / / www.image-net.org / challenges / LSVRC / 2017 / .

[0015]

[11] Y. Chen, W. Li, and L. Van Gool. ROAD: Reality oriented adaptation for semantic segmentation of urban scenes. In CVPR, 2018.

[0016]

[12] Y. Chen, W. Li, C. Sakaridis, D. Dai, and L. Van Gool. Domain adaptive Faster R-CNN for object detection in the wild. In CVPR, 2018.

[0017]

[13] Y. Zou, Z. Yu, B. Vijaya Kumar, and J. Wang. Unsupervised domain adaptation for semantic segmentation via class-balanced self-training. In ECCV, September 2018.

[0018]

[14] A. Geiger, P. Lenz, and R. Urtasun, “Are we ready for autonomous driving? The KITTI vision benchmark suite,” in CVPR, 2012.

[0019]

[15] A. Prakash, S. Boochoon, M. Brophy, D. Acuna, E. Cameracci, G. State, O. Shapira, and S. Birchfield. Structured domain randomization: Bridging the reality gap by context-aware synthetic data. arXiv preprint arXiv:1810.10093, 2018.

[0020]

[16] J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei. ImageNet: A Large-Scale Hierarchical Image Database. In CVPR09, 2009.

[0021]

[17] T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollar, and C. L. Zitnick. Microsoft COCO: Common objects in 'context'. In ECCV. 2014.

[0022]

[18] Shaoqing Ren, Kaiming He, Ross B. Girshick, and Jian Sun. Faster R-CNN: towards real-time object detection with region proposal networks. CoRR, abs / 1506.01497, 2015.

[0023]

[19] D.E. Rumelhart, G.E. Hinton, R.J. Williams, Learning internal representations by error propagation, Parallel distributed processing: explorations in the microstructure of cognition, vol.1: foundations, MIT Press, Cambridge, MA, 1986.

[0024]

[20] Jun-Yan Zhu, Taesung Park, Phillip Isola, and Alexei A. Efros. Unpaired image-to-image translation using cycle-consistent adversarial networks. CoRR, abs / 1703.10593, 2017.

[0025]

[21] M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter. Ganstrained by a two time-scale update rule converge to a local nash equilibrium. In Advances in Neural Information Processing Systems (NIPS), 2017.

[0026]

[22] C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich. Going deeper with convolutions. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 1–9, 2015.

[0027] Deep learning techniques have shown excellent performance in several fields such as 2D / 3D object recognition and detection, semantic segmentation, pose estimation, or motion capture (see [9, 10]). However, such techniques typically (1) require large amounts of training data to fully realize their potential and (2) generally require expensive manual labeling (also known as annotation):

[0028] (1) In fact, due to data scarcity or confidentiality reasons, it is difficult to collect hundreds of thousands of labeled data in several fields (e.g., the domains of manufacturing, health, or indoor motion capture). Publicly available datasets usually include objects in daily life, which are manually labeled via crowdsourcing platforms (see [16, 17]).

[0029] (2) In addition to the data collection problem, labeling is also very time-consuming. The annotation of the well-known ImageNet dataset (see

[16] ) took several years, where each image was labeled with “only” a single label. Pixel-by-pixel labeling per image on average takes more than one hour. Additionally, manual annotation may have some errors and lack accuracy. Some annotations require domain-specific expertise (e.g., manufacturing equipment labeling).

[0030] For all these reasons, in recent years, the use of virtual data to learn deep models has attracted increasing attention (see [1, 2, 3, 4, 5, 6, 7, 8, 15]). Virtual (also known as synthetic) data is computer-generated data created using CAD or 3D modeling software, rather than data generated by actual events. In fact, virtual data can be generated to meet very specific requirements or conditions that are not available in existing (real) data. This can be useful when privacy requirements limit the availability or use of data, or when the data required for the test environment simply does not exist. Note that in the case of virtual data, the labels are cost-free and error-free.

[0031] However, the inherent domain difference between virtual data and real data (also known as the reality gap) may make the learned model incompetent for real-world scenarios. This is mainly due to the overfitting problem, which causes the network to learn details that only exist in virtual data, and fail to generalize well and extract the information representation of real data.

[0032] Recently, especially for recognition problems related to autonomous vehicles becoming prevalent, one way is to collect photo-realistic virtual data, which can be automatically annotated at low cost, such as video game data.

[0033] For example, Richter et al. (see [2]) constructed a large-scale synthetic urban scene dataset for semantic segmentation from the GTAV game. Based on GTA data (see [2, 3, 4]), a large number of assets were used to generate various realistic driving environments that could be automatically labeled at the pixel level (about 25K images). Such photo-realistic virtual data usually does not exist in domains other than driving and urban contexts. For example, existing virtual manufacturing environments are non-photo-realistic CAD-based environments (there are no video games for the manufacturing domain). Manufacturing specifications do not require photo-realism. Additionally, having shadows and lightening in such virtual environments can be confusing. Thus, all work has focused on creating semantically and geometrically precise CAD models rather than photo-realistic CAD models. Moreover, due to the large variability between the layouts of different workshops and factories, creating photo-realistic manufacturing environments for training purposes is not relevant. Unlike urban contexts where urban scenes have strong idiosyncrasies (e.g., the size and spatial relationships of buildings, streets, cars, etc.), there is no "inherent" spatial structure for manufacturing scenes. Also note that some previous work (e.g., VKITTI (see [1])) created replicas of real-world driving scenes and thus was highly correlated with those original real data and lacked variability.

[0034] Furthermore, even with a high degree of realism, it is not obvious how to effectively use photo-realistic data to train neural networks to operate on real data. Typically, cross-domain adaptation components (see [11, 12, 13]) are used during the training of neural networks to narrow the gap between virtual and real representations.

[0035] In summary, photo-realistic virtual datasets require carefully designed simulation environments or the existence of annotated real data as a starting point, which is not feasible in many domains (e.g., manufacturing).

[0036] To alleviate this difficulty, Tobin et al. (see [7]) introduced the concept of domain randomization (DR), in which real-world rendering is avoided in favor of random variations. Their method randomly changes the texture and color of foreground objects, the background image, the number of lights in the scene, the poses of the lights, the camera position, and the foreground objects. The goal is to narrow the reality gap by generating virtual data with a sufficient amount of variation (the network treats real-world data as another variation). They use DR to train a neural network to estimate the 3D world positions of various shape-based objects relative to a robotic arm fixed to a table. Recent work (see [8]) demonstrated state-of-the-art performance of DR for 2D bounding box detection of cars in the real-world KITTI dataset (see

[14] ). However, in the case of such variations, DR requires a very large amount of data for training (about 100K per object category), the network often finds it difficult to learn the correct features, and the lack of context causes DR to fail to detect smaller or occluded objects. In fact, the results of this work are limited to (1) larger cars that are (2) fully visible (KITTI Easy dataset). When used for complex inference tasks such as object detection, this method requires a sufficient number of pixels within the bounding box for the learned model to make decisions without surrounding context. All these drawbacks were quantitatively confirmed in (see

[15] ).

[0037] To handle more challenging criteria (e.g., smaller cars, partial occlusions), the authors of

[15] proposed a variant of DR, in which they utilize urban context information and urban structured scenes by randomly placing objects according to a probability distribution generated by the specific problem at hand (i.e., car detection in an urban context) rather than the uniform probability distribution used in previous DR-based work.

[0038] Similarly, DR-based methods (see [5, 6, 7, 15]) mainly focus on "easy" object classes with little intra-class variability (e.g., the "car" class). Significant intra-class variability may lead to a significant drop in the performance of these methods because the amount of variation to be considered during data creation may grow exponentially. This is the case for objects that are spatially reconstructable, such as articulated objects.

[0039] Spatially reconstructable or articulated objects are objects that have components via joint attachments and can move relative to each other, which can give an infinite range of possible states for each object type.

[0040] The same problem may arise when dealing with multiple object classes (inter-class and intra-class variability).

[0041] In this context, there is still a need to use a scene dataset including spatially reconstructable objects to improve machine learning methods. Summary of the Invention

[0042] Accordingly, a computer-implemented machine learning method is provided. The method includes providing a test dataset of scenes. The test dataset belongs to a test domain. The method includes providing a domain adaptation neural network. The domain adaptation neural network is a neural network of machine learning that learns from data obtained from a training domain. The adaptation neural network is configured to infer spatially reconstructable objects in test domain scenes. The method further includes determining an intermediate domain. In terms of data distribution, the intermediate domain is closer to the training domain than the test domain. The method further includes inferring spatially reconstructable objects from scenes of the test domain transferred onto the intermediate domain by applying the domain adaptation neural network.

[0043] The method may include one or more of the following:

[0044] - For each scene of the test dataset, determining the intermediate domain includes:

[0045] · Transforming the scene of the test dataset into another scene that is closer to the training domain in terms of data distribution;

[0046] - The transformation of the scene includes:

[0047] · Generating a virtual scene based on the scene of the test dataset, and the other scene is inferred based on the scene of the test dataset and the generated virtual scene;

[0048] - Generating a virtual scene based on the scene of the test dataset includes applying a virtual scene generator to the scene of the test dataset, and the virtual scene generator is a neural network of machine learning configured to infer a virtual scene from scenes of the test domain;

[0049] - The virtual scene generator has been learned on a scene dataset that all includes one or more spatially reconstructable objects;

[0050] - Determining the intermediate domain further includes mixing the scene of the test dataset with the generated virtual scene, and the mixing produces another scene;

[0051] - The mixing of the scene of the test dataset and the generated virtual scene is a linear mixing;

[0052] - The test dataset includes real scenes;

[0053] - Each real scene of the test dataset is a real manufacturing scene that includes one or more spatially reconstructable manufacturing tools, and the domain adaptation neural network is configured to infer the spatially reconstructable manufacturing tools in the real manufacturing scene;

[0054] - The training domain includes a training data set of virtual scenarios, and each scenario includes one or more spatially reconstructable objects;

[0055] - Each virtual scenario in the virtual scenario data set is a virtual manufacturing scenario, which includes one or more spatially reconstructable manufacturing tools, and the domain adaptation neural network is configured to infer spatially reconstructable manufacturing tools in the real manufacturing scenario; and / or

[0056] - The data obtained from the training domain includes scenarios of another intermediate domain, and the domain adaptation neural network has been learned on another intermediate domain, and in terms of data distribution, another intermediate domain is closer to the intermediate domain than the training domain.

[0057] A computer program is also provided, which includes instructions for executing the method.

[0058] A device is also provided, which includes a data storage medium on which a computer program is recorded.

[0059] The device can form or be used as a non-transitory computer-readable medium, such as on SaaS (Software as a Service) or other servers, or on a cloud-based platform, etc. The device can alternatively include a processor coupled to the data storage medium. Thus, the device can form a computer system in whole or in part (for example, the device is a subsystem of the entire system). The system can further include a graphical user interface coupled to the processor. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] Embodiments of the present invention will now be described by way of non-limiting examples and with reference to the accompanying drawings, wherein:

[0061] - Figure 1 A flowchart of a computer-implemented process incorporating an example of the method is shown;

[0062] - Figures 2 to 4 A flowchart is shown, which shows an example of a computer-implemented process and / or an example of a method; <�

[0063] - Figure 5 An example of a graphical user interface of the system is shown;

[0064] - Figure 6 An example of the system is shown; and

[0065] - Figures 7 to 32 A process and / or method is shown. DETAILED DESCRIPTION

[0066] A computer-implemented machine learning method for utilizing a dataset of scenarios, the dataset of scenarios including objects that can be spatially reconstructed.

[0067] In particular, a first machine learning method is provided. The first method includes providing a dataset of virtual scenarios. The dataset of virtual scenarios belongs to a first domain. The first method further includes providing a test dataset of real scenarios. The test dataset belongs to a second domain. The first method further includes determining a third domain. In terms of data distribution, the third domain is closer to the second domain than the first domain. The first method further includes learning a domain-adaptive neural network based on the third domain. The domain-adaptive neural network is a neural network configured to infer objects that can be spatially reconstructed in real scenarios. Hereinafter, this first method may be referred to as the "domain-adaptive learning method".

[0068] The domain-adaptive learning method forms an improved machine learning method for utilizing a dataset of scenarios, the dataset of scenarios including spatially reconstructable objects.

[0069] It is noteworthy that the domain-adaptive learning method provides a domain-inference neural network that can infer objects that can be spatially reconstructed in real scenarios. This allows the domain-adaptive learning method to be used in the process of inferring objects that can be spatially reconstructed in real scenarios (e.g., in the process of inferring spatially reconstructable manufacturing tools (such as industrial articulated robots) in a real manufacturing scenario).

[0070] In addition, it may not be possible to directly learn the domain-adaptive neural network on the dataset of real scenarios, but the domain-adaptive neural network can still infer the actual objects that can be spatially reconstructed in real scenarios. This improvement is achieved especially by determining the third domain according to the domain-adaptive learning method. In fact, since it is composed of virtual scenarios, the first domain may be relatively far from the second domain (i.e., the domain of the test dataset of the domain-adaptive neural network composed of real scenarios) in terms of data distribution. Such a real-world gap may make it difficult and / or inaccurate to directly learn the domain-adaptive neural network based on the dataset of virtual scenarios. Such learning may indeed result in a domain-adaptive neural network that may produce inaccurate results in inferring objects that can be spatially reconstructed in real scenarios. On the contrary, the domain-adaptive learning method determines a third domain that is closer to the second domain than the first domain in terms of data distribution and bases the learning of the domain-adaptive neural network on this third domain. This makes it possible for the learning of the domain-adaptive neural network to accurately infer objects that can be spatially reconstructed in real scenarios, while the only dataset provided (i.e., given as the input to the domain-adaptive learning method) is actually a dataset of virtual scenarios. In other words, both the learning and the learned domain-adaptive neural network are more robust.

[0071] This is equivalent to saying that the dataset of virtual scenarios is theoretically intended to be the training set (or learning set) of a domain - adaptive neural network, and the domain - adaptive learning method includes: before actual learning, a pre - processing stage of the training set (i.e., determining the third domain) in order to learn a domain - adaptive neural network with the improvements discussed previously. In other words, the true training / learning domain of the domain - adaptive neural network is the third domain, and it stems from the pre - processing of the domain (i.e., the first domain) that was originally intended to be the training / learning domain of the domain - adaptive neural network. However, relying on the dataset of virtual scenarios as the theoretical training set (i.e., as described above, to be processed before being used for training) allows the use of a dataset with a large number of scenarios. In fact, usually there may be no or at least not many large datasets of real scenarios available that include spatially reconstructable objects, such as manufacturing scenarios of manufacturing tools that are spatially reconstructable. Due to privacy / confidentiality issues, such datasets can be really difficult to obtain: for example, manufacturing scenarios with spatially reconstructable manufacturing tools (e.g., robots) are usually confidential and thus not publicly available. Additionally, even if such a large dataset of real scenarios were available, it might contain many annotation (e.g., labeling) errors: annotating spatially reconstructable objects (e.g., manufacturing tools, e.g., robots) can be really difficult for someone who does not have the appropriate skills for it, because spatially reconstructable objects may have so many and / or such complex positions that their identification and manual annotation in a real scenario would require a certain amount of skills and knowledge. On the other hand, a large dataset of virtual scenarios can be easily obtained, for example, by using simulation software. Additionally, annotating objects (e.g., spatially reconstructable objects) in virtual scenarios is relatively easy, and most importantly, it can be performed automatically without errors. In fact, simulation software usually knows the specifications of the objects related to the simulation, so its annotation can be automatically performed by the software, and the risk of annotation errors is low.

[0072] In the examples, each virtual scenario of the dataset of virtual scenarios is a virtual manufacturing scenario that includes one or more spatially reconstructable manufacturing tools. In these examples, the domain - adaptive neural network is configured to infer spatially reconstructable manufacturing tools in a real manufacturing scenario. Additionally or alternatively, each real scenario of the test dataset is a real manufacturing scenario that includes one or more spatially reconstructable manufacturing tools. In this case, the domain - adaptive neural network is configured to infer spatially reconstructable manufacturing tools in a real manufacturing scenario.

[0073] In such examples, the domain - adaptive neural network is configured in any case to infer a spatially reconfigurable manufacturing tool (e.g., a spatially reconfigurable manufacturing robot) in a real manufacturing scenario. In these examples, particular emphasis is placed on the previously discussed improvements brought about by the domain - adaptive learning method. In particular, it is difficult to rely on a training set of a large number of real manufacturing scenarios that include spatially reconfigurable manufacturing tools, because such a set is typically confidential and / or publicly available but not adequately annotated. On the other hand, it is particularly convenient to provide a large dataset of virtual manufacturing scenarios that include spatially reconfigurable manufacturing tools as a training set, because such a dataset can be easily obtained and automatically annotated, for example, by using software capable of generating virtual simulations in a virtual manufacturing environment.

[0074] There is also provided a domain - adaptive neural network that can be learned according to the domain - adaptive learning method.

[0075] There is also provided a second computer - implemented machine - learning method. The second method includes providing a test dataset of scenarios. The test dataset belongs to a test domain. The second method includes providing a domain - adaptive neural network. The domain - adaptive neural network is a neural network of machine - learning that is learned based on data obtained from a training domain. The domain - adaptive neural network is configured to infer a spatially reconfigurable object in a test - domain scenario. The second method further includes determining an intermediate domain. In terms of data distribution, the intermediate domain is closer to the training domain than the test domain. The second method further includes inferring, by applying the domain - adaptive neural network, a spatially reconfigurable object from a scenario of the test domain transferred onto the intermediate domain. This second method can be referred to as the “domain - adaptive inference method”.

[0076] The domain - adaptive inference method forms an improved machine - learning method that utilizes a scenario dataset including spatially reconfigurable objects.

[0077] It is noteworthy that the domain - adaptive inference method allows for the inference of spatially configurable objects from scenarios. This allows the domain - adaptive inference method to be used in the inference process of spatially reconfigurable objects (e.g., in the inference process of a spatially reconfigurable manufacturing tool (e.g., an industrial articulated robot) in a manufacturing scenario).

[0078] In addition, the method allows for the accurate inference of spatially reconstructable objects from the scenarios of the test dataset, and the domain adaptation neural network used for such inference has been machine-learned based on data obtained from the training domain, which may be far from the test domain (the test domain to which the test dataset belongs) in terms of data distribution. This improvement is achieved in particular by determining an intermediate domain. In fact, in order to better use the domain adaptation neural network, before applying the domain adaptation neural network for inference, the domain adaptation inference method transforms the scenarios of the test domain into scenarios belonging to the intermediate domain, in other words, scenarios that are closer in terms of data distribution to the learning scenarios of the domain adaptation neural network. In other words, the domain adaptation inference method allows for not directly providing the scenarios of the test domain as input to the domain adaptation neural network, but rather providing as input to the domain adaptation neural network scenarios that are closer to the scenarios on which the domain adaptation neural network has been learned, in order to improve the accuracy and / or quality of the output of the domain adaptation neural network (each of one or more spatially reconstructable objects inferred in the corresponding scenarios of the test domain).

[0079] Therefore, in examples where the test domain and the training domain are relatively far from each other, the improvement brought by the domain adaptation inference method is particularly emphasized. For example, this may be the case when the training domain consists of virtual scenarios and the test domain consists of real scenarios. For example, this may also be the case when both the training domain and the test domain consist of virtual scenarios, but the scenarios of the test domain are more photo-realistic than the scenarios of the training domain. In all these examples, the method includes a stage of preprocessing the scenarios of the test domain (i.e., determining the intermediate domain), which are provided as input to the domain adaptation neural network (e.g., in the test phase of the domain adaptation neural network), in order to make them closer to the scenarios of the training domain in terms of photo-realism (which can be quantified in terms of data distribution). As a result, the photo-realism level of the scenarios provided as input to the domain adaptation neural network is close to the photo-realism of the scenarios on which the domain adaptation neural network has been learned. This ensures better output quality of the domain adaptation neural network (i.e., the inferred spatially reconstructable objects inferred on the scenarios of the test domain).

[0080] In an example, a training domain includes a training data set of virtual scenes, each virtual scene including one or more spatially reconstructable objects. Additionally, each virtual scene of the data set of virtual scenes can be a virtual manufacturing scene including one or more spatially reconstructable manufacturing tools, in which case the domain adaptive neural network is configured to infer spatially reconstructable manufacturing tools in a real manufacturing scene. Additionally or alternatively, a test data set includes real scenes. In an example of this case, each real scene of the test data set is a real manufacturing scene including one or more spatially reconstructable manufacturing tools, and the domain adaptive neural network is configured to infer spatially reconstructable manufacturing tools in a real manufacturing scene.

[0081] In these examples, in any case the domain adaptive neural network can be configured to infer spatially reconstructable manufacturing tools (e.g., spatially reconstructable manufacturing robots) in a real manufacturing scene. In these examples, the improvements brought by previous domain adaptive inference methods are particularly emphasized. In particular, it is difficult to rely on a training set of a large number of real manufacturing scenes including spatially reconstructable manufacturing tools, because such a set is usually confidential and / or publicly available but not adequately annotated. On the other hand, it is particularly convenient to use a large data set of virtual manufacturing scenes including spatially reconstructable manufacturing tools as a training set, because such a data set can be easily obtained and automatically annotated, for example, by using software capable of generating virtual simulations in a virtual manufacturing environment. At least for these reasons, the domain adaptive neural network may have been learned on a training set of virtual manufacturing scenes. Therefore, the preprocessing of the test data set discussed above is performed by determining an intermediate domain, so as to consider the training set of the domain adaptive neural network by cleverly making the real manufacturing scenes of the test data set closer in terms of data distribution to the scenes of the test scenario, in order to ensure better output quality of the domain adaptive neural network.

[0082] The domain adaptive inference method and the domain adaptive learning method can be performed independently. In particular, the domain adaptive neural network provided according to the domain adaptive inference method may have been learned according to the domain adaptive learning method or by any other machine learning method. Providing the domain adaptive neural network according to the domain adaptive inference method can include, for example, remotely accessing a data storage medium through a network, where the domain adaptive neural network has been stored on the data storage medium after it has been learned (e.g., according to the domain adaptive learning method) and retrieved from a database.

[0083] In the examples, the data obtained from the training domain includes scenes of another intermediate domain. In these examples, a domain-adaptive neural network has been learned on another intermediate domain. In terms of data distribution, the other intermediate domain is closer to the intermediate domain than the training domain. In such examples, the domain-adaptive neural network can be a domain-adaptive neural network that can be learned according to the domain-adaptive learning method. The test data set according to the domain-adaptive learning method can be equal to the test data set according to the domain-adaptive inference method, and the test domain according to the domain-adaptive inference method can be equal to the second domain according to the domain-adaptive learning method. The training domain according to the domain-adaptive inference method can be equal to the first domain according to the domain-adaptive learning method, and the data obtained from the training domain can belong to (e.g., form) the third domain determined according to the domain-adaptive learning method. The other intermediate domain can be the third domain determined according to the domain-adaptive learning method.

[0084] These examples combine the improvements brought by the previously discussed domain-adaptive learning method and domain-adaptive inference method: using a large data set of virtual scenes with correct annotations as the training set of the domain-adaptive neural network, a preprocessing stage of the training set performed by determining the third domain (or another intermediate domain), which improves the quality and / or accuracy of the output of the domain-adaptive neural network, and a test data set of real scenes performed by determining the intermediate domain, which further improves the quality and / or accuracy of the output of the domain-adaptive neural network.

[0085] Alternatively, the domain-adaptive learning method and the domain-adaptive inference method can be integrated into the same computer-implemented process. Figure 1 A flowchart illustrating the process is shown and the process will now be discussed.

[0086] The process includes an offline stage that integrates the domain-adaptive learning method. The offline stage includes providing S10 a data set of virtual scenes and a test data set of real scenes according to the domain-adaptive learning method. The data set of virtual scenes belongs to the first domain, which forms (theoretically, as described above) the training domain of the domain-adaptive neural network learned according to the domain-adaptive learning method. The test data set of real scenes belongs to the second domain, which forms the test domain of the domain-adaptive neural network learned according to the domain-adaptive learning method. The offline stage also includes determining S20 the third domain according to the domain-adaptive learning method. In terms of data distribution, the third domain is closer to the second domain than the first domain. The offline stage also includes learning 30 of the domain-adaptive neural network according to the domain-adaptive learning method.

[0087] After the offline phase, it is provided by the processing of the domain adaptation neural network learned according to the domain adaptation learning method. In other words, the domain adaptation neural network can form the output of the offline phase. After the offline phase, there can be a phase of storing the learned domain adaptation neural network, for example, storing it on the data storage medium of the device.

[0088] The process also includes an online phase, which integrates the domain adaptation inference method. The online phase includes providing, according to the domain adaptation inference method, the domain adaptation neural network learned during the offline phase S40. It is worth noting that providing the domain adaptation neural network S40 can include, for example, remotely accessing via a network, such as a data storage medium (on which the domain adaptation neural network has been stored at the end of the offline phase), and retrieving the domain adaptation neural network from the database. The online phase also includes determining, according to the domain adaptation inference method, an intermediate domain S50. In terms of data distribution, the intermediate domain is closer to the first domain than the second domain. It can be understood that, in the context of this process: the first domain of the domain adaptation learning method is the training domain of the domain adaptation inference method, the second domain of the domain adaptation learning method is the test domain of the domain adaptation inference method, the test data set of the domain adaptation learning method is the test data set of the domain adaptation inference method, and the data obtained from the training domain of the domain adaptation inference method belongs to (e.g., in the form of) a third domain. The online phase can also include inferring S60 of an object that can be spatially reconstructed according to the domain adaptation inference method.

[0089] The process combines the improvements brought by the previously discussed domain adaptation learning method and domain adaptation inference method: using a large data set with correctly annotated virtual scenes as the domain adaptation neural network training set, performing a preprocessing phase of the training set by determining a third domain, which improves the quality and / or accuracy of the output of the domain adaptation neural network, and a preprocessing phase of the test data set of the real scene performed by determining an intermediate domain, which further improves the quality and / or accuracy of the output of the domain adaptation neural network.

[0090] This process can constitute a process for inferring a spatially reconfigurable manufacturing tool (e.g., a spatially reconfigurable manufacturing robot) in a real manufacturing scenario. In fact, in an example of this process, the virtual scenario of the first domain can be each virtual manufacturing scenario including one or more spatially reconfigurable manufacturing tools, and the real scenario of the second domain can be each real manufacturing scenario including one or more spatially reconfigurable manufacturing tools. Thus, in these examples, the domain-adaptive neural network learned during the offline phase is configured to infer a spatially reconfigurable manufacturing tool in a real manufacturing scenario. The offline phase benefits from the previously discussed improvements provided by the domain-adaptive learning method in the specific case of manufacturing tool inference: virtual manufacturing scenarios are easy to obtain in large quantities and correctly annotated, and the preprocessing of the previously discussed training set can make the scenarios of the training set closer to the real scenarios of the test data set, which improves the robustness and accuracy of the learned domain-adaptive neural network. Combined with these improvements are the improvements brought by the domain-adaptive inference method integrated in the online phase of the process. These improvements include the preprocessing of the previously discussed test data set to make the real scenarios of the test data set closer to the scenarios of a third domain on which the domain-adaptive neural network has been learned during the offline phase, thereby improving the quality of the output of the domain-adaptive neural network. As a result, inferring a spatially reconfigurable manufacturing tool S60 in a real manufacturing scenario during the online phase is particularly accurate and robust. Thus, this process can constitute a particularly robust and accurate process for inferring a spatially reconfigurable manufacturing tool in a real manufacturing scenario.

[0091] The domain-adaptive learning method, the domain-adaptive inference method, and this process are computer-implemented. Now, the concept of a computer-implemented method (or process) is discussed.

[0092] "A method (or process) is computer-implemented" means that the steps (or substantially all steps) of the method (or process) are performed by at least one computer or any similar system. Thus, the steps of the method (or process) are performed by a computer, which may be fully automated or semi-automated. In an example, the triggering of at least some steps of the method (or process) can be performed through user-computer interaction. The required level of user-computer interaction can depend on the expected level of automation and be balanced with the need to fulfill the user's wishes. In an example, this level can be user-defined and / or pre-defined.

[0093] A typical example of the computer implementation of a method (or process) is to execute the method (or process) using a system suitable for this purpose. The system may include a processor coupled to a memory and a graphical user interface (GUI), on which a computer program is recorded, which includes instructions for executing the method (or process). The memory may also store a database. The memory is any hardware suitable for such storage and may include several physically distinct parts (e.g., one for the program and perhaps one for the database).

[0094] Figure 5 An example of the GUI of the system is shown, where the system is a CAD system.

[0095] The GUI 2100 can be a typical CAD-like interface with standard menu bars 2110, 2120 and bottom and side toolbars 2140, 2150. Such menus and toolbars contain a set of user-selectable icons, each icon associated with one or more operations or functions known in the art. Some of these icons are associated with software tools suitable for editing the 3D modeling object 2000 displayed in the GUI 2100 and / or working on the 3D modeling object 2000 displayed in the GUI 2100. The software tools can be grouped into workbenches. Each workbench contains a subset of the software tools. In particular, one of the workbenches is an editing workbench suitable for editing the geometric features of the modeled product 2000. In operation, a designer can, for example, pre-select a part of the object 2000 and then initiate an operation (e.g., change dimensions, color, etc.) or edit geometric constraints by selecting an appropriate icon. For example, a typical CAD operation is to model the punching or folding of a 3D modeling object displayed on the screen. The GUI can, for example, display data 2500 related to the product 2000 being displayed. In the example of this figure, the data 2500 shown as a "feature tree" and its 3D representation 2000 relate to a brake assembly including a brake caliper and a disc. The GUI can also show various types of graphical tools 2130, 2070, 2080, for example, for facilitating the 3D orientation of the object, for triggering simulations of operations on the edited product or for presenting various properties of the product 2000 being displayed. The cursor 2060 can be controlled by a haptic device to allow the user to interact with the graphical tools.

[0096] Figure 6 An example of the system is shown, where the system is a client computer system, such as a user's workstation.

[0097] The client computer of this example includes a central processing unit (CPU) 1010 connected to an internal communication bus 1000, and a random access memory (RAM) 1070 also connected to the bus. The client computer is also provided with a graphics processing unit (GPU) 1110, which is associated with a video random access memory 1100 connected to the bus. The video RAM 1100 is also known as a frame buffer in the art. A mass storage device controller 1020 manages access to a mass storage device (e.g., a hard disk drive 1030). Mass storage devices suitable for tangibly embodying computer program instructions and data include all forms of non-volatile memory, including, for example, semiconductor storage devices such as EPROM, EEPROM, and flash memory devices; magnetic disks such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM disks 1040. Any of the foregoing may be supplemented by, or incorporated in, a specially designed ASIC (application specific integrated circuit). A network adapter 1050 manages access to a network 1060. The client computer may also include a haptic device 1090, such as a cursor control device, a keyboard, etc. A cursor control device is used in the client computer to allow a user to selectively position a cursor at any desired location on a display 1080. In addition, the cursor control device allows the user to select various commands and input control signals. The cursor control device includes a plurality of signal generating devices for inputting control signals into the system. Generally, the cursor control device may be a mouse, and the buttons of the mouse are used to generate signals. Alternatively or additionally, the client computer system may include a sensitive pad and / or a sensitive screen.

[0098] The computer program may include instructions executable by a computer, the instructions including units for causing the above system to perform a domain adaptation learning method, a domain adaptation inference method, and / or a process. The program may be recordable on any data storage medium, including the memory of the system. The program may be implemented, for example, in digital electronic circuitry or in computer hardware, firmware, software, or in combinations thereof. The program may be implemented as an apparatus, for example, a product tangibly embodied in a machine-readable storage device for execution by a programmable processor. Processing / method steps (i.e., steps of the domain adaptation learning method, the domain adaptation inference method, and / or the process) may be performed by a programmable processor executing an instruction program to perform the functions of the process by operating on input data and generating output. Thus, the processor may be programmable and coupled to receive data and instructions from a data storage system, at least one input device, and at least one output device, and to send data and instructions to the data storage system, at least one input device, and at least one output device. If desired, the application program may be implemented in a high-level procedural or object-oriented programming language or in assembly or machine language. In any case, the language may be a compiled or interpreted language. The program may be a complete installation program or an update program. In any case, the application of the program on the system results in instructions for performing the domain adaptation learning method, the domain adaptation inference method, and / or the process.

[0099] The concept of providing a scenario data set will be discussed, relating to a domain adaptation learning method, a domain adaptation inference method, and a process. Before discussing this concept, the data structures involved will now be discussed. As will be recognized, in the case of providing a domain adaptation learning method, a domain adaptation inference method, and / or a process, the data structure definitions and examples provided herein may be applied to at least a part (e.g., all) of any data set.

[0100] In the context of the present disclosure, a scenario specifies the arrangement (e.g., disposition) of one or more objects in a background at a point in time. The background generally represents the physical environment of the real world, the objects generally represent physical objects in the real world, and the arrangement of the objects generally represents the deployment of real-world objects in the real-world environment at that point in time. This is equivalent to saying that a scenario is a representation of real-world objects in their real-world environment at a certain point in time. The representation may be real, geometrically real, semantically real, photo-realistic, and / or virtual. The concepts of real, geometrically real, semantically real, photo-realistic, and virtual scenarios will be discussed below. A scenario is generally (but not always) three-dimensional. For simplicity, hereinafter, no distinction will be made between a scenario (e.g., virtual or real) and its representation (e.g., represented by a virtual or real image).

[0101] In the context of the present disclosure, a dataset of scenarios typically includes a large number of scenarios, such as more than 1000, 10000, 100000, or 1000000 scenarios. Any dataset of scenarios of the present disclosure can consist of or substantially consist of scenarios from the same context (e.g., manufacturing scenarios, construction site scenarios, port scenarios, or apron scenarios). The scenarios involved in the domain adaptation learning method, the domain adaptation inference method, and / or the process can actually be all the scenarios from the same context (e.g., manufacturing scenarios, construction site scenarios, port scenarios, or apron scenarios).

[0102] Any scenario involved in the domain adaptation learning method, the domain adaptation inference method, and / or the process can include one or more spatially reconstructable objects. A spatially reconstructable object is an object that includes several (e.g., mechanical) parts, and is characterized in that there is one or more semantically realistic spatial relationships between the parts of the object. A spatially reconstructable object can typically be an (e.g., mechanical) assembly of parts that are physically connected to each other, and the assembly has at least one degree of freedom. This means that at least a first part of the object is movable relative to at least a second part of the object, while both the at least first part and the at least second part remain physically attached to the assembly. For example, the at least first part can translate and / or rotate relative to the at least second part. Thus, a spatially reconstructable object can have a very large number of spatial configurations and / or positions, which makes the object difficult to recognize for a person without appropriate skills, e.g., difficult to label. In the context of the present disclosure, a spatially reconstructable object can be a spatially reconstructable manufacturing tool (e.g., an industrial articulated robot) in a manufacturing scenario, a crane in a construction site scenario, an air passage in an apron scenario, or a port crane in a port scenario.

[0103] In the context of the present disclosure, a scenario can be a manufacturing scenario. A manufacturing scenario is a scenario that represents real-world objects in a real-world manufacturing environment. The manufacturing environment can be a factory or a part of a factory. The objects in the factory or a part of the factory can be manufactured products, i.e., products that have been manufactured through one or more manufacturing processes performed at the factory at a time point related to the scenario. The objects in the factory or a part of the factory can also be products being manufactured, i.e., products that are to be manufactured through one or more manufacturing processes performed at the factory at a time point related to the scenario. The objects in the factory or a part of the factory can also be (e.g., spatially reconstructable) manufacturing tools, which are tools involved in one or more manufacturing processes performed at the factory. The objects in the factory can also be objects that constitute the background of the factory.

[0104] In other words, a scene representing a factory or a part of a factory may thus include one or more products (each product being manufactured or having been manufactured by one or more manufacturing processes of the factory) and one or more manufacturing tools, for example each being involved in the manufacture of one or more of the said products. A product may generally be a (for example, mechanical) part or an assembly of parts (or equivalently an assembly of parts, since from the point of view of the present disclosure, an assembly of parts may be regarded as a part itself).

[0105] Mechanical parts may be part of a land vehicle (including, for example, automotive and light truck equipment, racing cars, motorcycles, truck and motor vehicle equipment, trucks and buses, trains), an aircraft (including airframe equipment, aerospace equipment, propulsion equipment, defense products, aviation equipment, space equipment), a naval vehicle (including naval equipment, commercial ships, marine equipment, yachts and workboats, ship equipment), general mechanical parts (including, for example, industrial manufacturing machinery, heavy mobile machinery or equipment, installed equipment, industrial equipment products, metal products, tire products, articulated and / or reconfigurable manufacturing equipment (such as robotic arms)), electromechanical or electronic parts (including, for example, consumer electronics, safety and / or control and / or instrumentation products, computing and communication equipment, semiconductors, medical devices and equipment), consumer products (including, for example, furniture, home and garden products, leisure products, fashion products, products of hard goods retailers, products of soft goods retailers), packaging (including food and beverage and tobacco, beauty and personal care, household product packaging).

[0106] Mechanical parts may also be one or a possible combination of the following: molded parts (i.e., parts manufactured by a molding manufacturing process), machined parts (i.e., parts manufactured by a machining manufacturing process), drilled parts (i.e., parts manufactured by a drilling manufacturing process), turned parts (i.e., parts manufactured by a turning manufacturing process), forged parts (i.e., parts manufactured by a forging manufacturing process), stamped parts (i.e., parts manufactured by a stamping manufacturing process) and / or folded parts (i.e., parts manufactured by a folding manufacturing process).

[0107] Manufacturing tools may be any of the following:

[0108] - machining tools (i.e., tools that perform at least part of a machining process), such as broaching machines, drilling machines, gear formers, hobbing machines, grindstones, lathes, screw machines, milling machines, sheet metal shears), sizing machines, saws, planers, Stewart platform milling machines, grinders, multi-tasking machines (e.g., having multiple axes that combine turning, milling, grinding and / or material handling into a highly automated machine tool);

[0109] - Compression molding machines (i.e., machines that perform at least part of the compression molding process), typically including at least one mold or molding matrix, such as quick plunger type molds, straight plunger type or floor plunger type molds;

[0110] - Injection molding machines (i.e., machines that perform at least part of the injection molding process), such as die casting machines, metal injection molding machines, plastic injection molding machines, liquid silicone injection molding machines or reaction injection molding machines;

[0111] - Rotary cutting tools (e.g., that perform at least part of a drilling process), such as drill bits, countersinks, flat bottom reamers, taps, dies, milling cutters, reamers or cold saw blades;

[0112] - Non-rotary cutting tools (e.g., that perform at least part of a turning process), such as pointed tools or profiling tools;

[0113] - Forging machines (i.e., that perform at least part of the forging process), such as mechanical forging presses (usually including at least one forging die), hydraulic forging presses (usually including at least one forging die), air-driven forging hammers, electric-driven forging hammers, hydraulically-driven forging hammers or steam-driven forging hammers;

[0114] - Die stamping machines (i.e., that perform at least part of the die stamping process), such as die stamping machines, mechanical presses, stamping presses, blanking presses, embossing presses, bending presses, flanging presses or die casting machines;

[0115] - Bending or folding machines (i.e., that perform at least part of the bending or folding process), such as box and pan brakes, press brakes, folder machines, plate bending machines or presses;

[0116] - Spot welding robots; or

[0117] - Electric paint spraying tools (i.e., that perform at least part of the product painting process), such as (e.g., electric) automotive paint spraying robots.

[0118] Any of the manufacturing tools listed above can be a spatially reconfigurable manufacturing tool, i.e., a manufacturing tool that is spatially reconfigurable, or can form at least part of a spatially reconfigurable manufacturing tool. For example, a manufacturing tool held by a tool holder of a spatially reconfigurable manufacturing tool forms at least part of a spatially reconfigurable manufacturing tool. A spatially reconfigurable manufacturing tool (e.g., any of the tools listed above) can be a manufacturing robot, such as (e.g., articulated) industrial robots.

[0119] Now discuss the concept of a real scenario. As previously mentioned, a scenario represents the arrangement of real-world objects in a real-world environment (i.e., the background of the scenario). A scenario is real when the deployed representation is a captured physical arrangement that is derived from the same physical arrangement in the real world. This capture can be performed by a digital image acquisition process that creates a digitally encoded representation of the visual characteristics of the physical arrangement of the objects in their environment. For simplicity, this digitally encoded representation may hereinafter be referred to as an "image" or a "digital image". The digital image acquisition process may optionally include the processing, compression, and / or storage of such images. Digital images can be created directly from the physical environment by a camera or similar device, a scanner or similar device, and / or a video camera or similar device. Alternatively, a digital image can be obtained from another image in a similar medium, such as a photograph, photographic film (e.g., the digital image is a snapshot of a film instant), or printed paper, by an image scanner or similar device. Digital images can also be obtained by processing non-image data, such as non-image data collected by tomographic equipment, scanners, side-scan sonar, radio telescopes, X-ray detectors, and / or photostimulable phosphor plates (PSP). In the context of the present disclosure, a real manufacturing scenario can, for example, originate from a digital image obtained by a camera (or similar device) located in a factory. Alternatively, a manufacturing scenario can originate from a snapshot of a film captured by one or more video cameras located in a factory.

[0120] Figures 7 to 12 Examples of real manufacturing scenarios are shown, each scenario respectively including articulated industrial robots 70 and 72, 80, 90, 92, 94, 96 and 98, 100, 102 and 104, 110 and 112, 120, 122 and 124.

[0121] Now discuss the concept of a virtual scene. When a scene is derived from computer-generated data, the scene is virtual. Such data can be generated by, for example, (3D) simulation software, (3D) modeling software, CAD software, or video games. This is equivalent to saying that the scene itself is computer-generated. The objects of a virtual scene can be placed at realistic locations within the virtual scene. Alternatively, the objects of a virtual scene can be placed at non-realistic locations within the virtual scene, such as when the objects are randomly placed within an already existing virtual scene, as will be discussed specifically below. This is equivalent to saying that a virtual scene represents the deployment of virtual objects (e.g., virtual representations of real-world objects) within a virtual background (e.g., a virtual representation of a real-world environment). For at least a portion of the virtual objects, the deployment can be physically realistic, i.e., the deployment of at least that portion of the virtual objects represents the real-world physical arrangement of real-world objects. Additionally or alternatively, at least another portion of the virtual objects can be arranged (e.g., relative to each other and / or relative to that at least portion of the objects) in a non-realistic deployment (i.e., not corresponding to a real-world physical arrangement), such as when at least that portion of the virtual objects are randomly placed or arranged, as will be discussed specifically below. Any object of a virtual scene can be automatically annotated or tagged. The concept of tagging objects will be discussed further below. A virtual manufacturing scene can be generated by 3D modeling software for a manufacturing environment, 3D simulation software for a manufacturing environment, or CAD software. Any object (e.g., a product or a manufacturing tool) within a virtual manufacturing scene can be virtually represented by a 3D modeling object that represents the skin (e.g., the outer surface) of an entity (e.g., a B-rep model). The 3D modeling object may have been designed by a user with a CAD system and / or may be sourced from a virtual manufacturing simulation performed by 3D simulation software for a manufacturing environment.

[0122] A modeling object is any object defined by data stored, for example, in a database. By extension, the expression "modeling object" refers to the data itself. Depending on the type of system that performs a domain adaptation learning method, a domain adaptation inference method, and / or a process, a modeling object can be defined by different kinds of data. The system can indeed be any combination of a CAD system, a CAE system, a CAM system, a PDM system, and / or a PLM system. In those different systems, a modeling object is defined by the corresponding data. Thus, one can refer to CAD objects, PLM objects, PDM objects, CAE objects, CAM objects, CAD data, PLM data, PDM data, CAM data, CAE data. However, these systems are not mutually exclusive, since a modeling object can be defined by data corresponding to any combination of these systems. It will be apparent from the definition of such a system provided below that the system can thus also be a CAD and PLM system.

[0123] A CAD system additionally means any system, such as CATIA, that is at least suitable for designing modeling objects based on a graphical representation of the modeling objects. In this case, the data defining the modeling objects includes data that allows the representation of the modeling objects. A CAD system can provide a representation of a CAD modeling object, for example, using edges or lines (in some cases with faces or surfaces). The lines, edges, or surfaces can be represented in various ways, such as non-uniform rational B-splines (NURBS). Specifically, a CAD file contains specifications from which geometries can be generated, which in turn allows the representation form to be generated. The specifications of the modeling objects can be stored in a single CAD file or multiple CAD files. The typical size of a file representing a modeling object in a CAD system is in the range of 1 MB per part. And a modeling object can typically be an assembly of thousands of parts.

[0124] A CAM solution additionally means any solution, software for hardware, that is suitable for managing the manufacturing data of a product. The manufacturing data typically includes data related to the product to be manufactured, the manufacturing process, and the required resources. A CAM solution is used to plan and optimize the entire manufacturing process of a product. For example, it can provide information to a CAM user about the feasibility, the duration of the manufacturing process, or the quantity of resources (e.g., a specific robot) that can be used at a specific step layer of the manufacturing process; and thus, allows a decision on the management or required investment. CAM is a subsequent process after the CAD process and a potential CAE process. Such CAM solutions are provided by Dassault Systèmes under the trademark Provided.

[0125] In the context of CAD, a modeling object can typically be a 3D modeling object, such as representing a product (e.g., a part or an assembly of parts), or possibly a product assembly. A "3D modeling object" means any object modeled by data that allows its 3D representation. The 3D representation allows the part to be viewed from all angles. For example, when 3D represented, a 3D modeling object can be manipulated and rotated around any of its axes or around any axis in the screen on which the representation is displayed. In particular, this does not include 2D icons that are not 3D modeled. The display of the 3D representation helps in the design (i.e., increases the speed at which designers statistically complete their tasks). Since the design of a product is part of the manufacturing process, this can speed up the manufacturing process in the industry.

[0126] Figures 13 to 17 A virtual manufacturing scenario is shown, each virtual scenario respectively including virtual articulated industrial robots 130, 140, 150 and 152, 160 and 162, and 170 and 172.

[0127] The virtual scene can be geometrically and semantically realistic, meaning that the geometry and functionality of virtual objects in the scene (e.g., products or manufacturing tools in a manufacturing scene) are realistic, e.g., realistically representing their corresponding geometry and functionality in the real world. The virtual scene can be geometrically and semantically true, rather than photo-realistic. Alternatively, the virtual scene can be geometrically and semantically realistic and can also be photo-realistic.

[0128] The concept of a photo-realistic image can be understood in terms of multiple visible features, such as: shading (i.e., how the color and brightness of a surface vary with illumination), texture mapping (i.e., the method of applying detail to a surface), bump mapping (i.e., the method of simulating small-scale bumps on a surface), fog / m participating media (i.e., how light darkens when passing through an opaque atmosphere or air), shadows (i.e., the effect of blocking light), soft shadows (i.e., the varying darkness caused by a partially occluded light source), reflections (i.e., specular or highly glossy reflections), transparency or opacity (i.e., the sharp transmission of light through a solid object), translucency (i.e., the highly scattered transmission of light through a solid object), refraction (i.e., the bending of light associated with transparency), diffraction (i.e., the bending, spreading, and interference of light rays passing through an object or hole that interferes with the rays), indirect illumination (i.e., the surface is illuminated by light reflected from other surfaces rather than directly from a light source (also known as global illumination)), caustics (i.e., a form of indirect illumination where light is reflected from a luminous object and / or focused through a transparent object to produce a bright highlight on another object), depth of field (i.e., an object appears blurry or out of focus at a distance too far in front of or behind the focused object), and / or motion blur (i.e., an object appears blurry due to high-speed motion or the motion of the camera). The presence of such characteristics makes such data (i.e., corresponding to or from which the virtual photo-realistic scene is derived) have a similar distribution to real data (from which the image representing the real scene is derived). In this context, the closer the virtual set distribution is to real-world data, the more photo-realistic it appears.

[0129] Now discuss the concept of data distribution. In the context of building a machine learning (ML) model, it is generally expected that the model is trained on data from the same distribution and tested against data from the same distribution. In probability theory and statistics, a probability distribution is a mathematical function that provides the probabilities of different possible outcomes in an experiment. In other words, a probability distribution is a description of a random phenomenon in terms of the probabilities of events. A probability distribution is defined with respect to a underlying sample space, which is the set of all possible outcomes of the observed random phenomenon. The sample space can be a set of real numbers or a higher-dimensional vector space, or it can be a non-numerical list. It can be understood that when training on virtual data and testing on real data, one faces the problem of distribution shift. When virtual images are modified with real images to address the shift problem, domain adaptation learning methods, domain adaptation inference methods, and / or processes can transform the distribution of virtual data to better align with the distribution of real data. To quantify how well two distributions are aligned, several methods can be used. For illustration, here are two examples of methods.

[0130] The first method is the histogram method. Consider the case of creating N pairs of images. Each pair of images includes a first created virtual image and a second created virtual image, where the second created virtual image corresponds to a real image representing the same scene as the first virtual image, and this real image is obtained by using a virtual-to-real generator. To evaluate the relevance of modifying the training data, domain adaptation learning methods, domain adaptation inference methods, and / or processes quantify the degree of alignment of the virtual image distribution with the real image data before and after the synthetic-to-real modification. To this end, domain adaptation learning methods, domain adaptation inference methods, and / or processes can consider a set of real images and calculate the Euclidean distance between the mean image of the real data and the virtual data before and after the synthetic-to-real modification. Formally, consider D v ={v1,...,v N} and D v′ ={v′1,...,v′ N}, deep features are derived from a pre-trained network, such as being trained on the ImageNet classification task (starting network, see

[22] ) for virtual images before and after the modification respectively. By using m real to represent the deep features derived from the mean image of the real dataset, the Euclidean differences are ED v ={v1 - m,...,v N - m} and ED v′ ={v ′ 1 - m,...,v ′ N - m}. To evaluate the alignment with the real data distribution, domain adaptation learning methods, domain adaptation inference methods, and / or processes use histogram comparison of EDv and ED v′ The distribution. A histogram is an accurate representation of the distribution of numerical data. It is an estimate of the probability distribution of a continuous variable. Figure 18 shows the ED v and ED v′ The histogram of the distribution of the Euclidean distance between. According to Figure 18 , compared with the distribution (blue) obtained when using the original virtual image ED v , when representing the ED from the modified virtual image v′ (red), the distribution is more concentrated around 0. This histogram validates the tested hypothesis, which assumes that in terms of data distribution, the modified virtual data is transferred to a new domain closer to the real-world data than the root virtual data. In terms of photo-reality, consider a test that asks people to visually guess whether a given image is artificial or real. It seems obvious that the modified image is more likely to fool humans visually compared to the original virtual image. This is directly related to the alignment of the distributions evaluated by the above histogram.

[0131] The second method is the Frechet Inception Distance (FID) method (see

[21] ). FID compares the distributions of the Inception embeddings of two sets of images (the activations from the penultimate layer of Inception, see

[22] ). Both of these distributions are modeled as multi-dimensional Gaussian models, parameterized by their respective means and covariances. Formally, let D A = {(x A )} be the first dataset of scenes (e.g., the first dataset of real scenes), and let D B = {(x B )} be the second dataset (e.g., the second dataset of virtual scenes). The second method first uses a feature extractor network derived from a pre-trained network (the initial network, see

[22] ) trained on, for example, the ImageNet classification task to extract the deep features v A and v B from x A and x B respectively. Next, the second method collects the features in v A . This will result in multiple C-dimensional vectors, where C is the feature dimension. Then, the second method is to fit a multivariate Gaussian to these feature vectors by calculating the mean and covariance matrix . Similarly, we have and The FID between these extracted features is:

[0132]

[0133] The smaller the distance, the DA and D B The closer the distribution is to D. From the perspective of photo-realism, when the first dataset is a dataset of real scenes and the second dataset is a dataset of virtual scenes, the smaller the distance, the more photo-realistic images from D B there are. This metric allows the evaluation of the relevance of modifying our data from one domain to another in terms of distribution alignment.

[0134] After the previous discussion about data distributions, the distance between two datasets of a scene can be quantified by distances such as the Euclidean distance or FID discussed above. Thus, in the context of the present disclosure, "the first dataset of a scene (e.g., a family, e.g., a domain, the concept of a domain discussed below) is closer to the second dataset of a scene (e.g., a family, e.g., a domain) than the third dataset of a scene (e.g., a family, e.g., a domain)" means that: the distance between the first dataset (or family or domain) and the second dataset (or family or domain) is less than the distance between the third dataset (or family or domain) and the second dataset (or family or domain). In the case where the second dataset is a dataset of real scenes and both the third and first datasets are datasets of virtual scenes, it can be said that the first dataset is more photo-realistic than the third dataset.

[0135] Now discuss the concept of "domain". A domain specifies the distribution of a family of scenes. Thus, a domain is also a family of distributions. If the scenes of a dataset and the scenes of a family are close in terms of data distribution, the dataset of the scene belongs to the domain. In an example, this means that there is a predetermined threshold such that the distance between the scene of the dataset and the family scene (e.g., Euclidean or FID as described above) is less than the predetermined threshold. It should be understood that a dataset of a scene can also form a domain itself, in which case the domain can be referred to as "the domain of the scene" or "the domain constituted by the scenes of the dataset". The first domain according to the domain adaptation learning method is a domain of virtual scenes, such as a domain of virtual scenes that are both geometrically and semantically realistic but not photo-realistic. The second domain according to the domain adaptation learning method is a domain of real scenes. The test domain according to the domain adaptation learning method can be a domain of real scenes. Alternatively, the test domain can be a domain of virtual scenes, e.g., geometrically and semantically realistic and optionally photo-realistic. The training domain according to the domain adaptation learning method can be a domain of virtual scenes, e.g., geometrically and semantically realistic and optionally photo-realistic. In an example, the training domain is a domain of virtual scenes that is less photo-realistic than the scenes of the test domain, e.g., the training domain is a domain of virtual scenes while the test domain is a domain of real scenes.

[0136] Now discuss the concept of providing a dataset of a scene.

[0137] In the context of the present disclosure, providing a dataset of scenarios can be performed automatically or by a user. For example, a user can retrieve a dataset from a memory storing the dataset. The user can further choose to complete the dataset by adding (e.g., one by one) one or more scenarios to the already retrieved dataset. Adding one or more scenarios can include retrieving them from one or more memories and including them in the dataset. Alternatively, a user can create at least a portion (e.g., all) of the scenarios of the dataset. For example, as previously discussed, when providing a dataset of virtual scenarios, a user can first obtain at least a portion (e.g., all) of the scenarios by using 3D simulation software and / or 3D modeling software and / or CAD software. More generally, in the context of a domain adaptation learning method, a domain adaptation inference method, and / or a process, the step of providing a dataset of virtual scenarios can be preceded by a step of obtaining (e.g., computing) the scenarios, such as by using 3D simulation software (e.g., for a manufacturing environment) and / or 3D modeling software and / or CAD software, as previously discussed. Similarly, in the context of a domain adaptation learning method, a domain adaptation inference method, and / or a process, the step of providing a dataset of real scenarios can be preceded by a step of obtaining (e.g., acquiring) the scenarios, such as by a digital image acquisition process as previously described, as previously discussed.

[0138] In an example of a domain adaptation learning method and / or process, each virtual scenario of a dataset of virtual scenarios is a virtual manufacturing scenario including one or more spatially reconfigurable manufacturing tools. In these examples, a domain adaptation neural network is configured to infer spatially reconfigurable manufacturing tools in a real manufacturing scenario. In these examples, the virtual scenarios of the dataset of virtual scenarios can all be sourced from a simulation. Thus, before the step S10 of providing the dataset of virtual scenarios, a step of performing a simulation can be carried out, for example, using 3D simulation software in a manufacturing environment, as previously discussed. In the example, the simulation can be a 3D experience (i.e., having virtual scenarios with precise modeling and having faithful (e.g., compliant, e.g., conforming to reality) behavior), where different configurations of a given scenario are simulated, such as different arrangements of objects in a given scenario and / or different configurations of certain objects (e.g., spatially reconfigurable objects) in a given scenario. The dataset can particularly include virtual scenarios corresponding to the different configurations.

[0139] In examples of domain adaptation learning methods and / or processes, each real scenario of the test data set is a real manufacturing scenario that includes one or more spatially reconstructable manufacturing tools. In these examples, a domain adaptation neural network is configured to infer the spatially reconstructable manufacturing tools in the real manufacturing scenario. In these examples, the real scenarios of the test data set can all be sourced from images acquired during a digital acquisition process. Thus, before providing the test data set of the S10 virtual scenario, the following steps can be: obtaining an image of a factory, from which the scenarios of the test data set are sourced, for example by using one or more scans and / or one or more cameras and / or one or more video cameras.

[0140] In examples of domain adaptation inference methods and / or processes, the test data set includes real scenarios. In these examples, each real scenario of the test data set can be a real manufacturing scenario that includes one or more spatially reconstructable manufacturing tools, and the domain adaptation neural network is optionally configured to infer the spatially reconstructable manufacturing tools in the real manufacturing scenario. Thus, in providing the test data set of the S10 virtual scenario, the following steps can be: obtaining an image of a factory, from which the scenarios of the test data set are sourced, for example by using one or more scans and / or one or more cameras and / or one or more video cameras.

[0141] Now, the determination of the third domain S20 is discussed.

[0142] The determination of the third domain S20 makes the third domain closer to the second domain than the first domain in terms of data distribution. Thus, the determination of the third domain S20 specifies the inference of another scenario for each virtual scenario of the data set of virtual scenarios belonging to the first domain, such that all the inferred other scenarios form (or at least belong to) a domain, i.e., the third domain, which is closer to the second domain than the first domain in terms of data distribution. The inference of the other scenario can include a transformation (e.g., conversion) of the virtual scenario to the other scenario. The transformation can include, for example, using a generator capable of transferring an image from one domain to another to transform the virtual scenario into the other scenario. It should be understood that such a transformation can only involve one or more parts of the virtual scenario: the one or more parts are transformed, for example using a generator, while the rest of the scenario remains unchanged. The other scenario can include a scenario in which one or more parts are transformed while the rest of the scenario remains unchanged. Alternatively, the transformation can include replacing the rest of the scenario with a part of some other scenario (e.g., belonging to the second domain). In such a case, the other scenario can include a scenario in which one or more parts are transformed while the rest of the scenario is replaced by the part of the other scenario. It should be understood that the inference can be similarly performed for all scenarios of the first domain, for example with the same paradigm applicable to all scenarios.

[0143] Now refer to Figure 2 an example of determining S20 of the third domain is discussed. Figure 2 A flowchart showing an example of determining S20 of the third domain is shown.

[0144] In the example, the determination S20 of the third domain includes, for each scene of the dataset of the virtual scene, extracting S200 one or more spatially reconstructable objects from the scene. In these examples, the determination S20 further includes, for each extracted object, converting S210 the extracted object into an object that is closer to the second domain in terms of data distribution than the extracted object.

[0145] As previously discussed, the virtual scene of the virtual scene dataset can be obtained, for example, by using 3D simulation software to calculate the virtual simulation from which the scene is derived, before the provision S10 of the virtual scene dataset. The calculation of the simulation can include the automatic positioning and annotation of each spatially reconstructable object in each virtual scene. For each virtual scene of the dataset, and for each spatially reconstructable object in the virtual scene, the positioning can include calculating a bounding box around the spatially reconstructable object. The bounding box is a rectangular box whose axes are parallel to the sides of the image representing the scene. The bounding box is characterized by four coordinates and is centered on the spatially reconstructable object at an appropriate scale and scale to enclose the entire spatially reconstructable object. In the context of the present disclosure, enclosing an object in a scene can mean including all parts of the object that are fully visible on the scene, which parts are the entire object or only a part of the object. The annotation can include assigning a label to each calculated bounding box, the label indicating that the bounding box contains a spatially reconstructable object. It should be understood that the label can indicate other information about the bounding box, such as the bounding box encloses a spatially reconstructable manufacturing tool, such as an industrial robot.

[0146] For each respective one of the one or more spatially reconstructable objects, extracting S200 the spatially reconstructable object from the virtual scene can specify any action that results in separating the spatially reconstructable object from the virtual scene, and can be performed by any method capable of performing such separation. In fact, extracting S200 the spatially reconstructable object from the virtual scene can actually include: detecting the bounding box enclosing the spatially reconstructable object, the detection being based on the label of the bounding box and optionally including accessing the label. The extraction S200 can further include separating the bounding box from the virtual scene, thereby also separating the spatially reconstructable object from the virtual scene.

[0147] The transformation S210 of the extracted object can specify any transformation that changes the extracted object into another object that is closer to the second domain in terms of data distribution. "Closer to the second domain" means that the object into which the extracted object is transformed is itself a scene that belongs to a domain that is closer to the second domain than the first domain. When the extraction S200 of the object includes separating the bounding box surrounding the object from the virtual scene as described above, the transformation S210 of the extracted object can include modifying the entire bounding box (and thus including the extracted object) or only the extracted object, where the modification of the entire bounding box or the extracted object only makes it closer to the second domain in terms of data distribution. The modification can include applying a generator that can convert a virtual image into a more photo-realistic image in terms of data distribution to the bounding box or the spatially reconstructable object with the bounding box.

[0148] In one embodiment, the transformation S210 can be performed by a CycleGAN virtual-to-real network (see

[20] ) trained on two sets of spatially reconstructable objects (virtual set and real set). It produces a virtual-to-real spatially reconstructable object (e.g., a robot) generator that is used to modify the spatially reconstructable object (e.g., a robot) in the virtual image. The CycleGAN model transfers images from one domain to another. It proposes a method that can learn to capture the special characteristics of an image set and figure out how to transform these features into a second image set without any paired training examples. In fact, this relaxation of having a one-to-one mapping makes this formulation very powerful. This is achieved through a generative model, specifically a generative adversarial network (GAN) called CycleGAN.

[0149] Figure 19 An example of the extraction S200 and transformation S210 of two virtual industrial articulated robots 190 and 192 is shown. Each of these robots 190 and 192 has been extracted S200 from the corresponding virtual scene. Then, the extracted robots 190 and 192 are transformed S210 into more photo-realistic robots 194 and 196, respectively.

[0150] Still referring to Figure 2 the flowchart of, the determination of the third domain can also include: for each transformed extracted object, placing S220 the transformed extracted object in one or more scenes, each of which is closer to the second domain than the first domain in terms of data distribution. In this case, the third domain includes each scene in which the transformed extracted object is placed.

[0151] Each of the one or more scenes can be a scene that does not include spatially reconstructable objects, such as a background scene (i.e., a scene including only the background, although the background itself can include background objects). Each of the one or more scenes can also be a scene from the same context as the scene to which the transformed extracted object originally belonged. In fact, this same context can be the context of all scenes of the dataset of the virtual scene.

[0152] Each of the one or more scenes is closer to the second domain than the first domain. "Closer to the second domain" means that each of the one or more scenes belongs to a domain that is closer to the second domain than the first domain. In an example, each of the one or more scenes is a real scene directly obtained from reality, as previously discussed. Alternatively, each of the one or more scenes can be generated from the corresponding virtual scene of the dataset of the virtual scene, one of the corresponding virtual scenes being, for example, the virtual scene to which the transformed extracted object belongs, such that the generated scene is more realistic and / or more photo-realistic than the corresponding virtual scene. Generating a scene from the corresponding virtual scene of the dataset can include applying a generator capable of converting a virtual image into a real and / or more realistic photo to the virtual scene. In one implementation, the generation of the scene can be performed by a CycleGAN virtual-to-real generator (see

[20] ) trained on two sets of backgrounds (e.g., factory backgrounds) (virtual set and real set). It produces a virtual-to-real background (e.g., factory background) generator that is used to modify the background (e.g., factory background) in the virtual image. The CycleGAN model (see

[20] ) transfers an image from one domain to another. It proposes a method that can learn to capture the special features of an image set and figure out how to transform these features into a second image set without any paired training examples. In fact, this relaxation of having a one-to-one mapping makes this formulation very powerful. This is achieved through a generative model, specifically a generative adversarial network (GAN) called CycleGAN.

[0153] The placement S220 of the transformed extracted object can be performed (e.g., automatically) by any method capable of pasting a single and / or multiple images of the object onto the background image. Placing S220 the transformed extracted object onto the one or more scenes can thus include applying such a method to paste the transformed extracted object onto each corresponding one of the one or more scenes. It should be understood that the scene on which the transformed extracted object is placed can include one or more other transformed extracted objects that have already been placed S220 on the scene.

[0154] From the examples discussed above, in these examples, the determination S20 of the third domain can particularly transform the spatially reconstructable objects of the scene and the background of the scene to which they belong separately, and then each transformed spatially reconstructable object can be placed into one or more transformed backgrounds. First, compared with the virtual scene of the dataset of the fully transformed (i.e., without separating the spatially reconstructable objects and their backgrounds) virtual scene, this allows for obtaining more transformed object configurations in the transformed background. Since the learning S30 of the domain adaptation neural network is based on the scene of the third domain, therefore, the learning is more robust because the third domain is closer to the test domain. Second, the virtual scene of the fully transformed virtual scene dataset may lead to insufficient results because there may be no fundamental relationship between the scenes of the test domain due to the actual background variability. For example, fully transforming a virtual scene including spatially reconstructable objects can result in a transformed scene where the spatially reconstructable objects cannot be distinguished from the background (e.g., mixed with the background). For all these reasons, the virtual scene of the fully transformed virtual scene dataset may lead to insufficient results, such as by generating a third domain that may not be closer to the second domain than the first domain. In addition, the scenes of the third domain may not be suitable for learning in this case because the characteristics of the objects of interest (spatially reconstructable objects) in these scenes may be affected by the transformation.

[0155] Figure 20 An example of the placement S220 of the transformed extracted object is shown. Figure 20 Examples of the placement S220 in the factory background scenes 200 and 202 are shown. Figure 19 of the transformed extracted industrial articulated robots 194 and 196.

[0156] Still referring to Figure 2 the flowchart of, in the example, the placement S220 of the transformed extracted object in one or more scenes is performed randomly. Such a random placement S220 of the transformed extracted object can be (e.g., automatically) performed by any method capable of randomly pasting a single and / or multiple images of the object into the background image. Therefore, randomly placing the transformed extracted object S220 in one or more scenes can include applying such a method to randomly paste the transformed extracted object onto each corresponding scene of the one or more scenes.

[0157] Randomly placing the transformed extraction objects in S220 improves the robustness of the learning S30 of the domain - adaptive neural network, which learning is based on a learning set constituted by the scenes of the third domain. In fact, randomly placing the transformed extraction objects ensures that during the learning S30, the domain - adaptive neural network does not focus on the positions of the transformed extraction objects, because the sparse range of such positions is too large for the learning S30 to rely on them and / or learn the correlation between the transformed extraction objects and their positions. Instead, the learning S30 can truly focus on the inference of objects that can be reconstructed spatially, regardless of their positions. Incidentally, not relying on the positions of objects that can be reconstructed spatially allows significantly reducing the risk that the domain - adaptive neural network misses spatially reconstructable objects that the domain - adaptive neural network should infer only because the objects do not occupy appropriate positions.

[0158] In the above example, for each transformed extraction object, the third domain includes each respective scene on which the transformed extraction object is placed S220. For example, the third domain is constituted by such scenes. In other words, for all scenes of the dataset of virtual scenes, the extraction S200, transformation S210, and placement S220 are performed similarly. For example, they are performed such that all scenes on which one or more transformed extraction objects are placed have substantially the same level of photo - realism. In other words, they belong to and / or form a domain, namely the third domain.

[0159] Still referring to Figure 2 the flowchart of, determining S20 that the third domain can include placing S230 one or more distractors in each respective one of the one or more scenes included in the third domain. A distractor is an object that is not an object that can be reconstructed spatially.

[0160] A distractor can be an object having substantially the same level of photo - realism as the transformed extraction objects. In fact, a distractor can be an object of a virtual scene of a dataset of virtual scenes, which object of the virtual scene is extracted from the virtual scene, transformed to be closer to an object of the second domain, and placed in the scenes of the third domain in a manner similar to that of spatially reconstructable objects of the virtual scene. Alternatively, the distractor can originate from another dataset of real or virtual distractors, in which case the determining S20 can include the step of providing such a dataset. It should be understood that in any case, the distractors have a certain level of photo - realism such that they can be placed S230 in the scenes of the third domain and after the placement S230, these scenes still belong to the third domain.

[0161] A distractor can be a mechanical part, for example, extracted and transformed from a virtual manufacturing scene as described above, or any other object (e.g., which is part of the background of a virtual (e.g., manufacturing) scene). Figure 21Shows three jammers 210, 212, and 214 that will be placed in one or more scenarios in the third domain in the context of a manufacturing scenario. The jammers 210, 212, and 214 are a safety helmet (which is a background object of the manufacturing scenario), a metal plate (which is a mechanical part), and an iron horseshoe forging (which is a mechanical part), respectively.

[0162] Placing the jammer S230, especially a jammer obtained by being similar to the transformed extraction object (as described above), can reduce the sensitivity of the neural network to numerical (e.g., digital, e.g., imaging) artifacts. In fact, as a result of the extraction S200 and / or the transformation S210 and / or the placement S220, it is possible that the scenario in the third domain including at least one transformed extraction object is characterized by one or more numerical artifacts at the position where there is at least one transformed extraction object in the placement S220. Placing the jammer in the scenario S230 may cause the same type of artifacts to appear at the position where the jammer is placed. Therefore, when being learned S30, the domain adaptation neural network will not focus on the artifacts, or at least will not link the artifacts to objects that can be reconstructed spatially, precisely because the artifacts are also caused by the jammer. This makes the learning S30 more robust.

[0163] Return to reference Figure 1 of the flowchart, now discuss the learning S30 of the domain adaptation neural network.

[0164] First, briefly discuss the general concepts of "machine learning" and "learning neural network".

[0165] It is known from the field of machine learning that the processing of an input by a neural network includes applying an operation to the input, which is defined by data including weight values. Therefore, the learning of a neural network includes determining the values of the weights based on a data set configured for such learning, and such a data set can be referred to as a learning data set or a training data set. For this purpose, the data set includes data segments, and each data segment forms a respective training sample. The training samples represent the diversity of the situations in which the neural network will be used after learning. Any data set involved in this article can contain more than 1000, 10000, 100000, or 1000000 training samples. In the context of the present disclosure, "learning a neural network based on a data set" means that the data set is the learning / training set of the neural network. "Learning a neural network based on a domain" means that the training set of the neural network is a data set belonging to and / or forming a domain.

[0166] Any neural network of the present disclosure, such as a domain adaptation neural network, or such as a detector, a teacher extractor, or a student extractor, which will be discussed later, may be a deep neural network (DNN) learned by any deep learning technique. Deep learning techniques are a powerful set of learning techniques in neural networks (see

[19] ), which is a biologically inspired programming paradigm that enables a computer to learn from observed data. In image recognition, the success of DNNs is attributed to their ability to learn rich intermediate media representations, rather than the hand-designed low-level features (Zernike moments, HOG, Bag-of-Word, SIFT) used in other image classification methods (SVM, Boosting, Random Forest). More specifically, DNNs focus on end-to-end learning based on raw data. In other words, by completing end-to-end optimization starting from raw features and ending with labels, they can stay as far away from feature engineering as possible.

[0167] Now discuss the concept of "configured to infer spatially reconstructable objects in a scene". The discussion applies in particular to a domain adaptation neural network according to any one of a domain adaptation learning method, a domain adaptation inference method, or a process, regardless of the way of learning the domain adaptation neural network. Inferring spatially reconstructable objects in a scene means inferring one or several spatially reconstructable objects in a scene that contains one or more such objects. This may mean inferring all spatially reconstructable objects in the scene. In this way, a neural network configured to infer spatially reconstructable objects in a scene takes as input the scene and output data related to the positions of one or more spatially reconstructable objects. For example, the neural network may output a (e.g., digital) image representation of the scene, where corresponding bounding boxes each enclose a corresponding one of the inferred spatially reconstructable objects. The inference may also include the calculation of such bounding boxes and / or the marking of the bounding boxes (e.g., enclosing the spatially reconstructable objects in the scene) and / or the marking of the inferred spatially reconstructable objects (as spatially reconstructable objects). The output of the neural network may be stored in a database and / or displayed on the screen of a computer.

[0168] The learning S30 of the domain adaptation neural network may be performed by any machine learning algorithm capable of learning a neural network configured to infer spatially reconstructable objects in a real scene.

[0169] The learning S30 of the domain - adaptive neural network is based on a third domain. Thus, the third domain is or at least contains the training / learning set of the domain - adaptive neural network. It is clear from the previous discussion of the determination S20 regarding the third domain that the training / learning set of the domain - adaptive neural network is derived from (e.g., obtained from) the virtual scene dataset provided in S10. In other words, the determination S20 can be considered as a pre - processing of the theoretical training / learning set, which is a dataset of virtual scenes, in order to transform it into an actual training / learning set, i.e., the scenes of the third domain, which are closer to the set of inputs that the neural network (once learned) will receive. As previously discussed, this improves the robustness of the learning S30 and the output quality of the domain - adaptive neural network.

[0170] The domain - adaptive neural network can include (e.g., can be constituted by) a detector, also called an object detector, which forms at least a part of the domain - adaptive neural network. Thus, the detector itself is a neural network, which is also configured to detect spatially reconstructable objects in real scenes. The detector typically can include millions of parameters and / or weights. The learning S30 of the domain - adaptive method can include training the detector, which sets the values of these parameters and / or weights. Such training typically can include updating these values, which includes continuously correcting these values for each input obtained by the detector, based on the output of the detector. During training, the input to the detector is a scene, and the output includes (at least after appropriate training) data related to the positions of one or more (e.g., all) spatially reconstructable objects in the scene (e.g., bounding boxes and / or annotations, as previously described).

[0171] The correction of the values can be based on the annotations associated with each input. An annotation is a set of data associated with a specific input, which allows evaluating whether the output of the model is true or false. For example, and as previously discussed, each spatially reconstructable object in a scene belonging to the third domain can be surrounded by a bounding box. As previously mentioned, in the context of the present disclosure, surrounding an object in a scene may mean surrounding all parts of the object that are fully visible in the scene. The bounding box can be annotated to indicate that it surrounds a spatially reconstructable object. The correction of the values can thus include evaluating such an annotation of the bounding box of the third - domain scene input to the detector, and based on the evaluated annotation, determining whether the spatially reconstructable object surrounded in the bounding box is indeed output by the detector. Determining whether the output of the detector includes a spatially reconstructable object surrounded in the bounding box can include evaluating the correspondence between the output of the detector and the spatially reconstructable object, e.g., by evaluating whether the output is true (e.g., whether it includes a spatially reconstructable object) or false (e.g., whether it includes a spatially reconstructable object).

[0172] The above-described manner of training the detector by using the annotations of the virtual dataset can be referred to as "supervised learning". The training of the detector can be carried out by any supervised learning method. After the detector is trained, the correction of the values stops. At this point, the detector is capable of processing new inputs (i.e., inputs not seen during the training of the detector) and returning detection results. In the context of the present disclosure, the new input to the detector is the scene of the test domain. After training, the detector will return two different outputs because the "detection" task means jointly performing the recognition (or classification) task and the localization task of spatially reconstructable objects in the real scene of the test dataset.

[0173] The localization task includes calculating a bounding box, each bounding box enclosing a spatially reconstructable object of the real scene input to the detector. As mentioned above, the bounding box is a rectangular box whose axes are parallel to the image sides and is characterized by four coordinates. In the example, for each spatially reconstructable object of the scene input to the detector, the detector returns a bounding box centered on the object at an appropriate ratio and scale.

[0174] The classification task includes labeling each calculated bounding box with a corresponding label (or annotation) and associating a confidence score with the label. The confidence score reflects the confidence of the detector that the bounding box indeed encloses a spatially reconstructable object of the scene provided as the input. The confidence score can be a real number between 0 and 1. In this case, the closer the confidence score is to 1, the greater the confidence of the detector that the label associated with the corresponding bounding box truly annotates the spatially reconstructable object.

[0175] Now refer to Figure 3 Discuss an example of the learning S30 of the domain adaptation neural network, Figure 3 A flowchart showing an example of the learning S30 of the domain adaptation neural network is shown.

[0176] Refer to Figure 3 According to the flowchart, the learning S30 of the domain adaptation neural network may include providing an S300 teacher extractor. The teacher extractor is a neural network of machine learning configured to output an image representation of the real scene. In this case, the learning S30 of the domain adaptation neural network further includes training an S310 student extractor. The student extractor is a neural network configured to output an image representation of the scene belonging to the third domain. The training S310 of the student extractor includes minimizing the loss. For each scene of one or more real scenes, the loss penalizes the difference between the result of applying the teacher extractor to the scene and the result of applying the student extractor to the scene.

[0177] The student extractor is a neural network configured to output an image representation of a scene belonging to a third domain. The student extractor can be part of a detector, and thus, the student extractor can be further configured to output the image representation of the scene belonging to the third domain to other parts of the detector. As previously discussed, at least one of such other parts can be configured for a localization task, and at least one of such other parts can be configured for a classification task. Thus, the student extractor can output the image representation of the scene of the third domain to the other parts such that the other parts can perform the tasks of localization and classification. Alternatively, the student extractor can be part of a domain adaptation neural network, and in this case can output the image representation to other parts of the domain adaptation neural network (e.g., the detector).

[0178] This amounts to saying that the student extractor supervises the output image representation of a scene belonging to the third domain, while at least one other part of the domain adaptation neural network (e.g., the detector discussed previously) supervises the detection (e.g., localization and classification) of spatially reconstructable objects in each scene of the third domain. The detection can be based on the image representation output by the student extractor. In other words, the image representation output by the student extractor is fed to the remaining parts of the domain adaptation neural network, e.g., the remaining parts of the detector if the student extractor forms part of the detector.

[0179] The student extractor is trained with the help of the teacher extractor provided at S300. The teacher extractor is a neural network of machine learning configured to output an image representation of a real scene. In an example, this means that the teacher extractor has been learned on real scenes to output an image representation of real scenes, e.g., by learning robust convolutional filters for the real data style. The neuron layers of the teacher extractor typically encode the representation of the real image to be output. The teacher extractor can be learned on any dataset of real scenes, e.g., obtainable from open source by any machine learning technique. Providing the teacher extractor at S300 can include accessing a database in which the teacher extractor has been stored after learning and retrieving the teacher extractor from the database. It should be understood that the teacher extractor may already be available, i.e., the domain adaptation learning method may not include learning the teacher extractor only by providing the teacher extractor that is already available at S300. Alternatively, the domain adaptation learning method can include: learning the teacher extractor by any machine learning technique before providing the teacher extractor at S300.

[0180] Still referring to Figure 2 the flowchart of, the training S320 of the student extractor is now discussed.

[0181] The training S320 includes minimizing a loss that penalizes the difference between the result of applying the teacher extractor to a scene and the result of applying the student extractor to the scene for each of one or more real-world scenarios. At least a portion (e.g., all) of the one or more real-world scenarios can be scenarios of a test dataset, and / or at least a portion (e.g., all) of the one or more real-world scenarios can be scenarios of another dataset of real-world scenarios belonging to a second domain, e.g., which are unannotated and not necessarily numerous. Minimizing the loss involves applying both the teacher extractor and the student extractor to the results of each respective real-world scenario in the one or more real-world scenarios, which means feeding the real-world scenarios into both the teacher extractor and the student extractor during the latter's training S310. Thus, while at least a portion of the domain-adaptive neural network is trained on virtual data (e.g., the previously discussed detector) belonging to a third domain, at least a portion of the domain-adaptive neural network, namely the student extractor, is trained based on real data.

[0182] The loss can be a quantity (e.g., a function) that measures the similarity and / or dissimilarity between the result of applying the teacher extractor to a scene and the result of applying the student extractor to the scene for each real-world scenario in the one or more real-world scenarios. For example, the loss can be a function that takes as an argument each scene of the one or more real-world scenarios, and takes as inputs the result of applying the teacher extractor to the scene and the result of applying the student extractor to the scene, and outputs a quantity (e.g., a positive real number) representing the similarity and / or dissimilarity between the result of applying the teacher extractor to the scene and the result of applying the student extractor to the scene. For each scene of the one or more real-world scenarios, the difference between the result of applying the teacher extractor to the scene and the result of applying the student extractor to the scene can be a quantification of the dissimilarity between the result of applying the teacher extractor to the scene and the result of applying the student extractor to the scene. Penalizing the difference can mean that the loss is an increasing function of the difference.

[0183] Minimizing the loss by penalizing the difference between the outputs of the teacher extractor and the student extractor improves the robustness of the learning S30 of the domain - adaptive neural network. In particular, the teacher extractor that has been learned and thus whose weights and / or parameters are not modified during the training S320 of the student extractor guides the student extractor. Indeed, since the difference between the outputs of the student extractor and the teacher extractor is penalized, the parameters and / or weights of the student extractor are modified so that the student extractor learns to output a real - image representation that is close to the image representation output by the teacher extractor in terms of data distribution. In other words, the student extractor is trained to mimic the teacher extractor in a way that outputs a real - image representation. Thus, the student extractor can learn the same or substantially the same robust convolutional filters as the teacher extractor during training for real - data styles. As a result, even if these scenarios may be relatively far from the second domain (constituted by real scenarios) in terms of data distribution, the student extractor outputs a real - image representation of the scenarios initially belonging to the third domain. This allows the domain - adaptive neural network, or in appropriate cases, its detector, to be invariant to data appearance, because the training S320 of the student extractor enables the domain - adaptive neural network to infer spatially reconstructable objects from real scenarios while being trained on annotated virtual scenarios of the third domain.

[0184] This loss can be referred to as the "distillation loss", and the training S320 of the student extractor using the teacher extractor can be referred to as the "distillation method". Minimization of the distillation loss can be carried out by any minimization algorithm or any relaxation algorithm. The distillation loss can be part of the loss of the domain - adaptive neural network, optionally also including the loss of the detector discussed previously. In one embodiment, the learning S30 of the domain - adaptive neural network can include minimizing the loss of the domain - adaptive neural network, and its type can be:

[0185] L training =L detector +L distillation

[0186] where L training is the loss corresponding to the learning S30 of the domain - adaptive neural network, L detector is the loss corresponding to the training of the detector, and L distillation is the distillation loss. Minimizing the loss L training of the domain - adaptive neural network can thus include minimizing the loss of the detector L detector and the distillation loss L distillation simultaneously or independently by any relaxation algorithm.

[0187] In the example, the result of applying the teacher extractor to the scene is the first Gram matrix, and the result of applying the student extractor to the scene is the second Gram matrix. The Gram matrix of an extractor (e.g., the student or the teacher) is a form that represents the data distribution on the image output by the extractor. More specifically, applying the extractor to the scene produces an image representation, and calculating the Gram matrix of the image representation produces an embedding of the distribution of the image representation, namely the so-called Gram-based representation of the image. This is equivalent to saying that in the example, the result of applying the teacher extractor to the scene is the first Gram-based representation of the image representation output by the teacher extractor, which is obtained by calculating the first Gram matrix of the image representation output by the teacher extractor, and the result of applying the student extractor to the scene is the second Gram-based representation of the image representation output by the student extractor, which is obtained by calculating the second Gram matrix of the image representation output by the student extractor. For example, the Gram matrix can contain non-localized information about the image output by the extractor, such as texture, shape, and weights.

[0188] In the example, the first Gram matrix is calculated on several neuron layers of the teacher extractor. The neuron layers include at least the last neuron layer. In these examples, the second Gram matrix is calculated on several neuron layers of the student extractor. The neuron layers include at least the last neuron layer.

[0189] The last neuron layer of the extractor (teacher or student) encodes the image output by the extractor. The first and second Gram matrices are both calculated on several layers of neurons, which are deeper layers of neurons. This means that they both contain more information about the output image than the image itself. This makes the training of the student extractor more robust because it increases the true style proximity between the outputs of the teacher extractor and the student extractor, which is the purpose of minimizing the loss. It is worth noting that the student extractor can thus infer the two-layer representation of the real image as accurately as its teacher.

[0190] In the example, the difference is the Euclidean distance between the result of applying the teacher extractor to the scene and the result of applying the student extractor to the scene.

[0191] Now discuss the implementation for minimizing the distillation loss. In this implementation, the style representation of the image output by the extractor (teacher or student) at certain layers of the neurons is encoded by the feature distribution of the image in that layer. In this implementation, the Gram matrix is used to encode the feature correlation, which is a form of feature distribution embedding. Formally, for a layer with a filter response of length the feature correlation is given by the Gram matrix ​Given, where is the layer inner product between the vectorized feature maps i and j:

[0192]

[0193] The intuition behind the proposed distillation method is to ensure that during the learning process, the weights of the student extractor are updated so that the data style representation of its output is close to the form generated by the teacher extractor based on the real data. To this end, learning S30 includes minimizing the Gram-based distillation loss L distillation . Given a real image, let and be the corresponding feature map outputs at a certain layer of the student extractor and the teacher extractor respectively. Therefore, the distillation loss for the proposed layer is of the following type:

[0194]

[0195] where is the corresponding Gram matrix and any distance (e.g., Euclidean distance) between. Considering distilling knowledge from layer , then the type of distillation loss is:

[0196]

[0197] As mentioned before, the learning S30 of the domain adaptation neural network can then include minimizing the following type of loss:

[0198] L training = L detector + L distillation .

[0199] Return the flowchart of reference Figure 1 . Now discuss S50 for determining the intermediate domain.

[0200] Determination of the intermediate domain S50 makes the intermediate domain closer to the training domain than the test domain in terms of data distribution. Thus, the determination of the intermediate domain S50 specifies the inference of another scenario for each scenario of the test data set of the scenario, such that all the inferred other scenarios form a domain that is closer to the training domain than the test domain in terms of data distribution, i.e., the intermediate domain. The inference of the other scenario may include the transformation (e.g., conversion) of the scenario to the other scenario. The transformation may include, for example, using a generator capable of transferring an image from one domain to another domain to transform the scenario into the other scenario. It should be understood that such a transformation may only involve one or more parts of the scenario: the one or more parts are transformed, for example, using a generator, while the rest of the scenario remains unchanged. The other scenario may include a scenario in which one or more parts are transformed while the rest of the scenario remains unchanged. Alternatively, the transformation may include replacing the rest of the scenario with a part of some other scenario (e.g., belonging to the training domain). In such a case, the other scenario may include a scenario in which one or more parts are transformed while the rest of the scenario is replaced by the part of some other scenario. It should be understood that the inference can be similarly performed for all scenarios of the test domain. For example, having the same paradigm applicable to all scenarios.

[0201] The domain - adaptive neural network is not learned on the training domain, but on the data obtained from the training domain. Such data can be the result of processing the data of the training domain. For example, the data obtained from the training domain can be a data set of scenarios, each scenario being obtained by transforming the corresponding scenario of the training domain. This is equivalent to saying that the data obtained from the training domain can form a data set of scenarios belonging to a domain that is the true training / learning domain of the domain - adaptive neural network. Such a domain can be precisely determined as the third domain by performing similar operations and steps. In any case, the determination of the intermediate domain S50 can be performed in such a way (e.g., by the transformation of the scenario discussed above) that the intermediate domain is closer to the domain formed by the data obtained from the training domain than the test domain. In an example, the data obtained from the training domain forms another intermediate domain, which is closer to the intermediate domain than the training domain in terms of data distribution. The other intermediate domain may be the third domain discussed previously. In all these examples, the determination of the intermediate domain S50 is equivalent to the pre - processing of the test data set so as to make it closer to the true training / learning set on which the domain - adaptive neural network has been learned. As mentioned before, this improves the output quality of the domain - adaptive neural network.

[0202] In an example, for each scenario of the test data set, the determination of the intermediate domain S50 includes transforming the scenario of the test data set into another scenario that is closer to the training domain in terms of data distribution.

[0203] Transforming a scene into another scene may include using a generator capable of transferring an image from one domain to another. It should be understood that such a transformation may only involve one or more parts of the scene: the one or more parts are transformed, e.g., using a generator, while the rest of the scene remains unchanged. The other scene may include a scene in which one or more parts are transformed while the rest of the scene remains unchanged. Alternatively, the transformation may include replacing the rest of the scene with a part of some other scene (e.g., belonging to the training domain). In such a case, the other scene may include a scene in which one or more parts are transformed while the rest of the scene is replaced by the part of some other scene. It should be understood that inferences can be made similarly for all scenes in the test domain, e.g., with the same paradigm for all scenes.

[0204] Now referring to Figure 4 an example of S50 for determining an intermediate domain is discussed. Figure 4 A flowchart illustrating an example of the determination S50 of an intermediate domain is shown.

[0205] Referring to Figure 4 the flowchart of, the transformation of a scene may include generating S500 virtual scenes from the scenes of a test dataset. In this case, another scene will be inferred based on the scenes of the test dataset and the generated virtual scenes.

[0206] Generating S500 can be performed by applying any generator capable of generating virtual scenes from the scenes of the test domain. Another scene can be inferred by mixing the scenes of the test dataset and the generated virtual scenes and / or by combining (e.g., mixing, seamless cloning, or merging (e.g., by a method of merging image channels)) the scenes of the test dataset and the generated virtual scenes.

[0207] Still referring to Figure 4 the flowchart of, generating S500 virtual scenes from the scenes of a test dataset may include applying a virtual scene generator to the scenes of the test dataset. In this case, the virtual scene generator is a neural network of machine learning, which is configured to infer virtual scenes from the scenes of the test domain.

[0208] The virtual scene generator takes the scenes of the test domain as input and outputs virtual scenes, e.g., virtual scenes with lower photo-reality than the scenes of the test domain. It should be understood that the virtual generator may have been trained / learned. Thus, generating S500 of virtual scenes may include providing the virtual scene generator, e.g., by accessing a memory storing the virtual scene generator after training / learning the virtual scene generator and retrieving the virtual scene generator from the memory. Additionally or alternatively, before providing the virtual scene generator, the virtual scene generator can be learned / trained by any machine learning technique suitable for such training.

[0209] In an example, a virtual scene generator has been learned on a dataset of scenes, where each scene includes one or more spatially reconstructable objects. In such an example, when a scene including one or more spatially reconstructable objects is fed to the virtual scene generator, the virtual scene generator is capable of outputting a virtualization of the scene, which is a virtual scene that also includes one or more spatially reconstructable objects, e.g., all located at the same location and / or having the same configuration. As previously mentioned, it should be understood that the domain adaptation inference method and / or process may only include providing the virtual scene generator (i.e., it has not been learned), or may also include learning of the virtual scene generator.

[0210] In an implementation of the domain adaptation inference method and / or process, where the test dataset consists of real manufacturing scenes containing manufacturing robots, the generator is a CycleGAN real-to-virtual network trained on two manufacturing scene datasets (a virtual dataset and a real dataset) containing manufacturing robots. In this implementation, the method and / or process includes training the CycleGAN network on the two datasets to adapt to the appearance of the robots in the test images of the test scene, thus being closer to the training set of the domain adaptation neural network, which in this implementation is the third domain according to the domain adaptation learning method and consists of virtual manufacturing scenes containing virtual manufacturing robots. The training produces a real-to-virtual robot generator that is capable of modifying the real test images to virtualize them.

[0211] Figure 22 [[ID=⑧]]An example of generating S500 using the CycleGAN generator is shown. In this example, the test dataset consists of real manufacturing scenes, each containing one or more manufacturing robots. Figure 22 The first real scene 220 and the second real scene 222 of the test dataset are shown. Both of these scenes are manufacturing scenes including articulated robots. Figure 22 The first virtual scene 224 and the second virtual scene 226 generated from the first scene 220 and the second scene 222 respectively by applying the CycleGAN generator are also shown.

[0212] Still referring to Figure 4 the flowchart of, determining S50 of the intermediate domain may include mixing S510 the scenes of the test dataset and the generated virtual scenes. In this case, the mixing results in another scene.

[0213] Formally, let I be the scenario of the test dataset, let I′ be the generated virtual scenario, and let I| be another scenario. The mixing S510 of I and I′ can refer to any method capable of mixing I and I′ so that the mixed result is I”. In an example, the mixing S510 of the scenario of the test dataset and the generated virtual scenario is a linear mixing. In one implementation, the linear mixing includes calculating another scenario I” through the following formula:

[0214] I” = α * I + (1 - α) * I′.

[0215] Figure 23 An example of the linear mixing is shown. Figure 23 A real manufacturing scenario 230 including a real manufacturing robot and a virtual scenario 232 generated from the real scenario 230 is shown. Figure 23 The result of the linear mixing, which is a virtual manufacturing scenario 234, is also shown.

[0216] Return to reference Figure 1 The flowchart of, now discuss inferring an object that can be reconstructed in the S60 space from the scenario of the test dataset of the test domain transferred on the intermediate domain.

[0217] First, inferring an object that can be reconstructed in the S60 space from a scenario can include: inferring an object that can be reconstructed in the space from one scenario, or inferring for each scenario of the input dataset of the scenario, an object that can be reconstructed in the space from each scenario of the input dataset of the scenario. It is noted that inferring S60 can constitute the test phase of a domain adaptive neural network, where (e.g., successively) all scenarios of the test dataset are provided as inputs to the domain adaptive neural network, and the domain adaptive neural network infers one or more objects that can be reconstructed in the space in each input scenario. Additionally or alternatively, inferring S60 can constitute the application phase of a domain adaptive learning method and / or a domain adaptive inference method and / or process, where the input dataset is a dataset of test domain scenarios, which is not equal to the test dataset. In this case, the application phase can include providing the input dataset, inputting each scenario of the input dataset into the domain adaptive neural network, and performing inferring S60 for each scenario. It should also be understood that inferring an object that can be reconstructed in the space in a scenario can include inferring (e.g., simultaneously or successively) one or more objects that can be reconstructed in the space in a scenario including one or more such objects, e.g., by inferring all objects that can be reconstructed in the space in the scenario.

[0218] Now discuss the inference S60 of an object that can be spatially reconstructed in a scenario. This inference can be part of the inference S60 of more objects that can be spatially reconstructed in the scenario. Inferring an object that can be spatially reconstructed in the S60 scenario by applying a domain - adaptive neural network includes inputting the scenario into the domain - adaptive neural network. Then, the domain - adaptive neural network outputs data related to the positions of the objects that can be spatially reconstructed. For example, the neural network can output a (e.g., digital) image representation of the scenario, where corresponding bounding boxes enclose the inferred objects that can be spatially reconstructed. The inference can also include the calculation of such bounding boxes and / or the marking of the bounding boxes (e.g., enclosing the objects that can be spatially reconstructed in the scenario) and / or the marking of the inferred objects that can be spatially reconstructed (as objects that can be spatially reconstructed). The output of the neural network can be stored in a database and / or displayed on the screen of a computer.

[0219] Figures 24 to 29 Shows the inference of a manufacturing robot in a real - world manufacturing scenario. Figures 24 to 29 Both show the corresponding real - world scenarios of the test domain fed as inputs to the domain - adaptive neural network, and the bounding boxes calculated and annotated around all manufacturing robots included in each corresponding scenario. As Figures 26 to 30 shown, multiple robots and / or at least partially occluded robots can be inferred in the real - world scenario.

[0220] In fact, the inference of an object that can be spatially reconstructed can be performed by a detector that forms at least part of the domain - adaptive neural network, which jointly performs the task of identifying (or classifying) and localizing objects that can be spatially reconstructed in the real - world scenarios of the test dataset. The localization task includes calculating bounding boxes, each of which encloses an object that can be spatially reconstructed in the scenario input to the detector. As mentioned before, a bounding box is a rectangular box whose axes are parallel to the image sides and is characterized by four coordinates. In the example, for each object that can be spatially reconstructed in the scenario input to the detector, the detector returns a bounding box centered on the object at an appropriate scale and size. The classification task includes marking each calculated bounding box with a corresponding label (or annotation) and associating a confidence score with the label. The confidence score reflects the confidence of the detector that the bounding box indeed encloses an object that can be spatially reconstructed in the scenario provided as input. The confidence score can be a real number between 0 and 1. In this case, the closer the confidence score is to 1, the greater the confidence of the detector that the label associated with the corresponding bounding box truly annotates the object that can be spatially reconstructed.

[0221] In an example where a domain - adaptive neural network is configured to infer objects in a real - world scenario, for instance when the test domain consists of real - world scenarios and the intermediate domain consists of virtual scenarios, the domain - adaptive neural network may also include an extractor. The extractor forms part of the domain - adaptive neural network, which is configured to output image representations of real - world scenarios and feed them to other parts of the domain - adaptive neural network, such as a detector or a part of the detector, e.g., to supervise the detection of spatially reconstructable objects. In the example, this means that the extractor can process the scenarios of the intermediate domain to output image representations of the scenarios, for example, by learning robust convolutional filters for the real - data style. The neuron layers of the extractor typically encode the representation of a real or photo - realistic image to be output. This is equivalent to saying that the extractor supervises the output image representation of the scenarios belonging to the intermediate domain, while at least one other part of the domain - adaptive neural network (e.g., the previously discussed detector) supervises the detection (e.g., localization and classification) of spatially reconstructable objects in each input scenario of the third domain. The detection can be based on the image representations output by the student extractor. This allows for the efficient processing and output of real data.

[0222] Each scenario input to the domain - adaptive neural network belongs to the test domain but is transferred on a determined intermediate domain. In other words, providing a scenario as input to the domain - adaptive neural network includes: transferring the scenario on the intermediate domain and feeding the transferred scenario as input to the domain - adaptive neural network. The transfer of the scenario can be performed by any method capable of transferring a scenario from one domain to another. In the example, transferring the scenario can include transforming the scenario into another scenario and optionally mixing the scenario with the transformed scenario. In this case, the transformation and optional mixing can be performed as in the example described in the flowchart previously referenced Figure 4 and optionally mixing. It is worth noting that the scenarios of the intermediate domain can actually be the scenarios of a test data set transferred on the intermediate domain.

[0223] Now, the implementation of this process is discussed.

[0224] This implementation proposes a learning framework in a virtual world to address the problem of insufficient manufacturing data. In particular, this implementation utilizes the efficiency of a virtual simulation data set for learning and an image - processing method to learn a domain - adaptive neural network, i.e., a deep neural network (DNN).

[0225] Virtual simulation constitutes a form of obtaining virtual data for training. The latter seems to have become a promising technology to solve the problem of the difficulty in obtaining training data required for machine - learning tasks. Virtual data is generated by an algorithm to mimic the characteristics of real data, while having automatic, error - free, and cost - free labeling.

[0226] Virtual simulation is proposed through this implementation manner to better explore the working mechanisms of different manufacturing devices. Formally, different manufacturing tools are considered through the different states that these tools will present during the running time in the learning process.

[0227] Deep neural network: A powerful set of learning techniques in neural networks

[19] , which is a biologically inspired programming paradigm that enables a computer to learn from observed data. In image recognition, the success of DNN is attributed to its ability to learn rich intermediate media representations, rather than the hand-designed low-level features (such as Zernike moments, HOG, bag of words, SIFT, etc.) used in other image classification methods (SVM, boosting, random forests, etc.). More specifically, DNN focuses on end-to-end learning based on raw data. In other words, by completing end-to-end optimization starting from raw features and ending with labels, they can stay as far away from feature engineering as possible.

[0228] Domain adaptation: A field related to machine learning and transfer learning. This scenario occurs when aiming to learn a model that performs well on a different (but related) target data distribution from the source data distribution. The purpose of domain adaptation is to find an effective mechanism to transfer or adapt knowledge from one domain to another. This is achieved by introducing a cross-domain bridging component to bridge the cross-domain gap between the training data and the test data.

[0229] The proposed implementation manner is particularly shown by Figure 31 which shows a diagram illustrating the offline and online phases of the process.

[0230] 1. Offline phase: This phase aims to train a model using virtual data, and the model should have the ability when inferring the real world. It contains two main steps. Note that this phase is transparent to the user.

[0231] 1) Conduct virtual simulation in a virtual manufacturing environment to generate representative virtual training data and automatically annotate it. The virtual manufacturing environment is geometrically realistic rather than photo-realistic. The type of annotation depends on the target inference task.

[0232] 2) Learn a neural network model based on the virtual data. It includes a model based on domain-adaptive DNN.

[0233] 2. Online phase: Given a real-world medium, the learned domain-adaptive model is applied to infer manufacturing equipment.

[0234] This implementation of the process particularly relates to the field of fully supervised object detection. Specifically, in this particular implementation, we are interested in a manufacturing environment, and the goal is to detect industrial robot arms in factory images. For this fully supervised field, the training dataset should contain up to hundreds of thousands of annotated data, which justifies the use of virtual data for training.

[0235] The latest object detectors are based on deep learning models. The values of millions of parameters cannot be set manually, and these are all features of the model. Therefore, learning algorithms must be used to set these parameters. When the learning algorithm updates the model parameters, the model is said to be in the "training mode". Due to the annotations associated with each input, it includes continuously "correcting" the model according to the output for each input. An annotation is a set of data associated with a specific input, which allows evaluating whether the output of the model is true or false. For example, training an object classifier to distinguish images of cats and dogs requires a dataset of annotated images of cats and dogs, and each annotation is "cat" or "dog". Therefore, if the object classifier outputs "dog" for an input cat image in its training mode, the learning algorithm will correct the model by updating its parameters. This method of supervising model training through an annotated dataset is called "supervised learning". After training the model, we will stop updating its parameters. Then, the model is only used to process new inputs (i.e., inputs not seen during the training mode) and return detection results, which is called the model being in the "testing mode".

[0236] In the context of this implementation, the object detector returns two different outputs because the "detection" task means jointly performing the recognition (or classification) task and the localization task.

[0237] 1. Localization output: Object localization can be performed because of the bounding box. The bounding box is a rectangular box whose axes are parallel to the image sides. It is characterized by four coordinates. Ideally, the object detector returns a bounding box centered on the object at an appropriate scale and proportion for each object.

[0238] 2. Classification output: Object classification is performed through class labels associated with the confidence scores of each bounding box. The confidence score is a real number between 0 and 1. The closer the value is to 1, the higher the confidence of the object detector in the class label associated with the corresponding bounding box.

[0239] This implementation is particularly shown by Figure 32 which shows a domain adaptation component used in both the offline and online phases.

[0240] It involves learning an object detection model using generated virtual data. Within both the offline and online phases, several domain adaptation methods are applied to address the gap between the virtual and real domains.

[0241] 1. Offline Phase: This phase aims to train a model using virtual data, which should be capable when inferring the real world. This phase consists of two main steps.

[0242] 1) Conduct virtual simulation in a virtual manufacturing environment to generate representative virtual training data, which are automatically annotated with bounding boxes and "robot" labels.

[0243] 2) Domain Adaptation Transformation

[0244] i. At the training set level: Transform the training set to a new training domain that is closer to the real world in terms of data distribution.

[0245] ii. At the model architecture level: Insert a domain adaptation module into the model to address the problematic virtual - to - real domain shift.

[0246] 2. Online Phase: Convert media in the real world to a new test domain, which is then better processed by the learned detection model and then inferred by the learned detection model. In other words, the new test domain is closer to the training domain than the original real domain.

[0247] It is important to note that a particularity of this implementation is that it does not encounter the problem of virtual - to - real image transformation accuracy. The intuition of the current work is to consider an intermediate domain between the virtual world and the real world rather than trying to migrate virtual data to the real world. Formally, the virtual training data and the real test data are respectively dragged to the training and test intermediate domains with the least domain shift.

[0248] According to the above framework, it can be understood that the current implementation under discussion is based on two main steps:

[0249] · Generation of virtual data for learning.

[0250] · Domain adaptation.

[0251] To address the scarcity of factory images, this implementation relies on a simulation - based virtual data generation method. The training set includes 3D - rendered factory scenes. These factory scenes are generated by the company's in - house design application and come with masks that segment objects of interest (robots) at the pixel level. Each scene is simulated into multiple frames to describe the various functional states that a robot may have.

[0252] This step generates a set of virtual images with unique and complex functions for detecting the tasks of robots in a manufacturing environment. Although this may seem like a "simple" single-object detection task, baseline detectors and state-of-the-art domain adaptation networks are not competent for this problem. Formally, the generated set is not photo-realistic but "cartoonified". The simplicity of the generated virtual scenes reflects the ease of generating our training data, which is the main advantage of this embodiment. At the same time, due to the complexity of the objects of interest, it is difficult to use the state-of-the-art methods for solving such virtual data, namely the domain randomization DR technique (see [7,15]).

[0253] In fact, domain randomization abandons photo-reality by randomly disturbing the environment in a non-photo-realistic way. The intuition behind this DR is to teach the algorithm to force the network to learn to focus on the basic features of the image and ignore the objects of no interest in the scene. More formally, the main real features missing in the virtual data are the low-level forms of the image, such as texture, lighting, and shading. DR ensures that these features are perturbed on the training dataset so that the algorithm does not rely on them to identify objects. Following the same logic, DR will not affect the common features between real data and virtual data. Within the scope of the virtual data generation used, one may learn that the correct features that the learned model expects to learn to rely on are shape and contour. This is important in the DR technique.

[0254] In addition, the object of interest in this embodiment is a robot, which has a configuration of many joints (usually 6 degrees of freedom), and these configurations are calibrated by a continuous value range. This gives an infinite range of possible states for each type of robot. Therefore, for such a complex industrial object, randomizing all unique features except the learned shape and contour is not the best choice. In fact, when the shape and contour of the object are complex, it is difficult to meet the variability required for them to represent the shape and contour of the object. Especially for the robot example, it is truly unimaginable to learn a model that can extract and focus on the unique shapes and contours of various types of robots with infinite joint positions to identify objects in an image.

[0255] The proposed embodiment establishes a learning problem that is competent when the data is complex (e.g., the current situation).

[0256] The proposed domain adaptation method can be considered a partial domain randomization technique. Formally, it proposes to enhance the correct features that the model is learning by deliberately randomizing or transforming the data features into a new domain that better represents the real world than the virtual environment. For example, it can better explore the specificity of texture features. To this end, three novel techniques are implemented in the embodiment currently under discussion:

[0257] 1. Virtual-to-real transformation

[0258] To correct the cross - domain offset, a common method in the prior art articles is to change the appearance of the virtual data so that it looks as if it were extracted from the real domain, i.e., the distribution they obtain can better represent the real world. For this purpose, the virtual images are globally modified using real scenarios without any correspondence between similar objects. However, to obtain the best image - to - image conversion performance, it might be guessed that transferring real robot textures to the robots in the virtual scene and using it for the background would be beneficial.

[0259] The current implementation is based on this targeted appearance transfer, which is achieved by applying different image - to - image conversion networks to the robots and factory backgrounds of each medium in the training set separately.

[0260] For this purpose, we use the CycleGAN model

[20] to convert images from one domain to another. A method is proposed that can learn to capture the special features of an image set and figure out how to transform these features into a second set of images without any paired training examples. In fact, this relaxation of the one - to - one mapping makes this representation very powerful. This is achieved through a generative model, specifically a generative adversarial network (GAN) called CycleGAN.

[0261] As described above, two different models are trained separately

[0262] · CycleGAN virtual - to - real network 1: It is trained on two factory background sets: the virtual set and the real set. It generates a virtual - to - real factory generator, which is used to modify the background in the virtual images.

[0263] · CycleGAN virtual - to - real 2: It is trained on two robot sets (virtual set and real set). This results in a virtual - to - real robot generator, which is used to modify the robots in the virtual images.

[0264] 2. Gram - based distillation

[0265] To better emphasize the model robustness regarding domain offset, we propose a novel knowledge distillation method implemented at the detector level. Note that most prior - art detectors (see [18, 9]) contain a feature extractor, which is a deep network of convolutional layers. The feature extractor outputs an image representation to feed the rest of the detector. The intuition of distillation is to guide the detection model to learn robust convolutional filters for the real - data style, which is consistent with the main goal of making the model invariant to the data appearance. <0>

[0266] Formally, this embodiment involves training a detector on virtual media while updating the weights of the detector's feature extractor to approximate the style encoded by the two-layer representation of real images fed into both our detector and a second detector pre-trained on real data from an open source.

[0267] Note that the style representation of an image at a certain layer is encoded by the feature distribution in that layer. For this embodiment, we use the Gram matrix that encodes feature correlations, which is a form of feature distribution embedding. Formally, for a layer with a filter response of length Feature correlations are given by the Gram matrix where is the inner product between the vectorized feature maps i and j in layer :

[0268]

[0269] The intuition behind the proposed distillation is to ensure that during the learning process, the weights of the detector's feature extractor are updated so that the style representation of the data it outputs is close to the style representation of the data that would be generated by another feature extractor pre-trained on real data. To this end, the currently discussed embodiment implements a Gram-based distillation loss L distillation .

[0270] Given a real image, we denote the corresponding feature map outputs of the detector's feature extractor and the pre-trained feature extractor for layer by and respectively. Thus, the proposed distillation loss at layer is the Euclidean distance between the corresponding Gram matrices and

[0271]

[0272] Considering distilling knowledge from layer the distillation loss is defined as

[0273]

[0274] The total training loss is

[0275] L training = L detector + L distillation

[0276] 3. Real-to-Virtual Conversion ​​

[0277] It may be expected that the task of converting virtual images into the real domain cannot be perfectly completed. This means that the modified data is closer to the real world than the virtual data, but does not exactly match the real distribution. Therefore, we assume that the modified virtual data is transferred to a new domain D1 located between the synthetic domain and the real domain. One contribution of this work is to extend the domain adaptation to real test data by dragging the real test data to another new domain D2. To use this method, in terms of distribution, D2 should be closer to D1 than the real domain; this is the hypothesis confirmed by comparing data representations in this application case. This form of adaptation is achieved by applying the real-to-virtual data conversion (i.e., the inverse operation previously applied to virtual data during the training phase) to the test data.

[0278] In the context of the present embodiment, the transformation is ensured by training a CycleGAN real-to-virtual network trained on two robot sets (virtual set and real set). This will produce a real-to-virtual robot generator that is used to modify real test images. At this level, we choose to train the CycleGAN network on the robot sets because we are mainly interested in adapting the appearance of the robot in the test images to be recognized by the detector. Note that more complex image transformations can also be performed on the test data by targeting regions individually, as was done for the training set. However, the improvement is limited, and we are satisfied with applying a CycleGAN network to the entire test image.

Claims

1. A computer-implemented method for domain adaptation inference, comprising: - Providing: - A test data set of real scenarios belonging to a test domain; And - A domain adaptation neural network, which is a neural network of machine learning that learns from data obtained from a training domain, the training domain includes a training data set of virtual scenarios, each virtual scenario includes one or more spatially reconstructable objects, and the domain adaptation neural network is configured to infer spatially reconstructable objects in the real scenarios of the test domain; - Determining an intermediate domain, in terms of data distribution, the intermediate domain is closer to the training domain than the test domain, wherein the determination of the intermediate domain includes, for each real scenario of the test data set, transforming the real scenario of the test data set into another scenario that is closer to the training domain in terms of data distribution by the following operations: - Generating a virtual scenario according to the real scenario of the test data set; and - Mixing the real scenario of the test data set with the generated virtual scenario, the mixing produces the other scenario, Wherein, all the transformed other scenarios for each real scenario of the test data set form the intermediate domain; and - By applying the domain adaptation neural network, inferring spatially reconstructable objects according to the real scenarios of the test domain transferred on the intermediate domain.

2. The method according to claim 1, wherein, Generating the virtual scenario according to the scenario of the test data set includes applying a virtual scenario generator to the scenario of the test data set, and the virtual scenario generator is a neural network of machine learning configured to infer a virtual scenario according to the scenario of the test domain.

3. The method according to claim 2, wherein The virtual scenario generator has been learned on a data set of scenarios, and each scenario includes one or more spatially reconstructable objects.

4. The method according to claim 1, wherein The mixing of the scenario of the test data set and the generated virtual scenario is a linear mixing.

5. The method according to any one of claims 1 - 4, wherein, Each real scenario of the test data set is a real manufacturing scenario including one or more spatially reconstructable manufacturing tools, and the domain adaptation neural network is configured to infer spatially reconstructable manufacturing tools in the real manufacturing scenario.

6. The method according to any one of claims 1 to 4, wherein Each virtual scenario of the data set of virtual scenarios is a virtual manufacturing scenario including one or more spatially reconstructable manufacturing tools, and the domain adaptation neural network is configured to infer spatially reconstructable manufacturing tools in the real manufacturing scenario.

7. The method according to any one of claims 1 to 4, wherein, The data obtained from the training domain includes scenarios of another intermediate domain, and the domain adaptation neural network has been learned on the another intermediate domain, and in terms of data distribution, the another intermediate domain is closer to the intermediate domain than the training domain.

8. A computer program product, comprising instructions for performing the method according to any one of claims 1 to 7.

9. An apparatus, comprising a data storage medium on which the computer program product according to claim 8 is recorded.

10. The apparatus according to claim 9, further comprising a processor coupled to the data storage medium.

Citation Information

Patent Citations

  • Adapting simulation data to real-world conditions encountered by physical processes

    US20180349527A1