Automatic method for 3D object recognition and localization
The method provides flexible and robust 3D object recognition and localization by using fixed sensor poses and simulated images for AI training, addressing the limitations of existing methods in dynamic environments.
Patent Information
- Application Number
- FR2024007887
- Authority / Receiving Office
- FR · FR
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-18
- Publication Date
- 2026-01-23
AI Technical Summary
Existing methods for object recognition and localization in real environments require complex sensor calibration, precise environmental modeling, and are ineffective if objects move, with tedious and expensive database creation processes.
A method using fixed sensor poses to provide probabilities of object occupancy in a spatial mesh, allowing 3D object recognition and localization without complete environmental modeling, utilizing simulated and real images for training AI models.
Enables flexible and robust object positioning in dynamic environments, improving reliability in inventory management, product tracking, and navigation systems.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Title of the invention: Automatic method for 3D object recognition and localization. Technical field
[0001] This disclosure falls within the domain of 3D object recognition and / or localization.
[0002] It relates more specifically to a method for recognizing and / or locating objects in 3D, a method for training an artificial intelligence model, as well as corresponding computing units, systems, computer programs and devices. Previous technique
[0003] Known methods for identifying and locating objects in a real environment often rely on computer vision techniques that require complex sensor calibration and precise modeling of the environment.
[0004] It is particularly known to use an artificial intelligence model to detect objects in a 2D image from an RGB camera and then to rely on a complete 3D model of the environment to link the objects detected on the image to their real position in the environment.
[0005] These known methods have significant limitations. They require knowing in advance the positions of the parts to be detected, which makes detection ineffective if the parts move.
[0006] Moreover, the creation of real and annotated databases for training the artificial intelligence model is a tedious and expensive task, requiring either a lot of time (about 3 minutes per image on average for some manual methods), or a complex and expensive specialized infrastructure.
[0007] There is a need for more flexible and robust solutions, capable of operating in dynamic environments without requiring prior full modeling. Summary
[0008] This disclosure improves the situation.
[0009] According to one aspect, a method implemented by at least a first computing unit is proposed, the method comprising: a supply, from at least two 2D images of a 3D environment including a 3D object, the 2D images being acquired by sensors whose relative poses are fixed, of at least one probability associated with at least one class of the 3D object representative of an occupation by the 3D object of at least one elementary volume of a spatial mesh of the 3D environment.
[0010] According to the embodiments, the 3D environment and the 3D object can be real or virtual.
[0011] The proposed technique thus helps to identify and retrieve the positions and / or orientations in space of objects whose shape and / or dimensions are known from streams of several cameras, even if the objects are moving in space, without requiring complete modeling of the environment. By providing at least one probability associated with the object's class and its occupancy within an elementary volume of the spatial mesh, the method helps to determine the precise position of the object in 3D space, based on the provided images. This approach differs from those that simply locate the object in a 2D image or that require prior modeling of the environment itself. Therefore, it allows for greater adaptability and accuracy in varied and dynamic environments.
[0012] Thus, the proposed method can contribute to improved reliability of object monitoring and management systems in various environments, real or virtual, for example industrial, logistical, or production environments. By helping to determine the object's location in 3D space, based solely on the input images, the proposed method can facilitate inventory management in a warehouse, and / or real-time product tracking on a production line, or even facilitate the navigation and manipulation of objects by robots in complex environments.
[0013] The features described in the following paragraphs may optionally be implemented independently of each other or in combination with each other.
[0014] In one example, the at least one probability also represents a probability of the position and / or orientation of the 3D object. This helps to refine the object's location in space and improve the relevance of automated actions based on this location.
[0015] In one example, the environment is a real environment and the method comprises generating a digital twin of the real environment, the digital twin taking into account at least one probability. The digital twin provides a real-time view of the monitored environment, thus facilitating automatic decision-making or decision-making by an operator or user.
[0016] In one example, the 2D images are derived from video streams acquired by the sensors, and the process is implemented for several sets of 2D images acquired in the same time window (simultaneously or almost simultaneously, for example) by the sensors. Such continuous image stream processing makes it possible to provide real-time updates on the positions and orientations of objects in the environment.
[0017] In one example, the method includes tracking the position and / or orientation of the 3D object, taking into account the probabilities obtained. This tracking, which can, for example, be performed in real time, can help improve the accuracy and reliability of automated systems in dynamic environments.
[0018] In one example, the spatial mesh is relative to a first position of at least one of the sensors. For example, the spatial mesh may originate from this first position.
[0019] In one example, the method includes determining the position of the 3D object in the 3D environment from at least one probability and the position of at least one of the sensors.
[0020] According to another aspect, a method implemented by at least one second computing unit is proposed, the method comprising: obtaining 2D views of at least one 3D object, the views being acquired or generated from viewpoints fixed relative to each other, and training an artificial intelligence model with 2D views to predict at least one class of the 3D object in elementary volumes of a spatial mesh of a 3D environment.
[0021] Model training can be based, according to embodiments, on real and / or simulated images (or a combination of real and simulated images). The possibility of using simulated views limits (or even eliminates) the often complex phase of acquiring and annotating images for model training. The generation of simulated 2D views of an object can be performed from at least one 3D model of the object, such as a CAD model or a 3D scan. This can eliminate the need for a complete physical or virtual environment to capture the views. This approach can therefore help simplify model training, as it only requires the 3D models of the objects to be detected and a suitable graphics setup.By limiting or avoiding the need for costly infrastructure and annotation time, the proposed method offers a more flexible and economical solution than some known methods.
[0022] The features described in the following paragraphs may optionally be implemented independently of each other or in combination with each other.
[0023] In one example, the method includes an alteration treatment of at least one visual feature of at least one 2D view. This makes it possible to generate variations in the training data and to help improve the robustness of the AI model to changes in the appearance of objects.
[0024] In one example, the method includes a preprocessing of at least one visual feature of at least one 2D view. The standardization of visual feature(s), such as lighting or color, can help limit unwanted variations in the training data and improve the performance of the AI model. In one example, the aforementioned alteration processing and / or standardization preprocessing are also applicable to 2D images acquired in the context of the method implemented by at least one first computing unit.
[0025] According to another aspect, a first computing unit is proposed configured to provide, from at least two 2D images of a 3D environment including a 3D object, the images being acquired by sensors whose relative poses to each other are fixed, at least one probability associated with at least one class of the 3D object representative of an occupation by the 3D object in an elementary volume of at least one spatial mesh of the 3D environment.
[0026] According to another aspect, a second computing unit is proposed configured to obtain 2D views of at least one 3D object, the views being acquired or generated from fixed relative viewpoints, and to train an artificial intelligence model with the 2D views to predict at least one class of the 3D object in elementary volumes of a spatial mesh of a 3D environment.
[0027] According to another aspect, an electronic device (for example, a terminal, a server, etc.) is proposed, comprising the first computing unit and / or the second computing unit. According to another aspect, a system is proposed, comprising a first electronic device comprising the first computing unit and a second electronic device comprising the second computing unit.
[0028] According to another aspect, a computer program is proposed that includes instructions for implementing all or part of a process as defined herein when this program is executed by a processor. According to another aspect, a non-transient, computer-readable recording medium is proposed on which such a program is recorded. Brief description of the drawings
[0029] Other features, details and advantages will become apparent from reading the detailed description below and from analyzing the accompanying drawings, in which: Fig. 1
[0030] [Fig.l] schematically represents a 3D object recognition and / or localization system in an example embodiment. Fig. 2
[0031] [Fig.2] schematically represents a set of image sensors configured to simultaneously acquire several images of an object in an example embodiment. Fig. 3
[0032] [Fig.3] schematically represents a set of images simultaneously acquired by the set of sensors of [Fig.2] in an example embodiment. Fig. 4
[0033] [Fig.4] represents an example of a three-dimensional mesh of an environment in an example embodiment. Description of the implementation methods
[0034] In the description that follows, identical reference numerals designate identical elements or elements having similar functions.
[0035] Some specific terms are now clarified for a better understanding of the proposed technique.
[0036] A computing unit refers to a hardware or software device capable of processing data and performing calculations. For example, one or more processors in one or more computers (such as terminals, smartphones, connected devices, etc.) and / or one or more servers. In the context of this document, the term "processor" does not imply any particular limitation in that it can refer to any type of processor having any architecture and any specific structural characteristics. It could be a CPU (Central Processing Unit), a GPU (Graphics Processing Unit) for massively parallel computing, or an FPGA (Field-Programmable Gate Array) for specific applications. This term also includes distributed computing devices, such as server clusters or cloud computing solutions.
[0037] Supply refers to the act of making available specific data or resources, such data or resources being required and / or necessary for a process. For example, this could involve providing images to an object recognition algorithm or providing a 3D model to a simulation system or an augmented or virtual reality system. This also includes, at least in some embodiments, making real-time data available for monitoring systems, delivering pre-trained models to machine learning systems, or providing access to databases for statistical analysis.
[0038] A real environment is a physical space in which one or more real objects to be identified and / or tracked are located. A real object is any tangible element present in this environment. This term excludes virtual or simulated objects which They do not physically exist. For example, a production room in a factory constitutes a real environment, and the machines, parts, and products in that room are real objects. A virtual object is something simulated or digitally represented, without any tangible physical existence. For example, a 3D model of a machine created using computer-aided design (CAD) software is a virtual object. Virtual objects can be positioned within a virtual environment. A virtual environment is a digital or simulated space created using computer software to reproduce real or imagined conditions. Unlike a real environment, it does not physically exist and is accessed via digital devices such as computer screens, virtual reality headsets, or simulation interfaces.For example, a virtual factory modeled to test the placement of new machines before their actual installation constitutes a virtual environment.
[0039] An image of an environment is a visual representation of that environment. An image can refer to a 2D view, such as a photograph taken by an RGB camera. In this application, a 2D view is considered not to contain depth information. It is therefore different from a 3D view or a volumetric model, which include depth or volume information. Images can be photographs, screenshots, computer-generated simulated images, or any other form of two-dimensional visual representation. Depth maps, although 2D, contain depth information and are therefore not considered 2D views in the strict sense of this definition. 2D views can be in color (RGB), grayscale, infrared, ultraviolet, or other spectra, depending on the sensors or simulation methods used to capture or generate these images.2D RGB views provide a color representation of the environment, while greyscale views represent light intensity without color distinction. Infrared and ultraviolet views capture information outside the visible spectrum, which can be useful in certain specialized applications.
[0040] An image acquisition sensor is a device capable of capturing one or more images. For example, an RGB camera that captures color images according to the three red, green, and blue channels. An acquired video stream is a continuous sequence of images captured by such a sensor. To specify that an image set of an environment or object includes images acquired simultaneously or almost simultaneously by several sensors means that these images were captured at the same time (i.e., within the same time window, for example, of less than 100 ms) by different sensors arranged in or around that environment. For example, several cameras arranged around a moving object can to capture images simultaneously to allow for an accurate reconstruction of its position and orientation at different times.
[0041] Sensors with fixed relative poses are arranged such that they remain stationary relative to each other, both in terms of position and orientation. The pose thus represents the sensor's "viewpoint." For example, two cameras mounted on a rigid support where the distance and angle between them do not change over time. This does not include configurations where the sensors can move independently. Such an arrangement ensures the stability of measurements and the accuracy of analyses, because the spatial relationships between the sensors remain constant.
[0042] By extension, in the context of generating simulated views of a modeled object, it is possible to use modeling or simulation software to generate these views of the modeled object from fixed viewpoints, or from viewpoints that are mobile but fixed relative to each other in terms of position, orientation, and optionally, field of view, so that the distance and angle between these viewpoints remain constant. For example, a 3D model of an object can be rendered from several distinct viewpoints to simulate shots from different virtual cameras arranged around the object in a fixed manner relative to each other. Such an approach makes it possible to create a coherent set of 2D views, thus facilitating the training of artificial intelligence models.
[0043] The class of an object designates its category or type, such as "chair," "table," or "robot." The pose of the object refers to its exact location in space, that is, its position and orientation. The three-dimensional position of an object can be defined in point mechanics by three coordinates expressed as values (for example, x, y, z in a Cartesian coordinate system), these coordinates being those of a reference point of the object, for example, its center of gravity. Several sets of coordinates can be defined for as many points of the same object in order to provide a representation of that object in rigid body mechanics. Such an approach can, for example, be advantageous in the case where the object in question is articulated or deformable. Orientation designates the angle or direction in which the object is oriented. The three-dimensional orientation of an object can be defined by three angular coordinates relative to three reference axes.For example, a chair can be classified under the "furniture" class, have a pose defined by three position coordinates (x, y, z) and an orientation defined by three angular coordinates (a, [3, Y)- .
[0044] A spatial mesh is a division of space into a three-dimensional grid composed of small elementary volumes, for example cubes or parallelepipeds. Each elementary volume is a small section of this mesh, like a 1 cm³ cube in a larger space. This term excludes two-dimensional meshes but can include irregular divisions of space, meaning that the shape and dimensions of the elementary volumes are not necessarily uniform throughout the entire environment. A spatial mesh allows for modeling space with precise granularity, useful for detailed analyses of how objects occupy space.
[0045] The occupation of an elementary volume by an object refers to the presence of an object within that volume. The presence or absence of an object within a volume is characterized by the application of a conventionally chosen criterion. For example, according to one possible criterion, every object may be considered to occupy a single elementary volume, determined to be the elementary volume where the object's center of gravity is located. Alternatively, according to another possible criterion, an object may be considered to occupy an elementary volume as soon as an overlap occurs, that is, as soon as at least a portion of the object extends over at least a portion of the elementary volume in question. Other examples of applicable criteria include the presence of a certain density of the object's material within the volume considered, or a percentage of the elementary volume occupied by the object that is greater (or less) than a certain numerical value, acting as a threshold.
[0046] An artificial intelligence (AI) model is an algorithm designed to perform specific tasks by processing data. Training is the process by which this model learns from data called "training data" to improve its performance. In particular, this training data can be labeled. This is known as guided learning. In deep learning, it is common to use training data to fine-tune a pre-trained model. This allows a generalist AI model to be adapted to a specific task or its behavior to be adjusted. Prediction is the application of the trained model to perform classifications or estimations on new data. For example, an AI model can be trained on images of chairs and tables to predict the class of objects in new images.This can include tasks such as object recognition, image segmentation, motion detection, etc.
[0047] The AI model can, for example, be an artificial neural network. Convolutional neural networks (CNNs) are particularly well-suited to object recognition in high-dimensional images and can learn to identify complex visual features through successive convolutional layers. Recurrent neural networks (RNNs) and their variants, such as Long Short-Term Memory (LSTM) networks, are often used to process sequential data and can thus be relevant for tracking moving objects, where a relationship Temporal relationships can be established between successively acquired images provided as input. Generative adversarial networks (GANs) can be used to generate realistic images and enrich a training dataset by creating variations of existing images, thereby improving model robustness. Transformers were initially designed for natural language processing and have recently shown promising performance in computer vision, particularly in segmentation and object detection tasks, thanks to their ability to capture long-term dependencies in image data.
[0048] A visual feature of an image is a specific aspect of the image, which may affect, for example, the color, texture, or shape of the visible object(s). For example, the red color of an apple or the rough texture of a wall. Other visual features may include contours, patterns, shadows, reflections, etc. Computer vision algorithms generally rely on these features to perform recognition and analysis tasks.
[0049] A visual feature alteration process for an image is designed to modify one or more of the image's visual characteristics. For example, applying a blur filter to alter sharpness or adjusting color saturation. This excludes processes that do not modify visual characteristics, such as lossless data compression. Other examples of processing include smoothing, contrast enhancement, geometric transformation (rotation, translation), etc.
[0050] A processing or preprocessing operation for the uniformity of a visual characteristic of an image is designed to standardize the visual characteristics of different images to make them more uniform. For example, adjusting the lighting for all images so that they have a similar brightness (or reducing luminance differences). This term excludes processing operations that aim to diversify or increase the differences between images. Other examples of processing or preprocessing include color normalization, gamma correction, noise reduction, etc.
[0051] It is present refers to [Fig.1], which represents an example of a 3D object recognition and / or localization system, according to one aspect of the present technique.
[0052] The 3D object recognition and / or localization system includes, in particular, the following modules: a module 3 for simulating viewpoints of a set of sensors, a module 4 for training an artificial intelligence model, a module 8 for implementing the artificial intelligence model, and a module 5 for determining the positions and / or orientations of objects.
[0053] In the example considered, the system is also connected to the following entities: a module 1 for supplying object models, a module 2 for supplying sensor parameters, the set 9 of sensors, and a module 10 for cataloging objects.
[0054] Any suitable hardware and / or software may be used for the practical implementation of said modules. In general, although aspects of the proposed technique may be described in this document as a process, device, system, procedure, method or method, it should be noted that the proposed technique may also cover computer memory that can be connected to a processor possibly connected to a communication interface, the memory storing instructions which, when executed by such a processor, enable the implementation of the processes, devices, systems, procedures, methods or methods described in this document.
[0055] The individual operation of the different modules thus listed is now detailed.
[0056] The object model supply module 1 "FOURN MDL OBJ" is configured to provide one or more 3D models, which are digital representations of physical objects to be detected, identified, located, and / or tracked. The 3D models can be provided in any format, for example, a standard format such as .ply or .obj. The 3D models can, for example, be stored in one or more 3D model libraries and made accessible by Module 1, which provides object models.
[0057] The sensor assembly 9 “ENS CAPT” comprises a plurality of image acquisition sensors as defined in this application. The sensor assembly may further comprise a structure for holding the sensors in fixed relative poses with respect to each other. The sensor assembly may be fixed or movable in translation and / or rotation.
[0058] By way of example, [Fig. 2] shows an array 9 of three image sensors bl, b-2, b-3 fixedly arranged in an environment and configured to simultaneously acquire three images of the same object 6 from three distinct angles. [Fig. 3] shows the three images 6-1, 6-2, 6-3 thus acquired.
[0059] The sensor parameter supply module 2, "FOURN PARAM CAPT", is configured to provide characteristics of the sensor set 9. These characteristics include, in particular, sensor placement data for the sensor set 9 (for example, for some or all of these sensors), such as the absolute position of at least one of these sensors (for example, of all of these sensors) in any reference frame, or the relative position of at least two of these sensors. The sensor characteristics may also include orientations, thus defining a complete pose for each sensor. For example, the position of a sensor can be defined by six coordinates: three for position and three for orientation, referred to as "6D Pose." Alternatively, the pose can be defined by three position coordinates if the orientation is fixed, referred to as "3D Pose," or by three position coordinates and one angular coordinate around a single axis, referred to as "4D Pose," or even by a single coordinate if the other five are known or fixed for a specific reason, such as when the sensor array is only translationally mobile along an axis. Characteristics can also include optical features such as focal length and principal points, as well as sensor type, image resolution and format, lighting parameters, and so on.For example, an RGB camera might have a focal length of 35mm, main points located in the center of the sensor, and a resolution of 1920x1080 pixels.
[0060] The "SIMUL PDV CAPT" viewpoint simulation module 3 is configured to use a 3D model of an object and the characteristics of the 9-sensor set to generate simulated 2D images of the real environment. These simulated views aim to provide a range of possible representations of the object from different viewpoints corresponding to those of the various sensors in the 9-sensor set, when the object is positioned in a variety of possible positions and orientations within the real environment. These simulations can also include various visual effects to create a diverse and realistic dataset necessary for AI training. For example, by simulating varying lighting conditions and poses of the object, module 3 can create diverse scenarios to improve the robustness of the AI model.One advantage of this module is its ability to generate simulated image sets of an object. These image sets are similar to those acquired by real sensors, preserving the same viewpoints and visual effects, so that AI training is aligned with the task of recognition and localization under real-world conditions. In a variation, the object's positions and orientations can be generated randomly to cover various possible configurations (ideally a large number), allowing training with a wide range of potential situations and thus helping to improve the reliability of the AI model.
[0061] The AI training module 4, "ENTR MOD IA", is configured to train or retrain the AI model using, for example, sets of simulated images. The training or retraining aims at least at classifying objects in the images and optionally at determining their position and orientation. This allows the AI to learn to identify and locate objects under various conditions. For example, using simulated images of mechanical parts. Under different angles and lighting conditions, AI can be trained to accurately recognize and locate these parts in real images.
[0062] The AI training module 4 can also be configured, as an alternative or in addition, to train the AI model using real shots.
[0063] Regardless of the origin of the images used for training, i.e. whether the images are acquired or simulated, these images are annotated (in other words labeled or tagged), i.e. the class of the object represented on the training images is known, as well as its position in 3D and optionally its orientation.
[0064] Consequently, the training process may involve the use of a three-dimensional spatial mesh of the environment, where each elementary volume of the mesh is associated with a probability of being occupied by a specific object class. The AI can, for example, be trained to predict the object class in each elementary volume. More generally, the AI can be trained to predict not only the class, but also the position and / or orientation of the object in space, using, for example, positional and / or angular coordinates. This type of detailed training enables the AI to generate accurate and robust predictions when confronted with new real-world data.
[0065] Taking up the example of [Fig. 3], the set of three images 6-1, 6-2, 6-3 simultaneously acquired of the object 6 by the image sensors 9-1, 9-2, 9-3 can, for example, be processed so as to form a tensor of size w * h * ncanai * ncamera, where: w is the width of an image (in pixels), h is the height of an image (in pixels), ncanai is the number of channels of an image (for example 3 if the image sensor used is an RGB camera), and nCamera is the number of images simultaneously acquired, each by a separate image sensor. The tensor thus obtained can be directly supplied as input to module 8 for the implementation of the artificial intelligence model.
[0066] Fig. 4 represents an example of a regular three-dimensional mesh 7 dividing the environment into elementary cubes.
[0067] The environment can for example be divided into elementary rectangular cubes aligned with the main axis of one of the cameras.
[0068] Knowing the camera positions, a three-dimensional mesh can be aligned with the principal axis of one of the cameras, which may be fixed or moving. Thus, this mesh can vary depending on the camera position. In this case, two approaches are possible. The different meshes, varying according to the camera position, can be transformed using position and / or orientation information from the The camera is relative to the environment. This transformation aligns all meshes, ensuring a common reference for determining object positions. Alternatively, object classes and positions can be trained and determined directly in the variable mesh associated with the moving camera. The results obtained in this variable mesh can then be transformed and projected to obtain a single reference, independent of the camera's pose.
[0069] In this example, the images 6-1, 6-2 and 6-3 of the object 6 - in this case a pair of pliers - were simultaneously acquired while the object 6 was located in one of the elementary cubes, with coordinates [1, 0, 0].
[0070] The AI can be trained from a finite number of classes, including for example a "Grip" class indicating the presence of a gripper-type object and an "Empty" class indicating the absence of an object, with the objective of predicting: for the elementary cube with coordinates [0, 0, 0] in which no object is located, a 100% probability of the "Empty" class, and for the elementary cube with coordinates [1, 0, 0] in which a gripper is located, a 100% probability of the "Grip" class.
[0071] AI can also be trained to predict, for the elementary cube with coordinates [1, 0, 0], in addition to the probability associated with the class of the object, a translation matrix (for example [0,5 ; 0,5 ; 0,5] when the center of gravity of the object is located at the center of the elementary cube) and / or a rotation matrix (typically a 3x3 matrix) reflecting respectively the position and orientation of the gripper within this elementary cube.
[0072] Module 8, "MOD AI," which implements the artificial intelligence model, represents the artificial intelligence algorithm that has been trained for the specific task of object recognition and, optionally, object localization. It is configured to use data provided by other modules to make predictions on new images. For example, once trained, the AI model can analyze real-time video streams to detect and identify objects in an industrial environment.
[0073] Module 5 for determining object positions and / or orientations, "DET POS OBJ", is configured to calculate the 3D positions and optionally the 3D orientations of objects detected by the AI model from images captured by the set of sensors 9. This module queries the AI implementation module 8, providing it with the acquired input images to obtain as output a probability of the presence of an object belonging to a given class in an elementary volume of the environment. For example, Module 5 can send time-stamped images captured by the sensors to the AI model, which analyzes each image to determine, for each elementary volume, the probability that this elementary volume contains an object and / or or the associated class probability. The AI model can also provide, for each elementary volume, the coordinates and orientation of the object within that elementary volume. This data can be used by Module 5 to update the object's position and / or orientation information by filtering the AI results based on the object's probability of presence. For example, it can verify that the probability of the object's class is indeed the highest probability among the different classes for each elementary volume. As another example, it is also possible to filter the probabilities determined for the elementary volumes based on at least one criterion.For example, it might be possible to compare the probability class of the object being sought within an elementary volume with a first value (used, for example, as a threshold) and determine, based on the result of the comparison, whether or not the detected object occupies that elementary volume. Other possible criteria include the spatial continuity of the elementary volumes occupied by the object or the consistency of the orientations determined for these elementary volumes. If, by applying such a criterion, it is determined that the detected object occupies several elementary volumes, it is possible to associate it with the average of the positions and orientations of the elementary volumes selected according to this criterion. This process allows for precise, real-time tracking of objects in the monitored environment.
[0074] The elementary mesh can be fixed relative to the environment; in this case, obtaining the position of the objects is direct. The position coordinates trained and recovered by the model can be: absolute (expressed as positional values, for example in meters, relative to a fixed point in the environment or relative to a point in the elementary volume, or relative, expressed as percentages of the width, length and / or height of the actual environment or elementary volume.
[0075] Alternatively, the elementary mesh may not be fixed relative to the environment but mobile relative to at least one of the cameras; that is to say, for example, the position and orientation of the elementary mesh may be slaved to the position and orientation of said at least one of the cameras. In this case, obtaining the 3D position of the objects may involve projecting the position coordinates trained and recovered by the model to position coordinates independent of the camera poses in the environment. In other words, the projection is equivalent to a change of reference frame, that is, the expression of coordinates—initially known in a frame of reference attached to a camera—into a frame of reference independent of the position of said camera. The projection may take into account the (known) position and orientation of the camera in the 3D environment. Alternatively, the projection can take into account the position and orientation of an elementary reference volume of the mesh relative to a reference element of the environment.
[0076] Position and / or orientation information can be used to update a digital twin, for example in real time. For example, by calculating the exact position of a product on a production line, the digital twin helps to track the product at each stage of the manufacturing process.
[0077] To this end, Module 10, the "RCS OBJ" object census module, is configured to store and / or update information relating to the identified objects and / or their position and / or orientation. This storage and / or updating can be implemented by Module 10 in a digital twin, designed to provide a 3D view of the monitored environment, which can, for example, be displayed directly and / or made accessible via an API for various industrial applications. For example, the digital twin can be used to visualize the layout and movement of machines in a factory in real time.
[0078] Module 1 provides the 3D model of the object to be detected, while Module 2 provides the relative characteristics and poses of the fixed cameras. Module 3 uses this information to simulate images of the object from different angles, thus creating a diverse dataset. These simulated images are then used by Module 4 to train the AI model. During training, the model learns to identify the object's class, position, and orientation from the images. For example, using 3D models of mechanical parts and the parameters of cameras installed around the assembly line, the system can simulate different views of the parts and train the AI to recognize them.
[0079] Once the AI model is trained, the fixed cameras b capture images of the real environment (continuously, periodically, or following an event, for example). The AI model analyzes these images to detect and identify the objects present. Module 5 calculates the positions and / or orientations of the detected objects and transmits this information, along with the object identifiers, to Module 10. Module 10 updates the digital twin with the data of the identified objects, thus providing a real-time 3D view of the monitored environment.
[0080] In summary, the 3D object recognition and localization system uses a series of interconnected modules to provide a solution that, in certain embodiments, can range from supplying object models and sensor parameters, to training the AI, to real-time detection and updating of a digital twin. This system can assist in the precise monitoring and management of objects in various real-world environments, thereby improving efficiency and the precision of the operations carried out or likely to be carried out in these environments.
[0081] It should be noted that certain modules may be replaced or modified, or be optional, depending on the embodiment (for example, according to the specific needs and constraints of the users and the desired applications of the proposed technique).
[0082] For example, the viewpoint simulation module 3 can be omitted if sufficient real-world data is available for training the AI. This real-world data may, for instance, have been previously acquired by sensor set b or by similar sensor sets in similar environments. As an example, sensor set b may consist of at least two sensors arranged in a fixed configuration within a single system such as an augmented, virtual, or mixed reality device. Since such a device is intended to be manufactured in large quantities and used by a large number of users in a variety of environments, it is expected that a wide variety of real-world data can be acquired by the device units and that data relating to the class and position and / or orientation of objects in the users' real-world environment will be used for training an AI model.To address the issue of personal data protection, it may be possible to implement appropriate learning techniques, such as federated learning techniques.
[0083] Similarly, the object position and / or orientation determination module 5 can be replaced by a module using real-time tracking techniques if the objects are in constant motion. For example, in an environment where objects follow predictable or repetitive trajectories, such as on an assembly line, a real-time tracking module may be more appropriate. This module can use motion-tracking algorithms to estimate the position and orientation of objects in real time from images captured by the sensor set 9. The tracking techniques may include Kalman filter algorithms, image correlation algorithms, or other visual tracking methods based on the characteristics of the objects.Furthermore, in applications where objects move quickly, such as in sports or fast-moving industrial processes, real-time tracking can help meet responsiveness and accuracy requirements related to the speed of object movement.
[0084] It is also possible to consider variants where certain modules are combined to simplify the system. For example, module 4 for AI training and module 8 for implementing the artificial intelligence model could be integrated into a single module that manages both training and application. of the model. This would reduce the complexity of the system and optimize the necessary hardware and software resources.
[0085] Another possible adaptation involves using mobile sensors while keeping them fixed relative to each other in sensor set b. For example, drones equipped with cameras could be used to capture images of the environment from different angles and perspectives. This approach would be particularly useful in environments where the objects or the monitored space change frequently, such as on a construction site or in natural environments. The drones can be programmed to follow specific trajectories and capture images systematically, thus providing varied data for training and applying the AI model.
[0086] In certain applications, it can be advantageous to use sensor fusion techniques to improve the accuracy and robustness of object recognition and localization. For example, by combining RGB image data with depth data captured by lidar sensors or depth cameras, the system can obtain a richer and more accurate representation of objects and their surroundings. This data fusion can be particularly useful in complex or cluttered environments, where visual information alone may not be sufficient for reliable object recognition.
[0087] Furthermore, the 3D object recognition and / or localization system can be adapted to work with heterogeneous data from different sources. For example, images captured by security cameras installed in a warehouse can be combined with location data provided by RFID tracking devices attached to objects. This data integration improves localization accuracy and provides a more complete and integrated view of the monitored environment.
[0088] Finally, it is possible to consider integrating the 3D object recognition and / or localization system with data management platforms to enable more in-depth analysis and optimized use of the collected information. For example, object recognition and localization data can be stored in a centralized database and analyzed using big data techniques to identify trends, optimize production processes, or improve inventory management. This integration with advanced analytics tools makes it possible to fully leverage the system's capabilities for a variety of industrial and commercial applications. Industrial application
[0089] The present technical solutions can be applied in various technical fields and environments.
[0090] By way of non-exhaustive notice, the following fields may be cited: manufacturing, logistics, inventory management, robotics, security and surveillance, transport management, embedded systems for example for automobiles, construction, agriculture, augmented and virtual reality, etc.
[0091] Possible environments can refer to both closed environments, for example a domestic environment, a warehouse, a store, a factory, a construction site, a hospital, or any other site, and more open environments, such as airports or train stations, farms, urban areas, etc.
[0092] Systems and devices, which can accommodate sensors whose relative positions with respect to each other are fixed, include aerial drones for example for monitoring large areas, mobile robots for example for inspecting industrial sites, fixed video surveillance systems for example for building security, autonomous vehicles for example for navigation and driving assistance, or even portable augmented reality devices.
[0093] Such systems and devices may further include a computing unit configured to provide, from the images captured by these sensors, at least one probability associated with at least one class of the real object representative of an occupation by the real object in an elementary volume of a spatial mesh of the real environment. Alternatively, such systems and devices may include a communication interface to such a computing unit.
[0094] Systems that can integrate a computing unit with the modules necessary for training a model on simulated or real images include data center servers dedicated to training AI models, development platforms, and simulation systems for training models in virtual environments. For example, a data center server can generate thousands of simulated images from 3D models to train an AI model. An AI development platform can be used by engineers to create and fine-tune models using simulated data before deploying them on embedded systems.
[0095] This disclosure is not limited to the examples described above, which are only examples, but encompasses all the variations that a person skilled in the art may consider in the context of the protection sought.
Claims
Demands
1. A method implemented by at least one first computing unit, the method comprising: providing, from at least two 2D images of a 3D environment including a 3D object, the 2D images being acquired by sensors whose relative poses to each other are fixed, at least one probability associated with at least one class of the 3D object representative of an occupation by the 3D object of at least one elementary volume of a spatial mesh of the 3D environment.
2. A method according to the preceding claim, wherein at least one probability further represents a probability of position and / or orientation of the 3D object.
3. A method according to the preceding claim, wherein the environment is a real environment and the method comprises a generation of a digital twin of the real environment, the digital twin taking into account at least one probability.
4. A method according to any one of the preceding claims, wherein the 2D images are derived from video streams acquired by the sensors and the method is implemented for several sets of 2D images acquired in the same time window by the sensors.
5. A method according to the preceding claim, comprising tracking a position and / or orientation of the 3D object taking into account the probabilities obtained.
6. A method according to any one of the preceding claims, wherein the spatial mesh is relative to a first position of at least one of the sensors.
7. A method according to the preceding claim, comprising determining the position of the 3D object in the 3D environment from at least one probability and the position of said first sensor.
8. A method implemented by at least one second computing unit, the method comprising: obtaining 2D views of at least one 3D object, the views being acquired or generated from viewpoints fixed relative to each other, and
9.
10. training an artificial intelligence model with 2D views to predict at least one class of the 3D object in elementary volumes of a spatial mesh of a 3D environment. Method according to the preceding claim, comprising an alteration treatment of at least one visual characteristic of at least one 2D view. A method according to any one of claims 1 to 7, comprising a pretreatment for standardizing at least one visual feature of at least one 2D view.
Citation Information
Patent Citations
Scene understanding using occupancy grids
US20240127538A1