Object recognition method with increased representativeness

DE602020061736T2Active Publication Date: 2025-11-05THALES SA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
DE602020061736
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-10-10
Filing Date
2020-10-08
Publication Date
2025-11-05
Estimated Expiration
2040-10-08

AI Technical Summary

Technical Problem

Current digital image databases for object recognition lack variability in viewing angles, occlusion, and noise levels, particularly in human recognition, limiting the effectiveness of supervised learning algorithms.

Method used

A method involving 3D volume reconstruction from multiple 2D images, generating diverse 2D images with varying angles, occlusions, and noise levels, and training a neural network on this expanded dataset to improve recognition confidence.

Benefits of technology

Enhances the representativeness of training datasets, improving recognition confidence even in degraded 2D images by generating a new plurality of 2D images from the reconstructed 3D volume, thereby increasing the reliability of object identification.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader
Need to check novelty before this filing date? Find Prior Art

Description

FIELD OF INVENTION

[0001] The invention relates to object recognition in digital imaging. It generally applies to the automatic recognition of objects within digital images taken under difficult conditions, in particular to the recognition of human beings from randomly taken two-dimensional (2D) digital images or of objects from digital images taken under difficult conditions (fog, large distance, occlusion of the object, shooting angle, low resolution image...). CONTEXT OF THE INVENTION

[0002] The field of Artificial Intelligence (AI) is currently experiencing exponential growth in multiple sectors. This growth is explained by the convergence of three concurrent elements: the development of learning algorithms called "Machine and / or Deep Learning"; the emergence of large databases on the internet ("big data"); and the increase in computing speed, enabling the training of learning algorithms.

[0003] In general, object recognition in AI relies on training datasets. In practice, each training dataset includes input data that leads to the creation of a model to provide an output called the image label. For example, in supervised learning (classification), the output is known, and the goal is for the algorithm to learn to respond automatically and provide the label of the object thus recognized in the image being processed. The article by Jun Liang et al., "Object Recognition Based on Three-Dimensional Model," in "Pervasive: International Conference on Pervasive Computing," January 31, 2012, Springer, ISBN: 978-3-642-17318-9, vol. 7202, pages 218-225, describes a method for recognizing an object of interest within a 2D digital image.

[0004] It is known that training a supervised learning algorithm requires a large amount of labeled input data. However, current digital image databases generally rely on labeled object images with relatively limited and / or rudimentary variability in terms of viewing angle (image transformation via rotations, shifts, noise addition / removal, etc.). Furthermore, variability in human recognition is relatively restricted (for example, images of the person to be recognized based solely on their face). SUMMARY OF THE INVENTION

[0005] The aim of the present invention is to improve the situation, in particular by providing a solution that at least partially alleviates the aforementioned disadvantages.

[0006] To this end, the present invention proposes a method for recognizing an object of interest within a degraded 2D digital image of said object.

[0007] According to a general definition of the invention, the process comprises the following steps: to first detect the object of interest within a 2D digital image and assign it a label; to reconstruct a three-dimensional (3D) volume of said labeled object from a plurality of available 2D digital images of said object of interest; to store in a database at least one record relating to said reconstructed and labeled 3D object; for each record thus stored, to generate a new plurality of 2D digital images according to a plurality of viewing modes from the reconstructed 3D volume of each object, the viewing modes including viewing modes at different levels of occlusion and / or noise addition; to train a neural network on a training set composed of an expanded set of 2D digital images thus generated and corresponding to the label of the object of interest to be recognized; from a degraded 2D digital image of said object of interest to be recognized;use the neural network thus trained to output the label of the object and a confidence index linked to the recognition of the object of interest. ;

[0008] Surprisingly, the Applicant observed that generating a new plurality of 2D digital images from the reconstructed 3D volume of the object increases the representativeness (variability) of the training datasets and thus improves the confidence index of recognition on a 2D image of the object to be recognized, even if the 2D image is degraded.

[0009] According to preferred embodiments, the invention comprises one or more of the following features which can be used separately or in partial combination with each other or in total combination with each other: If the confidence level is above a threshold, stop the recognition; otherwise, search for other elements to increase the success of the identification. As a non-limiting example, the 3D volume reconstruction is of the reflective tomography type. As a non-limiting example, the size of the reconstructed 3D volume of the object is 262 x 262 x 257 pixels. The plurality of 2D images resulting from the reconstructed 3D volume of the object belong to the group formed by 2D images at various angles (theta, phi, Phi...).); images at different distances; images with different occlusion rates, images with different noise levels; the plurality of 2D images from the reconstructed 3D volume for objects of interest, such as human beings, belong to the group formed by accessories, hat, eyeglasses, sunglasses and beard; as a non-limiting example, the resolution of the 2D digital images thus generated is 124 pixels x 253 pixels; as a non-limiting example, the neural network is a convolutional neural network of the type ResNet50, ResNet101 or ResNet152 (Residual Network with 50, 101 and 152 layers of neurons respectively).

[0010] The invention further relates to a computer program comprising program instructions for the execution of a process as previously defined, when said program is executed on a computer.

[0011] Other features and advantages of the invention will become apparent from the following description of a preferred embodiment of the invention, given by way of example and with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Other advantages and features of the invention will become apparent upon examination of the description and drawings in which: There figure 1 schematically represents the main steps of the recognition process according to the invention; The figure 2 schematically represents the sub-steps of the database creation step according to the method of the invention; The figure 3 schematically represents examples of 2D images of a boat taken in SWIR (Short Wave Infrared) in a horizontal plane from 9 different angles for the 3D reconstruction of the object; The figure 4 schematically represents the sub-steps of the step of generating an expanded database according to the method according to the invention; and The figure 5 schematically represents examples of 2D image database records for the object labeled "boat2E0A0" generated at various viewpoints and distances from the 3D volume of the object reconstructed from the 2D images of the figure 3 .

[0013] With reference to figures 1 à 5 We have represented the three main stages of the process of automatic object recognition in difficult conditions by an AI trained on a database of labeled images whose increased representativeness was achieved via 3D reconstructions of said objects.

[0014] The first main step 10 aims to create a database of objects already identified and reconstructed in 3D.

[0015] Step 10 begins with a substep in which the object of interest 11 (for example, a boat) is first detected. Then, a rapid acquisition 12 of 2D images (visible, infrared, active or passive) is performed, in a limited but sufficient number to create a 3D reconstruction of the object. Depending on the object's context, the 2D image acquisition can be carried out according to several scenarios, such as "ground-to-ground," "sea-to-sea," "air-to-ground," and "air-to-sea." For a boat, image acquisition can be carried out according to scenarios such as "sea-to-sea" and "air-to-sea." For example, with reference to the figure 3 Examples of 2D images of a boat taken in SWIR in horizontal plane (Scenario "sea-sea") were presented from 9 different angles for the 3D reconstruction of the object.

[0016] From the 2D images thus available ( figure 3), we proceed with the 3D reconstruction 13 of the object using an appropriate reconstruction method (reflective tomography, for example). We then obtain a three-dimensional (3D) volume 14 of the object (in voxels).

[0017] In practice, the three-dimensional volume can be obtained through a reconstruction process by transmission or fluorescence (Optical Projection Tomography, nuclear imaging or X-Ray Computed Tomography) or by reflection (back reflection of a laser wave or by solar reflection in the case of the visible band (between 0.4µm and 0.7µm) or near infrared (between 0.7µm and 1µm) or SWIR (between 1µm and 3µm) or by taking into account the thermal emission of the object (thermal imaging between 3µm and 5µm and between 8µm and 12µm), this three-dimensional reconstruction process is described in the patent "Optronic system and method for developing three-dimensional images dedicated to identification" (US8836762B2, EP2333481 B1).

[0018] We use the set of voxels from a three-dimensional reconstruction with the associated intensity, this reconstruction preferably having been obtained by back-reflection.

[0019] At the end of the 3D reconstruction, we obtain a database comprising records relating to objects already identified, i.e. {Volume3D_Object(n) Label_Object(n)}, n=1,2,..,N (N being the number of records of identified objects).

[0020] It should be noted that the database can be enriched with objects from modeling or simulations.

[0021] The second main step 20 of the process according to the invention consists of generating an expanded database of 2D images in various configurations and training a dedicated AI (Artificial Intelligence).

[0022] In practice, for each labeled object in the database, 21 2D images are generated from (views of) the 3D volume thus reconstructed.

[0023] In a set of embodiments of the invention, the 3D volume is externally delimited by a 3D surface, and, if the volume is incomplete, the 3D surface is open.

[0024] For example, views from the 3D volume are obtained from various angles (theta, phi, Phi) and at different distances. In several embodiments of the invention, the 3D volume can also be modified, for example, by applying different occlusion ratios and / or adding different noise.

[0025] In a set of embodiments of the invention, the addition of noise on the 3D surface, or of an occlusion, thus results in a modification of the initial 3D surface, generating new 2D images.

[0026] For faces, the views from the reconstructed 3D volume of the human to be identified can be of different kinds and with or without accessories, hat, vision glasses, sunglasses, beard, etc.

[0027] In a set of embodiments of the invention, the accessories are locally superimposed on elements of the 3D surface, which makes it possible to modify the 3D boundary of the reconstructed volume.

[0028] The plurality of 2D digital images thus generated according to a plurality of shooting modes from the modified or unmodified 3D volume of each object are then associated with the object's Label. In this way, a large number of 2D views, corresponding to different viewpoints of the 3D volume, and where applicable, its modifications, can be added to the training database.

[0029] We then obtain the following elements: Volume3D_Object(n) →{Image2D_Object(n, theta, phi, Phi, distance, Taux_occlusion,...), Label_Object(n)}

[0030] Finally, we proceed to choose a convolutional neural network, for example of the Residual Network type such as ResNet50, to train it on a training set composed of a set of 2D digital images {Images2D_Objet(n)} thus generated and corresponding to the labels {Labels_Objet(n)}, n = 1,2,3,...,N for all the objects N of interest.

[0031] The third main step 30 consists of recognizing an object of interest from a degraded 2D image of it.

[0032] For example, the prior detection of an ObjectX of interest consists of taking one or more 2D images (in visible, infrared, active or passive) under restrictive operational conditions (degraded weather, significant distance, occlusions of the object, any angle of capture...).

[0033] Next, the convolutional neural network thus trained is used to deliver as output the label of the object of interest and a confidence index (score) linked to the recognition of the object of interest.

[0034] If the Trust Index (Score) is high (above 95%, for example), it is planned to stop the recognition.

[0035] If the Confidence level (Score) is low, then the operator can look for other elements to increase the success of the identification.

[0036] As the database of already identified and reconstructed objects grows larger, the recognition reliability of the dedicated AI is strong and implicitly the identification of any object will be a success.

[0037] As a non-limiting example, the recognition process was applied to a boat labeled "boat2E0A0" using a single 2D image taken perpendicular to the sea surface (the "air-to-sea" scenario). This image was not part of the 2D training database. The image was resized to a resolution of 124 x 253 pixels to be compatible with the AI's querying process.

Claims

1. A method for recognising an object of interest in a degraded 2D digital image of said object, characterised in that it comprises the following steps: - detecting (11), beforehand, the object of interest in a 2D digital image and assigning it a label; - reconstructing (13) a 3D volume of said object thus labelled from a plurality of available 2D digital images (12) of said object of interest; - storing, in a database, a record relating to said object thus reconstructed in 3D form and labelled; - for each record thus stored, - generating (21) a new plurality of 2D digital images according to a plurality of viewing modes from the thus reconstructed 3D volume (14) of each object, the exposure modes comprising exposure modes with different levels of occlusion and / or of added noise; - training (23) a neural network on a learning set composed of an expanded set of 2D digital images thus generated and corresponding (22) to the label of the object of interest to recognise; - from a degraded 2D digital image of said object of interest to recognise, - using (30) the neural network thus trained to deliver as output the label of the object and a confidence index linked to the recognition of the object of interest.

2. The method according to claim 1, wherein, if the confidence index is above a threshold, provision is made to stop the recognition, and otherwise search for other elements to increase the success of the identification.

3. The method according to claim 1 or claim 2, characterised in that the 3D volume reconstruction (13) of the object belongs to the group formed by reflective tomography, transmission tomography.

4. The method according to any one of the preceding claims, characterised in that the plurality of 2D images derived from the reconstructed 3D volume of the object belong to the group formed by 2D viewing mode images from the 3D volume taken at various angles (theta, phi, Phi, etc.), images taken at different distances; images with different occlusion rates, images with different noises.

5. The method according to any one of claims 1 to 4, characterised in that the plurality of 2D images derived from the reconstructed 3D volume for objects of interest, of human being type, belong to the group formed by accessories, cap, spectacles, sunglasses and beard.

6. The method according to any one of the preceding claims, wherein the neural network is a convolutional neural network of the type belonging to the group formed by ResNet50, ResNet101, ResNet152.

7. A computer program comprising program instructions for the execution of a method according to one of the preceding claims, when said program is run on a computer.