Method and device for classifying and localising objects in image sequences, and associated system, computer program and storage medium

EP4552093A1Pending Publication Date: 2025-05-14SAFRAN SA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
EP2023751324
Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-07-04
Filing Date
2023-06-30
Publication Date
2025-05-14

Smart Images

  • Figure 1.1
    Figure 1.1
Patent Text Reader

Abstract

The present invention relates to a method and device (APP) for classifying and localising an object (OBJ) in an image sequence (IMG_SEQ). The proposed method comprises the steps of: • obtaining (S210) an image sequence (IMG_SEQ) of one or more images (IMG_1, …, IMG_N) acquired by a camera (CAM); and • determining (S220), by means of a classifier-localiser module (X_NN) from the image sequence (IMG_SEQ): o a class (ATR_CLAS_OBJ) assigned to the object (OBJ), the assigned class (ATR_CLAS_OBJ) being chosen from among a list of classes; and o an estimated position (EST_POS_OBJ) of the object (OBJ); the classifier-localiser module (X_NN) being configured, on the basis of reference image sequences (TR_DATA), to minimise a multi-objective loss function (F_LOSS) representative both of a classification objective and a localisation objective for classifying and localising, respectively, the objects in the reference sequences (TR_DATA).
Need to check novelty before this filing date? Find Prior Art

Description

Description Title of the invention: Method and device for classifying and locating objects in image sequences, associated system, computer program and information medium Technical field

[0001] The present invention relates to the fields of image processing and analysis, as well as to the field of computer vision. More particularly, the present invention relates to a method and a device for classifying and locating objects in a sequence of images, and a method for configuring a classifier-locator module for objects in a sequence of images, as well as an associated system, computer program and information medium. The present invention finds a particularly advantageous, although in no way limiting, application for the implementation of embedded systems, such as autonomous vehicles, surveillance systems or navigation systems. State of the prior art

[0002] Automatic object detection in an optical scene is a technical problem that has been studied for many years. In particular, accurately identifying an object in images is a critical task for many applications, including domestic, industrial, military, etc. The example of autonomous vehicles illustrates the importance of accurately identifying an object in a visual scene, particularly to prevent collisions.

[0003] There are several devices available for identifying an object. Radar systems are a common example of an object detection device. Waves emitted by a radar system are reflected by an object, then received and analyzed by the radar system to detect the presence and determine the position of the object. However, radar systems have the following disadvantages. They require specific equipment, which is generally bulky and expensive. Radar systems can detect the presence of an object and estimate its position, but the object to be detected must have certain specific physical properties. The object to be detected must reflect electromagnetic waves, which is not always the case.

[0004] It is also known to use one or more cameras to identify an object in an optical scene. For example, by exploiting stereovision methods comparing two images of the same object taken by two cameras, it is possible to obtain information on the position of an object. However, these methods require the use of several cameras simultaneously, which implies significant space requirements and costs. In addition, the estimates produced by these methods in the case of distant objects are not of satisfactory accuracy.

[0005] Furthermore, known solutions for identifying an object from one or more cameras exploit visual information related to the object. For example, existing methods rely on measuring the size of the object, while others rely on extracting characteristic points related to the object. Existing solutions thus require that the object be represented on several pixels in the images used. In the case of distant objects, these solutions do not allow for precise identification of an object and are not satisfactory. In addition, the accuracy of these solutions is significantly degraded by the effects of atmospheric attenuation, which are significant for distant objects.

[0006] There is therefore a need for a solution that can accurately identify an object in an optical scene, even a distant object. Statement of the invention

[0007] The present invention aims to remedy all or part of the drawbacks of the prior art, in particular those set out above.

[0008] To this end, according to one aspect of the invention, a method is proposed for classifying and locating at least one object in a sequence of images, said method comprising steps of: • obtaining a sequence of one or more images acquired by at least one camera; and • determination by a classifier-localizer module from the image sequence obtained of: o at least one class attributed to said at least one object, said at least one attributed class being selected from a list of attributable classes; and o at least one estimated position of said at least one object.

[0009] The classifier-localizer module used is configured from at least one reference image sequence to minimize a multi-objective loss function representative of both a classification objective and an object localization objective in said at least one reference sequence.

[0010] By "image" we mean a set of computer data representative of an optical scene.

[0011] By "position of an object", reference is made here to an absolute or relative position of an object present on an image. Absolute position is understood to be a position defined in relation to the terrestrial reference frame (e.g. geographic coordinates). In the context of the invention, a relative position is a position defined in relation to an observer position (e.g. a distance between the object and a camera) or in relation to a previous position of the object itself (e.g. a displacement of the object, a speed). It should be noted that, according to one embodiment, the proposed method determines an absolute position of an object from the sequence of images obtained and the position (e.g. geographic coordinates) of the device implementing the method.

[0012] By "loss function" we mean a function to be minimized when configuring a module or its parameters.

[0013] By "reference sequence" is meant herein image sequences used to configure the classifier-localizer module. More generally, the expression "reference data" is also hereinafter to refer to data used to configure the classifier-localizer module.

[0014] The present invention makes it possible to accurately identify one or more objects in an optical scene, even when the objects are far away. More particularly, the present invention enables this accurate identification by reliably assigning a class to an object present in an image sequence and by accurately locating this object.

[0015] Furthermore, the classifier-localizer module is configured to minimize a multi-objective loss function representative of both an objective of classifying one or more objects in a sequence of images and an objective of localizing these objects. The present invention thus implements a joint optimization of the classification and localization tasks, two intrinsically linked tasks. In one embodiment, it can thus be considered that the classification and localization tasks are learned in parallel by the classifier-localizer module from reference sequences during its configuration. The visual information linked to an object - for example, size, shape, movement, etc. - is exploited by the classifier-localizer module to carry out these two classification and localization tasks in parallel.As a result of the joint and parallel optimization described herein, the present invention synergistically improves both the reliability of the class assigned to an object and the accuracy of the estimated position of an object.

[0016] According to one embodiment of the invention, the classifier-locator module is configured by a configuration method in accordance with the invention.

[0017] According to one embodiment of the invention, said at least one estimated position comprises an estimated position of the object for each of the images of said sequence.

[0018] This particular embodiment makes it possible to improve the accuracy of localization of an object present in an optical scene, a position of the object being estimated for each of the images of a sequence. In particular, if the acquisition times of the images of a sequence are different, this embodiment makes it possible to obtain a precise localization in time of an object of the sequence, and thus to follow the movement of this object over time.

[0019] According to one embodiment of the invention, said step of determining the classification and localization method comprises sub-steps of: • determination of data encoded by an encoder from the sequence of images obtained; • determination of said at least one class assigned by a classifier from the encoded data; and • determination of said at least one position estimated by a locator from the encoded data.

[0020] This embodiment of the invention has the following advantages. A multidimensional intermediate representation of the optical scene is obtained in the present embodiment: said encoded data. The encoded data being provided as input to the classifier and the localizer, the two tasks performed use the same intermediate representation. This use of an intermediate representation common to the classification and localization tasks makes it possible to reinforce the synergy effect linked to the joint and parallel optimization of these two tasks. In combination with the aforementioned multi-objective optimization, the intermediate representation is both representative of visual characteristics of an object necessary for classification and of visual characteristics of this object necessary for localization.Thus, the classifier, to reliably classify an object, takes advantage of the visual characteristics related to the location (for example, the size and speed of movement of the object); and, correlatively, the localizer, to accurately estimate the position of an object, takes advantage of the visual characteristics related to the classification (for example, the shape and visual appearance of the object).

[0021] According to one embodiment of the invention, the encoder implements one or more convolutions between the sequence of images obtained and filters.

[0022] By "convolution" here we mean a convolution product, the mathematical operation usually denoted *.

[0023] By performing convolutions between the image sequence and filters in this embodiment, it is possible to extract spatial and / or temporal characteristics of the object in the image sequence. According to this embodiment, the encoder produces encoded data representative of spatio-temporal characteristics of an object. This encoded data constitute an intermediate representation provided as input to the classifier and the localizer. Consequently, this embodiment makes it possible to highlight spatio-temporal characteristics of an object in a sequence of images and thus to improve the reliability of the classification and the precision of the localization carried out jointly and in parallel.

[0024] According to one embodiment of the invention, said convolutions implemented by the encoder are convolutions in the spatial domain and in the temporal domain.

[0025] Using convolutions in both the spatial and temporal domains allows object-related features to be extracted in both domains. This embodiment makes it possible to obtain information related to the object in space and time (e.g., the movement of an object over time, its speed, shape variation, etc.) from the image sequence.

[0026] According to one embodiment of the invention, said step of obtaining a sequence of images comprises a sub-step of acquisition by at least one camera of said one or more images.

[0027] By "camera" we mean a device that converts optical images into electronic images.

[0028] According to one embodiment of the invention, the classifier-locator module, used to carry out said step of determining the method, comprises a neural network.

[0029] By "neural network" is meant here an artificial neural network comprising a set of artificial neurons connected to each other, an artificial neuron making it possible to determine an activation function of weighted and combined inputs.

[0030] Using a neural network to implement the classifier-localizer module allows for reliable classification of an object in an image sequence and precise localization of this object. In particular, a neural network is able to handle complex and non-linear cases and thus achieve significantly improved performance compared to analytical models. This last point is particularly advantageous for the classification and localization of distant objects, a scenario in which the effects of atmospheric attenuation are particularly significant.

[0031] In combination with the multi-objective optimization from reference image sequences described above, this embodiment has the advantage of not requiring a priori knowledge related to the physical environment and / or camera parameters. For information, no knowledge of the physical atmospheric attenuation model is required to classify an object and estimate its position, which is a major advantage compared to existing solutions.

[0032] This embodiment also makes it possible to exploit a significant amount of implicit knowledge linked to the reference image sequences. For example, the more the reference image sequences are representative of varied scenarios, the more the neural network used benefits from the knowledge linked to these reference sequences.

[0033] Furthermore, using a single neural network to jointly and in parallel implement the two tasks of classification and localization simplifies the practical implementation in a system, for example in an embedded system such as a drone or an autonomous car. Indeed, by exploiting a single neural network for these two tasks, the hardware and software constraints of implementation (e.g. memory size, necessary computing capacity, etc.) are relaxed, while allowing the implementation of a reliable classification of a precise location.

[0034] In the case where a forward propagation neural network is used, the network is a universal approximator and is thus able to approximate any continuous function on compact subsets of the set of real numbers, provided that it contains enough neurons. Exploiting a neural network thus makes it possible to implement various image processing and analysis functions.

[0035] According to one embodiment of the invention, said one or more images of the sequence obtained are consecutive. Furthermore, the step of obtaining the sequence of images comprises a sub-step of cropping said one or more acquired images, one of said objects being centered on a determined image (i.e. a certain image) of the sequence.

[0036] By "consecutive images" we mean images that are consecutive in time, and thus images whose acquisition times follow one another chronologically in time.

[0037] The cropping operation performed in this embodiment has the effect of providing, as input to the classifier-localizer module, information on the movement of an object within the sequence of images. Since the images are successive in time, the relative movement of an object within the sequence is representative of the speed of this object. Thus, the movement information of an object is highlighted by this cropping operation and then used by the classifier-localizer module to improve the reliability of the classification and the precision of the localization.

[0038] In addition, the cropping operation also makes it possible to reduce the size of the images processed by the method, while retaining the useful part of the images comprising an object to be identified and located. With a reduced image size, and without loss of useful information, the implementation in a practical system of classification and localization tasks is simplified on the hardware and software level, without performance degradation.

[0039] Compared to existing methods for automatically classifying an object in one or more images, the present invention further improves the reliability of object classification. Indeed, existing methods are more suited to classifying an object in a single image and cannot exploit information relating to the movement of an object over time.

[0040] According to another aspect of the invention, a method is proposed for configuring an object classifier-locator module in image sequences, said method comprising a step of initializing the classifier-locator module (in particular its parameters); and at least one iteration of the steps of: • determination by the classifier-localizer module from at least one sequence of reference images of: o at least one class attributed to at least one object in said at least one reference sequence, said at least one attributed class being selected from a list of classes; and o at least one estimated position of said at least one object; • evaluating a multi-objective loss function based on said at least one assigned class, said at least one estimated position, at least one known class and at least one known position of at least one object in said at least one reference sequence, the multi-objective loss function being representative of both a classification objective and an object localization objective; • reconfiguration of the classifier-localizer module to minimize the multi-objective loss function.

[0041] The present invention makes it possible to configure a classifier-localizer module making it possible to implement a reliable classification of one or more objects in a sequence of images and a precise localization of said objects.

[0042] The classifier-localizer module used is configured to minimize a multi-objective loss function representative of both an objective of classifying one or more objects in reference image sequences and an objective of localizing these objects. The present invention thus proposes to configure a module that jointly optimizes classification and localization. These two tasks being intrinsically linked, the proposed joint and parallel optimization has a synergistic effect, as described below. Compared to separate modules configured independently to optimize classification and localization, the classifier-localizer module configured by the present invention makes it possible to implement both more reliable classification and more precise localization.

[0043] According to one embodiment of the invention, said step of evaluating the multi-objective loss function comprises sub-steps of: • evaluation of a classification loss function from said at least one assigned class and at least one known class associated with said at least one object of said at least one reference image sequence; • evaluation of a localization loss function from said at least one estimated position and at least one known position associated with said at least one object of said at least one sequence of reference images; and • evaluation of the multi-objective loss function from the result of the classification loss function and the result of the localization loss function.

[0044] The loss function is said to be multi-objective because it is representative of both a classification objective and a localization objective. This embodiment makes it possible to evaluate, on the one hand, the reliability of the assigned classes and, on the other hand, the precision of the estimated positions, and to deduce a loss function common to these two tasks. Thus, this embodiment makes it possible to independently evaluate the classification objective and the localization objective, while ensuring that the classifier-localizer module jointly optimizes these two objectives.

[0045] According to one embodiment of the invention, the multi-objective loss function is defined based on a weighted sum of the classification loss function and the localization loss function, the weighting coefficients being non-zero coefficients.

[0046] Using a weighted sum makes it possible to adjust the importance given by the multi-objective loss function to one of the two classification and localization objectives relative to the other. This embodiment makes it possible to favor, or not, one objective over the other. However, since the multi-objective loss function is representative of these two objectives, the weighting coefficients are non-zero numbers. Furthermore, this embodiment makes it possible to adapt the multi-objective loss function to cases where the classification loss function and the localization loss function produce output values ​​with different orders of magnitude.

[0047] According to one embodiment, the classification loss function is evaluated using the expression: where L EC is the classification loss function, M is the number of assignable classes from said class list, δ( C i, C r ) is a binary indicator equal to 0 if for a said object the class C i of the said class list is different from class C r known to said object and equal to 1 otherwise, with P o,i a probability determined by the classifier-localizer module that the said object belongs to class C i .

[0048] According to one embodiment, the localization loss function is evaluated using the expression: where L IMSE is the localization loss function, N is the number of images of said at least one reference sequence a binary indicator equal to 1 if a said object is present in an image of index i of said at least one reference sequence and equal to 0 otherwise, with P o,i the estimated position of said object determined by the classifier-localizer module for the index image with P r,i the known position of said object for the image of index i.

[0049] According to one embodiment of the invention, said step of reconfiguring the classifier-localizer module is carried out using a gradient descent algorithm from the result of the multi-objective loss function obtained in said evaluation step.

[0050] In this embodiment, a gradient descent algorithm makes it possible to reconfigure the classifier-localizer module, in particular its parameters, in order to minimize the multi-objective loss function. This embodiment makes it possible to configure the classifier-localizer module on the basis of reference data and thus to implement a reliable classification of an object in a sequence of images and a precise localization of the latter.

[0051] According to one embodiment, the classifier-localizer module comprises a neural network and a gradient descent algorithm is used for the reconfiguration - in this case, reference is made to a gradient back-propagation method. This embodiment makes it possible to train the neural network on the basis of reference data in order to minimize the multi-objective loss function and thus to obtain a classifier-localizer module implementing a reliable classification of an object in an image sequence and a precise localization thereof.

[0052] For example, the reconfiguration step involves updating network parameters, such as the weights of each neuron in the network. The gradient backpropagation method aims to correct errors based on the magnitude of each element's contribution to them. Weights that contribute the most to an error will be modified more significantly than weights that cause a marginal error.

[0053] According to one embodiment of the invention, said at least one sequence of reference images comprises at least one of the elements of the following group: one or more images acquired by at least one camera; and one or more synthesized images.

[0054] By "synthesized" here we mean computer-generated images.

[0055] This embodiment makes it possible to obtain reference data for configuring the classifier-localizer module.

[0056] The embodiment, where the reference images are acquired by at least one camera, allows, among other things, the acquisition of reference images representative of observed real conditions. In this way, the classifier-localizer module obtained is effective in real conditions.

[0057] The embodiment in which reference images are synthesized allows for an increase in the volume of reference data used to configure the classifier-localizer module. Generating reference images by computer further allows for the creation of reference data representative of various scenarios, for example, different atmospheric attenuation levels, different weather conditions. Therefore, this allows for the configuration of a classifier-localizer module that minimizes a multi-objective loss function for various and numerous scenarios.

[0058] When the last two embodiments are taken in combination, this makes it possible to obtain a significant number of reference image sequences representative of both observed real conditions and varied scenarios. Thus, a classifier-localizer module configured from such reference data will implement a more reliable classification and a more precise localization.

[0059] According to one embodiment of the invention, the method for configuring a classifier-locator module comprises several iterations of said steps of determining, evaluating the multi-objective loss function, and reconfiguring the classifier-locator module.

[0060] Performing several iterations of said determination, evaluation and reconfiguration steps makes it possible to successively improve, as the iterations progress, the classifier-locator module to minimize the multi-objective loss function. Typically, the greater the number of iterations, the more the classifier-locator module used will minimize the multi-objective loss function. Consequently, this embodiment makes it possible to improve the classification reliability and the localization accuracy implemented by the classifier-locator module.

[0061] Consider here the embodiment in which several said iterations are performed and a gradient descent algorithm is used to reconfigure the classifier-localizer module. It is advantageous to perform several iterations on portions of the reference data, rather than a single iteration over the entire reference data set. Indeed, this makes it possible to improve, over the iterations, the convergence towards a classifier-localizer module which minimizes the multi-objective loss function. This embodiment also makes it possible to require fewer hardware and software resources - in terms of memory size, and processing resources, etc. - to configure the classifier-localizer module. Thus, the complexity of implementing the determination method on the hardware and software level is simplified by this embodiment.

[0062] According to another aspect of the invention, there is provided a device for classifying and locating at least one object in a sequence of images, said device comprising: • an obtaining module for obtaining a sequence of one or more images acquired by at least one camera; and • a classifier-localizer module for determining from the image sequence obtained: o at least one class attributed to said at least one object, said at least one attributed class being selected from a list of classes; and o at least one estimated position of said at least one object.

[0063] Said classifier-localizer module is configured from at least one sequence of reference images to minimize a multi-objective loss function representative of both a classification objective and an object localization objective in said at least one sequence of reference images.

[0064] The classification and localization device according to this embodiment has the advantages described above in connection with the proposed classification and localization method.

[0065] According to one embodiment of the invention, said classifier-locator module used by the device is configured by a configuration method in accordance with the invention.

[0066] This embodiment makes it possible to take advantage of the previously mentioned advantages in connection with the proposed configuration method within the classification and location device.

[0067] According to another aspect of the invention, a system is proposed comprising a classification and localization device according to the invention and at least one camera configured to acquire said one or more images of the sequence.

[0068] The advantage of this embodiment is that it is possible to perform both classification and localization tasks within a practical system with a single device.

[0069] This embodiment allows for an efficient practical implementation of a surveillance or navigation system, the classification of objects in an optical scene and the localization of these objects being critical tasks for these types of systems.

[0070] According to one embodiment, the proposed system comprises a single camera.

[0071] This embodiment makes it possible to require only a single camera to perform both the classification and the localization of an object in an optical scene, which is advantageous compared to existing solutions. In addition, a camera is a commonly integrated equipment in embedded systems, a camera being inexpensive and compact. Thus, reliable classification and precise localization of an object in an optical scene are, according to this embodiment, implemented in an economical and compact manner in a practical system, such as an embedded system.

[0072] According to one embodiment, the proposed system comprises a plurality of cameras.

[0073] This embodiment, by using a plurality of cameras, makes it possible to extend the total field of vision of the proposed system. This embodiment thus makes it possible to classify and locate objects in a wider geographical area. According to one embodiment, said system is a surveillance system, or a navigation system.

[0074] According to one embodiment, the proposed system is embedded (i.e. integrated) in a vehicle such as an aircraft, a ship, a railway vehicle, a road vehicle, etc.

[0075] According to another aspect of the invention, there is provided an aircraft comprising a system according to the invention.

[0076] According to one aspect of the invention, there is provided a computer program with instructions for implementing the steps of a method according to the invention, when the computer program is executed by at least one processor or computer.

[0077] The computer program may consist of one or more subparts stored in the same memory or in separate memories. The program may use any programming language, and may be in the form of source code, object code, or intermediate code between source code and object code, such as in a partially compiled form, or in any other desirable form.

[0078] According to one aspect of the invention, there is provided a computer-readable information medium comprising a computer program according to the invention.

[0079] The information carrier may be any entity or device capable of storing the program. For example, the carrier may include a storage medium, such as a non-volatile memory or ROM, for example a CD-ROM or a microelectronic circuit ROM, or a magnetic recording medium, for example a floppy disk or a hard disk. On the other hand, the storage medium may be a transmissible medium such as an electrical or optical signal, which may be conveyed via an electrical or optical cable, by radio or by a telecommunications network or by a computer network or by other means. The program according to the invention may in particular be downloaded onto a computer network. Alternatively, the information medium may be an integrated circuit in which the program is incorporated, the circuit being adapted to execute or to be used in the execution of the method in question. Brief description of the drawings

[0080] Other features and advantages of the present invention will become apparent from the description provided below of embodiments of the invention. These embodiments are given by way of illustrative example and are not intended to be limiting. The description provided below is illustrated by the attached drawings: [Fig. 1] Figure 1 schematically represents an example of a device for classifying and locating at least one object in a sequence of images according to one embodiment of the invention; [Fig. 2] Figure 2 schematically represents an example of a sequence of images obtained and processed by a device for classifying and locating at least one object in a sequence of images according to an embodiment of the invention; [Fig. 3] Figure 3 schematically represents an example of functional architecture of a device for classifying and locating at least one object in a sequence of images according to one embodiment of the invention; [Fig. 4] Figure 4 represents, in the form of a flowchart, steps of a method for classifying and locating at least one object in a sequence of images according to one embodiment of the invention; [Fig. 5] Figure 5 represents, in the form of a flowchart, steps of a method for configuring a classifier-localizer module to implement a classification and localization of at least one object in a sequence of images according to an embodiment of the invention; [Fig. 6] Figure 6 schematically represents an example of functional architecture of a device for classifying and locating at least one object in a sequence of images according to one embodiment of the invention; [Fig. 7] Figure 7 schematically represents an example of software and hardware architecture of a system for classifying and locating at least one object in a sequence of images according to one embodiment of the invention; [Fig. 8] Figure 8 schematically represents an example of functional architecture of a device for classifying and locating at least one object in a sequence of images according to one embodiment of the invention. Description of the embodiments

[0081] The present invention relates to a method and device for classifying and locating objects in an image sequence, and to a method for configuring an object classifier-locator module in image sequences, as well as to an associated system, computer program and storage medium.

[0082] Figure 1 schematically represents an example of a device for classifying and locating at least one object in a sequence of images according to one embodiment of the invention.

[0083] As illustrated in Figure 1, the classification and localization device APP is configured to receive as input a sequence of images IMG_SEQ and produce as output a class ATR_CLAS_OBJ assigned to an object OBJ as well as an estimated position EST_POS_OBJ of the object OBJ. The sequence of images IMG_SEQ typically comprises a plurality of images IMG_1, ..., IMG_N. The images IMG_1, ..., IMG_N are acquired by a camera CAM and are composed of a plurality of pixels. For example, the pixels of an image IMG_1 are encoded on one or more bits: one bit per pixel in the case of a monochrome image, 8 bits per pixel to have 256 colors, etc. The camera CAM is connected to the APP device such that the camera CAM can transmit the sequence of images IMG_SEQ to the APP device. The images IMG_1, ..., IMG_N of the sequence IMG_SEQ are representative of an optical scene comprising the material object OBJ.For example, the OBJ object can be a vehicle, an aircraft, a pedestrian, a house, etc. It should be noted that the OBJ object can be moving or not. The assigned class ATR_CLAS_OBJ to the OBJ object is selected from a finite list of assignable classes, such as a list of aircraft models, vehicle types, object categories, etc.

[0084] The APP device produces from the sequence IMG_SEQ an estimate EST_POS_OBJ of the position POSJDBJ of the object OBJ. The position POSJDBJ of the object OBJ is either an absolute position or a relative position. If the position POSJDBJ is absolute, then it designates a position defined relative to a terrestrial reference frame. It should be noted that, according to one embodiment, the APP device determines an absolute position of the object OBJ from the sequence IMG_SEQ and the position (eg geographic coordinates) of the APP device. According to an alternative embodiment, the absolute position POSJDBJ may correspond to geographic coordinates of the material object OBJ such that the position POSJDBJ of the object OBJ comprises one or more coordinates from the following set: latitude, longitude; and altitude.

[0085] According to one embodiment, the position POSJDBJ of the object is a relative position defined with respect to an observer position POSJDBS. The position POSJDBJ comprises, in one embodiment, one or more coordinates from the following set: an azimuth, a height, and a distance defined with respect to the observer position POSJDBS. According to an alternative embodiment, and as illustrated by FIG. 1, the estimated position EST_POS_OBJ is an estimate of the distance DISTJDBJ between the observer position POSJDBS and the object OBJ. Typically, the distance DISTJDBJ designates the distance between the material object OBJ and the camera CAM carrying out the acquisitions of the image sequence IMG_SEQ.

[0086] According to one embodiment, the position POSJDBJ of the object OBJ is a relative position defined with respect to a previous position of the object OBJ. In this embodiment, the estimated position EST_POS_OBJ by the APP device thus corresponds to a displacement of the material object OBJ, and comprises for example one or more coordinates from among: a longitudinal displacement, a lateral displacement, and a vertical displacement. According to an alternative embodiment, the APP device produces an estimate EST_POS_OBJ relating to the displacement of the object OBJ between two images IMG_1, IMG_2 of the sequence IMG_SEQ, and / or produces an estimate of the speed of the object OBJ.

[0087] The CAM camera is an IMG_SEQ image acquisition device, making it possible to transform optical images into digital images. To do this, the CAM camera comprises an electromagnetic radiation sensor, radiation whose wavelengths belong to the visible light spectrum. According to one embodiment of the invention, the CAM camera comprises an electromagnetic radiation sensor whose wavelengths are located beyond the visible light spectrum, such as an infrared sensor. The CAM camera can, according to this embodiment, acquire IMG_SEQ images in the infrared domain.

[0088] Obviously, no limitation is attached to the nature of the communication interface between the APP device and the CAM camera, which can be wired or wireless, and can implement any protocol known to those skilled in the art (Internet, IP, Ethernet, WiFi, Bluetooth, 3G, 4G, 5G, 6G, etc.). Furthermore, no limitation is attached to the format of the images IMG_1, ..., IMG_N of the sequence IMG_SEQ, which can implement any encoding known to those skilled in the art (JPG, PNG, TIFF, etc.).

[0089] Figure 2 schematically represents an example of a sequence of images obtained by a device for classifying and locating at least one object in a sequence of images according to one embodiment of the invention.

[0090] As illustrated in Figure 2, the image sequence IMG_SEQ comprises one or more images IMG_1, IMG_N. Typically, the image sequence IMG_SEQ comprises a plurality of images IMG_1, IMG_N.

[0091] According to one embodiment, the images IMG_1, ..., IMG_N of the sequence IMG_SEQ are representative of an optical scene comprising a moving object OBJ. According to one embodiment, the images IMG_1, ..., IMG_N of the sequence IMG_SEQ are consecutive in time, that is to say that the respective acquisition times of the images IMG_1, ..., IMG_N of the sequence IMG_SEQ follow one another in time. The first image IMG_1 thus corresponds to the oldest image, while the last image IMG_N corresponds to the most recent image. According to one embodiment, the classification and localization device APP performs a cropping operation of the images IMG_1, ..., IMG_N of the sequence IMG_SEQ such that the object OBJ is centered in the first image IMG_1. The cropping coordinates remain fixed for all images IMG_1, ..., IMG_N of the sequence IMG_SEQ.As shown in Figure 2, the cropping operation makes it possible to visually represent and highlight information on the movement of the object OBJ within the image sequence IMG_SEQ and over time.

[0092] According to one embodiment of the invention, the image sequence IMG_SEQ comprises a plurality of images IMG_1, ..., IMG_N acquired by a plurality of cameras CAM. According to one embodiment, the image sequence SEQ_IMG is representative of an optical scene comprising a plurality of material objects OBJ. Furthermore, according to one embodiment, the image sequence SEQ_IMG is representative of a plurality of optical scenes.

[0093] According to the embodiment described above, the APP device uses IMG_SEQ images acquired by a plurality of CAM cameras. It should then be mentioned that, according to this embodiment, the fields of vision of the different CAM cameras are distinct. In this way, an object OBJ of interest present in the field of vision of a first camera will then be observed by a second camera when the object OBJ leaves the field of vision of the first camera. Furthermore, according to this embodiment, the different CAM cameras have identical optics. This embodiment thus makes it possible to broaden the field of vision of the proposed APP device and, thus, makes it possible to classify and locate objects OBJ in a wider geographical area.

[0094] Furthermore, according to a variant of the invention, the APP device implements a classification and a localization of several objects OBJ from a sequence IMG_SEQ, at at least one assigned class ATR_CLAS_OBJ to the objects OBJ and at least one estimated position EST_POS_OBJ of the objects OBJ being provided as output by the APP device.

[0095] Figure 3 schematically represents an example of functional architecture of a device for classifying and locating at least one object in a sequence of images according to one embodiment of the invention.

[0096] As illustrated in Figure 3, the APP device receives as input a sequence IMG_SEQ of images IMG_1, IMG_N. The APP device determines, from the sequence IMG_SEQ, a class ATR_CLAS_OBJ assigned to the object and an estimated position EST_POS_OBJ of the object OBJ. The outputs ATR_CLAS_OBJ, EST_POS_OBJ of the APP device are determined by a classifier-localizer module X_NN based on the sequence IMG_SEQ. How to configure the classifier-localizer module X_NN (i.e. its parameters) to obtain reliable classification and precise localization is described in more detail below and illustrated in Figure 5.

[0097] According to one embodiment, the determination step performed by the classifier-localizer module X_NN comprises the following operations. The image sequence IMG_SEQ is provided as input to a CNN encoder for producing encoded data LS_DATA as output. The encoded data LS_DATA are on the one hand provided as input to a classifier CLA_NN and on the other hand to a locator LOC_NN. The classifier CLA_NN produces as output a class assigned ATR_CLA_OBJ to the object OBJ, while the locator LOC_NN produces as output an estimated position EST_POS_OBJ. The encoded data LS_DATA thus constitutes a multidimensional intermediate representation of the sequence IMG_SEQ common to the classification and localization tasks.

[0098] According to one embodiment, the X_NN classifier-localizer module comprises a parameterized function. More particularly, the X_NN classifier-localizer module comprises, according to one embodiment, a machine learning algorithm, such as a support vector machine, a Bayesian network, etc.

[0099] According to an alternative embodiment of the invention, the classifier-localizer module X_NN comprises (i.e. is implemented by) a neural network. Generally, a neural network comprises one or more layers of artificial neurons connected to each other. An artificial neuron performs the following operations. A neuron takes one or more values ​​as input and weights these inputs by coefficients called weights. The neuron combines the weighted inputs as well as a bias, typically this combination operation is a sum, or a norm. This combination is then provided as input to an activation function. Examples of commonly used activation functions are the sigmoid, ReLU (acronym for the English expression "Rectified Linear Unit"), hyperbolic tangent functions. Finally, the output of the artificial neuron is the result of the function activation. In the context of a neural network, the parameters of a neural network include a varied set of parameters, such as weights, biases, and activation functions used for the different neurons of the X_NN network. These parameters thus make it possible to optimize the X_NN neural network to implement reliable classification and precise localization. The types of neural networks capable of implementing an X_NN classifier-localizer module according to the invention are varied. In particular, mention should be made, in a non-limiting manner, of multilayer perceptrons, convolutional or convolutional neural networks, recurrent neural networks, etc.

[0100] As illustrated in Figure 3, and according to one embodiment of the invention, the neural network X_NN comprises: a convolutional neural network CNN; a classification neural network CLA_NN; and a localization neural network LOC_NN. In this particular embodiment, the encoder comprises the convolutional neural network CNN as described above for determining encoded data LS_DATA from the sequence IMG_SEQ. The classifier comprises the classification neural network CLA_NN and produces from the encoded data LS_DATA an assigned class ATR_CLAS_OBJ. The locator comprises the localization neural network LOC_NN and makes it possible to estimate the position EST_POS_OBJ from the encoded data LS_DATA. A more detailed example of implementation of the classifier-locator module X_NN by a neural network is provided below and illustrated in Figure 6.

[0101] It is important to note here that, in this embodiment, the neural network X_NN is determined globally to minimize the multi-objective loss function F_LOSS representative of both a classification objective and a localization objective. The neural networks CNN, CLA_NN and LOC_NN are thus jointly determined to minimize the loss function F_LOSS, and not independently. In one embodiment, it can be considered that the classification and localization tasks implemented by the module X_NN are learned in parallel.

[0102] Figure 4 represents, in the form of a flowchart, steps of a method for classifying and locating at least one object in a sequence of images according to one embodiment of the invention.

[0103] As illustrated in Figure 4, and according to one embodiment of the invention, the classification and localization method comprises the following steps and is implemented by an APP device. In the present description, the reference signs of the steps related to the classification and localization method begin with S2.

[0104] In step S210, a sequence IMG_SEQ of one or more images IMG_1, ..., IMG_N acquired by at least one camera CAM is obtained by the APP device.

[0105] In step S220, the APP device provides the image sequence IMG_SEQ obtained as input to a classifier-localizer module X_NN which thus determines: • at least one class assigned ATR_CLAS_OBJ to said at least one object OBJ, said at least one class assigned ATR_CLAS_OBJ being selected from a list of assignable classes; and • at least one estimated position EST_POS_OBJ of said at least one object OBJ.

[0106] The X_NN classifier-localizer module is configured from reference image sequences to minimize a multi-objective loss function representative of both a classification objective and an object localization objective in the reference sequences.

[0107] As illustrated by Figure 4, and according to a particular embodiment, step S210 of obtaining the method described above comprises at least one of the following sub-steps.

[0108] In substep S211, the APP device performs an acquisition by at least one camera of said one or more images IMG_1, ..., IMG_N of the sequence IMG_SEQ. For example, the APP device controls a CAM camera to acquire the images IMG_1, ..., IMG_N of the sequence IMG_SEQ. In particular, the images IMG_1, ..., IMG_N of the sequence IMG_SEQ can be acquired by a CAM camera with a fixed time step between the different acquisition times.

[0109] In substep S212, the images IMG_1, ..., IMG_N are cropped by the device APP to center an object OBJ in a determined image IMG_1 of the sequence IMG_SEQ. According to one embodiment, and as illustrated by FIG. 2, the object OBJ is centered on the first image IMG_1 of the sequence IMG_SEQ. In this embodiment, the images IMG_1, ..., IMG_N of the sequence IMG_SEQ are consecutive in time. The respective acquisition times of the images IMG_1, ..., IMG_N follow one another in time. Furthermore, the coordinates of the cropping remain fixed for all the images IMG_1, ..., IMG_N of the sequence IMG_SEQ, which makes it possible to represent the displacement of the object OBJ during the sequence. The cropping operation makes it possible to reduce the size of the processed images and thus to reduce the complexity of implementing the method, while retaining the information linked to the movement of the object OBJ within the image sequence IMG_SEQ.

[0110] As illustrated by Figure 4, and according to a particular embodiment, step S220 of determining the method described above comprises the following sub-steps.

[0111] In step S221, the APP device determines encoded data LS_DATA. The encoded data LS_DATA is produced by the CNN encoder taking the image sequence IMG_SEQ as input.

[0112] In step S222, the APP device provides as input to a CLA_NN classifier the encoded data LS_DATA which produces as output said at least one ATR_CLAS_OBJ class.

[0113] In step S223, said at least one estimated position EST_POS_OBJ is produced by a locator LOC_NN of the APP device taking as input the encoded data LS_DATA.

[0114] It should be noted that the order of steps S222 and S223 described here is in no way limiting. Steps S222 and S223 can be performed simultaneously (i.e. in parallel), one before the other or vice versa. The encoded data LS_DATA makes it possible to obtain a multidimensional intermediate representation common to the two classification and localization tasks and representative of spatio-temporal characteristics of the object OBJ in the image sequence IMG_SEQ.

[0115] Figure 5 represents, in the form of a flowchart, steps of a method for configuring a classifier-locator module to implement a classification and localization of at least one object in a sequence of images according to an embodiment of the invention.

[0116] As illustrated in Figure 5, and according to one embodiment of the invention, the proposed configuration method comprises the following steps. In the present description, the reference signs of the steps related to the configuration method begin with SI.

[0117] In step S120, the classifier-localizer module X_NN for implementing the classification and localization of objects OBJ in image sequences IMG_SEQ is initialized (e.g. its parameters). In the embodiment where the classifier-localizer module X_NN comprises a neural network, the weights of the neural network are, for example, initialized randomly.

[0118] In step S130, a sequence of reference images TR_DATA is provided as input to the classifier-locator module X_NN to determine: a class assigned ATR_CLAS_OBJ to the object OBJ from a list of assignable classes; and at least one estimated position EST_POS_OBJ of the object OBJ. According to an alternative embodiment, step S130 comprises substeps analogous to steps S221, S222 and S223 described above. According to a particular embodiment, the classifier-locator module X_NN produces, for a sequence of reference images TR_DATA as input, a set of values ​​PP_CLAS_1, ..., PP_CLAS_M and a set of estimated positions EST_POS_OBJ_1, ..., EST_POS_OBJ_N.For each image IMG_1, the X_NN classifier-localizer module provides an estimated position EST_POS_OBJ_1 of the object OBJ on the image IMG_1; and for each class in the list of attributable classes, the X_NN classifier-localizer module provides as output a PP_CLAS_1 value representative of a probability of belonging of the object OBJ to this class.

[0119] According to one embodiment of the invention, several reference image sequences TR_DATA are used during steps S130 and S140. Furthermore, a reference image sequence TR_DATA may comprise several material objects OBJ.

[0120] In step S140, a loss function F_LOSS is evaluated from the following inputs: said at least one assigned class ATR_CLAS_OBJ; said at least one estimated position EST_POS_OBJ; at least one known class of an object in the reference sequence TR_DATA; and at least one known position of an object in the reference sequence TR_DATA. The multi-objective loss function F_LOSS is representative of both a classification objective and a localization objective. In this embodiment, it should be noted that reference data TR_DATA is used to configure the classifier-localizer module such that the loss function F_LOSS is minimized.

[0121] In step S150, the classifier-localizer module X_NN is reconfigured (e.g. its parameters) to minimize the loss function F_LOSS. The result of the evaluation carried out in step S140 is, according to a variant of the invention, used to reconfigure the classifier-localizer module X_NN. According to one embodiment, this step consists of determining the parameter value of the classifier-localizer module X_NN to minimize the loss function F_LOSS.

[0122] According to one embodiment, the determination method comprises several iterations of steps S130, S140, and S150, which makes it possible to minimize the loss function F_LOSS as the iterations progress, and thus to improve the performance of the classifier-localizer module X_NN. Typically, the greater the number of iterations, the more the loss function F_LOSS will be minimized. For example, the number of iterations performed by the method can be determined in the following manner. If during an iteration, the evaluation of the loss function in step S140 produces a result below a certain threshold, then no additional iterations will be performed. In a different manner, the method for configuring the classifier-localizer module X_NN can also perform said iterations until the variation of the loss function between two iterations is below a threshold, i.e.the difference between two consecutive results of the loss function is less than the threshold.

[0123] As illustrated by Figure 5, and according to a particular embodiment, the proposed configuration method described above comprises the following step.

[0124] In step S1 10, reference data TR_DATA are obtained. More particularly, the reference data TR_DATA comprises: one or more reference image sequences; one or more known classes associated with objects OBJ of the reference sequences; one or more known positions associated with objects OBJ of the reference sequences.

[0125] According to a particular embodiment, step SI 10 of obtaining reference data comprises at least one of the following sub-steps, as shown in FIG. 5.

[0126] In sub-step SI 11, one or more reference image sequences TR_DATA_ACQ are acquired by at least one camera CAM. Step SI 11 further comprises, according to a variant, the determination of one or more known classes and positions associated with one or more objects OBJ of the acquired reference sequences TR_DATA_ACQ.

[0127] In substep S1 12, and according to a variant of the invention, one or more computer-synthesized TR_DATA_SYN image sequences are obtained. For example, the synthesized TR_DATA_SYN reference sequences may be the result of simulations. In this embodiment, the simulation tool may produce, in addition to the synthesized TR_DAT_SYN reference image sequences, the known classes and the known positions associated with the objects of these reference sequences.

[0128] As illustrated by FIG. 5, and according to a particular embodiment, step S140 of evaluating the multi-objective loss function F_LOSS of the method described above comprises the following sub-steps.

[0129] In step S141, a classification loss function F_LOSS_CLA is evaluated by taking as inputs said at least one assigned class ATR_CLAS_OBJ and at least one known class TR_DATA associated with said at least one object of the reference image sequence TR_DATA. This step makes it possible, among other things, to evaluate the reliability of the classifications carried out by the classifier-localizer module X_NN.

[0130] In step S142, a localization loss function F_LOSS_LOC is evaluated by taking as inputs said at least one estimated position EST_POS_OBJ and at least one known position TR_DATA associated with said at least one object of the reference image sequence TR_DATA. Step S142 thus makes it possible to evaluate the accuracy of the localizations carried out by the classifier-localizer module X_NN.

[0131] The order of sub-steps S141 and S142 described here is in no way limiting, these two steps being able to be carried out simultaneously, one after the other or vice versa.

[0132] In substep S143, the multi-objective loss function F_LOSS is evaluated from the result of the classification loss function F_LOSS_CLA obtained in substep S141 and the result of the localization loss function F_LOSS_LOC obtained in substep S142.

[0133] According to a variant of the invention, the classification loss function F_LOSS_CLA is a cross-entropy function. Let C be denoted r the known class TR_DATA associated with the object OBJ and p o the set of values ​​PP_CLAS_I, ..., PP_CLAS_M produced by the X_NN classifier-localizer module and representative of probabilities of the OBJ object's membership in each of the classes in the list. The classification loss function F_LOSS_CLA, noted here L EC , is then defined by: [Math 1]

[0134] where δ( C i , C r ) is a binary indicator equal to 0 if the class C i is different from class C r and equal to 1 otherwise, and P o,i is the PP_CLAS_I value representing a probability of belonging of the object OBJ to the class of index i in the list.

[0135] According to one embodiment, the result of the localization loss function F_LOSS_LOC is determined from the error between the estimated positions EST_POS_OBJ_1, ..., EST_POS_OBJ_N of the object OBJ and the known positions of the reference sequence TR_DATA. In particular, the localization error is taken into account only for the images in which the object OBJ is present. Some of the images in the sequence may not include the object, the latter being out of the frame. In this embodiment, the localization loss function F_LOSS_LOC is defined as follows. Let P be denoted by r the set of known positions of the object OBJ for the reference sequence TR_DATA, and P ois the set of estimated positions EST_POS_1, ..., EST_POS_OBJ_N of the object OBJ by the classifier-localizer module X_NN for each of the images IMG_1, ..., IMG_N of the reference sequence TR_DATA. To indicate the presence of the object OBJ in the reference sequence, the presence vector I = (I1, ... ,I N ) is used. If the image with index i includes the object OBJ, then the value of the component I i is equal to 1; and if this image does not include the object OBJ (eg the one outside the frame), then the component l i is equal to 0. The localization loss function F_LOSS_LOC, denoted here L IMSE , is then defined by: [Math 2]

[0136] is the L1 norm of the vector I, is a binary indicator equal to 1 if the object is present in the image of index i in the sequence and equal to 0 otherwise, P o,iis the position estimated by the classifier-localizer module X_NN for the image of index i in the list, and P r,i is the known position for the image of index i in the sequence. It can be noted that in this embodiment the localization loss function F_LOSS_LOC is expressed from quadratic errors. Thus, the proposed localization loss function F_LOSS_LOC is an extension of a mean squared error type loss function.

[0137] The expression of the localization loss function F_LOSS_LOC described above is defined for positions of the object OBJ of one dimension, for example a distance between the object OBJ and the camera CAM carrying out the acquisition of the images IMG_1, IMG_N. However, this expression is only an implementation variant. Such an expression of the localization loss function F_LOSS_LOC can easily be extended for positions comprising several coordinates, in particular by using the quadratic error between the estimated position and the known position on each component associated with the dimensions.

[0138] According to one embodiment, the estimated position of an object comprises a distance between an observer position and the object. According to this embodiment, the classifier-localizer module is configured to minimize the multi-objective loss function representative of the classification objective and the objective of locating objects in the reference sequence, the objective of locating an object comprising an objective of determining a distance between an observer position and said object.

[0139] The multi-objective loss function F_LOSS to be minimized is, according to an alternative embodiment, a weighted sum of the classification loss function F_LOSS_CLA and the localization loss function F_LOSS_LOC. The weighting coefficients of this sum are non-zero, a necessary condition for the multi-objective loss function F_LOSS to be representative of both a classification objective and a localization objective. In particular, the loss function F_LOSS, denoted here L, is expressed by: [Math 3]

[0140] where α and β are the positive non-zero weighting coefficients, For example, the value of α is equal to 1 and the value of β is equal to 10 -3 .

[0141] As illustrated by Figure 5, and according to a particular embodiment, step S150 of updating the method described above comprises the following sub-step.

[0142] In substep S151, the classifier-localizer module X_NN is using a gradient descent algorithm based on the result of the loss function F_LOSS obtained in step S140. In the embodiment where the classifier-localizer module comprises a neural network, gradient backpropagation is used to update the parameters of the network, and in particular, to determine the weights thereof.

[0143] . According to one embodiment, a gradient descent algorithm is used to configure the classifier-localizer module X_NN. In this embodiment, the reconfiguration step consists of updating the parameters of the classifier-localizer module X_NN. To do this, the gradient of the loss function F_LOSS is evaluated. Then, the parameters of the module X_NN are updated using the evaluated gradient. In particular, in In the case of a gradient descent algorithm, the parameters are updated by subtracting the value of the evaluated gradient multiplied by a positive real coefficient. Thus, a gradient descent algorithm aims to minimize the loss function.

[0144] According to one embodiment, the classifier-localizer module X_NN comprises a neural network and a gradient back-propagation method is used to update the neural network. In this embodiment, the reconfiguration step S151 of the X_NN network consists of determining the values ​​of the weights used by the artificial neurons of the X_NN network. The gradient back-propagation method uses the result of the loss function F_LOSS to update the weights of the X_NN network. In particular, the gradient back-propagation method consists of: propagating the result of the loss function through the different layers of the neural network, from the output layer to the input layer; and updating the weights of each layer from said propagated results.

[0145] In the previously described embodiment where the classifier-localizer module X_NN comprises a CNN encoder, a CLA_NN classifier and a LOC_NN locator, it is important to note that the CNN, CLA_NN and LOC_NN elements of the X_NN module are jointly determined. Indeed, the reconfiguration of the classifier-localizer module X_NN is carried out from the result of the multi-objective loss function F_LOSS, and not from the results of the classification loss function F_LOSS_CLA and the localization loss function F_LOSS_LOC. In other words, the CLA_NN classifier and the LOC_NN locator are configured in parallel (i.e. jointly) in order to optimize a common objective of reliable classification and precise localization, and not each independently to optimize their respective objective.

[0146] Figure 6 schematically represents an example of functional architecture of a device for classifying and locating at least one object in a sequence of images according to one embodiment of the invention.

[0147] As illustrated by Figure 6 and previously mentioned, the classification and localization device APP comprises a classifier-localizer module X_NN for determining, from a sequence of images IMG_SEQ, at least one class assigned ATR_CLAS_OBJ to said at least one object OBJ of the sequence IMG_SEQ and at least one estimated position EST_POS_OBJ of said at least one object OBJ.

[0148] In the particular embodiment illustrated by Figure 6, the classifier-localizer module X_NN produces as output an estimated position EST_POS_OBJ_1, ..., EST_POS_OBJ_N of the object OBJ for each of the images IMG_1, ..., IMG_N of the sequence IMG_SEQ. For example, the estimated position EST_POS_OBJ of the object OBJ by the classifier-localizer module X_NN is the average of these estimated positions EST_POS_OBJ_1, ..., EST_POS_OBJ_N, or the closest position, or the furthest position. In addition, the X_NN classifier-locator module produces, in addition to the assigned class ATR_CLAS_OBJ to the object OBJ, a set of values ​​PP_CLAS_1, ..., PP_CLAS_M. Each of the values ​​PP_CLAS_1, ..., PP_CLAS_M is representative of a probability of belonging of the object OBJ to a class in the list of assignable classes. In other words, for each class in the list, the X_NN classifier-locator module provides a PP_CLAS_1 value characterizing the probability that the object OBJ belongs to this class. For illustrative and non-limiting purposes, the X_NN classifier-locator module assigns to the object OBJ of the sequence IMG_SEQ the class whose PP_CLAS_1 value is the highest.

[0149] According to an embodiment of the invention, described above and illustrated by Figure 3, the classifier-locator module X_NN comprises: a CNN encoder providing encoded data LS_DATA as output from the sequence IMG_SEQ; a CLA_NN classifier whose output is at least one class assigned ATR_CLA_OBJ to the object OBJ and is determined from the encoded data LS_DATA; and a locator LOC_NN for determining an estimated position EST_POS_OBJ of the object on the basis of the encoded data LS_DATA. Figure 6 illustrates a particular embodiment of the invention in which the CNN encoder, the CLA_NN classifier and the locator LOC_NN are respectively implemented by a neural network.

[0150] In the embodiment described here, the CNN encoder is implemented using a convolutional neural network. In particular, the convolutional neural network is structured as follows. The CNN convolutional neural network, as shown in Figure 6, comprises: one or more convolution layers CNN_CONV_L1, CNN_CONV_L2; one or more subsampling layers CNN_POOL_L1, CNN_POOL_L2, so-called “pooling” layers; and at least one fully connected layer CNN_FC_L, the nodes of this layer all being connected to the nodes of the following layer. Of course, the CNN network comprises, according to other variants of the invention, other layers of neurons, no limitation being attached to the nature of the latter.

[0151] The function of a convolution layer is to perform convolution operations between the input data of the layer and filters, filters commonly called kernels (or with the Anglo-Saxon expression "kernels"). In this description, and for simplification, the term "convolution" is used to refer to a convolution product, the mathematical operation generally noted *. In the case of two-dimensional discrete data, the convolution of an image I and a filter F is defined by Let us take the example of an image I provided as input and a convolution layer using two filters F1 and F2, the convolutions I * F1 and I * F2 are evaluated to obtain two intermediate images. These intermediate images produced at the output of a convolution layer are commonly referred to as "feature maps". Thus, the parameters of a convolution layer are, among others the following: the number and size of filters, the control step size and the margin used by the convolution layer (these last two parameters are more commonly referred to as "stride" and "padding"). In the context of a convolution or pooling operation, the control step size and the margin refer respectively to the number of pixels by which the filter moves at each shift and to a technique consisting of adding pixels to the edge of the input image (e.g. zero-padding). Thus, by performing convolutions, a convolution layer makes it possible to extract spatio-temporal characteristics from input images.In this case, for the purpose of classifying and locating an object OBJ in a sequence of images IMG_SEQ, the convolution layers make it possible to highlight in the sequence of images IMG_SEQ information on the object OBJ such as information on size, shape, speed, movement, etc.

[0152] According to a particular embodiment of the invention, the convolution layers CNN_CONV_L1, CNN_CONV_L2 of the CNN encoding neural network perform convolutions between the sequence IMG_SEQ and filters in the spatial domain and in the temporal domain. Performing convolutions according to the three dimensions of the sequence of images IMG_SEQ, two dimensions for the spatial domain of the images and one dimension for the temporal domain of the sequence, makes it possible to extract spatio-temporal characteristics from the sequence of images IMG_SEQ. More particularly, according to a variant of the invention, the convolution layers perform convolutions called "3D pseudo-convolutions", the details of implementation of the 3D pseudo-convolutions are for example explained in the following document: Zhaofan Qiu, et al., “Learning spatio-temporal representation with pseudo-3d residual networks”, in proceedings of the IEEE International Conference on Computer Vision, pages 5533-5541, 2017. Exploiting 3D pseudo-convolutions allows extracting spatio-temporal characteristics from the IMG_SEQ image sequence, while reducing the complexity of the network, compared to classic three-dimensional convolutions.

[0153] A pooling layer, also called a pooling layer, allows for downsampling of the data provided as input to the layer. For example, taking an image, the input image is partitioned into a plurality of pixel rectangles, these rectangles being generally called tiles, and an output value is produced per tile. With tiles of size 2 x 2 pixels, the pooling layer allows for compression by a factor of 4 of the input data. For illustration purposes, the output value for a tile is the maximum value of the tile's data; such a pooling layer is commonly called "Max-Pool 2x2". In a different example, the output value associated with a tile is the minimum value of the tile's input data; in this case, the expression "Min-Pool 2x2" is used. Thus, a pooling layer allows for reducing the size of the data processed by the next layer of the neural network and thus reduce the complexity of the latter.

[0154] According to one embodiment, the convolutional network CNN comprises one or more successions of a convolution layer and a pooling layer. According to the particular embodiment illustrated by FIG. 6, the CNN network comprises two said successions and a fully connected layer, i.e. CNN: CNN_CONV_L1 > CNN_POOL_L1 > CNN_CONV_L2 > CNN_POOL_L2 > CNN_FC_L. According to this variant, the images IMG_1, ..., IMG_N of the sequence IMG_SEQ are provided as input to the first convolution layer CNN_CONV_L1, and the encoded data LS_DATA are produced as output from the fully connected layer CNN_FC_L. According to one embodiment variant, the CNN network also comprises layers of neurons, called correction layers, interposed between the aforementioned convolution and pooling layers.These correction layers apply an activation function to all pixels of the intermediate images, which allows the introduction of non-linear complexities and thus improves the processing performed by the CNN network. Compared to a fully connected multi-layer perceptron neural network, a convolutional neural network allows an efficient implementation of feature extraction from the input data. Indeed, the successions of convolution and pooling layers present minimal complexity and connectivity, and thus a simplified practical implementation with better performance. In addition, the fully connected multi-layer perceptrons used to implement such feature extractions are prone to overtraining problems.

[0155] According to one embodiment, the convolutional neural network CNN is implemented using a neural network of the ResNet type or one of its variants. The implementation of such a network is for example detailed in the above-mentioned document by Zhaofan Qiu, et al.

[0156] According to one embodiment of the invention, the localization neural network LOC_NN comprises one or more layers of neurons. For example, the network LOC_NN is a multi-layer perceptron. As illustrated in Figure 6, and according to a variant of the invention, the localization neural network LOC_NN is a perceptron comprising a fully connected layer LOC_FC_L whose inputs are the encoded data LS_DATA and whose outputs are a set of estimated positions EST_POS_OBJ_1, ..., EST_POS_OBJ_N. Each of the estimated positions EST_POS_OBJ_1 corresponds to an estimate of the position POSJDBJ of the object OBJ for an image IMG_1 of the sequence IMG_SEQ. According to one embodiment, an aggregated estimated position EST_POS_OBJ of the object OBJ is evaluated on the basis of the estimated positions EST_POS_OBJ_1, ..., EST_POS_OBJ_N, for example by taking the mean, the median, the minimum or the maximum of the latter.As an example, consider an embodiment where the localization neural network LOC_NN produces. a position estimate for each of the images in the sequence IMG_SEQ, and where an estimated position EST_POS_OBJ_1 is an estimate of a distance DISTJDBJ between an object OBJ and a camera CAM. In this embodiment, and for a sequence IMG_SEQ of N images, the localization neural network LOC_NN comprises an output layer with N neurons, each of which produces an estimate value EST_POS_OBJ_1 of the distance DISTJDBJ associated with an image IMG_1.

[0157] According to one embodiment of the invention, the classification neural network CLA_NN is a “classifier” type neural network. The CLA_NN network comprises one or more layers of neurons. For example, the CLA_NN network is a multi-layer perceptron. As illustrated in Figure 6, and according to a variant of the invention, the classification neural network CLA_NN is a perceptron comprising a fully connected layer CLA_FC_L whose inputs are the encoded data LS_DATA and whose outputs are a set of values ​​PP_CLAS_1, ..., PP_CLAS_M. These values ​​are representative of a probability of belonging of the object OBJ to the classes of the list of assignable classes. According to one embodiment, the class assigned ATR_CLAS_OBJ to the object is the class of the list whose PP_CLAS_1 value is the highest.According to an alternative embodiment, the CLA_NN network comprises a “Softmax” type layer for determining from the outputs of the CLA_FC_L layer the set of values ​​PP_CLAS_1, ..., PP_CLAS_M; the use of a Softmax layer makes it possible to normalize the output values, so that they are between 0 and 1 and thus representative of probabilities. More particularly, for inputs x = (x1, ..., x. K ), the outputs of the Softmax layer y = (y1...,y K ) are expressed by:

[0158] As an example, consider an embodiment where the classification neural network CLA_NN produces the set of values ​​PP_CLAS_1, ..., PP_CLAS_M. In this embodiment, and for a list of M assignable classes, the CLA_NN network comprises an output layer with M neurons, each of which produces a value PP_CLAS_1 representative of the probability of the object OBJ belonging to a class. In combination with the example described above of a network LOC_NN with N neurons as outputs, the neural network X_NN then comprises an output layer with M+N neurons.

[0159] Figure 7 schematically represents an example of software and hardware architecture of a system for classifying and locating at least one object in a sequence of images according to one embodiment of the invention.

[0160] As illustrated in Figure 7, the SYS system comprises an APP device and a CAM camera. The APP classification and localization device comprises in particular: a processing unit or processor PROC; and a memory MEM. Of course, the APP device comprises interfaces and a communication module for exchanging data with the CAM camera. The APP device has the hardware architecture of a computer, and as such includes a processor PROC, RAM, ROM MEM, and non-volatile memory.

[0161] In the embodiment described here, the memory MEM associated with the device constitutes an information or recording medium in accordance with the invention, readable by computer and by the processor PROC and on which a computer program in accordance with the invention is recorded. The computer program comprises instructions for implementing the steps of a method according to the invention, when the computer program is executed by the processor PROC. The computer program defines the functional modules represented by FIG. 8 of the APP device, which rely on or control the hardware elements of the latter.

[0162] For information purposes, the SYS system is embedded in a vehicle, for example in a land vehicle: car, truck, train, etc., or in a marine vehicle: boat, frigate, or even in an air vehicle: an aircraft, a helicopter, an airplane, a drone, etc. In particular, the SYS system is, according to an alternative embodiment, embedded in a so-called autonomous vehicle, such as an autonomous car or a drone. According to one embodiment of the invention, the SYS system constitutes a surveillance system or a navigation system.

[0163] Figure 8 schematically represents an example of functional architecture of a device for classifying and locating at least one object in a sequence of images according to one embodiment of the invention.

[0164] As illustrated in Figure 8, and according to one embodiment, the system SYS comprises a device APP for classifying and locating at least one object OBJ in a sequence of images IMG_SEQ and at least one camera CAM. Said at least one camera CAM is configured to acquire said one or more images IMG_1, ..., IMG_N of the sequence IMG_SEQ. In particular, the system SYS comprises, according to one embodiment, a single camera CAM. The device APP comprises the modules described below.

[0165] The term module can correspond to a software component as well as to a hardware component or a set of hardware and software components, a software component itself corresponding to one or more computer programs or subroutines or more generally to any element of a program capable of implementing a function or a set of functions as described for the modules concerned. In the same way, a hardware component corresponds to any element of a hardware assembly capable of implementing a function or a set of functions for the module concerned (integrated circuit, smart card, memory card, etc.).

[0166] As illustrated by Figure 8, according to a particular embodiment of the invention, the APP device for classifying and locating at least one object OBJ in a sequence of images IMG_SEQ comprises: • an obtaining module MODJDBT for obtaining a sequence IMG_SEQ of one or more images IMG_1, IMG_N acquired by at least one CAM camera; and • a classifier-locator module X_NN for determining from the image sequence IMG_SEQ: o at least one class assigned ATR_CLAS_OBJ to said at least one object OBJ, said at least one class assigned ATR_CLAS_OBJ being selected from a list of classes; and o at least one estimated position EST_POS_OBJ of said at least one object OBJ; the module being configured from a reference image sequence to minimize a multi-objective loss function F_LOSS representative of both a classification objective and an object localization objective in the reference image sequences TR_DATA.

[0167] As illustrated by Figure 8, and according to a particular embodiment of the invention, the classifier-locator module X_NN comprises: • a CNN encoder module to determine encoded data LS_DATA from the image sequence IMG_SEQ; • a CLA_NN classifier module for determining said at least one assigned class ATR_CLAS_OBJ from the encoded data LS_DATA; and • a locator module LOC_NN to determine said at least one estimated position EST_POS_OBJ from the encoded data LS_DATA.

[0168] As illustrated by Figure 8, and according to a particular embodiment of the invention, the MODJDBT obtaining module comprises: • an acquisition module MOD_ACQ to control the camera CAM and acquire said one or more images IMG_1, ..., IMG_N of the sequence IMG_SEQ; • a cropping module MOD_CRP for cropping said one or more images IMG_1, ..., IMG_N of the sequence IMG_SEQ to center the object OBJ in a determined image IMG_1 of the sequence IMG_SEQ. Typically, the object OBJ is centered on the first image IMG_1 of the cropped sequence IMG_SEQ.

[0169] It should be noted that the order in which the steps of a process as described above are followed, in particular with reference to the attached drawings, constitutes only an example of realization without any limiting character, variants being possible. Furthermore, the reference signs are not limiting of the scope of the protection, their sole function being to facilitate the understanding of the claims.

[0170] A person skilled in the art will understand that the embodiments and variants described above constitute only non-limiting examples of implementation of the invention. In particular, a person skilled in the art may envisage any adaptation or combination of the embodiments and variants described above in order to meet a very specific need.

Claims

Claims

1. Method (S200) for classifying and locating at least one object (OBJ) in an image sequence (IMG_SEQ), said method comprising steps of: • obtaining (S210) a sequence (IMG_SEQ) of one or more images (IMG_1, IMG_N) acquired by at least one camera (CAM); and • determination (S220) by a classifier-localizer module (X_NN) from the obtained image sequence (IMG_SEQ) of: o at least one assigned class (ATR_CLAS_OBJ) to said at least one object (OBJ), said at least one assigned class (ATR_CLAS_OBJ) being selected from a list of classes; and o at least one estimated position (EST_POS_OBJ) of said at least one object (OBJ), said at least one estimated position (EST_POS_OBJ) comprising at least one distance between an observer position (POS_OBS) and said at least one object (OBJ); said classifier-localizer module (X_NN) being configured from at least one reference image sequence (TR_DATA) to minimize a multi-objective loss function (F_LOSS) representative of both a classification objective and an object localization objective in said at least one reference sequence (TR_DATA).

2. Method according to claim 1 characterized in that said determining step (S220) comprises sub-steps of: • determination (S221) of encoded data (LS_DATA) by an encoder (CNN) from the obtained image sequence (IMG_SEQ); • determination (S222) of said at least one assigned class (ATR_CLAS_OBJ) by a classifier (CLA_NN) from the encoded data (LS_DATA); and • determination (S223) of said at least one estimated position (EST_POS_OBJ) by a locator (LOC_NN) from the encoded data (LS_DATA).

3. Method according to claim 2 characterized in that said encoder (CNN) implements one or more convolutions between the obtained image sequence (IMG_SEQ) and filters.

4. Method according to any one of claims 1 to 3 characterized in that said classifier-locator module (X_NN) comprises a neural network.

5. Method according to any one of claims 1 to 4 characterized in that said one or more images (IMG_1, IMG_N) of the obtained sequence (IMG_SEQ) are consecutive and in that the step of obtaining (S210) the sequence of images comprises a sub-step of cropping (S212) said one or more acquired images (IMG_1, IMG_N), a said object (OBJ) being centered on a determined image (IMG_1) of the sequence (IMG_SEQ).

6. Method (S100) for configuring a classifier-locator module (X_NN) of objects (OBJ) in image sequences (IMG_SEQ), said method comprising a step of initializing (S 120) said classifier-locator module (X_NN) and at least one iteration of the steps of: • determination (S130) by said classifier-localizer module (X_NN) from at least one reference image sequence (TR_DATA) of: o at least one assigned class (ATR_CLAS_OBJ) to at least one object (OBJ) in said at least one reference sequence (TR_DATA), said at least one assigned class (ATR_CLAS_OBJ) being selected from a list of classes; and o at least one estimated position (EST_POS_OBJ) of said at least one object (OBJ), said at least one estimated position (EST_POS_OBJ) comprising at least one distance between an observer position (POSJDBS) and said at least one object (OBJ); • evaluation (S140) of a multi-objective loss function (F_LOSS) on the basis of said at least one assigned class (ATR_CLAS_OBJ), of said at least one estimated position (EST_POS_OBJ), of at least one known class and of at least one known position of at least one object (OBJ) in said at least one reference sequence (TR_DATA), the multi-objective loss function (F_LOSS) being representative of both a classification objective and an object localization objective; • reconfiguration (SI 50) of said classifier-localizer module (X_NN) to minimize the multi-objective loss function (F_LOSS).

7. Method according to claim 6 characterized in that said step of evaluating (S140) the multi-objective loss function (F_LOSS) comprises sub-steps of: • evaluation (S 141) of a classification loss function (F_LOSS_CLA) from said at least one assigned class (ATR_CLAS_OBJ) and at least one known class (TR_DATA) associated with said at least one object (OBJ) of said at least one reference image sequence (TR_DATA); • evaluation (S142) of a localization loss function (F_LOSS_LOC) from said at least one estimated position (EST_POS_OBJ) and at least one known position (TR_DATA) associated with said at least one object (OBJ) of said at least one reference image sequence (TR_DATA); and • evaluation (S143) of the multi-objective loss function (F_LOSS) from the result of the classification loss function (F_LOSS_CLA) and the result of the localization loss function (F_LOSS_LOC).

8. Method according to claim 7 characterized in that said classification loss function (FLOSS_CLA) is evaluated using the expression: where L EC is the classification loss function (FLOSS_CLA), M is the number of assignable classes from said class list, δ( C i , C r ) is a binary indicator equal to 0 if for a said object (OBJ) the class C i ) of the said class list is different from class C r known to said object (OBJ) and equal to 1 otherwise, and with P o,i a probability determined by the classifier-locator module (X_NN) that said object (OBJ) belong to class C i .

9. Method according to claim 7 or 8 characterized in that said location loss function (FLOSS_LOC) is evaluated using the expression: where L IMSE is the localization loss function (FLOSS_LOC), N is the number of images of said at least one reference sequence (TR_DATA), a binary indicator equal to 1 if a said object (OBJ) is present in an image of index i of said at least one reference sequence (TR_DATA) and equal to 0 otherwise, P oi the estimated position of said object (OBJ) determined by the classifier-localizer module (X_NN) for the index image with P r,i the known position of said object (OBJ) for the image of index i.

10. Method according to any one of claims 6 to 9 characterized in that said reconfiguration step (SI 50) is carried out using a gradient descent algorithm from the result of the evaluation (S140) of the multi-objective loss function (F_LOSS).

11. Method according to any one of claims 6 to 10 characterized in that said at least one sequence of reference images (TR_DATA) comprises at least one of the elements of the following group: • one or more images acquired (TR_DATA_ACQ) by at least one camera (CAM); and • one or more synthesized images (TR_DATA_SYN).

12. Method according to any one of claims 6 to 11 characterized in that it comprises several iterations of said steps of determination (S130), evaluation (S140) of the multi-objective loss function (F_LOSS), and reconfiguration (S150).

13. Device (APP) for classifying and locating at least one object (OBJ) in a sequence of images (IMG_SEQ), said device comprising: • an obtaining module (MOD_OBT) for obtaining a sequence (IMG_SEQ) of one or more images (IMG_1, IMG_N) acquired by at least one camera (CAM); and • a classifier-localizer module (X_NN) for determining from the obtained image sequence (IMG_SEQ): o at least one assigned class (ATR_CLAS_OBJ) to said at least one object (OBJ), said at least one assigned class (ATR_CLAS_OBJ) being selected from a list of classes; and o at least one estimated position (EST_POS_OBJ) of said at least one object (OBJ), said at least one estimated position (EST_POS_OBJ) comprising at least one distance between an observer position (POS_OBS) and said at least one object (OBJ); said classifier-localizer module (X_NN) being configured from at least one reference image sequence (TR_DATA) to minimize a multi-objective loss function (F_LOSS) representative of both a classification objective and an object localization objective in said at least one reference image sequence (TR_DATA).

14. Device (APP) according to claim 13 characterized in that said classifier-locator module (X_NN) is configured by a configuration method (S100) according to any one of claims 6 to 12.

15. System (SYS) comprising a device (APP) according to claim 13 or 14 and at least one camera (CAM) configured to acquire said one or more images (IMG_1, ..., IMG_N) of the sequence (IMG_SEQ).

16. System (SYS) according to claim 15 characterized in that said system (SYS) is a surveillance system, or a navigation system.

17. Aircraft comprising a system (SYS) according to claim 15 or 16.

18. Computer program comprising instructions for implementing the steps of a method according to any one of claims 1 to 12, when said computer program is executed by at least one processor.