Vehicle environment scene classification device
The classification system for autonomous vehicles addresses the high human cost of manual annotation by using a neural network and automated image annotation, enhancing efficiency and accuracy in environmental scene classification.
Patent Information
- Application Number
- FR2023012185
- Authority / Receiving Office
- FR · FR
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-09
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2043-11-09
AI Technical Summary
Current autonomous vehicle systems require significant human effort and time for manual pixel-level annotation of environmental images, which is costly and subjective, especially when navigating unexplored scenes.
A classification system for environmental scenes of a vehicle, comprising a classification unit using a neural network to classify scenes from received images and a learning unit to train the network from a database, along with an image annotation unit that automates the annotation process by extracting images based on position and image parameter criteria, determining probability sets for each image, and merging these probabilities to generate annotations.
The system reduces the complexity and time required for image annotation, enabling efficient semantic segmentation of vehicle environments with minimal human intervention, even in unexplored scenes, and improves the accuracy and speed of autonomous vehicle deployment.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Title of the invention: Device for classifying a vehicle's environment scene Technical field
[0001] The invention relates generally to image processing systems and in particular to a device and a method for classifying a scene from the environment of an automated vehicle.
[0002] Driving system automation technologies have experienced a major boom in recent years, in the field of automated vehicles, such as autonomous vehicles, whether land, sea or underwater. Autonomous navigation is made possible by real-time recognition of the vehicle's environment, which has the advantage of requiring very little hardware. Indeed, generally, only an image acquisition device associated with a computing unit, whether onboard or not, is necessary. In existing solutions, the computing unit is configured to perform a semantic classification of the scenes in the environment. For this, the computing unit uses a semantic segmentation method which makes it possible to obtain information on the class of each of the pixels in an image.This is a complex computer vision task that classically relies on deep neural networks trained for the semantic classification of scenes in the vehicle's environment.
[0003] Thus, when the vehicle moves in its environment, it is able to analyze the environment in detail to detect obstacles, avoid them and optimize its movement in accordance with the environment in which it is moving.
[0004] In current systems, neural networks are trained using a database of thousands of semantically annotated images. Neural networks learn to automatically recognize the class of pixels in an image based on these annotations.
[0005] However, a significant obstacle to the deployment of such systems is the high human cost linked to the implementation of these approaches, called semantic segmentation. Indeed, such systems require manual annotation, at the level of each pixel of an image, among the thousands of images in the database to enable robust classification. However, these semantic annotations are necessary to be able to implement the learning phase of artificial intelligence. Thus, the deployment of autonomous vehicles on scenes that are still unexplored requires significant manual re-annotation work, leading to significant delays in the exploitation of these new environments. In addition, the precise annotation of these images is sometimes subjective due to the environment studied where the boundary between two classes is not clearly defined (for example, boundary between sand and the water column in an underwater environment).
[0006] There is thus a need for an improved device and method for classifying scenes in the environment of an automated vehicle.
[0007] General definition of the invention
[0008] The invention improves the situation by proposing a device for classifying a scene from the environment of a vehicle, comprising a classification unit using a neural network configured to classify a scene in response to a received image representing the scene and a learning unit configured to train the neural network, from a set of images stored in a database. Advantageously, the classification device comprises an image annotation unit comprising: - an image pre-processing module configured to extract a set of thumbnails from a received image as a function of a position criterion in the image and at least one criterion relating to an image parameter, - a classification module configured to determine, for each image, a set of membership probabilities, each associated with a class from a set of predefined classes, and - an image annotation module configured to determine an annotation corresponding to the image from the sets of probabilities determined for all the thumbnails, the annotation being transmitted to the database for storage in association with the image in the database.
[0009] According to one aspect of the invention, the image annotation module can be configured to merge the membership probabilities determined for all classes and for all thumbnails and to reconstruct an annotation, pixel by pixel, for the received image, which provides the annotation.
[0010] In one embodiment, each set of probabilities determined for a thumbnail image may be represented by a probability vector having a set of components corresponding to the membership probabilities, each associated with one of the classes.
[0011] In one embodiment, an image parameter may be image frequency information, the criterion relating to the image parameter being a frequency level.
[0012] The image criterion used for extraction may be relative to ranges of values of the image parameter.
[0013] In one embodiment, the preprocessing unit uses a fixed or variable step to implement a sliding window extraction of the thumbnails.
[0014] In one embodiment, the image annotation module can be configured to select for each image pixel, the class having a maximum membership probability among the membership probabilities obtained for all the thumbnails.
[0015] In one embodiment, the pre-processing module can be configured to implement, prior to determining the set of probabilities for an image, an operation of resizing the image according to the level of its image parameter.
[0016] In one embodiment, the pre-processing module can be configured to add, for each image of size KxK among the P images extracted from the input image of size MxN, the C membership probabilities obtained for the image in a patch of size KxK, corresponding to the image, on a matrix of size MxNxC corresponding to the input image.
[0017] In one embodiment, the probability matrix Mp can be defined according to the equation:
[0018] Mp(i, J,n) -Mp(i, j, n) + P(n)*?^'
[0019] where P denotes the probability vector P at the output of the classification unit and D the vector of the different dimensions of extraction of the images.
[0020] In one embodiment, the probability matrix Mp can be defined according to the equation:
[0021] jj, + g(^
[0022] P denotes the probability vector P at the output of the classification unit, G(L j) denotes a Gaussian window and Æ denotes the vector of the different dimensions of extraction of the images.
[0023] According to one aspect of the invention, the Gaussian window G(i, j) can be defined by:
[0024] G(ÎJ) =eO)xe(4X3)pourtG [0, Z]
[0025] The annotation of the received image can then be determined by selecting, at each point of the probability matrix Mp, the index having the maximum value in the probability vector associated with each pixel.
[0026] There is further provided a method for classifying a scene from the environment of a vehicle, comprising a classification step using a neural network configured to classify a scene in response to a received image representing the scene and a learning step configured to train the neural network, from a set of images stored in a database. Advantageously, the method comprises: - pre-processing of the image consisting of extracting a set of thumbnails from a received image as a function of a position criterion in the image and at least one criterion relating to an image parameter, - a classification step consisting of determining, for each image, a set of membership probabilities each associated with a class from a set of predefined classes, and - an image annotation step of determining an annotation corresponding to the image from the sets of probabilities determined for all the thumbnails, the annotation being transmitted to the database for storage in association with the image in the database.
[0027] The embodiments of the invention thus make it possible to annotate the pixels of the images in the learning database, with reduced complexity and processing time. Brief Description of Figures
[0028] Other characteristics, details and advantages of the invention will emerge from reading the description given with reference to the appended drawings given by way of example and which represent, respectively:
[0029] [Fig-1] - [Fig.l] schematically represents a vehicle implementing a control system, according to embodiments of the invention.
[0030] [Fig.2] - [Fig.2] schematically represents the structure of the annotation unit according to embodiments of the invention.
[0031] [Fig.3] [Fig.3] describes the image annotation method implemented by the image annotation unit according to embodiments of the invention.
[0032] [Fig.4] [Fig.4] represents an example of a probability matrix and patches, according to an exemplary embodiment.
[0033] [Fig.5] [Fig.5] illustrates a change in the size of the sliding window depending on the class.
[0034] [Fig.6] [Fig.6] illustrates the method of determining image annotations, in an embodiment where a single-scale approach is used.
[0035] [Fig.7] [Fig.7] illustrates the method of determining image annotations, in an embodiment where a multi-scale approach is used.
[0036] [Fig.8] [Fig.8] illustrates the method of determining image annotations, applied to a sequence of images, using the pre-trained classifier.
[0037] Detailed description of the application
[0038] [Fig.l] schematically represents a vehicle 100 implementing a vehicle control system 10, according to the embodiments of the invention.
[0039] The control system 10 is configured to recognize in real time the environment of the vehicle 100 and to adapt the driving of the vehicle 100 dynamically according to the recognition carried out in order to allow automated navigation, such as for example autonomous navigation.
[0040] The control system 10 comprises a device for classifying scenes from the environment of the vehicle 1 (also more simply called a “classification system”) and a perception system 2.
[0041] The classification device 1 is configured to perform a semantic classification of the scenes of the environment of the vehicle 100 from the images detected by the perception system 2 of the vehicle 100.
[0042] The vehicle 100 may be an automated or semi-automated vehicle capable of moving in an environment, such as a land, sea or underwater vehicle. The environment may therefore be a land, sea or underwater environment.
[0043] The environment of the vehicle 100 is associated with a set of predefined classes, depending on the application of the invention, such as for example the following set of classes: [roadway, sidewalk, vegetation, vehicle]. In embodiments, the classes associated with the environment can be modified (adding, deleting classes for example) by implementing a relearning of the system.
[0044] The control system 10 may be a driving assistance system configured to optimize the driving of the vehicle and ensure its safety (detection of obstacles for example) based on the information determined by the perception system 2.
[0045] The perception system 2 comprises a set of detection devices arranged on the vehicle 100 and configured to provide signals representative of an image of the environment in which the vehicle 100 is moving. The detection devices 20 may comprise one or more cameras. A camera may be equipped with a photosensitive sensor sensitive to radiation of a light spectrum included in the infrared and / or in the visible. The cameras may further comprise one or more acoustic and / or laser cameras (in applications of the invention in an underwater environment for example).
[0046] The detection devices 20 may be arranged on the vehicle so as to acquire images of the environment of the vehicle located in front of the vehicle 100, on the sides, above and / or below the vehicle 100 depending on the application of the invention. For example, in a terrestrial environment, the detection devices may be attached to the front of the vehicle and / or in the passenger compartment of the vehicle and / or on the roof of the vehicle. The detection devices may further comprise a lidar (Laser Detection And Ranging), and / or a radar.
[0047] Other types of sensors may be used by the perception system 2 such as an ultrasonic sensor, a steering wheel angle sensor, a wheel speed sensor, a brake pressure sensor, a transverse yaw rate and acceleration sensor, or a combination thereof.
[0048] The perception system 2 is configured to determine images of the environment from the information detected by the set of detection devices. In embodiments, the perception system 2 may further comprise a processing device 22 configured to implement multi-sensor fusion algorithms capable of combining information from the different detection devices 20 to determine the images of the environment.
[0049] The classification device 1 comprises a learning unit 11 for pre-training the neural network used on real data sets stored in an image database 3. The database contains for each data item an image / annotation pair.
[0050] The classification device 1 further comprises an image annotation unit 12 configured to automatically annotate one or more images from the image database 3 used for the learning phase. The image annotation unit 12 is thus configured to generate the annotations stored in the image database 3. An image annotation is represented by a matrix of size identical to the image where each element has a value between 0 and N1 with N designating the number of classes (membership of a pixel to a class).
[0051] Such annotations are used by the learning unit 11 to train the neural network(s) used.
[0052] The classification device 1 comprises a classification unit 14 configured to semantically segment each image detected by the perception system 2 in real time, using the previously trained neural network.
[0053] To facilitate understanding of the embodiments of the invention, definitions or concepts relating to neural networks are detailed below in relation to the embodiments of the invention.
[0054] A neural network comprises neurons interconnected with each other by synapses which can be implemented in the form of digital memories. A neural network can comprise a set of successive layers, comprising an input layer carrying the input signal and an output layer carrying the result of the prediction made by the neural network (network result), and one or more intermediate layers. Each layer of a neural network takes its inputs from the outputs of the previous layer. The number of neurons on each layer is equal to the number of inputs of the neurons of the following layer. A layer The neural network data thus includes a set of neurons taking their inputs from the neurons of the previous layer.
[0055] The signals propagated at the input and output of the layers of the network can be digital values (information coded in the value of the signals), or electrical pulses in the case of pulse coding (information coded temporally according to the order of arrival of the pulses or according to the frequency of the pulses).
[0056] A neural network comprises a set of input data (or 'input coefficients') and output data (or 'output coefficients').
[0057] The output values are calculated from the inputs and the synaptic weights by applying an activation function to the input coefficients.
[0058] Each neuron in the neural network is configured to calculate a weighted sum of its input coefficients using a combination function and the synaptic weights and then applying an activation function to the resulting weighted sum to produce its output:
[0059] The synaptic weights of a neural network are determined by learning in a learning phase implemented by the learning unit. Random values are initially assigned to the weights of the neural network and then a set of real data from the database 3 are used to carry out the learning. Learning a neural network consists of determining the optimal values of the synaptic weights, for each neuron of the neural network, from the last layer of the network to the first, using a learning function.
[0060] The training phase may implement a plurality of iterations of the training function, each iteration comprising a forward propagation step and a backward propagation step to correct errors between the outputs obtained in the forward propagation phase and the expected outputs for the input sample under consideration.
[0061] The learning phase thus makes it possible to compare the output obtained with the expected output (in the case of a supervised method), and based on this comparison, to update the connections between the neurons represented by the synaptic weights to improve the final result.
[0062] In the forward propagation phase, input data sets from the database 3 are used to implement the learning. Each data set forms a sample x associated with desired values (or expected values). The signal corresponding to the input sample is propagated forward in the layers of the neural network from the first layer, from a layer (k-1) to the next layer (k) until the last layer. In the phase of forward propagation, the activation function q> and the synaptic weights connecting neurons from a previous layer (k-1 ) and a next layer (k) are used.
[0063] When the forward propagation is completed, a result y is obtained at the output.
[0064] In the backpropagation phase, any errors obtained by a neuron are backpropagated to its synapses and to the neurons connected to it. Backpropagation can be, for example, gradient backpropagation to modify synaptic weights by taking into account their impact on the errors generated. Thus, synaptic weights that contribute to generating a significant error can be modified more significantly than weights that have generated a less significant error.
[0065] The duration of the learning phase may depend on the size of the database 3 storing the samples used for learning and the size of the network.
[0066] After the learning phase, a faster so-called generalization phase is implemented by the classification unit 14. In the generalization phase, the weights learned from the learning phase are used (static neural network in which the weights are fixed). Input data corresponding to the images of the scene detected in real time by the perception system 2 are presented to the neural network which provides a response from the neural network representing the probability of each pixel belonging to the predefined classes.
[0067] The control system 2 may implement a set of driving applications using the classification determined by the classification device 1 such as a lane change application to assist the driver in changing lanes or a collision warning application.
[0068] [Fig.2] schematically represents the structure of the annotation unit 12 according to embodiments of the invention.
[0069] The image annotation unit 12 is configured to determine an annotation to be assigned to an image representative of a scene of the environment to be analyzed, in response to the reception of this image determined by the perception system 2. The classification device 1 can then store the image in association with the determined annotation in the database 3.
[0070] The scene classification device 1 can thus perform a semantic annotation of the images making it possible to produce a semantic segmentation of the environment of the vehicle 100 in minimal time, even in the case of scenes that are still unexplored where no annotation is available.
[0071] In one embodiment, as illustrated in [Fig.2], the image annotation unit 12 may comprise: - an image preprocessing module 120 configured to preprocess the images and extract sub-images (or thumbnails) from the received image (input image); - a classification module 122 (also called a “classifier”) configured to perform a classification of the thumbnail images extracted by the preprocessing unit 102, using the neural network, to determine the probabilities of each thumbnail image belonging to each of the classes of the vehicle's environment; - An annotation module 124 configured to merge the responses of the classification module 122 obtained for each image and to reconstruct a semantic annotation, pixel by pixel, for the complete image considered, which provides the annotation to be assigned to the received image.
[0072] The image preprocessing module 120 can be configured to collect the local information of the image considered, according to several levels and to extract a set of thumbnails according to an image position criterion and at least one image criterion relating to an image parameter. A thumbnail can thus correspond to position information in the image and to image criteria defined from the scale levels of the image parameter considered (i.e. ranges of values of the image parameter).
[0073] The image criterion used for extracting the thumbnails may relate to different types of image parameters, such as for example the image frequency, the image texture, the image contours.
[0074] The scale levels of the image parameter considered for the extraction may be, for example, a scale pyramid of a given factor, such as, for example, a factor of 2. In an example of application of the invention where the image criterion used for the extraction relates to the image parameter corresponding to the maximum image resolution, for a maximum resolution of 224x224 pixels, and a minimum resolution of 14x14 pixels, the scale pyramid of factor 2 used comprises the levels: 224x224; 112x112; 56x56; 28x28; and 14x14.
[0075] The maximum resolution may depend on the resolution of the perception system (2). The thumbnails are extracted for each scale level.
[0076] The preprocessing unit 120 can take into account the low and high frequency information inherent to each of the classes of the scene of the environment considered. The preprocessing unit 120 can use a fixed step to implement a multi-scale sliding window extraction. The step can be predefined. The step can be applied to the pixel grid of the image (vertically and horizontally). The step and the scale levels make it possible to configure the fineness of the semantic annotations that will be automatically generated.
[0077] The classifier 122 is configured to determine, for each extracted image, a set of membership probabilities each associated with a class from a set of predefined classes.
[0078] In one embodiment, the classifier 122 is configured to assign, to each image extracted by the image preprocessing module 120, a characteristic vector (also called a “probability vector”) comprising components, each component being associated with one of the classes of the scene of the environment considered and having a value corresponding to the probability of the image belonging to the associated class. Thus, the probability vector assigned to an image represents the probability of the image belonging to each of the classes of the scene of the environment considered. The classifier 122 uses the neural network, which may in particular be a deep neural network, pre-trained beforehand on a suitable data set, during a learning phase, to enable robust classification.
[0079] The annotation module 124 is configured to merge the responses of the classifier 122 (represented by the probability vectors obtained for each image) so as to aggregate the membership probabilities obtained for each image, corresponding to different image positions and image criteria. The annotation module 124 is in particular configured to assign to each pixel of the initial image a probability vector indicating its degree of membership to the classes present in the scene considered. The annotation module 124 thus associates with the image considered a semantic annotation by selecting for each image pixel the class having a maximum membership probability. As used herein, the term “annotation” is associated with an image and refers to a set of class information, each class information being associated with a pixel of the image and corresponding to the determined maximum probability class.
[0080] The image annotation unit 12 can then store the determined annotations in association with the corresponding image in the image database 3.
[0081] The learning unit 11 can then implement learning of the semantic segmentation neural network not using annotations of the same type.
[0082] To carry out the learning, the learning unit 11 can use the annotations thus generated.
[0083] It should be noted that the annotations determined by the reconstruction unit 124 may not be perfect and form a set of so-called degraded annotations. However, the construction of the neural network of the modules 11 and 14 being of the encoder / decoder type, this type of network will naturally smooth the annotations provided as inputs and thus limit the impact of the imperfection of such degraded annotations.
[0084] The classifier 122 can be trained to recognize the classes on thumbnails of a given size, such as for example images of sizes 224x224. The classifier 122 can be configured to predict the probabilities of belonging to the predefined classes on each thumbnail extracted from the image, at several scale levels. In a mode of particular embodiment, the classifier 122 can in particular be configured to predict the majority class (i.e. having a maximum probability of belonging) on each image extracted from the image, at several scale levels.
[0085] Such predictions can be obtained by applying an underlying principle of multi-scale sliding window. Sliding window is a technique for analyzing an image by dividing it into small regions called windows and gradually moving this window across the entire image to perform local analysis. The sliding window process involves three steps: - a window initialization step in position and size. The window can be square and its initial position located at the origin of the image. The size of the window depends on the scale fixed in the multi-scale approach. - a step of analyzing the pixels of the window by performing a classification of the window to extract its probabilities of belonging to the classes of the system. - a window movement step, after analyzing the current region. In this step, the window is moved by a certain number of pixels, called "step". This step is repeated until the window has covered the entire image. The step can be fixed and equal to the resolution of the minimum scale (for example, for an image size of 14x14, the step is equal to 14).
[0086] A multi-scale approach allows for as many sliding windows as desired sizes (scale pyramid).
[0087] The fusion of the class predictions available for each pixel, carried out by the reconstruction unit 124, makes it possible to obtain a semantic annotation of the image considered which can be degraded. Each annotation thus obtained enriches the semantic annotation database 3 and can be used for the implementation of the learning by the learning unit 11, which makes it possible to use a strong semantic segmentation neural network.
[0088] [Fig.3] describes the image annotation method implemented by the image annotation unit according to embodiments of the invention.
[0089] In step 300, an input image of size MxN is received.
[0090] In step 302, the input image is pre-processed. This step includes the extraction of sub-images (or thumbnails) of the input image. The image can notably be a color image. Each thumbnail can be associated with an image resolution.
[0091] In step 302, the characteristics of the input image received in step 300 may be stored in a thumbnail memory (e.g., a thumbnail catalog), in which the different thumbnails extracted from the image are associated with fixed image positions and resolutions, which may be different. The thumbnails may advantageously be defined according to several scale levels. Taking into account the multi-scale information of the thumbnails allows the description of the image both at the level of the textures present in it, and characterizing each class, but also at the level of the structures and contours visible in the image. In one embodiment, the method can for example implement a multi-frequency approach allowing the extraction of low and high frequency information simultaneously.
[0092] In step 304, the probabilities of each extracted image belonging to each of the classes of the vehicle's environment are determined by implementing a classification (i.e. a class prediction) of the image based on the neural network, providing a vector of probabilities for each class (this step can be implemented by the classifier 122).
[0093] Step 304 (classification step) consists of determining, from the neural network previously trained on annotated thumbnails having a fixed size (for example 224x224), the probabilities of belonging to the different classes locally present in each thumbnail extracted in step 302, and the probability vector grouping these probabilities of belonging for the different classes. The neural network may be a deep neural network, such as for example a network of the AlexNet or SqueezeNet type. The neural network is used to determine the probabilities of belonging of each extracted thumbnail to each of the C classes, C being the number of classes present in the considered scene of the environment. These classes may be predefined during the training (i.e. the learning) of the image classification network.In embodiments, step 304 may comprise, before determining the probability vector of an image, an operation of resizing an image according to its scale level, in order to make the input of the classifier 122 compatible with the size of the image.
[0094] Step 304 thus provides, for each image, C membership probabilities which locally describe the content of the original image at a given position and scale of the image corresponding to a image (the C membership probabilities being able to constitute the C components of the probability vector which is therefore then of size C).
[0095] In step 306, a semantic annotation is determined for each pixel of the image using the probability vectors obtained for the different thumbnails from the input image. Step 306 can be implemented by the image annotation reconstruction unit 124.
[0096] In step 306, a semantic annotation is determined for each image from the probabilities of belonging to the classes determined in step 304, for all the thumbnails from the input image. Thus, considering P thumbnails extracted in step 302, step 306 uses the P probability vectors determined for the P thumbnails.
[0097] Step 306 comprises, for each image of size KxK among the P images extracted from the input image of size MxN, an addition of the C membership probabilities obtained for the image in a patch of size KxK (a patch is a square centered on an image pixel), corresponding to the image, on a matrix of size MxNxC corresponding to the input image. Such a matrix therefore represents the probabilities that each of the classes appears at any point of the original image. Step 302 of extracting the images according to position and scale level criteria makes it possible to completely construct the probability matrix.
[0098] Figure 4 represents an example of a probability matrix 40 and patches 400, according to an exemplary embodiment. Figure 4 illustrates a change in the pitch of the sliding window which extracts the thumbnails, the offset between two thumbnails being fixed at the start of the method and remaining unchanged throughout its application. It should be noted that it is possible to go down to a pitch of 1 pixel. However, the smaller the pitch, the greater the number of operations. For example, for a pitch of 224 / x pixels, the number of operations is greater than for a pitch of 224 pixels.
[0099] In one embodiment, a class may be constrained to a given image resolution to take into account the variety of environments in which autonomous vehicles operate. Indeed, in such environments, it is common for certain classes to represent texture information rather than structural or contour information and vice versa. According to the application of the invention, it is possible to observe the results at different extraction dimensions (i.e. scales) and to choose the most consistent dimension for each class.
[0100] [Fig.5] illustrates a change in the size of the sliding window depending on the class. As illustrated in [Fig.5], a different factor can be applied for each class to account for the varying size of the patches added to the probability matrix.
[0101] On the same image, thumbnails of different sizes are extracted and make it possible to establish the probabilities of belonging to specific classes according to the information they contain (for example, frequency level information, low or high frequencies).
[0102] Considering a probability matrix Mp, a probability vector P at the output of the neural network-based classifier 122 and the vector D of the different extraction dimensions (i.e. scale levels), the probability matrix can be defined according to the following equation (1):
[0103] Mp{iJ,n) = Mp(i,j,n) + P(n)*?ÿ^
[0104] For a given image dimension, only the probabilities on the classes concerned by this dimension are added in step 302.
[0105] In one embodiment, the step set in step 302 of extracting the thumbnails (step of the sliding window) can make it possible to control the fineness of the semantic annotation of the final image. In embodiments, a Gaussian filtering type smoothing can also be implemented to improve the readability of the information present at the boundaries of the classes, which are defined by the size of the extraction step, according to equation (3). Indeed, with a given step, it is generally not possible to obtain different predicted classes on a surface smaller than such a pixel step. In embodiments, each probability matrix resulting from step 304 can be weighted by a two-dimensional Gaussian window, for each scale level used.Thus, instead of directly adding the probability of the class considered on a ZxZ surface, the results of a two-dimensional Gaussian window of size ZrZ multiplied by the probability of the class are added according to the following equation (2): . [01°6] Mp(i,j,n) = Mp(i,j,n) + (2)
[0107] In equation (2), the Gaussian window G(iy j) is defined by:
[0108] G(i,j) =e(BW^) for [QZ], i [0, Z] (3)
[0109] The semantic annotation of the initial image can then be determined by selecting, at each point of the probability matrix Mp, the index between 0 and C, of the class having the highest probability. Thus, each pixel of the input image is assigned the class for which its membership is most probable.
[0110] [Fig.6] illustrates the method of determining image annotations, in an embodiment where a single-scale approach is used.
[0111] In the example of Figure 6, the method according to the embodiments of the invention determines deteriorated annotations for C classes. The original image is divided, for example, into 224x224 thumbnails, and the classifier 122 is applied to each thumbnail to form a vector of probabilities of belonging to each of the classes. The vector thus obtained is added to the part corresponding to the extracted thumbnail in a probability matrix of the same size as that of the input image, but having a depth equal to C, instead of the depth of the input image (which may be an RGB image for 'Red Green Blue') equal to 3. In one embodiment, in step 308, each pixel of the probability matrix Mp may be associated with a color to determine annotations suitable for visualization in a human-machine interface. Furthermore, the index having the maximum probability (i.e.index having the maximum value in the probability vector associated with each pixel) can be extracted to automatically generate the semantic annotation of the initial image.
[0112] [Fig.7] illustrates the method for determining image annotations, in an embodiment where a multi-scale approach is used. In such an embodiment, thumbnails of different resolutions are extracted from the initial image to encapsulate the scale level information (e.g. high and low frequency information) and thus extract not only the contours of the image but also the image texture. The different thumbnails can then be adjusted in size to be classified in step 304 before being merged in step 306 in order to determine the final probability matrix for automatically constructing the semantic annotation of the input image considered.
[0113] [Fig.8] illustrates the method for determining image annotations, applied to a sequence of images, using the pre-trained classifier 122 (block 81) from a classifier model 82 and annotated images 80A for each class composing the scene studied, a sequence of input images 80B for which a semantic annotation is to be determined. The method implements a multi-scale classification approach by weighted sliding window (block 83) to establish a membership probability matrix (block 85) of each pixel to a given class to determine (block 84) the semantic annotation of the initial image, which makes it possible to obtain annotated images 85.
[0114] The embodiments of the invention are thus particularly suitable for any application using a semantic segmentation neural network and more particularly its relearning from new data, such as for example a seabed segmentation application. They allow the deployment of AUVs in such environments with a reduced semantic annotation cost. In such an example application, the embodiments of the invention can be used for risk detection. Indeed, they make it possible to automatically or dynamically determine semantically annotated images on a predetermined number of classes representative of the seabed, to highlight the proportion of each class in an image and to analyze whether a part of the image does not belong to these classes, which indicates the probability of being faced with an unknown element.According to another example, the system and method according to the embodiments of the invention can be used to calculate the proportion of the representation of the class "water" in the image (i.e. the class representing water in the images) in order to determine whether an AUV (acronym for "Autonomous Underwater System") is at an acceptable depth to navigate in its environment using technologies such as SLAM (Acronym for "simultaneous localization and mapping") for example.
[0115] Those skilled in the art will understand that the system or subsystems according to the embodiments of the invention may be implemented in various ways by hardware, software, or a combination of hardware and software, in particular in the form of program code that may be distributed as a program product, in various forms. In particular, the program code may be distributed using computer-readable media, which may include computer-readable storage media and communication media. The methods described in the present disclosure may in particular be implemented in the form of computer program instructions executable by one or more processors in a computer computing device. These computer program instructions may also be stored in a computer-readable medium.
[0116] The invention is not limited to the embodiments described above by way of non-limiting example. It encompasses all the variant embodiments which may be envisaged by those skilled in the art.
Claims
Claims
1. Device for classifying a scene from the environment of a vehicle, comprising a classification unit (14) using a neural network configured to classify a scene in response to a received image representing the scene and a learning unit (11) configured to train said neural network, from a set of images stored in a database (3), characterized in that it comprises an image annotation unit comprising: - an image pre-processing module (120) configured to extract a set of thumbnails from a received image as a function of a position criterion in the image and at least one criterion relating to an image parameter, - a classification module (122) configured to determine, for each thumbnail, a set of membership probabilities each associated with a class from a set of predefined classes,and - an image annotation module (124) configured to determine an annotation corresponding to said image from the sets of probabilities determined for all said thumbnails, said annotation being transmitted to said database (3) for storage in association with said image in the database.,
2. A classification device according to claim 1, wherein the image annotation module (124) is configured to merge the membership probabilities determined for all classes and for all thumbnails and to reconstruct an annotation, pixel by pixel, for said received image, thereby providing said annotation.
3. Classification device according to one of claims 1 and 2, in which each set of probabilities determined for a thumbnail is represented by a probability vector having a set of components corresponding to said probabilities of membership, each associated with one of said classes.
4. Classification device according to one of the preceding claims, in which an image parameter is image frequency information, the criterion relating to said image parameter being a frequency level.
5. Classification device according to one of the preceding claims, in which the pre-processing unit uses a fixed or variable step to implement an extraction of the thumbnails by sliding window.
6. Classification device according to one of the preceding claims, in which the image annotation module (124) is configured to select for each image pixel, the class having a maximum membership probability among the membership probabilities obtained for all the thumbnails.
7. Classification device according to one of the preceding claims, in which the pre-processing module is configured to implement, prior to determining the set of probabilities for an image, an operation of resizing the image according to the level of its image parameter.
8. Classification device according to one of the preceding claims, in which the pre-processing module is configured to add, for each image of size KxK among the P images extracted from the input image of size MxN, the C membership probabilities obtained for the image in a patch of size KxK, corresponding to the image, on a probability matrix of size MxNxC corresponding to the input image.
9. Classification device according to claims 3 and 8, in which the probability matrix Mp is defined according to the equation: Mp ( 4 j, n ) = Mp ( z, j, n ) + P ( n ) ' where P denotes the probability vector P at the output of the classification unit and P the vector of the different extraction dimensions of the thumbnails.
10. Classification device according to claims 3 and 8, in which the probability matrix Mp is defined according to the equation: Mp(iJ,n) = Mp(i,j,n) + G(iJ)*P(n)*'^^ where P denotes the probability vector P at the output of the classification unit, G(i, j) a Gaussian window and P the vector of the different dimensions of extraction of the thumbnails.
11. Classification device according to claim 10, in which the Gaussian window G(L j ) is defined by: G(i, j) = for[0,Z], i [0, Z], the annotation of the received image then being determined by selecting, at each point of the probability matrix Mp, the index having the maximum value in the probability vector associated with each pixel.
12. Method for classifying a scene from the environment of a vehicle, comprising a classification step using a neural network configured to classify a scene in response to a received image representing the scene and a learning step configured to train said neural network, from a set of images stored in a database (3), characterized in that the method comprises: - a pre-processing of the image (120) consisting of extracting a set of thumbnails from a received image as a function of a position criterion in the image and at least one criterion relating to an image parameter, - a classification step (122) consisting of determining, for each thumbnail, a set of membership probabilities each associated with a class from a set of predefined classes,and - an image annotation step (124) of determining an annotation corresponding to said image from the sets of probabilities determined for all said thumbnails, said annotation being transmitted to said database (3) for storage in association with said image in the database.,