Method, system and apparatus for classification of environmental conditions of a vehicle

The integration of gated camera and LiDAR sensors with convolutional neural networks enables precise classification of environmental conditions, improving vehicle control in adverse weather and lighting, addressing the limitations of existing systems.

WO2026154215A1PCT designated stage Publication Date: 2026-07-23TEKNOLOGIAN TUTKIMUSKESKUS VTT OY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
TEKNOLOGIAN TUTKIMUSKESKUS VTT OY
Filing Date
2026-01-14
Publication Date
2026-07-23

AI Technical Summary

Technical Problem

Existing systems fail to accurately classify environmental conditions such as weather and road conditions for autonomous vehicles, leading to reduced sensor performance and inadequate adaptation of vehicle controls.

Method used

A method using a combination of gated camera and LiDAR sensors, processed by convolutional neural networks, to classify environmental conditions, including lighting, rain, and road states, enabling adaptive vehicle control.

Benefits of technology

Enhances the accuracy of environmental condition detection, allowing vehicles to adapt driving controls effectively in various weather and lighting conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FI2026050013_23072026_PF_FP_ABST
    Figure FI2026050013_23072026_PF_FP_ABST
Patent Text Reader

Abstract

An automated vehicle comprises a gated camera system (16-1), a long-range LiDAR system (14), and a control system (12) controlling an operation of the vehicle. The vehicle is further provided with a convolutional neural network, CNN, based controller for detecting environmental conditions of an operational environment of an automated vehicle. A Lidar image from the LiDAR sensor system is encoded into first feature vector (205) by a first pre-trained CNN (204), while a gated image and gated camera capture parameters from the gated camera system are encoded into a second feature vector (203) by a second pre-trained CNN (202). The first and second feature vectors (203, 205) are combined into a combined feature vector (207) that is inputted to two or more pre-trained environmental condition classifiers (210, 212, 214). Each classifier is configured to classify one environmental condition that can be used for controlling the vehicle.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] METHOD, SYSTEM AND APPARATUS FOR CLASSIFICATION OF ENVIRONMENTAL CONDITIONS OF A VEHICLE

[0002] FIELD OF THE INVENTION

[0003] The present invention relates to classification of environmental conditions of a vehicle.

[0004] BACKGROUND OF THE INVENTION

[0005] Modern motor vehicles are often provided with technologies that assist drivers of motor vehicles with the safe operation of a vehicle or enables semi-au-tonomous driving or autonomous self-driving vehicles. Driver assistance systems may have safety features that are designed to avoid crashes and collisions by alerting the driver to problems, implementing safeguards, and taking control of the vehicle, if necessary. Adaptive features may automate lighting, provide adaptive cruise control, assist in avoiding collisions, alert drivers to possible obstacles, assist in lane departure and lane centering, etc. The vehicles are typically equipped with various types of sensors in order to detect objects in the surroundings and on the road ahead, such as other road users, traffic signs, lane markings, or obstacles on the road. The sensors scan and record data from surroundings of the vehicle, and a computing device of the vehicle may use the sensor data to detect objects and their respective characteristics (position, shape, size, heading, speed, etc.). The information collected from the sensors may be merged with the vehicle’s direction and speed to assess whether an impact with an obstacle is possible. Once an impending collision is detected, these systems may provide a warning to the driver. When the collision becomes imminent, they may take the required safety action autonomously without any driver input (e.g., by braking or steering or both). Example sensors may include cameras (e.g., gated cameras or stereo cameras), radio frequency-based detection and ranging (RADAR) units, or laser rangefinder and / or light detection and ranging (LiDAR) units.

[0006] Awareness of surrounding environmental state is an important safety factor for an automated vehicle. In some cases, autonomous vehicles or driver assistance systems may not be able to operate or may not operate well when encountering certain weather conditions in the automated vehicle's operational environment. Weather conditions such as rain (light rain, moderate rain, heavy rain), wet roads, fog, sleet, snow, hail, and so on, can cause the behaviour of the autonomous vehicle to be modified. Such weather conditions may be referred to herein as adverse weather conditions as they may result in reduced sensor and / or operationalcomponent performance of the automated vehicle. Different factors affect different aspects of vehicle safety. The visibility of cameras changes depending on the lighting conditions (e.g , amount of natural / ambient light, daylight / dark at night). Low / poor lighting conditions, such as darkness in nighttime, hinders the view of RGB cameras and can require parameter changes for other camera types as well. Rain and snowfall also hinder or change the sensor views in cameras and Lidars, and having knowledge of the rain state would allow the vehicle to change its driving processing parameters to "rain mode". Similarly, rain and snowfall can hinder a cameras view, but also adds noise to lidar point clouds. If the road surface is covered with snow or water, different steering parameters may be used to assure safety. As such, the automated vehicle should provide for detection of various weather conditions of the operational environment of the vehicle.

[0007] Prior art solutions mainly aim to improve the object detection in adverse weather, but they do not actually classify the environmental state in itself. They try to generalize the detection models to every weather, which is good for object detection, but it does not solve the problem of knowing what the road and rain conditions are to adapt the vehicle steering controls appropriately.

[0008] Saket S Chaturvedi etal, Pay "Attention" to Adverse Weather: Weather-aware Attention-based Object Detection, 2022, disclose an improved object detection in adverse weather conditions.

[0009] George Sebastian et al, RangeWeatherNet for LiDAR-only weather and road condition classification, 2021 IEEE Intelligent Vehicles Symposium (IV), 2021, p. 777-784, disclose an architecture for weather and road condition classification based on LiDAR point clouds. The presentation of the LiDAR cloud is fed to a convolutional encoder, at whose end there are two classification heads attached: weather classification head and road condition classification head. The approach involves problems relating to the accuracy of the classification as well as relating to a low-light state estimation and classification.

[0010] BRIEF DESCRIPTION OF THE INVENTION

[0011] An object of the present invention is to provide a method and a control system for implementing the method that is able to accurately estimate and detect relevant environmental conditions of an automated vehicle. The object of the invention is achieved by a method, a control system, a vehicle, and an apparatus recited in the independent claims. The preferred embodiments of the invention are disclosed in the dependent claims.An aspect of the invention is a method for detecting environmental conditions of an operational environment of an automated vehicle, comprising receiving Lidar image from a LiDAR sensor system,

[0012] encoding the Lidar image into first feature vector by a first pre-trained convolutional neural network, CNN,

[0013] receiving gated image and gated camera capture parameters from a gated camera system,

[0014] encoding the gated image and the gated camera capture parameters into a second feature vector by a second pre-trained CNN,

[0015] combining the first and second feature vectors into a combined feature vector,

[0016] inputting the combined feature vector to two or more pre-trained environmental condition classifiers, each classifier being configured to classify one environmental condition, and

[0017] controlling a vehicle based on a classified environmental condition. In an embodiment, the second pre-trained CNN comprises several convolution layers, and wherein the method further comprises inserting the gated camera capture parameters each to different one of the several con-volution layers in the second pre-trained CNN.

[0018] In an embodiment, the receiving of the gated camera capture parameters comprises receiving a gated camera parameter array including the gated camera capture parameters for each of n laser pulse series fired by the gated camera for each gated camera image, and forming a gated camera capture parameter vector of length n and including the same parameter of each of the n laser pulse series, and wherein each of the gated camera capture parameter vectors is inserted to different one of the several convolution layers in the second pre-trained CNN.

[0019] In an embodiment, the inserting comprises expanding the gated camera capture parameter vector of length n to dimensions that match to dimensions of convolution features at the respective convolution layer which the gated camera capture parameter vector is inserted to.

[0020] In an embodiment, the expanding comprises linear mapping of parameter values of the respective gated camera capture parameter vector to values in the matching dimension.

[0021] In an embodiment, the inserting comprises summing the gated camera capture parameters to the convolution features at the respective convolution layer.In an embodiment, the gated camera capture parameters include one or more of the following: laser is active or not, number of gate pulses, number of laser pulses, total time of laser illumination, total time of gate on.

[0022] In an embodiment, the gated camera capture parameters include information on the autoexposure settings used by the gated camera system.

[0023] In an embodiment, the controlling the vehicle based on a classified environmental condition comprises adapting driving control schemes or algorithms of the vehicle according to the classified environmental condition.

[0024] In an embodiment, the classified environmental condition is notified to a driver or passenger of the vehicle, preferably displayed on a user interface of the vehicle.

[0025] In an embodiment, the controlling the vehicle based on a classified environmental condition comprises controlling one or more components and / or functions of the vehicle, preferably including one or more of a vehicle controller, brakes, engine, lighting, steering.

[0026] In an embodiment, the classified environmental condition comprises one or more of: a lighting condition, a rain state, a road state (generally a terrain state).

[0027] A second aspect of the invention is a control system configured to implement the steps of the method according to embodiments of the first aspect for detecting and classifying environmental conditions of an operational environment of an automated vehicle.

[0028] A third aspect of the invention is a vehicle, comprising

[0029] a gated camera system,

[0030] a long-range LiDAR system,

[0031] a control system controlling an operation of the vehicle, and means for implementing the steps of the method according to embodiments of the first aspect.

[0032] A fourth aspect of the invention is an apparatus comprising at least one processor and at least one non-transitory memory including computer program code instructions, the computer program code instructions being configured to, when executed, cause the apparatus to carry out the steps of the method according to embodiments of the first aspect.

[0033] BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In the following the invention will be described in greater detail bymeans of exemplary embodiments with reference to the attached drawings, in which

[0035] Fig. 1 is a simplified block diagram of an exemplary sensor and control arrangement for an automated vehicle in which embodiments of the invention may be applied;

[0036] Figure 2 shows simplified examples of an image view of a gated camera divide into distance ranges (image slices), an area of interest (R01) bounding an object, and a LiDAR scanning view focused to the R01;

[0037] Figure 3 shows a simplified example of a 2D image slice that contains an object;

[0038] Figure 4 shows a simplified functional block diagram illustrating an exemplary convolutional neural network (CNN) system for detecting environmental conditions of an operational environment of an automated vehicle according to principles of the present invention;

[0039] Figure 5A shows a simplified functional block diagram illustrating an exemplary deep convolutional encoder;

[0040] Figure 5B shows a simplified functional block diagram illustrating an exemplary classification stage;

[0041] Figure 6 shows a simplified functional block diagram illustrating an exemplary implementation of the convolutional encoders for deep combination of gated camera images, gated camera capture parameters, and LiDAR images; and Figure 7 shows a simplified flow diagram illustrating an example of expanding a gated camera capture parameter vector Xi 2to match a feature map at a respective convolution layer.

[0042] DETAILED DESCRIPTION OF THE INVENTION

[0043] The present invention relates generally to detection of environmental conditions of an operational environment of an automated vehicle, e.g., weather, lighting conditions (e.g., dark at night / daylight) and / or a condition of a road / ter-rain surface, etc. Embodiments of the invention may be used both in manually driven or human-driven vehicles, and in autonomous or self-driving vehicles, for driving assistance or automated driving. The primary objective is to provide accurate and reliable detection of all interesting environmental conditions so that the vehicle can be controlled according to the detected environmental condition.

[0044] Automated vehicles may be vehicles that use multiple sensors to sense the environment and move without a human driver (autonomous vehicles) orassist the human driver (vehicles with driver assistance system, DAS). An example automated vehicle can include various sensors, such as a camera sensor, a light detection and ranging (LIDAR) sensor, and a radio detection and ranging (RADAR) sensor. Camera sensors may include standard visual cameras, such as RGB cameras, or gated cameras. The sensors collect data and measurements that the automated vehicle can use for operations such as navigation. The sensors can provide the data and measurements to an internal computing system of the automated vehicle, which can use the data and measurements to control driving parameters and / or various driving-related system of the automated vehicle, such as a vehicle propulsion system, a braking system, a steering system, lightning system, etc. A road object detection can be based on the fusion of a gated camera vision and Li-DAR technologies. Object detection is a key technology behind driver assistance systems (DAS) and autonomous vehicles that enable cars to detect driving lanes or objects on the road ahead (such as pedestrians, other vehicles, lost cargo dropped on the road, etc.) detection to prevent accidents.

[0045] Awareness of surrounding environmental state is an important safety factor for an automated vehicle. In some cases, autonomous vehicles or driver assistance systems may not be able to operate or may not operate well when encountering certain environmental or weather conditions in the automated vehicle's operational environment. The environmental conditions we may need to observe may include generally lighting conditions (e.g., daylight, dark at night), a terrain / road state, and a rain state. The rain state may include different forms of rain, e.g. one or more of following states: no rain, rain, light rain, moderate rain, heavy rain, snowfall, sleet, hail, etc. The terrain / road conditions may include different forms of road states, e.g. one or more of following states: a dry road, a wet road, a snowy road, an icy road, and so on

[0046] To accurately estimate these environmental conditions, we need sensor data which covers all the states we are interested in. In other words, sensors where there is a (visual) difference in the data depending on changes in the environmental states. For example, rain can be seen in most LiDARs as either grainy noise, or occluded spots in the LiDAR view (from water drops stuck on the LiDAR). However, discerning natural / ambient light state (e.g., daylight vs night) from a LiDAR is trickier, since most of the reflective environment does not change depending on darkness or light. From an RGB camera view we can discern the "terrain state" since the difference between a snow-covered road, a rain covered road, and a dry road can be quite clear. However, if we want to be able to make this classification in lowlight conditions or in dark (e.g., at nighttime), then the RGB camera will be much more unreliable. Thus, solutions using a standard camera, such as RGB camera, or LiDAR camera or both will face the problem relating to a low-light state (e.g., in dark at nighttime) estimation and knowing the road and rain conditions so that the automated vehicle can adapt the vehicle steering controls appropriately.

[0047] Fig. 1 is a simplified block diagram of an exemplary sensor and control arrangement in which embodiments of the invention may be applied. The exemplary sensor and control arrangement is mounted on a vehicle for the purpose of driving assistance or automatic driving and arranged to detect environmental conditions existing around the vehicle. The object detection system includes, at least one LiDAR sensor 14, at least one gated camera sensor 16-1...16-N, and a control system 10. The control system 10 may include a controller or processor 102, and a memory 104. The control system 10 may be a controller, or part of the controller, of a vehicle. The control system 10 may include a controller of the LiDAR sensor system and / or a controller of the gated camera sensor system 16-1...16-N. The memory 104 may include instructions executable by the processor 102 and may also store operational data. The control system 10 may be configured to receive sensor data 15 from the LiDAR sensor(s) 14 and sensor data 17 from the gated camera sensor(s) 16-1...16-N, and data from possible other sensors, such as an RGB camera sensor 13, and it may also control their operation. The control system 10 may also receive operational data of the vehicle, such as speed, GPS location, weather information, etc. The control system 10 may be configured to work in an interconnected fashion with various components of the vehicle, such as a vehicle controller, brakes, engine, lighting, steering, etc., designated generally as a vehicle control system 12 in the example of Fig. 1, to enable driving assistance or automated driving functions.

[0048] In gated imaging technology, the gated camera sensor 16-1 includes two key components: a pulsed light source 160 and a gated image sensor 162 with controlled opening and closing (gating) of the gated image sensor 162 in synchronism with the pulsed light source 160. The pulsed light source 160 emits a light pulse at time tO, while the gated image sensor 162 is closed. At time tl the light pulse is reflected from the targeted object 18. At time t2, after a certain time delay (t2 - ti) that corresponds to the distance of interest, the camera sensor 162 is opened (starts exposure) for a short period of time (At) corresponding to the desired depth of view. Only the photons that arrive within the right period of time (a temporal slice) contribute to the resulting image. Therefore, the resulting image consists ofinformation only from reflected photons at the distance of interest. In other words, the gated camera divides a shooting range into a plurality of distance ranges. Therefore, the technology is also called range gating. The time delay or gating delay (t2 - ti) determines the position of the range in the scene (e.g., dl, d2, d3 in Fig. 2), and the camera gating time or exposure time (At) will define the depth of the range (e.g., Ad in Fig. 2). Thus, a two-dimensional (2D) image of a 3D space is obtained for each range, each slice image including only objects present in the corresponding range. In the example of Fig. 2, three image slices 1, 2, and 3 within the imaging view 22 of the gated camera sensor 16 are depicted, the image slice 3 containing the object 18. Fig. 3 illustrates a simplified example of the 2D image slice 3 that contains the object 18.

[0049] In embodiments, the control system 10 or the like may process the image slices. In the example of Fig. 2, the distance ranges or image slices 1 and 2 contain an adverse weather phenomenon 20, such as water vapor from rain, fog, or snow, which may reflect the emitted light pulses in front of the image slice 3. However, these reflections do not interfere the image slice 3, because they arrive too early to the gated image sensor 162 which is closed until time t2. Only reflections from the range 3 at the distance ds arrive at or after time ts when the gated image sensor 162 is gated open and imaging. Thereby, the backscatter from fog, rain or snow is gated out, and the image slice 3 is a clear image of the scene at distance ds.

[0050] Fig. 1 further shows a simplified LiDAR sensor 14 that includes a pulsed light source 160 (e.g., laser or light diode) and a photo sensor 142 (such as photodiode). For example, the LiDAR sensor 14 may include a laser range finder reflected by a rotating mirror, for example, and the laser is scanned around a scene being digitized, in one or two dimensions, gathering distance measurements at specified angle intervals. LiDAR sensors measure the time-of-flight (ToF) of pulses of light emitted by the light source 140 into the scene and returned to the photo sensor 142 along a path that is sequentially scanned across a scene. The returned reflections can be used to create 3D point clouds that precisely record the external surfaces of objects and scenes. In other words, a point cloud is a collection of points of data plotted in 3D space, using a 3D pulsed light scanning. As a result, the LiDAR sensor 14 delivers a precise depth with high spatial resolution at short distances. The existing LiDAR systems are developed to scan a predefined area with fixed sensor parameters and scanning pattern. However, they suffer from quadratically decreasing spatial resolution at longer distances, prohibiting detection tasks for far objects, e.g., resulting in a few measurement points for pedestrians at 100 mdistance. The LiDAR systems also fail in the presence of strong back-scatter. An approach to improve object detection at short distances is the fusion of LiDAR data onto camera images. By combining depth information from LiDAR with image information from cameras, the object detection can be enhanced.

[0051] According to an aspect of the invention, a method for detecting environmental conditions of an operational environment of an automated vehicle is provided that is based on a deep combination of gated camera images, gated camera capture parameters, and LiDAR images. The gated camera is designed for all conditions visibility. The design itself enables view in low-light conditions or in dark (e.g., in nighttime) , and it is designed to "see through" moderate amounts of fog, rain and snowfall. The optical data provided by the gated camera is different from the optical data obtained from an RGB camera. Although the gated camera in itself is designed to filter out these particles we are interested in, combining the gated camera view with the LiDAR provides us with a setup that can observe a change in all of the environmental conditions we are interested in. In other words, if one of the environmental conditions we are interested in changes, we will observe it in one of these two sensors. Gated camera capture parameters, or gating parameters, provide additional data about the image capture event, such as exposure and gating times, whether the lasers were used or not, and the power lasers were used at. This provides extra information about the environmental conditions, particularly about the lighting conditions. Detection in low or poor lighting conditions is improved compared to the detection that is based on the RGB vision.

[0052] The aspect of the invention enables the automated vehicle to be aware of the weather condition, road condition and lighting conditions using a single neural network model. Thereby, the automated vehicle is able to adapt its driving control schemes or algorithms to each condition.

[0053] Fig.4 shows a simplified functional block diagram illustrating an exemplary convolutional neural network (CNN) system for detecting environmental conditions of an operational environment of an automated vehicle according to principles of the present invention. The exemplary CNN system has two encoder channels with different inputs. A pre-trained convolutional encoder 204 is arranged to receive Lidar image stream 15 from a LiDAR sensor system, e.g., the LiDAR system 14 shown in Fig. 1. The encoder 204 acts as a feature extractor that encodes the Lidar image (extracts most salient features from the image) and outputs the extracted features in LiDAR feature maps or a Lidar feature vector. Another pre-trained encoder 202 is arranged to receive a gated camera image stream17-1 as one input and a set of gated camera capture parameters 17-2 as another input from a gated camera sensor, such as the gated camera sensor 16-1 shown in Fig. 1. The encoder 202 encodes the gated image and the gated camera capture parameters (extracts most salient features) and outputs the extracted features in gated-camera feature maps or a gated camera feature vector. Next, in general terms, the extracted Lidar and gated-camera features are combined, and the combined features are outputted to a classification stage. In the illustrated example, the gated-camera and Lidar feature maps or feature vectors outputted from encoders 202 and 204 are combined, and the combined extracted gated-camera and LiDAR features are outputted to the classification stage.

[0054] In the exemplary CNN system illustrated in Fig. 4, the encoder 202 is configured to flatten the gated-camera feature maps into and output a gated-camera feature vector 203, and the encoder 204 is configured to flatten the LiDAR feature maps into and output a LiDAR feature vector 205. Thereafter, in the illustrated exemplary embodiment, the outputted gated-camera and LiDAR feature vectors 203 and 205 are combined (e.g., concatenated or summed) in a combining block 206 to form a feature vector that is outputted as the combined feature vector 207 to the subsequent classification stage. Fully connected layers in the classification stage require a one-dimensional input, and the flattening is a process that converts the multi-dimensional feature maps into a one-dimensional vector.

[0055] In some embodiments, the encoder 202 may be configured to output gated-camera feature maps, and the encoder 204 is configured to output LiDAR feature maps, and the gated-camera feature maps and the LiDAR feature maps may be combined to form combined feature maps. The combined feature maps may then be flattened to form a combined feature vector that is outputted as the combined feature vector to the subsequent classification stage.

[0056] A deep convolutional encoder typically comprises several convolutional layers, as well as pooling layers, and possibly activation layers, as illustrated in a simplified block diagram shown in Fig. 5A. A first layer of a convolutional encoder is typically so-called input layer (not shown) which preprocesses or transforms the input data (of 1, 2, 3 or more input channels) into a numerical format in a 2 or 3-dimensional array suitable for the convolutional layers. A convolutional layer is configured to apply convolution operations to the input array (e.g., input image or a feature map from the previous convolution layer) using filters (or kernels) to detect features, such as edges, textures, and patterns. The filter is a two-dimensional (2-D) array of weights that represents the feature to be extracted, andit is also called a feature detector that represents part of the input array. The filter can vary in size NxM, which also determines the size of the resulting field. The filter size may be a 3x3 matrix. The filter is then applied to an area of the input array (e.g. image), and a dot product is calculated between the input array values and the filter. This dot product is then fed into an output array. Afterwards, the filter shifts by a stride, repeating the process until the kernel has swept across the entire input array. Stride is the distance, or number of cells (e.g. pixels), that the kernel moves over the input array; the stride is typically 1. The final output from the series of dot products from the input and the filter is referred to as a feature map. There can be several distinct filters which are used separately and provide different feature maps. The number of filters affects the depth of the output. For example, three distinct filters would yield three different feature maps, creating a depth of three. After each convolution operation, there may be an activation layer that introduces nonlinearity to the model, such as a Rectified Linear Unit (ReLU) transformation which replaces negative values with zeros in the feature maps. To further reduce the size of the feature map (i.e., the width and height of the map) generated from a convolution layer, a pooling (also referred to as subsampling) may be applied before further processing. Pooling is the process of summarizing the features within a group of cells in the feature map. This summary of cells can be acquired by taking the maximum, minimum, or average within a group of cells. Each of these methods is referred to as min, max, and average pooling, respectively. A pooling layer may be provided after each convolutional layer. Finally, the feature maps from the last convolution or pooling layer may be flatten into an output feature vector.

[0057] The subsequent classification stage is configured and pretrained to classify or detect one or more environmental conditions based on the combined extracted gated-camera and LiDAR features from the two encoders 202 and 204. In other words, the classification stage of the CNN makes the final decision about the class or category of the input data (image). Generally, the classification stage may assign probabilities to each class, and the class with the highest probability is chosen as the prediction. Fig. 5B shows a simplified functional block diagram illustrating basic elements of an exemplary classification stage. The classification stage receives a feature vector as an input. The classification stage may include fully connected layers of neural network so that every node in one layer is connected to every neuron in the next layer. The fully connected layers may abstract the feature vector into a decision vector of the same size as the number of classes. The last fully connected layer may be followed by an activation function (e.g., Softmax) thatconverts the output from the fully connected layers into class probabilities, i.e. assigns probabilities to each class, and the class with the highest probability is chosen as the class prediction (e.g., the class “Wet” in Fig. 5B). It shall be appreciated that the exact structure or operation of the classification stage is not essential to the primary invention, but it can be implemented in various ways which are readily obvious for a person skilled in the art upon reading the present application. In embodiments, the classification stage may include one or more pre-trained environmental condition classifiers, e.g., classifiers 210, 212, and 214 illustrated in Fig. 4. Each classifier 210, 212, and 214 may be pretrained to classify or detect one environmental condition. In the illustrated example, classifier 210 may output a natu-ral / ambient light state (e.g., daylight vs night), classifier 220 may output a rain state (e.g., no rain, rain, light rain, moderate rain, heavy rain, snowfall), and classifier 230 may output a road or terrain state (e.g., a dry road, a wet road, a snowy road, an icy road). The vehicle can be controlled based on the detected or classified environmental condition. In the illustrated example the detected or classified environmental condition or conditions may be provided to the vehicle control system 12 adapted to control various components and / or functions of the vehicle, such as a vehicle controller, brakes, engine, lighting, steering, etc.

[0058] Fig. 6 shows a simplified functional block diagram illustrating an exemplary implementation of the convolutional encoders 202 and 204 for deep combination of gated camera images, gated camera capture parameters, and LiDAR images. The convolutional encoder 204 is arranged to receive a 3D LiDAR image, having an array size 800x480x3 (distance, intensity, height), for example. The exemplary encoder 204 carries out four convolution and pooling operations or layers 204-1, 204-2, 204-3 and 204-4 one by one in that order. The first convolution and pooling layer 204-1 outputs 8 feature maps of size 33x54 that are inputted to the second convolution and pooling layer 204-2. The second layer 204-2 outputs 16 feature maps of size 167x27 that are inputted to the third convolution and pooling layer 204-3. The third layer 204-3 outputs 32 feature maps of size 84x14 that are inputted to the fourth convolution and pooling layer 204-4. The fourth layer 204-4 outputs 64 feature maps of size 42x7 that are flattened into a 1-dimensional LiDAR feature vector 205.

[0059] The encoder 202 is arranged to receive a gated camera image, having an array size 800x400x1 (width x height x depth), for example. The exemplary encoder 202 carries out six convolution and pooling operations or layers 202-1, 202-2, 202-3, 202-4, 202-5 and 202-6 one by one in that order. In accordance with theprinciples of the present invention, the exemplary encoder 202 further receives a set of gated camera capture parameters X and encodes them together with the gated camera image.

[0060] The gated camera parameter input contains information about the settings used for capturing the gated camera image by a gated camera sensor, and thereby about the ambient lighting and lighting conditions. In embodiments, a parameter array is received that represents a some or all the parameters of one single gated camera sensor unit, such as sensor unit 16-1 shown in Fig. 4. There may be other gated camera sensor units, such as 16-2...16-N, in the system, but they are not needed in the exemplary embodiments. Further, in the illustrated exemplary embodiment, the single gated camera sensor unit may be set to run in an autoexposure mode in which the camera unit automatically adjusts the exposure according to ambient lighting so as to achieve optimal exposure in varying lighting conditions. Thus, when the gated camera parameter input contains the autoexposure settings used for capturing the gated camera image by a gated camera sensor, the parameter values represent and change according to the prevailing ambient lighting conditions. An exemplary gated camera capture parameter array X for a single gated camera unit is illustrated in Table 1. The numerical values in the array are merely illustrating examples. In the exemplary embodiment, the single gated camera unit fires 4 light (laser) pulse sequences with different parameters to capture a single gated camera image. In Table 1, the parameter values of each fired laser sequence are presented on one line in the array, resulting in four lines of values.

[0061] Table 1.

[0062]

[0063] Columns in the Table 1 are following:

[0064] Series index i: Serial number of the fired laser sequence (line in the array) Xj,o laser on: Indicates whether the laser is active or not (1 or 0) Xj,i series control: 0: inactive; 1: laser unused, only ambient light exposure; 5: laser illumination + ambient light exposure, UNUSEDXi, 2 gate pulses: Number of gating pulses in the sequence. Higher number means more sensitive exposure

[0065] Xi, 3 laser pulses: Number of laser pulses in the sequence. Higher number means more laser illumination

[0066] x^laser time: Total time of laser illumination in the sequency. Higher number means more laser illumination

[0067] Xi, 5 gate time: Total gating on time in the sequence. Higher number means longer exposure and more sensitivity.

[0068] Gated camera capture parameter Xj,i is not used in the exemplary embodiment, but it could be used in some other embodiments. Each of the gated camera capture parameters Xj.o, Xj,2, Xj,3, Xj,4, andXi,s is inputted to the encoder 202 at different one of the several convolution layers 202-1, 202-2, 202-3, 202-4, 202-5 and 202-6. The parameter data from the parameter array X is fed column by column, meaning that each input xi nconsists of a parameter vector of length 4, i.e. vector size is 4x1. In other words, the entire column of array X with values for all 4 sequences is a single input vector xi n. In the example shown in Table 1 and Fig. 6, the input vector Xj,o includes “laser on” parameters of all 4 sequences i, namely Xj,o = (1, 1, 1, 1). Similarly, “gate pulses” parameter vector would be Xj,2 = (0, 485, 400, 115), a “laser pulses” parameter vector would be Xj,3 = (0, 485, 400, 115), a “laser time” parameter vector would be Xj,4 = (70, 70, 70, 70), and a “gate time” parameter vector would be Xj.s = (90, 125, 150, 160).

[0069] In the example shown in Fig.6, the input vector Xj,o is fed to the encoder 202 at the input of the convolution layer 202-2, the parameter vector Xj,2 is fed at the input of the convolution layer 202-4, the parameter vector Xj,3 is fed at the input of the convolution layer 202-5, the parameter vector Xj,4 is fed at the input of the convolution layer 202-5 , and the parameter vector Xj.s is fed at the output of the encoder.

[0070] In embodiments, each parameter vector xi nis combined with the feature maps outputted from the preceding convolution layer, and the combined feature maps are inputted to the next convolution layer, or to the encoder output. The combining operations, e.g. summing, are depicted by sum operators 201-1, 201-2, 201-3, 201-4, and 201-5 in the example shown in Fig. 6.

[0071] In embodiments, each parameter vector xi nis expanded to match the dimensions of the image convolution feature at each insertion point. In the example shown in Table 1 and Fig. 6, each parameter vector xi nhaving the size 4x1 must be expanded to the size of the feature maps it is combined with. For example, Xi Qmaybe expanded to the size 400x240x8, Xi2to 100x60x32, Xi 3to 50x30x64, Xi4to 25x15x128, and Xi 5to 12x7x256.

[0072] In embodiments, the expanding may be performed by mapping gated camera parameter values of each parameter vector Xi,nof size 4x1 to output values of an expanded meta data array having the required convolution dimension that can be can be combined with the corresponding convolution feature map. For example, gated camera parameter values of the 4x1 parameter vector may be mapped to values in a 100x60x32 meta data array. In embodiments, the mapping may be a linear layer (e.g., fully connected layer) having a gated camera parameter vector as an input and a meta data array of a predetermined size as an output. The linear layer mapping may be trained..

[0073] Referring to Fig. 7, let us examine an example of expanding or mapping the parameter vector X)2from size 4x1 to size 100x60x32. Firstly (step 71), a fully connected (linear) layer mapping of the parameter vector Xi 2of size 4x1 into a parameter vector f=

[0074]

[0075] ... f30f31of length 32 is carried out. This matches the third dimension of intermediate network output of size 100x60x32 at the convolution layer 202-4. The fully connected layer may comprise 32 neurons with trained weights. Then (step 72), the newly acquired vector f of length 32x1 is expanded to a matrix of shape 100x60x32 by cloning the values of the vector f. For example, in the new expanded matrix, the first channel (having 100 rows and 60 columns) may be filled with the first value fo of the vector f, second channel is filled with second value fr, etc. Now the sizes or shapes of the feature maps and the expanded parameter vector match, and we can simply element-wise sum the intermediate encoder output at the layer 202-4 and the expanded parameter vector in sum operation 201-2.

[0076] The combined feature maps at the output of the encoder 202 may be flattened into a 1-dimensional feature vector 203 that may then be combined 206 with the 1-dimensional feature vector 205 to provide the combined (e.g., concatenated or summed) feature vector 207. In an exemplary embodiment, the feature vectors 203 and 205 may be of equal size and directly summed element by element to form a summed feature vector 207. The classification stage may then provide one or more classifications as described above.

[0077] Training of the exemplary CNN system may be carried out in various ways. For example, the CNN system can be trained using an Adam optimizer [] with an initial learning rate of 0.001. The Adam optimizer is disclosed e.g. in Adam: A Method For Stochastic Optimization, D.P. Kingma et al, 3rd InternationalConference for Learning Representations, San Diego, 2015,

[0078]

[0079] The loss function can be a triplet loss observing the output of each classification head and giving each of them the same weight in the loss. A cross entropy loss [see for example, https: / / pytorch. org / docs / siable / gener-

[0080]

[0081] can be applied individually to each classification output, and they their evenly weighted sum can be used as the loss value:

[0082] loss = 0.333 * lossjighting + 0.333 * loss_terrain + 0.333 * loss_rain The encoding, classification and control functions described herein may be implemented by various means. For example, these techniques may be implemented in hardware [one or more devices), firmware [one or more devices), software [one or more modules), or combinations thereof. For a firmware or software, implementation can be through modules [e.g., procedures, functions, and so on) that perform the functions described herein. The software codes may be stored in any suitable, processor / computer-readable data storage medium[s) or memory unit[s) and executed by one or more processors / computers. The data storage medium or the memory unit may be implemented within the processor / computer or external to the processor / computer, in which case it can be communicatively coupled to the processor / computer via various means as is known in the art. Additionally, components of systems described herein may be rearranged and / or complimented by additional components in order to facilitate achieving the various aspects, goals, advantages, etc., described with regard thereto, and are not limited to the precise configurations set forth in a given figure, as will be appreciated by one skilled in the art.

[0083] The description and the related drawings are only intended to illustrate the principles of the present invention by means of examples. Various alter-native embodiments, variations and changes are obvious to a person skilled in the art on the basis of this description. The present invention is not intended to be limited to the examples described herein but the invention may vary within the scope and spirit of the appended claims.

Claims

CLAIMS1. A method for detecting environmental conditions of an operational environment of an automated vehicle, comprisingreceiving Lidar image from a LiDAR sensor system,encoding the Lidar image into first feature vector by a first pre-trained convolutional neural network, CNN,receiving gated image and gated camera capture parameters from a gated camera system,encoding the gated image and the gated camera capture parameters into a second feature vector by a second pre-trained CNN,combining the first and second feature vectors into a combined feature vector,inputting the combined feature vector to two or more pre-trained environmental condition classifiers, each classifier being configured to classify one environmental condition, andcontrolling a vehicle based on a classified environmental condition.

2. The method as claimed in claim 1, wherein the second pre-trained CNN comprises several convolution layers, and wherein the method further comprises inserting the gated camera capture parameters each to different one of the several con-volution layers in the second pre-trained CNN.

3. The method as claimed in claim 2, whereinthe receiving of the gated camera capture parameters comprises receiving a gated camera parameter array including the gated camera capture parameters for each of n laser pulse series fired by the gated camera for each gated camera image, and forming a gated camera capture parameter vector of length n and including the same parameter of each of the n laser pulse series, and wherein each of the gated camera capture parameter vectors is inserted to different one of the several convolution layers in the second pre-trained CNN.

4. The method as claimed in claim 3, wherein the inserting comprises expanding the gated camera capture parameter vector of length n to dimensions that match to dimensions of convolution features at the respective convolution layer which the gated camera capture parameter vector is inserted to.

5. The method as claimed in claim 4, wherein the expanding comprises linear mapping of parameter values of the respective gated camera capture parameter vector to values in the matching dimension.

6. The method as claimed in any one of claims 2-5, wherein the insertingcomprises summing the gated camera capture parameters to the convolution features at the respective convolution layer.

7. The method as claimed in any previous claim, wherein the gated camera capture parameters include one or more of the following: laser is active or not, number of gate pulses, number of laser pulses, total time of laser illumination, total time of gate on.

8. The method as claimed in any previous claim, wherein the gated camera capture parameters include information on the autoexposure settings used by the gated camera system.

9. The method as claimed in any previous claim, wherein the controlling the vehicle based on a classified environmental condition comprises adapting driving control schemes or algorithms of the vehicle according to the classified environmental condition.

10. The method as claimed in any previous claim, wherein the classified environmental condition is notified to a driver or passenger of the vehicle, preferably displayed on a user interface of the vehicle.

11. The method as claimed in any previous claim, wherein the controlling the vehicle based on a classified environmental condition comprises controlling one or more components and / or functions of the vehicle, preferably including one or more of a vehicle controller, brakes, engine, lighting, steering.

12. The method as claimed in any previous claim, wherein the classified environmental condition comprises one or more of: lighting condition, a rain state, a road or terrain state.

13. A control system configured to implement the steps of the method as claimed in any one of claims 1-12 for detecting and classifying environmental conditions of an operational environment of an automated vehicle.

14. A vehicle, comprisinga gated camera system,a long-range LiDAR system,a control system controlling an operation of the vehicle, and means for implementing the steps of the method as claimed in any one of claims 1-12.

15. An apparatus comprising at least one processor and at least one non-transitory memory including computer program code instructions, the computer program code instructions being configured to, when executed, cause the apparatus to carry out the steps of the method as claimed in any one of claims 1-12.