Method for providing free space estimation using motion data

By applying optical flow technology and depth map projection methods in image sequences, the problems of high computational cost and time noise in the prior art are solved, and efficient free space estimation is realized when executed online within the vehicle.

CN120198467APending Publication Date: 2025-06-24ZENSEACT AB
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411892872.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-12-22
Filing Date
2024-12-20
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The prior art, when using motion data to perform free space estimation, has high computational costs and is prone to introduce time noise, making it difficult to achieve ideal performance when performed online within a vehicle.

Method used

The motion data of the 3D point is determined by applying optical flow techniques in the image sequence and then projecting it into the 3D space, combining the depth map to determine the motion data of the 3D point and allocating it to the free space estimation.

Benefits of technology

This method finds a balance between performance and computational cost, reducing temporal noise in motion data, providing more accurate free space estimation, and reducing computational complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198467A_ABST
    Figure CN120198467A_ABST
Patent Text Reader

Abstract

The invention relates to a method for providing free space estimation using motion data. A computer-implemented method (100) performed in a vehicle equipped with an automated driving system comprises: obtaining (S102) a sequence of images captured by an image capture device of the vehicle, wherein the sequence of images comprises a plurality of images for depicting a scene at a respective one of a plurality of time instances; obtaining (S104) a set of 3D points based on a depth map of the scene depicted in the sequence of images, wherein each 3D point in the set of 3D points is associated with a three-dimensional position of the 3D point within the scene; determining (S106) motion data associated with each 3D point of the set of 3D points, wherein the motion data indicates an estimated motion of an object in the scene associated with the 3D point; and assigning (S116) a set of 3D points with associated motion data to a free-space estimate of the scene based on the three-dimensional position associated with each 3D point.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present inventive concept relates to the field of autonomous vehicles. In particular, the present inventive concept relates to methods and apparatuses for providing free space estimation using motion data. Background Art

[0002] In recent years, with the development of technology, image capture and processing technologies have been widely applied in different technical fields. In particular, vehicles produced today are usually equipped with some form of vision system or perception system to enable new functions. In addition, more and more modern vehicles have advanced driver assistance systems (ADAS) to improve vehicle safety and, more generally, road safety. ADAS (which can be represented, for example, by adaptive cruise control (ACC), collision avoidance systems, forward collision warnings, lane support systems, etc.) are electronic systems that can assist the driver of a vehicle. Today, research and development are being carried out in numerous technical fields associated with both the ADAS field and the autonomous driving (AD) field. ADAS and AD can also be collectively referred to as autonomous driving systems (ADS), and ADS systems correspond to all different levels of automation as defined, for example, by the SAE J3016 levels (0 - 5) of driving automation.

[0003] One of the challenges faced by autonomous or semi - autonomous vehicles (i.e., vehicles that support ADS) is the ability to accurately detect their surrounding environment and navigate within it. To achieve this ability, a vehicle needs to be able to detect and evaluate the free space around it, which means the drivable area in the environment that is not occupied by any obstacles. This is important for both ensuring safety and providing efficient navigation. Free space estimation is generally an instantaneous static view of the world obtained from sensor data collected by the vehicle, without dynamic information. Since the driving environment is highly dynamic, it is beneficial to additionally have information related to the speed of objects in the vehicle's surrounding environment.

[0004] Deriving free space estimation using speed information is not an easy task and is a task that must be performed online within the vehicle, using limited computational resources while still achieving desirable performance. Therefore, there is a need for a new and improved solution in this field. Summary of the Invention

[0005] The techniques disclosed herein seek to alleviate, mitigate, or eliminate one or more of the above - identified deficiencies and drawbacks in the prior art to address various problems related to free space estimation. More specifically, the techniques of the present disclosure provide techniques for estimating the motion of objects in the vehicle's surrounding environment for free space estimation and for improving free space estimation.

[0006] The inventors have achieved a new and improved method for accomplishing this goal, which provides a balance between performance and computational cost. The techniques of the present disclosure are at least partially based on the concept of using visual changes in optical flow in two-dimensional images for motion estimation and combining it with projection into three dimensions for subsequent use in free space estimation. Prior methods of combining motion data with free space estimation were based on motion estimation of 3D points (e.g., lidar point clouds) by calculating the rate of change of the 3D points' own past positions or by tracking how the individual elements of the free space estimation move over time. However, this not only incurs a high computational cost, but also introduces temporal noise between the 3D point position estimates at different time instances when calculating the rate of change of the historical positions. The techniques disclosed herein can provide improvements in both aspects.

[0007] Aspects and embodiments of the disclosed invention are defined in the following and the appended independent and dependent claims.

[0008] According to a first aspect, there is provided a computer-implemented method performed in a vehicle equipped with an autonomous driving system. The method includes obtaining a sequence of images captured by an image capture device of the vehicle. The sequence of images includes a plurality of images for depicting a scene at respective time instances among a plurality of time instances. The method further includes obtaining a set of 3D points based on a depth map of the scene depicted in the sequence of images. Each 3D point in the set of 3D points is associated with a three-dimensional position of the 3D point in the scene. The method further includes: determining motion data associated with each 3D point in the set of 3D points. The motion data indicates an estimated motion of an object in the scene associated with the 3D point. The motion data associated with each 3D point is determined by the following steps: obtaining a 2D point corresponding to the 3D point in the image plane of an image in the sequence of images; applying an optical flow between the image and a subsequent image in the sequence of images to the 2D point, thereby obtaining a subsequent 2D point in the image plane of the subsequent image; determining a subsequent 3D point by projecting the subsequent 2D point based on the depth map of the scene; and determining the motion data based on a difference between the three-dimensional positions of the 3D point and the subsequent 3D point. The method further includes: assigning the set of 3D points with associated motion data to a free space estimation of the scene based on the three-dimensional position associated with each 3D point.

[0009] As mentioned above, the disclosed technology can lie in that it provides a computationally advantageous method that is still capable of achieving the desired performance. It can further provide improvements in real-time processing and a reduction in computational complexity. More specifically, the image sequence constitutes a temporally smoother source for estimating motion (compared to the current motion estimation method for free space estimation). Therefore, by using the 2D image pixel motion on the image sequence to determine the motion data, any temporal noise in the estimated motion data can be reduced. Later, projecting the estimated motion data into 3D by using depth estimation can provide a more accurate fusion with free space estimation.

[0010] Another possible advantage of some embodiments lies in that it utilizes off-the-shelf technologies such as optical flow, making its implementation and execution in vehicles equipped with ADS more efficient. It can be achieved by reusing the dedicated hardware or common software libraries that already exist in the vehicle.

[0011] Another possible advantage of some embodiments lies in that by introducing constraints on the estimated motion, the motion estimation can be further improved. The constraint can be to include only motion parallel to the estimated ground plane. The reason for this is that by applying the assumption that all objects in the surrounding environment typically move in a direction parallel to the ground plane, the possible noise originating from 3D depth estimation can be further reduced. Utilizing this constraint, the motion estimation can obtain further robustness and balance the 3D depth estimation noise by effectively filtering past depth estimates to be closer to the current depth estimate.

[0012] According to a second aspect, there is provided a computer program product comprising instructions which, when the program is executed by a computing device, cause the computing device to perform the method according to any embodiment of the first aspect.

[0013] According to a third aspect, there is provided a (non-transitory) computer-readable storage medium. The non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a processing system, the one or more programs including instructions for performing the method according to any embodiment of the first aspect. When applicable, any of the above-mentioned features and advantages of the first aspect also apply to the second and third aspects. To avoid unnecessary repetition, reference may be made to the above content.

[0014] According to a fourth aspect, there is provided an apparatus including a control circuit. The control circuit is configured to obtain an image sequence captured by an image capturing device of a vehicle. The image sequence includes a plurality of images for depicting a scene at respective time instances among a plurality of time instances. The control circuit is further configured to obtain a set of 3D points based on a depth map of the scene depicted in the image sequence. Each 3D point in the set of 3D points is associated with a three-dimensional position of the 3D point in the scene. The control circuit is further configured to: determine motion data associated with each 3D point in the set of 3D points. The motion data indicates an estimated motion of an object in the scene associated with the 3D point. The motion data associated with each 3D point is determined by the following steps: obtaining a 2D point corresponding to the 3D point in the image plane of an image in the image sequence; applying an optical flow between the image and a subsequent image in the image sequence to the 2D point, thereby obtaining a subsequent 2D point in the image plane of the subsequent image; determining a subsequent 3D point by projecting the subsequent 2D point based on the depth map of the scene; and determining the motion data based on a difference between the three-dimensional positions of the 3D point and the subsequent 3D point. The control circuit is further configured to: assign the set of 3D points with associated motion data to a free space estimation of the scene based on the three-dimensional position associated with each 3D point. When applicable, the above-mentioned features and advantages of the foregoing aspects also apply to this fourth aspect. To avoid unnecessary repetition, reference may be made to the above content.

[0015] According to a fifth aspect, there is provided a vehicle equipped with an autonomous driving system. The vehicle includes an image capturing device. The vehicle further includes an apparatus according to any embodiment of the fourth aspect. When applicable, the above-mentioned features and advantages of the foregoing aspects also apply to this fifth aspect. To avoid unnecessary repetition, reference may be made to the above content.

[0016] As used herein, the term "non-transitory" is intended to describe a computer-readable storage medium (or "memory") that excludes propagating electromagnetic signals, but is not intended to otherwise limit the type of physical computer-readable storage devices covered by the term "computer-readable medium or memory". For example, the term "non-transitory computer-readable medium" or "tangible memory" is intended to cover types of storage devices that do not necessarily store information permanently, including for example random access memory (RAM). Program instructions and data stored in a non-transitory form on a tangible computer-accessible storage medium may further be transmitted via a transmission medium or signals such as electrical, electromagnetic, or digital signals, which may be conveyed via communication media such as a network and / or a wireless link. Thus, the term "non-transitory" as used herein is a limitation on the medium itself (i.e., tangible, rather than a signal), rather than a limitation on data storage persistence (e.g., RAM versus ROM).

[0017] The disclosed aspects and preferred embodiments can be appropriately combined with each other in any manner that would be obvious to any person of ordinary skill in the art, such that one or more features or embodiments disclosed with respect to one aspect can also be considered to be disclosed with respect to another aspect or an embodiment of another aspect.

[0018] Further embodiments are defined in the dependent claims. It should be emphasized that, when used in this specification, the term "comprising / including" is used to indicate the presence of the recited features, integers, steps or components. It does not preclude the presence or addition of one or more other features, integers, steps, components or groups thereof.

[0019] These and other features and advantages of the disclosed technology will be further clarified with reference to the embodiments described in the following context. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] When combined with the drawings, the above aspects, features and advantages of the disclosed technology will be more fully understood by reference to the following illustrative and non - limiting detailed description of example embodiments of the present disclosure, in which:

[0021] Figure 1 is a schematic flowchart representation of a method according to some embodiments.

[0022] Figure 2 is a schematic illustration of an apparatus according to some embodiments.

[0023] Figure 3 is a schematic illustration of a vehicle according to some embodiments.

[0024] Figure 4A The surrounding environment of the vehicle is illustrated in perspective view by way of example.

[0025] Figure 4B The surrounding environment of the vehicle is illustrated in side view by way of example. DETAILED DESCRIPTION

[0026] The present technology will now be described in detail with reference to the drawings, which show some example embodiments of the disclosed technology. However, the disclosed technology can be embodied in other forms and should not be construed as limited to the disclosed example embodiments. The disclosed example embodiments are provided to fully convey the scope of the disclosed technology to those skilled in the art. Those skilled in the art will understand that the steps, services and functions explained herein can be implemented using separate hardware circuits, using software running in conjunction with a programmed microprocessor or a general - purpose computer, using one or more application - specific integrated circuits (ASICs), using field - programmable gate arrays (FPGAs) and / or using one or more digital signal processors (DSPs).

[0027] It will also be understood that when the present disclosure is described in the form of a method, the present disclosure can also be embodied as a device including one or more processors and one or more memories coupled to the one or more processors, in which computer code is loaded to implement the method. For example, in some embodiments, the one or more memories may store one or more computer programs that, when executed by the one or more processors, cause the device to perform the steps, services, and functions disclosed herein.

[0028] It should also be understood that the terms used herein are for the purpose of describing particular embodiments only and are not intended to be limiting. It should be noted that, as used in this specification and the appended claims, the articles "a", "an", "the", and "said" are intended to mean that there is one or more elements, unless the context clearly dictates otherwise. Thus, for example, in certain contexts, a reference to "a unit" or "the unit" may refer to more than one unit, and so on. Furthermore, the words "comprising", "including", "containing" do not exclude other elements or steps. It should be emphasized that when used in this specification, the term "comprising / including" is used to specify the presence of the recited features, integers, steps, or components. It does not exclude the presence or addition of one or more other features, integers, steps, components, or groups thereof. The term "and / or" should be construed to mean "both" and each as an alternative.

[0029] It will also be understood that although the terms "first", "second", etc. may be used herein to describe various elements or features, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, the first element may be referred to as the second element, and similarly, the second element may be referred to as the first element, without departing from the scope of the various embodiments. Both the first element and the second element are elements, but they are not the same element.

[0030] As used herein, the phrase "one or more" in a group of elements (such as "one or more of A, B, and C" or "at least one of A, B, and C") should be construed as conjunctive logic or disjunctive logic. In other words, it may refer to all elements, one element, or a combination of two or more elements in a group of elements. For example, the phrase "one or more of A, B, and C" may be construed as A or B or C, A and B and C, A and B, B and C, or A and C.

[0031] The technology of the present disclosure generally relates to the aggregation of motion data and free space estimation in a vehicle equipped with an autonomous driving system (ADS). Now reference will be made to Figures 1 to 4B to describe the methods, devices, and vehicles capable of implementing the technology of the present disclosure.

[0032] In the context of an autonomous (or semi-autonomous) vehicle, "free space estimation" refers to the process of determining the availability and characteristics of unoccupied areas within the vicinity of the vehicle. This estimation can be used in the vehicle's navigation and decision-making systems to ensure safe and efficient operation.

[0033] Thus, "free space" can be understood as the portion of a road (or any road-like area) that is not occupied by any other object and is thus "free". In other words, free space refers to the area in the vehicle's surroundings that is free of obstacles such as other vehicles, pedestrians, buildings, etc. Thus, free space can be regarded as the portion of a road (e.g., the vehicle's lane, or an adjacent lane in the same direction of travel) that is not occupied by any other object and on which the vehicle can and / or is permitted (e.g., considering safety requirements or traffic rules) to drive. This can also be referred to as "drivable free space". In other words, in some embodiments, "free space" can be interpreted as "drivable free space".

[0034] Free space estimation is typically determined online in the vehicle as the vehicle travels around to evaluate the surrounding environment and any objects therein. The goal can be to create a detailed map or representation of the unoccupied space, taking into account factors such as the distance, size, and geometry of the various objects. This process typically involves collecting sensor data from sensors such as lidar, radar, cameras, or other sensing devices. The collected sensor data is then processed by various algorithms to distinguish between free space and non-free space (i.e., obstacles). Some traditional free space estimation or determination methods rely on using different sensors of the vehicle to determine free space in a three-dimensional representation of the world. One existing method is to generate a depth image of the surrounding environment and thereby determine the free space that appears to belong to a flat surface (which is the ground). Another method involves using a machine learning model trained for image segmentation.

[0035] In some cases, "undrivable free space" can be used to refer to the portion of a road that is free but on which vehicles are not permitted to drive. Some examples of undrivable free space include oncoming lanes, bike lanes, closed lanes, gravel shoulders, and grass shoulders. Thus, in some embodiments, free space can include both drivable free space and undrivable free space. In some embodiments, when "free space" is used to refer to the space described above as "drivable free space", the space described above as "undrivable free space" can be considered "non-free space". In other words, undrivable free space can be part of non-free space.

[0036] The estimated "free space" can be divided into "drivable free space" and "non-drivable free space" by some further steps, such as using map data or analyzing image or other types of sensor data to determine which traffic rules apply in this situation, or what type of area the space is (i.e., whether it is a lane, shoulder, safety island, etc.). For example, image analysis techniques can be used to determine which are the roads and which are the non-drivable areas beside the roads. Such steps can be part of some embodiments, as will be further explained below.

[0037] The area corresponding to the free space can be represented in different ways, such as a list of points describing the boundary between the drivable space and the non-drivable space, or an occupancy grid in the ground plane including a plurality of grid cells, where each cell carries some information about the driving performance of that particular cell (e.g., the probability that the cell is free / unoccupied).

[0038] Free space estimation can also consider dynamic elements such as moving objects to adapt to changes in the surrounding environment in real time. In particular, free space estimation can include the motion data (e.g., speed) of any object in the surrounding environment. This motion data provides information about the estimated motion behavior of the object. If assigned to free space estimation, this motion data can, for example, indicate the dynamic behavior of any free space boundary. This information, in turn, can be used to distinguish between static and dynamic objects, slow and fast moving, or objects moving towards or away from the vehicle, and can subsequently be used during trajectory or braking decisions. For example, a free space boundary corresponding to a vehicle located in front of the ego vehicle and having a similar relative motion to the ego vehicle may allow for a softer braking, while a free space boundary corresponding to a decelerating vehicle, a stationary vehicle (or any other object), or even a vehicle moving head-on towards the ego vehicle may trigger a more severe braking. Thus, in addition to the position information of the free space, this additional information about the dynamic behavior can also provide improvements, for example, in terms of functionality, safety, comfort, and capabilities of any subsequent tasks that utilize free space estimation.

[0039] Figure 1 is a schematic flowchart representation of method 100 according to some embodiments. More specifically, method 100 can be a method 100 for providing a free space estimation using the estimated motion data. In other words, method 100 can be a method 100 for estimating free space using motion data in a vehicle equipped with ADS, or a method 100 for assigning motion data to free space estimation.

[0040] Next, the different steps of method 100 are described in more detail. Even though shown in a specific order, the steps of method 100 can be performed in any suitable order and multiple times. Thus, although Figure 1 a specific order of method steps may be shown, the order of these steps can be different from the depicted order. Additionally, two or more steps can be performed simultaneously or partially simultaneously. For example, the steps represented as S102 and S104 can be performed in any order or at any point in time based on a particular implementation. Such variations will depend on the selected software and hardware systems as well as the designer's choices. All such variations are within the scope of the present invention. Similarly, the software implementation can be accomplished using standard programming techniques with rule-based logic and other logic to perform the various steps. Other variations of method 100 will be apparent in light of the present disclosure. The embodiments mentioned and described above are given only as examples and should not limit the present invention. Other solutions, uses, purposes, and functions within the scope of the present invention as claimed in the patent claims described below should be apparent to those skilled in the art. It should be understood that Figure 1 the steps included in dashed lines in

[0041] are examples of multiple optional steps that can be part of multiple alternative embodiments. These optional steps do not need to be performed in order. Additionally, it should be understood that not all steps need to be performed. The example steps can be performed in any order and in any combination.

[0042] Method 100 includes obtaining (S102) an image sequence captured by an image capture device of a vehicle. The image sequence includes multiple images for depicting a scene at a corresponding time instance among multiple time instances. The image sequence (which can also be referred to as a sequence of images or a series of images) can be regarded as a stream of image frames of a scene (e.g., a video stream). The image sequence can include at least two images for depicting a scene at a corresponding time instance. More specifically, the image sequence can include a first image for depicting a scene at a first time instance and a second image for depicting a scene at a second time instance. The first time instance and the second time instance are two consecutive time instances.

[0043] The term "obtain" is broadly interpreted throughout this disclosure and encompasses receiving, retrieving, collecting, acquiring, etc., directly and / or indirectly between two entities configured to communicate with each other or further with other external entities. However, in some embodiments, the term "obtain" should be interpreted as determining, deriving, forming, calculating, etc.

[0044] For example, an image sequence can be directly received from an image capture device. Alternatively or in combination, the image sequence can be received from a memory (e.g., an intermediate storage), where the image sequence captured by the image capture device is stored.

[0045] A scene can be understood as a part of the vehicle's surrounding environment depicted in the images of the image sequence. In other words, a scene can be understood as at least a part of the vehicle's surrounding environment. Thus, a scene includes any object around the vehicle. The vehicle's surrounding environment should be understood as the general area around the vehicle, in which objects (such as other vehicles, landmarks, obstacles, etc.) can be detected and recognized by vehicle sensors (radar, lidar, cameras, etc.), i.e., detected and recognized within the sensor range of the vehicle.

[0046] Method 100 further includes obtaining (S104) a set of 3D points based on the depth map of the scene depicted in the image sequence. In other words, the 3D points in the set of 3D points can be defined by the depth map. Each 3D point in the set of 3D points is associated with the three-dimensional position of the 3D point within the scene. The three-dimensional position can be a position within the reference frame of the vehicle or the vehicle's image capture device. Thus, the three-dimensional position can be in a 3D coordinate space.

[0047] The 3D points can also be referred to as reference points. Each 3D point can correspond to a sub-part of the image, e.g., one pixel or a subset of pixels in the image. Additionally, each 3D point can be associated with an object in the scene (i.e., an object depicted in the images of the image sequence). It should be noted that one or more 3D points in the set of 3D points can be associated with different parts of the same object.

[0048] Once obtained, a set of numerical metrics (such as positions in the image plane) can be used in combination with depth estimation to define the 3D points. Alternatively or in combination, the 3D points can be defined by three coordinates in a 3D reference frame. Then, the identification and tracking of the 3D points can allow the tracking of the objects associated with the 3D points over time.

[0049] The depth map should be interpreted herein as a representation of the depth of an image or an image sequence. Thus, the depth map can be regarded as adding a third dimension (to the two existing dimensions of the 2D image plane). Thus, the depth herein refers to the distance between the image capture device and the depicted object.

[0050] The depth map may include depth values for each pixel of the image or for a subset of the pixels, which depth values indicate how far the corresponding object is from the camera. For the purposes of the present technology, the depth map includes at least depth information of a set of 3D points. Thus, the depth map may provide information related to the spatial relationships and distances between different objects in the depicted scene. Determining the depth data may be accomplished using any conventional technique, such as stereo vision using a stereo camera setup, lidar, radar, computer vision algorithms, or neural networks.

[0051] For example, the depth map may be based on the output of a machine learning model that is configured to determine the depth map based on a sequence of images as input. Thus, the depth map may be obtained as the output of a machine learning model supplied with the sequence of images or its images, and thus a set of 3D points is obtained. Thus, the step of obtaining (S104) the set of 3D points may include generating the set of 3D points by supplying the sequence of images to the machine learning model. The machine learning model may be trained using the images and optionally in combination with lidar data.

[0052] Alternatively or in combination, the depth map may be based on the lidar point cloud of the scene. In other words, the set of 3D points may correspond to the lidar point cloud of the scene. Thus, when the vehicle is traveling along the road, by collecting relevant sensor data (i.e., the images captured by the image capture device and the lidar point cloud captured by the vehicle's lidar sensor), the step of obtaining (S102) the sequence of images and the step of obtaining (S104) the set of 3D points may be performed simultaneously.

[0053] It should be understood that combinations of different techniques for obtaining the depth map may be used. For example, stereo images may be used in combination with the techniques described above to provide further signals regarding the depth of the scene.

[0054] Method 100 further includes determining (S106) motion data associated with each 3D point in the set of 3D points. The motion data indicates the estimated motion of the object in the scene associated with the 3D point. The motion data may include information related to the speed of the object in the scene. In other words, the motion data may be speed data. The speed data may indicate the speed of the object in the scene. Alternatively or in combination, the motion data may include information related to the acceleration of the object in the scene. The motion data may include information regarding the direction of motion and / or the magnitude of the motion (such as a speed value or an acceleration value).

[0055] In the broadest sense, motion data associated with each 3D point can be determined based on the motion of an object in a scene determined in a two-dimensional manner from subsequent images in an image sequence. In other words, the motion data can be determined by the two-dimensional optical flow between at least two subsequent images. More specific examples of how the motion data can be determined will be given below.

[0056] For each 3D point in the set of 3D points, motion data (S106) associated with the 3D point is determined by the following steps represented as S108 to S114. First, a 2D point corresponding to the 3D point in the image plane of the images in the image sequence is obtained (S108). This 2D point should be understood in this context as a representation of the 3D point, but in two dimensions. More specifically, the 2D point can be represented by an x coordinate and a y coordinate in the image plane of the image. The image plane refers to the 2D coordinate system spanned by the two axes of the 2D image. In other words, it is the 2D space that represents the visual information of the image.

[0057] Next, the optical flow between the image in the image sequence and a subsequent image is applied (S110) to the 2D point. In other words, the image motion vector determined for the image and the subsequent image can be applied to the 2D point. The optical flow or image motion vector can be interpreted in this context as information for describing the motion of an object in a scene between subsequent images in the image sequence. By applying the optical flow to the 2D point, a subsequent 2D point in the image plane of the subsequent image can be obtained. The subsequent 2D point should be understood in this context as the 2D point in the subsequent image corresponding to the (previous) 2D point mentioned above. The 2D point and the subsequent 2D point can be considered to correspond to each other in a sense because they are both associated with the same object in the scene. Therefore, applying the optical flow (or image motion vector) can be regarded as transforming the 2D point from the image to the subsequent image.

[0058] The optical flow in this article refers to a technique commonly used to describe image motion. It is typically applied to a series (or sequence) of images that have a small time step between them (such as video frames). The optical flow calculates the motion (e.g., velocity) for points within the image and provides an estimate of where these points may be located in the next image sequence. More specifically, it can use two consecutive camera images to match pixel points or groups of pixels between the images, and then provide a motion vector indicating the amount and direction of motion in the image based on how the groups of pixels move. This can be done by locally searching for the best match (e.g., by looking at the intensity in the image). Given the displacement (i.e., direction and distance) of a group of pixels and knowing the time frame between the images, the velocity can be calculated. To reduce the number of calculations required, a maximum displacement of the points can be set. Then, the calculated optical flow between the two images can be represented by a vector field across the image plane. The resolution of the vector field can be coarser than the original image because groups of pixels can be matched instead of individual pixels to reduce the computational cost. Therefore, applying (S110) the optical flow to 2D points can be precisely expressed as applying (S110) the optical flow vector field to 2D points. As is easily understood, there are several methods for implementing optical flow calculations, and any of these methods can be applied to this technology.

[0059] Next, the subsequent 3D points are determined (S112) by projecting the subsequent 2D points based on the depth map of the scene. In other words, the depth map can be used to project or transform the subsequent 2D points into three dimensions. Therefore, the subsequent 3D points can be determined by adding a third coordinate (e.g., the z coordinate) to the subsequent 2D points to describe the depth in the image. In other words, the subsequent 2D points can be projected from the 2D image plane into the 3D coordinate space. The resulting subsequent 3D points can be regarded as the 3D points in the subsequent image, corresponding to the 3D points in the above-mentioned (previous) image.

[0060] Next, the motion data is determined (S114) based on the difference between the three-dimensional positions of the 3D points and the subsequent 3D points. In other words, the three-dimensional motion of the object associated with the 3D points and the subsequent 3D points can be determined.

[0061] In some embodiments, determining (S106) (or the step represented as S114) the motion data further includes obtaining an estimate of the ground plane of the scene depicted in the image sequence. The estimate of the ground plane (or ground plane estimate) can be obtained by any suitable known technique for estimating a plane in one or more images, including but not limited to, stereovision techniques, lidar point clouds, and machine learning-based techniques. The motion data can then be determined (S106) as the motion parallel to the estimated ground plane. For example, in the above step represented as S114, the difference between the three-dimensional positions of a 3D point and a subsequent 3D point can be determined as the difference in a plane parallel to the estimated ground plane. Requiring the motion to be parallel to the estimated ground plane can serve as a constraint on the determined motion data. By restricting the motion data to the motion in a plane parallel to the estimated ground plane, the noise in the determined motion data can be reduced.

[0062] In some embodiments, the motion data is used to indicate an estimated motion of an object in the scene associated with the 3D points relative to the vehicle. Thus, information about whether and how the object moves relative to the vehicle can be obtained. This can be advantageous because it may require less complex calculations since the motion of the vehicle itself does not need to be considered.

[0063] In some embodiments, the motion data is further determined (S106) based on the vehicle motion data of the vehicle. In other words, when determining the motion data of the 3D points in the set of 3D points associated with an object in the scene, the motion of the vehicle can be considered. The motion data can then further indicate an estimated motion of the object in the scene associated with the 3D reference points relative to the ground.

[0064] The vehicle motion data of the ego vehicle can include information related to the angular (or turning) motion of the vehicle (i.e., yaw, pitch, and / or roll). The vehicle motion data can further include information related to the translational motion (e.g., forward motion and backward motion).

[0065] By determining the motion relative to the ground, it is allowed to distinguish dynamic objects from static objects in the scene. A dynamic object herein refers to an object that has motion relative to both the ego vehicle and the ground. Examples of dynamic objects can be moving vehicles, cyclists, pedestrians, or any other road users, as well as animals or other moving inanimate objects. In turn, a static object herein refers to an object that has motion relative to the ego vehicle (due to the motion of the ego vehicle itself) but is stationary (or has near-zero motion) relative to the ground. Examples of static objects can be parked cars, roadblocks, road debris, etc. This can improve any subsequent tasks that utilize the motion data because different measures can be taken depending on whether the object is a static object or a dynamic object.

[0066] Method 100 further includes assigning (S116) a set of 3D points with associated motion data to a free space estimate of the scene. The set of 3D points is assigned (S116) to the free space estimate based on the three-dimensional positions of each of the 3D points. In other words, each 3D point can be assigned to the cell in which the 3D point is located.

[0067] The free space estimate can be regarded as a 2D representation in 3D coordinates, which means that it is defined by coordinates in 3D form but is subject to a plane constraint. Figure 4A and Figure 4B This is further illustrated by way of example. Additionally, as described above, the free space estimate can be represented as (or include) a list of points describing the boundary (also referred to as the free space boundary) between drivable space and non-drivable space, and / or be represented as an occupancy grid in the ground plane. The occupancy grid can be formed by a two-dimensional grid in a plane parallel to the ground plane. The grid can include a plurality of grid cells (or simply referred to as "cells"). Then, the free space boundary can be formed by a list of cells of the occupancy grid where the boundary between free space and non-free space occurs. Alternatively, the list of points of the free space boundary can be independent of the cells of the occupancy grid. In this case, the points (or boundary points) can be represented by their 2D or 3D coordinates of their positions within the scene. Depending on the specific expression of the free space estimate, the step of assigning (S116) the set of 3D points to the free space estimate can be done differently.

[0068] In some embodiments, assigning (S116) a set of 3D points with associated motion data to a free space estimate of the scene can include assigning (S118) each 3D point in the set of 3D points to a cell among a plurality of cells in the occupancy grid of the free space estimate based on the three-dimensional position of the 3D point. Then assigning (S116) the set of 3D points with associated motion data to a free space estimate of the scene can further include assigning (S120) aggregated motion data to each cell in the occupancy grid based on the motion data associated with the 3D points assigned to the corresponding cell. The aggregated motion data of a cell can be determined, for example, as the average or median of the motion data associated with the 3D points assigned to the cell. The motion data can be averaged, for example, by averaging their directions and magnitudes separately. However, it should be understood that other ways of forming the aggregated motion data can also be used. As an example, a weighted average can be used. Additionally, any outliers among the 3D points assigned to a cell can be removed to obtain improved results. Further, the standard deviation of the 3D points within a cell can be used as a confidence score.

[0069] In some embodiments, assigning (S116) 3D points with associated motion data to a free space estimate of a scene can include selecting (S122) a subset of a set of 3D points corresponding to a free space boundary of the free space estimate. Then, assigning (S116) 3D points with associated motion data to the free space estimate of the scene can further include assigning (S124) aggregated motion data to the free space boundary based on the subset of 3D points. In other words, the aggregated motion data can be assigned to the free space boundary based on the motion data of the 3D points belonging to the free space boundary. The aggregated motion data of the free space boundary can be determined, for example, as an average or a median of the motion data associated with the 3D points belonging to the free space boundary. However, it should be understood that other ways of forming the aggregated motion data can also be used. As an example, a weighted average can be used. Additionally, any outliers among the 3D points in the subset of 3D points can be removed to obtain improved results. Further, the standard deviation of the 3D points within the subset of 3D points can be used as a confidence score.

[0070] Method 100 can further include providing (S126) the free space estimate to a trajectory planning module configured to generate candidate trajectories of a vehicle. Thus, the candidate trajectories can be generated based on the free space estimate.

[0071] Optionally, executable instructions for performing these functions are included in a non-transitory computer-readable storage medium or in other computer program products configured to be executed by one or more processors.

[0072] Generally, a computer-accessible medium can include any tangible or non-transitory storage medium or storage media, such as electronic media, magnetic media, or optical media (e.g., a disk or a CD / DVD-ROM coupled to a computer system via a bus). As used herein, the terms “tangible” and “non-transitory” are intended to describe a computer-readable storage medium (or “memory”) that excludes propagating electromagnetic signals, but are not intended to otherwise limit the types of physical computer-readable storage devices covered by the term “computer-readable medium or memory”. For example, the term “non-transitory computer-readable medium” or “tangible memory” is intended to cover types of storage devices that do not necessarily store information permanently, including, for example, random access memory (RAM). Program instructions and data stored in a tangible computer-accessible storage medium in a non-transitory form can further be transmitted via a transmission medium or signals such as electronic signals, electromagnetic signals, or digital signals, which can be conveyed via communication media such as a network and / or a wireless link.

[0073] Figure 2is a schematic illustration of an apparatus 200 according to some embodiments. The apparatus 200 (which may also be referred to as a computing device) may refer to any general-purpose computing device configured to perform the techniques described herein. The apparatus 200 may be configured, for example, to perform the method 100 as described in conjunction with Figure 1 as described.

[0074] Even though the apparatus 200 is illustrated herein as a single apparatus, the apparatus 200 may be a distributed computing system formed by multiple different computing devices.

[0075] The apparatus 200 includes control circuitry 202. The control circuitry 202 may physically include a single circuit device. Alternatively, the control circuitry 202 may be distributed across several circuit devices.

[0076] As Figure 2 shown in the example of, the apparatus 200 may further include a transceiver 206 and a memory 208. The control circuitry 202 is communicatively coupled to the transceiver 206 and the memory 208. The control circuitry 202 may include a data bus, and the control circuitry 202 may communicate with the transceiver 206 and / or the memory 208 via the data bus.

[0077] The control circuitry 202 may be configured to perform overall control of the functions and operations of the apparatus 200. The control circuitry 202 may include a processor 204, such as a central processing unit (CPU), a microcontroller, or a microprocessor. The processor 204 may be configured to execute program code stored in the memory 208 to perform the functions and operations of the apparatus 200. The control circuitry 202 is configured to perform the steps of the method 200 as described above in conjunction with Figure 2 as described. These steps may be implemented in one or more functions stored in the memory 208.

[0078] The transceiver 206 is configured to enable the apparatus 200 to communicate with other entities, such as vehicles or other devices. The transceiver 206 may transmit data to the apparatus 200 and may receive data from the apparatus 300.

[0079] The memory 208 may be a non-transitory computer-readable storage medium. The memory 208 may be one or more of a buffer, a flash memory, a hard disk drive, a removable medium, a volatile memory, a non-volatile memory, a random access memory (RAM), or other suitable devices. In a typical arrangement, the memory 208 may include non-volatile memory for long-term data storage and volatile memory that serves as the system memory of the apparatus 200. The memory 208 may exchange data with the circuitry 202 via the data bus. There may also be accompanying control lines and address buses between the memory 208 and the circuitry 202.

[0080] The functions and operations of the apparatus 200 may be implemented in the form of executable logic routines (e.g., lines of code, software programs, etc.) that are stored on a non-transitory computer-readable recording medium (e.g., the memory 208) of the apparatus 200 and executed by the circuitry 202 (e.g., using the processor 204). In other words, when it is stated that the circuitry 202 is configured to perform a particular function, the processor 204 of the circuitry 202 may be configured to execute a portion of program code stored on the memory 208, where the stored portion of program code corresponds to the particular function. Additionally, the functions and operations of the circuitry 202 may be a stand-alone software application or form part of a software application for performing additional tasks related to the circuitry 202. The described functions and operations may be regarded as methods that the respective apparatus is configured to execute, such as the method 100 discussed above in connection with Figure 1 Furthermore, although the described functions and operations may be implemented in software, such functions may also be implemented via dedicated hardware or firmware or some combination of one or more of hardware, firmware, and software. In the following, the functions and operations of the apparatus 200 are described.

[0081] The control circuitry 202 is configured to obtain an image sequence captured by an image capture device of a vehicle. The image sequence includes a plurality of images for depicting a scene at respective time instances of a plurality of time instances. This may be performed, for example, by executing a first obtaining function 210.

[0082] The control circuitry 202 is further configured to obtain a set of 3D points based on a depth map of the scene depicted in the image sequence. Each 3D point in the set of 3D points is associated with a three-dimensional position of the 3D point within the scene. This may be performed, for example, by executing a second obtaining function 212. It should be understood that the first obtaining function 210 and the second obtaining function 212 may be implemented by a common obtaining function. Such variations depend on the specific implementation.

[0083] The control circuitry 202 is further configured to determine motion data associated with each 3D point in the set of 3D points. The motion data is used to indicate an estimated motion of an object in the scene associated with the 3D point. This may be performed, for example, by executing a determining function 214. The motion data associated with each 3D point is determined by the following steps: obtaining a 2D point corresponding to the 3D point in the image plane of an image in the image sequence; applying an optical flow between the image and a subsequent image in the image sequence to the 2D point to obtain a subsequent 2D point in the image plane of the subsequent image; determining a subsequent 3D point by projecting the subsequent 2D point based on the depth map of the scene; and determining the motion data based on a difference between the three-dimensional positions of the 3D point and the subsequent 3D point.

[0084] The control circuit 202 is further configured to assign a set of 3D points with associated motion data to a free space estimation of the scene based on the three-dimensional positions associated with each 3D point. This can be performed, for example, by executing the assignment function 216.

[0085] The control circuit 202 can further be configured to output the free space estimation to a trajectory planning module configured to generate candidate trajectories for the vehicle. This can be performed, for example, by executing the output function 218.

[0086] It should be noted that the principles, features, aspects, and advantages of the method 100 described above also apply to the apparatus 200 described herein. To avoid unnecessary repetition, reference may be made to the above content. Figure 1 is a schematic illustration of a vehicle 300 according to some embodiments. The vehicle 300 is equipped with an autonomous driving system (ADS) 310. As used herein, "vehicle" refers to any form of motorized conveyance. For example, the vehicle 300 can be any road vehicle, such as a car (as illustrated herein), motorcycle, (cargo) truck, bus, smart bicycle, etc.

[0087] Figure 3 The vehicle 300 includes a plurality of elements (e.g., different systems and modules represented by the boxes illustrated in

[0088] which are common in autonomous or semi-autonomous vehicles. It will be understood that the vehicle 300 can have Figure 3 any combination of the various elements shown in Figure 3 In addition, the vehicle 300 can include more elements than those shown in Figure 3 Although the various elements are shown herein as being located inside the vehicle 300, one or more elements can be located outside the vehicle 300. Further, even though the various elements are depicted herein in a particular arrangement, the various elements can be implemented in different arrangements, as will be readily understood by those skilled in the art. It should be further noted that the various elements can be communicatively connected to each other in any suitable manner. Figure 3 The vehicle 300 of

[0089] Vehicle 300 further includes a control system 302. The control system 302 is configured to perform overall control of the functions and operations of the vehicle 300. The control system 302 includes a control circuit 304 and a memory 306. The control circuit 304 may physically include a single circuit device. Alternatively, the control circuit 304 may be distributed across a number of circuit devices. As an example, the control system 302 may share its control circuit 304 with other components of the vehicle. The control circuit 304 may include one or more processors, such as a central processing unit (CPU), a microcontroller, or a microprocessor. The one or more processors may be configured to execute program code stored in the memory 306 in order to perform the functions and operations of the vehicle 300. The one or more processors may be or include any number of hardware components for performing data processing or signal processing or for executing computer code stored in the memory 306. In some embodiments, the control circuit 304 or some of its functions may be implemented on one or more so-called system-on-chips (SoCs). As an example, the ADS 310 may be implemented on an SoC. The memory 306 optionally includes high-speed random access memory, such as dynamic random access memory (DRAM), static random access memory (SRAM), double data rate random access memory (DDR RAM), or other random access solid-state memory devices; and optionally includes non-volatile memory, such as one or more disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. The memory 306 may include database components, object code components, script components, or any other type of information structure for supporting the various activities in this description.

[0090] In the illustrated example, the memory 306 further stores map data 308. The map data 308 may be used, for example, by the ADS 310 of the vehicle 300 in order to perform autonomous functions of the vehicle 300. The map data 308 may include high-definition (HD) map data. It is contemplated that the memory 308, even though illustrated as a separate element from the ADS 310, may be provided as an integrated element of the ADS 310. In other words, according to some embodiments, any distributed or local memory device may be used to implement the concepts of the present invention. Similarly, the control circuit 304 may be distributed, such that one or more processors of the control circuit 304 are provided as an integrated element of the ADS 310 or any other system of the vehicle 300. In other words, according to some embodiments, any distributed or local control circuit device may be used to implement the concepts of the present invention.

[0091] Vehicle 300 further includes a sensor system 320. The sensor system 320 is configured to acquire sensing data regarding the vehicle itself or its surrounding environment. The sensor system 320 may include, for example, a Global Navigation Satellite System (GNSS) module 322 (e.g., GPS), which is configured to collect geographical location data of the vehicle 300. The sensor system 320 may further include one or more sensors 324. The one or more sensors 324 may be any type of on-vehicle sensors, such as image capture devices (e.g., one or more cameras), lidar and radar, ultrasonic sensors, gyroscopes, accelerometers, odometers, etc. It should be understood that the sensor system 320 may also provide the possibility of acquiring sensing data directly or via dedicated sensor control circuitry in the vehicle 300.

[0092] Vehicle 300 further includes a communication system 326. The communication system 326 is configured to communicate with external units, such as with other vehicles (i.e., via vehicle-to-vehicle (V2V) communication protocols), remote servers (e.g., cloud servers, cluster servers, etc.), databases, or other external devices (i.e., vehicle-to-infrastructure (V2I) or vehicle-to-everything (V2X) communication protocols). The communication system 326 may use one or more communication technologies to communicate. The communication system 326 may include one or more antennas. Cellular communication technologies may be used for remote communication, such as to a remote server or a cloud computing system. Additionally, if the cellular communication technology used has low latency, it may also be used for V2V, V2I, or V2X communication. Examples of cellular radio technologies include Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Enhanced Data Rates for GSM Evolution (EDGE), Long Term Evolution (LTE), 5th Generation Mobile Communication Technology (5G), 5G NR (New Radio), etc., and also include future cellular solutions. However, in some solutions, short- and medium-range communication technologies, such as wireless local area network (LAN) (e.g., IEEE 802.11-based solutions), may be used to communicate with other vehicles near the vehicle 300 or with local infrastructure elements. The European Telecommunications Standards Institute (ETSI) is developing cellular standards for vehicle communication, and 5G is considered to be a suitable solution, for example, due to its low latency and efficient handling of high bandwidth and communication channels.

[0093] The communication system 326 can further provide the possibility of sending outputs to a remote location (e.g., a remote operator or a control center) via one or more antennas. Additionally, the communication system 326 can further be configured to allow various components of the vehicle 300 to communicate with each other. As an example, the communication system can provide local network setups such as CAN bus, I2C, Ethernet, and fiber optic, etc. The local communication within the vehicle can also be of the wireless type with protocols such as Wireless Fidelity (WiFi), Long Range Radio (LoRa), Zigbee, Bluetooth, or similar medium / short-range technologies.

[0094] The vehicle 300 further includes a maneuvering system 328. The maneuvering system 328 is configured to control the maneuvering of the vehicle 300. The maneuvering system 328 includes a steering module 330 configured to control the driving direction of the vehicle 300. The maneuvering system 328 further includes a throttle module 332 configured to control the actuation of the throttle of the vehicle 300. The maneuvering system 328 further includes a braking module 334 configured to control the actuation of the brakes of the vehicle 300. The various modules of the maneuvering system 328 can receive manual inputs from the driver of the vehicle 300 (i.e., from the steering wheel, the accelerator pedal, and the brake pedal respectively). However, the maneuvering system 328 can be communicatively connected to the vehicle's ADS 310 to receive instructions on how the various modules should act. Thus, the ADS 310 can control the maneuvering of the vehicle 300.

[0095] As described above, the vehicle 300 includes an ADS 310. The ADS 310 can be part of the vehicle's control system 302. The ADS 310 is configured to perform functions and operations in the autonomous (or semi-autonomous) functions of the vehicle 300. The ADS 310 can include multiple modules, where each module is assigned a different function of the ADS 310.

[0096] The ADS 310 can include a positioning module 312 or positioning block / system. The positioning module 312 is configured to determine and / or monitor the geographical location and driving direction of the vehicle 300, and can utilize data from the sensor system 320 such as data from the GNSS module 322. Alternatively or in combination, the positioning module 312 can utilize data from one or more sensors 324. However, the positioning system can alternatively be implemented as Real-Time Kinematic (RTK) GPS to improve accuracy.

[0097] The ADS 310 may further include a perception module 314 or a perception block / system. The perception module 314 may refer to any well-known module and / or function, such as being included in one or more electronic control modules and / or nodes of the vehicle 300, and may be adapted to and / or configured to interpret sensing data related to the driving of the vehicle 300 to identify, for example, obstacles, lanes, relevant markers, appropriate navigation paths, etc. Thus, the perception module 314 may be adapted to rely on multiple data sources (e.g., automotive imaging, image processing, computer vision, and / or in-vehicle networking, etc.) and obtain inputs from multiple data sources, and combine with, for example, the sensing data from the sensor system 320. The perception module 314 may be configured to determine a free space estimate based on the sensing data. In addition, the perception module 314 may be configured to perform the functions of the method 100 as described above in connection with Figure 1 These functions may be implemented in a separate device provided in the vehicle 300 (such as the device 200 described above in connection with Figure 2 . Alternatively, as will be readily understood by those skilled in the art, these functions may be distributed across one or more modules, systems, or components of the vehicle 300. For example, the control circuit 304 of the control system 302 may be configured to perform the steps of the method 100. Thus, the device 200 may be distributed within the vehicle 300 and, for example, may share its control circuit with the control circuit 304 of the control system 302.

[0098] The positioning module 312 and / or the perception module 314 may be communicatively connected to the sensor system 320 to receive sensor data (e.g., images) from the sensor system 320 (e.g., from the camera of the sensor system 320). The positioning module 312 and / or the perception module 314 may further transmit control instructions to the sensor system 320.

[0099] The ADS may further include a path planning module 316 (also referred to as a trajectory planning module). The path planning module 316 is configured to determine a planned path (or candidate trajectory) of the vehicle 300 based on the perception and position of the vehicle determined by the perception module 314 and the positioning module 312, respectively. The planned path determined by the path planning module 316 may be sent to the maneuvering system 328 for execution. The path planning module 316 may utilize the free space estimate with the assigned motion data to determine the planned path.

[0100] The ADS may further include a decision and control module 318. The decision and control module 318 is configured to perform the control of the ADS 310 and make decisions. For example, the decision and control module 318 may decide whether the planned path determined by the path planning module 316 should be executed. The decision and control module 318 may further be configured to detect any avoidance actions of the vehicle, such as deviations from the planned path or the expected trajectory of the path planning module 316. This includes both avoidance actions performed by the ADS 310 and avoidance actions performed by the vehicle driver.

[0101] It should be understood that parts of the described solutions may be implemented in the vehicle 300, in a system located outside the vehicle, or in a combination of inside and outside the vehicle; for example, in a so-called cloud solution in a server communicating with the vehicle. Different features and steps of the embodiments may be combined in other combinations than the described combinations. In addition, the elements (i.e., systems and modules) of the vehicle 300 may be implemented in combinations different from those described herein.

[0102] Figure 4A By way of example, the surrounding environment of the vehicle 402 is illustrated in perspective view. The vehicle 402 may also be referred to as the ego vehicle 402. The vehicle 402 may be the vehicle 300 as described above in connection with Figure 3 what has been described.

[0103] Figure 4B By way of example, the surrounding environment of the vehicle 402 is illustrated in side view. More specifically, Figure 4A and Figure 4B illustrate examples of how free space estimation may be represented and provide examples of some of the principles of the present technology. These illustrations are mainly for providing a better understanding of the disclosed technology and should not be regarded as limiting. For this reason, these illustrations should be regarded as simplified examples of real-world scenarios.

[0104] The vehicle 402 includes an image capture device, illustrated herein as camera 404. In this example, the camera 404 has a field of view 420 located in front of the vehicle. However, it should be understood that the principles of the present disclosure may be applied to any direction of the vehicle. For example, rear or side cameras may also be used. This allows for free space estimation in any space around the vehicle.

[0105] In the illustrated example, the camera 404 is provided outside the vehicle 402. However, it should be understood that the camera 404 may be provided at any suitable location inside the vehicle 402 (e.g., integrated into the vehicle body). In addition, the vehicle 402 may include additional cameras as well as other types of sensors, as explained above for example in connection with Figure 3 what has been described.

[0106] The camera 404 is configured to capture an image of a scene of the environment around the vehicle 402. More specifically, the camera 404 is configured to capture an image sequence as described above. In this example, the vehicle 402 travels along a two-lane road defined by a first road boundary 418a, a second road boundary 418b, and a lane separator 418c.

[0107] In the surrounding environment in front of the ego vehicle 402, there are a first vehicle 416a and a second vehicle 416b. The first vehicle 416a and the second vehicle 416b herein represent possible obstacles that limit the free space of the ego vehicle 402. It should be noted that the obstacles can be any type of stationary object or moving object such as other vehicles, cyclists, pedestrians, animals, road construction signs, traffic cones, buildings, etc.

[0108] Figure 4A A representation of free space estimation is further shown. The free space expression in the illustrated example includes an occupancy grid 406. The occupancy grid 406 includes a plurality of cells (or grid cells) 410. The occupancy grid 406 is located in the estimated ground plane of the depicted surrounding environment. For illustrative purposes, a part of the occupancy grid 406 is shown as a dashed line herein. As a way of observation, the occupancy grid 406 divides the ground plane into a plurality of adjacent cells 410. It should be noted that the occupancy grid 406 can span an area larger or smaller than the illustrated part of the occupancy grid 406.

[0109] For illustrative purposes, the cells 410 of the occupancy grid 406 are depicted herein as rectangles of uniform size. However, it should be understood that the size and shape of the cells 410 are not limited to those depicted herein. For example, the cells 410 can have any polygonal shape. In addition, this example illustrates that the occupancy grid 406 can have a uniform cell size. However, in some embodiments, the occupancy grid 406 can have varying cell sizes. As an example, the cell size can increase with the distance from the vehicle 402. In other words, compared with the cells 410 closer to the camera 404, the occupancy grid 406 can have cells 410 of larger size farther from the camera 404. A smaller cell size can improve the accuracy and resolution of the estimated free space, but also result in a larger number of cells and thus an increased computational resource requirement. Having smaller-sized cells near the vehicle and larger-sized cells far from the vehicle 402 can be advantageous because it provides relatively high accuracy / resolution near the vehicle 402, which is very important when determining the route of the vehicle 402. At the same time, by allowing lower accuracy / resolution (which is less relevant) at farther distances, the computational resource requirement can be reduced.

[0110] To describe free space and non-free space, the cells 410 of the occupancy grid 406 can be classified as corresponding to free space or corresponding to non-free space. For illustrative purposes, this is shown by tick marks (corresponding to free space or drivable space) and cross marks (corresponding to non-free space or occupied space) for some of the cells 410. How to determine whether a cell 410 is free can be done in several different ways known in the art, such as by sensor data interpretation, machine learning-based segmentation techniques, etc. Since the present technology can be implemented using any suitable technique, it will not be further elaborated herein.

[0111] Now turning to Figure 4B , which illustrates the surrounding environment in a side view. More specifically, Figure 4B illustrates the principle of the present technology related to assigning motion data to free space estimation. Figure 4B Shows a ego vehicle 402 with a front camera 404. There is another vehicle 416 in front of the ego vehicle 402. Thus, the vehicles 402, 416 illustrated herein are traveling in the direction from left to right.

[0112] Figure 4B Further illustrates a set of 3D points 412 superimposed in the scene. Each 3D point 412 is associated with motion data represented by an arrow connected to the corresponding 3D point. Thus, the arrow used to represent motion data can indicate the direction (the direction of the arrow) and magnitude (the length of the arrow) of the motion. The plurality of 3D points 412 should be understood as 3D points with associated motion data, and the associated motion data is determined by method 100 as described previously.

[0113] Figure 4B The occupancy grid 406 is further shown in, represented by a dashed line herein. The vertical dashed line divides the occupancy grid into multiple cells 410. Since Figure 4B is shown in a side view, the occupancy grid 406 is shown in a one-dimensional manner. However, it should be understood that the occupancy grid 406 can be two-dimensional or more-dimensional, such as as shown in Figure 4A . In addition, the multiple cells 410 of the occupancy grid 406 can be classified as free (tick mark) space or non-free (cross) space.

[0114] The set of 3D points 412 with associated motion data can be assigned to free space estimation (in this case, assigned to the occupancy grid) based on their three-dimensional positions. In the one-dimensional case illustrated herein, this means that the 3D points are assigned to the cells among the multiple cells 410 corresponding to the distance from the camera 404 to the 3D point.

[0115] In the case where more than one 3D point is assigned to a cell, aggregated motion data can be assigned to the cell. Thus, the aggregated motion data is an aggregation of motion data associated with the 3D points assigned to the cell. The aggregated motion data is illustrated herein by an arrow in a dashed line manner.

[0116] In some embodiments, the motion data is restricted to be parallel to the estimation of the ground plane 408. Additionally, the ground plane can be estimated as a number of individual ground planes in order to capture any horizontal variations in the surrounding environment. This is illustrated herein by the aggregated motion data being parallel to the ground plane 408 at positions corresponding to the respective cells, as illustrated by the dashed arrows.

[0117] In some embodiments, the motion data is assigned to the free space boundary 414, which is represented herein by a vertical line with a dotted pattern. The free space boundary can be regarded as the boundary between free space and non-free space. In some embodiments, the free space boundary 414 is defined by a line at a specific coordinate (or a plane perpendicular to the ground plane in the case where the occupancy grid 406 is 2D). In this case, the aggregated motion data can be determined based on a subset of 3D points located within a defined distance from the free space boundary 414. In another example, the free space boundary 414 can be defined by cells located at the boundary between free space and non-free space. In this case, the aggregated motion data can be determined based on a subset of 3D points assigned to the cells.

[0118] In the illustrated example, the motion data indicates the relative motion between an object in the scene and the ego vehicle 402. Thus, the motion data of the 3D points associated with the ground will indicate a motion more or less the same as the motion of the ego vehicle 402. For example, if the ego vehicle 402 is traveling at a speed of 40 km / h, the 3D points will appear to move towards the camera 404 at approximately the same speed. Another vehicle 416 in this example is traveling in the same direction as the ego vehicle 402 but at a lower speed. For instance, the other vehicle 416 is traveling at a speed of 30 km / h. Therefore, any 3D points associated with the other vehicle 416 may appear to be moving towards the camera 404 of the ego vehicle 402 at a speed of approximately 10 km / h. Thus, the arrow representing the motion data of the 3D points associated with the other vehicle 416 appears shorter than the arrow representing the motion data of the 3D points associated with the ground. However, it should be understood that by considering the motion of the ego vehicle 302, the motion data can indicate the relative motion between the object and the ground, as has been explained above.

[0119] As can be readily understood by those skilled in the art, the inventive concept is in no way limited by Figure 4A and Figure 4BLimitations of the illustrative examples. For example, free space estimation can be performed in any driving scenario such as one-way streets, one-way roads with oncoming vehicles, multi-lane roads, parking lots, intersections, roundabouts, etc. Additionally, the size and shape of the illustrated elements may not represent real-world scenarios but are considered non-limiting examples for illustrative purposes.

[0120] It should be understood that, even though illustrated in the physical surroundings of vehicle 402, the set of 3D points 412, occupancy grid 406, and free space boundary 414 are elements that are virtually implemented in any device implementing the techniques of method 100. In Figure 4A and Figure 4B these elements are shown superimposed on the physical surroundings only for illustrative purposes.

[0121] The present technology has been presented above with reference to specific embodiments. However, other embodiments besides those described above are also feasible and are within the scope of the present invention. Within the scope of the present invention, method steps different from those described above may be provided, and the method may be implemented by hardware or software. Thus, according to an exemplary embodiment, a non-transitory computer-readable storage medium storing one or more programs is provided, the one or more programs being configured to be executed by one or more processors of a vehicle control system, the one or more programs including instructions for performing the method according to any one of the embodiments discussed above. Alternatively, according to another exemplary embodiment, a cloud computing system may be configured to execute any method presented herein. The cloud computing system may include distributed cloud computing resources that jointly execute the methods presented herein under the control of one or more computer program products.

[0122] It should be noted that any reference numerals do not limit the scope of the claims, and the present invention may be implemented at least in part in both hardware and software, and the same hardware item may represent several "devices" or "units".

Claims

1. A computer-implemented method (100) performed in a vehicle equipped with an automated driving system, the method (100) comprising: obtaining ( S102 ) an image sequence captured by an image capture device of the vehicle, wherein the image sequence includes a plurality of images for depicting a scene at corresponding time instances of a plurality of time instances; obtaining (S104) a set of 3D points based on a depth map of the scene depicted in the sequence of images, wherein each 3D point in the set of 3D points is associated with a three-dimensional position of the 3D point within the scene; determining (S106) motion data associated with each 3D point in the set of 3D points, wherein the motion data indicates an estimated motion of an object in the scene associated with the 3D point, wherein the motion data associated with each 3D point is determined (S106) by: Obtaining (S108) a 2D point corresponding to the 3D point in an image plane of an image in the image sequence; applying (S110) an optical flow between the image and a subsequent image in the image sequence to the 2D point, thereby obtaining a subsequent 2D point in an image plane of the subsequent image; determining (S112) a subsequent 3D point by projecting the subsequent 2D point based on the depth map of the scene; and determining (S114) the motion data based on a difference between the three-dimensional positions of the 3D point and the subsequent 3D point; and The set of 3D points with associated motion data are assigned (S116) to a free space estimate of the scene based on the three-dimensional position associated with each 3D point.

2. The method (100) according to claim 1, wherein: Determining (S106) the motion data further comprises obtaining an estimate of a ground plane of the scene depicted in the sequence of images, and The motion data is determined (S106) as motion parallel to the estimated ground plane.

3. The method (100) according to claim 1, wherein: The depth map is based on an output of a machine learning model, which is configured to determine a depth map based on a sequence of images as input.

4. The method (100) according to claim 1, wherein: The depth map is based on a lidar point cloud of the scene.

5. The method (100) of claim 1, wherein: The motion data indicates an estimated motion of the object in the scene associated with the 3D point relative to the motion of the vehicle.

6. The method (100) of claim 1, wherein: The motion data is further determined (S106) based on vehicle motion data of the vehicle, and The motion data further indicates an estimated motion of the object in the scene associated with a 3D reference point relative to a ground surface.

7. The method (100) of claim 1, wherein: The motion data includes information about the velocity of the objects in the scene.

8. The method (100) of claim 1, wherein: Assigning (S116) the set of 3D points with associated motion data to the free space estimate of the scene comprises: assigning (S118) each 3D point in the set of 3D points to a cell in a plurality of cells in an occupancy grid of the free space estimate based on the three-dimensional position of the 3D point, and Aggregate motion data is assigned (S120) to each cell in the occupancy grid based on the motion data associated with the 3D points assigned to the corresponding cells.

9. The method (S100) according to claim 1, wherein: Assigning (S116) the 3D points with associated motion data to the free space estimate of the scene comprises: selecting (S122) a subset of the set of 3D points corresponding to a free space boundary of the free space estimate, and Aggregate motion data is assigned (S124) to the free space boundary based on the subset of the 3D points.

10. The method (100) of claim 1, further comprising: The free space estimate is provided (S126) to a trajectory planning module configured to generate candidate trajectories for the vehicle.

11. A computer program product comprising instructions which, when executed by a computing device, cause the computing device to perform the method (100) according to any one of claims 1 to 10.

12. An apparatus (200) comprising a control circuit (202), the control circuit (202) being configured to: obtaining a sequence of images captured by an image capture device of the vehicle, wherein the sequence of images includes a plurality of images depicting a scene at respective ones of a plurality of time instances; obtaining a set of 3D points based on a depth map of the scene depicted in the sequence of images, wherein each 3D point in the set of 3D points is associated with a three-dimensional position of the 3D point within the scene; determining motion data associated with each 3D point in the set of 3D points, wherein the motion data indicates an estimated motion of an object in the scene associated with the 3D point, wherein the motion data associated with each 3D point is determined by: Obtaining a 2D point corresponding to the 3D point in an image plane of an image in the image sequence; applying an optical flow between the image and a subsequent image in the sequence of images to the 2D point, thereby obtaining a subsequent 2D point in an image plane of the subsequent image; determining a subsequent 3D point by projecting the subsequent 2D point based on the depth map of the scene; and determining the motion data based on a difference between the three-dimensional positions of the 3D point and the subsequent 3D point; as well as The set of 3D points with associated motion data are assigned to a free space estimate of the scene based on the three-dimensional position associated with each 3D point.

13. The device (200) according to claim 12, wherein: The control circuit (202) is further configured to output the free space estimate to a trajectory planning module configured to generate candidate trajectories for the vehicle.

14. A vehicle (300) equipped with an automatic driving system, the vehicle (300) comprising: an image capture device, and The device (200) according to claim 12 or 13.