Method and device for generating images for vehicles comprising a stereoscopic vision system.

The method of cropping and predicting depths in stereoscopic vision systems for vehicles addresses the inefficiencies in existing systems, enhancing ADAS performance by reducing processing time and resource use.

FR3155610B1Active Publication Date: 2026-02-20STELLANTIS AUTO SAS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
FR2023012768
Authority / Receiving Office
FR · FR
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-11-21
Publication Date
2026-02-20
Estimated Expiration
2043-11-21

AI Technical Summary

Technical Problem

Existing stereoscopic vision systems in vehicles require extensive image processing time and resources to determine depth and distance, which is not feasible due to limited onboard resources and the need for rapid data updates in ADAS systems.

Method used

A method for generating images using a stereoscopic vision system in vehicles that involves cropping images based on common field of view and predicting initial depths, eliminating the need for time-consuming rectification steps.

Benefits of technology

Facilitates faster and more efficient processing of stereoscopic images, improving the reliability and speed of ADAS systems by reducing processing time and resource consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000014_0000
    Figure 00000014_0000
  • Figure 00000030_0000
    Figure 00000030_0000
  • Figure 00000031_0000
    Figure 00000031_0000
Patent Text Reader

Abstract

A method or device implementing an image generation method for a vehicle equipped with a stereoscopic vision system comprising a first camera and a second camera. The method includes receiving representative data from a first image (31a) and a second image (32a) acquired by the cameras at the same time instant, predicting depths associated with pixels in the second image, determining coordinates in the second image of pixels corresponding to pixels in the first image, and selecting a subset of pixels from the first image corresponding to pixels forming part of an edge of the second image. A cropped first image (31b) and a cropped second image (32b) are generated from the first and second images based on the coordinates associated with the pixels in the subset, to generate images of the same size. Figure 3 (for the abbreviation)
Need to check novelty before this filing date? Find Prior Art

Description

Title of the invention: Method and device for generating images for vehicles comprising a stereoscopic vision system. technical field

[0001] The present invention relates to methods and devices for generating images for vehicles comprising an embedded stereoscopic vision system, for example, in a motor vehicle. The present invention also relates to a method and device for determining a common field of view in two images acquired by a stereoscopic vision system embedded in a vehicle. The present invention further relates to a method and device for controlling one or more ADAS systems embedded in a vehicle from a depth determined from images generated for a stereoscopic vision system embedded in the same vehicle. Technological background

[0002] Many modern vehicles are equipped with Advanced Driver-Assistance Systems (ADAS). Such ADAS systems are passive and active safety systems designed to eliminate human error in driving all types of vehicles. ADAS use advanced technologies to assist the driver while driving and thus improve performance. ADAS use a combination of sensor technologies to perceive the environment around a vehicle and then provide information to the driver or act on certain vehicle systems.

[0003] There are several levels of ADAS, such as reversing cameras and blind spot sensors, lane departure warning systems, adaptive cruise control or automatic parking systems.

[0004] ADAS systems embedded in a vehicle are powered by data obtained from one or more on-board sensors such as, for example, cameras. These cameras make it possible, in particular, to detect and locate other road users or any obstacles present around a vehicle in order, for example: - to adapt the vehicle's lighting according to the presence of other users; - to automatically regulate the vehicle's speed; - to act on the braking system in case of risk of impact with an object.

[0005] Determining depth or distance from images acquired by a stereoscopic vision system is performed using images acquired by that vision system. In order to allow the prediction of great depths or For distances exceeding 200 meters, a stereoscopic vision system installed in a vehicle typically includes at least two cameras positioned several tens of centimeters apart. This distance results in a difference in the field of view of the stereoscopic vision system's cameras, and the images they produce must then be processed for usability. This processing most often consists of rectifying these images.

[0006] Image processing, however, requires considerable time and resources, whereas an ADAS powered by depth or distance data requires rapid updating of this data while maintaining its accuracy. The resources allocated to the stereoscopic vision system are not unlimited, however, as an onboard system must be compact and / or lightweight to be integrated into a vehicle.

[0007] Thus, the quality of the data emitted and the speed of image processing by a vision system determine the proper functioning of driving aid devices using this data. Summary of the present invention

[0008] One object of the present invention is to solve at least one of the problems of the technological background described above.

[0009] Another object of the present invention is to reduce the processing time of images acquired by a stereoscopic vision system embedded in a vehicle.

[0010] Another object of the present invention is to improve road safety, in particular by improving the reliability of AD AS systems powered by data obtained from at least one camera.

[0011] According to a first aspect, the present invention relates to a method for generating images for a vehicle, the vehicle carrying a stereoscopic vision system comprising a first camera and a second camera, each configured to acquire an image of a three-dimensional scene from a different point of view, the method being characterized in that it comprises the following steps: - receipt of representative data from: • a first image acquired by the first camera at a first instant of acquisition, • a second image and a third image acquired by the second camera respectively at the first acquisition time instant and at a second acquisition time instant, the second acquisition time instant being prior to the first acquisition time instant; prediction, from the second and third images, of initial depths associated with a second set of pixels of the second image by a depth prediction model associated with a monoscopic vision system formed by the second moving camera, at least one part of the second set of pixels forming part of an edge of the second image; - projection of a first set of pixels from the first image into the three-dimensional scene to obtain a set of points, each corresponding to a pixel from the first set of pixels, based on the first coordinates of the pixel in the first image, a second depth associated with the pixel, and intrinsic parameters of the first camera. the second coordinates of the points of the set of points being defined in a reference frame associated with the first camera; - determination of third coordinates of each point of the set of points in a reference frame associated with the second camera as a function of the second coordinates and extrinsic parameters of the stereoscopic vision system; - determination of fourth pixel coordinates in the second image corresponding to the points as a function of the third coordinates and intrinsic parameters of the second camera; - selection of a subset of pixels from the first set of pixels based on the fourth coordinates and the first depths, the subset of pixels corresponding to the second set of pixels; - generation of a first cropped image by cropping the first image and generation of a second cropped image by cropping the second image, the cropping of the first image being a function of averages of the first coordinates associated with the pixels of the subset of pixels and the cropping of the second image being a function of a width and a height of the first cropped image.

[0012] Such a device makes it possible to crop the images acquired by two cameras of a stereoscopic vision system so as to retain only pixels associated with a field of view common to both cameras. Determining depths from the cropped images is thus facilitated and does not require rectification, which represents a time-consuming and resource-intensive step.

[0013] According to one variant of the method, the second set of pixels comprises a fourth set of pixels associated with a first edge of the second image, - the first edge corresponding to the left edge when the second camera is placed to the right of the first camera, and - the first edge corresponding to the right edge when the second camera is placed to the left of the first camera.

[0014] According to yet another variant of the method, the second set of pixels comprises a fifth set of pixels associated with a second edge of the second image, - the second edge corresponding to the upper edge when the second camera is placed below the first camera, and - the second edge corresponding to the lower edge when the second camera is placed above the first camera.

[0015] According to another variant of the process, the cropping of the second image is carried out in such a way as to generate the second cropped image with a width and a height equal to respectively the width and height of the first cropped image.

[0016] According to yet another variant of the process, the width and height of the first cropped image are each divisible by 32.

[0017] According to a further variant of the method, the first depths are predicted by a convolutional neural network as a function of a displacement of the vehicle between the second and first time instants of acquisition.

[0018] According to a second aspect, the present invention relates to a vehicle image generation device incorporating a stereoscopic vision system, the device comprising a memory associated with at least one processor configured for the implementation of the steps of the process according to the first aspect of the present invention.

[0019] According to a third aspect, the present invention relates to a vehicle, for example of the automobile type, comprising a device as described above according to the second aspect of the present invention.

[0020] According to a fourth aspect, the present invention relates to a computer program which includes instructions adapted for carrying out the steps of the process according to the first aspect of the present invention, in particular when the computer program is executed by at least one processor.

[0021] Such a computer program may use any programming language and be in the form of source code, object code, or an intermediate form between source code and object code, such as in a partially compiled form, or in any other desirable form.

[0022] According to a fifth aspect, the present invention relates to a computer-readable recording medium on which is recorded a computer program comprising instructions for carrying out the steps of the process according to the first aspect of the present invention.

[0023] On the one hand, the recording medium can be any entity or device capable of storing the program. For example, the medium can include a storage means, such as a ROM, a CD-ROM or a microelectronic circuit-type ROM, or a magnetic recording means or a hard disk drive.

[0024] On the other hand, this recording medium can also be a transmissible medium such as an electrical or optical signal, such a signal being able to be transmitted via an electrical or optical cable, by conventional or radio frequency or by beam. self-steering laser or by other means. The computer program according to the present invention can in particular be downloaded onto an Internet-type network.

[0025] Alternatively, the recording medium may be an integrated circuit in which the computer program is incorporated, the integrated circuit being adapted to execute or to be used in the execution of the process in question. Brief description of the figures

[0026] Other features and advantages of the present invention will become apparent from the description of the particular and non-limiting embodiments of the present invention below, with reference to the attached Figures 1 to 4, in which:

[0027] [Fig-1] schematically illustrates a stereoscopic vision system embedded in a vehicle, according to a particular and non-limiting example of the present invention;

[0028] [Fig.2] illustrates a flowchart of the different stages of an image generation process for the vehicle of [Fig.1], according to a particular and non-limiting example of the present invention;

[0029] [Fig.3] schematically illustrates images acquired and images generated from a stereoscopic vision system embedded in the vehicle of [Fig.1], according to a particular and non-limiting embodiment of the present invention;

[0030] [Fig.4] schematically illustrates a device configured for image generation for a stereoscopic vision system embedded in the vehicle of [Fig.1], according to a particular and non-limiting embodiment of the present invention. Description of examples of achievements

[0031] A method and device for generating images for a vehicle carrying a stereoscopic vision system will now be described in what follows with joint reference to Figures 1 to 4. The same elements are identified with the same reference signs throughout the description that follows.

[0032] The terms "first," "second" (or "firsts," "seconds"), etc., are used in this document by arbitrary convention to allow for the identification and distinction of different elements (such as operations, means, etc.) implemented in the embodiments described below. Such elements may be distinct or correspond to a single element, depending on the embodiment.

[0033] According to a particular and non-limiting example of an embodiment of the present invention, a method for generating images for a vehicle carrying a stereoscopic vision system is for example implemented by a computer of the vehicle's embedded system controlling this stereoscopic vision system.

[0034] The stereoscopic vision system comprises a set of cameras, at least two of which are arranged so that each acquires an image of a three-dimensional scene from a different viewpoint, the optical axes representing an orientation of the field of view of each camera being contained in two parallel planes, for example oriented non-parallel. The stereoscopic vision system is thus said to be non-parallel.

[0035] To this end, the image generation method for a vehicle carrying a stereoscopic vision system includes receiving data representing a first image acquired by the first camera at a first time instant of acquisition, as well as receiving data representing a second image and a third image acquired by the second camera at respectively the same first time instant of acquisition and at a second time instant of acquisition prior to the first time instant of acquisition.

[0036] The method also includes predicting depths associated with pixels of the second image, determining coordinates in the second image of pixels corresponding to pixels of the first image and selecting a subset of pixels of the first image corresponding to pixels forming part of an edge of the second image.

[0037] A first cropped image and a second cropped image are then generated from the first and second images according to the coordinates associated with the pixels of the subset of pixels and to generate images of the same size.

[0038] Fig. 1 schematically illustrates a stereoscopic vision system equipping a vehicle, according to a particular and non-limiting embodiment of the present invention.

[0039] Such an environment 1 corresponds, for example, to a road environment consisting of a network of roads accessible to the vehicle 10.

[0040] In this example, vehicle 10 corresponds to a vehicle with an internal combustion engine, an electric motor(s), or a hybrid vehicle with an internal combustion engine and one or more electric motors. Vehicle 10 thus corresponds, for example, to a land vehicle such as a car, a truck, a bus, or a motorcycle. Finally, vehicle 10 corresponds to an autonomous or non-autonomous vehicle, that is to say, a vehicle operating according to a predetermined level of autonomy or under the total supervision of the driver.

[0041] The vehicle 10 advantageously comprises several onboard cameras 11, 12, each configured to acquire images of a scene in the environment of the vehicle 10. This set of cameras 11, 12 forms the stereoscopic vision system. Two cameras 11 and 12 are illustrated in [Fig. 1]. The present invention is not, however, limited to a stereoscopic vision system comprising two cameras but extends to any stereoscopic vision system comprising 2 or more cameras, for example 2, 3, 4 or 5 cameras.

[0042] The two cameras 11, 12 have known intrinsic parameters. These parameters include, in particular: - the focal length of the first camera 11; - the focal length of the second camera 12; - distortions which are due to imperfections in the optical system of each camera; - the direction Cl of the optical axis of the first camera 11; - the C2 direction of the optical axis of the second camera 12; and - the respective resolutions of cameras 11, 12.

[0043] The intrinsic parameters characterize the transformation that associates, for an image point, the camera coordinates to the pixel coordinates, in each camera. These parameters do not change if the camera is moved.

[0044] The intrinsic matrix of the first camera 11 is defined by:

[0045] [Math.l]

[0046] With: • K, the intrinsic matrix of the first camera 11, • f J the focal length of the first camera 11, • Cx is the abscissa of the optical center of an image acquired by the first camera 11, and • Cy is the ordinate of the optical center of an image acquired by the first camera 11.

[0047] Similarly, the intrinsic matrix of the second camera 12 is defined by:

[0048] [Math.2] 'fr Cx' Kr= fr CTy .00 1 .

[0049] With: • K the intrinsic matrix of the second camera 12, • the focal length of the second camera 12, • 0% the abscissa of the optical center of an image acquired by the second camera 12, and • Cy is the ordinate of the optical center of an image acquired by the second camera 12.

[0050] The optical axes of the first and second cameras are arranged horizontally, that is to say in two parallel horizontal planes. These two planes are either separated by a height H, H having a value of zero if the planes coincide.

[0051] Distortions, which are due to imperfections in the optical system such as defects in the shape and positioning of camera lenses, will deflect the light beams and thus induce a positioning error for the projected point relative to an ideal model. It is then possible to complete the camera model by introducing the three distortions that generate the most significant effects, namely radial, decentering, and prismatic distortions, induced by defects in lens curvature, parallelism, and coaxiality of the optical axes. In this example, the cameras are assumed to be perfect, meaning that distortions are either not taken into account or their correction is addressed during image acquisition.

[0052] These two cameras 11, 12 are arranged so that each acquires an image of a scene from a different viewpoint. The first viewpoint is, for example, located on or in the left-hand rearview mirror of the vehicle 10 or at the top of the windshield of the vehicle 10. The second viewpoint is, for example, located on or in the right-hand rearview mirror of the vehicle 10 or at the top of the windshield of the vehicle 10. If both cameras are located at the top of the windshield of the vehicle, they are then positioned at a certain distance. In this example, the first camera 11 is located at the top of the windshield of the vehicle 10, and the second camera 12 is located in the right-hand rearview mirror of the vehicle 10.

[0053] A first marker is associated with the first camera 11: - the direction of the x-axis is defined as horizontal and normal to the optical axis of the first camera 11. The distance B separating the optical center of the first camera 11 from the projection of the optical center of the second camera 12 onto the horizontal plane passing through the optical center of the first camera 11 is called the reference basis (in English "baseline"); - the direction of the y-axis is defined as vertical and normal to the optical axis of the first camera 11; - The direction of the z-axis is defined as orthogonal to the directions of the x and y axes. The three axes x, y and z thus form an orthonormal coordinate system.

[0054] The extrinsic parameters related to the position of cameras 11, 12 are the following parameters: - three translations in the x, y, and z directions: Tx, Ty, and Tz, constituting the translation vector T, with Tx = B, Ty = H, and Tz = 0; and - a rotation around the y-axis by an angle 0 called the yaw angle.

[0055] A matrix for going from the reference frame of the first camera 11 to the reference frame of the second camera 12 is thus defined by:

[0056] [Math.3] [ÏR.T] = ' COS0 0 - sin# 0 sin0 B ' 10 AM 0 cos0 0 ,

[0057] With: • [R, T] the matrix to go from the frame of reference of the first camera 11 to the frame of reference from the second camera 12, • 0 the yaw angle, • -B the reference base, and • H the difference in height between the first 11 and second 12 cameras.

[0058] The extrinsic parameters are determined, for example, during a calibration phase of the stereoscopic vision system.

[0059] A key constraint of stereoscopic vision systems used in automobiles is, for example, the large distance between the two cameras. Indeed, to cover a measurement range of 200 meters, the reference base must be 60 cm for cameras commonly used in this field.

[0060] The two cameras 11, 12 acquire images of a scene located in front of the vehicle 10, the first camera 11 alone covering a first acquisition field 13, the second camera 12 alone covering a second acquisition field 14 and the two cameras 11, 12 both covering a third acquisition field 15. The first and third acquisition fields 13, 15 thus allow a monoscopic view of the scene by the first camera 11, the second and third acquisition fields 14, 15 allow a monoscopic view of the scene by the second camera 12 and the third acquisition field 15 allows a stereoscopic view of the scene by the stereoscopic vision system composed of the two cameras 11, 12.

[0061] An obstacle 18 is placed in the acquisition field of the cameras, for example in the third acquisition field 15. The presence of the obstacle 18 defines an occlusion field for the stereoscopic vision system composed here of the three fields 16, 17 and 19.

[0062] Among these three fields, field 16 is visible from the second camera 12. The part of the scene present in this field 16 is therefore observable using the monoscopic vision system composed of the second camera 12.

[0063] The field 17 is visible from the first camera 11. The part of the scene present in this field 17 is therefore observable with the monoscopic vision system composed of the first camera 11.

[0064] Finally, field 19 is not visible from any of the cameras. The part of the scene present in this field 19 is therefore not observable.

[0065] It is evident that it is possible to use such a stereoscopic vision system to take images of scenes located on the sides or behind the vehicle 10 by equipping it with cameras placed and oriented differently.

[0066] The images acquired by the cameras 11, 12 at a given acquisition time are in the form of data representing pixels characterized by: - ​​coordinates in each image; and - data relating to the colours and brightness of objects in the observed scene in the form of, for example, RGB colourimetric coordinates (from the English "Red Green Blue", in French "Rouge Vert Bleu") or HSL (Tone, Saturation, Luminosity).

[0067] The images acquired by cameras 11, 12 represent views of the same scene taken from different viewpoints, the camera positions being distinct. This scene includes, for example: - buildings; - road infrastructure; - other stationary users, for example a parked vehicle; and / or - other mobile users, for example another vehicle, a cyclist or a moving pedestrian.

[0068] These images are sent to a computer of a device equipping the vehicle 10 or stored in a memory of a device accessible to a computer of a device equipping the vehicle 10.

[0069] An image generation process for the vehicle 10 is advantageously implemented by the vehicle 10, i.e. by a computer or a combination of computers of the vehicle 10's embedded system, for example by the computer or computers in charge of the stereoscopic vision system of the vehicle 10.

[0070] In a first operation, the computer receives initial data representative of: • a first image 31a acquired by the first camera 11 at a first temporal instant of acquisition and • a second image 32a and a third image acquired by the second camera 12 respectively at the first time instant of acquisition and at a second time instant of acquisition, the second time instant of acquisition being prior to the first time instant of acquisition.

[0071] The first 31a and second 32a images form a first pair of stereoscopic images, that is to say images acquired by two separate cameras at the same time instant.

[0072] In a second operation, first depths associated with a second set of pixels of the second image are predicted from the second and third images, by a depth prediction model associated with a monoscopic vision system formed from the second moving camera 12.

[0073] At least a part of the second set of pixels forms part of an edge of the second image.

[0074] According to a first particular embodiment, the second set of pixels comprises a fourth set of pixels associated with a first edge of the second image 32a, - the first edge corresponding to the left edge when the second camera 12 is placed to the right of the first camera 11, and - the first edge corresponding to the right edge when the second camera 12 is placed to the left of the first camera 11.

[0075] According to a second particular embodiment, the second set of pixels comprises a fifth set of pixels associated with a second edge of the second image 32a, - the second edge corresponding to the upper edge when the second camera 12 is placed below the first camera 11, and - the second edge corresponding to the lower edge when the second camera 12 is placed above the first camera 11.

[0076] The first and second particular embodiments are, according to a third particular embodiment, combined. Thus, the second set of pixels includes both the fourth and fifth sets of pixels, for example when the second camera 12 is below and to the right of the first camera 11.

[0077] According to a particular embodiment, the first depths are predicted by a convolutional neural network as a function of a displacement of the vehicle 10 between the second and first time instants of acquisition, by any depth prediction model associated with a monoscopic vision system such as packnet® which achieves metric accuracy with supervision of the motion prediction model of the second camera with the speed of the vehicle.

[0078] The initial depths are predicted in the frame of reference of the second camera 12. They are then transcribed into the frame of reference of the first camera 11 by the following function:

[0079] [Math.4]

[0080] With: IV w ■1 coordinates of a point in the three-dimensional scene in the reference frame from the first camera 11, • [ -R - T ] a matrix associated with the change of reference frame, from the reference frame of the second camera 12 to the reference frame of the first camera 11, defined as the opposite matrix to the matrix for passing from the reference frame of the first camera 11 to the reference frame of the second camera 12 previously defined, • the inverse of the intrinsic matrix of the second camera 12, • ' ' a homogeneous vector comprising the coordinates of a pixel of the second

[0081]

[0082]

[0083]

[0084]

[0085]

[0086] 1 image 32a having xr as the x-coordinate and Yr as the y-coordinate, and • a first depth associated with the pixel of the second image. In a third operation, a first set of pixels from the first image 31a is projected into the three-dimensional scene to obtain a set of points, each corresponding to a pixel from the first set of pixels, as a function of first coordinates of the pixel in the first image 31a, a second depth associated with the pixel, and intrinsic parameters of the first camera 11. Second coordinates of the points in the point set are defined in a reference frame associated with the first camera 11. According to one particular implementation example, the first set of pixels includes all the pixels of the first image. According to another particular embodiment example, the first set of pixels includes a portion of the pixels from the first image, for example the lower half of the first image 31a if the first camera 11 is located above the second camera 12. For example, a pixel from the first image 31a is projected into the three-dimensional scene by the following function: [Math.5] i w w WI -1 With : Xf fl ■1 _ W~ i ■1 IV i

[0087]

[0088]

[0089]

[0090]

[0091]

[0092] ^W. coordinates of a point in the three-dimensional scene in the reference frame from the first camera 11, jq! the inverse matrix of the intrinsic matrix of the first camera 11, yi a homogeneous vector comprising the coordinates of a pixel of the first image 31a having xl as the abscissa and X as the ordinate, * Z^y a second depth associated with the pixel of the first image, • Cx is the abscissa of the optical center of an image acquired by the first camera 11, and • Cy is the ordinate of the optical center of an image acquired by the first camera 11. It should be noted that if the origin of a reference frame associated with the first image 31a is located in the top left of the first image 31a, then the coordinates of the optical center of the first image 31a are defined by: [Math.6] And [Math.6] CI = h cy 2 With : • c^x the abscissa of the optical center of the first image 31a, • Cy, the ordinate of the optical center of the first image 31a, • WJ the width of the first image 31a, and • hj the height of the first image 31a. Similarly, for the second image 32a, if the origin of a reference frame associated with the second image 32a is located in the upper left corner of the second image 32a, then the coordinates of the optical center of the second image 32a are defined by: [Math.7] And [Math.7] With : • Cx is the abscissa of the optical center of the second image 32a, • Cy, the ordinate of the optical center of the second image 32a, • wr the width of the second image 32a, and • hr the height of the second image 32a.

[0093]

[0094] In a fourth operation, the third coordinates of each point of the set of points are determined in a reference frame associated with the second camera 12 as a function of the second coordinates and extrinsic parameters of the stereoscopic vision system. The transition from second coordinates expressed in the frame of reference of the first camera at third coordinates expressed in the second frame of reference camera 12 is obtained by the following function:

[0095] [Math.8] 1 Y^ j Yrw = [Æ,T] Al ' cos0 0 sin0 B 0 1 0 H -sin0 0 cos0 0 yt Y1 1 W 1 x(vcos0+zjysin© + B Y^+H - XjVsin0+zjyCosS either [Math.8] ^^+7^0+8 ■r W + Z^cos0

[0096] With : coordinates of a point in the three-dimensional scene in the reference frame from the second camera 12, • [R, T] the matrix for going from the frame of reference of the first camera 11 to the frame of reference of the second camera 12, a homogeneous vector comprising the coordinates of a point in the scene 1 W i 1 i three-dimensional in the frame of reference of the first camera 11, • 0 the yaw angle, • B the reference base, and • H the difference in height between the first 11 and second 12 cameras.

[0097]

[0098]

[0099] In a fifth operation, fourth pixel coordinates are determined in the second image 32a corresponding to the points as a function of the third coordinates and intrinsic parameters of the second camera 12. The fourth coordinates are obtained using the following function: [Math.9]

[0100] xjq xw y? = Kr Yrw [L 0 cD 0 fp cy L 0 0 11 xw 1 w zw fr [ + H ) + CÇ ( - —+ Z[vcosd ) With : ry™ ■ < a homogeneous vector comprising the homogeneous coordinates of a pixel of the second image 32a having a homogeneous abscissa, y™ for ordinate homogeneous and z? for homogeneous depth, • Kr the intrinsic matrix of the second camera 12, coordinates of a point in the three-dimensional scene in the reference frame from the second camera 12, • the focal length of the second camera 12, • Cx is the abscissa of the optical center of the first image 31a, • Cy, the ordinate of the optical center of the first image 31a, • Cx is the abscissa of the optical center of the second image 32a, • Cy, the ordinate of the optical center of the second image 32a, • 0 the yaw angle, • B the reference base,

[0101]

[0102]

[0103]

[0104]

[0105]

[0106] • H the difference in height between the first 11 and second 12 cameras, and * Z^ a depth of a point in the frame of the first camera 11. The following function allows you to convert homogeneous coordinates into coordinates in the second image 32a: [Math. 10] । f lz^ÿcose+zl in0+J z y} z Cy With : [xr ' the coordinates of a pixel in the second image having xr for abscissa and laugh yr for ordinate, • Xr has a homogeneous abscissa, a homogeneous ordinate, and zr a homogeneous depth, • the focal length of the second camera 12, • c% the abscissa of the optical center of the first image 31a, • Cy, the ordinate of the optical center of the first image 31a, • the abscissa of the optical center of the second image 32a, • Cy, the ordinate of the optical center of the second image 32a, • 0 the yaw angle, • B the reference base, • H the difference in height between the first 11 and second 12 cameras, and * Zw a depth of a point in the frame of the first camera 11. In a sixth operation, a subset of pixels from the first set of pixels is selected based on the fourth coordinates and the first depths, the subset of pixels corresponding to the second set of pixels. If the first camera 11 is placed to the left of the second camera 12, then the second set of pixels includes a fourth set of pixels belonging to the right edge of the second image 32a. The right edge of the second image 32a is defined as the first vertical row of pixels in the second image 32a. These pixels of the fourth set of pixels are defined by the following equation: [Math. 11] xr = 0 With : [Math. 11] xr the x-coordinate of a pixel in the second image 32a.

[0107] By integrating this equation into the previous equations, we obtain the following equation to associate the pixels of the first image 31a with the right edge of the second image 32a:

[0108] [Math. 12] Z}wf rsin6+frB+Z{,xc cos0 Xl_min= *"Cx Z- SlU\ / “£ COSt^

[0109]

[0110] [YES] With : • xL_min is the abscissa of a pixel in the first image 31a associated with the right edge of the second image 32a, • fj the focal length of the first camera 11, • the focal length of the second camera 12, • the abscissa of the optical center of the first image 31a, • Cx is the abscissa of the optical center of the second image 32a, • 0 the yaw angle, • B the reference base, * Z^ a depth of a point in the frame of the first camera 11. Similarly, if the first camera 11 is placed below the second camera 12, then the second set of pixels includes a fifth set of pixels belonging to the bottom edge of the second image 32a. The bottom edge of the second image 32a is defined as the last horizontal row of pixels in the second image 32a. These pixels of the fifth set of pixels are defined by the following equation: [Math. 13] Year ~ Year max With : [Math. 13] yr the ordinate of a pixel in the second image 32a, and [Math. 13] yrinax the ordinate of a pixel from the bottom edge of the second image 32a.

[0112] By integrating this equation into the previous equations, we obtain the following equation to associate the pixels of the first image 31a with the lower edge of the second image 32a:

[0113] [Math. 14] I J-----7--r ri ........... ............ / ■.......................... _L — tz cosfl) T y r.mar

[0114] With: • Yr max is the ordinate of a pixel from the bottom edge of the second image 32a, • yi inax the ordinate of a pixel of the first image 31a associated with the lower edge of the second image 32a, • fj the focal length of the first camera 11, • the focal length of the second camera 12, • Cx is the abscissa of the optical center of the first image 31a, • Cy, the ordinate of the optical center of the first image 31a, • Cy, the ordinate of the optical center of the second image 32a, • 0 the yaw angle, • H the difference in height between the first 11 and second 12 cameras, and * Z[y a depth of a point in the frame of the first camera 11.

[0115] Conversely, if the first camera 11 is placed above the second camera 12, then the second set of pixels includes a fifth set of pixels belonging to the upper edge of the second image 32a. The upper edge of the second image 32a is defined as the first horizontal row of pixels of the second image 32a. These pixels of the fifth set of pixels are defined by the following equation:

[0116] [Math. 15] yr=o With : [Math. 15] yr the ordinate of a pixel in the second image 32a.

[0117] By integrating this equation into the previous equations, we obtain the following equation to associate the pixels of the first image 31a with the upper edge of the second image 32a:

[0118] [Math. 16]

[0119] With: • Yi mjn, the ordinate of a pixel from the first image 31a associated with the upper edge of the second image 32a, • fJ the focal length of the first camera 11, • fr the focal length of the second camera 12, • Cx the abscissa of the optical center of the first image 31a, • Cy the ordinate of the optical center of the first image 31a, • Cy the ordinate of the optical center of the second image 32a, • 0 the yaw angle, • H the difference in height between the first 11 and second 12 cameras, and * a depth of a point in the frame of the first camera 11.

[0120] In a seventh operation, a first cropped image 31b is generated by cropping the first image 31a and a second cropped image 32b is generated by cropping the second image 32a. The cropping of the first image 31a is a function of averages of the first coordinates associated with the pixels of the subset of pixels and the cropping of the second image 32a is a function of a width and a height of the first cropped image 31b.

[0121] As illustrated in [Fig. 3], the first cropped image 31b includes a portion of the first image 31a. During cropping, a left portion 311 of the first image 31a is cropped, as well as an upper or lower portion 312 of the first image 31a, depending on the relative position of the first camera 11 with respect to the second camera 12. Similarly, the second cropped image 32b includes a portion of the second image 32a. During cropping, a right portion 321 of the second image 32a is cropped, as well as a lower or upper portion 322 of the first image 31a, depending on the relative position of the first camera 11 with respect to the second camera 12. The illustrated example thus corresponds to a first camera 11 positioned below the second camera 12.

[0122] The left limit of the portion of the first cropped image 31b is defined as the average of the abscissas of the subset of pixels of the first image 31a equal

[0123]

[0124]

[0125]

[0126]

[0127]

[0128]

[0129]

[0130] at xJ min, each of these pixels being associated with the fourth set of pixels of the second image 32a. Similarly, if the first camera 11 is below the second camera 12 (Figure 3), the lower limit of the first cropped image 31b is defined as the average of the ordinates of the subset of pixels in the first image 31a, each of these pixels being associated with the fifth set of pixels in the second image 32a equal to y*. In other words, according to this particular embodiment, the first cropped image 31b comprises the pixels of the first image 31a whose abscissa is greater than xf and whose ordinate is less than y*. Conversely, if the first camera 11 is above the second camera 12, the upper limit of the first cropped image 31b is defined as the average of the ordinates of the subset of pixels of the first image 31a, each of these pixels being associated with the fifth set of pixels of the second image 32a equal to y* .fl. In other words, according to this particular embodiment, the first cropped image 31b comprises the pixels of the first image 31a whose abscissa is greater than r, ■ and whose ordinate is greater than y* .. 1 1 J limn According to a particular embodiment, the cropping of the second image 32a is carried out in such a way as to generate the second cropped image 32b with a width and a height equal to respectively the width and height of the first cropped image 31b. Thus, if the first cropped image 31b has a width w' and a height h', then the second cropped image 32b has the same dimensions, that is to say that the second cropped image 32b has a width w' and a height h'. According to a particular embodiment example, the width and height of the first cropped image 31b are each divisible by 32. This makes it possible, in particular, to subsequently apply certain depth prediction models for a stereoscopic system using an optical flow calculation method. Thus, the x and y coordinates corresponding to the previously determined cropping limits are determined by the following functions: [Math. 17] X1 min = + (32- I Xj inin I %32) y* = |y, 1 + (32-1 y, l%321 J Imax [J l inax J ( [JI max J ) if the first camera 11 is below the second camera 12, or [Math. 18] y*. =73-(1^-7 / 1 + (32-1 hy, . |%32n J limn ([ J Immi ( [ Jlmmj )) if the first camera 11 is above the second camera 12.

[0131] With: xl_min an abscissa corresponding to the lower edge of the first cropped image 31b, 1 max an ordinate corresponding to the left edge of the first cropped image 31b, y* an ordinate corresponding to the right edge of the first cropped image 31b, [.] an operator allowing obtaining the first integer less than its argument, and %32 corresponds to the remainder of a division of an argument by 32.

[0132] According to a particular embodiment, the determination of the abscissas and ordinates corresponding to the cropping limits is performed for several sets of images acquired by the first 11 and second 12 cameras; that is, this process is repeated several times. At each iteration of the process, representative data of a first image from a set of images Ei acquired by the first camera 11 at a third acquisition time instant and of the second and third images from the same set of images Ei acquired by the second camera 12 respectively at the third acquisition time instant and at an acquisition time instant prior to the third acquisition time instant are obtained. The second through seventh operations are then iterated for each of the sets of images Ei independently, making it possible to obtain for each set of images Ei an abscissa Xj m7n(Ei) and an ordinate y*max.(Ei) when the first camera 11 is below the second camera 12 or y* (Ei) when the first camera 11 is above the second camera 12 corresponding to the limits allowing the cropping of the first image of each set of images Ei. .

[0133] In order to standardize the cropping of images acquired by the stereoscopic vision system, generalized limits (x'; y') are determined and are, for example, equal to the respective means of the abscissas Xj m,„(Ei) and the ordinates y* (Ei) or y* . (Ei) obtained by the following functions: J l_max ' J Inunv ' 1

[0134] [Math. 19] x — average^ ^(£1)^+ (32- [average^ ^(Ei) ) ] %32))

[0135] [Math.20] y' = average^y*max(Ei)j+ (32- ^mean^} max(Ei) ) J %32 when the first camera 11 is below the second camera 12, or

[0136]

[0137]

[0138]

[0139]

[0140]

[0141]

[0142] [Math.21] y1 = mean\y^ niin (Ei) j + (32 - [mean (y^ ^ (Ei))] %32 when the first camera 11 is above the second camera 12 With : x is a generalized abscissa, y' is a generalized ordinate, Xj (Ei) corresponding to the lower edge of the first cropped image of a set of images Ei, y* (Ei) an ordinate corresponding to the left edge of the first cropped image of a set of images Ei, y* ^^Ei) an ordinate corresponding to the right edge of the first cropped image of a set of images Ei, [.] an operator that allows you to obtain the first integer less than its argument, and %32 corresponds to the remainder of a division of an argument by 32. According to another example, in order to obtain cropped images whose widths and heights are divisible by 32, generalized limits (x' ;y') are determined and are, for example, equal to the respective averages of the abscissas x^ rnni(Ei) and the ordinates y*(Ei) or y*(Ei) obtained by the following functions: [Math.22] x = mean^x^ min(Ei) ) [Math.23] y' = mean ( yf max (Ei)) when the first camera 11 is below the second camera 12, or [Math.24] y' = mean ( y] min (Ei) With : • x' a generalized abscissa, • there is a generalized ordinate, * min( Ei) corresponding to the lower edge of the first cropped image of a set of images Ei, (Ei) an ordinate corresponding to the left edge of the first image max cropped from a set of images Ei. • y* ( Ei) an ordinate corresponding to the right edge of the first cropped image of a set of images Ei.

[0143] The generalized limits are then used to automatically reframe images acquired by the first 11 and second 12 cameras of the on-board stereoscopic vision system.

[0144] Thus, the image generation process for a vehicle equipped with a stereoscopic vision system makes it possible to generate pairs or couples of cropped stereoscopic images, that is, images in which the pixels of each image in the stereoscopic image pair are associated with the same part of a three-dimensional scene seen by the two cameras forming the stereoscopic vision system. The pixels of the first 31b and second 32b cropped images are thus associated with objects in the three-dimensional scene that are present both in the field of view of the first camera 11 and in the field of view of the second camera 12.

[0145] A depth prediction process associated with the stereoscopic vision system on board the vehicle 10 can for example use these cropped images 31b, 32b to predict depths associated with the pixels of these images without having to resort to a rectification operation, making the execution of this depth prediction process faster.

[0146] An AD AS receiving the depth data predicted by this process of Depth prediction, combined with the vehicle's onboard stereoscopic vision system, provides rapidly updated information and is therefore more responsive. Its operation is thus safer and more precise.

[0147] Figure [Fig. 2] illustrates a flowchart of the different stages of a method 2 for generating images for a vehicle carrying a stereoscopic vision system, for example for the vehicle 10 of [Fig. 1], according to a particular and non-limiting embodiment of the present invention.

[0148] The method 2 is for example implemented by one or more processors of one or more computers embedded in the vehicle 10, for example by a computer controlling the stereoscopic vision system.

[0149] In step 21, representative data of: • a first image 31a acquired by the first camera 11 at the first temporal instant of acquisition, and • a second image 32a and a third image acquired by the second camera 12 at respectively the first temporal instant of acquisition and a second instant temporal acquisition, the second temporal instant of acquisition being prior to the first temporal instant of acquisition.

[0150] In a step 22, initial depths associated with a second set of pixels from the second image are predicted from the second and third images by a depth prediction model associated with a monoscopic vision system formed by the second moving camera 12. At least a portion of the second set of pixels then forms part of an edge of the second image.

[0151] In a step 23, a first set of pixels from the first image 3la is projected into the three-dimensional scene to obtain a set of points, each corresponding to a pixel from the first set of pixels, as a function of first coordinates of the pixel in the first image 31a, a second depth associated with the pixel and intrinsic parameters of the first camera 11, the second coordinates of the points of the set of points being defined in a reference frame associated with the first camera 11.

[0152] In a step 24, third coordinates of each point of the point set are determined in a reference frame associated with the second camera 12 as a function of the second coordinates and extrinsic parameters of the stereoscopic vision system.

[0153] In a step 25, fourth pixel coordinates are determined in the second image 32a corresponding to the points as a function of the third coordinates and intrinsic parameters of the second camera 12.

[0154] In a step 26, a subset of pixels from the first set of pixels is selected according to the fourth coordinates and the first depths, the subset of pixels corresponding to the second set of pixels.

[0155] In a step 27, a first cropped image 31b is generated by cropping the first image 31a and a second cropped image 32b is generated by cropping the second image 32a, the cropping of the first image 31a being a function of averages of the first coordinates associated with the pixels of the subset of pixels and the cropping of the second image 32a being a function of a width and a height of the first cropped image 31b.

[0156] Figure 4 schematically illustrates a device 4 configured for generating images for a vehicle equipped with a stereoscopic vision system, for example, for the vehicle 10 of Figure 1, according to a particular and non-limiting embodiment of the present invention. The device 4 corresponds, for example, to a device embedded in the vehicle 10, for example, a computer.

[0157] Device 4 is configured for example to carry out the operations described opposite Figures 1 and 3 and / or the steps described opposite [Fig.2]. Examples of such a device 4 include, but are not limited to, embedded electronic equipment such as a vehicle's on-board computer, an electronic control unit such as an ECU (Electronic Control Unit), a smartphone, a tablet, or a laptop computer. The elements of the device 4, individually or in combination, may be integrated into a single integrated circuit, into several integrated circuits, and / or into discrete components. The device 4 may be implemented as electronic circuits, software (or computer) modules, or a combination of electronic circuits and software modules.

[0158] The device 4 comprises one (or more) processor(s) 40 configured to execute instructions for carrying out the steps of the process and / or for executing instructions from the software embedded in the device 4. The processor 40 may include integrated memory, an input / output interface, and various circuits known to those skilled in the art. The device 4 further comprises at least one memory 41, for example, volatile and / or non-volatile memory, and / or includes a memory storage device that may include volatile and / or non-volatile memory, such as EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disk, or optical disk.

[0159] The computer code of the embedded software(s), including the instructions to be loaded and executed by the processor, is for example stored on memory 4L

[0160] According to various particular and non-limiting embodiments, the device 4 is coupled in communication with other similar devices or systems (for example other computers) and / or with communication devices, for example a TCU (Telematic Control Unit), for example via a communication bus or through dedicated input / output ports.

[0161] According to a particular and non-limiting embodiment, the device 4 includes a block 42 of interface elements for communicating with external devices. The interface elements of the block 42 include one or more of the following interfaces: - radio frequency RF interface, for example of the Wi-Fi® type (according to IEEE 802.11), for example in the 2.4 or 5 GHz frequency bands, or of the Bluetooth® type (according to IEEE 802.15.1), in the 2.4 GHz frequency band, or of the Sigfox type using UBN (Ultra Narrow Band) radio technology, or LoRa in the 868 MHz frequency band, LTE (Long-Term Evolution), LTE-Advanced; - USB interface (from the English "Universal Serial Bus" or "Universal Serial Bus" in French); HD MI interface (from the English "High Definition Multimedia Interface", or "High Definition Multimedia Interface" in French); - LIN interface (from the English "Local Interconnect Network", or in French "Réseau interconnecté local").

[0162] According to another particular and non-limiting embodiment, the device 4 includes a communication interface 43 which enables communication with other devices (such as other computers in the embedded system) via a communication channel 430. The communication interface 43 corresponds, for example, to a transmitter configured to transmit and receive information and / or data via the communication channel 430. The communication interface 43 corresponds, for example, to a wired network of the CAN (Controller Area Network), CAN FD (Controller Area Network Flexible Data-Rate), FlexRay (standardized by ISO 17458) or Ethernet (standardized by ISO / IEC 802-3) type.

[0163] According to a particular and non-limiting embodiment, the device 4 can provide output signals to one or more external devices, such as a display screen 440, touch or not, one or more loudspeakers 450 and / or other peripherals 460 (projection system) via the output interfaces 44, 45, 46 respectively. According to a variant, one or more of the external devices is integrated into the device 4.

[0164] Of course, the present invention is not limited to the embodiments described above but extends to a method for determining a common field of view in two images acquired by a stereoscopic vision system mounted in a vehicle, which would include secondary steps without departing from the scope of the present invention. The same would apply to a device configured for implementing such a method.

[0165] The present invention also relates to a vehicle, for example an automobile or more generally an autonomous land-powered vehicle, comprising the device 4 of [Fig.4].

Claims

1. Demands A method for generating images for a vehicle, said vehicle (10) carrying a stereoscopic vision system comprising a first camera (11) and a second camera (12), each configured to acquire an image of a three-dimensional scene from a different point of view, said method being characterized in that it comprises the following steps: - receipt (21) of data representative of: • a first image (31a) acquired by the first camera (11) at a first temporal instant of acquisition, • a second image (32a) and a third image acquired by the second camera (12) respectively at the first time instant of acquisition and at a second time instant of acquisition, the second time instant of acquisition being prior to said first time instant of acquisition; - prediction (22), from the second and third images, of first depths associated with a second set of pixels of the second image by a depth prediction model associated with a monoscopic vision system formed of the second camera (12) in motion, at least a part of said second set of pixels forming part of an edge of the second image; - projection (23) of a first set of pixels from the first image (31a) into the three-dimensional scene to obtain a set of points each corresponding to a pixel of said first set of pixels as a function of first coordinates of said pixel in the first image (31a), of a second depth associated with said pixel and of intrinsic parameters of the first camera (11), the second coordinates of the points of the set of points being defined in a reference frame associated with the first camera (11); - determination (24) of third coordinates of each point of said set of points in a reference frame associated with the second camera (12) as a function of said second coordinates and extrinsic parameters of the stereoscopic vision system; - determination (25) of fourth pixel coordinates in the second image (32a) corresponding to said points as a function of third coordinates and intrinsic parameters of the second camera (12); - selection (26) of a subset of pixels from the first set of pixels according to the fourth coordinates and the first depths, said subset of pixels corresponding to the second set of pixels; - generation (27) of a first cropped image (31b) by cropping the first image (31a) and generation of a second cropped image (32b) by cropping the second image (32a), the cropping of the first image (31a) being a function of averages of the first coordinates associated with the pixels of said subset of pixels and the cropping of the second image (32a) so as to generate the second cropped image (32b) with a width and a height equal to respectively the width and height of the first cropped image (31b),. the second set of pixels comprising a fourth set of pixels associated with a first edge of the second image (32a), - the first edge corresponding to the left edge when the second camera (12) is placed to the right of the first camera (11), and - the first edge corresponding to the right edge when the second camera (12) is placed to the left of the first camera (11), a left limit of the part of the first cropped image (31b) being defined as an average of the abscissas of the subset of pixels of the first image (31a), each of these pixels being associated with the fourth set of pixels of the second image (32a).

2. A method according to claim 1, wherein the second set of pixels comprises a fifth set of pixels associated with a second edge of the second image (32a), - the second edge corresponding to the upper edge when the second camera (12) is placed below the first camera (H), and - the second edge corresponding to the lower edge when the second camera (12) is placed above the first camera (H).

3. A method according to any one of claims 1 or 2, wherein the width and height of the first cropped image (31b) are each divisible by 32.

4. A method according to any one of claims 1 to 3, wherein the first depths are predicted by a convolutional neural network as a function of a displacement of the vehicle (10) between the second and first time instants of acquisition.

5. A computer program comprising instructions for carrying out the method according to any one of the preceding claims, when such instructions are executed by a processor.

6. Computer-readable recording medium on which is recorded a computer program comprising instructions for carrying out the steps of the process according to claims 1 to 4.

7. Image generation device (4) for vehicle (10), said device (4) comprising a memory (41) associated with at least one processor (40) configured for carrying out the steps of the method according to any one of claims 1 to 4.

8. Vehicle (10) comprising the device (4) according to claim 7.