Method and device for generating images for a vehicle comprising a stereoscopic vision system.
The method of cropping images from a vehicle's stereoscopic vision system to maintain a common field of vision addresses the inefficiencies in existing systems, enabling faster and more accurate depth determination for improved ADAS performance and road safety.
Patent Information
- Application Number
- FR2023012768
- Authority / Receiving Office
- FR · FR
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-21
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2043-11-21
AI Technical Summary
Existing stereoscopic vision systems in vehicles face challenges in processing images efficiently, leading to delays in data updates for Advanced Driver-Assistance Systems (ADAS), which can compromise road safety.
A method and device for generating images by cropping images from two cameras of a stereoscopic vision system to maintain only pixels associated with a common field of vision, allowing for depth determination without rectification, thereby reducing processing time.
The proposed solution facilitates faster and more accurate depth determination, enhancing the operational safety and responsiveness of ADAS systems by reducing the time required for image processing.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Title of the invention: Method and device for generating images for a vehicle comprising a stereoscopic vision system. Technical field
[0001] The present invention relates to methods and devices for generating images for a vehicle comprising an on-board stereoscopic vision system, for example in a motor vehicle. The present invention also relates to a method and a device for determining a common field of vision in two images acquired by a stereoscopic vision system on-board a vehicle. The present invention also relates to a method and a device for controlling one or more AD AS systems on-board a vehicle from a depth determined from images generated for a stereoscopic vision system on-board this same vehicle. Technological background
[0002] Many modern vehicles are equipped with so-called AD AS (Advanced Driver-Assistance System). Such AD AS systems are passive and active safety systems designed to eliminate the element of human error in driving vehicles of all types. AD AS use advanced technologies to assist the driver while driving and thus improve their performance. AD AS use a combination of sensor technologies to perceive the environment around a vehicle, then provide information to the driver or act on certain vehicle systems.
[0003] There are several levels of ADAS, such as rearview cameras and blind spot sensors, lane departure warning systems, adaptive cruise control and automatic parking systems.
[0004] The AD AS on board a vehicle are supplied with data obtained from one or more on-board sensors such as, for example, cameras. These cameras make it possible in particular to detect and locate other road users or possible obstacles present around a vehicle in order, for example: - to adapt the vehicle's lighting depending on the presence of other users; - to automatically regulate the vehicle speed; - to act on the braking system in the event of a risk of impact with an object.
[0005] The determination of depth or distance, from images acquired by a stereoscopic vision system is carried out from images acquired by this vision system. In order to allow the prediction of great depths or distances, for example greater than 200 meters, a stereoscopic vision system on board a vehicle generally comprises at least two cameras several tens of centimeters apart. This distance creates a difference in the field of vision of the cameras of the stereoscopic vision system, the images obtained by the latter must then be processed to enable their use. Such processing most often consists of the rectification of these images.
[0006] Image processing, however, requires a lot of time and resources, while an AD AS powered by depths or distances requires a rapid update of this data while keeping it accurate. The resources allocated to the stereoscopic vision system are however not scalable, as an embedded system must be compact and / or lightweight in order to be integrated into a vehicle.
[0007] Thus, the proper functioning of the driving assistance peripherals using this data depends on the quality of the data transmitted and the speed of image processing by a vision system. Summary of the present invention
[0008] An object of the present invention is to solve at least one of the problems of the technological background described above.
[0009] Another object of the present invention is to reduce the processing time of images acquired by a stereoscopic vision system on board a vehicle.
[0010] Another object of the present invention is to improve road safety, in particular by improving the operational safety of AD AS systems supplied by data obtained from at least one camera.
[0011] According to a first aspect, the present invention relates to a method for generating images for a vehicle, the vehicle carrying a stereoscopic vision system comprising a first camera and a second camera each configured to acquire an image of a three-dimensional scene from a different point of view, the method being characterized in that it comprises the following steps: - receipt of data representative of: • a first image acquired by the first camera at a first acquisition time instant, • a second image and a third image acquired by the second camera respectively at the first acquisition time instant and at a second acquisition time instant, the second acquisition time instant being prior to the first acquisition time instant;- prediction, from the second and third images, of first depths associated with a second set of pixels of the second image by a depth prediction model associated with a monoscopic vision system formed of the second moving camera, at least a part of the second set of pixels forming a part of an edge of the second picture ; - projection of a first set of pixels of the first image into the three-dimensional scene to obtain a set of points each corresponding to a pixel of the first set of pixels as a function of first coordinates of the pixel in the first image, a second depth associated with the pixel and intrinsic parameters of the first camera, second coordinates of the points of the set of points being defined in a reference frame associated with the first camera; - determination of third coordinates of each point of the set of points in a reference system associated with the second camera as a function of the second coordinates and extrinsic parameters of the stereoscopic vision system; - determination of fourth coordinates of pixels in the second image corresponding to the points as a function of the third coordinates and intrinsic parameters of the second camera; - selecting a subset of pixels from the first set of pixels based on the fourth coordinates and the first depths, the subset of pixels corresponding to the second set of pixels; - generation of a first cropped image by cropping the first image and generation of a second cropped image by cropping the second image, the cropping of the first image being a function of averages of the first coordinates associated with the pixels of the subset of pixels and the cropping of the second image being a function of a width and a height of the first cropped image.
[0012] Such a device makes it possible to crop the images acquired by two cameras of a stereoscopic vision system in order to keep only pixels associated with a field of vision common to the two cameras. The determination of depths from the cropped images is thus facilitated and does not require rectification, which represents a step requiring a lot of time and resources.
[0013] According to a variant of the method, the second set of pixels comprises a fourth set of pixels associated with a first edge of the second image, - the first edge corresponding to the left edge when the second camera is placed to the right of the first camera, and - the first edge corresponding to the right edge when the second camera is placed to the left of the first camera.
[0014] According to another variant of the method, the second set of pixels comprises a fifth set of pixels associated with a second edge of the second image, - the second edge corresponding to the top edge when the second camera is placed below the first camera, and - the second edge corresponding to the bottom edge when the second camera is placed above the first camera.
[0015] According to another variant of the method, the cropping of the second image is carried out so as to generate the second cropped image with a width and a height equal to respectively the width and the height of the first cropped image.
[0016] According to yet another variant of the method, the width and height of the first cropped image are each divisible by 32.
[0017] According to an additional variant of the method, the first depths are predicted by a convolutional neural network as a function of a movement of the vehicle between the second and the first acquisition time instants.
[0018] According to a second aspect, the present invention relates to an image generation device for a vehicle carrying a stereoscopic vision system, the device comprising a memory associated with at least one processor configured for implementing the steps of the method according to the first aspect of the present invention.
[0019] According to a third aspect, the present invention relates to a vehicle, for example of the automobile type, comprising a device as described above according to the second aspect of the present invention.
[0020] According to a fourth aspect, the present invention relates to a computer program which comprises instructions adapted for executing the steps of the method according to the first aspect of the present invention, in particular when the computer program is executed by at least one processor.
[0021] Such a computer program may use any programming language and be in the form of source code, object code, or intermediate code between source code and object code, such as in a partially compiled form, or in any other desirable form.
[0022] According to a fifth aspect, the present invention relates to a computer-readable recording medium on which is recorded a computer program comprising instructions for carrying out the steps of the method according to the first aspect of the present invention.
[0023] On the one hand, the recording medium may be any entity or device capable of storing the program. For example, the medium may comprise a storage means, such as a ROM memory, a CD-ROM or a microelectronic circuit type ROM memory, or a magnetic recording means or a hard disk.
[0024] Furthermore, this recording medium may also be a transmissible medium such as an electrical or optical signal, such a signal being able to be conveyed via an electrical or optical cable, by conventional or hertzian radio or by self-directed laser beam or by other means. The computer program according to the present invention may in particular be downloaded from an Internet-type network.
[0025] Alternatively, the recording medium may be an integrated circuit in which the computer program is incorporated, the integrated circuit being adapted to perform or to be used in performing the method in question. Brief description of the figures
[0026] Other characteristics and advantages of the present invention will emerge from the description of the particular and non-limiting exemplary embodiments of the present invention below, with reference to the appended figures 1 to 4, in which:
[0027] [Fig-1] schematically illustrates a stereoscopic vision system embedded in a vehicle, according to a particular and non-limiting exemplary embodiment of the present invention;
[0028] [Fig.2] illustrates a flowchart of the different steps of a method for generating images for the vehicle of [Fig.l], according to a particular and non-limiting exemplary embodiment of the present invention;
[0029] [Fig.3] schematically illustrates images acquired and images generated from a stereoscopic vision system on board the vehicle of [Fig.l], according to a particular and non-limiting exemplary embodiment of the present invention;
[0030] [Fig.4] schematically illustrates a device configured for generating images for a stereoscopic vision system on board the vehicle of [Fig.1], according to a particular and non-limiting exemplary embodiment of the present invention. Description of examples of implementation
[0031] A method and a device for generating images for a vehicle carrying a stereoscopic vision system will now be described in the following with joint reference to FIGS. 1 to 4. The same elements are identified with the same reference signs throughout the description which follows.
[0032] The terms "first(s)", "second(s)" (or "first(s)", "second(s)"), etc. are used in this document by arbitrary convention to enable different elements (such as operations, means, etc.) implemented in the embodiments described below to be identified and distinguished. Such elements may be distinct or correspond to a single element, depending on the embodiment.
[0033] According to a particular and non-limiting example of embodiment of the present invention, a method for generating images for a vehicle carrying a stereoscopic vision system is for example implemented by a computer of the on-board system of the vehicle controlling this stereoscopic vision system.
[0034] The stereoscopic vision system comprises a set of cameras of at least two cameras arranged so as to each acquire an image of a three-dimensional scene from a different point of view, the optical axes representing a orientation of the field of vision of each camera being included in two parallel planes, for example oriented in a non-parallel manner. The stereoscopic vision system is thus said to be non-parallel.
[0035] For this purpose, the method for generating images for a vehicle carrying a stereoscopic vision system comprises the reception of data representative of a first image acquired by the first camera at a first acquisition time instant as well as the reception of data representative of a second image and a third image acquired by the second camera at respectively the same first acquisition time instant and at a second acquisition time instant prior to the first acquisition time instant.
[0036] The method also includes predicting depths associated with pixels of the second image, determining coordinates in the second image of pixels corresponding to pixels of the first image, and selecting a subset of pixels of the first image corresponding to pixels forming part of an edge of the second image.
[0037] A first cropped image and a second cropped image are then generated from the first and second images based on the coordinates associated with the pixels of the subset of pixels and to generate images of the same size.
[0038] [Fig.l] schematically illustrates a stereoscopic vision system equipping a vehicle, according to a particular and non-limiting exemplary embodiment of the present invention.
[0039] Such an environment 1 corresponds, for example, to a road environment formed of a network of roads accessible to the vehicle 10.
[0040] In this example, the vehicle 10 corresponds to a vehicle with a thermal engine, with an electric motor(s) or even a hybrid vehicle with a thermal engine and one or more electric motors. The vehicle 10 thus corresponds, for example, to a land vehicle such as an automobile, a truck, a bus, a motorcycle. Finally, the vehicle 10 corresponds to an autonomous vehicle or not, that is to say a vehicle traveling according to a determined level of autonomy or under the total supervision of the driver.
[0041] The vehicle 10 advantageously comprises several on-board cameras 11, 12, each configured to acquire images of a scene in the environment of the vehicle 10. This set of cameras 11, 12 forms the stereoscopic vision system. Two cameras 11 and 12 are illustrated in [Fig.l]. The present invention is however not limited to a stereoscopic vision system comprising two cameras but extends to any stereoscopic vision system comprising 2 or more cameras, for example 2, 3, 4 or 5 cameras.
[0042] The two cameras 11, 12 have known intrinsic parameters. These parameters consist in particular of: - the focal length of the first camera 11; - the focal length of the second camera 12; - distortions which are due to imperfections in the optical system of each camera; - the direction Cl of the optical axis of the first camera 11; - the direction C2 of the optical axis of the second camera 12; and - the respective resolutions of cameras 11, 12.
[0043] The intrinsic parameters characterize the transformation which associates, for an image point, the camera coordinates with the pixel coordinates, in each camera. These parameters do not change if the camera is moved.
[0044] The intrinsic matrix of the first camera 11 is defined by:
[0045] [Math.l] L 0 0 1
[0046] With: • K the intrinsic matrix of the first camera 11, • / , the focal length of the first camera 11, • the abscissa of the optical center of an image acquired by the first camera 11, and • Cy the ordinate of the optical center of an image acquired by the first camera 11.
[0047] Similarly, the intrinsic matrix of the second camera 12 is defined by:
[0048] [Math.2]
[0049] With: • K the intrinsic matrix of the second camera 12, • / the focal length of the second camera 12, • C the abscissa of the optical center of an image acquired by the second camera 12, and • this is the ordinate of the optical center of an image acquired by the second camera 12.
[0050] The optical axes of the first and second cameras are arranged horizontally, that is to say in two horizontal parallel planes. These two planes are either separated by a height H, H having a zero value if the planes are the same.
[0051] The distortions, which are due to imperfections in the optical system such as defects in the shape and positioning of the camera lenses, will deflect the light beams and therefore induce a positioning deviation for the projected point. compared to an ideal model. It is then possible to complete the camera model by introducing the three distortions that generate the most effects, namely radial, decentering and prismatic distortions, induced by defects in curvature, parallelism of the lenses and coaxiality of the optical axes. In this example, the cameras are assumed to be perfect, that is to say that the distortions are not taken into account or that their correction is processed at the time of image acquisition.
[0052] These two cameras 11, 12 are arranged so as to each acquire an image of a scene from a different point of view, the first point of view is for example located on or in the left rearview mirror of the vehicle 10 or at the top of the windshield of the vehicle 10, the second point of view is for example located on or in the right rearview mirror of the vehicle 10 or at the top of the windshield of the vehicle 10. In the case where the two cameras are located at the top of the windshield of the vehicle, they are then placed at a certain distance. In this example, the first camera 11 is located at the top of the windshield of the vehicle 10, the second camera 12 is located in the right rearview mirror of the vehicle 10.
[0053] A first marker is associated with the first camera 11: - the direction of the x axis is defined horizontal and normal to the optical axis of the first camera 11. The distance B separating the optical center of the first camera 11 from the projection of the optical center of the second camera 12 on the horizontal plane passing through the optical center of the first camera 11 is called the reference base (in English “baseline”); - the direction of the y axis is defined vertical and normal to the optical axis of the first camera 11; - the direction of the z axis is defined orthogonal to the directions of the x and y axes. The three axes x, y and z thus form an orthonormal reference frame.
[0054] The extrinsic parameters linked to the position of the cameras 11, 12 are the following parameters: - three translations in the x, y and z directions: Tx, Ty and Tz constituting the translation vector T with Tx = B, Ty = H and Tz = 0; and - a rotation around the y axis by an angle 0 called the yaw angle.
[0055] A matrix for moving from the reference frame of the first camera 11 to the reference frame of the second camera 12 is thus defined by:
[0056] [Math.3] ' COS# 0 sin# B' 0 1 0 H . - sin# 0 cos# 0.
[0057] With: • [R, T] the matrix to move from the frame of reference of the first camera 11 to the frame of reference of the second camera 12, • 0 yaw angle, • B the reference base, and • H the height difference between the first 11 and second 12 cameras.
[0058] The extrinsic parameters are determined, for example, during a calibration phase of the stereoscopic vision system.
[0059] A main constraint of the stereoscopic vision system used in the automobile is, for example, the large distance between the two cameras. Indeed, to be able to cover a measurement range of 200 meters, the reference base must reach 60cm for the cameras commonly used in this field.
[0060] The two cameras 11, 12 acquire images of a scene located in front of the vehicle 10, the first camera 11 covering only a first acquisition field 13, the second camera 12 covering only a second acquisition field 14 and the two cameras 11, 12 both covering a third acquisition field 15. The first and third acquisition fields 13, 15 thus allow a monoscopic vision of the scene by the first camera 11, the second and third acquisition fields 14, 15 allow a monoscopic vision of the scene by the second camera 12 and the third acquisition field 15 allows a stereoscopic vision of the scene by the stereoscopic vision system composed of the two cameras 11, 12.
[0061] An obstacle 18 is placed in the acquisition field of the cameras, for example in the third acquisition field 15. The presence of the obstacle 18 defines an occlusion field for the stereoscopic vision system composed here of the three fields 16, 17 and 19.
[0062] Among these three fields, field 16 is visible from the second camera 12. The part of the scene present in this field 16 is therefore observable using the monoscopic vision system composed of the second camera 12.
[0063] The field 17 is visible from the first camera 11. The part of the scene present in this field 17 is therefore observable using the monoscopic vision system composed of the first camera 11.
[0064] Finally, field 19 is not visible from any of the cameras. The part of the scene present in this field 19 is therefore not observable.
[0065] It is obvious that it is possible to use such a stereoscopic vision system to take images of scenes located on the sides or behind the vehicle 10 by equipping it with differently placed and oriented cameras.
[0066] The images acquired by the cameras 11, 12 at an acquisition time instant are presented in the form of data representing pixels characterized by: - coordinates in each image; and - data relating to the colors and brightness of objects in the observed scene in the form, for example, of RGB (from the English “Red Green Blue”) or TSL (Tone, Saturation, Brightness) colorimetric coordinates.
[0067] The images acquired by the cameras 11, 12 represent views of the same scene taken from different viewpoints, the positions of the cameras being distinct. On this scene are for example: - buildings; - road infrastructure; - other stationary users, for example a parked vehicle; and / or - other mobile users, for example another vehicle, a cyclist or a moving pedestrian.
[0068] These images are sent to a computer of a device equipping the vehicle 10 or stored in a memory of a device accessible to a computer of a device equipping the vehicle 10.
[0069] An image generation process for the vehicle 10 is advantageously implemented by the vehicle 10, i.e. by a computer or a combination of computers of the on-board system of the vehicle 10, for example by the computer(s) in charge of the stereoscopic vision system of the vehicle 10.
[0070] In a first operation, the calculator receives first data representative of: • a first image 31a acquired by the first camera 11 at a first acquisition time instant and • a second image 32a and a third image acquired by the second camera 12 respectively at the first acquisition time instant and at a second acquisition time instant, the second acquisition time instant being prior to the first acquisition time instant.
[0071] The first 31a and second 32a images form a first pair of stereoscopic images, that is to say images acquired by two separate cameras at the same time instant.
[0072] In a second operation, first depths associated with a second set of pixels of the second image are predicted from the second and third images, by a depth prediction model associated with a monoscopic vision system formed from the second moving camera 12.
[0073] At least a portion of the second set of pixels forms a portion of an edge of the second image.
[0074] According to a first particular exemplary embodiment, the second set of pixels comprises a fourth set of pixels associated with a first edge of the second image 32a, - the first edge corresponding to the left edge when the second camera 12 is placed to the right of the first camera 11, and - the first edge corresponding to the right edge when the second camera 12 is placed to the left of the first camera 11.
[0075] According to a second particular exemplary embodiment, the second set of pixels comprises a fifth set of pixels associated with a second edge of the second image 32a, - the second edge corresponding to the upper edge when the second camera 12 is placed below the first camera 11, and - the second edge corresponding to the lower edge when the second camera 12 is placed above the first camera 11.
[0076] The first and second particular embodiments are, according to a third particular embodiment, combined. Thus, the second set of pixels comprises both the fourth set of pixels and the fifth set of pixels, for example when the second camera 12 is below to the right of the first camera 11.
[0077] According to a particular exemplary embodiment, the first depths are predicted by a convolutional neural network as a function of a movement of the vehicle 10 between the second and the first time instants of acquisition, by any depth prediction model associated with a monoscopic vision system such as packnet® which achieves metric precision with supervision of the prediction model of the movement of the second camera with the speed of the vehicle.
[0078] The first depths are predicted in the frame of reference of the second camera 12. They are then transcribed into the frame of reference of the first camera 11 by the following function:
[0079] [Math.4]
[0080] Y1 7l With : coordinates of a point in the three-dimensional scene in the reference frame of the first camera 11, • [ -7? - T ] a matrix associated with the change of reference frame, from the reference frame of the second camera 12 to the reference frame of the first camera 11, defined as the matrix opposite to the matrix for moving from the reference frame of the first camera 11 to the reference frame of the second camera 12 previously defined, * K") the inverse of the intrinsic matrix of the second camera 12, a homogeneous vector comprising the coordinates of a pixel of the second image 32a having xr for abscissa and ordinate, and • a first depth associated with the pixel of the second image.
[0081] In a third operation, a first set of pixels of the first image 31a is projected into the three-dimensional scene to obtain a set of points each corresponding to a pixel of the first set of pixels as a function of first coordinates of the pixel in the first image 31a, a second depth associated with the pixel and intrinsic parameters of the first camera 11. Second coordinates of the points of the set of points are defined in a reference system associated with the first camera 11.
[0082] According to a particular exemplary embodiment, the first set of pixels comprises all the pixels of the first image.
[0083] According to another particular exemplary embodiment, the first set of pixels comprises a portion of the pixels of the first image, for example the lower half of the first image 31a if the first camera 11 is located above the second camera 12.
[0084] A pixel of the first image 31a is for example projected into the three-dimensional scene. ional by the following function:
[0085] [Math.5] -vi
[0086] With: coordinates of a point in the three-dimensional scene in the reference frame Y1 1 w of the first camera 11, • K' / 1 the inverse matrix of the intrinsic matrix of the first camera 11, a homogeneous vector comprising the coordinates of a pixel of the first 11 image 31a having xi for abscissa and ordinate, • Z!w a second depth associated with the pixel of the first image, • dx the abscissa of the optical center of an image acquired by the first camera 11, and • the ordinate of the optical center of an image acquired by the first camera 11.
[0087] It should be noted that if the origin of a reference frame associated with the first image 31a is located at the top left of the first image 31a, then the coordinates of the optical center of the first image 31a are defined by:
[0088] [Math.6] And [Math.6]
[0089] With: • dx the abscissa of the optical center of the first image 31a, • c{ the ordinate of the optical center of the first image 31a, • wi the width of the first image 31a, and • hj the height of the first image 31a.
[0090] Similarly, for the second image 32a, if the origin of a reference frame associated with the second image 32a is located at the top left of the second image 32a, then the coordinates of the optical center of the second image 32a are defined by:
[0091] [Math.7] And [Math.7]
[0092] With: • cx the abscissa of the optical center of the second image 32a, • d the ordinate of the optical center of the second image 32a, • the width of the second image 32a, and • hr the height of the second image 32a.
[0093] In a fourth operation, third coordinates of each point of the set of points are determined in a reference system associated with the second camera 12 as a function of the second coordinates and extrinsic parameters of the stereoscopic vision system.
[0094] The transition from second coordinates expressed in the frame of reference of the first camera to third coordinates expressed in the frame of reference of the second
[0095] camera 12 is obtained by the following function: [Math. 8] VV vl y / 1 W 1 . ■ cosf? 0 . - sin 7 X w cos8+Z^sin# + B either [Math. 8] ) cos# Z + ZyySin# + B Æ W Yr IV f,
[0096] With : coordinates of a point in the three-dimensional scene in the reference frame Yw of the second camera 12, • [R, T] the matrix to move from the frame of reference of the first camera 11 to the frame of reference of the second camera 12, Z' 7 1 a homogeneous vector comprising coordinates of a point of the scene W V 11'
[0097]
[0098] three-dimensional in the frame of reference of the first camera 11, • 0 yaw angle, • B the reference base, and • H the height difference between the first 11 and second 12 cameras. In a fifth operation, fourth pixel coordinates are determined in the second image 32a corresponding to the points as a function of the third coordinates and intrinsic parameters of the second camera 12. The fourth coordinates are obtained by the following function:
[0099]
[0100] [Math.9] With : Yw fr 0 0 fr 0 0 here 1 X\v Yw Z'w +jf / j . H ..( .41.^^ + Zlwcosd\ « \ A / - r ^■r a homogeneous vector comprising the homogeneous coordinates of a pixel of the second image 32a having X? for homogeneous abscissa, for homogeneous ordinate and zr for homogeneous depth, • Kr the intrinsic matrix of the second camera 12, coordinates of a point in the three-dimensional scene in the reference frame Zw. of the second camera 12, • f the focal length of the second camera 12, • dx the abscissa of the optical center of the first image 31a, • cx the ordinate of the optical center of the first image 31a, • the abscissa of the optical center of the second image 32a, • Cy the ordinate of the optical center of the second image 32a, • 0 yaw angle, • B the reference base, • H the height difference between the first 11 and second 12 cameras, and • Zlw a depth of a point in the frame of the first camera 11.
[0101]
[0102]
[0103] The following function converts homogeneous coordinates into homogeneous coordinates in the second image 32a: [Math. 10] With : the coordinates of a pixel in the second image having xr as abscissa and yr for ordinate, • Xf a homogeneous abscissa, y* a homogeneous ordinate and zr a homogeneous depth, • f the focal length of the second camera 12, • Clx the abscissa of the optical center of the first image 31a, • c(, the ordinate of the optical center of the first image 31a, • c* the abscissa of the optical center of the second image 32a, • Cy the ordinate of the optical center of the second image 32a, • 0 yaw angle, • B the reference base, • H the height difference between the first 11 and second 12 cameras, and • Z!w a depth of a point in the frame of the first camera 11.
[0104] In a sixth operation, a subset of pixels of the first set of pixels is selected based on the fourth coordinates and the first depths, the subset of pixels corresponding to the second set of pixels.
[0105] If the first camera 11 is placed to the left of the second camera 12, then the second set of pixels comprises a fourth set of pixels belonging to the right edge of the second image 32a. The first vertical row of pixels of the second image 32a is called the right edge of the second image 32a. These pixels of the fourth set of pixels are defined by the following equation:
[0106] [Math. 11] xr = 0 With: [Math. 11] the abscissa of a pixel of the second image 32a.
[0107] By integrating this equation into the previous equations, we obtain the following equation to associate the pixels of the first image 31a with the right edge of the second image 32a:
[0108] [Math. 12] 2^y-sin&Ycos0 |
[0109] With: • Xl min the abscissa of a pixel of the first image 31a associated with the right edge of the second image 32a, • the focal length of the first camera 11, • f the focal length of the second camera 12, • dx the abscissa of the optical center of the first image 31a, • d the abscissa of the optical center of the second image 32a, • 0 yaw angle, • B the reference base, • Zlw a depth of a point in the frame of the first camera 11.
[0110] Similarly, if the first camera 11 is placed below the second camera 12, then the second set of pixels comprises a fifth set of pixels belonging to the lower edge of the second image 32a. The last horizontal row of pixels of the second image 32a is called the lower edge of the second image 32a. These pixels of the fifth set of pixels are defined by the following equation: [YES] [Math. 13] y = y - r rjnax With : [Math. 13] v. the ordinate of a pixel of the second image 32a, and [Math. 13] y - r_max the ordinate of a pixel of the lower edge of the second image 32a.
[0112] By integrating this equation into the previous equations, we obtain the following equation to associate the pixels of the first image 31a with the lower edge of the second image 32a:
[0113] [Math. 14] + Cy = y^max
[0114] With: • yr_max the ordinate of a pixel of the lower edge of the second image 32a, • yimax the ordinate of a pixel of the first image 31a associated with the lower edge of the second image 32a, • / ; the focal length of the first camera 11, • f the focal length of the second camera 12, • dx the abscissa of the optical center of the first image 31a, • Cy the ordinate of the optical center of the first image 31a, • This is the ordinate of the optical center of the second image 32a, • 0 yaw angle, • H the height difference between the first 11 and second 12 cameras, and • Zlw a depth of a point in the frame of the first camera 11.
[0115] Conversely, if the first camera 11 is placed above the second camera 12, then the second set of pixels comprises a fifth set of pixels belonging to the upper edge of the second image 32a. The first horizontal row of pixels of the second image 32a is called the upper edge of the second image 32a. These pixels of the fifth set of pixels are defined by the following equation:
[0116] [Math. 15] v =0 'r With : [Math. 15] » the ordinate of a pixel of the second image 32a.
[0117] By integrating this equation into the previous equations, we obtain the following equation to associate the pixels of the first image 31a with the upper edge of the second image 32a:
[0118] [Math. 16]
[0119] With: • the ordinate of a pixel of the first image 31a associated with the upper edge of the second image 32a, • f} the focal length of the first camera 11, • f the focal length of the second camera 12, • here the abscissa of the optical center of the first image 31a, • the ordinate of the optical center of the first image 31a, • Cy the ordinate of the optical center of the second image 32a, • 0 yaw angle, • H the height difference between the first 11 and second 12 cameras, and • Z!w a depth of a point in the frame of the first camera 11.
[0120] In a seventh operation, a first cropped image 31b is generated by a cropping the first image 31a and a second cropped image 32b is generated by cropping the second image 32a. The cropping of the first image 31a is a function of averages of the first coordinates associated with the pixels of the subset of pixels and the cropping of the second image 32a is a function of a width and a height of the first cropped image 31b.
[0121] As illustrated in [Fig. 3], the first cropped image 31b comprises a portion of the first image 31a. During cropping, a left portion 311 of the first image 31a is cropped as well as an upper or lower portion 312 of the first image 31a depending on the relative position of the first camera 11 relative to the second camera 12. Similarly, the second cropped image 32b comprises a portion of the second image 32a. During cropping, a right portion 321 of the second image 32a is cropped as well as a lower or upper portion 322 of the first image 31a depending on the relative position of the first camera 11 relative to the second camera 12. The illustrated example thus corresponds to a first camera 11 placed below the second camera 12.
[0122] The left limit of the part of the first cropped image 31b is defined as the average of the abscissas of the subset of pixels of the first image 31a equal to xï min, each of these pixels being associated with the fourth set of pixels of the second image 32a.
[0123] Similarly, if the first camera 11 is below the second camera 12 (figure 3), the lower limit of the first cropped image 31b is defined as the average of the ordinates of the subset of pixels of the first image 31a, each of these pixels being associated with the fifth set of pixels of the second image 32a equal to y^ max. In other words, according to this particular exemplary embodiment, the first cropped image 31b comprises the pixels of the first image 31a whose abscissa is greater than xf min and whose ordinate is less than y^ mav.
[0124] Conversely, if the first camera 11 is above the second camera 12, the upper limit of the first cropped image 31b is defined as the average of the ordinates of the subset of pixels of the first image 31a, each of these pixels being associated with the fifth set of pixels of the second image 32a equal to y^ min . In other words, according to this particular exemplary embodiment, the first cropped image 31b comprises the pixels of the first image 31a whose abscissa is greater than xf min and whose ordinate is greater than y '^ min.
[0125] According to a particular exemplary embodiment, the cropping of the second image 32a is carried out so as to generate the second cropped image 32b with a width and a height equal to respectively the width and the height of the first cropped image 31b. Thus, if the first cropped image 31b has a width w' and a height h', then the second cropped image 32b has the same dimensions, that is to say that the second cropped image 32b has width w' and height h'.
[0126] According to a particular exemplary embodiment, the width and height of the first cropped image 31b are each divisible by 32. This makes it possible in particular to subsequently apply certain depth prediction models for a stereoscopic system using an optical flow calculation method.
[0127] Thus, the abscissas and ordinates corresponding to the previously determined cropping limits are determined by the following functions:
[0128] [Math. 17] + (^2- J ^32^ And
[0129] y* = lv 1 + (32-1 y 1 %32) s' 'a First camera 11 is below Ijnax L - !_max J \ ~ l_max j / of the second camera 12, or
[0130] [Math. 18] y*. = h-(\hy. . | + (32-| / ?-y; . |%32H • Ijnm \ L Ùwn J \ L i_jnm J / / if the first camera 11 is above the second camera 12.
[0131] With: an abscissa corresponding to the lower edge of the first cropped image 31b, rnax an ordinate corresponding to the left edge of the first cropped image 31b, y* an ordinate corresponding to the right edge of the first cropped image 31b, [J an operator allowing to obtain the first integer less than its argument, and %32 corresponding to the remainder of a division of an argument by 32.
[0132] According to a particular embodiment, the determination of the abscissas and ordinates corresponding to the cropping limits are determined for several sets of images acquired by the first 11 and second 12 cameras, that is to say that this process is repeated several times. At each iteration of the process, data representative of a first image of a set of images Ei acquired by the first camera 11 at a third acquisition time instant and of second and third images of the same set of images Ei acquired by the second camera 12 respectively at the third acquisition time instant and at an acquisition time instant prior to the third acquisition time instant. The second to seventh operations are then iterated for each of the sets of images Ei independently, making it possible to obtain for each set of images Ei an abscissa xî min (Ei) and an ordinate y* (Ei) when the first camera 11 is below of the second camera 12 or y* ( Ei ) when the first camera 11 is above - l_min v of the second camera 12 corresponding to the limits allowing the cropping of the first image of each set of images Ei.
[0133] In order to standardize the cropping of the images acquired by the stereoscopic vision system, generalized limits (x';y') are determined and are for example equal to the respective averages of the abscissas Xj min(Ei) and the ordinates y* (Ei) or y* . (Ei) obtained by the following functions: l_min v
[0134] [Math. 19] x' = mean^ min ( Ei ) ) + ( 32 - [ mean (x].,„n(®))J%32))
[0135] [Math.20] y' = mean^* ( Ei ) ) + ( 32 -1 mean(y* (Ei)] I %32 ) j V l—Wiax ■ 7 / \ L " v l_maxk 7 / j / / when the first camera 11 is below the second camera 12, or
[0136] [Math.21] y = mean^y* min(£ / )^+^32-^ mean( y* min ( Ei ) jj %32) j when the first camera 11 is above the second camera 12
[0137] With: x a generalized abscissa, Y a generalized ordinate, xî mm(Ei) corresponding to the lower edge of the first cropped image of a set of images Ei, y* (Ei) an ordinate corresponding to the left edge of the first cropped image of a set of images Ei, y* . ( Ei ) an ordinate corresponding to the right edge of the first cropped image of a set of images Ei, [J an operator allowing to obtain the first integer less than its argument, and %32 corresponding to the remainder of a division of an argument by 32.
[0138] According to another example, in order to obtain cropped images whose widths and heights are divisible by 32, generalized limits (x';y') are determined and are for example equal to the respective averages of the abscissas x*min(Ei) and the ordinates y* f Ei) or y* (Ei) obtained by the following functions: -7 l_maxv 7 !_mm■ 7
[0139] [Math.22] x' = average^ mà^Ei) )
[0140] [Math.23] y' — averageï y* ( Ei) ) \ - l_max ' when the first camera 11 is below the second camera 12, or
[0141] [Math.24] y' = mean ( y* . (Ei ) )
[0142] With: • x a generalized abscissa, • y a generalized ordinate, * xîmtn^^) corresponding to the lower edge of the first cropped image of a set of images El, • y* (Ei) an ordinate corresponding to the left edge of the first cropped image of a set of images Ei, • y* tj (Ei) an ordinate corresponding to the right edge of the first cropped image of a set of images Ei.
[0143] The generalized limits are then used to automatically crop images acquired by the first 11 and second 12 cameras of the on-board stereoscopic vision system.
[0144] Thus, the process of generating images for a vehicle carrying a stereoscopic vision system makes it possible to generate pairs or couples of cropped stereoscopic images, that is to say, the pixels of each of the images of the pair of stereoscopic images are associated with the same part of a three-dimensional scene seen by the two cameras forming the stereoscopic vision system. The pixels of the first 31b and second 32b cropped images are thus associated with objects of the three-dimensional scene present both in the field of vision of the first camera 11 and in the field of vision of the second camera 12.
[0145] A depth prediction process associated with the stereoscopic vision system on board the vehicle 10 can for example use these cropped images 31b, 32b to predict depths associated with the pixels of these images without having to resort to a rectification operation, making the execution of this depth prediction process faster.
[0146] An AD AS receiving the depth data predicted by this depth prediction process associated with the stereoscopic vision system on board the vehicle 10 then has rapidly updated information and is then more responsive. Its operation is then safer and more precise.
[0147] [Fig.2] illustrates a flowchart of the different steps of a method 2 for generating images for a vehicle carrying a stereoscopic vision system, by example for the vehicle 10 of [Fig.l], according to a particular and non-limiting exemplary embodiment of the present invention.
[0148] Method 2 is for example implemented by one or more processors of one or more computers on board the vehicle 10, for example by a computer controlling the stereoscopic vision system.
[0149] In a step 21, data representative of: • a first image 31a acquired by the first camera 11 at the first time instant of acquisition, and • a second image 32a and a third image acquired by the second camera 12 at respectively the first acquisition time instant and a second acquisition time instant, the second acquisition time instant being prior to the first acquisition time instant.
[0150] In a step 22, first depths associated with a second set of pixels of the second image are predicted from the second and third images, by a depth prediction model associated with a monoscopic vision system formed from the second moving camera 12. At least a part of the second set of pixels then forms a part of an edge of the second image.
[0151] In a step 23, a first set of pixels of the first image 31a is projected into the three-dimensional scene to obtain a set of points, each corresponding to a pixel of the first set of pixels, as a function of first coordinates of the pixel in the first image 31a, of a second depth associated with the pixel and of intrinsic parameters of the first camera 11, second coordinates of the points of the set of points being defined in a reference system associated with the first camera 11.
[0152] In a step 24, third coordinates of each point of the set of points are determined in a reference system associated with the second camera 12 as a function of the second coordinates and extrinsic parameters of the stereoscopic vision system.
[0153] In a step 25, fourth pixel coordinates are determined in the second image 32a corresponding to the points as a function of the third coordinates and intrinsic parameters of the second camera 12.
[0154] In a step 26, a subset of pixels of the first set of pixels is selected according to the fourth coordinates and the first depths, the subset of pixels corresponding to the second set of pixels.
[0155] In a step 27, a first cropped image 31b is generated by cropping the first image 31a and a second cropped image 32b is generated by cropping the second image 32a, the cropping of the first image 31a being a function of averages of the first coordinates associated with the pixels of the sub- set of pixels and the cropping of the second image 32a being a function of a width and a height of the first cropped image 31b.
[0156] [Fig. 4] schematically illustrates a device 4 configured for generating images for a vehicle carrying a stereoscopic vision system, for example for the vehicle 10 of [Fig. 1], according to a particular and non-limiting exemplary embodiment of the present invention. The device 4 corresponds for example to a device on board the vehicle 10, for example a computer.
[0157] The device 4 is for example configured for the implementation of the operations described with regard to figures 1 and 3 and / or the steps described with regard to [Fig.2]. Examples of such a device 4 include, but are not limited to, on-board electronic equipment such as an on-board computer of a vehicle, an electronic calculator such as an ECU (“Electronic Control Unit”), a smartphone, a tablet, a laptop. The elements of the device 4, individually or in combination, can be integrated into a single integrated circuit, into several integrated circuits, and / or into discrete components. The device 4 can be produced in the form of electronic circuits or software (or computer) modules or even a combination of electronic circuits and software modules.
[0158] The device 4 comprises one (or more) processor(s) 40 configured to execute instructions for carrying out the steps of the method and / or for executing the instructions of the software(s) embedded in the device 4. The processor 40 may include integrated memory, an input / output interface, and various circuits known to those skilled in the art. The device 4 further comprises at least one memory 41 corresponding for example to a volatile and / or non-volatile memory and / or comprises a memory storage device which may comprise volatile and / or non-volatile memory, such as EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic or optical disk.
[0159] The computer code of the embedded software(s) comprising the instructions to be loaded and executed by the processor is for example stored in the 4L memory.
[0160] According to various particular and non-limiting embodiments, the device 4 is coupled in communication with other similar devices or systems (for example other computers) and / or with communication devices, for example a TCU (from the English “Telematic Control Unit” or in French “Telematic Control Unit”), for example via a communication bus or through dedicated input / output ports.
[0161] According to a particular and non-limiting exemplary embodiment, the device 4 comprises a block 42 of interface elements for communicating with external devices. The interface elements of the block 42 comprise one or more of the interfaces following: - RF radio frequency interface, for example Wi-Fi® type (according to IEEE 802.11), for example in the 2.4 or 5 GHz frequency bands, or Bluetooth® type (according to IEEE 802.15.1), in the 2.4 GHz frequency band, or Sigfox type using UBN (Ultra Narrow Band) radio technology, or LoRa in the 868 MHz frequency band, LTE (Long-Term Evolution), LTE-Advanced; - USB interface (from the English “Universal Serial Bus” or “Universal Serial Bus” in French); HD MI interface (from the English “High Definition Multimedia Interface” or “High Definition Multimedia Interface” in French); - LIN interface (from the English “Local Interconnect Network”).
[0162] According to another particular and non-limiting exemplary embodiment, the device 4 comprises a communication interface 43 which makes it possible to establish communication with other devices (such as other computers of the on-board system) via a communication channel 430. The communication interface 43 corresponds for example to a transmitter configured to transmit and receive information and / or data via the communication channel 430. The communication interface 43 corresponds for example to a wired network of the CAN (from the English "Controller Area Network" or in French "Réseau de contrôles") type, CAN FD (from the English "Controller Area Network Flexible Data-Rate" or in French "Réseau de contrôles à débit de données flexible"), FlexRay (standardized by the ISO 17458 standard) or Ethernet (standardized by the ISO / IEC 802-3 standard).
[0163] According to a particular and non-limiting exemplary embodiment, the device 4 can provide output signals to one or more external devices, such as a display screen 440, touch-sensitive or not, one or more speakers 450 and / or other peripherals 460 (projection system) via the output interfaces 44, 45, 46 respectively. According to a variant, one or other of the external devices is integrated into the device 4.
[0164] Of course, the present invention is not limited to the exemplary embodiments described above but extends to a method for determining a common field of vision in two images acquired by a stereoscopic vision system on board a vehicle, which would include secondary steps without thereby departing from the scope of the present invention. The same would apply to a device configured for implementing such a method.
[0165] The present invention also relates to a vehicle, for example an automobile or more generally an autonomous land-powered vehicle, comprising device 4 of [Fig.4].
Claims
1. Claims Method for generating images for a vehicle, said vehicle (10) carrying a stereoscopic vision system comprising a first camera (11) and a second camera (12) each configured to acquire an image of a three-dimensional scene from a different point of view, said method being characterized in that it comprises the following steps: - reception (21) of data representative of: • a first image (31a) acquired by the first camera (11) at a first acquisition time instant, • a second image (32a) and a third image acquired by the second camera (12) respectively at the first acquisition time instant and at a second acquisition time instant, the second acquisition time instant being prior to said first acquisition time instant; - prediction (22), from the second and third images, of first depths associated with a second set of pixels of the second image by a depth prediction model associated with a monoscopic vision system formed of the second camera (12) in motion, at least a part of said second set of pixels forming a part of an edge of the second image; - projection (23) of a first set of pixels of the first image (31a) into the three-dimensional scene to obtain a set of points each corresponding to a pixel of said first set of pixels as a function of first coordinates of said pixel in the first image (31a), of a second depth associated with said pixel and of intrinsic parameters of the first camera (11), second coordinates of the points of the set of points being defined in a reference system associated with the first camera (11); - determination (24) of third coordinates of each point of said set of points in a reference system associated with the second camera (12) as a function of said second coordinates and extrinsic parameters of the stereoscopic vision system; - determination (25) of fourth coordinates of pixels in the second image (32a) corresponding to said points as a function of the third coordinates and intrinsic parameters of the second camera (12); - selection (26) of a subset of pixels from the first set of pixels as a function of the fourth coordinates and the first depths, said subset of pixels corresponding to the second set of pixels; - generation (27) of a first cropped image (31b) by cropping the first image (31a) and generation of a second cropped image (32b) by cropping the second image (32a), the cropping of the first image (31a) being a function of averages of the first coordinates associated with the pixels of said subset of pixels and the cropping of the second image (32a) being a function of a width and a height of the first cropped image (31b).
2. The method of claim 1, wherein the second set of pixels comprises a fourth set of pixels associated with a first edge of the second image (32a), - the first edge corresponding to the left edge when the second camera (12) is placed to the right of the first camera (11), and - the first edge corresponding to the right edge when the second camera (12) is placed to the left of the first camera (11).
3. Method according to claim 1 or 2, for which the second set of pixels comprises a fifth set of pixels associated with a second edge of the second image (32a), - the second edge corresponding to the upper edge when the second camera (12) is placed below the first camera (11), and - the second edge corresponding to the lower edge when the second camera (12) is placed above the first camera (11).
4. Method according to one of claims 1 to 3, for which the cropping of the second image (32a) is carried out so as to generate the second cropped image (32b) with a width and a height equal to respectively the width and the height of the first cropped image (31b).
5. Method according to one of claims 1 to 4, for which the width and the height of the first cropped image (31b) are each divisible by
6. JL. Method according to one of claims 1 to 5, for which the first depths are predicted by a convolutional neural network as a function of a movement of the vehicle (10) between the second and the first acquisition time instants.
7. Computer program comprising instructions for implementing implementation of the method according to any one of the preceding claims, when these instructions are executed by a processor.
8. A computer-readable recording medium on which is recorded a computer program comprising instructions for carrying out the steps of the method according to claims 1 to 6.
9. Device (4) for generating images for a vehicle (10), said device (4) comprising a memory (41) associated with at least one processor (40) configured for implementing the steps of the method according to any one of claims 1 to 6.
10. Vehicle (10) comprising the device (4) according to claim 9.
Citation Information
Patent Citations
Method and device for retargeting a 3D content
US9743062B2