Method and device for learning a depth prediction model associated with a pinhole vision system embedded in a vehicle
The method optimizes the learning process for depth prediction in vehicle ADAS systems by using a convolutional neural network with stereoscopic vision, reducing computation time and enhancing prediction accuracy through image resizing and error minimization.
Patent Information
- Application Number
- FR2024005463
- Authority / Receiving Office
- FR · FR
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-05-28
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2044-05-28
AI Technical Summary
Learning a depth prediction model for vehicle ADAS systems using onboard cameras is tedious and requires significant time and high-quality training data, necessitating an optimization of the training phase to reduce duration while improving reliability.
A method for determining depth using a convolutional neural network with a stereoscopic vision system, involving image resizing and pixel set determination to minimize loss errors, optimizing computation time and enhancing prediction accuracy.
The method efficiently predicts depth by minimizing errors and optimizing computation time, resulting in a reliable depth prediction model for vehicle ADAS systems.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Title of the invention: Method and device for learning a depth prediction model associated with a pinhole vision system embedded in a vehicle. Technical field
[0001] The present invention relates to methods and devices for determining depth using a vision system mounted in a vehicle, for example, in a motor vehicle. The present invention also relates to a method and device for measuring the distance between an object and a vehicle equipped with a vision system. Technological background
[0002] Many modern vehicles are equipped with Advanced Driver-Assistance Systems (ADAS). Such ADAS systems are passive and active safety systems designed to eliminate human error in driving all types of vehicles. ADAS systems use advanced technologies to assist the driver while driving and thus improve performance. ADAS systems use a combination of sensor technologies to perceive the environment around a vehicle and then provide information to the driver or act on certain vehicle systems.
[0003] There are several levels of ADAS, such as reversing cameras and blind spot sensors, lane departure warning systems, adaptive cruise control or automatic parking systems.
[0004] The AD AS systems embedded in a vehicle are powered by data obtained one or more onboard sensors, such as cameras. These cameras make it possible to detect and locate other road users or potential obstacles around a vehicle in order to, for example: • to adapt the vehicle's lighting according to the presence of other road users; • to automatically regulate the vehicle's speed; • to act on the braking system in case of risk of impact with an object.
[0005] Learning a depth prediction model associated with these cameras is tedious and requires both a large amount of high-quality training data and time. In order to learn the depth prediction model from data acquired by the vision system itself, the training phase should be optimized to reduce its duration while improving its reliability. Summary of the present invention
[0006] One object of the present invention is to solve at least one of the problems of the technological background described above.
[0007] Another object of the present invention is to improve the quality of data from the processing of an image acquired by a vision system, in particular by a depth prediction model implemented by a neural network associated with a stereoscopic vision system.
[0008] Another object of the present invention is to improve road safety, in particular by improving the reliability of AD AS systems powered by data obtained from a vision system.
[0009] According to a first aspect, the present invention relates to a method for determining the depth of a pixel of an image by a depth prediction model implemented by a convolutional neural network associated with a vision system embedded in a vehicle, the vision system comprising a first and a second camera so as to each acquire an image of a three-dimensional scene from a different point of view, said method being implemented by at least one processor, and being characterized in that the depth prediction model is learned in a learning phase comprising the following steps: - reception of representative data of a first image and a second image acquired by respectively the first camera and the second camera at the same time instant of acquisition; - determination of a width and height of a pixel search window as a function of image coordinates in the second image of a pixel of the first image determined from an analytical function and depths between a minimum depth and a maximum depth; - determination, for each first pixel of a set of pixels of the first image, of a first target set of pixels in the second image as a function of the coordinates of the first pixel and the pixel search window, and determination, for each second pixel of a set of pixels of the second image, of a second target set of pixels in the first image as a function of the coordinates of the second pixel and the pixel search window; - analytical determination of an abscissa in the first image of a first target pixel corresponding to a first reference pixel belonging to a lateral edge of the second image, the first target pixel being included in a target set of pixels associated with the first reference pixel, and of an ordinate in the second image of a second target pixel corresponding to a second reference pixel belonging to a horizontal edge of the first image, the second target pixel being included in a target set of pixels associated with the second reference pixel, as a function of a determined depth and geometric parameters, said geometric parameters including intrinsic and extrinsic parameters of the first and second cameras; - generation of a first image resized by cropping the first image and a second image resized by cropping the second image according to the x-coordinate of the first target pixel and the y-coordinate of the second target pixel, the first resized image and the second resized image being of the same definition; - prediction of a depth associated with each first pixel by the depth prediction model from the first resized image and the first target set and prediction of a depth associated with each second pixel by the depth prediction model from the second resized image and the second target set; - analytical generation of a third image from the first resized image, the depths associated with the first pixels and the geometric parameters and generation of a fourth image from the second resized image, the depths associated with the second pixels and the geometric parameters; - determination of a first error associated with each first pixel by comparing pixels from the first resized image and the fourth image, and of a second error associated with each second pixel by comparing pixels from the second resized image and the third image; - learning the depth prediction model by minimizing a loss error determined from the first and second errors.
[0010] Such a learning process gains efficiency through image resizing to process only pixels located within a field of view common to both cameras and by searching for corresponding pixels in target sets of pixels of limited size. The depth prediction model thus learned makes it possible to accurately predict the depth of a pixel in an image acquired by one of the cameras of the stereoscopic vision system, that is, the distance separating the vehicle carrying the camera from a physical object in the three-dimensional scene associated with that pixel.
[0011] According to a variant of the learning process, an analytical correspondence between a pixel of the first image and a pixel of the second image is determined by the following function: j I f -1----------------- 4- / ^ / Z^ÿsm» t \ y + 2^y COSvJ ^r^_i7l , JJ j- +2^'COSt / j With : • xr an abscissa of the pixel of the second image, • A - the y-coordinate of the pixel in the second image, • xi, an abscissa of the pixel in the first image, • y, an ordinate of the pixel in the first image, • an angle between an optical axis of the first camera and an optical axis of the second camera, • H is a height separating the optical axis of the first camera and the optical axis of the second camera, • f is a focal length associated with the first camera, • fr is a focal length associated with the second camera, and * a depth associated with the pixel of the first image.
[0012] According to another variant of the learning method, the coordinates of the first and second target pixels are determined analytically by the following functions when the first camera is to the left and below the second camera, the lateral edge of the second image corresponding to the left edge of the second image and the horizontal edge of the first image corresponding to the upper edge of the first image: •^410« — With : • x4io is the x-coordinate of the first target pixel, • -^42(½) is the y-coordinate of the second target pixel, • Aii is the x-coordinate of the horizontal edge of the first image, • an x-coordinate of an optical center of the first image, • cly is a y-coordinate of an optical center of the first image, • an x-coordinate of an optical center of the second image, • cy is a y-coordinate of an optical center of the second image, • xi is the x-coordinate of a pixel of the first image, • H is the height separating the optical axis of the first camera and the optical axis of the second camera. And , JA f, y4^0#—1 V f, sin#-y cos#) +ZyyCOS# • ft a focal length associated with the first camera, • fr a focal length associated with the second camera, and * Zlw 'a determined depth.
[0013] According to another variant of the learning method, the coordinates of the first and second target pixels are determined analytically by the following functions when the first camera is to the left above the second camera, the lateral edge of the second image corresponding to the left edge of the second image and the horizontal edge of the first image corresponding to the bottom edge of the first image: and h] X410«- \ + cx O - ----—r+rf, Zid-FSint^T-COSy 7 / , I '-y "Vi h / I----j----+Zlycosy| With : • x4ioa is the x-coordinate of the first target pixel, • ^411 the x-coordinate of the horizontal edge of the first image, • d- an abscissa of an optical center of the first image, • A- an ordinate of an optical center of the first image, • cx is an abscissa of an optical center of the second image, • cy, an ordinate of an optical center of the second image, • x < the x-coordinate of a pixel in the first image, • H is the height separating the optical axis of the first camera and the optical axis of the second camera. • fl a focal length associated with the first camera, • fr a focal length associated with the second camera, and * Zlw 'a determined depth.
[0014] According to yet another variant of the learning process, the determined depth corresponds to an average depth of a depth prediction range.
[0015] According to yet another variant of the learning process, the first and second errors are photometric errors determined by the following function: L*(p) = Ep[ (!-«)• \Hp)-HP) I +« • ( ï-lSSIM(Hp). ï(p) ) ) ] With: • L*(p) the first reconstruction error denoted L](p), respectively the second reconstruction error denoted L2(p), p being a pixel defined by its coordinates in an image, • l[p] a pixel value in the first resized image, respectively second resized image, * l(p) a pixel value in the fourth image, or third image respectively, • SSIM a function that takes into account a local structure, and • has a weighting factor depending in particular on the type of environment.
[0016] According to yet another variant of the learning process, the loss error is determined by the following function: L' = With : • the loss error, • L^p) the first reconstruction error for a pixelp of the first resized image, and • L2(p) the second reconstruction error for a pixelp of the second resized image corresponding to the pixelp of the first resized image.
[0017] According to yet another variant of the learning process, the third and fourth images are generated using the following function: With : • Ps a pixel of a generated image corresponding to the third image, respectively the fourth image, • A function to convert from homogeneous coordinates to pixel coordinates by removing one dimension from a vector, • 1 an intrinsic matrix associated with the second camera, respectively with the first camera, • K' is an intrinsic matrix associated with the first camera, respectively with the second camera, • T is an extrinsic matrix of the stereoscopic vision system, • 0 a projection function in the three-dimensional scene of a pixel as a function of its depth, and * L^p ) is a predicted depth for a pixel pt of the first image, respectively of the second image, determined by the convolutional neural network.
[0018] According to a second aspect, the present invention relates to a device for learning a depth prediction model associated with a vision system embedded in a vehicle, the device comprising a memory associated with at least one processor configured for the implementation of the steps of the process according to the first aspect of the present invention.
[0019] According to a third aspect, the present invention relates to a vehicle, for example of the automobile type, comprising a device as described above according to the second aspect of the present invention.
[0020] According to a fourth aspect, the present invention relates to a computer program which includes instructions adapted for carrying out the steps of the process according to the first aspect of the present invention, in particular when the computer program is executed by at least one processor.
[0021] Such a computer program may use any programming language and be in the form of source code, object code, or an intermediate form between source code and object code, such as in a partially compiled form, or in any other desirable form.
[0022] According to a fifth aspect, the present invention relates to a computer-readable recording medium on which is recorded a computer program comprising instructions for carrying out the steps of the process according to the first aspect of the present invention.
[0023] On the one hand, the recording medium can be any entity or device capable of storing the program. For example, the medium can include a storage means, such as a ROM, a CD-ROM or a microelectronic circuit-type ROM, or a magnetic recording means or a hard disk drive.
[0024] On the other hand, this recording medium can also be a transmissible medium such as an electrical or optical signal, such a signal being able to be transmitted via an electrical or optical cable, by conventional or radio frequency, by self-directing laser beam, or by other means. The computer program according to the present invention can, in particular, be downloaded from an Internet-type network.
[0025] Alternatively, the recording medium may be an integrated circuit in which the computer program is incorporated, the integrated circuit being adapted to execute or to be used in the execution of the process in question. Brief description of the figures
[0026] Other features and advantages of the present invention will become apparent from the description of the particular and non-limiting embodiments of the present invention below, with reference to the attached Figures 1 to 7, in which:
[0027] [Fig-1] schematically illustrates a vision system equipping a vehicle, according to a a particular and non-limiting example of a realization of the present invention;
[0028] [Fig.2] illustrates a flowchart of the different steps in a process for determining the depth of a pixel in an image using a prediction model depth associated with a vision system embedded in the vehicle of the [Fig.1], according to a particular and non-limiting embodiment of the present invention;
[0029] [Fig.3] illustrates a flowchart of the different stages of a method for learning the depth prediction model used in the method of [Fig.2], according to a particular and non-limiting example of the present invention;
[0030] [Fig.4] schematically illustrates images received and / or generated during the learning process of [Fig.3], according to a particular and non-limiting example of the present invention;
[0031] [Fig.5] schematically illustrates a device configured to determine the depth of a pixel of an image acquired by a vision system on board the vehicle of [Fig.1], according to a particular and non-limiting embodiment of the present invention;
[0032] [Fig. 6] schematically illustrates a pixel search window in an image acquired by the vision system embedded in the vehicle of [Fig. 1], according to a particular and non-limiting embodiment of the present invention; and
[0033] [Fig.7] schematically illustrates areas for searching for corresponding pixels in images acquired by the vision system on board the vehicle of [Fig.1], according to a particular and non-limiting embodiment of the present invention. Description of examples of achievements
[0034] A method and device for determining the depth of a pixel of an image by a neural network associated with a vision system embedded in a vehicle will now be described in what follows with joint reference to Figures 1 to 7. The same elements are identified with the same reference signs throughout the description that follows.
[0035] The terms "first," "second" (or "firsts," "seconds"), etc., are used in this document by arbitrary convention to allow for the identification and distinction of different elements (such as operations, means, etc.) implemented in the embodiments described below. Such elements may be distinct or correspond to a single element, depending on the embodiment.
[0036] According to a particular and non-limiting example of an embodiment of the present invention, a method for determining the depth of a pixel of an image by a depth prediction model implemented by a convolutional neural network associated with a vision system comprising several cameras.
[0037] Indeed, the depth prediction model is learned in a learning phase comprising resizing images acquired by the vision system in order to process only parts of images corresponding to parts of the three-dimensional scene observed by two cameras of the vision system and the searching for matching pixels between images by limiting the search area to sets of target pixels of limited size.
[0038] Depths associated with the pixels of the resized images are then predicted and images are generated from the images acquired by the vision system, from the predicted depths and from intrinsic and extrinsic parameters of the cameras of the vision system, thus allowing a first and second error to be determined by comparing pixels of the resized images and the generated images.
[0039] The depth prediction model is learned by minimizing a loss error determined from the first and second errors.
[0040] Fig. 1 schematically illustrates a vision system equipping a vehicle, according to a particular and non-limiting embodiment of the present invention.
[0041] Such an environment 1 corresponds, for example, to a road environment consisting of a network of roads accessible to the vehicle 10.
[0042] In this example, vehicle 10 corresponds to a vehicle with an internal combustion engine, an electric motor(s), or a hybrid vehicle with an internal combustion engine and one or more electric motors. Vehicle 10 thus corresponds, for example, to a land vehicle such as a car, a truck, a bus, or a motorcycle. Finally, vehicle 10 corresponds to an autonomous or non-autonomous vehicle, that is to say, a vehicle operating according to a predetermined level of autonomy or under the total supervision of the driver.
[0043] The vehicle 10 advantageously comprises at least two on-board cameras, a first camera 11 and a second camera 12, configured to acquire images of a three-dimensional scene unfolding in the environment of the vehicle 10 from distinct observation positions. The first camera 11 and the second camera 12 form a stereoscopic vision system when used together, as illustrated in [Fig. 1]. The first camera 11 forms a monoscopic vision system when used alone, and similarly, the second camera 12 forms another monoscopic vision system when used alone. The present invention, however, extends to any vision system comprising at least two cameras, for example, 2, 3, or 5 cameras.
[0044] The intrinsic parameters of the first camera 11 characterize the transformation which associates, for an image point, hereafter called "point", its three-dimensional coordinates in the reference frame of the first camera 11 with the pixel coordinates in an image acquired by the first camera 11. These parameters do not change if the first camera 11 is moved. The intrinsic parameters of the first camera 11 include in particular a first focal length fl associated with the first camera 11.
[0045] The intrinsic parameters of the second camera 12 characterize, for their part, the transformation which associates, for an image point, its three-dimensional coordinates in the reference frame of the second camera 12 with the pixel coordinates in an image acquired by the second camera 12. These parameters do not change if the second camera 12 is moved. The intrinsic parameters of the second camera 12 include in particular a second focal length f2 associated with the second camera 12.
[0046] Distortions, which are due to imperfections in the optical system such as defects in the shape and positioning of camera lenses, will deflect the light beams and thus induce a positioning error for the projected point relative to an ideal model. It is then possible to complete the camera model by introducing the three distortions that generate the most significant effects, namely radial, decentering, and prismatic distortions, induced by defects in lens curvature, parallelism, and coaxiality of the optical axes. In this example, the cameras are assumed to be perfect, meaning that distortions are not taken into account, and their correction is addressed during image acquisition or calibration.
[0047] These two cameras 11, 12 are arranged so that each acquires an image of a scene from a different viewpoint. The first viewpoint is, for example, located on or in the left-hand rearview mirror of the vehicle 10 or at the top of the windshield of the vehicle 10. The second viewpoint is, for example, located on or in the right-hand rearview mirror of the vehicle 10 or at the top of the windshield of the vehicle 10. If both cameras are located at the top of the windshield of the vehicle, they are then positioned at a certain distance. In this example, the first camera 11 is located at the top of the windshield of the vehicle 10, and the second camera 12 is located in the right-hand rearview mirror of the vehicle 10.
[0048] A first marker is associated with the first camera 11: - the direction of the x-axis is defined as horizontal and normal to the optical axis of the first camera 11. The distance B separating the optical center of the first camera 11 from the projection of the optical center of the second camera 12 onto the horizontal plane passing through the optical center of the first camera 11 is called the reference basis (in English "baseline"); - the direction of the y-axis is defined as vertical and normal to the optical axis of the first camera 11; - The direction of the z-axis is defined as orthogonal to the directions of the x and y axes. The three axes x, y and z thus form an orthonormal coordinate system.
[0049] The extrinsic parameters related to the position of cameras 11, 12 are the following parameters: - three translations in the x, y, and z directions: Tx, Ty, and Tz, constituting the translation vector T; and - three rotations in the x, y and z directions: 0x, 0y and 0z.
[0050] An extrinsic matrix of the vision system then includes the extrinsic parameters previously defined.
[0051] Note that the first and second cameras 11,12 are arranged so that the optical axis Cl of the first camera 11 is in a first plane parallel to a second plane comprising the optical axis C2, these two planes being separated by a distance H along the vertical direction z. Thus, the angle 0 of rotation is defined between the two optical axes and includes only the 0z component.
[0052] The extrinsic parameters are determined, for example, during a calibration phase of the stereoscopic vision system comprising the first camera 11 and the second camera 12.
[0053] A key constraint of stereoscopic vision systems used in automobiles is, for example, the large distance between the two cameras. Indeed, to cover a measurement range of 200 meters, the reference base must be 60 cm for cameras commonly used in this field.
[0054] The two cameras 11, 12 acquire images of a scene located in front of the vehicle 10, the first camera 11 alone covering a first acquisition field 13, the second camera 12 alone covering a second acquisition field 14 and the two cameras 11, 12 both covering a third acquisition field 15. The first and third acquisition fields 13, 15 thus allow a monoscopic view of the scene by the first camera 11, the second and third acquisition fields 14, 15 allow a monoscopic view of the scene by the second camera 12 and the third acquisition field 15 allows a stereoscopic view of the scene by the stereoscopic vision system composed of the two cameras 11, 12.
[0055] An obstacle 18 is placed in the acquisition field of the cameras, for example in the third acquisition field 15. The presence of the obstacle 18 defines an occlusion field for the stereoscopic vision system composed here of the three fields 16, 17 and 19.
[0056] Among these three fields, field 16 is visible from the second camera 12. The part of the scene present in this field 16 is therefore observable using the monoscopic vision system comprising the second camera 12.
[0057] The field 17 is visible from the first camera 11. The part of the scene present in this field 17 is therefore observable using the monoscopic vision system comprising the first camera 11.
[0058] Finally, field 19 is not visible to any of the cameras. The part of the scene present in this field 19 is therefore not observable.
[0059] According to one particular embodiment, the field of view of the second camera 12 covers at least half of the field of view of the first camera 11.
[0060] It is evident that it is possible to use such a stereoscopic vision system to take images of scenes located on the sides or behind the vehicle 10 by equipping it with cameras placed and oriented differently.
[0061] The images acquired by the cameras 11, 12 at a given acquisition time are in the form of data representing pixels characterized by: - coordinates in each image; and - data relating to the colours and brightness of objects in the observed scene in the form of, for example, RGB colourimetric coordinates (from the English "Red Green Blue", in French "Rouge Vert Bleu") or HSL (Tone, Saturation, Luminosity).
[0062] Each pixel of the acquired image represents an object in the three-dimensional scene present in the camera's field of view. Indeed, a pixel of the acquired image is the smallest visible unit and corresponds to a point of light resulting from the emission or reflection of light by a physical object present in the three-dimensional scene. When light strikes this object, photons are emitted or reflected, which are captured by a photosensitive sensor of the camera after passing through its lens. This sensor divides the three-dimensional scene into a grid of pixels. Each pixel records the light intensity at a specific location, thus capturing visual details. The combination of millions of pixels creates an image that faithfully represents the physical object observed by the camera. An image point described above is therefore a point on the surface of an object in the three-dimensional scene.
[0063] The images acquired by cameras 11, 12 represent views of the same scene taken from different viewpoints, the camera positions being distinct. This scene includes, for example: - buildings; - road infrastructure; - other stationary users, for example a parked vehicle; and / or - other mobile users, for example another vehicle, a cyclist or a moving pedestrian.
[0064] According to a particular embodiment, a field of view of the first camera 11 covers at least half of a field of view of the second camera 12, and a field of view of the second camera 12 covers at least half of a field of view of the first camera 11. In other words, more than half of the pixels of an image acquired by the first camera 11 correspond to an object in the three-dimensional scene seen by the second camera 12, and pixels of an image acquired by the second camera 12 also correspond to this object in the scene. three-dimensional. Similarly, more than half of the pixels of an image acquired by the second camera 12 correspond to an object of the three-dimensional scene seen by the first camera 11, pixels of an image acquired by the first camera 11 also corresponding to this object of the three-dimensional scene.
[0065] The images acquired by the first camera 11 and by the second camera 12 are sent to a computer of a device equipping the vehicle 10 or stored in a memory of a device accessible to a computer of a device equipping the vehicle 10.
[0066] A method for determining depth by a vision system on board the vehicle 10 is advantageously implemented by the vehicle 10, i.e. by a processor, a computer or a combination of computers of the on-board system of the vehicle 10, for example by the computer or computers in charge of the vision system of the vehicle 10.
[0067] Figure 2 illustrates a flowchart of the different steps of a method 2 for determining the depth of a pixel in an image by means of a depth prediction model implemented by a convolutional neural network associated with a vision system embedded in a vehicle, for example in vehicle 10 of Figure 1, according to a particular and non-limiting embodiment of the present invention. Method 2 is implemented, for example, by a device of the vision system embedded in vehicle 10 or by device 5 of Figure 5.
[0068] In a step 21, representative data of an image acquired by the first camera 11 and of an image acquired by the second camera 12 are received.
[0069] In a step 22, depths associated with a set of pixels of one of the received images are predicted by the depth prediction model from the two received images.
[0070] Each determined depth then corresponds to a distance separating the vehicle 10 or a part of the vehicle 10 from an object in the three-dimensional scene to which a pixel is associated, the determination of a depth of a pixel then corresponding to a measurement of a distance separating an object from the vehicle carrying the vision system.
[0071] If the ADAS uses these depths or distances as input data to determine the distance between a part of the vehicle 10, for example the front bumper, and another road user, the ADAS is then able to determine this distance precisely. For example, if the ADAS's function is to activate the braking system of the vehicle 10 in the event of a risk of collision with another road user, and the distance between the vehicle 10 and that same road user decreases sharply, then the ADAS is able to detect this approach. suddenly and to act on the vehicle's braking system 10 to avoid a possible accident.
[0072] Figure 3 illustrates a flowchart of the different stages of a method for learning the depth prediction model used in a method for determining the depth of a pixel of an image, for example in method 2 of Figure 2, according to a particular and non-limiting embodiment of the present invention.
[0073] The learning process 3 is for example implemented by the device on board the vehicle 10 implementing the process of determining a depth by vision system on board a vehicle or by the device 5 of the [Fig.5].
[0074] In a step 31, representative data of a first image 41 and a second image 42 are received, the first image 41 being acquired by the first camera 11 at an acquisition time instant and the second image 42 being acquired by the second camera 12 at the same acquisition time instant.
[0075] According to a particular embodiment, the first image and second image are of the same definition, that is to say they have the same number of pixels, have the same number of pixels according to their height and the same number of pixels according to their width.
[0076] According to another particular embodiment, the first and second images are not of the same resolution. An additional step then consists of resizing or cropping them to obtain a first and second image of the same resolution.
[0077] In a step 32 and as illustrated in [Fig.6], a width LF and a height HF of a pixel search window F are determined as a function of image coordinates in the second image 42 of a pixel 41 Or of the first image 41 determined from an analytical function and depths between a minimum depth and a maximum depth.
[0078] According to a particular embodiment, the first camera 11 is positioned to the left and below the second camera 12. The pixel 410r of the first image 41 is then located at the bottom right of the first image 41, with coordinates (L4i ;H4[) if the origin of the images is their top left corner.
[0079] According to another example, if the second camera 12 is located below and to the right of the first camera 11, then the pixel 41 of the first image 41 is the one located at the top right of the first image 41, with coordinates (L4i ;0) if the origin of the images is their top left corner.
[0080] According to a particular embodiment, an analytical correspondence between a pixel of the first image 41 and a pixel of the second image 42 is determined by the following function: [Math.l] / i I* y \ i
[0081] With: • xr an abscissa of the pixel of the second image 42, • an ordinate of the pixel of the second image 42, • xi an abscissa of the pixel of the first image 41, • y{ an ordinate of the pixel of the first image 41, • 6 an angle between an optical axis of the first camera 11 and an optical axis of the second camera 12, • H is a height separating the optical axis of the first camera 11 and the optical axis of the second camera 12, • ,fl a focal length associated with the first camera 11, • fr a focal length associated with the second camera 12, and * Zlw a depth associated with the pixel of the first image 41.
[0082] The objective of this step 32 is to determine which are the candidate pixels of the The second image 42 corresponds to pixel 410r of the first image 41, based on possible depths, that is, the depths between a minimum and a maximum depth. The candidate pixels then form a set of candidate pixels I in the second image 42. The minimum and maximum depths are, for example, equal to the limits of the measurement range of the vision system; for example, the minimum depth is two meters (2m), while the maximum depth is one hundred meters (100m). Note that the minimum depth is of great importance; indeed, a disparity or optical flux determined for a pixel is maximal when a predicted depth for that pixel is minimal. A minimum depth of 0 would therefore result in target sets of pixels of infinite size and would be of no use because it would not reduce the search time for a pixel corresponding to that pixel.
[0083] The width LF of the pixel search window F is, for example, equal to the maximum value of the absolute value of the difference between an abscissa in the second image 42 of a pixel in the pixel set I and the abscissa in the first image 41 of pixel 410r. Similarly, the height HF of the pixel search window F is, according to this same example, equal to the difference between an ordinate in the second image 42 of a pixel of the pixel set I and the ordinate in the first image 41 of pixel 410r.
[0084] This gives us:
[0085] [Math 2] • Lf = max(lx1-x4i0rl), and • Hf = max(lyry410rl),
[0086] With: • Lp is the width of the pixel search window F, • Hp the height of the pixel search window F, • X! the abscissa of a pixel of the set of candidate pixels I, • y! the ordinate of a pixel from the set of candidate pixels I, • x4ior the abscissa in the first image 41 of pixel 410r, and • y4i0r the ordinate in the first image 41 of pixel 410r.
[0087] In a step 33 and as illustrated in [Fig.7], a first target set of pixels 422 in the second image 42 is determined for each first pixel 412 of a set of pixels of the first image 41 as a function of coordinates of the first pixel 412 and of the search window.
[0088] Similarly, a second target set of pixels 413 in the first image 41 is determined for each second pixel 423 of a pixel set in the second image 42 as a function of the coordinates of the second pixel 423 and the search window.
[0089] Each first target set of pixels 422 then corresponds to a rectangular area whose dimensions are equal to those of the pixel search window F previously defined and whose position in the second image 42 is defined by the coordinates in the first image 41 of a first pixel 412.
[0090] Similarly, each second target set of pixels 413 then corresponds to a rectangular area whose dimensions are equal to those of the pixel search window F previously defined and whose position in the first image 41 is defined by the coordinates in the second image 42 of a second pixel 423.
[0091] These target sets allow for faster searching of pixels corresponding to a first or second pixel by restricting the search area, i.e. by restricting the number of pixels that can correspond to the first or second pixel 412, 423. The search for a corresponding pixel in each of the first and second images 41, 42 is thus faster.
[0092] In a step 34, an abscissa in the first image 41 of a first target pixel 410a corresponding to a first reference pixel 420a belonging to a lateral edge 421v of the second image 42 is determined analytically, the first target pixel 410a being included in a target set of pixels associated with the first reference pixel 420a. Similarly, an ordinate in the second image 42 of a second target pixel 420b corresponding to a second reference pixel 410b belonging to a horizontal edge 41Ih of the first image 41 is also determined analytically, the second target pixel 420b being included in a target set of pixels associated with the second reference pixel 410b.These analytical determinations are a function of a determined depth and geometric parameters, the geometric parameters including intrinsic parameters of the first and second cameras 11, 12 and extrinsic parameters of the vision system, these extrinsic parameters being representative of the relative positions of the first and second cameras 11, 12.
[0093] The particular embodiment illustrated in [Fig. 4] corresponds to a specific relative position of the first and second cameras 11, 12, the first camera 11 being positioned to the left and below the second camera 12. The lateral edge 421v of the second image 42 then corresponds to the left edge of the second image 42 and the horizontal edge 41Ih of the first image 41 then corresponds to the upper edge, or top edge, of the first image 4L. Coordinates of the first and second target pixels 410a, 420b are determined analytically by the following functions when the frames associated with the images are defined so as to place the origin of each frame in the upper left of each image: [Math.3] ZH^sin^ / ^+Z^c^cos# f, +^l X4l0“= +4 and = +cî'
[0094] With: • x4iou the abscissa of the first target pixel 410a, • 420b the ordinate of the second target pixel 420b, • Zm the abscissa of the horizontal edge 41 Ih of the first image 41, • 4- an abscissa of an optical center of the first image 41, • Cy an ordinate of an optical center of the first image 41, • an abscissa of an optical center of the second image 42, • an ordinate of an optical center of the second image 42, • x> the abscissa of a pixel of the first image 41, • H is the height separating the optical axis of the first camera 11 and the optical axis of the second camera 12, • f{ a focal length associated with the first camera 11, • f a focal length associated with the second camera 12, and * Z^y 'a determined depth.
[0095] According to another particular embodiment, when the first camera 11 is to the left above the second camera 12, the lateral edge 421v of the second image 42 corresponds to the left edge of the second image, and the horizontal edge 41Ih of the first image 41 corresponds to the bottom edge of the first image. The coordinates of the first and second target pixels 410a, 420b are determined analytically by the following functions when the frames associated with the images are defined so as to place the origin of each frame in the upper left corner of each image: [Math.4] Z^yf^inff+f^+Z^^cosf) / •^4io« ~ 7t . f, Â + cx and 0= / \ + cy Z^ sin^ cos0] r-“F—+-2 / n-cos0j
[0096] With: • x4Uta the x-coordinate of the first target pixel 410a, • J411 the abscissa of the horizontal edge 41 Ih of the first image 41, • 4 an abscissa of an optical center of the first image 41, • c!y an ordinate of an optical center of the first image 41, • an abscissa of an optical center of the second image 42, • c'y an ordinate of an optical center of the second image 42, • x> the abscissa of a pixel of the first image 41, • H is the height separating the optical axis of the first camera 11 and the optical axis of the second camera 12, • a focal length associated with the first camera 11, • fr a focal length associated with the second camera 12, and * Z,y 'a determined depth.
[0097] According to a particular embodiment, the determined depth corresponds to an average depth within a depth prediction range. For example, if the depth of the observed three-dimensional scene objects is between 2 and 100 m, then the determined depth has a value of 51 m.
[0098] According to other particular embodiments, the determined depth corresponds to a median value of predicted depths for different three-dimensional scenes observed or to an average value of predicted depths for the reference pixel during subsequent observations.
[0099] In a step 35, as illustrated in [Fig. 4], a first resized image 41' is generated by cropping the first image 41 and a second resized image 42' is generated by cropping the second image 42, the cropping being a function of the abscissa of the first target pixel 410a and the ordinate of the second target pixel 420b and the first resized image 41' and the second resized image 42' and so as to obtain first and second resized images 41' and 42' of the same definition.
[0100] The first and second resized images 41', 42' thus have a width determined from the abscissa x4i0a of the first target pixel 410a. Therefore, the width L4n of the first resized image 41' is equal to the width L42i of the second resized image 42' and is equal to the difference between the width L4i of the first image 41, the latter being equal to the width L42 of the second image 42, and the abscissa x4i0a of the first target pixel 410a. We thus obtain: L4i = L42 and L411 = L421 = L4rx4i0a.
[0101] Similarly, the first and second resized images 41', 42' have a height determined from the y420b coordinate of the second target pixel 420b. Thus, the height H4n of the first resized image 41' is equal to the height H42 of the second resized image 42' and is equal to the difference between the height H4 of the first image 41, the latter being equal to the height H42 of the second image 42, and the y420b coordinate of the second target pixel 420b. We thus obtain: H4i = H42 and H42i = H4h = H42 - y42ob -
[0102] The first resized image 41' is generated from the first image 41, with pixels of the first resized image 41' corresponding to pixels of the first image 41, of which: • an abscissa is located between the abscissa x4i0a of the first target pixel 410a and an abscissa of the right edge 41 Iv of the first image 41, and • an ordinate is located between the ordinate y4n equal to the height H4n and an ordinate of the upper edge 41 Ih of the first image 4L
[0103] The pixels of the first image 41 retained to generate the first resized image 41' then correspond to objects of the three-dimensional scene observable by both the first camera 11 and the second camera 12, that is to say, objects of the three-dimensional scene located in the common field of view of the first and second cameras 11, 12, the third acquisition field 15. Indeed, some areas of the first image 41 include pixels corresponding to objects located outside the field of view of the second camera 12 and therefore do not have a corresponding pixel in the second image 42. Thus, the first resized image 41' includes only pixels each having a corresponding pixel in the second image 42 in the absence of occlusion.
[0104] Similarly, the second resized image 42' is generated from the second image 42, with pixels in the second resized image 42' corresponding to pixels in the second image 42, of which: • an abscissa is located between an abscissa x42i equal to the width L42i and the abscissa of the left edge 421v of the second image 42, and • an ordinate is between the ordinate y420b of the second target pixel 420b and the ordinate of the lower edge 42Ih of the second image 42.
[0105] The pixels of the second image 42 retained to generate the second resized image 42' then correspond to objects of the three-dimensional scene observable by both the first camera 11 and the second camera 12, that is to say, objects of the three-dimensional scene located in the common field of view of the first and second cameras 11, 12, the third acquisition field 15. Indeed, some areas of the second image 42 include pixels corresponding to objects located outside the field of view of the first camera 11 and therefore do not have a corresponding pixel in the first image 41. Thus, the second resized image 42' includes only pixels each having a corresponding pixel in the first image 41 in the absence of occlusion.
[0106] The use of a predetermined depth and cropping the first and second images 41, 42 horizontally and vertically is not intended to remove every pixel from one image that does not have a corresponding pixel in the other image, but rather to retain a large proportion of pixels corresponding to objects in the three-dimensional scene observed by the two cameras 11, 12. For example, the pixels in one image that have a corresponding pixel in the other image represent more than 90 or 95% of the pixels in a resized image. Any error induced by the presence of pixels in one image without a corresponding pixel in the other image is then rendered negligible in the subsequent steps.
[0107] In a step 36, a depth associated with each first pixel 412 is predicted by the depth prediction model from the first resized image 41' and the first target set 422. Similarly, a depth associated with each second pixel 423 is predicted by the depth prediction model from the second resized image 42' and the second target set 413.
[0108] Such a depth prediction model, implemented by a convolutional neural network, is known to those skilled in the art and is presented for example in the paper "Unifying Flow, Stereo and Depth Estimation" written by Haofei Xu, Jing Zhang, Jianfei Cai, Hamid Rezatofighi, Fisher Yu, Dacheng Tao and Andreas Geiger, published in July 2023.
[0109] Thus, the depths are predicted for each pixel of a resized image 41', 42', the definition of a depth map is then equal to the definition of each resized image and avoids any approximation or interpolation to determine depths associated with each pixel of a resized image 41', 42', thus saving computation time.
[0110] Thus, by comparing the positions of the pixels in the first and second resized images 41', 42', the depth associated with each first or second pixel is determined, in particular with the first analytical function described above.
[0111] In a step 37, a third image 43 is analytically generated from: • the first resized image 41', • the depths associated with the first pixels 412, and • the geometric parameters.
[0112] Similarly, a fourth image 44 is analytically generated from: • the second resized image 42', • the depths associated with the second pixels 423, and • the geometric parameters.
[0113] Generating an image from a resized image acquired by a camera of the vision system consists of reprojecting a pixel of the resized image into the three-dimensional scene as a point, and then projecting this point into the image plane of another camera of the vision system, so as to obtain an image corresponding to a view of the three-dimensional scene from the viewpoint of the other camera. The image plane of a camera corresponds to a plane defined in the camera's frame of reference, normal to the camera's optical axis and located at the camera's first focal length. Thus, the third image 43 generated from the first resized image 41' is comparable to the second resized image 42'. Similarly, the fourth image 44 generated from the second resized image 42' is comparable to the first resized image 41'.Because objects can obscure other objects in the scene, the generated images are not identical to acquired images. Furthermore, depth prediction, like the models used to generate the images, is not error-free. Therefore, comparing a generated image to an acquired and resized image allows for an evaluation of the accuracy of the different models used.
[0114] According to a particular embodiment, the third and fourth images are generated using the following function:
[0115] [Math.5]
[0116] With: • Ps a pixel of a generated image corresponding to the third image 43, respectively the fourth image 44, • A function to convert from homogeneous coordinates to pixel coordinates by removing one dimension from a vector, • K is an intrinsic matrix associated with the second camera 12, respectively with the first camera 11, • K' is an intrinsic matrix associated with the first camera 11, respectively with the second camera 12, • T is an extrinsic matrix of the stereoscopic vision system, • 0 a projection function in the three-dimensional scene of a pixel as a function of its depth, and * D^p J is a predicted depth for a pixel Pt of the first resized image 41', respectively of the second resized image 42', determined by the convolutional neural network.
[0117] Note that the extrinsic matrix is not the same for the generation of the two images, indeed, a first extrinsic matrix allows us to go from a reference frame associated with the first camera 11 to a reference frame associated with the second camera 12 during the generation of the third image, while a second extrinsic matrix allows us to go from a reference frame associated with the second camera 12 to a reference frame associated with the first camera 11 during the generation of the fourth image.
[0118] In a step 38, a first error associated with each first pixel 412 is determined by comparing pixels of the first resized image 41' and pixels of the fourth image 44.
[0119] Similarly, a second error associated with each second pixel 423 is determined by comparing pixels of the second resized image 42' and the third image 43.
[0120] According to a first particular embodiment, the first and second errors are photometric errors as presented in the document "Digging Into Self-Supervised Monocular Depth Estimation" by Clément Godard, Oisin Mac Aodha, Michael Firman and Gabriel Brostow published in August 2019 and are determined by the following function: [Math.6] L*(.P) = Ep[ ( 1-«) ■ k(p) -î(p) | +«• ( l-iSSIMlj(p), ï(p) ) ) ]
[0121] With: • L*(p) the first reconstruction error denoted L^p}, respectively the second reconstruction error denoted L2(p) being a pixel defined by its coordinates in an image, • l(p) a value of the pixelp in the first resized image 41', respectively second resized image 42', * / (p) a value of pixelP in the fourth image 44, respectively third image 43, • SSIM is a function that takes into account a local structure, and • has a weighting factor depending in particular on the type of road environment in which the vehicle travels 10.
[0122] According to a second particular embodiment, the first and second errors include a determination of a reconstruction error of a generated image determined, furthermore, by the following function:
[0123] [Math.7] W, O ) =E P1 E to .dW(p ; )|VjD sZ ( Pl )| <rw,i)
[0124] With: • Lsmooth( D( p VF, o) a first error for a pixel p of the third image 43, respectively a second error for a pixel p of the fourth image 44 , • is a depth associated with a pixel pt, • VF is a parameter matrix, • ° is the order of a smoothing gradient, • an L1 norm of the second-order depth gradients is calculated with VF = 1, and ° = 2, • x and y are the dimensions of the third image, and the fourth image respectively. • is a hyperparameter dependent on the road environment in which vehicle 10 is traveling, and • ït(P{ ) is a value of the pixel pt in the third image 43, respectively fourth image 44.
[0125] This second function is generally used to deal with discontinuity at the boundary of objects.
[0126] The reconstruction error is thus defined, for example, from the photometric errors and reconstruction errors previously defined.
[0127] In a step 39, the depth prediction model is learned by minimizing a loss error determined from the first and second errors.
[0128] According to a particular embodiment, the loss error is determined by the following function: [Math.8] L' = Ypmin(Lï(p\L2(p)')
[0129] With: • The loss error, • (p) the first reconstruction error for a pixelp of the first image resized to 41 inches, and • L2{p) the second reconstruction error for a pixelp of the second resized image 42' corresponding to the pixel p of the first resized image 41'.
[0130] The training of the depth prediction model consists of adjusting the input parameters of the convolutional neural network in order to minimize the previously calculated loss error.
[0131] Thus, the depth prediction model used for predicting the depth of a pixel in an image acquired by the first camera 11 or the second camera 12 is made more reliable through this learning process. Furthermore, the data used for this learning are obtained from the stereoscopic vision system itself; they therefore correspond to data that are perfectly representative of the use of the vision system mounted on the vehicle 10.
[0132] The steps of resizing the received images and determining the target sets of pixels make it possible in particular to optimize computation times, the learning process is thus faster to implement than other known learning processes, reducing in particular the time to search for corresponding pixels and not requiring depth interpolation steps, the depths being predicted for each pixel of each of the resized images.
[0133] Figure 5 schematically illustrates a device 5 configured for learning a depth prediction model associated with a vision system embedded in a vehicle 10 and / or for predicting a depth associated with a pixel of an image acquired by the vision system embedded in a vehicle 10, according to a particular and non-limiting embodiment of the present invention. The device 5 corresponds, for example, to a device embedded in the first vehicle 10, for example, a computer associated with the stereoscopic vision system.
[0134] Device 5 is, for example, configured to carry out the operations described opposite Figures 1 and 4 and / or steps described opposite Figures 2 and 3. Examples of such a device 5 include, but are not limited to, embedded electronic equipment such as a vehicle's on-board computer, an electronic control unit such as an ECU (Electronic Control Unit), a smartphone, a tablet, or a laptop computer. The device elements Devices 5, individually or in combination, can be integrated into a single integrated circuit, into several integrated circuits, and / or into discrete components. Device 5 can be implemented as electronic circuits, software (or computer) modules, or a combination of electronic circuits and software modules.
[0135] The device 5 comprises one (or more) processor(s) 50 configured to execute instructions for carrying out the steps of the process and / or for executing instructions from the software embedded in the device 5. The processor 50 may include integrated memory, an input / output interface, and various circuits known to those skilled in the art. The device 5 further comprises at least one memory 51, for example, volatile and / or non-volatile memory, and / or includes a memory storage device that may include volatile and / or non-volatile memory, such as EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disk, or optical disk.
[0136] The computer code of the embedded software(s) including the instructions to be loaded and executed by the processor is for example stored on memory 51.
[0137] According to various particular and non-limiting embodiments, the device 5 is coupled in communication with other similar devices or systems (for example other computers) and / or with communication devices, for example a TCU (Telematic Control Unit), for example via a communication bus or through dedicated input / output ports.
[0138] According to a particular and non-limiting embodiment, the device 5 includes a block 52 of interface elements for communicating with external devices. The interface elements of the block 52 include one or more of the following interfaces: - radio frequency RF interface, for example of the Wi-Fi® type (according to IEEE 802.11), for example in the 2.4 or 5 GHz frequency bands, or of the Bluetooth® type (according to IEEE 802.15.1), in the 2.4 GHz frequency band, or of the Sigfox type using UBN (Ultra Narrow Band) radio technology, or LoRa in the 868 MHz frequency band, LTE (Long-Term Evolution), LTE-Advanced; - USB interface (from the English "Universal Serial Bus" or "Universal Serial Bus" in French); HD MI interface (from the English "High Definition Multimedia Interface", or "High Definition Multimedia Interface" in French); - LIN interface (from the English "Local Interconnect Network", or in French "Réseau interconnecté local").
[0139] According to another particular and non-limiting embodiment, the device 5 includes a communication interface 53 which enables communication with other devices (such as other computers in the embedded system) via a communication channel 530. The communication interface 53 corresponds, for example, to a transmitter configured to transmit and receive information and / or data via the communication channel 530. The communication interface 53 corresponds, for example, to a wired network of the CAN (Controller Area Network), CAN FD (Controller Area Network Flexible Data-Rate), FlexRay (standardized by ISO 17458) or Ethernet (standardized by ISO / IEC 802-3) type.
[0140] According to a particular and non-limiting embodiment, the device 5 can provide output signals to one or more external devices, such as a display screen 540, touch or not, one or more speakers 550 and / or other peripherals 560 via the output interfaces 54, 55, 56 respectively. According to a variant, one or more of the external devices is integrated into the device 5.
[0141] Of course, the present invention is not limited to the embodiments described above but extends to a method for measuring the distance between an object and a vehicle equipped with a vision system, which would include secondary steps without falling outside the scope of the present invention. The same would apply to a device configured for implementing such a method.
[0142] The present invention also relates to a vehicle, for example an automobile or more generally an autonomous land-powered vehicle, comprising the device 5 of [Fig.5].
Claims
1. Demands Method for determining the depth of a pixel of an image by a depth prediction model implemented by a convolutional neural network associated with a vision system embedded in a vehicle (10), the vision system comprising a first and a second camera so as to each acquire an image of a three-dimensional scene from a different point of view, said method being implemented by at least one processor, and being characterized in that the depth prediction model is learned in a learning phase comprising the following steps: - reception (31) of data representative of a first image (41) and a second image (42) acquired by respectively the first camera (11) and the second camera (12) at the same time instant of acquisition; - determination (32) of a width and height of a pixel search window (F) as a function of image coordinates in the second image (42) of a pixel of the first image (41) determined from an analytical function and depths between a minimum depth and a maximum depth; - determination (33), for each first pixel (412) of a set of pixels of the first image (41), of a first target set of pixels (422) in the second image (42) as a function of coordinates of the first pixel (412) and of the pixel search window (F), and determination, for each second pixel (423) of a set of pixels of the second image (42), of a second target set of pixels (413) in the first image (41) as a function of coordinates of the second pixel (423) and of the pixel search window (F);- analytical determination (34) of an abscissa in the first image (41) of a first target pixel (410a) corresponding to a first reference pixel (420a) belonging to a lateral edge (42 Iv) of the second image (42), the first target pixel (410a) being included in a target set of pixels associated with the first reference pixel (420a), and of an ordinate in the second image (42) of a second target pixel (420b) corresponding to a second reference pixel (410b) belonging to a horizontal edge (41 Ih) of the;
2. first image (41), the second target pixel (420b) being included in a target set of pixels associated with the second reference pixel (410b), according to a determined depth and geometric parameters, said geometric parameters including intrinsic and extrinsic parameters of the first and second cameras; - generation (35) of a first resized image (41') by cropping the first image (41) and of a second resized image (42') by cropping the second image (42) as a function of the abscissa of the first target pixel (410a) and the ordinate of the second target pixel (420b), the first resized image (41') and the second resized image (42') being of the same definition; - prediction (36) of a depth associated with each first pixel (412) by the depth prediction model from the first resized image (41') and the first target set (422) and prediction of a depth associated with each second pixel (423) by the depth prediction model from the second resized image (42') and the second target set (413); - analytical generation (37) of a third image (43) from the first resized image (41'), the depths associated with the first pixels (412) and the geometric parameters and generation of a fourth image (44) from the second resized image (42'), the depths associated with the second pixels (423) and the geometric parameters; - determination (38) of a first error associated with each first pixel (412) by comparison of pixels of the first resized image (41') and fourth image (44) and of a second error associated with each second pixel (423) by comparison of pixels of the second resized image (42') and third image (43); - learning (39) of the depth prediction model by minimizing a loss error determined from the first and second errors. A method according to claim 1, wherein an analytical correspondence between a pixel of the first image (41) and a pixel of the second image (42) is determined by the following function:
3. With : —:-------------- ---2-J----'-- 1 / IZ^xr^gi \ TL >' (------7;----- • xr an abscissa of the pixel of the second image (42), • ?r an ordinate of the pixel of the second image (42), • xi an abscissa of the pixel of the first image (41), • an ordinate of the pixel of the first image (41), • 0 an angle between an optical axis of the first camera (11) and an optical axis of the second camera (12), • H a height separating the optical axis of the first camera (11) and the optical axis of the second camera (12), • a focal length associated with the first camera (11), • a focal length associated with the second camera (12), and * Z!w a depth associated with the pixel of the first image (41). Method according to claim 2, wherein the coordinates of the first and second target pixels are determined analytically by the following functions when the first camera (11) is to the left and below the second camera (12), the lateral edge (42 Iv) of the second image (42) corresponding to the left edge of the second image and the horizontal edge (411 h) of the first image (41) corresponding to the upper edge of the first image: —~+c< jyCOS# With : • x4iOa the abscissa of the first target pixel (410a), • ^420* the ordinate of the second target pixel (420b), • Ai i the abscissa of the horizontal edge (41 Ih) of the first image (41), • 4 an abscissa of an optical center of the first image (41), • c? an ordinate of an optical center of the first image (41), • cx an abscissa of an optical center of the second image (42), • cy an ordinate of an optical center of the second image (42), • xi the abscissa of a pixel of the first image (41), x410a “ Av / rs >n ^+ / f®+A tA cos# sin^-ycos#) And •^420# • H the height separating the optical axis of the first camera (11) and the optical axis of the second camera (12), • a focal length associated with the first camera (11), • fr a focal length associated with the second camera (12), and * Z!w 'a determined depth.
4. A method according to claim 2, wherein the coordinates of the first and second target pixels are determined analytically by the following functions when the first camera (11) is to the left above the second camera (12), the lateral edge (421v) of the second image (42) corresponding to the left edge of the second image and the horizontal edge (411h) of the first image (41) corresponding to the lower edge of the first image: and ffë^+H\ x4lGa~ f \ + 4 n = — zU / smA-y-cos# J;\ ^y nVi h 1- f +ZwcosHj With: • x4iOa the abscissa of the first target pixel (410a), • the abscissa of the horizontal edge (41 Ih) of the first image (41), • Cx an abscissa of an optical center of the first image (41), • an ordinate of an optical center of the first image (41), • cx an abscissa of an optical center of the second image (42), • cy an ordinate of an optical center of the second image (42), • xi the abscissa of a pixel of the first image (41), • H the height separating the optical axis of the first camera (11) and the optical axis of the second camera (12), • fj a focal length associated with the first camera (11), • fr a focal length associated with the second camera (12), and * Zlw 'a determined depth.;
5. A method according to any one of claims 1 to 4, wherein said determined depth corresponds to an average depth of a depth prediction range.
6. A method according to any one of claims 1 to 5, wherein the first and second errors are photometric errors determined by the following function: L~(p) = Ep[ ( la ) ■ k(p) -hp) 1 + « ■ ( 1-^SSIM{I(p), I(p) ) ) ] With: • L*(p) the first reconstruction error denoted Lx(p), respectively the second reconstruction error denoted L2(p), p being a pixel defined by its coordinates in an image, • l(p) a value of the pixel p in the first resized image (41'), respectively second resized image (42'), * l(p) a value of the pixel p in the fourth image (44), respectively third image (43), • SSIM a function which takes into account a local structure, and • has a weighting factor depending in particular on the type of environment.
7. A method according to any one of claims 1 to 6, wherein the loss error is determined by the following function: L' = p), L2(p) ) With: • L' the loss error, • L] ( p ) the first reconstruction error for a pixelp of the first resized image (21'), and • L2(p) the second reconstruction error for a pixelp of the second resized image (22') corresponding to the pixelp of the first resized image (21').
8. A computer program comprising instructions for carrying out the method according to any one of the preceding claims, when such instructions are executed by a processor.
9. Device (5) for determining depth by means of a vision system mounted in a vehicle (10), said device (5) comprising a memory (51) associated with at least one processor (50) configured for carrying out the steps of the method according to any one of claims 1 to 7.
10. Vehicle (10) comprising the device (5) according to claim 9.
Citation Information
Patent Citations
Method, computer program, device for identifying profiles
EP4325442A1
Apparatuses and methods for machine vision system including creation of a point cloud model and / or three dimensional model
US10755428B2