Method and device for determining depth information of a scene captured by at least two camera
By comparing velocity data in images captured by the camera to identify target points, the problem of large computational load and long time consumption in existing technologies is solved, and more efficient and accurate depth information calculation is achieved.
Patent Information
- Application Number
- CN202380095238.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-21
- Publication Date
- 2025-11-21
AI Technical Summary
Existing technologies are computationally intensive and time-consuming when calculating scene depth information, and the results are not very accurate, especially when the imaging scene is complex.
By comparing the velocity data of the target point in the first image with the velocity data in the second image, the target point in the second image is identified. The velocity data is used to reduce computing resources and time, thereby improving the accuracy of identification.
This reduces the computational resources and time required to identify target points in the second image, improves the accuracy and reliability of identification, and generates more accurate depth information.
Smart Images

Figure CN121002533A_ABST
Abstract
Description
[0001] This invention relates to a method for determining depth information of a scene based on at least two images of the scene captured by at least two cameras. It also relates to an apparatus and system for implementing this method. Furthermore, it relates to the use of the depth information in various applications, such as driving vehicles, identifying objects in a scene, or tracking objects in the scene.
[0002] The field of this invention is to determine depth information associated with one or more points in a scene based on several images of the scene. Background Technology
[0003] The depth of a point in the scene image corresponds to the distance between that point and the viewpoint from which the scene image was acquired. This distance or depth information is important and useful in various applications. For example, it can be used to assist driving or autonomous vehicles based on the location of objects in their environment. It can also be used for object detection or for tracking objects over time.
[0004] There are schemes for calculating this depth information based on at least two images of a scene captured by cameras located at different capture positions. By identifying the position of the same point in the scene on the images, and knowing the camera's position, it is then possible to calculate the depth of that point on each image of the scene using triangulation.
[0005] However, these schemes are computationally intensive and time-consuming, and the accuracy of the results is not always satisfactory, especially when the composition of the imaging scene is complex.
[0006] One objective of this invention is to overcome at least one disadvantage of the prior art.
[0007] Another objective of this invention is to provide a scheme for determining depth information based on an image of a scene, which requires less computational resources and / or takes less time and / or produces more accurate results. Summary of the Invention
[0008] The present invention proposes a method for determining depth information of a scene in a first image of a scene captured by a first camera located at a first capture position to achieve at least one of the above objectives, the method comprising a processing phase for target points of the scene in the first image, the processing phase comprising the following steps:
[0009] - Identify the target point in a second image, the second image being captured by a second camera located at a second capture position different from the first position, and
[0010] - Calculate the depth associated with the target point based on its position in each of the first and second images;
[0011] The feature is that the identification step includes comparing the velocity data of the target point in the first image with the velocity data of at least one point in the second image, wherein the velocity data is previously calculated.
[0012] Therefore, in a manner similar to the prior art, the present invention proposes to use the position of the target point in at least two images of a scene captured by cameras located at different capture positions to determine the depth associated with the target point in the scene.
[0013] However, and in a completely innovative way, unlike existing technologies, this invention proposes using the velocity data of the target point in the first image to identify the target point in the second image. This requires fewer computational resources and takes less time to identify the target point in the second image. Furthermore, the velocity information associated with the target point helps reduce or even avoid errors in identifying the target point in the second image, resulting in more accurate and reliable results.
[0014] In other words, for scene points in an image, and especially for the image pixels corresponding to those scene points, velocity data enriches the data associated with those points / pixels. Thus, more data can be used to identify the scene points on different images of a scene captured from different capture locations.
[0015] In this application, "image" refers to a digital image, particularly a digital raster image. Each point of the matrix corresponds to a pixel of the image and includes one or more numerical values representing different components of the image. For example, in the case of an RGB image, each pixel may include three values, each representing one of the RGB components of the digital image. Of course, each pixel may include other values representing other attributes, such as brightness, hue, saturation, etc.
[0016] In this application, a point in a scene corresponds to a pixel in an image of that scene. When a scene is imaged by two cameras located at different capture positions (preferably spaced as far apart as possible), the same point in the scene can correspond to two pixels at different locations in each of the two images. Preferably, the camera planes can be parallel and the central axes of the cameras can be common to reduce computation, although this is not required.
[0017] In the following text, UV refers to the reference frame associated with the image sensor. This reference frame is a two-dimensional reference frame that defines a plane, and specifically the plane of the image sensor used for image acquisition.
[0018] In the following text, XY is the reference frame associated with the scene. Furthermore, Z is the direction perpendicular to the image plane (i.e., the XY plane) in the scene's reference frame.
[0019] For a point in the image, i.e., for a target point and / or for at least one point in the second image, the velocity data used in the target point identification step includes at least the velocity of the point in the image plane.
[0020] The velocity of a point in the image plane can be represented in a reference frame (denoted as UV) connected to the image sensor. In this case, the velocity of the point in the image plane (denoted as V) uv ) includes the first value (denoted as V) u The two-dimensional velocity vector gives the velocity in the U direction in the image plane, and the second value (denoted as V). v ), which gives the velocity in the V direction in the image plane.
[0021] According to the present invention, the first image may be captured at a first moment (referred to as the first current moment), and the second image may be captured at a second moment (referred to as the second current moment).
[0022] The first current moment and the second current moment are the same or different but very close in time, such that the displacement of the scene point between the two moments is negligible.
[0023] Without loss of generality, and to avoid lengthy wording, the first and second current moments will be considered the same in the following text and referred to as the current moment.
[0024] The processing stage, or the method according to the invention, may include the step of calculating the velocity of the target point in the first image.
[0025] Alternatively, the velocity of the target point in the first image can be calculated in a preliminary step not part of the method according to the invention. For example, the first image can be an enhanced image that already includes the velocity associated with the target point. The velocity data of the target point can be a portion of the data of the pixels corresponding to the target point in the first image.
[0026] For at least one point in the second image, the processing stage, or the method according to the invention, may include the step of calculating the velocity of the point in the second image.
[0027] Alternatively, for at least one point in the second image, the velocity of that point in the second image can be calculated in a preliminary step that is not part of the method according to the invention. For example, the second image can be an enhanced image that already includes the velocity associated with the point. For example, the velocity data of the point can be part of the data of the pixel corresponding to the point in the second image.
[0028] According to the present invention, velocity data associated with points in an image can be calculated in any known manner, and the present invention is not limited to the specific method of calculating said velocity data.
[0029] According to some embodiments, velocity data of points in an image can be calculated based on the image and based on at least one image (referred to as a past image), which was captured by the same camera at a past moment prior to the current moment when the image was captured.
[0030] For example, the velocity of a target point in a first image can be calculated using images captured by a first camera at a past moment, prior to the current moment when the first image was captured. The past image may or may not be an image immediately preceding the first image (e.g., within the image stream captured by the first camera).
[0031] Alternatively or additionally, the velocity of points in the second image can be calculated using images captured by the second camera at past moments, prior to the current moment when the second image was captured. These past images may or may not be images immediately preceding the second image (e.g., within the image stream captured by the first camera).
[0032] According to a non-limiting example, the steps for calculating velocity data of points in an image may include the following:
[0033] -In past images, it may not be the point mentioned;
[0034] - Calculate the distance between the positions of the point in the current image and past images; and
[0035] - Calculate speed data based on the distance and the time interval between the current moment and the past moment.
[0036] For example, the velocity of a target point in the first image can be calculated by dividing the distance between the target point's position in the first image and a past image captured by the first camera by the time interval between the capture times of the first and past images. This produces a two-dimensional velocity vector that gives the velocity of the target point in the plane of the first image.
[0037] Alternatively or additionally, for at least one point in the second image, the velocity of the point can be calculated by dividing the distance between the position of the point in the second image and a past image captured by the second camera by the time interval between the capture times of the second image and the past image. This produces a two-dimensional velocity vector that gives the velocity of the point in the plane of the second image.
[0038] When identifying points in the current image from past images, if velocity data is known, velocity data associated with points in the past images can be considered. In fact, by knowing the velocities of at least some points in the past images, it is possible to determine which of those points might be located in the current image. Thus, the possible location of that point in the past images can be determined. The search for that point in the past images can then begin at that possible location.
[0039] This method can be used to determine the velocity of target points in the first image.
[0040] This method can be used to determine the velocity of points in the second image.
[0041] Typically, this scheme can be used to identify points in a current image captured by a camera from past images captured by the camera at a past moment.
[0042] Furthermore, during the step of calculating the velocity of a point in the current image, identifying the point in the current image from past images can be achieved by comparing the pixel corresponding to that point in the current image with at least one pixel in the past image. This comparison may involve at least one of the following corresponding data for the pixel:
[0043] -Color data,
[0044] -Texture data,
[0045] -Brightness data,
[0046] - Hue data,
[0047] -Saturation data,
[0048] -RGB data,
[0049] -Attributes calculated from at least one of these data,
[0050] -wait.
[0051] Optionally, the context of the pixel can also be considered during the comparison to improve the reliability of the comparison, such as the pixels belonging to similar geometric patterns (e.g., lines with similar orientations) and / or similar color patterns around the pixel.
[0052] According to some embodiments, the target point identification step in the second image may further include comparing acceleration data associated with the target point in the first image with acceleration data of at least one point in the second image, the acceleration data being previously calculated.
[0053] The processing phase, or the method according to the invention, may include a step for calculating the acceleration of the target point in the first image.
[0054] Alternatively, the acceleration of the target point in the first image can be calculated in a preliminary step that is not part of the method according to the invention. For example, the first image can be an enhanced image that already includes the acceleration associated with the target point. The acceleration data of the target point can, for example, be a portion of the data of the pixels corresponding to the target point in the first image.
[0055] The processing phase, or the method according to the invention, may include the step of calculating the acceleration of at least one point in the second image in the second image.
[0056] Alternatively, for at least one point in the second image, the acceleration of said point in the second image can be calculated in a preliminary step that is not part of the method according to the invention. For example, the second image can be an enhanced image that already includes the acceleration associated with said point. For example, the acceleration data of said point can be part of the data of the pixel corresponding to said point in the second image.
[0057] For a point in the current image, acceleration data can be determined by any known method.
[0058] According to some embodiments, the acceleration of a point in the current image can be determined based on the following:
[0059] - The velocity of the point in the current image;
[0060] - The velocity of the point in a past image captured by the same camera at a time past the capture time of the current image (referred to as the current time); and
[0061] - The time elapsed between the current moment and the past moment.
[0062] The velocity of a point in the current image can be calculated as described above. Similarly, the velocity data of a point in a past image can be calculated using a so-called previous image captured by the same camera at a so-called previous moment before the capture of the past image.
[0063] In other words, the acceleration of the target point in the first image can be calculated based on the following:
[0064] - The velocity of the target point in a past image captured by the first camera at a past moment before the current moment of capturing the first image;
[0065] - The velocity of the target point in the first image; and
[0066] - The time between the past moment and the present moment.
[0067] Alternatively or additionally, for at least one point in the second image, the acceleration of said point in the second image can be calculated based on the following:
[0068] - The velocity of the point in a past image captured by the second camera at a past moment before the current moment of capturing the second image;
[0069] - The velocity of the point described in the second image; and
[0070] - The time between the past moment and the present moment.
[0071] As described above, according to the present invention, a target point is identified in a second image by comparing the velocity data and optionally acceleration data of the target point in the first image with the velocity data and possibly acceleration data of one or more points in the scene in the second image.
[0072] According to some embodiments, the step of identifying a target point in a second image may further include comparing a pixel in the first image corresponding to the target point with at least one pixel in the second image, the comparison involving at least one of the following data:
[0073] -Color data,
[0074] -Texture data,
[0075] -Brightness data,
[0076] - Hue data,
[0077] -Saturation data,
[0078] -RGB data,
[0079] -Data derived from these data,
[0080] -wait.
[0081] Alternatively, the environment of the target point can be considered during the comparison to improve the credibility of the comparison, such as the target point belonging to similar geometric patterns (e.g., lines with similar directions) between images, and / or similar color patterns around the target point.
[0082] Therefore, when comparing the target point with points in the second image, the comparison is based on the corresponding data:
[0083] -speed,
[0084] - and optionally acceleration and / or any combination of the data listed above.
[0085] This allows for more accurate identification of target points on the second image.
[0086] According to some embodiments, the search for a target point in the second image can begin at any location in the second image.
[0087] Alternatively, the search for a target point in the second image can begin at the location where that target point is located in the first image.
[0088] According to other embodiments, the processing stage may include selecting a target region in the second image where the target point may be located, the selection being based on:
[0089] - The velocity in the image plane associated with the target point at a time prior to the current time, for example, in a past image captured by the first camera or in a past image captured by the second camera; and
[0090] - The location of the target point in the first image, and / or in an image captured by the first camera at a time prior to the current time, and / or in an image captured by the second camera at a time prior to the current time.
[0091] In this case, the search for target points in the second image can begin from the region of interest.
[0092] In practice, knowing the location of the target point in the first image, the target point can be found in past images captured by the first camera at a past time. The target point may have already been identified in past images captured by the second camera at the same past time, for example, during previous iterations of the method according to the invention. The velocity of the target point may also have been determined in past images captured by the first camera or by the second camera, for example, during previous iterations of the method according to the invention. Each of these data can be used alone or in combination with another of the data to estimate the possible location of the target point in the second image. In the identification step, the search for the target point in the second image may begin at the possible location or within a defined target area surrounding the possible location.
[0093] The invention has been described with reference to the target points of the scene appearing in the first image.
[0094] Of course, the method according to the invention may include a processing phase that iterates sequentially or simultaneously over several target points of the first image. In this way, depth data can be determined for several target points of the scene appearing in the first image.
[0095] Specifically, the processing phase can be performed on each point of the first image. In this case, each point in the first image is a target point.
[0096] Alternatively, the processing phase can be performed on pre-selected target points within the first image. In this case, the method according to the invention may include a phase of selecting target points in the first image.
[0097] Selecting a target point in the first image can be done using any technique / method or relationship.
[0098] According to some embodiments, the stage of selecting a target point in the first image may include the following steps:
[0099] - Extract the spatial gradient at each point of the first image, particularly by taking the derivative of the image, and even more particularly by Sobel filtering;
[0100] - Calculate the gradient norm at each point, especially the Euclidean norm; and
[0101] - Select points whose gradient norm is greater than a predefined threshold as target points.
[0102] Thus, the points in the first image to be processed are those of interest based on the norm of the spatial gradient. These points are typically the boundaries of various objects in the scene and provide a good indication of how the scene evolves. Therefore, the method according to the invention avoids processing all points of the scene in the first image, but only those of interest, thereby optimizing the processing of the first image in terms of time and computing resources.
[0103] Spatial gradients can be computed based on any parameters stored for each pixel. Alternatively, spatial gradients can be computed based on parameters derived from at least one parameter stored for each pixel. For example, spatial gradients can be computed based on luminance, which can correspond to a weighted sum of the R, G, and B components of each pixel in an RGB image.
[0104] According to some embodiments, the method according to the invention may include storing depth data of at least one target point associated with the target point in a first image and / or a second image.
[0105] Depth data can be stored together with pixel data corresponding to the target point in a first image and / or a second image.
[0106] Furthermore, depending on optional advantages, the method according to the invention may include storage:
[0107] - Speed data, and / or
[0108] -Acceleration data,
[0109] The velocity data and / or acceleration data are optionally calculated for the target point during the processing phase implemented for the target point and stored in the first image and / or the second image, for example, stored together with the pixel data corresponding to the target point.
[0110] Of course, the method according to the invention can be implemented to determine the depth information of the scene in a timely manner at several moments based on at least two image streams of the scene captured by at least two cameras located at different positions, each image stream including at least one image of the scene captured at each of the several moments.
[0111] At least at a current moment, the method according to the invention can be implemented to determine the depth of points for which depth data was not determined at past moments. This allows for the consideration of new points in the scene.
[0112] Alternatively or additionally, at least at a current moment, the method according to the invention can be implemented to determine the depth of a target point in a scene, the target point having had its depth data determined at least at a past moment. Thus, at the current moment, the depth of the target point can be known at the current moment and at least at a past moment, making it possible to track changes in the depth of the target point over time.
[0113] Furthermore, advantageously, by knowing the depth of the same target point at different times, it is possible to calculate:
[0114] - The velocity of the target point in a direction perpendicular to the image plane, for example, in a reference frame connected to the scene;
[0115] - and possibly, the acceleration of the target point in that direction.
[0116] The processing phase may also include steps for calculating the velocity of at least one target point, and / or calculating the acceleration of at least one target point.
[0117] The method or processing stage according to the invention may further include storing the velocity and / or the acceleration in the first image and / or the second image, particularly in the data of the pixels corresponding to the point in the image.
[0118] According to another aspect of the invention, an enhanced image of a scene obtained by the method according to the invention is provided and stored on a storage device.
[0119] The enhanced image according to the present invention is a digital image, preferably a raster.
[0120] The storage medium can be any type (portable or non-portable), such as memory cards, USB flash drives, computers, servers, telephones, cameras, etc.
[0121] According to another aspect of the invention, it is proposed to use the enhanced image according to the invention in at least one of the following applications:
[0122] - Detect objects in the scene at a given time.
[0123] - Track objects in the scene over time
[0124] - Driver assistance or autonomous or semi-autonomous driving for vehicles.
[0125] In practice, at least one depth data point associated with a target point can be used to detect objects in the image, since points with the same depth (and optionally the same velocity and / or acceleration) are likely to belong to the same object. Therefore, objects can be distinguished and / or detected in the enhanced image according to the invention based on depth data (and optionally velocity and / or acceleration data).
[0126] Furthermore, depth data associated with a target point at several points in time can be used to track that target point over time, thereby tracking the object to which the target point belongs over time.
[0127] Furthermore, as mentioned above, the depth data of the target point at two different times can be used to calculate at least the velocity of that point in the depth direction, i.e., in the direction perpendicular to the image plane. This velocity data can be used to estimate or predict the possible position of the target point (and thus the object to which the target point belongs) in that depth direction in the near future (especially at subsequent times). This improves object tracking in both time and space. Additionally, it can be used to assist in the driving of vehicles or robots.
[0128] According to another aspect of the invention, a computer program is provided comprising executable instructions that, when executed by a computer, implement all the steps of the method according to the invention.
[0129] Computer programs can be written in any computer language, such as machine language, C, C++, JAVA, Python, etc.
[0130] According to another aspect of the invention, an apparatus is proposed comprising means configured to implement all the steps of the method according to the invention.
[0131] The device according to the invention can be any type of device programmed to implement the method according to the invention, such as a server, computer, tablet computer, calculator, processor, computer chip, or camera.
[0132] For example, the device according to the invention may be equipped with a computer program according to the invention.
[0133] Specifically, the processing device may be integrated into one of the first camera and the second camera, or may be one of the first camera and the second camera, and more generally, one of the cameras used to image the scene and provide a processed image.
[0134] According to another aspect of the invention, a system is proposed, the system comprising:
[0135] - At least two camera modules are used to image the scene from two different capture positions, and
[0136] - Devices, including means configured to implement all the steps of the method according to the invention, or devices according to the invention.
[0137] Each camera module can include at least one lens and at least one image sensor (such as a CCD or CMOS sensor).
[0138] According to another aspect of the invention, a vehicle equipped with at least one imaging system according to the invention is proposed.
[0139] According to some embodiments, the means of transport can be an automatic or non-automatic land vehicle, such as a car.
[0140] According to some embodiments, the vehicle can be an autonomous or non-autonomous aircraft, such as a drone, airplane, or helicopter.
[0141] According to some embodiments, the means of transport can be an automatic or non-automatic maritime vehicle, such as a ship or submarine.
[0142] According to another aspect of the present invention, a method for assisting in driving a vehicle is proposed, comprising at least one iteration of the following steps:
[0143] - Determine depth or distance data associated with at least one target point in the scene imaged by the vehicle (particularly associated with the object to which the at least one target point belongs) by means of the method according to the invention, and
[0144] - Generate at least one instruction regarding driving the vehicle based on the depth data and / or distance data.
[0145] According to another aspect of the invention, a user device equipped with an imaging system according to the invention is proposed.
[0146] User devices can be cameras, smartphones, tablets, virtual reality headsets, augmented reality headsets, etc.
[0147] According to another aspect of the invention, a medical imaging apparatus equipped with an imaging system according to the invention is provided.
[0148] Medical imaging devices can be endoscopes, especially disposable endoscopes.
[0149] Description of the Drawings and Detailed Description of the Embodiments
[0150] Other advantages and features will become apparent when considering the detailed description of the completely non-limiting embodiments and the accompanying drawings, wherein:
[0151] - Figure 1 This is a schematic diagram of a non-limiting exemplary embodiment of the target point selection stage in an image that can be implemented in this invention;
[0152] - Figure 2 This is a schematic diagram of a non-limiting exemplary embodiment of the speed calculation steps that can be implemented in this invention;
[0153] - Figure 3 This is a schematic diagram of a non-limiting exemplary embodiment of depth computing that can be implemented in this invention;
[0154] - Figure 4 This is a schematic representation of a non-limiting exemplary embodiment of the method according to the present invention;
[0155] - Figure 5 This is a schematic representation of a non-limiting exemplary embodiment of the method according to the present invention applied to an image stream;
[0156] - Figure 6 This is a schematic representation of a non-limiting example embodiment of the device according to the present invention;
[0157] - Figure 7 This is a schematic diagram of a non-limiting exemplary embodiment of an imaging system according to the present invention; and
[0158] - Figure 8 This is a schematic representation of a non-limiting exemplary embodiment of a car according to the present invention.
[0159] It should be clearly understood that the embodiments described below are in no way limiting. In particular, variations of the invention may be conceived that include only such selections of features separate from the other disclosed features, provided that the selection of features disclosed herein is sufficient to provide technical benefits or distinguish the invention from the prior art. Such selections include at least one preferred functional feature that has no structural details or only partial structural details, provided that only that portion is sufficient to provide technical benefits or distinguish the invention from the prior art.
[0160] In particular, all the described variations and embodiments can be combined with each other, provided that there are no technical obstacles to such combination.
[0161] In the rest of the accompanying drawings and description, the same reference numerals are used for common features in several of the drawings.
[0162] Figure 1 This is a schematic diagram of a non-limiting exemplary embodiment of the target point selection stage in an image that can be implemented in this invention.
[0163] Figure 1 The selection phase 100 can be used to select a target point in a first image of a scene captured by a first camera positioned at a first capture position, the depth of which is determined by the method according to the invention using a second image of the scene captured by a second camera positioned at a second capture position.
[0164] Figure 1 The purpose of the selection phase 100 shown is to select points of interest in the first image. Points of interest are typically the boundaries of various objects in the scene and provide a good indication of how the scene evolves. Therefore, the goal of selection phase 100 is to avoid processing all points of the scene in the first image in the rest of the method, thereby optimizing the processing of the first image in terms of time and computing resources.
[0165] During selection phase 100, one or more scene points in the image can be selected based on any logic or relationship.
[0166] In the following text, the first image of the scene is referred to as IM1.
[0167] According to some embodiments, the selection phase 100 includes step 102 of extracting the spatial gradient (e.g., brightness-related) at each point of the first image IM1 through Sobel filtering.
[0168] Then, in step 104, for each point of the first image IM1, the norm of each gradient calculated in step 102, in particular the Euclidean norm, is calculated.
[0169] Then, in step 106, each point with a gradient norm greater than a predefined value is selected as the target point.
[0170] The selection threshold can be determined through testing. Alternatively, the selection threshold can be calculated based on the Euclidean norm calculated for all points in the first image. For example, the selection threshold could correspond to the average of the Euclidean norms calculated in step 104.
[0171] Thus, the points in the first image to be processed are those points of interest based on the spatial gradient norm. As mentioned above, these points are typically the boundaries of various objects in the scene and provide a good indication of how the scene evolves.
[0172] Figure 2This is a schematic diagram of a non-limiting exemplary embodiment of the speed calculation step that can be implemented in this invention.
[0173] Figure 2 Step 200 can be used to calculate the velocity associated with the target point in the image plane (denoted as plane UV in a reference frame associated with the image sensor) for a target point appearing in the image (referred to as the current image, denoted as IMC). The current image IMC is captured by the camera at a time referred to as the current moment.
[0174] Velocity calculations use images captured by the same camera at past moments relative to the current moment (referred to as past images, denoted as IMPs). Preferably, the camera is stationary. If the camera moved between the past and current moments, the camera displacement will be taken into account in the velocity calculation.
[0175] Calculation step 200 includes step 202 of identifying a target point in the current image from past images. Identification is performed by comparing the pixel corresponding to the target point in the current image with at least one pixel in a past image. The comparison involves at least one of the following data:
[0176] -Color data,
[0177] -Texture data,
[0178] -Brightness data,
[0179] - Hue data,
[0180] -Saturation data,
[0181] -RGB data,
[0182] - Parameters calculated from at least one of these data,
[0183] -wait.
[0184] The target point can be compared with all points in past images to identify the target point in past images.
[0185] Alternatively, identification step 202 may include an optional step 204 of determining a target region in a past image, wherein a target point may be located within that target region. The target region can be identified based on velocity data associated with points in the past image when such velocity data is known. In practice, by knowing the velocities of at least some points in the past image, it is possible to determine which of these points might be located in the current image as the target point. Thus, the possible location of the target point in the past image can be determined. A target region of, for example, a predetermined size / area can be defined around the possible locations. The search for the target point in the past image can begin at the possible locations or within the target region.
[0186] In step 206, the target point is compared with one or more points in the past images:
[0187] - If optional step 204 is not performed, then start at any position, or
[0188] - If optional step 204 is performed, start from the possible location or target area.
[0189] If the comparison fails to identify the target point, for example because there are no points in past images that are the same as or sufficiently similar to the target point, then step 200 terminates.
[0190] Otherwise, proceed to step 200.
[0191] In step 208, the velocity VC of the target point in the image plane is calculated, for example, using the following formula. uv This velocity is represented in the UV reference frame associated with the image sensor:
[0192] VC uv =(PPC) uv PPP uv ) / (I c -I p )
[0193] in
[0194] -PPC uv The position of the target point in the current image is represented in the UV reference frame associated with the image sensor;
[0195] PPP uv The location of the target point in past images is represented in the UV reference frame associated with the image sensor;
[0196] -I c For the current moment, and
[0197] -I p This refers to a past moment.
[0198] The velocity of the target point in the UV of the image plane (denoted as VP) uv If the past image IMP is known, then step 200 can include calculating the acceleration (denoted as AC) of the target point in the image plane UV for the current image IMC, for example, using the following relationship. uv This acceleration is represented in a reference frame connected to the image sensor:
[0199] AC uv =(VC) uv -VP uv ) / (I c -I p )
[0200] In optional step 212, the velocity VC associated with the target point in the current image can be stored. uv and / or acceleration AC uv For example, it can be stored directly in the current image, or in the data of the pixels corresponding to the target point in the current image IMC.
[0201] Figure 3 This is a schematic diagram of a non-limiting exemplary embodiment of a depth calculation step for points in a scene that appear in two images of the scene, which can be implemented in this invention.
[0202] Figure 3 A configuration is shown in which a scene is imaged by a first camera 302 providing a first image IM1 of the scene and a second camera 304 providing a second image IM2 of the scene. Each of cameras 302 and 304 is shown schematically.
[0203] The corresponding positions of cameras 302 and 304 are known, such that when capturing images IM1 and IM2, the distance between the camera positions (denoted as D) is... cam ) is known.
[0204] Point 306 in the scene imaged by camera 302 is located at position P1 = {u1, v1} on the sensor of camera 302, which is represented in the reference frame UV associated with the image sensor: therefore, point 306 corresponds to the pixel located at position P1 in image IM1. When the scene is imaged by camera 304, the same point 306 in the scene is located at position P2 = {u2, v2} on the sensor of camera 304: therefore, point 306 corresponds to the pixel located at position P2 in image IM2. This is denoted as d = P1 - P2.
[0205] Assume that cameras 302 and 304 have the same focal length f, and approximate it as follows:
[0206] - The sensors of cameras 302 and 304 are located in the same plane, and
[0207] - Compared to the distance D between each of the cameras 302 to 304 in the direction Z perpendicular to the image sensor plane and point 306. z The focal length f is very small;
[0208] The distance D z It can be calculated based on the following relationship:
[0209] D z =D cam .f / d
[0210] Because compared to D z The focal length f is very small, and D z Corresponding to the distance between the plane of each sensor and point 306 in the scene, therefore D z The depth corresponding to the point in each of images IM1 and IM2.
[0211] refer to Figure 3 The described technique is by no means limiting. It can be used to calculate the depth associated with a target point in the scene using a first and a second image of the scene.
[0212] Figure 4 This is a schematic diagram of a non-limiting exemplary embodiment of the method according to the present invention.
[0213] Figure 4 Method 400 can be used to determine depth information of at least one point (referred to as a target point) of a scene in a first image IM1 of the scene captured by a first camera from a second capture position different from the first capture position, using a second image IM2 of the scene captured by a second camera from the first capture position.
[0214] Images IM1 and IM2 can be captured at different but very close moments, such that the displacement of objects in the scene between said different moments is negligible. Preferably, images IM1 and IM2 are captured at the same moment (hereinafter referred to as the current moment).
[0215] Method 400 may optionally include step 402, for calculating a velocity and optionally acceleration associated in the image plane UV with at least one point (particularly several points) of the scene appearing in the second image IM2, the velocity and optionally acceleration being expressed in a reference frame associated with the image sensor. The velocity and possibly acceleration in the image plane UV for at least one point in the second image IM2 can be calculated using any technique. In a non-limiting example embodiment, step 402 may be... Figure 2Step 200. In this case, for at least one point in the second image IM2, the velocity and, possibly, the acceleration are calculated using past images captured by the second camera at past times relative to the current time when the second image was captured.
[0216] Alternatively, the velocity and possibly acceleration associated with at least one point in the second image IM2 in the image plane UV can be previously calculated or measured and provided in association with the second image, for example, stored in the second image, particularly in the data of the pixel corresponding to said point.
[0217] Optionally, but particularly advantageously, method 400 may include a stage 404 for selecting one or more target points in a first image IM1, the depth information of which is determined by the method according to the invention. For example, stage 404 may be... Figure 1 Phase 100.
[0218] Alternatively, depth information can be determined for each point in the first image IM1.
[0219] Method 400 includes a processing phase 410 for calculating depth information of the target points in the first image IM1 for target points of a scene appearing in the first image IM1 using the second image IM2.
[0220] Phase 410 is executed individually for each target point. When several target points are pending processing, phase 410 is executed sequentially or simultaneously for each of the target points.
[0221] Processing stage 410 may include determining the velocity VC of the target point in the image plane UV in the first image IM1. uv Optional step 412. Speed VC uv This can be determined using any known technique. In a non-limiting example embodiment, the speed VC uv By / According to Figure 2 Step 200 is determined. In this case, the velocity of the target point in the UV plane of the image plane is calculated using a past image captured by the first camera at a past time relative to the current time when the first image was captured.
[0222] Alternatively, the velocity VC of the target point in the UV of the image plane uv It can be calculated or measured in advance. In this case, the velocity VC uv It can be provided in association with the first image, for example, stored in the first image, particularly in the data of the pixels corresponding to the target point in the first image.
[0223] Processing stage 410 may include determining the acceleration (denoted as AC) of the target point in the image plane UV in the first image IM1. uv Optional step 414. Acceleration AC uv This can be determined using any known technique. In a non-limiting example embodiment, reference can be made as described above. Figure 2 Step 210: Determine the acceleration AC uv .
[0224] Alternatively, the acceleration AC of the target point in the UV of the image plane uv It can be previously calculated or measured. In this case, the acceleration AC uv It can be provided in association with the first image, for example, stored in the first image, particularly in the data of the pixels corresponding to the target point in the first image.
[0225] In some embodiments, optional steps 412 and 414 can be combined into a single step, such that the same step calculates both the velocity and acceleration of the target point in the first image. Such a step can be equivalent to... Figure 2 Step 200.
[0226] Processing stage 410 then includes step 416 of identifying target points in the second image.
[0227] The identification step 416 may include an optional step 418, which identifies the target region in the second image IM2 where the target point may be located. This identification can be performed based on the following:
[0228] - The velocity associated with the target point at a time prior to the current time, such as in a past image taken by the first camera or in a past image taken by the second camera; and
[0229] - The location of the target point in the first image, and / or in a past image captured by the first camera at a past moment before the current moment, and / or in an image captured by the second camera at a past moment before the current moment.
[0230] In practice, by knowing the location of the target point, it can be found in past images captured by the first camera at a past time. Furthermore, if the target point has already been identified by the second camera in past images captured at the same past time, for example, during previous iterations of the method, the velocity of the target point may also have been determined in past images captured by the first camera or the second camera, for example, during previous iterations of the method according to the invention. At least one of these data can be used to estimate the possible location of the target point in the second image.
[0231] It is possible to define a target area of a predetermined size around the thus determined possible location.
[0232] The identification step 416 includes step 420, in which a target point is searched in the second image IM2. This search is performed by comparing the target point with at least one point in the second image.
[0233] According to the present invention, the comparison includes comparing the velocity VC of the target point in the first image. uv The velocity VC of the point in the second image IM2 uv Compare them.
[0234] Optionally, the comparison may include comparing the acceleration AC of the target point in the first image. uv The acceleration AC of the point in the second image IM2 uv Compare them.
[0235] Optionally, but preferably, the comparison may include comparing pixels in the first image corresponding to the target point with pixels in the second image associated with the point being compared to the target point, the comparison involving at least one of the following data:
[0236] -Color data,
[0237] -Texture data,
[0238] -Brightness data,
[0239] - Hue data,
[0240] -Saturation data,
[0241] -RGB data,
[0242] -Data derived from these data,
[0243] -wait.
[0244] If the search performed in step 420 does not identify any point in the second image IM2 that is the same as or sufficiently close to the target point in the first image IM1, then the search fails to identify the target point in the second image IM2. Then, processing phase 410 for that target point stops.
[0245] When the search performed in step 420 identifies a point in the second image that is the same as or sufficiently close to the target point in the first image, the search can be stopped and the target point can be identified in the second image IM2.
[0246] When optional step 418 is performed, the search for the target point in the second image IM2 at step 420 can start from the possible location or the location within the target area identified in optional step 418.
[0247] When optional step 418 is not performed, the search for the target point in the second image IM2 at step 420 can start at a random location, or at the same location where the target point has been in the first image IM1, etc.
[0248] At the end of identification step 416, and when the target point has been identified in the second image IM2, the location of the target point is known:
[0249] - In the first image IM1, represented in a reference frame associated with the image sensor of the first image IM1, and
[0250] - Represented in the reference frame associated with the image sensor of the second image IM2.
[0251] Processing phase 410 then includes step 424, for calculating the depth associated with the target point in the first image IM1 based on the position of the target point in the first image IM1 and the position of the target point in the second image IM2.
[0252] This depth calculation can be performed using conventional computational techniques, particularly triangulation:
[0253] -The location of the target point in the first image IM1
[0254] -The location of the target point in the second image IM2
[0255] - The distance between the first camera and the second camera when capturing the first image IM1 and the second image IM2, and
[0256] - The focal length of the camera.
[0257] In particular, it is possible to use references Figure 3 The described technical computational depth.
[0258] When calculating the depth associated with a target point, processing phase 410 for that target point can terminate. A new iteration of processing phase 410 can then be performed for another target point in the first image IM1. And so on.
[0259] Method 400 may include an optional step 430 for storing the following items in the first image IM1 for at least one target point:
[0260] -Depth data calculated for the target point in step 424
[0261] - and optionally, velocity data and / or acceleration data calculated in optional step 412 and / or optional step 414.
[0262] This data, or these data, can be stored specifically along with pixel data corresponding to the target points. Therefore, step 430 can provide a first enhanced image (denoted as IM1*) with the data of one or more target points in the image.
[0263] Optional step 430 can also store at least one of the following data for at least one target point in the second image IM2:
[0264] -Depth data calculated for the target point in step 424
[0265] - and optionally, velocity and / or acceleration data calculated for the target point in optional steps 412 and / or 414.
[0266] This data, or these data, can be stored together with the pixel data in the second image corresponding to the target point.
[0267] Option step 430 may also store at least one of the following data in the second image IM2 for at least one point in the second image IM2:
[0268] - Speed data, and / or
[0269] -Acceleration data;
[0270] The velocity and / or acceleration data are calculated for the point in the second image IM2 in optional step 402. This data, or these data, may be stored together with the data of the pixels corresponding to the point in the second image. Therefore, step 430 can provide a second enhanced image (denoted as IM2*) with the data.
[0271] Figure 5 This is a schematic diagram of a non-limiting exemplary embodiment of using an image stream to determine depth information.
[0272] in other words, Figure 5 The method for determining depth information based on two image streams 502 and 504 (specifically, two videos) provided by two separate cameras for the same scene is shown (and in particular) Figure 4 Example of method 400).
[0273] Image stream 502 includes image 5021 captured at time t1, image 5022 captured at time t2 after t1, image 5023 captured at time t3 after t2, and so on, up to image 502 captured at time tm. m .
[0274] Image stream 504 includes image 5041 captured at time t1, image 5042 captured at time t2 after t1, image 5043 captured at time t3 after t2, and so on, up to image 504 captured at time tm. m .
[0275] exist Figure 5 In the example shown, image 5021 was not processed by the method according to the invention. Similarly, image 5041 was also not processed by the method according to the invention.
[0276] Execute the first iteration of method 400 (denoted as 4001):
[0277] - Treat time t2 as the current time and time t1 as a past time;
[0278] - Image 5022 is considered as the first image, and image 5021 is considered as a past image relative to the first image 5022; and
[0279] - Image 5042 is considered as the second image, and image 5041 is considered as a past image relative to the first image 5042.
[0280] The first iteration 4001 of method 400 provides depth data, both associated with target points in the scene appearing in the first image 5022 and the second image 5042. Optionally, the depth data associated with the target points, and possibly the velocity data, are stored in the first enhanced image denoted as 5022* and / or the second enhanced image denoted as 5042*.
[0281] The second iteration 4002 of method 400 is executed:
[0282] - Treat time t3 as the current time and time t2 as a past time;
[0283] - Consider image 5023 as the first image, and consider image 5022 or preferably enhanced image 5022* as a past image relative to the first image 5023; and
[0284] - Image 5043 is regarded as the second image, and image 5042 or preferably enhanced image 5042* is regarded as a past image relative to the second image 5043.
[0285] The second iteration 4002 of method 400 provides depth data, both associated with target points in the scene appearing in the first image 5023 and the second image 5043. Optionally, the depth data associated with the target points, and possibly velocity data and / or acceleration data, may be stored in the first enhanced image denoted as 5023* and / or the second enhanced image denoted as 5043*.
[0286] Method 400 can therefore be repeated until the remaining images from streams 502 and 504 have been processed.
[0287] Optionally, images 5021, 5022* to 502 m *Can be stored as a first enhanced stream 502*. Images 5041, 5042* through 504 m *Can be stored as a second enhanced stream 504*.
[0288] Each enhanced image obtained by the method according to the invention, i.e., image IM1* or IM2*, or images 5022* to 502*, is... m Each of the images in *5042* to 504 m Each of the options * can be used to identify objects in the image using depth data calculated for the image.
[0289] In addition, images 5022* to 502* from stream 502* m *, or images 5042* to 504m* from stream 504*, can be used to track objects in these images using calculated depth data.
[0290] Furthermore, in image streams 502* and 504* respectively, the change in the depth of a target point over time can be used to predict the depth of the target point (and therefore the object to which the target point belongs) at future times. This prediction can be used, for example, to generate commands or setpoints within a vehicle. Such commands or instructions could be driving instructions, emergency braking or obstacle avoidance maneuvers, or trajectory modifications, etc.
[0291] Figure 6 This is a schematic representation of a non-limiting embodiment of the device according to the present invention.
[0292] Device 600 is configured to implement a method according to the invention for determining depth information of at least one target point of a scene appearing in a first image of the scene captured by a first camera using a second image of the scene captured by a second camera, and in particular... Figure 4 Method 400.
[0293] Device 600 may optionally include module 602 for processing the second image IM2 and calculating the velocity and, optionally, acceleration of at least one point of the scene appearing in the second image in the image plane XY. Specifically, module 602 may be configured to implement... Figure 1 Step 402 of method 100.
[0294] The device 600 may optionally include a module 604 for preprocessing a first image IM1 to select target points of interest in the first image IM1. Specifically, module 604 may be configured to implement... Figure 4 Method 400, step 404.
[0295] The device 600 includes a module 610 for processing at least one target point in the first image to determine depth data associated with the target point. The processing module 610 can be configured to implement processing phases of the method according to the invention, and in particular... Figure 4 Method 400, stage 410.
[0296] Module 610 may optionally include module 612 for determining the velocity associated with the target point in the image plane. Specifically, module 612 may be configured to implement... Figure 1 Step 412 of method 400.
[0297] Module 610 may optionally include module 614 for determining the acceleration associated with the target point in the image plane. Specifically, module 614 may be configured to implement... Figure 1 Method 400, step 414.
[0298] Module 610 may optionally include module 618 for determining a target region in the second image IM2. Specifically, module 618 may be configured to implement... Figure 4 Method 400, step 418.
[0299] Module 610 includes module 620 for identifying target points in the second image IM2. Specifically, module 620 can be configured to implement... Figure 4 Method 400, step 420.
[0300] Processing module 610 includes module 622 for calculating the depth of target points in image IM1. Specifically, module 622 can be configured to implement... Figure 4 Method 400, step 424.
[0301] The device 600 may optionally include an enhancement module 630 for adding depth data, and possibly velocity data and / or acceleration data, associated with at least one point of a scene appearing in the first image IM1 and / or the second image IM2. Specifically, module 630 may be configured to implement... Figure 4 Method 400, step 430.
[0302] At least one of modules 602 to 604, 610 to 614, 618 to 622, and 630 may be a separate module independent of the other modules. Alternatively, at least two of these modules may be integrated into the same module. In particular, all modules of device 600 may be integrated into the same module.
[0303] At least one of modules 602 to 604, 610 to 614, 618 to 622 and 630 may be a hardware module, such as a processor, chip, graphics card, calculator, computer, server, etc.
[0304] Alternatively, at least one of modules 602 to 604, 610 to 614, 618 to 622 and 630 may be a software module executed by at least one computing device.
[0305] According to another alternative, at least one of modules 602 to 604, 610 to 614, 618 to 622 and 630 may be a combination of at least one hardware module and at least one software module.
[0306] The configuration of one or each of modules 602 to 604, 610 to 614, 618 to 622 and 630 may be a software configuration (e.g., an application or computer program loaded in the module) and / or a hardware configuration (e.g., a specific hardware architecture).
[0307] Figure 7 This is a schematic diagram of a non-limiting example embodiment of an imaging system according to the present invention.
[0308] System 700 can be configured to implement the method according to the invention, and in particular Figure 4 Method 400.
[0309] System 700 includes a first camera module 702, which includes at least one lens and at least one image sensor (e.g., a CCD or CMOS sensor). The first camera module 702 captures a first image or a first image stream of the scene.
[0310] System 700 includes a second camera module 704, which includes at least one lens and at least one image sensor (e.g., a CCD or CMOS sensor). The second camera module 704 captures a second image or a second image stream of the scene.
[0311] The imaging system 700 also includes a processing unit 706, which may be used to determine depth information of at least one target point of a scene appearing in the images captured by modules 702 to 704. Figure 6 The equipment is 600.
[0312] At least one of the camera modules 702 to 704 and the processing unit 706 may be incorporated into a single unit. Alternatively, the processing unit 706 and at least one camera module 702 to 704 may not be incorporated into a single unit and may be independent of each other. In this case, the processing unit and the at least one camera module communicate directly or indirectly, via wired or wireless means.
[0313] Figure 8 This is a schematic diagram of a non-limiting example embodiment of a vehicle according to the present invention.
[0314] Figure 8 The vehicle 800 shown can be any type of land vehicle, such as a car, truck, etc.
[0315] The means of transport 800 may or may not be a partially or fully autonomous means of transport.
[0316] Figure 8 The front view of vehicle 800 is shown.
[0317] The vehicle 800 is equipped with an imaging system according to the present invention, in particular Figure 7 System 700. Typically, vehicle 800 may be equipped with:
[0318] - At least two camera modules, such as camera modules 702 and 704;
[0319] - At least one processing unit 706.
[0320] Of course, camera modules 702 to 704 and processing unit 706 can be located in any suitable position within the vehicle. In the example shown, camera modules 702 to 704 are located at the front of the vehicle to image the scene in front of the vehicle 800.
[0321] The processing unit 706 can be a standalone unit. Alternatively, the processing unit 706 can be integrated into the calculator or management unit (not shown) of the vehicle 800.
[0322] The vehicle 800 may include a driving assistance unit 802, which receives enhanced images, depth values, or any other attributes determined from depth values from the processing unit 706, and provides driving instructions to the vehicle driver or directly to vehicle components.
[0323] Of course, the present invention is not limited to the examples described above.
Claims
1. A method (400) for determining depth information of a scene in a first image (IM1) of the scene, the first image (IM1) of the scene being captured by a first camera (302) located at a first capture position, the method (400) comprising a processing phase (410) for target points of the scene in the first image (IM1), the processing phase (410) comprising the following steps: - Identify (416) the target point in the second image (IM2), the second image (IM2) being captured by a second camera (304) located at a second capture position different from the first position, and -Based on the position of the target point in each of the first and second images (IM1, IM2), calculate (424) the depth associated with the target point; The feature is that the identification step (416) includes comparing the velocity data of the target point in the first image (IM1) with the velocity data of at least one point in the second image (IM2), the velocity data being previously calculated.
2. The method (400) according to the preceding claim, characterized in that, The velocity data associated with points in the image is calculated based on the image and based on at least one image referred to as a past image, which was captured by the same camera at a past moment prior to the moment the image was captured, referred to as the current moment.
3. The method (400) according to the preceding claim, characterized in that, The steps (200) for calculating the velocity data of points in an image include the following: - Identify the point (202) in the past image; - Calculate the distance between the positions of the point in the current image and the past image; and - Calculate speed data based on the distance and the time interval between the current moment and the past moment.
4. The method (400) according to any one of the preceding claims, characterized in that, The target point identification (416) step in the second image includes comparing acceleration data associated with the target point in the first image (IM1) with acceleration data of at least one point in the second image (IM2), the acceleration data being previously calculated.
5. The method (400) according to any one of the preceding claims, characterized in that, The step of identifying (416) the target point in the second image (IM2) further includes comparing at least one pixel in the first image (IM1) corresponding to the target point with at least one pixel in the second image (IM2), the comparison involving at least one of the following data: -Color data, -Texture data, -Brightness data, - Hue data, -Saturation data, -RGB data, -Data derived from these data, -wait.
6. The method (400) according to any one of the preceding claims, characterized in that, The step of identifying (416) the target point in the second image (IM2) includes selecting a target region in the second image (IM2) where the target point may be located, the selection being based on: - The velocity associated with the target point in the image plane at a time prior to the current time; and - The location of the target point in the first image (IM1), and / or in an image captured by the first camera (302) at a time prior to the current time, and / or in an image captured by the second camera (304) at a time prior to the current time.
7. The method (400) according to any one of the preceding claims, characterized in that, The method includes an iterative processing phase (410) for several target points of the first image (IM1).
8. The method (400) according to any one of the preceding claims, characterized in that, The method includes a phase (100; 404) of selecting at least one target point in the first image (IM1), the selection step (100; 404) including the following steps: - Extract (102) spatial gradients at each point of the first image (IM1), particularly by taking the derivative of the image, and even more particularly by Sobel filtering; - Calculate the gradient norm, especially the Euclidean norm, for each point (104); - Retain points whose gradient norm is greater than a predefined threshold (106) as target points.
9. The method (400) according to any one of the preceding claims, characterized in that, The method is implemented to determine the depth of at least one point of the scene at several moments based on at least two image streams (502, 504) of the scene captured by two cameras located at different positions, each image stream including at least one image of the scene captured at each of the several moments.
10. The method (400) according to any one of the preceding claims, characterized in that, The method further includes, for at least one target point, storing (430) the depth data of the target point associated with the target point in the first image (IM1) and / or the second image (IM2).
11. An enhanced image (IM1*, IM2*) of a scene, stored on a storage device, obtained by the method according to the preceding claims.
12. An enhanced image of the scene according to claim 11 (IM1*, IM2*; 5022* to 502...) m *;5042* to 504 m *) For use in at least one of the following applications: - Detect objects in the scene at a given time. - Track objects in the scene over time - Driver assistance or autonomous or semi-autonomous driving for vehicles.
13. A computer program comprising executable instructions that, when executed by a computer, implement all the steps of the method (400) according to any one of claims 1 to 10.
14. An apparatus (600) comprising means configured to perform all steps of the method (400) according to any one of claims 1 to 10.
15. A system (700), comprising: - At least two camera modules (702, 704) are used to image the scene from two different capture positions, and - Apparatus (706), comprising means configured to perform all steps of the method (400) according to any one of claims 1 to 10.
16. A user device comprising an imaging system (700) according to the preceding claim.
17. A means of transport (800) comprising the imaging system (700) according to claim 15.