Method and system for providing a continuous view of a ground underneath a moving vehicle
The method addresses the discontinuous view issue by using feature matching and homography matrix alignment to provide a seamless, continuous, and undistorted view of the ground underneath the vehicle, enhancing driving efficiency and safety.
Patent Information
- Application Number
- GB2023018181
- Authority / Receiving Office
- GB · GB
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-29
- Publication Date
- 2025-06-11
AI Technical Summary
Conventional methods provide a discontinuous view of the ground underneath a moving vehicle due to inaccurate motion data conversion from vehicle to image coordinate systems, leading to potential errors and distortions.
A computer-implemented method involving feature matching, alignment correction using a homography matrix, and image stitching to create a continuous and seamless view of the ground underneath the vehicle, utilizing deep convolutional neural networks for enhanced accuracy and uniformity.
Enables a precise, continuous, and undistorted view of the ground underneath the vehicle, aiding efficient maneuvering and preventing damage by accurately aligning and harmonizing images.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Various embodiments relate to method and system for providing to the driver, a continuous view of a ground underneath a moving vehicle. BACKGROUND
[0002] While driving on some terrains, drivers feel the need to have a view of the ground underneath their vehicle. This view may assist the drivers to maneuver the vehicle efficiently. The view may also help avoid damages to the vehicle, because the driver can view the ground underneath and can avoid potholes, bumps and stones on the ground which could seriously damage a vehicle, in particular in regions with less road maintenance. Conventionally the methods that are used to render a view of the ground underneath the vehicle, provide a discontinuous view due to inaccurate motion data received from the vehicle. This data may include data about location and orientation of the vehicle. The data from vehicle is calculated in vehicle coordinate system and stitching happens in image coordinate system. The vehicle coordinate system axes are fixed in a reference frame attached to the vehicle. An image coordinate system defines the spatial reference in terms of a primary image. The conversion from vehicle coordinate system to image coordinate system may sometimes create errors.
[0003] In view of the above, there is a need for an improved method for providing a continuous view of a ground underneath a moving vehicle that can address the abovementioned problems. SUMMARY
[0004] In order to satisfy the need described above, there is provided a computer implemented method for continuously providing a view of the ground underneath a moving vehicle. The method may include performing a feature matching between a current top view image, and a previous top view image. The previous top view image may comprise the ground in front of the moving vehicle, which would later be underneath the moving vehicle in the current top view image. The feature matching may help to identify the corresponding features in the two images. For instance, a stone at coordinate (x,y) in an image, corresponds to the stone at coordinate (x’,y’) in the other image.
[0005] The method may further include identifying and correcting an alignment difference between each pixel of the current top view image and each pixel of the previous top view image. The alignment may be identified based on the feature matching. This also enables getting a precise alignment difference of each pixel between the two top view images. The method may also include mapping the previous top view image to the current top view image. This may result in a modified previous top view image which is now aligned with the current top view image. The method further includes extracting a region of interest (Rol) from the modified previous top view image and stitching this extracted region of interest to the current top view image. After stitching, this region of interest maybe rendered as the ground underneath the vehicle in the current top view image. The position of the vehicle in the current top view image may be displayed by an outline of the region of interest.
[0006] In an alternative combination of features which may be combined with alternative combinations described above, the alignment difference between the images is identified by calculating rotation and translation of each pixel of the current top view image using a homography matrix and the feature matching. This enables precisely calculating relations between the images so that precise alignment correction, mapping and stitching can be done. In one combination of features, a variable for the ratio of rotation to translation can be calculated based on vehicle data, such as steering wheel rotation data, acceleration or deceleration data. This may help to find the right amount of rotation and translation for the alignment difference more quickly and to save time or computing power.
[0007] In an alternative combination of features, for performing homography calculation, an appropriate deep convolutional neural network (CNN) may be activated. In one example homographyNet may be used. HomographyNet is a Visual Geometry Group (VGG) style CNN which produces the homography that can relate two images. HomographyNet may be in two types, classification and regression. Regression model may directly estimate the real-valued homography parameters, and classification model may produce a distribution over quantized homographies.
[0008] In an alternative combination of features which may be combined with alternative combinations described above, correcting the alignment difference comprises calculating a mathematical result using the pixel coordinates and the homography matrix. Alignment differences are corrected so that the images of the ground underneath the vehicle do not appear distorted to the driver. The driver will be able to interpret the path correctly.
[0009] In an alternative combination of features which may be combined with alternative combinations described above, extracting the region of interest comprises calculating a dimension of the region of interest, based on a dimension of the vehicle and the scale factor of the image. This enables efficiently identifying the area between the wheels which should be provided to the driver in the view of the ground underneath. The region of interest will be different for different vehicles and will depend on the wheel settings of the vehicle being driven.
[0010] In an alternative combination of features which may be combined with alternative combinations described above, extracting the region of interest is done for each pixel. Initially the region of interest maybe identified as the portion of the image between the wheels of the vehicle. Each pixel of the identified Rol maybe copied from the aligned previous top view image and pasted to the current top view image in the identified portion of the images.
[0011] In an alternative combination of features which may be combined with alternative combinations described above the images are harmonized after stitching the extracted region of interest to the current top view image. This ensures that the colors and brightness and other such features are uniform across all portions of the image and the driver does not feel any difference in the rendered view.
[0012] In an alternative combination of features, for performing image harmonization a deep convolutional neural network (CNN) called Context aware Encoder-decoder can be activated. This model can capture both the context and semantic information of the composite images during harmonization. In this encoder-decoder model architecture the CNN acts as an encoder and a second CNN acts as a decoder. The deep CNN model that consists of an encoder maybe trained to capture the context of the input image and a decoder to reconstruct the harmonized image using the learned representations from the encoder.
[0013] In an alternative combination of features which may be combined with alternative combinations described above, the harmonized current top view image maybe transferred to the storage device in the vehicle. This stored image may be fetched to be used as previous top view image, when a next current top view image gets generated.
[0014] The above-described advantageous aspects of a computer implemented method also hold for all aspects of a below-described system for providing a continuous view of a path under a vehicle. All below-described advantageous aspects of a system for providing a continuous view of a path under a vehicle also hold for all aspects of an above-described computer implemented method.
[0015] The invention also relates to a system for continuously providing a view of a ground underneath a vehicle, comprising a memory and at least one processor coupled to the memory and configured to carry out the method as described in the above paragraphs.
[0016] In an alternative combination of features which may be combined with alternative combinations described above the system has multiple cameras mounted on the vehicle to generate top view images.
[0017] In an alternative combination of features which may be combined with alternative combinations described above, there is provided a non-transitory computer-readable storage medium comprising instructions, which when executed by a processor, performs the process as explained in the above paragraphs. BRIEF DESCRIPTION OF DRAWINGS
[0018] In the drawings, like reference characters generally refer to the same parts throughout the different views. The drawings are not necessarily to scale, emphasis instead generally being placed upon illustrating the principles of the invention. In the following description, various embodiments are described with reference to the following drawings, in which: Fig 1 relates to an embodiment of a general purpose computing machine for implementing the present disclosure; Fig 2 relates to a flowchart of an embodiment of the process as disclosed in the present disclosure; Fig 3 relates to an example of the view underneath the vehicle rendered to the driver. DESCRIPTION
[0019] The combination of features described below in context of the devices are analogously valid for the respective methods, and vice versa. Furthermore, it will be understood that the combinations of features described below may be considered together, for example, a part of one combination may be joined with a part of another combination.
[0020] It will be understood that any property described herein for a specific device may also hold for any device described herein. It will be understood that any property described herein for a specific method may also hold for any method described herein. Furthermore, it will be understood that for any device or method described herein, not necessarily all the components or steps described must be enclosed in the device or method, but only some (but not all) components or steps may be enclosed.
[0021] In an alternative combination of features, which may be combined with alternative combinations described above, the present disclosure relates to providing a continuous view of a ground underneath a moving vehicle. This may help the driver move the vehicle more efficiently and avoid damage to the vehicle.
[0022] Throughout this document the word ‘stitching’ may have the meaning of a process of merging images or portions of images with narrow and overlapping fields of view, to create a larger image with a bigger field of view. Image stitching may be done using specific tools configured to sense areas of overlap and combine them. The overlaps need to be as exact as possible and the exposures should be identical to provide a final image with no visible seams.
[0023] Throughout this document, the word ‘homography’ may have the meaning of identifying the transformation between two images having same view, using matrix calculations. Homography maybe estimated by identifying the keypoints in both the images and then estimating the transformation between the two views. The matrix used for the calculations maybe referred to as a homography matrix.
[0024] FIG. 1 illustrates a generalized example of a suitable computing environment 100 which maybe used during implementation of the present disclosure. The computing environment 100 is not intended to suggest any limitation as to scope of use or functionality, as the technologies may be implemented in diverse general-purpose or special-purpose computing environments.
[0025] With reference to FIG. 1, the computing environment 100 includes at least one processing unit 110 coupled to memory 120. In FIG. 1, this basic configuration 130 is included within a dashed line. The processing unit 110 executes computer-executable instructions. In a multi-processing system, multiple processing units execute computer-executable instructions to increase processing power. The memory 120 may be non-transitory memory, such as volatile memory (e g., registers, cache, RAM), non-volatile memory (e g., ROM, EEPROM, flash memory, etc.), or some combination of the two. The memory 120 can store software 180 implementing any of the technologies described herein.
[0026] A computing environment may have additional features. For example, the computing environment 100 includes storage 140, one or more input devices 150, one or more output devices 160, and one or more communication connections 170. An interconnection mechanism (not shown) such as a bus, controller, or network interconnects the components of the computing environment 100. Typically, operating system software (not shown) provides an operating environment for other software executing in the computing environment 100, and coordinates activities of the components of the computing environment 100.
[0027] The storage 140 may be removable or non-removable, and includes magnetic disks, magnetic tapes or any other non-transitory computer-readable media which can be used to store information and which can be accessed within the computing environment 100. The storage 140 can store software 180 containing instructions for any of the technologies described herein.
[0028] The input device(s) 150 may be a touch input device such as a keyboard, touchscreen, a voice input device, a scanning device, or another device that provides input to the computing environment 100. The output device(s) 160 may be a display, speaker, or another device that provides output from the computing environment 100. Some input / output devices, such as a touchscreen, may include both input and output functionality.
[0029] The communication connection(s) 170 enable communication over a communication mechanism to another computing entity. The communication mechanism conveys information such as computer-executable instructions, audio / video or other information, or other data. By way of example, and not limitation, communication mechanisms include wired or wireless techniques implemented with an electrical, optical, RF, infrared, acoustic, or other carrier.
[0030] The techniques herein can be described in the general context of computer-executable instructions, such as those included in program modules, being executed in a computing environment on a target real or virtual processor. Generally, program modules include routines, programs, libraries, objects, classes, components, data structures, etc., that perform particular tasks or implement particular abstract data types. The functionality of the program modules may be combined or split between program modules as desired in various combinations. Computer-executable instructions for program modules may be executed within a local or distributed computing environment.
[0031] Any of the storing actions described herein can be implemented by storing in one or more computer-readable media (e.g., non-transitory computer-readable storage media or other tangible media). Any of the things described as stored can be stored in one or more computer-readable media (e.g., computer-readable storage media or other tangible media).
[0032] Any of the methods described herein can be implemented by computer-executable instructions in (e.g., encoded on) one or more computer-readable media (e.g., non-transitory computer-readable storage media or other tangible media). Such instructions can cause a computer to perform the method. The technologies described herein can be implemented in a variety of programming languages.
[0033] Any of the methods described herein can be implemented by computer-executable instructions stored in one or more non-transitory computer-readable storage devices (e.g., memory, CD-ROM, CD-RW, DVD, or the like). Such instructions can cause a computer to perform the method.
[0034] In the present combination of features which may be combined with alternative combinations, the computing environment (100) may be implemented within a vehicle, which may be appropriately configured with memory, processors, input - device and other components as explained above including software instructions. In another combination of features, the computing environment may be implemented on a system on chip (SoC) or other embedded systems.
[0035] In an alternative combination of features, which may be combined with alternative combinations described above, the process of the present patent application will be explained along with the description of FIG 2.
[0036] In one combination of features, which may be combined with alternative combinations described above, a current top view image and a previous top view image of the vehicle may be used for this process. A feature matching may be done between the current top view image and the previous top view image (200). Performing a feature matching may provide a relation between the two images. The process of feature matching may include detecting patterns in the features of the image and matching the detected patterns in the features. The accuracy of the feature matching may depend on the quality and complexity of the images.
[0037] In one combination of features which may be combined with alternative combinations described above, a homography matrix calculation may be performed for the images (201) using the feature matching (200). This enables identifying the alignment difference for each pixel in the previous and the current top view images. In one combination of features, the alignment difference may be calculated by identifying the rotation and translation of each pixel of the current top view image, with respect to the previous top view image. This may also help calculating the precise displacement of each pixel in any direction. For images it may be a matrix specifying the transformation between two views of similar images. The transformation may be estimated by identifying the keypoints in both the images and then estimating the transformation between the two images based on the keypoints.
[0038] A homography matrix may also be a transformation matrix in a coordinates space that is mapped between two planar projections of an image. These transformations may include rotation, translation, scaling, skewing operations or a combination of such transformation of pixels. A homography matrix may also be understood as providing different perspective in the image. Further details of implementation of homography matrix can be found at https: / / medium.com / swlh / image-processing-with-python-image-warping-using-homography-matrix-22096734K)9a
[0039] In one combination of features which may be combined with alternative combinations described above, homography calculation (201) may further include a mapping between points or pixels on one surface or plane, to points on another.
[0040] An example of homography matrix is provided below. A homography matrix maybe a 3X3 matrix.
[0041] In one combination of features which may be combined with alternative combinations described above, the alignment difference is corrected using the homography matrix and the pixel coordinates (202). In one combination of features, a mathematical result may be calculated between pixel coordinates of the previous top view image and the homography matrix. The mathematical result may include calculating a product between pixel coordinates of the previous top view image frame and the homography matrix. This may provide the alignment correction between the two images, for each pixel. Correcting the alignment difference may enable the images to be aligned and seamless. The driver in the vehicle will not get a distorted view of the ground underneath. In one combination of features, if a point in one image (x’,y’) is multiplied by the homography matrix, it may give the corresponding point (x,y) from an the other image. Mathematically, homographic calculation may be represented as such: where (x,y) may represent pixel coordinates in one image, (x5, y’) may represent pixel coordinates in another image and H may be the homography matrix represented as this 3x3 matrix:
[0042] In one combination of features which may be combined with alternative combinations described above, the previous top view image is mapped to the current top view image based on the calculated identified alignment difference. This provides a seamless and perfectly aligned top view images.
[0043] In one combination of features which may be combined with alternative combinations described above, a Region Of Interest (Rol) may be identified in the aligned top view image (203). In one combination of features the region of interest is determined as the area underneath the vehicle, that may be required by the driver. This area may be separate for each type of vehicle. The area between the wheels of the vehicle may provide the view of the ground underneath the vehicle to the driver. This may be based on the dimension of the vehicle. In another combination of features, the Rol may also be based on the scale factor. The scale factor may be calculated as the ratio between the real-world distance represented in the top view image, and the total number of pixels.
[0044] In one combination of features which may be combined with alternative combinations described above, the extracted Rol from the aligned top view mapped image is stitched to the current top view image (204). Therefore as the driver watches the current top view image, he may also see the Rol on this image. This Rol may include the ground which is underneath the vehicle.
[0045] In one combination of features which may be combined with alternative combinations described above, image harmonization may be done for Rol with respect to the rest of the current top view image (205). Image harmonization may enable uniform image features such as color intensity, and brightness across the top view image. As a result, the view of the ground underneath the vehicle may show uniformly with the rest of the image, and not as a separate patch having different image features. This provides the driver with a uniform image stream, with no variation in colors and other image features. In one combination of features, the processes that can be used for image harmonization may include color correction, tone mapping, and style transfer. These techniques can be applied using image processing software or libraries. The position of the vehicle in the current top view image may be displayed by an outline of the region of interest. In another combination of features, machine learning techniques may be used for image harmonization. For example, a neural network can be trained on a dataset of images to learn how to adjust the appearance of an image so that it is more consistent with a reference image.
[0046] In one combination of features which may be combined with alternative combinations described above, the current top view image which also has the Rol harmonized with the rest of image may be transferred to the storage device. This image may then be used as a previous top view image in the next instant when a next top view image is generated.
[0047] In one combination of features which may be combined with alternative combinations described above, the process as defined above is performed till a driving control unit of the vehicle is turned off by the driver.
[0048] In one combination of features which may be combined with alternative combinations described above, when the vehicle starts, a current top view image is generated which may not have any view of the ground underneath. There may be black pixels in the portions of the image which should show the ground underneath. The black pixels may be corrected by pixels from a surround view image of the vehicle. In one combination of features, this correction can be enabled using the homography matrix. The current top view image with the corrected pixels is then saved in the storage device. This can be used as the previous top image when the next current top view image is generated.
[0049] An example of the view underneath the vehicle rendered to the driver, will be explained using FIG 3. The image shown in fig 3 is the view that may be rendered to the driver. The rectangular area (301) maybe the ground underneath the vehicle. This may be generated using the top view images of the vehicle, and performing feature matching, alignment correction, mapping, stitching and harmonization as explained in the above paragraphs. The portion underneath the vehicle (301) may be generated using the process explained in the above paragraphs.
[0050] A line (302) is shown on the ground where the vehicle is driving. When the vehicle passes over the line, it is shown to the driver in the view of the ground underneath the vehicle. As shown in the pic, the part of line outside the vehicle and the part underneath the vehicle is aligned. There is no distortion. The view of the ground underneath the vehicle is harmonized with the rest of the vehicle.
[0051] In one combination of features which may be combined with alternative combinations described above, a vehicle maybe configured appropriately to implement the method as explained above. The vehicle may have one or more cameras mounted on it. These cameras may be used to generate top view image of the vehicle. One example of generating top view image can be read from patent US10152816B2 at pages 19-20, columns 4 to 6.
[0052] In one combination of features, the vehicle may be configured to process the top view images in the manner as explained in the above paragraphs. The vehicle may be implemented with on or more image processors which may be configured to perform feature matching between a current top view image, and a previous top view image. The image processors may be configured to identify and correct alignment difference between each pixel of the current top view image and each pixel of the previous top view image based on the feature matching, and then mapping the previous top view image to the current top view image based on the identified alignment difference resulting in a combined top view image. In one combination of features of this disclosure, the image processors may be further configured to extract a region of interest from the combined top view image, and continuously stitching the extracted region of interest to the current top view image for generating a continuous view of the ground underneath the moving vehicle.
[0053] In one combination of features which may be combined with alternative combinations described above, the vehicle may be implemented with storage means configured to receive and transfer the top view images as explained in the above paragraphs. The storage means may also be configured to store the images captured by the one or more camera mounted on the vehicle.
[0054] While embodiments of the invention have been particularly shown and described with reference to specific embodiments, it should be understood by those skilled in the art that various changes in form and detail may be made therein without departing from the spirit and scope of the invention as defined by the appended claims. The scope of the invention is thus indicated by the appended claims and all changes which come within the meaning and range of equivalency of the claims are therefore intended to be embraced. It will be appreciated that common numerals, used in the relevant drawings, refer to components that serve a similar or the same purpose.
[0055] It will be appreciated to a person skilled in the art that the terminology used herein is for the purpose of describing various embodiments only and is not intended to be limiting of the present invention. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0056] It is understood that the specific order or hierarchy of blocks in the processes / flowcharts disclosed is an illustration of exemplary approaches. Based upon design preferences, it is understood that the specific order or hierarchy of blocks in the processes / flowcharts may be rearranged. Further, some blocks may be combined or omitted. The accompanying method claims present elements of the various blocks in a sample order, and are not meant to be limited to the specific order or hierarchy presented.
[0057] The previous description is provided to enable any person skilled in the art to practice the various aspects described herein. Various modifications to these aspects will be readily apparent to those skilled in the art, and the generic principles defined herein may be applied to other aspects. Thus, the claims are not intended to be limited to the aspects shown herein, but is to be accorded the full scope consistent with the language claims, wherein reference to an element in the singular is not intended to mean “one and only one” unless specifically so stated, but rather “one or more.” The word “exemplary” is used herein to mean “serving as an example, instance, or illustration.” Any aspect described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other aspects. Unless specifically stated otherwise, the term “some” refers to one or more. Combinations such as “at least one of A, B, or C,” “one or more of A, B, or C,” “at least one of A, B, and C,” “one or more of A, B, and C,” and “A, B, C, or any combination thereof’ include any combination of A, B, and / or C, and may include multiples of A, multiples of B, or multiples of C. Specifically, combinations such as “at least one of A, B, or C,” “one or more of A, B, or C,” “at least one of A, B, and C,” “one or more of A, B, and C,” and “A, B, C, or any combination thereof’ may be A only, B only, C only, A and B, A and C, B and C, or A and B and C, where any such combinations may contain one or more member or members of A, B, or C. All structural and functional equivalents to the elements of the various aspects described throughout this disclosure that are known or later come to be known to those of ordinary skill in the art are expressly incorporated herein by reference and are intended to be encompassed by the claims.
Claims
1. A computer implemented method for providing a continuous view of a ground underneath a moving vehicle, comprising,performing a feature matching between a current top view image, and a previous top view image;identifying and correcting an alignment difference between each pixel of the current top view image and each pixel of the previous top view image based on the feature matching;mapping the previous top view image to the current top view image based on the identified alignment difference, resulting in an aligned previous top view image;extracting a region of interest from the aligned previous top view image; andstitching the extracted region of interest from the aligned previous top view image to the current top view image.
2. The method as claimed in claim 1, wherein the alignment difference is identified by calculating a rotation and translation of each pixel of the current top view image, using a homography matrix and the feature matching.
3. The process as claimed in claims 1 &2, wherein the step of correcting the alignment difference comprises calculating a result from the pixel coordinates of the current top view image and the homography matrix.
4. The method as claimed in claims 1- 3, wherein extracting the region of interest comprises calculating a dimension of the region of interest based on a dimension of the vehicle and a scale factor of the previous top view image.
5. The method as claimed in claims 1-4, wherein extracting the region of interest is done for each pixel.
6. The method as claimed in claims 1—5, wherein the current top view image is harmonized after stitching the extracted region of interest.
7. The method as claimed in claims 1 - 6, wherein an outline of the stitched region of interest is displayed in the harmonized current top view image.
8. The method as claimed in claims 1-7, comprising continuously performing the method as claimed above until a driving control unit in the vehicle is turned off.
9. The method as claimed in claims 1- 8, wherein the harmonized current top view image is transferred to the storage device in the vehicle.
10. A computer implemented system for providing a continuous view of a ground underneath a moving vehicle, comprisinga memory; andat least one processor coupled to the memory and configured to carry out the method of any one of claims 1 to 9.
11. The computer implemented system as claimed in claim 10, wherein one or more cameras mounted on the vehicle are used to generate a top view image.
12. A non-transitory computer-readable storage medium comprising instructions, which when executed by a processor, performs the process of any one of claims 1 to 9.
Citation Information
Patent Citations
Vehicle chassis area image generation method, electronic device and storage medium
CN114493990A
Apparatus and method for displaying information
GB2559759A
Method for image restoration in a computer vision system
US20110273582A1
Image stitching with an adaptive three-dimensional bowl model of the surrounding environment for surround view visualization
WO2023192754A1