Methods and apparatus of determining a road width

WO2026202142A1PCT designated stage Publication Date: 2026-10-01JAGUAR LAND ROVER LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2026/058528
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-25
Filing Date
2026-03-25
Publication Date
2026-10-01

Smart Images

  • Figure EP2026058528_01102026_PF_FP_ABST
    Figure EP2026058528_01102026_PF_FP_ABST
Patent Text Reader

Abstract

Aspects relate to a computer-implemented method (200) of determining a width of a road in a direction of travel of a vehicle. The method comprises receiving (202), as input from an image capture device, an image within an image viewport (304) of the image capture device, wherein the image comprises a view in the direction of travel of a vehicle, the view capturing an area of ground, and wherein the image viewport comprises a plurality of image pixels; classifying (204) the image to determine a portion of the image pixels representative of the road in the direction of travel and a further portion of the image pixels representative of non-road background; determining (206) a real-world width of the road from the classified image in dependence on a predetermined mapping relationship of the image viewport and a ground level plane (316) of the area of ground captured in the image; and outputting (208) the determined real-world width of the road.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] METHODS AND APPARATUS OF DETERMINING A ROAD WIDTH

[0002] TECHNICAL FIELD

[0003] The present disclosure relates to determining the width of a road in a direction of travel of a vehicle. Aspects of the invention relate to a computer implemented method, an apparatus comprising a processor and memory, a system, a vehicle, and computer readable instructions.

[0004] BACKGROUND

[0005] It is known to detect the location of a road marked road or lane, using a camera. There is a challenge in the art to detect the location, and in particular to determine the dimensions, of a road in a camera image which is not marked using road or lane markings.

[0006] It is an aim of the present invention to address one or more of the disadvantages associated with the prior art.

[0007] SUMMARY OF THE INVENTION

[0008] Aspects and embodiments of the invention provide a computer implemented method, an apparatus, a system, a vehicle, and computer readable instructions executable by one or more processors, as claimed in the appended claims

[0009] According to an aspect of the present invention there is provided a computer-implemented method of determining a road dimension from an image of the road, the method comprising: receiving an image as input from an image capture device, the image formed from a plurality of image pixels and comprising a view capturing an area of ground; classifying the image to determine a portion of the image pixels representative of a road and a further portion of the image pixels representative of non-road background; determining a real-world dimension of the road from the classified image in dependence on a predetermined mapping relationship of the image capture device and a ground level plane of the area of ground captured in the image; and outputting the determined real-world dimension of the road. Advantageously, a real-world dimension of a road ahead can be determined through computer implemented image recognition and mathematically mapping the road portion of the image to the corresponding road portion in the real world. The mathematical mapping may be computationally quick to perform so the overall process including image capture, image classification, and mapping, may be relatively quick to allow near real time determination of the road dimension from the image of the area including the road.

[0010] According to an aspect of the present invention there is provided a computer-implemented method of determining a width of a road in a direction of travel of a vehicle, the method comprising: receiving, as input from an image capture device, an image within an image viewport of the image capture device, wherein the image comprises a view in the direction of travel of a vehicle, the view capturing an area of ground, and wherein the image viewport comprises a plurality of image pixels; classifying the image to determine a portion of the image pixels representative of the road in the direction of travel and a further portion of the image pixels representative of non-road background; determining a real-world width of the road from the classified image in dependence on a predetermined mapping relationship of the image viewport and a ground level plane of the area of ground captured in the image; and outputting the determined real-world width of the road.

[0011] The image viewport is the area of the image capture device which captures the image of the view. It comprises the plurality of image pixels which have a physical size. As a simple example, for an image viewport 100pm x 100pm in area comprising 10,000 square pixels, each pixel has dimensions of 1pm x 1pm in a 100 x 100 grid of pixels. The image viewport may equivalently be called a pixel viewport, or simply, a viewport.

[0012] Advantageously, a real-world dimension of a road ahead can be determined through computer implemented image recognition to determine which portion of an image shows a road, and through mathematically mapping the portion of theimage showing the road to a real world dimension of the road via a predetermined mapping relationship. In this way, the real world size of a road ahead can be automatically determined from an image of the road.

[0013] The vehicle may be a road vehicle.

[0014] Classifying the image may comprise identifying the image pixels representative of road in the absence of detectable road markings. Advantageously, the road portion of an image may be identified in an image via image classification, even if the road is not clearly delineated by standard road markings such as painted road lines and / or curbs. This may be particularly useful if the vehicle is travelling off-road.

[0015] The predetermined mapping relationship may comprise one or more of: a known height of the image capture device from ground level, a focal length of the image capture device, and a size of the image pixels. The size of the image pixels is the physical size of each pixel in the image viewport. Advantageously, mapping from the captured image of the road to a real-world dimension of the road may be performed using known and well-defined dimensions independent of the particular road.

[0016] The predetermined mapping relationship maps the image viewport to the corresponding ground level plane captured in the image by transforming coordinates in a plane of the image viewport to corresponding coordinates in the ground level plane in dependence on the known height of the image capture device from ground level and the focal length of the image capture device. Advantageously, mapping from the captured image of the road to a real-world dimension of the road may be performed using a geometric transformation, which can be performed efficiently computationally.

[0017] The predetermined mapping relationship may comprise an approximation that a plane of the image viewport is perpendicular to the corresponding ground level plane. The predetermined mapping relationship may include the assumption (i.e. approximation) that the plane of the image viewport (e.g. in a vertical plane) is perpendicular to the corresponding ground level plane (i.e. a horizontal plane). An optical axis of the image capture device may be parallel with the corresponding ground level plane. The optical axis of the image capture device may be perpendicular with (normal to) the plane of the image viewport. Advantageously, by making the simple assumption of the image viewport being perpendicular to the ground level plane, the geometric transformation between the two may be performed computationally efficiently.

[0018] The predetermined mapping relationship may comprise the relationship:

[0019] h / h \

[0020] (xr,zr) = f(xv,yv) = (xvd —,fl 1—7 - 11)

[0021]

[0022] JV VV d

[0023] where xr= real world coordinate in x direction; zr= real world coordinate in z direction; xv= view port coordinate in x direction; yv= view port coordinate in y direction; d = pixel size; h = height of image capture device from ground level; fl = focal length of the image capture device. Advantageously, a simple mathematical relationship may be used to determine the real world coordinates of the road based on the captured image of the road. Because the processing involved to map from the captured image to the real world road coordinates is relatively simple, the determination may be computationally made very quickly.

[0024] The size of the image pixels may be determined by one or more of: calculation based on a known number of image pixels of the image viewport and size of an image sensor of the image capture device; and calibration of the image capture device by capturing an image of a calibration device comprising features of known dimensions positioned on the optical axis of the image capture device. Advantageously, the characteristics of the pixels of the image capture device may be determined in different ways, and may allow for the method to be performed using an image capture device which captures images having some optical distortion, such as those captured by a fisheye lens.Classifying the image may comprise using a Neural Network (e.g. a Convolutional Neural Network, CNN) trained on images comprising roads to identify the image pixels representative of road and the image pixels representative of non-road background. Advantageously, the road portion captured in the imae may be discriminated from the nonroad background portion of the image using a neural network. This may allow for accurate discrimination of location of the road and the image

[0025] Outputting the determined real-world width of the road may comprise overlaying an indication of the road width on a displayed image of the view in the direction of travel of the vehicle. Advantageously, a driver may be presented with an indication of the dimensions of the road ahead in a simple and intuitive way.

[0026] The image of the view in the direction of travel of the vehicle may be a live image feed captured by the image capture device. The live image feed may include a shorter time delay or time lag between image capture and image display when compared with previous systems. Advantageously, because the method can be computationally performed quickly, at least in part because the transformation from the image viewport to the real world ground level may be performed using a mathematical relationship which may be calculated quickly, the calculated indication of the road width may be displayed on the image as captured in real time, for example as the vehicle is driving along the road. In this way, for example, a driver receives information in real-time as to whether the vehicle may fit on the road ahead for example on a narrow track or off-road track having obstacles to one or either side of the road.

[0027] Outputting the determined width of the road may comprise overlaying an indication of the road width on the live image capture feed to provide an augmented reality image of the road in the direction of travel of the vehicle. Classifying the image may comprise identifying the image pixels representative of road comprising gravel, slate, tarmac, water, soil, and mud. Advantageously, the road portion of an image may be identified in an image via image classification for various different types of road surface. The image may be pre-processed by resizing using a nearest neighbour method prior to classification. Advantageously, pre-processing may improve image classification.

[0028] The image capture device may comprise a fish eye lens camera. The method may comprise: capturing a calibration image captured using the fish eye lens camera; correcting for distortion in the calibration image to determine a correction function for images captured by the fish eye lens camera; and performing correction of the image of the view in the direction of travel of a vehicle using the correction function prior to classifying the image. Advantageously, captured images may be corrected using a prior calibration of the image capture device, i.e. a fisheye lens camera, to allow for accurate determination of road dimensions in the real world.

[0029] The method may further comprise transmitting the determined width of the road to a storage means for subsequent retrieval. The storage means may comprise one or more of: a portable electronic device, a remote server, the cloud, and a storage means located at the vehicle. Advantageously, the results of the process may be stored for later retrieval, for example if the same route is subsequently travelled by the vehicle (or another vehicle with access to the stored results) then the results may be retrieved and provided without necessarily recalculating the road width.

[0030] In an aspect of the present invention, there is provided an apparatus comprising at least one processor; and a memory coupled to the at least one processor, the memory storing instructions that when executed by the at least one processor cause the apparatus to implement any method disclosed herein. Advantageously, the method may be performed computationally, for example on an on-board computer of the vehicle.

[0031] The apparatus may comprise one or more controllers collectively comprising at least one electronic processor having an electrical input for receiving an input signal; and at least one memory device electrically coupled to the at least one electronic processor and having instructions stored therein; and wherein the at least one electronic processor is configured to access the at least one memory device and execute the instructions thereon so as to implement any method disclosed herein.In an aspect of the present invention, there is provided a system for a vehicle, the system comprising: any apparatus disclosed herein, an image capture device; and a display screen configured to display the image of the view in the direction of travel of the vehicle and display an indication of the determined real-world width of the road. Advantageously, the system of apparatus, along with an image capture device and display screen, may be installed in a vehicle, or these system elements may already be present in a vehicle and adapted for use as described herein. Further, the system may be retrofitted to an existing vehicle.

[0032] The image capture device may be located at one or more of a bonnet of the vehicle, a windshield of the vehicle, a tail portion of the vehicle, a front bumper, and a rear bumper.

[0033] In an aspect of the present invention, there is provided a vehicle comprising: any apparatus disclosed herein configured to perform any method disclosed herein; or any system disclosed herein configured to perform any method disclosed herein.

[0034] In an aspect of the present invention, there are provided computer readable instructions which, when executed by one or more processors, cause the one or more processors to perform any method disclosed herein.

[0035] Within the scope of this application it is expressly intended that the various aspects, embodiments, examples and alternatives set out in the preceding paragraphs, in the claims and / or in the following description and drawings, and in particular the individual features thereof, may be taken independently or in any combination. That is, all embodiments and / or features of any embodiment can be combined in any way and / or combination, unless such features are incompatible. The applicant reserves the right to change any originally filed claim or file any new claim accordingly, including the right to amend any originally filed claim to depend from and / or incorporate any feature of any other claim although not originally claimed in that manner.

[0036] BRIEF DESCRIPTION OF THE DRAWINGS

[0037] One or more embodiments of the invention will now be described, byway of example only, with reference to the accompanying drawings, in which:

[0038] Figure 1 shows a vehicle comprising an apparatus configured to perform a method of determining the width of the road in an image in accordance with examples of the invention;

[0039] Figure 2 shows a method of determining the width of the road in an image in accordance with examples of the invention; Figures 3A and 3B show an apparatus configured to perform the method of determining the width of the road in an image in accordance with examples of the invention;

[0040] Figure 4 shows a geometric model of an image capture device and captured scenery according to which the width of the road captured in an image may be determined in accordance with examples of the invention;

[0041] Figure 5 shows an example classified image in accordance with examples of the invention;

[0042] Figure 6 shows an example visual output of an image captured of the road ahead of a vehicle with a determined road width indicated thereon in accordance with examples of the invention;

[0043] Figure 7 shows an example system in accordance with examples of the invention;

[0044] Figure 8 illustrates a method of training a machine learning model to classify an image in accordance with examples of the invention; and

[0045] Figure 9 illustrates an example overall method of obtaining and providing a road width from an image in accordance with examples of the invention.

[0046] DETAILED DESCRIPTION

[0047] It is known to detect the location of a road which includes road markings by using a camera. There is a challenge in the art to detect the location, and dimensions, of a road in a camera image which is not marked using road or lane markings. Moreover, it is a challenge in the art to obtain information about road location and dimensions of the road on the fly as a vehicle travelsalong the road due to the computational overheads involved in processing captured image data and providing timely information about the road. Furthermore, it is a challenge in the art to determine real-world dimensions based on images captured by one camera (monocular image capture) as opposed to a stereo camera setup, or using radar.

[0048] While there are solutions for lane detection and lane width detection which use standard road markings to demarcate a lane, there remains a problem in the art that when a vehicle is travelling on a road or track which does not have a clear road marking. For example, if driving off-road, or for example if the road markings have worn away over time or become obscured by rain, gravel, or other reason, no clear road markings are present to demarcate the location of the road or track.

[0049] Being able to determine the width of the road or track allows a driver of the vehicle to be informed of the width of the road and make a decision as to how to travel along the road, or even whether the vehicle is too wide to pass along the road. If the driver is able to receive a warning indication that the width of the road or track ahead appears too narrow for the vehicle to pass this may be advantageous when driving, e.g. off road, or along narrow country lanes, for example. This also applies in respect of determining whether the road ahead is wide enough to pass if there is an oncoming vehicle also travelling along the road; that is, whether both vehicles are able to pass each other on the road or whether there is not enough room for them to do so because the road is too narrow.

[0050] Examples disclosed herein process captured images of the road ahead to identify portions of the image which show road, and a calculation may be performed to identify the width of the road identified in the image at one or more different distances ahead of the vehicle. The term “ahead” is intended to mean the ground onto which a vehicle is about to move; if the vehicle is driving forwards then the road ahead is the road which can be seen through the front windscreen of the vehicle, and if the vehicle is reversing, then the road ahead is the space behind the vehicle onto which the vehicle will reverse.

[0051] Figure 1 illustrates a vehicle 100 according to examples of the present invention. The vehicle 100 comprises an apparatus 300 as illustrated in Figure 3 configured to perform a method 200 of determining the width of the road in an image as illustrated in Figure 2. The vehicle 100 comprises an image capture device 340, such as a camera, (e.g. a monocular camera, a fish eye lens camera), or other image capture device. In other examples the vehicle 100 may comprise plural such image capture devices 340 (for example, one at the front and one at the rear of the vehicle, or one on the left and one on the right (i.e. driver / passenger sides) of the vehicle).

[0052] Broadly, in the first stage of the method 200 of determining the width of the road in an image, a road surface in an image of landscape ahead is identified. In an embodiment, a convolutional neural network may be used to classify pixels in the image as either showing road surface, or non-road background. In the second stage of the method, a width of the road is identified from a segmented (road vs non-road) image in which road surface has been identified using a mathematical relationship. This can be achieved for different positions along the road ahead and the determined width(s) can be provided to a vehicle user.

[0053] Figure 2 shows a computer-implemented method 200 of determining the width of the road in a direction of travel of a vehicle in an image in accordance with examples of the invention. The method 200 may be performed by the apparatus 300 illustrated in Figure 3 or Figure 4. In particular, with reference to Figure 4 the memory 130 may comprise computer-readable instructions which, when executed by the processor 120, perform the method 200 according to examples of the invention.

[0054] The method 200 comprises: receiving at step 202, as input from an image capture device, an image 210 within an image viewport of the image capture device. The image 210 comprises a view in the direction of travel of a vehicle. The vehicle may be a road vehicle. The view captures an area of ground. The image viewport comprises a plurality of image pixels. An example image is illustrated in Figure 6. The image viewport is the area of the image capture device 340 which captures the image of the view, and it comprises the plurality of image pixels, which each have a physical pixel size. For example, the image viewport may be an image capturing portion of a charged-coupled device (CCD) camera comprising an array of pixels. As a simpleexample, for an image viewport 100pm x 100pm in area comprising 10,000 square pixels, each pixel has dimensions of 1 pm x 1 pm in a 100 x 100 grid of pixels.

[0055] The method 200 comprises at step 204, classifying the image to determine a portion of the image pixels representative of the road in the direction of travel and a further portion of the image pixels representative of non-road background. In this way the image may be computationally analysed to determine which portions of the image show road (which may be, for example, a tarmac-surfaced road, a gravel-surfaced road, a hard earth surfaced road or track, or any other path along which a vehicle may potentially travel), and which portions of the image do not contain road (e.g. containing sky, buildings, footpath, hedgerows, grass, etc). Figure 6 illustrates boundaries 608 separating the road pixel region from the non-road pixel region.

[0056] The method 200 comprises determining at step 206, a real-world width of the road from the classified image in dependence on a predetermined mapping relationship of the image viewport and a ground level plane of the area of ground captured in the image. The predetermined mapping relationship can be used to convert (map) the size (e.g. width) of the portion of the image corresponding to road to the equivalent size (width) of the road in the real world.

[0057] The method 200 then comprises outputting at step 208, the determined real-world width of the road 220. In this way, a real-world dimension of a road ahead can be determined through computer implemented image recognition to determine which portion of an image shows a road, and through mathematically mapping the portion of the image showing the road to a real world dimension of the road via a predetermined mapping relationship. In this way, the real world size of the road ahead can be automatically determined from an image of the road. Advantageously, by using (typically computationally intensive) computerised image recognition to identify portions of the image which indicate road, butthen using (typically computationally simple) mathematics to map the size of the road in the image to the size of the road in real life, the overall process from image capture to road size determination can be computationally performed quickly, i.e. quickly enough for the size to be provided as the vehicle is being driven along the road, such that a real time indication of the road width can be provided to the driver, for example to assist the driver in understanding how much room / space they have either side of their vehicle to drive their vehicle along the road. If the road would be too narrow to navigate along, the driver can be informed of this and choose to take an alternative path, for example.

[0058] “Real-time” should be understood to mean “near real time” to account for the short time taken for signals to be provided from the image capture device to the display device and any associated calculations, for example, categorising the image pixels into road or non-road, and mapping the dimensions of the image of the road to the real world dimensions of the road, but with a small enough delay that the provided visual output is an accurate representation of the actual real world view at the time of output.

[0059] This disclosure covers computer readable instructions (for example, stored on a computer readable medium which may be a non-transitory computer readable medium) which, when executed by one or more processors, cause the one or more processors to perform any method disclosed herein such as that discussed in relation to Figure 2.

[0060] Figures 3A and 3B show an apparatus 300 configured to perform the method 200 of determining the width of the road in an image in accordance with embodiments of the invention. The apparatus 300 comprises at least one processor 120; and a memory 130 coupled to the at least one processor 120, the memory 130 storing instructions that when executed by the at least one processor 120 cause the apparatus 300 to implement any method 200 disclosed herein. Advantageously, the method 200 may be performed computationally, for example on an on-board computer of the vehicle 100.

[0061] Figure 3A illustrates an apparatus 300 configured to receive at block 202, as input from an image capture device 340, an image 210 within an image viewport of the image capture device. The image 210 comprises a view in the direction of travel of a vehicle. The view captures an area of ground. The image viewport comprises a plurality of image pixels. The apparatus300 is configured to classify the image 210 at block 204, to determine a portion of the image pixels representative of the road in the direction of travel and a further portion of the image pixels representative of non-road background. The apparatus 300 is configured to determine at block 206, a real-world width of the road from the classified image in dependence on a predetermined mapping relationship of the image viewport and a ground level plane of the area of ground captured in the image. The apparatus 300 is configured to output at block 208, the determined real-world width of the road 220, for example, to a display device 350 as an output device 350 to indicate the determined real-world width of the road 220 to a user of the vehicle 100.

[0062] Figure 3B shows an apparatus 300 for a vehicle 100. The apparatus 300 comprises one or more controllers 110. The apparatus 300 is configured to receive an input signal 210 indicative of an image provided by an image capture device 340. The apparatus 300 may then output an output signal 220 indicative of the real-world width of the road to an output device 350 such as a display screen (which may also display the image). The apparatus 300 as illustrated in Figure 3B comprises one controller 110, although it will be appreciated that this is merely illustrative. The apparatus 300 comprises processing means 120 and memory means 130. The processing means 120 may be one or more electronic processing device 120 which operably executes computer-readable instructions. The memory means 130 may be one or more memory device 130. The memory means 130 is electrically coupled to the processing means 120. The memory means 130 is configured to store instructions, and the processing means 120 is configured to access the memory means 130 and execute the instructions stored thereon.

[0063] The controller 110 comprises an input means 140 and an output means 150. The input means 140 may comprise an electrical input 140 of the controller 110. The output means 150 may comprise an electrical output 340 of the controller 110. The input 140 is arranged to receive an input signal 210 indicative of an image provided by an image capture device 340. The output 150 is arranged to output a signal 220 indicative of the determined real-world width of the road, for example, to a display device 350.

[0064] The method 200 of Figure 2 and the apparatus 300 of Figures 3A and 3B may also comprise the following features.

[0065] Classifying the image may comprise identifying the image pixels representative of road in the absence of detectable road markings. That is, it is not necessary for the image of the road to capture demarcations indicative of the presence of a road, such as kerbstones, painted road markings, cats eyes, or other intentional road demarcation. The image classification may work for roads which are not intentionally marked as such - for example, tracks over grass, gravel, or soil, or roads which are marked by the presence of hedgerows or long grass either side of the roadway. In some examples the road may comprise two wheel tracks having a “wild” area between, such as grass, gravel, or loose dirt, as illustrated in Figure 6. Thus the present invention may be advantageously employed if driving “off-road”, since the road portion of an image may be identified via image classification, even if the road is not clearly delineated by standard road markings. Roads and pathways of different appearances may be processed to identify road and non-road portions using machine learning models trained on appropriate road images (e.g. to accurately identify areas of images showing road under wet conditions, training images of roads taken in similar wet conditions may be used to train the machine learning model). It should be noted that the present invention may equally be applicable to a road with markings or to a road with no clear markings.

[0066] In some examples, the determined width of the road may be transmitted to a storage device / storage means for subsequent retrieval. The storage means may comprise one or more of: a portable electronic device, a remote server, the cloud, and a storage means located at the vehicle. Advantageously, the results of the road width determination process may be stored for later retrieval, for example if the same route is subsequently travelled by the vehicle (or another vehicle with access to the stored results) then the results may be retrieved and provided without necessarily recalculating the road width. The calculated road width results may be stored alongside GPS coordinates of the location of vehicle at the time of image capture, or the time of road width determination, for example, so they may be identified and retrieved by a subsequent vehicle located at thesame GPS coordinates (e.g. within some geofence tolerance) or as part of a process carried out at a back end computer by entering a location identified (e.g. GPS coordinates) as input and retrieving the calculated road width(s) forthat location.

[0067] Figure 4 shows a geometric model of an image capture device and captured scenery according to which the width of the road captured in an image may be determined in accordance with examples of the invention. Such a model may be used to perform the method step of determining 206 a real-world width of the road from the classified image in dependence on a predetermined mapping relationship of the image viewport and a ground level plane of the area of ground captured in the image.

[0068] Figure 4 illustrates how light rays pass from locations (objects) 312, 316 in the real world, to the image capture device image viewport 304a, 304b, according to the focal length 308a, 308b of the image capture device. The image viewport 304a, 304b provides a pixelized representation of what the image capture device detects when it captures an image. As illustrated, the “initial” image viewport 304a that is shown in front of the focal point 302 at the focal distance 308a. The equivalent “far” image viewport 304b is the viewport on which the image of the environment 312, 316 forms at the focal distance 308b from the focal point 302. This “far” image viewport 304b was used in the example calculations discussed below, though it is equivalent to using the “initial” image viewport 304a. That is, the “initial” image viewport 304b positioning can be used because in reflection about a x-y plane containing the focal point 302 it is the same distance away, 308a, 308b as the “far” image viewport 304a (that is, the focal distance 308a is the same as the focal distance 308b). So, the calculations end up being the same whichever image viewport 304a, 304b is chosen. Image viewport 304a is technically the correct place where the image is formed in a physical camera / image capture device; in this example image viewport 304b is the viewport used for the calculation example below (and consequently this is where the origin of the coordinate system was set in this example).

[0069] In this example, the method of determining road width is based on a projective transformation and some initial assumptions. The initial assumptions used to determine the width of the road include that the road is flat, and that the road is parallel to a plane constructed from the camera having an optical axis parallel to the plane of the four wheels of the vehicle. In other words the image captured by the image capture device is assumed to capture a plane which is perpendicular to a flat ground plane on which the vehicle is travelling. Each assumption can be tested and adjusted for using information from additional vehicle signals and a simple chi-squared test to determine the error between observed and expected optical flow measurements. The chi-squared test may be based on optical flow (i.e. a stream of images provided by the image capture device) which identifies corresponding points in successive captured image frames, and determines the change in road width between successive frames.

[0070] In this example model, the image capture device is modelled as a pinhole camera at focal point 302, to provide a single focal point 320 in space at a height A 310 from the road level. The image capture device captures an image in an image viewport 304a, 304b, which is an area of pixels of the image capture device which capture the image of the view in front of the image capture device. In this example, the captured image includes a region of the road 316 illustrated to indicate how the mapping relationship may be implemented. The optical axis 314 of the image capture device is illustrated and assumed to be parallel with the plane of the ground, along which the x-axis lies as shown.

[0071] Four points 312 on the road surface are indicated which correspond to four corners of a portion of a road 316 captured in the image viewport 304a, 304b of the image capture device. Using a known width of a pixel in the image as captured by the image capture device (e.g. from knowledge of the specifications of the image capture device which sets out the number of pixels wide by the number of pixels high the image viewport is, and the size of each of those pixels in the image viewport), a projection may be made to determine the size of a portion of real-world space 316 corresponding to a pixel width in the captured image. The properties of the image capture device, i.e. the focal length fl 308a, 308b, can be determined in an image capture device calibration step as described below. In this example, the focal point 302 of the image capture device is illustrated having an image viewport 304a in front (i.e. the image viewport 304a is located between the focal point 302 andthe captured view 316), and an equivalent image viewport 304b of the image capture device behind (obtained as a mirror image of the image viewport 304b behind the focal point 302 in the plane of the image viewports 304a, 304b).

[0072] The model assumes the real-world coordinates of points on the road 312 are of interest and it is from these the width of the road can be determined. This model assumes that the road 316 is flat (as shown, it is in the x-z plane), and perpendicular to the image viewport 304a, 304b (as shown, they are in the x-y plane). This assumption means that one Cartesian axis / plane can be eliminated from mathematical consideration making the calculations more computationally efficient to perform. Assuming that the height 310 of the image capture device 302 from the road level 316, h, the focal length 308a, 308b of the camera, and the real world size represented by each pixel of the image viewport 304a, 304b, d, that each pixel in the image capture device sensor represents, the relationship can be solved using linear algebra to determine the real world positions of objects / regions 316 captured in the image sensed by the image capture device.

[0073] At each point 306 (xv,yv) where a light ray from an object in front of the image capture device sensor hits the image capture device sensor (the image viewport 304a, 304b), the colour of the pixel in that location represents the colour of the light that hits it from the real-world object. The object 312, 316 in the real world can be designated 3D coordinates A = (xr,yr,zr). Processing is performed to obtain the three-dimensional coordinates of the object 312, 316 in the real world from the known two-dimensional coordinates of the equivalent point in the image viewport 304a, 304b.

[0074] The following variables can be considered:

[0075] fl - focal length 308a, 308b of the camera

[0076] h - height 310 of the camera from the road

[0077] d - real world pixel size (i.e. the size of a pixel of the image viewport)

[0078] ( v,yv) - coordinates of a pixel in the image viewport 304a, 304b, with the origin located at the point where the camera’s optical axis 314 meets the plane of the image viewport 304a, 304b.

[0079] In Figure 4, the location / position of the focal point 308a, the image capture device 302 may be represented by the vector C = 0

[0080] h . The real world coordinate system may be selected such that the x coordinates describes the displacement from the -fl

[0081] image capture device’s optical axis 314 parallel to the road, the y coordinate is the perpendicular height from the road, and the z coordinate is the horizontal displacement of the image in the image viewport 304a parallel with the ground. The location xv

[0082] of a point in the image viewport 304a may be represented by the vector yv, which corresponds to a location of an equivalent -0 - xr

[0083] point in real-world space which may be represented by the vector 0

[0084] zr

[0085] Figure 4 shows that the image viewport 304a frame of reference is displaced from the origin of the real-world equivalent “frame” 316 by a distance h 310 along the -axis. A transformation may be performed from the viewport 304a coordinate system to the real world 316 coordinate system by adding the image capture device height h 310 and multiplying by the real-world pixel size d so that:

[0086]

[0087] which maps a point on the road 312 into the image captured in the image viewport 304a by the image capture device. The variable .v, for the x coordinate position in the real world may be assigned to the relation xvd, representing the x position in the image viewport multiplied by the height of the image capture device <7310. Similarly the variable yrfor they coordinate positionin the real world may be assigned to the relation ^ / representing they position in the image viewport multiplied by the height of the image capture device d 310. The term arrepresents a variable corresponding to a point in the real world.

[0088] A straight-line L connecting the image capture device focal point 302 to a pixel 306 may therefore be described by the following equation:

[0089] 0 + (xrt — 0)t xrt

[0090] h - (( / i + yr) - h)t h — yrt h — yrt

[0091] -fl + (0 - (- / 7))t. -fl + ft .fl(t - 1)

[0092]

[0093] where t is a parameter indicating how far from the image capture device, the point on the line (x;,y;,z;) is. In order to find the coordinates on the road that the pixel corresponds to, the point where the straight lineL meets the road needs to be identified. Since by definition, they coordinate represents the height above the road, this can be equated to 0 to find t.

[0094]

[0095] Now a substitution may be made back into the equation of the line L to obtain the real-world coordinates:

[0096] h

[0097] fK- - l)

[0098]

[0099] Vr

[0100] Since a mathematical model has been constructed in a way that the coordinate is constant for avand the z coordinate is constant for A, this may be performed for all pixels. The term avrepresents a variable corresponding to a point in the image and Arrepresents a variable corresponding to a point in the real world. The variable «, may be given in terms of image pixel dimensions, and may be given in terms of real-world measurements (e.g. metres).

[0101] For any road pixel av= (xv,yv), the real-world coordinates of the equivalent part of the road may be computed using a function f : R2-> R2-.

[0102]

[0103] where xr= real world coordinates in x direction; zr= real world coordinate in z direction; xv= image viewport coordinate in x direction; yv= image viewport coordinate in y direction; d = pixel size; h = height of image capture device from ground level; fl = focal length of the image capture device. This relationship may be considered to be a predetermined mapping relationship transforming the plane of the image viewport 304a to the equivalent plane on ground level 316 where the road is located.

[0104] Thus it can be seen that the predetermined mapping relationship may comprise a known height h of the image capture device from ground level, a focal length fl of the image capture device, and a size of the image pixels. The relatively simple mapping from the captured image of the road per the image viewport 304a, to a real-world dimension of the road 316, may thus be performed using known and well-defined dimensions which are independent of the particular road, and the simple relationship means the mapping is a computationally efficient process. The predetermined mapping relationship maps the image viewport 304a, 304b to the corresponding ground level plane 316 captured in the image by transforming coordinates in a plane of the image viewport 304a, 304b to corresponding coordinates in the ground level plane 316 in dependence on the known height 310 of the image capture device from ground level and the focal length 308a, 308b of the image capture device. The calculation may be performed in real time, allowing for an indication of the width of the road to be provided to a vehicle user as they approach that portion of the road in the vehicle, allowing for timely information to be provided.As noted above, the predetermined mapping relationship may comprise an approximation that a plane of the image viewport 304a, 304b is perpendicular to the corresponding ground level plane 316. The predetermined mapping relationship may include the assumption (i.e. approximation) that the plane of the image viewport 304a, 304b (e.g. in a vertical plane) is perpendicular to the corresponding ground level plane 316 (i.e. a horizontal plane). An optical axis 314 of the image capture device may be parallel with the corresponding ground level plane 316. The optical axis 314 of the image capture device may be perpendicular with (normal to) the plane of the image viewport 304a, 304b. By making the simple assumption of the image viewport 304a being perpendicular to the ground level plane 316, the geometric transformation between the two may be performed computationally efficiently.

[0105] In considering the image viewport 304a, 304b and the size of the pixels in the image viewport, the size of the image pixels may be determined by calculation based on a known number of image pixels of the image viewport and size of an image sensor of the image capture device. Such information may be obtained from the image capture device manufacturer. In some examples, the size of the image pixels may be determined by calibration of the image capture device by capturing an image of a calibration device comprising features of known dimensions positioned on the optical axis of the image capture device. For example, a chessboard / chequerboard pattern or grid calibration device comprising features (gridlines and grid squares) of known dimension, may be imaged from different angles. From the known calibration device feature dimensions, and the position of the image capture device from the calibration device, the pixel size may be obtained. For example, using captured calibration images, a matrix may be constructed identifying the position of the optical axis, distortion coefficients, and focal lengths. Such a matrix may be used to perform compensation for images captured by the camera for more accurate determination of the road width. Advantageously, the characteristics of the pixels of the image capture device may be determined in different ways, and may allow for the method to be performed using an image capture device which captures images having some optical distortion, such as those captured by a fisheye lens.

[0106] In some examples, the image capture device may comprise a fish eye lens camera (alone as a monocular imaging system, or as an image capture device of a plurality of such devices in a multi-device imaging system). Therefore, in examples in which a fisheye lens or other lens providing a distorted view of the environment are used as the image capture device, a calibration image may be captured using the fish eye lens camera. The captured calibration image may be corrected for distortion to determine a correction function for images captured by the fish eye lens camera. Then correction of a subsequently captured image, such as the image of the view in the direction of travel of a vehicle may be performed using the correction function prior to classifying the image.

[0107] Figure 5 shows an example classified image 500 in accordance with examples of the invention. An image 500 of the environment in the direction of travel of the vehicle is illustrated. The central illustration 510 indicates a “true mask” whereby a human has manually defined pixels in the image as being road regions 514 or non-road background 512. The illustration on the right 520 indicates a “predicted mask” obtained by running a computer implemented classification of pixels in the image according to the pixels being identified as road regions 524 or non-road background 522. It can be seen that the computer implemented prediction of the location of the road 524 compared to non-road regions 522 of the image closely matches the classification performed manually illustrated in the central image 510. Figure 5 thus illustrates an example result of the method step of classifying 204 the image 500 to determine a portion of the image pixels 524 representative of the road in the direction of travel and a further portion of the image pixels 522 representative of non-road background.

[0108] A worked example of the classification step will now be explained. In this example, the results of which are shown in Figure 5 as performed by the inventors, a convolutional neural network model was trained by first manually segmenting for example, 1692 training images which were extracted from a video feed. The video was recorded on a 20 minute journey in a vehicle driving along various surfaces including gravel, slate, tarmac, water and mud. A 20% test set was created by selecting every fifth captured image. The test set was manually partitioned into road surface and non-road background by the person who originally drove the route, since they had good knowledge of accurately differentiating between road surface and non-roadbackground in the images. An example is shown in the centre panel of Figure 5. Classifying the image may thus be performed using a Neural Network, NN (e.g. a Convolutional Neural Network, CNN) trained on images comprising roads to identify the image pixels representative of road and the image pixels representative of non-road background. In this way, the road portion captured in the image may be discriminated from the non-road background portion of the image using a neural network, allowing for accurate discrimination of location of the road and the image. Depending on the training data used for the NN model, accurate road identification may be achieved for different environmental conditions (e.g. rain, snow, frost, dry, dusty, etc) by training the model using images captured in similar environmental conditions

[0109] Using the neural network method described above, the inventors found that it is possible to identify the portion of an image which shows a road surface for over 98.5% of pixels in images of the tracks investigated by the inventors. Furthermore, in clear driving conditions, the method disclosed herein can provide accurate distance determinations within an error of ± 5 cm for a track width of 3.5 m, and is able to do so for a range of surfaces including gravel, mud, water, rock, slate, and tracks with a central grass portion. Classifying the image may thus comprise identifying the image pixels representative of road comprising, for example, gravel, slate, tarmac, water, soil, and mud. Advantageously, the road portion of an image may be identified in an image via image classification for various different types of road surface.

[0110] In a pre-processing stage, captured images and label data were resized to 128 x 128 pixels using a nearest neighbour method to preserve hard edges for segmentation. By pre-processing the image, for example by resizing using a nearest neighbour method prior to classification, the image classification may be improved. The images were then normalised to a range of between zero and one for binary classification of each pixel into either road or non-road background. An augmentation layer was applied which provided a random transformation to the dataset, to improve the variability in the training data and thus improve the model’s ability to generalise from the data it was trained on to better identify road and non-road background in images in use.

[0111] To train the convolutional neural network, a U-Net convolutional neural network model was used. Using a pre-trained MobileNetV2 as an encoder, the model extracted features from several intermediate layers of the MobileNetV2 model which were passed through an up sampling stack to recover the spatial resolution, reconstructing the original image with an output mask pixel classification generated by the U-Net model. The model was then refined, by presenting a further for example, 1353 images containing roads and non-road background to the convolutional neural network. After each epoch the accuracy of the model was evaluated against the test set. The training and validation cycle continued until the percentage change in accuracy was less than 0.1%. The accuracy was represented by the relation

[0112]

[0113] > I I — I I 1 \ + 1 , where TP means true positive, TN means true negative, FP means false positive and FN means false negative.

[0114] Figure 6 illustrates an example of visual output 600 provided by a display device 350 in receipt of the output from the apparatus described above, the output providing the determined real-world width of the road 220. The display device 350 may be a display screen in the vehicle, or of a portable display device (e.g. a vehicle user’s smartphone, rear view mirror or sat nav system) located in a vehicle and in communication with the apparatus 300 of the vehicle. In this example the display device 350 also displays the images captured by the image capture device of the vehicle of the road ahead of a vehicle. The images may be displayed in real time as the vehicle travels through the environment.

[0115] The width of the road 606a-606c, determined to be within the boundaries 608 either side of the road as determined in an image classification step, may be illustrated by overlaying an indication of the determined width of the road, in one or more positions along the road. In the illustrated example, the determined width of the road is illustrated in three different positions 606a, 606b, 606c along the road. Outputting the determined real-world width of the road may thus comprise overlaying an indication of the road width 606a-c on a displayed image 600 of the view in the direction of travel of the vehicle.Advantageously, a driver may be presented with an indication of the dimensions of the road ahead in a simple and intuitive way.

[0116] Additionally or alternatively, an indication of the road ahead 602 as determined according to method disclosed herein may be displayed not overlaying the road and / or image of the road, for example at the top of the display screen as illustrated. In other examples, the indication of the width of the road ahead may be indicated via separate output means to the display 350 displaying the image of the road ahead. For example, the determined width of the road ahead may be displayed on a separate display means, and / or may be indicated to the driver via audio output, for example.

[0117] As shown, the image 600 of the view in the direction of travel of the vehicle may be a live image feed captured by the image capture device 340. The image capture device 340 may be located at a bonnet of the vehicle, a windshield of the vehicle, a tail portion of the vehicle, a front bumper, and / or a rear bumper. As the height of the image capture device above the road / ground level decreases the error in estimated road width is likely to increase so it may be advantageous to locate the image capture device higher on the vehicle, such as toward the top of the vehicle, rather than lower on the vehicle such as on a bumper. In some examples, such as if using dashcam images, part of the captured image which capture the vehicle bonnet may be masked out to help prevent the image recognition model from accidentally determining the vehicle bonnet to be the road at ground level.

[0118] The live image feed 600 including the indication of the road width ahead 606a-606c may include a shorter time delay or time lag between image capture and image display when compared with previous systems because of the relatively fast computation as discussed in relation to Figure 4 mapping the captured image dimensions to the real world dimensions. Advantageously, the calculated indication of the road width may be displayed on the image 600 as captured in real time, for example as the vehicle 100 is driving along the road. In this way, for example, a driver receives information in real-time as to whether the vehicle may fit on the road ahead for example on a narrow track or off-road track having obstacles to one or either side of the road.

[0119] As illustrated, outputting the determined width of the road 606a-606c may comprise overlaying an indication of the road width on the live image capture feed 600 to provide an augmented reality image of the road in the direction of travel of the vehicle.

[0120] Figure 7 shows an example system 700 in accordance with examples of the invention. The system 700 is for a vehicle 100 and the system 700 comprises any apparatus 300 such as that discussed in relation to Figures 3 and 4; an image capture device 340 (e.g. a camera); and a display screen 350 such as that illustrated in Figure 6, which is configured to display the image 600 of the view in the direction of travel of the vehicle and display an indication 602, 606a-c of the determined real-world width of the road. Thus the system, for a vehicle, comprises any apparatus 300 disclosed herein, an image capture device 340; and a display screen 350 configured to display the image 600 of the view in the direction of travel of the vehicle and display an indication 606a-606c of the determined real-world width of the road. Advantageously, the system 700 may be installed in a vehicle 100, or, the system elements 300, 340, 350 may already be present in a vehicle and adapted for use as described herein.

[0121] Figure 8 illustrates a method 800 of training a machine learning model to classify an image and thereby determine a portion of the image pixels which represent road, and another portion of the image pixels which represent non-road background. It will be appreciated that this specific implementation may be used, but that other implementations are possible and still achieve image classification as disclosed herein.

[0122] In step 802, image pre-processing takes place. For example, captured images and label data of the images (e.g. from a video stream) are each resized to 128 x 128 pixels using a nearest neighbour method to preserve hard edges for segmentation. Then, each image is normalised so that each pixel is within a value range of between zero and one, so that binary classificationof each pixel may be performed to classify a pixel as either road , or non-road background . In some examples, a custom augmentation layer may be applied to an image to perform random transformations in the dataset of pixels of the image, which can improve the variability in the training data and increase the machine learning model’s ability to generalise (and therefore more accurately identify road from non-road background in subsequent images being analysed by the model).

[0123] Step 804 is a stratification process. The pre-processed images are provided and a determination at step 806 is performed to check if there are more images available. If the answer is no (808), the process moves onto the training and validation process 822. If the answer is yes (810), the process moves to a determination step 812 of whether an image is a fifth image or not. If the answer is yes (814), the fifth image is added to a test set at step 818. If the answer is no (816), the image is added to a training set at step 820. The process then returns to the determination step 806 to check if there are more images to process.

[0124] Step 822 is a training and validation process. Once all images have passed through the stratification process 804, a machine learning model is trained 824 on the training set obtained in step 820. The machine learning model may be, for example, a neural network, such as a convolutional neural network, e.g. a U-Net model (a convolutional neural network suitable for image segmentation). After the model has been trained, in step 826, an evaluation is made of the trained model using the test set added in step 818. In step 828, a determination is made regarding whether the accuracy of the trained model, based on evaluation using the test set in step 826, is above a predetermined accuracy threshold (e.g. an accuracy above 90%, 95%, 98%, or other selected accuracy threshold). If the accuracy is below the predetermined accuracy threshold, the process returns (830) to training the model. If the accuracy is above the predetermined accuracy threshold, the process completes (832). For example, at each epoch, if the relevant metrics have improved, the weights of that model are saved, if a threshold number (e.g. four) epochs have passed with no improvement, the training process may be halted and the best performing model may be restored.

[0125] Figure 9 illustrates an example overall method 900 of obtaining and providing a road width from an image. It will be appreciated that while this specific implementation may be used, other implementations and variations are possible and still determine and provide a road width as disclosed herein. In step 902, an image of the road is acquired, for example from a video stream captured by a video camera. In step 904, a trained machine learning model such as those described herein is applied to the acquired image to identify pixels in the image showing road, and pixels in the image showing non-road regions. In step 906, horizontally contiguous pixels in the portion of the image identified as showing road are identified (horizontally contiguous pixels may refer to a connected row of pixels along a horizontal, or left to right, direction in the image oriented with the ground towards the bottom of the image in the sky towards the top of the image).

[0126] In step 908, based on the height, or vertical position in the image, of the identified contiguous row of pixels, and on the number of pixels in the row, the real world width represented by the contiguous row of pixels can be computed, as discussed above. In step 910, the image acquired from the video stream may be augmented by displaying the estimation of the road width, for example by overlaying the estimated road width on or near a location on the image where the width applies (see Figure 5 for an example).

[0127] It should be noted that the methods described in the figures are for illustrative purpose only and it may be possible for some steps in the method to be omitted.

[0128] It will be appreciated that various changes and modifications can be made to the present invention without departing from the scope of the present application.

Claims

CLAIMS1. A computer-implemented method of determining a width of a road in a direction of travel of a vehicle, the method comprising:receiving, as input from an image capture device, an image within an image viewport of the image capture device, wherein the image comprises a view in the direction of travel of a vehicle, the view capturing an area of ground, and wherein the image viewport comprises a plurality of image pixels;classifying the image to determine a portion of the image pixels representative of the road in the direction of travel and a further portion of the image pixels representative of non-road background;determining a real-world width of the road from the classified image in dependence on a predetermined mapping relationship of the image viewport and a ground level plane (316) of the area of ground captured in the image; and outputting the determined real-world width of the road.

2. The computer-implemented method of claim 1, wherein classifying the image comprises identifying the image pixels representative of road in the absence of detectable road markings.

3. The computer-implemented method of any preceding claim, wherein the predetermined mapping relationship comprises one or more of:a known height of the image capture device from ground level,a focal length of the image capture device, anda size of the image pixels.

4. The computer-implemented method of claim 3, wherein the predetermined mapping relationship maps the image viewport to the corresponding ground level plane captured in the image by transforming coordinates in a plane of the image viewport to corresponding coordinates in the ground level plane in dependence on the known height of the image capture device from ground level and the focal length of the image capture device.

5. The computer-implemented method of claim 3 or claim 4, wherein the predetermined mapping relationship comprises an approximation that a plane of the image viewport is perpendicular to the corresponding ground level plane.

6. The computer-implemented method of any of claims 3 to 5, wherein the predetermined mapping relationship comprises the relationship:(xr,zr) = f(xv,yf) = (xvd —;,fl (—7-1)yv vvwhere xr= real world coordinate in x direction; zr= real world coordinate in z direction; xv= view port coordinate in x direction; yv- view port coordinate in y direction; d = pixel size; h = height of image capture device from ground level; fl = focal length of the image capture device.

7. The computer-implemented method of any of claims 3 to 6, wherein the size of the image pixels is determined by one or more of:calculation based on a known number of image pixels of the image viewport and size of an image sensor of the image capture device; andcalibration of the image capture device by capturing an image of a calibration device comprising features of known dimensions positioned on the optical axis of the image capture device.

8. The computer-implemented method of any preceding claim, wherein classifying the image comprises using a Neural Network trained on images comprising roads to identify the image pixels representative of road and the image pixels representative of non-road background.

9. The computer-implemented method of any preceding claim, wherein outputting the determined real-world width of the road comprises overlaying an indication of the road width on a displayed image of the view in the direction of travel of the vehicle.10 The computer-implemented method of any preceding claim, wherein the image of the view in the direction of travel of the vehicle is a live image feed captured by the image capture device.

11. The computer-implemented method of any preceding claim, wherein the image capture device comprises a fish eye lens camera, and wherein the method comprises:capturing a calibration image captured using the fish eye lens camera;correcting for distortion in the calibration image to determine a correction function for images captured by the fish eye lens camera; andperforming correction of the image of the view in the direction of travel of a vehicle using the correction function prior to classifying the image.

12. An apparatus comprising:at least one processor; anda memory coupled to the at least one processor, the memory storing instructions that when executed by the at least one processor cause the apparatus to implement the method of any of claims 1 to 11.

13. A system for a vehicle, the system comprising:the apparatus of claim 12;an image capture device; anda display screen configured to display the image of the view in the direction of travel of the vehicle and display an indication of the determined real-world width of the road.

14. A vehicle comprising:the apparatus of claim 12 configured to perform the method of any of claims 1 to 11; orthe system of claim 13 configured to perform the method of any of claims 1 to 11.

15. Computer readable instructions which, when executed by one or more processors, cause the one or more processors to perform the method according to any of claims 1 to 11.