Damage detection

Automated image-based damage detection using machine learning models addresses the challenge of undetected tray damage in large vehicles, enhancing operational efficiency and reducing maintenance costs.

GB2635539BActive Publication Date: 2025-12-19MOTION METRICS INTERNATIONAL CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
GB2023017536
Authority / Receiving Office
GB · GB
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-11-15
Publication Date
2025-12-19
Estimated Expiration
2043-11-15

AI Technical Summary

Technical Problem

Large load carrying vehicles, such as haul trucks, experience damage to their trays due to harsh operating conditions, which goes undetected and worsens over time, leading to increased repair costs and downtime in 24/7 operations like mining, where manual inspections are unsafe or infrequent.

Method used

A method and system using image sensors and machine learning models to identify and segment images of load carrying containers, processing them to detect damage regions, and determine the presence of damage in real-time, enabling automated damage detection and reporting.

Benefits of technology

Enables safe and frequent detection of damage, reducing repair costs and downtime by providing real-time alerts and reports, ensuring vehicles remain operational in demanding environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000001_0000
    Figure 00000001_0000
  • Figure 00000002_0000
    Figure 00000002_0000
  • Figure 00000003_0000
    Figure 00000003_0000
Patent Text Reader

Abstract

Detecting damage in a load carrying container of a load carrying vehicle, such as truck tray of a truck 300. Images of vehicles are collected as they move around a worksite 301 and regions of interest
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to load earn ing vehicles. In a particular form the present disclosure relates to remote detection of damage to load carrying vehicles. BACKGROUND ART

[0002] Large vehicles are commonly used to transport a payload in an open carrying container of the vehicle, typically referred to as a bed or tray. As an example, in mining operations, mining shovels and excavators load earthen material into the tray of a haul truck for transportation to a processing location or stockpile. Such haul trucks repeatedly perform such hauling activities between the loading location and the processing or unloading location, and due to the nature of the loading and loads carried, may occasionally incur damage to the truck tray. Haul trucks trays are large and expensive underlying equipment, and experience a harsh and often abrasive environment. In some cases, the truck trays are protected from damage from loads by situating w ear parts atop the truck tray to protect the underlying equipment from the environment. Further, due to the nature of the loads carried by such haul trucks, the truck tray and / or wear parts become damaged by the operation in the harsh environment and left unchecked, the size and extent of damage will typically increase over time. Hus m turn increases the cost and time taken to repair the damage or replace the wear part. Many mining operations are 24 / 7 operations, and thus there is considerable desire to keep the vehicle in a workable condition to maximize the efficiency of the mining operation. Scheduled maintenance requires taking the vehicle out of production operations for the duration of the inspection and maintenance. This represents wasted time if no damage is detected or if too much damage has occurred and / or there is damage to the underlying equipment. Due to the size, structure, and / or operation of the underlying equipment, manual inspections may not be performed safely or frequently. SUMMARY

[0003] According to a first aspect, there is provided a method for detecting damage in a load carrying container of a load carrying vehicle, the method including: identifying a container region of interest (ROI) in one or more images wherein the container ROI is estimated to contain an empty- load carrying container of a vehicle; segmenting, for each of the one or more images, the respective image into a plurality of overlapping sub-images, where each sub-image includes at least two overlapping portions where each overlapping portion including a plurality of contiguous pixels shared with an adjacent sub-image; processing each subimage with a trained damage region object detection model, wherein the trained damage region object detection model is trained to identify damage to the load carrying container in one or more damage regions in the sub-image, each damage region including a plurality of pixels associated with damage within a boundary; and determining, using the output of the trained damage region object detection model, the presence of at least one damaged region in the load carrying container of the vehicle in one or more images.

[0004] According to a second aspect, there is provided a computational apparatus configured to detect damage in a load carrying container of a load carrying vehicle including: at least one memory; at least one processor configured to: identify a container region of interest (ROI) in one or more images wherein the container ROI is estimated to contain an empty load carrying container of a vehicle; segment, for each of the one or more images, the respective image into a plurality of overlapping sub-images, where each sub-image includes at least two overlapping portions where each overlapping portion including a plurality of contiguous pixels shared with an adjacent sub-image; process each sub-image with a trained damage region object detection model, wherein the trained damage region object detection model is trained to identify damage to the load carrying container in one or more damage regions in the sub-image, each damage region including a plurality of pixels associated with damage within a boundary; and determine, using the output of the trained damage region object detection model, the presence of at least one damaged region in the load carrying container of the vehicle in one or more images.

[0005] According to a third aspect, there is provided a system configured to detect damage in a load carrying container of a load carrying vehicle including: at least one image data capture system including at least one image sensor; a computational apparatus including: at least one memory; at least one processor configured to: identify a container region of interest (ROI) in one or more images wherein the container ROI is estimated to contain an empty load carrying container of a vehicle; segment, for each of the one or more images, the respective image into a plurality of overlapping sub-images, where each sub-image includes at least two overlapping portions where each overlapping portion including a plurality of contiguous pixels shared with an adjacent sub-image; process each sub-image with a trained damage region object detection model, wherein the trained damage region object detection model is trained to identify damage to the load carrying container in one or more damage regions in the sub-image, each damage region including a plurality of pixels associated with damage within a boundary; and determine, using the output of the trained damage region object detection model, the presence of at least one damaged region in the load carrying container of the vehicle in one or more images.

[0006] According to a fourth aspect, there is provided a computational apparatus configured to detect damage in a load carrying container of a load carrying vehicle including: at least one memory: at least one processor configured to implement the method for detecting damage in a load carrying container of a load carrying vehicle of the first aspect, and any examples as described herein.

[0007] According to a fifth aspect, there is provided a system configured to detect damage in a load carrying container of a load carrying vehicle including: at least one image data capture system including at least one image sensor; a computational apparatus configured to implement the method for detecting damage in a load carrying container of a load carrying vehicle of the first aspect, and any examples as described herein.

[0008] In some examples identifying a container region of interest (ROI) in one or more images includes obtaining or accessing one or more images generated or captured by one or more image sensors on a worksite The one or more images sensors are configured to generate the one or more images from a captured image data set. The image sensors may be part of an image data capture system or a vehicle monitoring system. The image sensors may be a camera (e.g. a single optical image sensor with associated optical assembly), a stereo camera (e.g. two optical image sensors and associated optical assemblies), a LIDAR apparatus or a RADAR apparatus configured to generate a point cloud representation of a field of view from an image data set captured by the one or more image sensors at the worksite. The captured image dataset may include two successive or simultaneous images captured by image sensors such as a stereo camera or multiple cameras in different locations, or the captured image data set may be captured by a ranging based image sensor, such as a LIDAR or RADAR apparatus, configured to scan the field of view and capture range data at a plurality of points in the field of view. The image sensors may be mounted and orientated with a downward field of view configured to capture images of a base surface of the load carrying container as it passes the image sensor, or to capture image data which can be used to generate an image of the base surface of the load carry ing container as it passes the image sensor. In some examples, the method further includes generating and sending an electronic damage alert based on the determination of the presence of at least one damaged region in the load carrying container of the vehicle in one or more images. In some examples, the electronic damage alert is an electronic damage report further including a location in the load carrying container of the vehicle for each of the one or more damaged regions. In some examples, the method includes determining a vehicle identifier of the vehicle. In some examples, the vehicle identifier is included in the electronic damage alert, or the vehicle identifier is used to determine one or more recipients of the electronic damage alert. The vehicle identifier may be determined by interfacing with a fleet management system, processing an image to detect a vehicle identifier located on the vehicle, and / or by determining if the image of the vehicle matches a vehicle in a database of previously identified vehicles.

[0009] Other aspects and features will become apparent to those ordinarily skilled in the art upon review of the following description of specific disclosed examples in conjunction with the accompanying figures. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Examples of the present disclosure will be discussed with reference to the accompanying drawings wherein:

[0011] Figure 1A is a perspective view of an image data capture system of a damage detection system configured to detect damage in a load carrying container of a load carrying vehicle 106 according to an example;

[0012] Figure IB is a perspective view of a load carrying vehicle passing an image sensor mounted on a post in a system configured to detect damage in a load carrying container of the load carrying vehicle according to another example;

[0013] Figure IC is a perspective view of a load carrying vehicle passing under a pair of image sensor each mounted on a respective post in a system configured to detect damage in a load carrying container of the load carrying vehicle according to another example;

[0014] Figure ID is a perspective view of an example of a camera incorporating an image sensor which may be used in an image data capture system, such as the examples illustrated in Figures 1A, IB or IC;

[0015] Figure IE is a perspective view of an example of a stereo camera which may be used in an image capture system and incorporates two spaced apart image sensors which may be used to obtain a Three Dimensional point cloud representation of a field of view;

[0016] Figure 2 is a block diagram of a computational system configured to detect damage in a load carrying container of a load carrying vehicle according to an example;

[0017] Figure 3 is a flowchart of an example of a method for detecting damage in a load carrying container of a load carrying vehicle, which may be implemented in the computational system illustrated in Figure 2;

[0018] Figure 4A is an image of a load carrying vehicle illustrating the identified truck boundary and the identified container region of interest (RO1) according to an example;

[0019] Figure 4B is an enlarged portion of Figure 4A showing a damaged region delineated by a boundary according to an example;

[0020] Figure 4C is a heat map representation of the container ROI of Figure 4A showing a heat map representation of an identified damage region according to an example;

[0021] Figure 4D is a representation of a container ROI image and a grid of overlapping sub-images in which each overlapping portion includes one half of the sub-image in the adjacent row or adjacent column according to an example;

[0022] Figure 4E is a representation of a container ROI image and a grid of overlapping sub-images in which each overlapping portion includes one quarter of the sub-image in the adjacent row or adjacent column according to another example;

[0023] Figure 4F is a representation of a sub-image illustrating edges of wear plates and a damage region according to an example;

[0024] Figure 5A is a block diagram of a machine learning training process for training a machine learning model according to an example;

[0025] Figure 5B is an architectural diagram of a deep neural network with one flatten layer, multiple dense layers and an output layer according to an example;

[0026] Figure 5C is a flow chart of a method for parallel identification of a load carrying container ROI and one or more damage regions in the load carrying container of a load carrying vehicle and generate a heat map representation of damage to the load canying container of a load carrying vehicle according to an example;

[0027] Figures 6A, 6C and 6E are three images of the same damaged load carrying vehicle captured at three different time points, and Figures 6B, 6D, and 6F are masked images of Figures 6A, 6C, and 6E wherein each mask indicates the boundary of the vehicle in the respective image according to an example;

[0028] Figure 7A is an illustration of a software dashboard of a computational system for detecting damage in a load carrying container of a load canying vehicle displaying an electronic report including three images of the same vehicle and an estimate indicating that the load cany ing container of the vehicle is not damaged according to an example;

[0029] Figure 7B is another illustration of a software dashboard of a computational system for detecting damage in a load carrying container of a load canying vehicle displaying an electronic report including three images of the same vehicle and an estimate of the presence of the damage to the load carrying container of the vehicle including an estimate of the likelihood of the damage, the damage location and damage size, a heat map representation of the damage, and a close up of the damage region and estimated damage region bounding box according to an example; and

[0030] Figure 7C is a representation of an electronic damage report of a computational system for detecting damage in a load carrying container of a load carry ing vehicle showing a set of 4 images showing close up images of estimated damage regions and damage region bounding boxes, with the associated vehicle identifier, damage location, damage size and likelihood of damage according to an example.

[0031] In the foilowing description, like reference characters designate like or corresponding parts throughout the figures. DETAILED DESCRIPTION

[0032] Figure 1A is a perspective view of an example of an image data capture system 100 of a damage detection system 200 configured to detect damage in a load carrying container 104 of a load carrying vehicle 106. The image data capture system 100 is mounted and configured to generate one or more images by one or more image sensors 110 on a worksite 126. Tire one or more images sensors 110 are configured to generate the one or more images from a captured image data set. The one or more images may include a load carrying vehicle 106 as it passes through the field of view 112 of the one or more image sensors 110. In this example, the image sensor 110 is mounted on a cross truss of a truss 108 at an image data capture location 126. The vehicle 106 is used to transport payloads at a worksite such as a mine or quarry and passes through one or more image data capture locations 126. In this example, the load carrying container 104 is empty such that the base surface 102 of the load carrying container 104 is visible from locations above the vehicle and thus is in the field of view 112 of the image sensor 110 located on the truss 108. The field of view 112 may include the entire base surface including the side walls. The image data capture system 100 may be used to capture one or more images of one or more vehicles at one or more capture locations 126 for computational analysis by the damage detection system.

[0033] The image data capture system 100 includes one or more image sensors 110. The one or more image sensors are configured to capture an image data set which is processed to generate one or more images. The image sensors may be a camera (e.g., a single optical image sensor with associated optical assembly), a stereo camera (e.g., two optical image sensors and associated optical assemblies), a LIDAR apparatus or a RADAR apparatus configured to generate a point cloud representation of a field of view. The image data may include pixel data in one or more channels or frequency bands, ranging data (e.g., for LIDAR / RADAR image sensors) over a field of view, and metadata. The metadata may include sensor configuration, optical assembly, capture time, capture location, and processing related information. In some examples the images generated from the image data may be a single charnel image such as a gray-scale image, or multi-channel image such as RGB image. In some examples the images may be an image generated by an image sensor including an array of pixels, or the image may be a pseudo image generated by combining images from multiple cameras or optical image sensors, or by processing a point cloud representation of the field of view to generate a two dimensional image. In this example the image sensor 110 is incorporated in a camera 130 mounted on the truss 108 and disposed to successively capture images of vehicles traversing a field of view 112 of the camera 130. In some examples, the camera 130 is configured to capture images of vehicles passing through the field of view 112 of image sensor. Accordingly in some examples, the camera is configured to record images as a video stream which are analyzed to identify one or more frames containing a vehicle which are then flagged for additional processing. In some configurations, the images that are flagged are those of the entire vehicle or complete load container within the field of view. In some examples, the camera 130 is configured to capture images at a predefined rate. The predefined rate may be determined based on the likely speed of vehicles passing under the truss and the physical size of the field of view on the ground to ensure one or more images are captured as the vehicle passes through the field of view. For example, the capture rate may be in the range of 0.25Hz - 10Hz. although a value outside of this range may be selected based on the specific application and operational constraints (e.g., speed of vehicles, available processing power, etc.). As noted above a video stream may be captured at a frame rate such as 24Hz, and then separated into individual images. In some embodiment images may be sampled from the video stream at a predetermined rate, such as every second image (e.g., an effective rate of 12Hz) or every fourth image e.g., an effective rate of 6Hz). Each captured image is analyzed to identify if a vehicle is present, in which case the image is flagged for additional processing. In some examples, the camera 130 is under the control of another computing apparatus and is configured to take one or more images on receipt of a capture signal. The one or more images may be captured as a video stream or as a burst of images at predefined rate which are analyzed to identify frames or individual images containing a vehicle. A sensor may be used to detect the proximity of a vehicle to the truss and used to trigger capture of one or more images. Alternatively, a tracking system, for example as part of a fleet management system may identify when a vehicle is within a predefined proximity range of the camera and alert the local computing apparatus or the camera directly of the proximity of a vehicle to the camera. For example, a global navigation receiver may be fitted to each vehicle and continuously or periodically report the location of the vehicle to the fleet management system. Tire fleet management system may store the location of cameras 130 and determine proximity.

[0034] In some examples, the image sensor 110 may include a stereo camera, a LIDAR apparatus, a RADAR apparatus, or similar apparatus configured to capture an image data set which is used to generate a point cloud representation of a field of view, from which a two dimensional image or two dimensional pseudo image may be generated. The point cloud representation may be generated by combining two successive or simultaneously captured images obtained from a stereo camera or from two or more image sensors each in a different location, or it may be generated by a ranging based image sensor configured to scan the field of view and capture range data at a plurality of points in the field of view. The ranging based image sensor may be appropriately configured apparatus such as LIDAR or RADAR which scans the field of view and transmits and receives signals to capture range data across the field which is processed to generate a three dimensional (3D) point cloud representation. In some examples, the field of view 112 could be the sampling size of a LIDAR or RADAR.

[0035] In this example, the image sensor 110 is mounted above the vehicle and the field of view 112 is oriented downward to capture images of the base surface 102 of the load carrying container 104 as it passes the capture location. When the load carrying container 104 is carrying a payload, the base surface 102 will be partially or completely obscured by the payload. Junction box 118 is mounted at a truss upright member 120 and may be used to provide power, signal, a secondary sensor, and control cables to the image sensor 110. The junction box 118 may include or house a computing apparatus and may be connected to a network over a wired or wireless link.

[0036] In this example, the load carrying vehicle 106 is a mine haul truck used to transport earthen material around a worksite 126, such as from a mining location to a processing location, and the load carrying container 104 is the truck tray. In other examples, the load carrying container 104 may be associated with another type of load carrying vehicle such as a railcar, a barge, or other seaborne transport container. Alternatively, the load cany ing container 104 may be a trolley or skip, such as used in a quarry or underground mining operations. In some examples, the vehicle 106 may be an automated driverless vehicle. For example, self-navigating vehicles may be used in some sites and vehicle 106 would thus automatically navigate through the truss 108.

[0037] In the illustrated example of Figure 1 A, illuminators 114 and 116 are located on the truss and are directed downwardly to illuminate the field of view 112, but a single illuminator may be used in other configurations. Junction box 118 may also provide power to the illuminators 114 and 116 and may be used to control the operation of the illuminators. Hie illuminators 114 and 116 may be implemented using ruggedized light emitting diode based light sources to provide visible light. In some examples, the light sources could be single frequency, use a filter, or be configured to emit a specific band of wavelengths, and the image sensor 110 may be configured to detect the specific wavelengths or bands (including through the use of filters). The illuminators 114 and 116 may be used to constantly illuminate the field of view 112 (e.g., 24 hours per day), or their use may be limited to low light conditions such as at night, in which case the operation may be controlled by a timer circuit or a light sensor. In some examples, the illuminators 114 and 116 are switched on as required for capturing images. For example, based on a proximity sensor detecting a nearby vehicle or on command from a tracking system which identifies a vehicle is within a predefined proximity range of the camera. In some examples, no illuminators may be used on the truss and images may be captured under natural lighting conditions. In some examples, one or more illuminators may be located on the truck, such as over the truck cabin to illuminate the truck tray (regardless of whether illuminators are present on the truss or not).

[0038] In other examples, the image sensor 110 may be located in other locations and multiple image sensors may be used to capture images. Figure IB illustrates another example of an image data capture system 150 in which the image sensor 152 is a camera which is mounted on a column truss 154 with a downward looking field of view 156 to capture an image dataset from which an image can be generated of the load carrying container (truck tray) of the load carrying vehicle (load haul truck) 106. In alternative examples, the image sensor 152 may be a ranging based image sensor such as a LIDAR apparatus or RADAR apparatus located on the column truss 154 to scan the field of view 156 below the mounting location on the truss though which a vehicle may pass to capture image data set which may be processed to generate a 3D point cloud representation 170 of a field of view 156. The 3D point cloud representation represents each data point as an (x, y, z) data point in a coordinate system with a predefined origin (or reference axes) 172, for example located on the ground surface at the left hand rear side of the vehicle as shown in Figure IC. Each data point represents a point on a surface of an object in the field of view. The 3D point cloud representation may be used to generate a 2D image looking downward from a predefined reference point of view, such as a reference point of view above the vehicle. In some examples the image sensor 110 may be mounted on an unmanned aerial vehicle (UAV) which may be navigated to a location above the vehicle. In some examples the image sensor is a camera which is mounted to the UAV using a stabilized mount such as a gimbal mount to allow control of the field of view independent of the motion of the UAV. In some examples the UAV may be tethered by a cable used to provide power, send control signals, and receive images. The cable may be connected to a junction box.

[0039] Figure IC is a perspective view of another example similar to that illustrated in Figure IB, in which a first sensor 152, which in this example is a first camera, is mounted on a first column truss 154 and a second image sensor 162, which in this example is a second camera, is mounted on a second column truss 164. Illuminators may be optionally located on the trusses 154 164. The first column truss is spaced apart from the second column truss and the first image sensor 152 (first camera) is configured or orientated to capture a first field of view 156 and the second image sensor 162 (second camera) is configured or orientated to capture a second field of view 166 such that the first field of view 156 and second field of view 166 have an overlapping region 168 through which a vehicle may pass. For example, tire two comer trusses may be located on either side of a road. Each camera may be connected to a junction box 118, for example by one or more cables (although a wireless link may be used), and the junction box 118 may be configured to send a trigger signal to each camera to obtain simultaneous respective first and second 2D images from the two different perspective viewpoints. The two images may be combined to generate a 3D point cloud representation 170 of the overlapping or common field of view 168. In some examples the two images may be mosaiced to generate a 2D pseudo image of the complete vehicle. In some examples the first image sensor 152 and second image sensor 162 are both ranging image sensors, such as LIDAR or RADAR apparatus, and the range data from the ranging image sensors is combined to generate the point cloud representation 170 of the combined fields of view 156 158, or of the overlapping or common field of view 168.

[0040] The 3D point cloud representation 170 can be analyzed to identify surfaces or visual features in the field of view. In some examples, the container of the vehicle may be identified by using points with heights above a minimum height 174, or be located within an expected height range for the container of the vehicle. In these examples the range data may be converted to a height with respect to a reference surface such as the ground. The 3D point cloud representation may also be used to generate a two dimensional image from a predefined reference point of view, such as above the vehicle, where each pixel represents a position in a plane (e.g., the (x.y) plane). Each pixel value is based on the image data set associated with the point cloud data point. In some examples, the image data is the time of arrival indicating height (range) of the reflecting object, and the height values may be mapped to a color or intensity range. In some examples, the image data (or an image data point in the image data set) may be an intensity in one or more color channels (e.g., RGB intensity), or it may be a pseudo intensity based on a return signal indicative of the reflective or absorption properties of the surface with respect to the transmitted radiation. A height may also be associated with each pixel if the pixel value is not based on height value. The point cloud representation, or a 2D image generated from the point cloud representation, may be used to determine the presence of a vehicle within the field of view, as well as the vehicle features such as outline of the vehicle and / or the vehicle container. For example, a truck tray surface may have height values in a narrow range indicating a flat surface, an inclined or sloped surface, or a gently curved surface and be surrounded be height values corresponding to the top of the walls, and the truck tray surface height values would be higher than the surrounding background surface (e.g., the road). These vehicle features may also be used to aid in vehicle identification.

[0041] The height of the image sensors and choice of support structure (e.g., truss geometiy) may be selected based on the size of the vehicles and speed at which vehicles may pass the capture point. Larger heights may be used for faster, taller, or longer trucks to increase the physical size of the field of view 112 to ensure the complete load carrying container or the entire vehicle is captured in a single image. In some examples the image sensors are located at heights in a range between 12m and 22m, although heights outside of this range may be used based on the specific implementation requirements. The aspect ratio of the image obtained by image sensors may be commonly used aspect ratio such as 1:1, 3:2, 4:3, 13:9, 11:8, 16:9, any reverse (or inverse) of these, or a custom ratio may be used. Similarly, the resolution of the images generated from the image sensor may be a commonly used pixel resolution such as 1080x1080, 1080x1920, 1920x1080, 1560x1440, 2048x1536, 2048x2048 pixels, 2840x2840 pixels, 3840x2160, 4096x2160, 4096x3072, or greater, any reverse (or inverse) orientation, or another custom resolution. The height and resolution of the image sensor (or camera) may be selected to increase the likelihood that the entire vehicle is contained in a captured image, and may be selected based on one or more factors. For example, one factor may be the likely speed of the vehicle as it passes the capture location. Another example of a factor may be whether a proximity sensor is used to at the capture location to trigger capturing of an image, and if so, the type and accuracy of the proximity sensor, and whether multiple images are to be captured of a vehicle as it passes the capture location. The height and resolution of the image sensor (or camera) may be selected such that a predefined vehicle occupies at least a predefined percentage of pixels in the captured image in one or more images. For example, the predefined percentage may be set to 20% or 40% to increase the likelihood that at least one image is captured that includes the entire vehicle, or to increase the likelihood that multiple images may be captured each containing the entire vehicle.

[0042] Figure ID is a perspective view of an example of a camera 130 incorporating an image sensor 110 shown in Figure 1A. The camera 130 includes a lens assembly 132 and an image sensor 110 such as a CCD or CMOS located in ruggedized housing 134 which is configured to capture images and to store or send the captured images to another computer apparatus or storage location. In some examples, the camera may continuously capture images. In some examples, the camera 130 is under the control of another computing apparatus and is configured to take one or more images on receipt of a capture signal. In some examples, the image sensor may be sensitive to visible light wavelengths, while in other examples, the image sensor may be configured to be sensitive to thermal wavelengths or other wavelengths ranges or bands outside of the visible spectrum. The various components of the truck may have different thermal properties compared to background objects (e.g., the ground) and thus an image sensor sensitive to infrared or thermal wavelengths may be used to aid in identifying the truck against the background. Additionally, the image sensor or optical assembly may include filters specific to a particular wavelength or a band of wavelengths. In other examples, the camera 130 is a LIDAR or RADAR apparatus configured to generate one or more images from an image data set captured by the LIDAR or RADAR apparatus.

[0043] Figure IE is a perspective view of an example of a stereo camera 131 which may be used as the sensor in the example shown in Figure IF. The stereo camera 131 includes a first optical assembly 132 associated with a first image sensor 110a and a second optical assembly 133 associated with a second image sensor 110a, where the first image sensor 110a and second image sensor 110b are offset or spaced apart a lateral distance D to capture respective first and second 2D images from different perspective viewpoints. The first and second 2D images can be combined using a stereoscopic image processing method to generate a 3D point cloud or disparity representation of a three dimensional volume 170 formed by the overlapping or common field of view 112. Stereoscopic image processing methods use the known separation, orientation, and focal points of the two image sensors 110a, 110b to process the pixels in the two images to identify which pixel (or pixels) in the second image that corresponds to a pixel (or pixels) in the first image (i .e., solve the correspondence problem for all pixels). The first and second image sensors 110a 110b, and the first and second optical assemblies 132 133 of stereo camera 131 may be mounted in a ruggedized housing 134. The first and second image sensors 110a 110b may be full color sensors (e.g., capture 3 channel (RGB) signals at each pixel), monochrome, or configured to capture a single wavelength or a band of wavelengths (or frequencies).

[0044] Figure 2 is a block diagram of an example of the damage detection system 200 including an image data capture system 100, a remote computing apparatus 230, and a user computing apparatus 250 configured to perform detection of damage in a load carrying container of a load carrying vehicle. Tire computational system 200 may be used for implementing a method 300 for detecting damage in a load carrying container of a load carrying vehicle as is shown in Figure 3. In the illustrated example, the computational system 200 includes an image data capture system 100, including an image data capture apparatus such as a camera 130 including an image sensor 110, a junction box 118, illuminators 114 and 116, and RFID reader 122, but in other examples other configurations including a different image capture system or image capture apparatus could be used. For example the image data capture apparatus may be a LIDAR or RADAR in which image sensor 110 is a ranging based image sensor. Junction box 118 distributes operating power to the image sensor 110 via a power conductor 202, and also provides power and control for the illuminators 114 and 116, and the RFID reader 122. In this example, the image data capture apparatus is a camera 130 including one or more image sensors 110, one or more processors 204 in communication with one or more memories 206, and an input / output (I / O) interface 138, all mounted within the housing 134 (shown in Figure ID, IE). In other examples the image capture apparatus is a LIDAR or RADAR apparatus in which imaging sensor 110 is a ranging based imaging sensor and also includes one or more processors 204 in communication with one or more memories 206, and an input / output (I / O) interface 138 within a housing 134. Memory 206 provides storage for instructions for directing the processor 204 to capture the successive images and also provides storage for the captured image data set and the processor and memory may be combined in a microcontroller. The processor(s) 204 may be embedded processor for controlling the operation of the camera. The memory 206 may also store instructions for configuring the processor to process the captured images, and / or to transfer stored images or results to a central storage location such as a mass storage unit 240, including to a network or cloud storage location, virtual, or to another computing apparatus via a wired or wireless connection. In some examples the processor(s) 204 are configured to perform analysis of captured images, for example to identify the presence of a vehicle in an image, and / or identify a container ROI.

[0045] The I / O interface 208 is in communication with the processor 204 and implements a sensor interface 210 that includes inputs 212 for receiving image data from the sensor 110, and any additional sensors present. The I / O interface 208 further includes a communications interface 214, which may be a wired interface such as an Ethernet interface or a wireless communications interface such as a Bluetooth, Wi-Fi, cellular or other wireless interface. In this example, the communications interface 214 has a port 216, which is connected via a data cable 218 routed back to the junction box 118, although in other configurations a wireless connection could be used. Junction box 118 may include a modem, router, or other network equipment that facilitates a data connection to a network 220. Network 220 may be a local area network (LAN) implemented for local data communications within worksite 126. Alternatively, junction box 118 may route signals on the data cable 218 to a wide area network such as the Internet. In some examples, where there is no wired connection available to the network 220 the junction box 118 may include a wireless communications interface for wireless connection to the network 220 such via a cellular, satellite, WiMAX, or similar transceiver (or terminal) implementing a wireless connection protocol.

[0046] In the example shown in Figure 2, the computational system 200 also includes a remote processor apparatus 230, which includes one or more processors 232, and one or more memories 238, in communication with an input / output (I / O) interface 234. The I / O interface 234 implements a communications interface 236 for transmitting and receiving data over the network 220, such as for receiving images from one or more image data capture systems 100 and for sending electronic damage alerts and damage reports. In some examples an electronic report may be sent to a user computing apparatus 250, such as a user computer apparatus 250 of a fleet management system or in the respective vehicle to alert the driver of the damage. Hie one or more processors 232 are in communication with one or more memories 238 for storing data and instruction codes. In this example processor 232 is also in communication with a mass storage unit 240 for storing image data and for storing analysis results. In the example shown the remote processor circuit 230 further provides processing via a graphics processing unit (GPU) 242, which may be used to provide additional processing power for image processing intensive tasks. Processor 232 may thus be configured as a GPU or the remote processor circuit 230 may further include a GPU co-processor for offloading some processing tasks from the processor.

[0047] In examples where the network 220 is a local area network, the remote processor apparatus 230 may be disposed at an operations center associated with the worksite 126. In other examples where the network 220 is a wide area network the remote processor apparatus 230 may be located at a remote processing center set up to process images for multiple worksites. Alternatively, the remote processor apparatus 230 may be provided as on-demand cloud computing platform, made available by vendors such as Amazon Web Services (AWS) or Microsoft Azure. In some examples, the computational system may be a distributed system in which multiple remote processor apparatus 230 are located within, or adjacent to, each junction box 118 at the image data capture location to perform local processing of captured images which may then be transferred to a central server, cloud server or computational system via network 220. In other configurations, the remote processor apparatus 230 is not included.

[0048] The system 200 may further include a user computing apparatus 250 including one or more processor 252, one or more memories 254, and an I / O interface 256. The I / O interface 256 implements a communications interface 258, which is able to receive data (e.g., alerts and / or reports) via network 220. The I / O interface 256 also includes interface 260 for displaying a user interface on a display device 264 or sending an electronic report 262 to a third party. The user computer apparatus 250 may be located at the operations center of worksite 126, where electronic reports of damage assessment can be displayed or reviewed by operators, in another working location. The user computer apparatus could be a mobile computing apparatus allowing a user to view or access a damage report in any location with network access. The user computer apparatus could also be a computing apparatus in the vehicle allowing a driver to receive an electronic report including an alert of damage to the vehicle. In self-navigating or other driverless vehicles such as a railway load carrying container 104, the remote computer apparatus may send an electronic report, which may be an alert signal, which is processed by the user computing apparatus onboard the vehicle to cause the vehicle to be diverted or flagged so that further action can be taken (e.g., repair).

[0049] Whilst the example of the system 200 shown in Figure 2 includes a processor 204 within the camera 130 and a separate remote processor apparatus 230, in other examples these could be combined into a single computer apparatus, for example the processor 204 in camera 130, maybe configured to perform capture and analysis of images (e.g. processor 204 also performs the functions of processor 232), or an image sensor could be wired to, and directly controlled by, the remote processor apparatus 230 (e.g. processor 204 may be omitted), or processor 204 may be a slave processor acting as an embedded processor under the control of processor 232 which perfonns processing of images captured by image sensor 110. Functions described below as being performed by the remote processor apparatus 230 may thus be performed by the embedded processor 204 or any other combination of central processor apparatus. That is computational system 200 may also be implemented as a cloud based systems in which images are sent directly to cloud storage, and cloud servers processor the stored images processing and generate electronic reports which may be received or viewed on local user apparatus including mobile phones, tablets, laptops and desktops, using a local an app, browser, or local software interface which is configured to receive and exchange data with the cloud based servers. In some examples an additional capture computing apparatus may be located injunction box 118. The capture computing apparatus may be configured to interface with the camera 130 or an image sensor 110 and to store the images captured by the image sensor 110, control operation of the camera 130, perform additional processing of the images, and / or transfer the images or any processed results to the remote processor apparatus 230 which may store and perform additional processing of the captured images. In some examples the I / O interface 208 in the camera may be configured to wirelessly transfer the images to the capture computing apparatus or to the remote computing apparatus 250 for storage and additional processing.

[0050] It will also be understood that the various computing apparatus, include the image sensor 110, remote processor apparatus 230 and user computer apparatus 250 may include one or more processors and one or more memories, and one or more input and output devices, such as mouse and keyboard for receiving user input and a display screen for displaying images or a user interface. The one or more processors may include a central processing unit (CPU) including an Arithmetic and Logic Unit (ALU) and a Control Unit and Program Counter element, and CPU may be a single CPU (core) or multiple CPU’s (multiple core). The one or more processors may include one or more CPUs. The various computing apparatus may also include one or more graphical processing units (GPUs). The one or more processors may be multi-core processors, parallel processors, vector processors, and may be part of a distributed or cloud computing apparatus. The one or more memories may be operatively coupled to one or more processors and may include RAM and ROM components, and secondary or non-volatile storage components such as solid-state disks, hard disks, and / or USB storage devices which may be provided within or external to the computing apparatus. An external storage device may be a directly connected device, or a network storage device (i.e., connected over a network interface) external to the computing apparatus, including cloud-based storage devices. The one or more memories may include instructions to cause the one or more processors to execute an example of a method described herein. The one or more memories may be used to store the operating system, any additional software modules, and instructions for implementing various software blocks for implementing examples of the method described herein to detect the presence of damage in a load carrying container of a load carrying vehicle. The one or more processors may be configured to load and execute the software modules and instructions stored in the one or more memories.

[0051] Figure 3 is a flowchart of an example of method 300 for detecting damage in a load carrying container of a load carrying vehicle. Examples of the method 300 may be computationally implemented in the computational apparatus illustrated in Figure 2 to provide the damage detection system 200. To assist in understanding the system 200, the method for detecting damage in a load carrying container of a load carrying vehicle will first be outlined with reference to Figure 3, followed by a discussion of the how the method may be varied and implemented in various examples of the system 200.

[0052] Block 302 broadly includes identifying a container region of interest (ROI) in one or more images. The container ROI is defined by a boundary and is estimated / determined / identified to contain an empty load carry ing container 104 of a vehicle 106. For example, with reference to Figure 1A the container ROI is the empty truck tray 104 of truck 106 which can be determined by identifying the base surface 102 in the image of the truck tray 104. The boundary of the container ROI may be a boundary containing the empty truck tray and adjacent portions of the truck, or it may be boundary containing the entire truck which includes the empty truck tray. The ROI may be identified in a single image which captures the entire empty truck tray. In some examples the ROI may be identified in a pseudo 2D image which captures the entire empty truck tray. The pseudo 2D image generated from mosaicking or combining multiple images captured by cameras in different locations such as illustrated in Figure 2B, each of which capture part of the truck tray. In some examples the pseudo 2D image may be generated from a point cloud representation.

[0053] Block 304 broadly includes segmenting, for each of the one or more images, the respective image into a plurality of overlapping sub-images. Each sub-image includes at least two overlapping portions where each overlapping portion includes a plurality of contiguous pixels shared with an adjacent sub-image. In some examples the sub-images generated by the segmentation are square images or rectangular images with a pre-defined sized. In some examples the sub-images may be regular polygons such as hexagons, octagons, diamonds, or a pixelated approximation of a circle or an ellipse. In some examples the sub-images could be an irregular shape. In some examples the size of the sub-image may be determined based on the size of a physical feature of the vehicle or a part such as wear plate. In some examples the shape of the sub-image may be constant for all sub-images. In some examples the shape of the sub-image may van across the image. In some examples the size of the overlapping portion may be constant for all sub-images. In some examples the size of the overlapping portion may vary. The size of the overlapping portion may be determined based on the size of expected damage regions, or a distribution of the size of damage regions. Block 306 broadly includes processing each sub-image with a trained damage region object detection model. The trained damage region object detection model is trained to identify damage to the load carrying container in one or more damage regions in the sub-image, as damage may occur in multiple locations within the same sub-image. Each damage region includes a plurality of pixels associated with damage (e.g., estimated / classified / identified to represent damage to the load carrying container) within a boundary. The boundary may be a rectangular bounding box or another shape including regular polygons and irregular shapes. The boundary- may- outline only the damage itself.

[0054] Block 308 broadly includes determining, using the output of the trained damage region object detection model, the presence of at least one damaged region in the load carrying container of the vehicle in the one or more images. In some examples, the determination may be a binary determination of the presence or absence of damage, e.g., a damage alert, which may be determined by comparing an estimated likelihood of damage against a threshold. The determination may be made by combining the output of the individual sub-images to make a determination of damage to the load carrying container as a whole, or to refine estimates of the damage region, for example where the damage is present in two or more sub-images. The detennination may also be performed by analyzing, combining, or comparing multiple sub-images taken at different time points, for example, to detect a change or build confidence in detections by the trained object detection module. In other examples, in some images mud or dust may obscure damage in the sub-images, or increase the difficulty in identifying small regions of damage, and thus using multiple images taken at different time points may be used to build robustness in detection. The determination may also include an estimation of one or more damage parameters including the location, size, shape, extent of damage (e.g., minor, major), duration (e.g., if present in images taken at multiple time points), and temporal progression of damage (e.g., changes in damage parameters overtime).

[0055] In a computational implementation the blocks in Figure 3 represent functional blocks of software code, or instructions, that may be stored in, and read from, one or more memories to direct or configure one or more processors in the computational apparatus illustrated in Figure 2 to implement the functionality illustrated in the respective block. The actual code to implement each block may be written in any suitable programming language, such as C, C++, C#, Java, Python, and / or assembly code, and stored in the memory in machine or processor readable format. One or more software libraries may be used for performing tasks such as image analysis, feature detection, numerical analysis and machine learning including SciPy, NumPy, GPL, IMSL, MATLAB, Pandas, Matplotlib, Skleam, Keras, PyTorch, TensorFlow, and Theano.

[0056] It will also be understood that the method broadly outlined in Figure 3 and discussed above may be varied in other examples as discussed herein, and may include additional features (or blocks) as well as variations. In the following discussion examples may focus on variations of a specific feature, and it is to be understood that the various examples relating to different features may be combined as required. Where alternative examples are provided either one or both alternatives may be implemented (depending upon the context), and in some cases several “alternatives” may all be implemented, for example to build robustness, add redundancy, and / or to allow a consensus decision to be made.

[0057] For example, in some examples block 302 may include accessing or obtaining one or more stored images. Obtaining one or more images may include capturing one or more images at one or more capture locations 126 (e.g., across a worksite). Captured images may undergo immediate processing, for example to identify an empty container ROI, or the images may be stored for later processing. In some examples, the images may undergo initial processing by the capture system which then stores the partially processed images. For example, the capture system 100 may process a captured image to identify the presence of a truck in a captured image and store the images in which a vehicle is identified for later processing. Accessing one or more stored images may include accessing images stored on a mass storage device or a cloud storage device which were stored or provided by an image data capture system 100. Accessing one or more images may also include accessing a third-party sy stem such as a fleet management system or a vehicle monitoring system configured to capture one or more images. Metadata relating to the image may be stored with, or associated with the images, for example in a database.

[0058] In some examples the method may include (optional) block 310 broadly including generating and sending an electronic damage alert based on the determination of the presence of at least one damaged region in the load carrying container of the vehicle in one or more images. The electronic damage alert may be an electronic damage report including at least a location in the load carrying container of the vehicle for each of the one or more damaged regions. The electronic damage report may include additional damage parameters as well as one or more images of the damage. In some examples the method may include (optional) block 312 which broadly includes determining a vehicle identifier of the vehicle in an image. The vehicle identifier may be included in the electronic damage alert (or report) or used to determine the recipient of the electronic damage alert or electronic report. Other variations and features may be implemented as described herein.

[0059] It is also to be understood that the various blocks illustrated in Figure 3 may be combined or performed in parallel, and the order illustrated in Figure 3 is illustrative and is not a strictly temporal order. When processing multiple images some blocks may be repeated and may be performed in a different order. For example, the determination step 308 could be an iterative or repeated step performed using the output from multiple images captured at different time points. For example, a first set of images may be analyzed in blocks 302-306, and an initial determination made in block 308. However, if the confidence of damage detection is low, additional images of the same vehicle may be analyzed, that is blocks 302-306 are repeated on additional images, before determination step 308 is continued making use of the additional output of repeated step 306. It will also be understood that the images of a vehicle may be collected from multiple capture locations, and the analysis as outlined in blocks 302-306 may be performed on the images collected from different capture locations, and thus the determination block 308 may be performed on the results obtained from the images from the multiple capture locations.

[0060] With reference to Figure 4, examples of the method for detecting damage in a load carrying container of a load carrying vehicle include identifying a container region of interest (ROI) in one or more images where the container ROI is estimated to contain an empty load carrying container of a vehicle. In the context of the specification, a region of interest (ROI), is defined as a set of contiguous pixels in the image surrounded by a boundary which is estimated to contain a target object. That is the boundary will not necessarily outline the target object, and other objects may also be present within the boundary. Additionally, as the boundary is an estimate (e.g., a likelihood or probability), the object is not guaranteed to be within or not wholly within the boundary. The boundary, and thus the shape of the ROI may be a rectangular bounding box, a regular polygon, a (pixelated) approximation of a circle or ellipse, a composite shape (e.g., formed from two or more shapes) or have an irregular shape (e.g., outlining the damage).

[0061] In some examples, the boundary forms part of the ROI so that the boundary pixels are part of the pixels “within” the ROI (i.e., the boundary is the line dividing adjacent pixels). This allows the boundary pixels to be part of the boundary of the target object. In other examples, the boundary pixels may be excluded from the ROI. Whether boundary pixels are part of the ROI or not will typically depend upon the specific method and associate configuration used to identify the boundary. For the sake of convenience, the use of “within a boundary” is to be interpreted broadly to include pixels forming the boundary unless they are explicitly excluded.

[0062] In the case of regular shapes, such as bounding boxes, the location may be geometrically specified using a reference location such as a center or edge, and dimensions of the edges forming the boundary such a length and width. In some examples, the location may be stored as the set of pixels contained within the ROI (as defined by the boundary). A binary image mask (or simply mask) may also be generated using the pixels within the boundary. The image mask may be an image of the same size as the original image (in which the object is detected) in which pixels within the boundary’ are set to a value of 1, and the pixels outside of the boundary are set to 0. Applying a mask image to an original image may include a bitwise AND operation between the original image and the mask image.

[0063] In some examples, the ROI may be estimated using computer vision techniques such as edge detectors, blob detectors, generalized Hough transforms, or other feature extractors to identify a boundary. The boundary of the ROI may be formed from joining multiple edges and boundaries of shapes estimated using one or more of these computer vision methods. Identification of the ROI may further include an analysis of the pixel values, locations and / or height data (if available from a point cloud representation) to estimate if the pixels within the boundary-’ match the target object. For example, summary' parameters could be estimated overall, or over part, of an image and these summary' parameters compared with expected values for the target object. The summary parameters may include central estimators (averages, medians), variances, percentile ranges, quartile ranges, gradients, etc., of pixel intensities, contrasts, textures, or other pixel or image properties. These may be based on analysis of reference (training images), and / or other information such as a 3D model of the target vehicle, texture maps, and material properties of the vehicle surface (e.g., reflectivity) combined with the capture time. For example, a metal container may be expected to be highly reflective, whereas an earthen material filled container will be less reflective. These methods may be used to build up a model of appearance of the target object in an image, and correlation and statistical modelling techniques may then be used to estimate whether a ROI contains the target object (e.g., based on an expectation generated from the model). That is the distribution of pixel data may be analyzed to determine if they match an expected distribution for an empty container.

[0064] In some examples, the ROI may be estimated using a trained machine learning (ML) object detector. Hie ML object detector is first trained to identify ROI containing the target object using a training dataset, and the trained ROI may then be applied to (new) images to estimate if a target object is within the image. ML object detection models are a class of ML models configured and trained to detect objects in images. The output of a ML object detector is a region (including a boundary) and a classification of the object within the boundary. The boundary may be a regular polygon such as a rectangular bounding box with a height, a width and a reference point (e.g., center or edge), although the boundary may be other shapes and may be an irregular boundary. The object detector may be trained to detect a single class of object or multiple classes of objects. The classification may be a binary classification of whether the specific object (the class) is present or not, or the classification may be a likelihood the object is contained within the boundary. Object detectors may detect multiple instances of objects within an image (e.g., of the same class or multiple classes). Examples of ML object detection models include region-based Convolutional Neural Networks (R-CNN) which are typically two-stage models in which the first stage identifies object regions and the second stage classifies the object in each region and refines a bounding box, and the YOLO (You-Only-Look-Once) family of object detectors which performs region identification and classification simultaneously (i.e., on the whole image or in a single pass through the data). The YOLO family includes both anchor based and anchor free methods for estimation of boundary boxes. Anchor based object detectors use a set of predefined bounding boxes of various sizes and aspect ratios over an image which are then refined to identify specific objects within an image. Anchor free object detectors identify clusters of key points from which a bounding box can be estimated. For example, the YOLOv3-v5 are anchor based object detectors whilst YOLOvl, YOLOX, and DAMO-YOLO are anchor-free object detectors.

[0065] Identifying a container ROI in an image may be performed using a computer vision based method or a trained ML object detector. The identification may include joint estimation of the ROI boundary and estimation that an empty load carrying container of a vehicle is within the boundary, or joint estimation of the ROI boundary and estimation that a load carrying container of a vehicle is within the boundary followed by a determination of whether the load carrying container is empty, or the identification may be sequential process including estimation of the ROI boundary, and evaluation of the pixels within the boundary to determine that an empty load carrying container of a vehicle is within the boundary which may include joint estimation of the load carrying container and whether it is empty or estimation of the load carrying container followed by an estimation of whether the load carrying container is empty. In some examples, a trained ML object detector may be used to identify the container ROI, and computer vision and / or statistical analysis methods may be used to analyze the pixels with the boundary to determine if the load carrying container is empty. In other examples, the ML object detector may be trained to identify if a load carrying container is empty by training on images of vehicles with empty load carrying containers, and vehicles with partially or fully fdled load carrying containers.

[0066] It is further noted that load carrying container is within (and may form part of) the vehicle boundary when viewed from above. Thus, identification of a container ROI may include identification of the entire load carrying vehicle in an image such that the boundary of the container ROI is a vehicle boundary containing the entire vehicle. In other examples, the boundary of the container ROI is a boundary containing at least a part of the vehicle that includes the load carrying container. That is the container ROI may be a vehicle ROI (where the vehicle is a load carrying vehicle). In some examples, identification of a load container ROI may use a trained ML object detector model trained to identify a load carrying vehicle. In some examples, identification of a load container ROI may use a trained ML object detector model trained to identify a load carrying container or a vehicle. In some examples, identification of a load container ROI may use a trained ML object detector model trained to identify both a load earn ing vehicle and the load carrying container of a vehicle (i.e., a multi-class object detector). Similarly, to the discussion above, the determination of whether the load carrying container is empty may be performed jointly with estimation of the vehicle boundary (e.g., a ML object detector may be trained to identify vehicles with empty load carrying containers), or a determination that a load carrying container is empty may be performed after identification of the vehicle boundary.

[0067] In some examples, such as those where a 3D point cloud representation is obtained, height values may be used to determine the boundary and / or whether the load carrying container is empty. In one example a point cloud representation of a field of view is generated using at least two successive or simultaneous images obtained from a stereo camera or a plurality of image sensors each in a different location, or generating point cloud representation of a field of view from a LIDAR apparatus or a RADAR apparatus. The presence of a vehicle in the point cloud representation is first identified (or determined) and then the load container ROI of the vehicle is obtained or determined from the point cloud representation. Determination of the boundary and / or if the container is empty may be performed using depth information obtained from the point cloud representation by identifying the presence of a base surface of a container and one or more surrounding walls extending proximally from the base surface. In other examples, the height values may be required to be within an expected height range for the container of the vehicle, or a plane or surface may be fitted and compared with an expected surface for the container, and / or side walls, of the vehicle. An expected height range or surface(s) may be obtained from a 3D model of the vehicle or from measurements of vehicles.

[0068] An example of identification of a container region of interest (ROI) is further illustrated in Figure 4A. An image 400 containing a mine haul load truck is shown. In this example, a trained ML object detector was used to identify both a truck boundary 404, separating the truck from the surrounding background area 402 of the image, and an empty truck tray (load carrying container) 406 of the truck. The boundary of the truck tray ROI is indicated by solid line and the boundary of interior of the load container ROI is indicated by a dashed line. In this example, illuminators were not used and shadows 408 are created on the truck tray surface due to the truck tray walls. A truck identifier (“1234”) 410 is also visible and located over the cabin portion of the truck. In other examples, ML object detector may be used to just identify the truck boundary or just the empty truck tray. As can be seen in Figure 4A, the truck tray 406 contains an approximately semicircular damaged region 430. Figure 4B is an enlarged portion of Figure 4A showing the damaged region 430 delineated by a rectangular boundary 432, and Figure 4C is a heat map representation of the container ROI 440 (of Figure 4A) showing a heat map representation 442 of the identified damage region 430 in the image, one or both of w hich may be included in an electronic damage report. A coordinate grid is also shown in the heat map representation by dotted lines.

[0069] As described herein, examples of methods for detecting damage are configured to segment each image into sub-images and process each sub-image to identify damage regions in one or more sub-images, which may be combined to determine the presence of one or more damage regions. This may be used to generate an electronic alert, which may include an electronic report.

[0070] Segmentation of an image into sub-images is performed such that each sub-image includes at least two overlapping portions where each overlapping portion includes a plurality of contiguous pixels shared with an adjacent sub-image. Segmenting may include creating a sub-image of a predetermined size (e.g., 512x512 pixels, and then copying contiguous pixels of the original image file into a sub image file). When the sub-image is created the pixel values may be set to a predetermined value, such as zero, or if the set of contiguous pixels copied is not of the same size as the sub-image (e.g., less than 512x512 pixels) any remaining (i.e., pixels) not copied over may be set to the predetermined pixel value such as zero. Tire image segmented may be the same image in which the container ROI is identified (i.e., the original image), or it may be an image generated from image in which the load container ROI is identified, such as by cropping the original image to the load container ROI boundary or by masking pixels outside of the load container ROI boundary. If the boundary is not a rectangular boundary, then the load container ROI image 406 may be converted to a predetermined size or shape, e.g., a rectangle, by the addition of padding pixels. Segmentation of the image into overlapping sub-images may be performed on the whole image or on a portion of an image, for example based on the truck boundary 404 or container boundary 406. In the case of segmenting the whole image, any detected damage outside of the container ROI may be ignored (discarded or masked). This allow s parallel implementation in which segmentation and container region detection are performed in parallel, with the results of the region detection used to filter out any damage regions detected outside of the container ROI. In some examples, the sub-images tile only a portion of the total image such that some portions of the image are not included in any sub-images (effectively the non-included pixels are masked). That is segmenting an image into sub-images may include segmenting only the container ROI portion of the image, segmenting a portion of the image based on the boundary of the container ROI (e.g. the portion includes the container ROI and some adjacent pixels, or a portion of the image within the boundary of the container ROI), or segmenting the entire image. In figure 4D, the image is masked down to the truck boundary 404, and which contains the container boundary 406 designated by dashed line, and is segmented into a grid of 70 overlapping sub-images 420 including 10 rows and 7 columns in which each overlapping portion 422 includes a fraction of the sub-image (e.g. a half, a quarter, or one-sixteenth, and the like). Each of the sub-images 420 shown is labelled by its (row, column) location in the image 406 incrementing from the top left hand comer. In this example, each overlapping portion 422 includes one half of the next subimage in the same row (overlapping portion 422a), one half of the sub-image in the same column but next row (overlapping portion 422b), and one quarter (overlapping portion 422c) with the sub-image in the next row and next column (the diagonally adjacent sub-image). In this example the truck boundary 406 is represented as a rectangle with rounded comers to provide a small clearance with edges of the sub-images to assist with identification of the multiple sub-images. Other configurations are also possible. In some examples the truck boundary 406 may be a rectangular box such that the top edge of the truck boundary align with the top edges of the first row of sub-images, the left edge of the truck boundary aligns with the left edges of the first column of sub-images, and the top left comer of the first sub-image (1,1) is also the top left comer of the truck boundary 406. In this example, each subsequent row overlaps with the bottom half of the previous row, and each subsequent column overlaps with right half of the previous column. Comer subimages will include three overlapping portions and for other sub-images there will be at least 4 overlapping portions. For example, with reference to Figure 4D, sub-image (1,1) has a first overlapping portion 422a defined by vertical lines with the adjacent sub-image (1,2) in the same row and next column, a second overlapping portion 422b defined by horizontal lines with the adjacent sub-image (2,1) in the same column and next row, and a third overlapping portion 422c defined by vertical and horizontal lines. Third overlapping portion 422c is the intersection of the first overlapping portion 422a and the second overlapping portion 422b, and with the adjacent sub-image (2,2) in the next column and next row (the diagonally adjacent sub-image). Thus image (1,1) includes one half of image (1,2), one half of image (2,1) and one quarter of image (2,2). Image (2,2) also shares one half with image (1,2) and one half with image (2,1). The sub-images 420 in the last column and last row extend beyond the boundary of the container ROI406, and pixels in these extension regions 424 are filled with padding values (e.g., 0).

[0071] The use of an overlap assists with detecting damage located in the edges of sub-images, and which otherwise may be spread over two adjacent images. In the example, shown in Figure 4D the overlap is 50% of the pixels, however, in other examples, the amount or percentage of overlap between adjacent images may be varied. Selection of the percentage of overlap includes a trade-off in confidence of detection vs processing time. Using a larger percentage increases confidence of detection although the trade-off is an increase in the processing time due to the additional number of sub-images to be processed. Figure 4E is similar to Figure 4D and shows segmentation of the image shown in Figure 4A cropped to the truck boundary 404 (solid line), and which contains container ROI boundary 406 (dashed line), into a grid of overlapping sub-images in which each overlapping portion includes one quarter of the sub-image (25%) in the adjacent row or column (and one sixteenth of the diagonally adjacent sub-image). This segmentation creates 35 images in 7 rows and 5 columns - a reduction of 35 images (i.e., half) compared to Figure 4D. In some examples, a heat map representation may be estimated from each damage region identified in a sub-image, and these individual heat map representations from all sub-images may be combined or merged. This can assist in identifying damage occurring near edges of sub-images as the heat maps representations from the two images are merged and rejoined. This allows the use of smaller overlaps, and thus fewer images and reduced processing time, whilst maintaining confidence that damage regions are reliably detected. The percentage of overlap may range between 10% and 50%, although values of more than 50% may also be used, and the choice may be based on the requirements of the specific implementation such as how frequently are images captured, available processing power, etc.

[0072] Each sub-image is processed with a trained ML damage region object detection model. The trained ML damage region object detection model is an object detector which lias been trained to identify damage to the load carrying container. The trained ML damage region object detection model is configured to identify any damage regions of interest in a sub-image. A damage region of interest will also be referred to as damage region or damage ROE Each damage region identified includes a plurality of pixels in the sub-image associated with damage within a boundary'. The boundary defines a set of pixels in the sub-image which the model estimates indicate the presence of damage to the load carrying container. There may be no damage regions identified (no output or an output indicating a negative detection), or there may be one damage region identified, or multiple damage regions may be identified. In some examples, the output of the ML damage region object detection model is a brnarv indication of the likely presence of a damage region in the sub-image (e.g., a binary presence / absence output) or an estimate of the likelihood (or probability, or confidence) that a damage region is present is output which can be compared with a threshold (e.g., 75%, 90%, 95%) to make a determination that damage is present. In other examples, the output of ML damage region object detection model is configured to output one or more damage regions which may be provided as a boundary defining a damage region. In some examples, the boundary is provided as a bounding box (e.g., a length and a width) with an associated location reference point (e.g., center or comer point). In some examples an estimate of the likelihood the pixels within the damage region represent damage may also be provided.

[0073] The operation of a trained damage region ML object detection model is further illustrated in Figure 4F, which is a representation of a sub-image 420 illustrating edges of wear plates 438 and a damaged region 430 of the surface of the load carrying container. In this example, during processing of the sub-image 420 by the trained damage region ML object detection model, the model identifies candidate bounding boxes 432a, 432b, 432c, 432d of different sizes, aspect ratios, and locations in order to find the best fit bounding box 434 of damage region 430. The best fit bounding box 434 is output along with the damage likelihood estimate 436 that the bounding box contains a damage region (“damage: 98.4%”). As these wear plates 438 become damaged, they may be removed and replaced, or repaired and refitted. These may be tiled over parts of the load carrying container and the edges of adjacent wear plates 438 may appear as lines in sub-images. The trained damage region ML object detector may identify section of the edges of the wear plates as a possible damage region as shown by candidate bounding box 439. However, as the likelihood estimate that the join represents a damage region is low (e.g., 0.7%) the model will either reject and not output the bounding box, or the bounding box will be reported with the associated low likelihood estimate (0.7%) which will then be rejected when compared with a predetermined threshold.

[0074] As noted above, the trained damage region ML object detection model is generated by training an object detection model on a training set of training images. These may include multiple sub-images obtained from multiple images of a load carrying container of multiple vehicles. The training images may be labelled with damage containing regions. Further sub-images include sub-images with damaged regions and subimages without damaged regions. The object detection model may be an object detection model based on a region-based Convolutional Neural Networks (R-CNN), a YOLO-based model, or another object detector model. The training process is further illustrated in Figure 5A which is a schematic diagram of a machine learning (ML) training method 500 according to an example and Figure 5B is a schematic diagram of the architecture of a neural network model which may be trained using the machine learning training method shown in Figure 5A. There are five main processes involved in the ML training method 500. First input data 502 undergoes pre-processing 504 which may include numerical procedures such as data normalization, data mapping, data augmentation, data serialization, and / or feature detection to generate a pre-processed (or normalized) dataset 506. Normalization may include scaling the data to ensure that all data is over the same range (e.g., 0 to 1, or 0 to 255). Mapping may include mapping an image to a predefined standard size such as 256x256 pixels or 512x512 pixels, and may include addition of padding pixels to fill missing pixel values. Data augmentation may include generation of artificial training data based on the input data, for example by inverting reflecting, scaling, and / or modifying the image or pixel values. Data serialization may convert an array or image into a linear vector. Feature detection may include determining one or more features of the data, including average values, variances, as well as computer vision based feature detection such as shapes. The pre-processed input dataset. A training / testing split 508 in which the pre-processed dataset 506 is then split into a training dataset 510 and a test dataset 512. In this example, the training dataset includes 70% of the pre-processed dataset 506 and the test dataset 224 includes the remaining 30%, however other training:test splits may be used, such as 80:20 split.

[0075] One or more machine learning training epochs are then performed. In each epoch, model fitting 514 is performed of an initial (or input) machine learning model 516 using the training data 510 to generate a trained machine learning model 518. The test set is then provided to the trained machine learning model 518 to evaluate 520 the performance of the trained machine learning model 518. Hypcrparamctcr adjustment 522 may then be performed to adjust the ML model and generate an adjusted ML model 516 for the next training epoch. In the next epoch, the model fitting 514 processes trains the adjusted ML model 516 on the training data 510 to generate the trained ML model 518 which is then evaluated 520. This ML training process is repeated for a number of epochs until a final trained ML model 518' is obtained. Stopping of the training process may be based on one or more of an accuracy criteria, the number of training epochs, or the improvement or difference compared to one or more previous epochs. Some ML models only require a single training epoch, m which case, hyperparameter adjustment is not performed. Typically, deep learning and neural network based models are trained over many epochs (e.g., 10’s to 100’s of epochs). As additional data becomes available, ML models may be retrained on a new or expanded dataset (i.e., including the additional data). In some examples, a third blind validation dataset may also be split from the pre-processed dataset 506 in data split 508. The training:test:validate split may be 60:30:10 or 70:20:10. The performance of the final trained ML model 518' is then evaluated 520 against this validation dataset. This is also referred to as a blind dataset as unlike the test dataset, the data is completely withheld from the training process and is only used to validate the final model 518' (e.g., determine accuracy on a completely unseen dataset).

[0076] Figure 5B is an architectural diagram 530 of a deep neural network with one flatten layer 532, multiple dense layers 534 and an output layer 536 according to an example. In this example, the neural network model architecture included one flatten layer (FL) at the beginning, multiple hidden dense layers (HDL) in the middle, and an output layer at the end. In figure 5B m indicates the number of neurons in the Mi HDL. In each neuron, the input quantities are first aggregated through summation and then an activation function (fA) is applied, such as rectified linear unit (ReLU). The output layer 536 then processes the last layer of the hidden dense layers to generate an output, such as an object detection or classification. Hie object detection may be a boundary such as boundary box, a location m the image, and probability’ the object is within the boundary. Training of ML models may be performed using software libraries and packages such as Keras, TensorFlow, NumPy, Skleam python library, and Pandas. Hyperparameter adjustment is typically performed based on a loss function, an optimization function and a learning rate and may include adjusted model and training parameters such as the learning rate, the number of neurons within a dense layer, number of dense layers, activation function, number of epochs, etc.

[0077] Training encodes the model into the model architecture by adjusting and configuring the model architecture, such as by determining the appropriate number, configuration and weights between neurons to enable the model to robustly identify the target object in an image. The training process thus allows the model to learn a complex signature that can identify the target object, i.e., distinguish the target object from other features in the image. Whilst it may be difficult to identify exactly how a model makes a decision, the effectiveness of the training can be assessed using a metric such as accuracy from applying the model to a test and / or validation set.

[0078] In some examples, the ML object detection model used for detecting damage regions is an example of a DAMO-YOLO object detector (Xu et al, “DAMO-YOLO : A Report on Real-Time Object Detection Design”, arXiv:2211.15444, available at arxiv.org, the contents of which is hereby incorporated by reference). In one example, a DAMO-YOLO-S model is used where S represents the size small (other sizes include tiny (“-T), medium (“-M”) and large C -L ”). The architecture of the YOLO family includes three parts - the backbone, neck and the head. The features of the input image are extracted by the backbone. These features are passed through the neck, where aggregation of multiscale features takes place. The head uses these feature maps to output localization and classification scores.

[0079] In DAMO-YOLO, the architecture includes Maximum Entropy (MaE) Neural Architecture Search (NAS) backbones, an efficient Reparametrized Generalized FPN (RepGFPN) neck, and a Zero Head in which AlignedOTA label assignment and distillation enhancement is performed. NAS is neural architecture search which adjusts 3 parameters for each network, while maximizing test set performance above a threshold, and minimizing resource use below a threshold. It uses Pareto Optimization to do this. The 3 parameters are: Width: How many features (network weights) per layer ; Depth: How many layers for the network; and Resolution: Input resolution. The size (e.g., Small, Medium) represents different scales of backbones. The RepGFPN neck uses a generalized FPN in which multi-scale features are fused from previous and current layers. In GFPN the number of features is the same for all different scales. In RepGFPN, NAS is run on the width parameter only per scale, to determine the optimal number of features per scale. Regarding the Zero Head, many object detection networks have separate "decoupled" prediction heads in which small convolution networks are applied to region of interest (ROI) crops of the image. ZeroHead is called ZeroHead as no small convolution networks are applied to ROI crops (i.e., there are no heads), and instead the ROI is flattened, and a direct bounding box prediction is made. The default optimizer is a Stochastic Gradient Descent (SGD) optimizer which is configured to optimize two losses . The first loss is a cross entropy loss classifying the truck tray damage, and the second loss is a bounding box regression loss. In some examples the default settings of the model are used apart from the optimizer where the Adam optimizer is used instead of the default SGD optimizer, and the loss function is modified to increase the weight of the class loss by doubling it. The ZeroHead includes task projection layers for each loss.

[0080] In one example, identifying the container ROI in the image, and segmenting the image are performed in parallel. Figure 5C is a flowchart 550 of an example method for parallel identification of a load carrying container ROI and one or more damage regions in the load carrying container of a load carrying vehicle. Each of the damage regions is filtered or masked to exclude pixels outside of the container ROI, and may be used to generate a heat map representation of damage to the load carrying container of a load cariying vehicle.

[0081] In the illustrated example, the input 552 is an RGB image of a load carrying vehicle. The image is separately processed, i.e., in parallel, to identify the ROI in the image and to segment the image and identify damage regions. The first parallel processing pathway 554, which we will refer to as the vehicle identification path, is shown in the top row of flowchart 550 and includes a ROI detector module 556 and a binary mask module. The second parallel processing pathway 560, which we will refer to as the segmentation pathway is shown the bottom row of flowchart 550. The ROI detector 556 includes a trained object detection model configured to processes the image 552 to identify a container region of interest (ROI). This is a region in the image, defined by a boundary that is estimated to contain at least the vehicle container. In some examples, the ROI detector, and specifically the trained object detection model, is configured (or trained) to identify the vehicle boundary. That is the boundary of the container ROI may be the vehicle boundary (as this also contains the vehicle container). Similarly in some embodiments, the truck tray (or container) may be the only part of the vehicle visible when viewed from above, in which case the container ROI is then also the vehicle boundary. In some examples, the ROI detector 556, and specifically the trained object detection model, is configured (or trained) to identify both the vehicle boundary and the load container boundary (within the vehicle boundary). In the event that the load container boundary cannot be identified then the vehicle boundary may be used as the boundary for the container ROI. In some examples where the truck tray (or container) is the only part of the vehicle visible when viewed from above, different portions of the truck tray may be identifiable. For example, the truck tray may contain a load carrying portion (the main hole or cavity in which the load sits) and a cabin protection portion - such as substantially flat shield portion extending from a forward wall of the load carrying portion. The ROI detector may thus be configured to identify both the boundary of the truck tray, and a boundary of the load carrying portion. In the event that the boundary of the load carrying portion cannot be identified, then the boundary of the truck tray may be used as the boundary of the container ROI. The output of the ROI detector is a boundary and location in the image and is used to generate a binary mask 558. Identification of the container ROI and generation of binary mask is further illustrated in Figures 6A to 6F. Figures 6A, 6C and 6E are three images of the same damaged load carrying vehicle captured at three different time points, and Figures 6B, 6D, and 6F are masked images of Figures 6A, 6C, and 6E wherein each mask indicates the boundary of the truck tray in the respective image. In this example the only part of the vehicle visible from above is the truck tray (i.e., the vehicle container), and the boundary of the container ROI is also the vehicle boundary.

[0082] In the segmentation pathway (second parallel processing pathway) 560 the image is segmented (or divided) into a plurality of overlapping sub-images 562. A damage detector 564, including a trained damage region object detection model processes each sub-image. The trained damage region object detection model is trained to identify damage to the load carrying container by identifying one or more damage regions in the respective sub-image. The trained damage region object detection model generates a damage boundary prediction for each damage region 566. In some examples the damage boundary prediction is a bounding box prediction (e.g. the damage region object detection model uses a boundary box predictor) that includes a rectangular damage boundary (a bounding box), a location of the bounding box in the image, and a likelihood estimate that the pixels within the boundary box represent a damaged region of the vehicle container. Heatmap converter 568 processes each bounding box prediction 566 into a heatmap representation of each bounding box. In one example, the heatmap representation is a Gaussian heatmap of the bounding box. For example, if the bounding box has dimensions (2X. 2Y) centered at location (x0, y0), a 2D gaussian is generated using: f(x, y) = A exp + (1)

[0083] That is the Gaussian representation is centered on the bounding box location and the X (half width) and Y (half height) dimensions are used as the variance (az = X, oy = Y). The amplitude A may be set to the maximum value of an intensity or desired color range. For example, a greyscale representation may be generated by setting A to 255, such that the damage region is an ellipse with grayscale values from 0 to 255. The individual Gaussians representations from all of the sub-images are then combined, for example by geometric addition, to generate a combined heatmap representation 570. After addition, the values may be rescaled back to a desired range, such as 0-255. The two parallel pathways are then recombined (+) in which the binary mask 558 is applied to the heatmap representation 570 to obtain a masked heatmap 574. Tire masked heatmap 574 is then cropped around the vehicle or container ROI boundary 576, and the cropped image is then resized 578 to a predetermined fixed size, and the resultant heatmap image is output 580. In this example, the ROI detector 556 and damage detector 564 may use trained DAMO-YOLO based object detection models.

[0084] In other examples, identification of the container ROI region and segmentation may be performed sequentially. ROI detector 556 may be used to identify the container ROI (or the vehicle boundary) and generate a binary mask 558. The binary mask 556 may then be applied to the input image 552, and this masked image is then segmented 562 into sub images. The damage detector 564 is used to predict boundary boxes 556 which are converted to heatmap representations 568 and combined to generate a combined heat map 570. As the vehicle boundary is known this can be used to set the boundaiy of the combined heat map 570 (and thus cropping step 576 can be omitted). The combined heatmap can be resized (if required) and output 580.

[0085] In some examples, each bounding box identified in a sub-image may be filtered or masked to exclude pixels outside of the container ROI. In some examples a quality filter may be applied and used to reject (i.e., discarding) a bounding box if more than a threshold number of pixels are outside of the boundary of the container ROI. In some examples, identifying a container region of interest (ROI) includes processing an image using a trained object detection model trained to detect the load carrying container of a vehicle.

[0086] In some examples, after segmenting an image and processing the sub-images, a determination is made using the output of the trained damage region object detection model of the presence of at least one damaged region in the load carrying container of the vehicle in the image. If damage is detected, an electronic damage alert may be generated and sent. This may be generated based on an analysis of a single image, or it may be generated and sent based on the analysis of multiple images of the same vehicle captured at different times.

[0087] The electronic damage alert may be an alert of the presence of damage and may be sent directly to the vehicle. For example, if image data captured from a vehicle is analyzed immediately after capture, for example, by a computing apparatus at the capture location, the electronic damage alert may be wirelessly sent to the vehicle, for example, over a Bluetooth link or other wireless communications protocol. In some examples, the electronic damage report may be sent to a central worksite monitoring system fleet management system which tracks the location of vehicles on the yy orksite. The electronic damage alert may be sent with or include an identifier of the capture location and time. In some examples, electronic damage alert is an electronic damage report further including a location in the load carrying container of the vehicle for each of the one or more damaged regions. In some examples, the electronic damage alert is an electronic report including a damage representation of the load carrying container of the vehicle indicating one or more damaged regions in the load carrying container of the vehicle. In some examples, the electronic damage report further includes a heat map representation of the vehicle load carrying container shoyving each damaged region. As outlined above, the heatmap representation may be generated by combining one or more heat map representations of damage regions.

[0088] In some examples, a plurality of images of a vehicle are captured at multiple time points, and the steps of identifying a container ROI, segmenting the image and processing each sub-image are performed separately for each of the plurality of images. This example may further include identifying the vehicle before the damage detection or by the damage detection system. Determining the presence of at least one damaged region may then include determining if a damaged region is present in a plurality of images of the same vehicle obtained at different time points, e.g., by increasing confidence in alert. In some examples, the multiple detection threshold may be based on several parameters. For example, positive detections may be required in at least a predetermined percentage of images (e.g., 50% of images), or a minimum number of consecutive images (e.g., X images in a set of Y images). Time limits may also be applied such as requiring that the images are captured at a predetermined number of time points (e.g., at least X% of images captured at Z or more time points). For example, a positive detection may be required over at least three time points or at least three positive detections in nine images captured over three time points.

[0089] An electronic damage alert may then be generated based on the consistent detection of damage over time. In some examples, the multiple images may be captured as a burst of images, for example over a few seconds as the vehicle passes under camera. This provides multiple views of the vehicle at different angles with respect to the camera and can thus reduce the likelihood of false detections due to lighting or other effects. In some examples, the multiple images may be obtained from multiple passes through the same image data capture location at spaced apart time points, such as, spaced by an hour or more (including multiple hours, days or weeks) or from multiple image data capture locations spaced around the worksite. This may provide robustness against variations in the image due to lighting, mud or dust which may obscure damage, or increase the difficulty in identifying small regions of damage, and thus using multiple images taken at different time points may be used to build robustness and confidence level in detection. In some examples, the electronic damage alert further includes a damage report indicating the damage through time to the vehicle.

[0090] In some examples, of the method may further include determining a vehicle identifier of the vehicle, and including the vehicle identifier in the electronic damage alert or using the vehicle identifier to determine one or more recipients of the electronic damage alert. Each vehicle operating on the worksite may have an associated unique vehicle identifier associated with the vehicle and a vehicle identification system may be used to identify a vehicle and be used to associate a unique vehicle with a captured image. The vehicle identification system may store details of the vehicle identifier and the associated vehicle details and includes an apparatus configured to identify a vehicle identifier at one or more locations in the worksite 126. Similarly, a target recipient (or electronic address for a target recipient) may be associated with the vehicle identifier which can be looked up using the vehicle identification system to determine where to send an electronic alert to. The target recipient may be a supervisor, vehicle owner, maintenance coordinator and an email address, mobile phone number, MAC address, or other electronic address may be stored and associated with the vehicle identifier. The vehicle identification system may also be used to track a vehicle on the worksite, and may be a standalone system or it may be a part of a fleet management system used to manage the fleet of vehicles on the worksite. The vehicle identification system may be a wireless identification system, an image recognition based system, or some other system using satellite, Internet of Things (loT) or tracking based technologies. At each location where an image is captured, the vehicle identification system may use a single methodology or systems to identify the vehicle, or the vehicle identification system may use multiple methodologies to obtain multiple estimates of the vehicle identifier which are combined or compared to increase the confidence of the unique vehicle identification. For example, a vehicle identifier from a wireless identification system could be compared with an identifier from an image recognition based system and the vehicle only classed as identified if both identifiers match.

[0091] In the example shown in Figure 1A, a wireless identification system is used including a radiofrequency identification (RFID) reader 122 is located on the truss upright member and is powered by the junction box 118, and is in communication with the vehicle identification system. A RFID tag 124 is affixed to each vehicle and encodes the vehicle identifier, or another identifier that the fleet management system can use to determine the vehicle identifier. The RFID reader 122 is configured to read an RFID tag 124 of a passing vehicle 106 and identify the vehicle. The vehicle identifier, and optionally a capture time, may then be reported to the vehicle identification system, and / or be associated with image data captured by the image sensor 110.

[0092] In other examples, other vehicle identification systems may be used. In some examples, the wireless identification system may include a vehicle mounted communications transmitter or transceiver, such as a Bluetooth or Wi-Fi transceiver, which is configured to broadcast a vehicle identifier, and a corresponding receiver or transceiver located on the truss 108 which is configured to receive and decode the vehicle identifier in transmissions from vehicle mounted communications transmitter or transceiver as the vehicle passes through the truss. The vehicle-based communications transmitter or transceiver may periodically transmit the vehicle identifier or provide it on demand in response to a request message from the truss mounted receiver or transceiver. The vehicle identifier may be the MAC address of the vehicle mounted transceiver, or the MAC address of the transceiver may be associated with the vehicle identified in the vehicle identification system, or the vehicle identifier may be encoded in an identification message. The vehicle mounted communications transmitter may communicate using a defined communications protocol such as Bluetooth, Wi-Fi, Cellular or other IEEE protocol, and may be part of an existing vehicle system, such as audio system, or part of a fleet management or communication system which provides additional functionality beyond identification of the vehicle (e.g., for monitoring vehicle status).

[0093] In some examples, the vehicle may include a position estimation apparatus, such as a global navigation system receiver. A data logger to store position records including a location of the vehicle and an associated timestamp for later uploading to the vehicle identification system or a communication system may be used to periodically transmit the current location of the vehicle to the fleet management system. The vehicle identification system may be configured to store the location of the truss and be configured to analyze the position records to determine the time that a vehicle passed under the truss, and associate the vehicle identifier, and thus the vehicle, with a respective image.

[0094] In some examples, determining the vehicle identifier may include interfacing with a fleet management system and receiving a vehicle identifier based on a capture time and a location of the image. In some examples, the fleet management system may include a sensor located in proximity to the camera (e.g., mounted on the same or and a nearby truss on which the camera is mounted) and which senses a vehicle identifier within a predefined time window of a capture time of the image.

[0095] In some examples, the vehicle identification system may include an image-based identification system. In these examples, a vehicle identifier may be painted on, affixed to, or mounted on one or more locations on the vehicle, such as over the cabin of the vehicle and / or on the sides of the vehicle such that it is visible in the field of view of the image sensor 110. The vehicle identifier may be an alphanumeric string such as “T1234” or another visual representation such as barcode or a 2D QR code encoding the vehicle identifier. In these examples, image data captured by the image sensor 110 may be analyzed to recognize and extract the vehicle identifier from the image. In some examples, a machine learning object detector may be trained to recognize the location of the visual representation in the image, such as a YOLO based object detector, and once identified a character recognition algorithm or similar object recognition algorithm used to recognize and extract the vehicle identifier at the identified location.

[0096] In other examples, the image-based identification system may be configured to identify a truck in an image, and to generate an image signature of the truck. In some examples, the container ROI may be used to generate the vehicle signature, or an object detector may be trained to identify a vehicle ROI. The image signature may include a set of multiple features including geometric parameters such as relative dimensions of different parts of the vehicle, and pixel-based parameters such as averages, variances or ranges of pixels in different parts of the vehicle or in the overall vehicle. The vehicle signature may be stored in data store, such as a database, and a unique vehicle identifier may be associated with each new vehicle signature identified. Each time an image of a vehicle is obtained, an image signature is generated from the image and compared with the signatures in the data store. If the signature does not match an existing signature, a new vehicle identifier is created, and the new vehicle identifier and the signature are stored. If the signature matches an existing signature, then the associated stored vehicle identifier may be retrieved and associated with the image. A match may be determined using similarity-based metrics such as by defining a matching range for each parameter and scoring the similarity of each parameter to generate an overall score that can be compared with a threshold. In some examples, a vehicle can be determined based on the damage previously detected. Other approaches such as correlation-based methods, machine learning methods and object detection models may also be used to determine similarity of a vehicle to an existing vehicle.

[0097] In some examples, multiple estimates of the vehicle identifier may be obtained and used to detennine the vehicle identifier. That is the same identification methods may be used on multiple images, or multiple identification methods may be used on a single image, or multiple identification methods may be used on multiple images.

[0098] In some examples, a software based reporting system may be developed to allow supervisor or fleet managers to review the results of the damage detection system. Referring to Figure 7A, the software dashboard 700 shows an electronic report including three images 702, 704, 706, of the same vehicle captured at three different time points. A vehicle information panel 710 provides the vehicle identifier (“Truck ID: T1234”) and an estimate indicating that the load carrying container of the vehicle is not damaged (“Damage Present: No”). Figure 7B shows three images of the same vehicle 702, 704, 706, displayed along with capture time points. In this example, the damage detection system has identified that damage is present which is reported in the vehicle information panel. Again, the vehicle identifier is present along with an estimate of the likelihood of the damage (95%), along with the damage location and damage size in a coordinate system which in this example is centered on the lower left comer of the vehicle when view from above. A heat map representation of the damage 712 is presented along with a close-up image 714 of the damage region and estimated damage region bounding box. Referring to Figure 7C, a set of four images 722, 724, 726, 728 are presented each showing close-up images of estimated damage regions and damage region bounding boxes, with the associated vehicle identifier, damage location, damage size and likelihood of damage.

[0099] In one example, a DAM0-Y0L0 based object detection model was trained on a training dataset including 2119 images (1926 positive samples, 193 negative samples), and a test set of 770 images (all positive samples). The images were collected over multiple months and included both daytime and nighttime images. The images were captured by a single camera with a resolution of 2048x1536 pixels, and subimages were 512x512 pixels with a 50% overlap (256 pixels) between adjacent sub-images. The training was performed using software written in Python 3.10 utilizing Pytorch, and NumPy libraries and executed on a computer configured with a GeForce RTX 3080 Ti (12Gb) GPU. Results for DAMO-YOLO on entire images yielded an accuracy of 99.4%, True Positive: 59, False Negative: 1, True Negative: 127, False Positive: 0 and AP50 of 0.72. The smallest damage region detected comprised 286 pixel region within an 11x26 bounding box. This region is 0.1% of the total area of the sub-image area (512x512 pixels), and only 0.009% of the total area of the original image (2048x1536). This demonstrates the ability of examples of the system and method described herein to identify damage regions in vehicles whilst they are still small, and thus cheaper and easier to fix. A YOLOX implementation (Ge et al, “YOLOX: Exceeding YOLO Series in 2021”, arXiv:2107.08430v2, available at arxiv.org) yielded an accuracy of 87%.

[00100] As a comparison, a YOLOv7 object based detector was trained on whole images (i.e., with no segmentation and creation of sub-images) of trucks. The original image sizes of the truck were approximately 900x600 pixels which were resized to 640x640 pixel training images. However, accuracy of this approach was low at around 42%. An attempt to train a Patchcore Anomaly Detection model (Roth et al, “Towards Total Recall in Industrial Anomaly Detection”, arXiv:2106.08265, available at arxiv.org) to identify damage regions in whole images failed at initial attempts, and due to these poor results was not trained on large scale data.

[00101] These results thus show that by segmenting the original image into overlapping sub-images and using a trained ML object detection model on the sub-images, increased reliability and accuracy in identification of damage regions in vehicles can be achieved. Robustness and confidence can be increased through the use of multiple detections - in either multiple images collected in a burst or from images collected over multiple time points (e.g., hours or days apart). Some examples allowed damage regions to be detected whilst they are still small (e.g., occupying a small proportion of an image), which then would allow scheduling of maintenance before the damage becomes more significant and costly to remedy. Further, the detection is robust to the wide range of lighting conditions that may be found on a worksite, such as a mining worksite. Examples of the system, methods, and associated apparatus described herein may thus be used to identify damage to load carrying containers of load carrying vehicles.

[00102] The reference to any prior art in this specification is not, and should not be taken as, an acknowledgement or any form of suggestion that such prior art forms part of the common general knowledge.

[00103] In some cases, a single example may, for succinctness and / or to assist in understanding the scope of the disclosure, combine multiple features. It is to be understood that in such a case, these multiple features may be provided separately (in separate examples), or in any other suitable combination. Alternatively, where separate features are described in separate examples, these separate features may be combined into a single example unless otherwise stated or implied. Uris also applies to the claims which can be recombined in any combination. That is a claim may be amended to include a feature defined in any other claim. Further a phrase referring to “at least one of’ a list of items refers to any combination of those items, including single members. As an example, “at least one of: a, b, or c” is intended to cover: a, b, c, a-b, a-c, b-c, and a-b-c.

[00104] It will be appreciated by those skilled in the art that the disclosure is not restricted in its use to the particular application or applications described. Neither is the present disclosure restricted m its preferred example with regard to the particular elements and / or features described or depicted herein. It will be appreciated that the disclosure is not limited to the example or examples disclosed, but is capable of numerous rearrangements, modifications and substitutions without departing from the scope as set forth and defined by the following claims.

[00105] Please note that the following claims are provisional claims only and are provided as examples of possible claims and are not intended to limit the scope of what may be claimed in any future patent applications based on the present application. Integers may be added to or omitted from the example claims at a later date so as to further define or re-define the scope.

Claims

1. A method for detecting damage in a load carrying container of a load carrying vehicle, the method including:identifying a container region of interest (ROI) in one or more images wherein the container ROI is estimated to contain an empty load carrying container of a vehicle;segmenting, for each of the one or more images, the respective image into a plurality of overlapping sub-images, where each sub-image includes at least two overlapping portions where each overlapping portion including a plurality of contiguous pixels shared with an adjacent sub-image;processing each sub-image with a trained damage region object detection model, wherein the trained damage region object detection model is trained to identify damage to the load carrying container in one or more damage regions in the sub-image, each damage region including a plurality of pixels associated with damage within a boundary; anddetermining, using the output of the trained damage region object detection model, the presence of at least one damaged region in the load carrying container of the vehicle in one or more images.

2. The method as claimed in claim 1, further including:generating and sending an electronic damage alert based on the determination of the presence of at least one damaged region in the load carrying container of the vehicle m one or more images.

3. The method as claimed in claim 2, wherein electronic damage alert is an electronic damage report further including a location in the load carrying container of the vehicle for each of the one or more damaged regions.

4. The method as claimed in claim 2 further including determining a vehicle identifier of the vehicle, and including the vehicle identifier in the electronic damage alert or using the vehicle identifier to determine one or more recipients of the electronic damage alert.

5. The method as claimed in any preceding claim, wherein the trained damage region object detection model is trained on attaining set of training images, wherein the set includes a plurality of sub-images obtained from a plurality of images of a load earn ing container of a plurality of vehicles, and the training images are labelled with damage containing regions, and the plurality of sub-images includes sub-images with damaged regions and sub-images without damaged regions.

6. The method as claimed in any preceding claim, further including obtaining or accessing one or more images generated by one or more image sensors on a worksite wherein the one or more images sensors are configured to generate the one or more images from a captured image data set.

7. The method as claimed in claim 6, where generating the one or more images from a captured image data set includes:generating a point cloud representation of a field of view from an image data set captured by the one or more image sensors at the worksite;generating an image from the point cloud representation of the field of view, wherein the image is generated as a two dimensional image from a predefined reference point;8. The method as claimed in claim 7, wherein the captured image dataset includes at least two successive or simultaneous images captured by the one or more image sensors.

9. The method as claimed in claim 8, wherein the one or more image sensors include at least one stereo camera or a plurality of image sensors each in a different location.

10. The method as claimed in claim 7, wherein the captured image data set is captured by at least one ranging based image sensor configured to scan the field of view and capture range data at a plurality of points in the field of view.

11. The method as claimed in any preceding claim, wherein identifying the container ROI in the image and segmenting the image are performed in parallel, and the method further includes filtering or masking each damage region identified in a sub-image to exclude pixels outside of the container ROI.

12. The method as claimed in any preceding claim, wherein identifying the container ROI in an image is performed prior to segmenting the image, and segmenting the image includes segmenting the container ROI of the image.

13. The method as claimed in any preceding claim, wherein identifying the container ROI includes processing an image to detect the load carrying container of a vehicle and determining if the load carry ing container is empty.

14. The method as claimed in any preceding claim, wherein identifying a container region of interest (ROI) includes processing an image using a trained object detection model trained to detect the load carrying container of a vehicle.

15. The method as claimed in claim 9, wherein the trained object detection model is trained to detect an empty load carrying container of a vehicle.

16. The method as claimed in claim 7, wherein identifying a container region of interest (ROI) in one or more images includes:determining the presence of a vehicle in the point cloud representation and then determining the container ROI of the vehicle and depth information obtained from the point cloud representation,and the step of determining if the container is empty is performed using depth information obtained from the point cloud representation by identifying the presence of a base surface of a container and one or more surrounding walls extending proximally from the base surface.

17. The method as claimed in claim 16, wherein each pixel in the image is assigned an intensity value and a depth value obtained from the point cloud representation, or the intensity of the pixel is determined using the depth value obtained from the point cloud representation.

18. The method as claimed in any preceding claim, wherein each of the sub-images are generally of the same size.

19. The method as claimed in claim 18, wherein each overlapping portion includes one half of the subimage.

20. The method as claimed in claim 2, wherein the electronic damage alert includes an electronic report including a damage representation of the load carrying container of the vehicle indicating one or moredamaged regions in the load carrying container of the vehicle wherein a location is determined using the boundary of one or more damaged regions identified in one or more sub-images.

21. The method as claimed in claim 20, further including generating a heat map representation of each damaged region, and the damage representation is a heatmap representation generated by combining one or more heat map representations.

22. The method as claimed in claim 20, further including:capturing a plurality of images of the vehicle, and the steps of identifying a container ROI, segmenting the image and processing each sub-image are performed separately for each of the plurality of images; andthe damage representation is generated by combining one or more damage regions identified in one or more sub-images of one or more images.

23. The method as claimed in claim 2, further including:capturing a plurality of images of the vehicle at multiple time points, and the steps of identifying a container ROI, segmenting the image and processing each sub-image are performed separately for each of the plurality of images; anddetermining the presence of at least one damaged region includes determining if a damaged region is present in a plurality of images obtained at different time points and generating an electronic damage alert further includes a damage report indicating the damage through time to the vehicle.

24. The method as claimed in any preceding claim, further including determining a vehicle identifier of the vehicle in the image.

25. The method as claimed in claim 24, wherein determining the vehicle identifier includes interfacing with a fleet management system and receiving a vehicle identifier based on a capture time and a location of the image.

26. The method as claimed in claim 25, wherein the fleet management system includes a sensor located in a proximity of the camera and which detects a vehicle identifier within a predefined time window of a capture time of the image.

27. The method as claimed in claim 24, wherein determining a vehicle identifier of the vehicle in the image includes processing the image to detect a vehicle identifier located on the vehicle.

28. The method as claimed in claim 24, wherein determining a vehicle identifier of the vehicle in the image includes identifying a vehicle region of interest (ROI), and determining if the vehicle ROI matches a vehicle in a database of previously identified vehicles, and if there is a match extracting a vehicle identifier from the database, and if there is not a match then generating a vehicle identifier and storing the generated vehicle identifier with a representation of the vehicle in the database.

29. The method as claimed in claim 24, wherein a plurality of estimates of the vehicle identifier are obtained using one or more identification methods on one or more images, and the plurality of estimates are used to determine the vehicle identifier.

30. The method as claimed in any preceding claim, wherein the image is an overhead image data captured when the vehicle is located under a camera.

31. A computational apparatus configured to detect damage in a load carrying container of a load carrying vehicle including:at least one memory';at least one processor configured to:identify a container region of interest (RO1) in one or more images wherein the container ROI is estimated to contain an empty' load carrying container of a vehicle;segment, for each of the one or more images, the respective image into a plurality of overlapping sub-images, where each sub-image includes at least tw o overlapping portions where each overlapping portion including a plurality of contiguous pixels shared with an adjacent subimage;process each sub-image with a trained damage region object detection model, wherein the trained damage region object detection model is trained to identify damage to the load carrying container in one or more damage regions in the sub-image, each damage region including a plurality of pixels associated with damage within a boundary; anddetermine, using the output of the trained damage region object detection model, the presence of at least one damaged region in the load carrying container of the vehicle in one or more images.

Citation Information

Patent Citations

  • Container body damage detection method and system

    CN115527200A

  • Wear and loss detection system and method using bucket-tool templates

    US20220136218A1