Bulk volume monitoring and estimation of material

A machine learning system using image data from load containers addresses the challenge of inconsistent bulk volume estimation by identifying key points and calculating a three-dimensional profile, enhancing resource planning and operational efficiency in heavy equipment operations.

GB2635203BActive Publication Date: 2026-03-20MOTION METRICS INTERNATIONAL CORP
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
GB · GB
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-11-03
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

Existing technologies struggle to provide consistently accurate bulk volume estimations of materials in load containers of heavy equipment due to factors like material characteristics, bucket fill factor, material compaction, and loading technique variations, leading to inefficiencies in resource planning and operational management.

Method used

A machine learning-based system that utilizes image data from load containers during operation to estimate bulk volume by identifying key points, estimating occluded points, and calculating a three-dimensional profile, employing neural networks and reference planes to determine the pose and volume of materials within the load container.

Benefits of technology

Enables precise and consistent bulk volume estimation without additional sensors, allowing for improved resource planning, operator productivity assessment, and enhanced operational efficiency across various industries using heavy equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000001_0000
    Figure 00000001_0000
  • Figure 00000002_0000
    Figure 00000002_0000
  • Figure 00000003_0000
    Figure 00000003_0000
Patent Text Reader

Abstract

The volume of materials held by a load container during a digging cycle and / or held by a load container is estimated. Image data of a load container during a digging cycle are collected. Key points an
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] This disclosure relates generally to bulk volume estimation, fill factor, and monitoring of materials excavated or loaded into a load container of heavy equipment, and more specifically, for performing a bulk volume estimation of a load of material held within a load container by performing a machine learning analysis on image data relating to a digging cycle of the heavy equipment. DESCRIPTION OF THE RELATED TECHNOLOGY

[0002] Heavy equipment monitoring tools are advanced technological systems designed to track and monitor the performance, location, and health of heavy machinery used in various industries such as construction, mining, agriculture, and transportation. These tools leverage a combination of hardware, software, and sensors to provide real-time data and insights to equipment operators, fleet managers, and maintenance teams.

[0003] Some examples of objectives of heavy equipment monitoring tools are to enhance operational efficiency, optimize productivity, reduce downtime, improve safety, and lower maintenance costs. By capturing and analyzing data from the equipment, these tools enable businesses to make informed decisions, streamline operations, and increase the utilization of their assets. SUMMARY

[0004] Introduced here is a system and method that uses techniques / technologies to provide a bulk volume estimation of material held by a load container using image data relating to the load container while it is in operation. The system trains machine learning models (e.g., neural networks) using relevant features extracted from the image data to estimate the bulk volume of the material. Upon system deployment, an estimation can occur while there is load inside the load container. Based on the bulk volume estimations, the system can also provide quantitative assessments of the load container and operator. It should be noted the system can perform these estimations using image data (e.g., point clouds, depth maps, RGB images, etc.) of the load container while it is in use and without the need for additional sensors.

[0005] Examples of the present disclosure are directed to providing mechanisms, including a method and non-transitory computer storage mediums for using image data in a machine-learning-based framework to estimate bulk volumes of material held by the load container while in use. In some examples, the method collects image data depicting a load 2 container during a digging cycle and determines an operational phase of the load container based on the images. While the load container is in operation and contains material, the method identifies key points of the load container using the image data representing the load container. The method also estimates occluded key points of the load container using the identified key points and the image data estimated from the prior model, which are continuously updated. The method further calculates a three dimensional profile of the material held by the load container to separate the load from the load container, fits a model with the key points and / or the occluded key points to the three dimensional profile of the material to determine pose, and estimates a bulk volume of the material based on the three dimensional profile of the material and the pose. This pose estimation may also be accomplished by matching key points with a reference 3-dimensional model, such as a CAD model.

[0006] Examples of the present disclosure include a method of bulk volume estimation of material held by a load container. The method includes collecting image data of a load container during a digging cycle and identifying key points of the load container from the image data representing the load container. The method also includes estimating occluded key points of the load container using the image data representing the load container and the key points. The method also includes calculating a three dimensional profile of the material held by the load container. The method further includes a pose estimation method of fitting a model with the key points and the occluded key points to the three dimensional profile of the material to determine pose. The method also includes estimating a bulk volume of the material based on the three dimensional profile of the material and the pose.

[0007] Additional examples of the present disclosure include non-transitory computer readable storage medium including instructions stored thereon which, when executed by a processor, cause the processor to collect image data of a load container during a digging cycle and identify key points of the load container from the image data representing the load container. The instructions also cause the processor to estimate occluded key points of the load container using the image data representing the load container and the key points and calculate a three dimensional profile of the material held by the load container. The instructions also cause the processor to fit a model with the key points and the occluded key points to the three dimensional profile of the material to determine pose. The instructions further cause the processor to estimate a bulk volume of the material based on the three dimensional profile of the material and the pose. Optionally, the processor fits a reference plane to the load container based on a normal vector associated with the key points and the occluded key points, wherein the normal vector represents the pose of the load container.

[0008] According to another embodiment, a method for bulk volume estimation of material held by a load container includes collecting stereo image data of a load container during a digging cycle, collecting two dimensional image data of the load container during the digging cycle, identifying key points of the load container using the stereo image data and the two dimensional image data, and estimating occluded key points of the load container using the key points, the stereo image data, and the two dimensional image data. The method further includes estimating a material region using the key points and the occluded key points, estimating pose of the load container, and estimating a bulk volume of the material using the estimated material region and the estimated pose.

[0009] According to still another embodiment, a method for bulk volume estimation of material held by a load container includes collecting image data of a load container during digging cycles; dynamically creating a reference correlating poses of the load container and three dimensional profiles of material held by the load container during the digging cycles to bulk volume of the material held by the load container during the digging cycles; collecting image data of the load container during a subsequent digging cycle; determining a three dimensional profile of the material held by the load container in the subsequent digging cycle; and estimating bulk volume of the material held by the load container in the subsequent digging cycle using the reference, the image data of the load container during the subsequent digging cycle, and the three dimensional profile of the material held by the load container in the subsequent digging cycle. In some embodiments, the estimating step does not require the load container to be in any particular pose.

[0010] Further examples are directed to a system for bulk volume estimation of material held by a load container and configured to perform the methods described above.

[0011] At a high level, the technology relates to using machine learning techniques to analyze image data associated with a load container while in operation to estimate the bulk volume of material held by the load container. During the training of the system, training data may be compiled using data preprocessing techniques on a sequence of image data of the load container while in operation. From the training data, machine learning models are trained to determine the operational phases of the load container, the key points associated with the load container, the occluded key points associated with the load container, pose estimation statistical model, and a load segmentation of the material.

[0012] Upon deployment, the trained machine learning models can be provided image data of a load container in operation. Using the derived information provided by the models, the system can utilize a three-dimensional profile of the load segmentation of the material and a reference plane fitted to the key points and occluded key points to estimate the bulk volume of the material based on measurements of the material below and above the reference plane.

[0013] This summary is intended to introduce a selection of concepts in a simplified form that is further described in the detailed description section of this disclosure. The summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended as an aid in determining the scope of the claimed subject matter. Additional objects, advantages, and novel features of the technology will be set forth in the following description and, in part, will become apparent to those skilled in the art upon examination of the disclosure or learned through practice of the technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] These and other features, aspects, and advantages of the examples of the disclosure will become better understood with regard to the following description, appended claims, and accompanying drawings where:

[0015] FIG. 1 illustrates a schematic view of a mining shovel, with the mining shovel in an operational phase, in accordance with examples of the present disclosure.

[0016] FIG. 2 illustrates an exemplary plan view of landmarks associated with a digging element, in accordance with examples of the present disclosure.

[0017] FIG. 3 A-F illustrate exemplary schematic views of a digging element of the mining shovel represented in FIG. 1 in various operational phases of a digging cycle, in accordance with examples of the present disclosure.

[0018] FIG. 4 illustrates a front view of a stereoscopic camera, in accordance with examples of the present disclosure.

[0019] FIG. 5 illustrates a diagram of a machine-learning framework estimating bulk volume of material held by a load container using image data depicting the load container in operation, in accordance with examples of the present disclosure.

[0020] FIG. 6 illustrates an example heavy equipment monitoring system, in accordance with examples of the present disclosure.

[0021] FIG. 7 illustrates a schematic diagram of a bulk volume estimation system in accordance with examples of the present disclosure.

[0022] FIG. 8 illustrates a schematic representation of a feature segmentation of material held by a digging element, in accordance with examples of the present disclosure.

[0023] FIG. 9 illustrates a training and deployment cycle of a bulk volume estimation system, in accordance with examples of the present disclosure.

[0024] FIG. 10 illustrates a flowchart of estimating a bulk volume of material held by a load container from image data depicting the load container in operation using machine learning models in accordance with examples of the present disclosure.

[0025] FIG. 11 illustrates a schematic diagram of an exemplary environment in which an bulk volume estimation system can operate, in accordance with examples of the present disclosure.

[0026] FIG. 12 illustrates a block diagram of an exemplary computing device, in accordance with examples of the present disclosure.

[0027] While the present disclosure is amenable to various modifications and alternative forms, specifics thereof, have been shown by way of example in the drawings and will be described in detail. It should be understood, however, that the intention is not to limit the particular examples described. On the contrary, the intention is to cover all modifications, equivalents, and alternatives falling within the scope of the present disclosure. Like reference numerals are used to designate like parts in the accompanying drawings. DETAILED DESCRIPTION

[0028] This disclosure relates generally to bulk volume estimation, fill factor percentage, and monitoring of materials excavated or loaded into a load container operated by heavy equipment. More specifically, this disclosure relates to performing a bulk volume estimation of a load of materials within a load container by performing a machine learning analysis on image data retrieved during the operational phases of the heavy equipment. The following description is directed to some particular examples for the purpose of describing innovative aspects of this disclosure. However, a person having ordinary skill in the art will readily recognize that the teachings herein can be applied in a multitude of different ways. Overview

[0029] Earth excavation may be a process that involves the removal of earthen material from a site to form a cavity or open area. The purpose of earth excavation may include preparing a site for construction, creating space for infrastructure, or removing earthen material for processing elsewhere. The choice of excavation method can depend on the project requirements and the nature of the earthen material. Manual excavation can be used for smaller jobs or in areas where heavy equipment may be unable to operate. For larger jobs, mechanical methods are typically used. These involve heavy equipment such as excavators, backhoes, bulldozers, draglines, front loaders, haul trucks, skid-steer loaders, trenchers, cranes, and the like.

[0030] The use of heavy equipment provides a multitude of benefits during the earth excavation process including the use of heavy equipment monitoring tools. These monitoring tools generally refer to digital and electronic devices, systems, and software that track and analyze the performance, efficiency, and safety of workers, heavy machinery like excavators, bulldozers, loaders, haul trucks, and the like. These tools gather valuable data that help in predictive maintenance, location tracking, fuel consumption, and overall operational performance.

[0031] For instance, estimating the volume of material in a load container can be challenging for several reasons. These challenges include, but are not limited to, material characteristics, bucket fill factor, material compaction, variation in bucket shape and size, loading technique, material spillage, and bucket wear and tear. For instance, the type of material being excavated greatly influences the volume that fits in a bucket. Loose soil, packed soil, gravel, rocks, or mixed earthen material all have different densities and can fill the bucket differently. Additionally, not all materials fill a bucket to their full capacity. For instance, larger rocks or boulders may leave empty spaces in the bucket, while smaller materials like sand, fines, or gravel may fill it completely. Technologies such as load sensing systems and onboard weighing systems can provide rough volume estimations but fail to provide consistently accurate estimations.

[0032] Examples of the present disclosure include a bulk volume estimation method and system, including a neural network trained to estimate activity labels representing various operational phases of the digging cycle to be associated with image data relating to the load container. As described further within, these activity labels include, without limitation, “background,” “excavate,” “swing full,” “dump truck,” “dump no truck,” “swing empty,” “idle,” “tramming or hauling,” “maintenance,” and combinations and intermediaries of same. In some examples, the activity labels may also include “maintenance.” Once trained, the neural network inputs the image data into a neural network trained to classify operational phases of the load container, analyzes the image data by the neural network, and outputs an operational phase estimation of the load container.

[0033] In some examples, the bulk volume estimation method and system include a volume estimation module having a computer model (e.g., a conditional probability model) trained to estimate occluded key points in image data depicting a load container. The load container may be, for example, in a “swing full” operational state. The conditional probability module can be trained from image data taken of a load container during a “swing empty” state and image data taken of a load container during a “swing full” state. By training on both sets of image data, the conditional probability model can estimate pose and potentially occluded key points occluded by the material during a “swing full” state of a load container during a digging cycle. This pose estimation method may utilize observed points including key points, with a reference point cloud.

[0034] In some examples, the conditional probability model can continuously be trained to refine its estimations of the occluded key points. In such instances, image data depicting the load container in a “swing empty” state can be utilized to train and refine the parameters of the conditional probability model until convergence criteria are met.

[0035] In some examples, the system aggregates bulk volume estimations and associates the volume estimations with an operator. Aggregated bulk volume totals can be compiled and reported on a per-operator basis. For instance, aggregated bulk volume estimations of a load container can be compiled during an entire shift of an operator, per truck being filled, or on a perdigging cycle basis.

[0036] In some examples, the system performs the bulk volume estimations of the material by retrieving a load segmentation of the material held by the load container in the “swing full” state and determining a three-dimensional profile of the material using the load segmentation as a region of interest. The system can then estimate the bulk volume of the material based on the three-dimensional profile of the material in relation to the reference plane. In some instances, the volume estimation be performed using a discrete volume summation methodology. Using this approach, the pose of load container is utilized and determined using the normal vector of the reference plane. The pose can be used to estimate the bulk volume of the material by providing a position and an orientation of the load container.

[0037] In other example instances, the volume estimations can be performed using a continuous profile model fit on the material region. And in still other example instances, nominal struck volume is used to estimate heap volume to estimate fill factor and bulk volume. Among other ways, the manufacturer specifications for the load container model can be used to estimate bulk volume from heap volume.

[0038] Having briefly described an overview of aspects of the present disclosure, various terms used throughout this description are provided. Although more details regarding various terms are provided throughout this description, general descriptions of some terms are included below to provider a clearer understanding of the ideas disclosed herein:

[0039] Heavy equipment generally refers to machinery designed for construction, mining, earthmoving, and other heavy-duty tasks. Examples of heavy equipment include, but are not limited to, excavators, backhoe loaders, bulldozers, wheel loaders, skid-steer loaders, trenchers, motor graders, rock breakers, and the like. Heavy equipment may typically be used in the construction and mining industries, enabling tasks such as digging trenches, excavating foundations, clearing land, loading trucks, and moving earthen materials.

[0040] Digging element generally refers to a piece of machinery or an attachment designed for digging and loading material. Digging elements, such as a shovel or bucket, can be used to lift and move bulk materials, such as earthen material, soil, coal, gravel, snow, rock, and sand. Additionally, digging elements may come in various shapes and sizes for different applications and various replaceable components such as teeth, wear plates, lip shrouds, wing shrouds, and reinforcing ribs.

[0041] Bulk volume generally refers to the ratio of the volume of materials after excavation or load from stockpile compared to the volume of materials in its natural state. Bulk volume, also known as “bulk factor,” can be a measurement that accounts for the increase in volume that occurs when a material may be excavated or loaded from stockpile and becomes loosened or less compact by a load container. A bulk volume measurement can impact the amount of material to be moved or hauled and can significantly influence cost and resource planning. The exact bulk volume can vary based on the type of material (e.g., sand, clay, rock), the degree of compaction, the moisture content, the load container, and other factors.

[0042] Fill factor generally refers to the ratio of the actual volume of the material contained in the bucket compared to its maximum capacity, and can be expressed as a decimal or percentage. The fill factor can quantify the extent to which the bucket is filled with material. The actual volume of material can refer to the volume occupied by the material within the bucket, considering factors such as the shape of the material pile or heap, any void spaces, and the distribution of the material within the bucket. The maximum capacity of the bucket can be determined by its design, which can include factors such as the shape, dimensions, and any physical constraints that affect the capacity of the bucket. As such the fill factor can provide insight into how efficiently the bucket is being utilized in holding material. For instance, a fill factor of 1.0 or 100% can mean the bucket is filled to its maximum capacity, while a fill factor less than 1.0 can indicate that the bucket is not fully utilized and has unused capacity. It should be noted that the fill factor can vary depending on various factors, including the shape of the bucket, the properties of the material being handled, and any design or operational constraints.

[0043] As used herein, a neural network generally refers to a machine-learning model that learns to approximate unknown functions by analyzing example (e.g., training) data at different levels of abstraction. Generally, neural networks can model complex non-linear relationships by generating hidden vector outputs along a sequence of inputs. In particular, a neural network can include a model of interconnected digital neurons that communicate and learn to approximate complex functions and generate outputs based on a plurality of inputs provided to the model. A neural network can include a variety of machine learning models, including convolutional neural networks, recurrent neural networks, deep neural networks, and deep stacking networks, to name a few examples. A neural network may include or otherwise make use of one or more machine learning algorithms to learn from training data. In other words, a neural network can include an algorithm that implements machine learning to attempt to model high-level abstractions in data. An example implementation may include a convolutional neural network, including convolutional layers, pooling layers, and / or other layer types.

[0044] As such, various aspects of the present disclosure relate generally to heavy equipment, and more particularly, to bulk volume estimation of material held by a load container operated by heavy equipment during a digging cycle. Some aspects specifically relate to a machine learning framework that determines the digging cycle of the load container, identifies key points (e.g., landmarks) associated with the load container and surrounding area, estimates the three dimensional profile of the material being held by the load container, and calculates a bulk volume of the material during the digging cycle. In some examples, a neural network is used to detect the operational phase of the load container (e.g., “excavate,” “swing full,” “dump,” “swing empty,” “idle”). An additional neural network may be used to estimate the key points, or landmarks (e.g., tooth, shroud, side cutters, wear plates). During the swing full operational phase, a neural network can be used to estimate the three dimensional profile of the material stored by the load container. A computer model (e.g., a conditional probability model) can be used to infer the location of occluded key points and to estimate the pose of the load container using a reference plane fitted to the detected key points and occluded key points. Using the three dimensional profile of the materials in relation to the reference plane and pose of the load container, a bulk volume estimation of the material being held by a load container may be performed.

[0045] Particular aspects of the subject matter described in this disclosure can be implemented to realize one or more of the following potential advantages. The bulk volume estimation machine learning framework described herein provides greater utilization flexibility because examples are generalizable to all types of load containers used by various heavy equipment. As such, the techniques described are not specific to particular buckets or shovels, nor are they specific to the heavy equipment. As such, this provides greater compatibility with other monitoring systems across a spectrum of industries that may not utilize the same sensors for their equipment. Other additional advantages include the ability to aggregate the bulk volume of material across multiple digging cycles to determine certain metrics such as quantizing the productivity of a particular digging configuration, the productivity of the operator, the performance of the operator, and monitoring the fill factor of bucket or haul truck. Example Heavy Equipment Operating Environment

[0046] Referring to FIG. 1, an example heavy equipment operating environment 100 is shown. As shown, in the operating environment 100, the heavy equipment may be a mining cable shovel 102 generally configured for excavating material from a mining area 114. The mining area 114 can include open-pit mines, strip mines, quarries, oil sands, coal mines, metal mines, and other material piles shown generally at 110. In other examples, the heavy equipment can be any other type of equipment with a digging element capable of engaging the mining area 114 or any other type of equipment designed to transfer loads from one location to another, such as a dragline shovel, a backhoe excavator, hydraulic face shovel, and the like.

[0047] As shown, the mining shovel 102 includes a frame 120 pivotably mounted on a track 122, a boom 124 mounted to the frame 120, a handle 125 pivotably mounted to the boom 124, a digging element 126 mounted to the handle 125, and a control mechanism shown generally at 128 for controlling the mining shovel 102 to perform an excavate operation of the mining area 114 or a dump operation of a load excavated from the mining area 114 or a material pile. The track 122 enables the mining shovel 102 to traverse across the open pit mine. The frame 120 is coupled with a power source (not shown) for powering the track 122 and the control mechanism 128 and an operator station 121 for housing an operator. The operation of the heavy equipment may be autonomous, semi-autonomous, or non-autonomous. The pivotal mounting 131 between the frame 120 and the track 122 allows the frame 120 to rotate about a y-axis of a coordinate axes 130 relative to the track 122.

[0048] The boom 124 has a distal end 132 and a proximal end 134. The distal end 132 may be fixedly mounted (or pivotably mounted in some examples) to the frame 120 such that rotation and movement of the frame 120 translates into corresponding rotation and movement of the boom 124. The pivotal mounting 133 between the boom 124 and the handle 125 may be generally at a midway point between the distal and proximal ends 132 and 134 of the boom 124. Similarly, the handle 125 has a distal end 136 and a proximal end 138. The pivotal mounting 133 between the boom 124 and the handle 125 may allow the handle 125 to rotate about a z-axis of the coordinate axes 130 relative to the boom 124 and also may allow a mounting point between the boom 124 and the handle 125 to translate along a length of the handle 125 between the distal and proximal ends 136 and 138 of the handle 125. The digging elementl26 may be pivotally mounted to the distal end 136 of the handle 125. The digging element 126 may be an excavator bucket. Excavator buckets are a type of load container and are a crucial component of the machine that digs and transports material. While the design of the bucket can vary based on the task, the basic components of an excavator bucket include a bucket shell, teeth, shrouds, digging edge, wing shrouds, bucket linkage, wear plates, and reinforcing ribs and buttons. The individual components can potentially be monitored to ensure the bucket remains operable and in working condition Rotation and movement of the handle 125 generally translates into corresponding rotation and movement of the digging elementl26. The pivotal mounting 135 between the handle 125 and the digging elementl26 may also allow the digging elementl26 to rotate about the z-axis of the coordinate axes 130 relative to the handle 125.

[0049] As also shown in FIG. 1, the control mechanism 128 may include a motorized pulley 140 coupled to the distal end 132 of the boom 124, at least one cable 142 having a first end coupled to the frame 120 via a frame motorized spool 144 and a second end coupled to the digging element 126 via an implement motorized spool 146, a motorized pivot 148 as the pivotal mounting 133 between the boom 124 and the handle 125, and a door control mechanism 149 of a bottom wall 162 of a load container 154, but other configuration are possible. The frame motorized spool 144 and implement motorized spool 146 may allow the digging element 126 to rotate by reeling in and spooling out the at least one cable 142 during different stages of the operational phases of the mining shovel 102, as described below. The motorized pivot 148 may allow the mounting point between the boom 124 and the handle 125 to translate along the length of the handle 125 between the distal and proximal ends 136 and 138 of the handle 125. The motorized pulley 140 may support the extension and retraction of the cable 142 and the translation of the handle 125. The door control mechanism 149 may open and close a bottom wall 162 of the load container 154, as described below.

[0050] The mining shovel 102 also includes a visual sensing system 180 mounted to the mining shovel 102 via a bracket 182 coupled to the boom 124, and may be proximate the motorized pulley 140. The visual sensing system 180 is generally positioned such that a field-of-view 184 of the visual sensing system 180 captures signals, light, images, and / or video of the digging element 126 as data, and particularly the load container 154, the shrouds, the teeth, and the load, if any, and a portion of the mining area 114 lying behind the digging element 126 during the various stages of the digging cycle. It should be noted, in some examples, the visual sensing system 180 can be mounted to other locations or other components of the heavy equipment to capture the digging element 126, the load, and a portion of the mining area 114 depending on the type of heavy equipment. For example, a visual sensing system 180 may be mounted to a handle of a hydraulic face shovel. Additionally, or alternatively, in other examples, the field-of-view 184 may be selected to capture other components of the heavy equipment or other background features associated with the digging element 126 and / or mining area 114. For example, the field-of-view 184 may be selected to capture signals, light, images, and video of the mining area 114 or earthen material pile before, during, and after each excavation stage. The visual sensing system 180 may also be mounted at other locations or to other components of the heavy equipment depending on the desired field of view.

[0051] In some examples, the visual sensing system 180 includes a Light Detection and Ranging (LiDAR) sensor. LiDAR sensors use pulses of light from a laser to send and receive signals. The LiDAR sensor measures the amount of time it takes for the light to return after bouncing off an object. This information can be used to create a detailed three-dimensional (3D) representation of the object because the LiDAR sensor can calculate the direction and angle at which each pulse is sent out and returned and can use this distance measurement to create a precise 3D coordinate for each point that reflected a light pulse. The 3D representation capabilities of the LiDAR sensor allow for 3D mapping and surveying. As such, the LiDAR sensor can generate high-resolution maps of the shape and location of an object (e.g., the material held by the load container 156).

[0052] In some examples, the visual sensing system 180 includes a Radio Detection and Ranging (radar) sensor. The radar sensor can use radio waves to detect objects. Just like with LiDAR sensor, the radar sensor sends out a signal and measures the time it takes for that signal to return after reflecting off an object. The time of flight measurements can be used to determine the distance of an object, but also its size and its relative velocity. In some examples, the radar sensor is a Synthetic Aperture Radar (SAR) configured to create 3D reconstructions of an object (e.g., the material held by the load container 156). SAR works by moving the radar sensor over the target area, and taking a series of radar readings as it travels. By combining these readings, the SAR can create a detailed image of the object. In some examples, the radar sensor is encompassed within a phased array radar system that electronically steers the direction of the radar beam, 02 05 24 allowing the system to scan an area in multiple directions and thus create a three-dimensional representation of the object space.

[0053] In some embodiments, the visual sensing system 180 includes stereoscopic camera capable of capturing images or video of the field-of-view 184 during the operational phases of the mining shovel 102. An example of a stereoscopic camera is shown in FIG. 4 at 190 and may be the Multi Sense S30 (RTM) stereocamera from Carnegie Robotics, LLC. The stereoscopic camera 190 may be capable of using stereo-imaging to obtain a three-dimensional (3D) point cloud of the images and frames of the video. In the example shown, the stereoscopic camera 190 includes a first imager 192 and a second imager 194 mounted spaced apart by a first fixed distance 195 within a housing 196. The first and second imagers 192 and 194 capture first and second images of the field-of-view 184 from a slightly different perspective (such as left and right or up and down). Combining the first and second images from slightly different perspectives and knowing the fixed distance 195 can allow an internal edge processor of the stereoscopic camera 190 or an embedded processor circuit of the mining shovel 102 described below to calculate three-dimensional coordinates for pixels associated with features found in both the first and second images and to estimate a depth value of that pixel relative to the stereoscopic camera 190. Additionally, in the example shown, the first and second imagers 192 and 194 may capture the first and second images in monochrome, and the visual sensing system 180 may further include a third imager 198 which captures a third image of the field-of-view 184 in color (red, green and blue). The third imager 198 may be separated from the first imager 192 by a second fixed distance 197 and may be separated from the second imager 194 by a third fixed distance 199.

[0054] In some examples, the stereoscopic camera 190 captures a video of the field-of-view 184, and the video captured by one or more of the first, second, and third imagers 192, 194, 198 may be captured substantially continuously during the operational phases of the digging element 126. As an example, the video may be at a frame rate of up to 3 frames per second (FPS), a resolution height of 1920 pixels, and a resolution width of 1188 pixels. In other examples, the captured video may be up to 10 FPS, up to 30 FPS, or up to 120 FPS and may have a resolution height ranging between 50 and 2000 pixels and a resolution length ranging between 50 and 3000 pixels, and only limited by the capabilities of the stereoscopic camera 190. In examples where the stereoscopic camera 190 captures images, one or more of the first, second, and third imagers 192, 194, and 198 may capture sequential images at specific time intervals to generate a plurality of sequential images at different points in time during an operational phase. The time intervals between sequential images may be 20 seconds, and the images may have a resolution height of 1920 pixels and a resolution width of 1200 pixels. In some examples, the time intervals may be intervals of nanoseconds, microseconds, milliseconds, and / or in seconds depending on the configuration. The images may have a resolution height ranging between 50 and 2000 pixels and a resolution length ranging between 50 and 3000 pixels. The frames of the video or the images of sequential images of the field-of-view 184 and any depth values calculated by the internal edge processor may then be transmitted by the stereoscopic camera 190 to the embedded processor circuit for subsequent processing.

[0055] In some examples, the visual sensing system 180 may be implemented as a single two-dimensional (2D) camera that captures sequential images or a video of the field-of-view 184 during the operating cycle of the mining shovel 102. In yet other examples, the visual sensing system 180 may include a camera, a radar, a LiDAR, and / or any combination thereof. Example Digging Element

[0056] FIG. 2 depicts an excavator bucket 200, including a plurality of ground-engaging teeth 210, lip shrouds 220, upper and lower wing shrouds 230, 240, heel shrouds 250, and a load container 260. The excavator bucket 200, similar to the load container 126 in FIG. 1, is configured for various digging and loading tasks in construction, mining, and other earth-moving operations.

[0057] The ground-engaging teeth 210 are components of the excavator bucket 200 that are removable and replaceable tips attached to the cutting edge of the excavator bucket 200. The ground-engaging teeth 210 are designed to enhance the digging performance and improve the ability of the excavator bucket 200 to break through materials. The ground-engaging teeth 210 can come in various shapes, such as pointed, flat, or spade-like, depending on the application of the excavator bucket 200. As such, the ground-engaging teeth 210 can be changed or replaced to accommodate different digging conditions or material types.

[0058] The lip shrouds 220 are components of the excavator bucket 200 that are removable and replaceable covers for the cutting edge of the excavator bucket 200. The lip shrouds 220 can be positioned along the lower most edge of the excavator bucket 320 that comes in contact with the ground or material being excavated. The lip shrouds 220 can be segmented between the ground-engaging teeth so as to reinforce the cutting edge and facilitate efficient digging.

[0059] The upper and lower wing shrouds 230, 240 are components of the excavator bucket 200 that are additional plates or edges attached to the side of the excavator bucket 200 positioned perpendicular to the cutting edge to provide protection to the sides against wear, impact, and damage. The upper and lower wing shrouds 230, 240 can provide extra protection to the sides of the excavator bucket 200 and improve its cutting ability. Additionally, the upper and lower wing shrouds 230, 240 can help prevent wear and damage to the excavator bucket 200 and also aid in defining the profile of the excavator bucket 200 during excavation.

[0060] In addition to protection, the upper and lower wing shrouds 230, 240 can provide additional structural reinforcement to the excavator bucket 200. The upper and lower wing shrouds 230, 240 can act as braces, reinforcing the side walls of the excavator bucket 200 and preventing flexing or bending under heavy loads. This added strength increases the overall durability and robustness of the excavator bucket 200, allowing it to withstand demanding applications.

[0061] The heel shrouds 250 are components of the excavator bucket 200 coupled to the base of the excavator bucket 200 and used to protect the heel area from excessive wear. The heel is the curved portion at the back of the bucket that connects to the bucket linkage or stick of the excavator. This area may be subject to intense abrasion and impact when digging and loading materials. Heel shrouds 250 act as a shield, preventing direct contact between the bucket's heel and the ground or other abrasive surfaces, thereby reducing wear and extending the lifespan of the excavator bucket 200. Operational Phases of a Digging Cycle

[0062] FIGS. 3A-F illustrate the cable mining shovel 102 of FIG. 1 performing the operational phases of a digging cycle, but similar cycles can be performed by loaders, back hoes, drag-lines, hydraulic shovels, and other non-cable type buckets. In operation, the operator, as shown in FIG. 1, within the operator station 121 of the frame 120 will cause the mining shovel 102 to go through a series of activities that make up a digging cycle of the mining shovel 102. The digging cycle may include at least the following operational phases: (1) excavate stage, (2) swing full stage, (3) dump stage, (4) swing empty stage, and (5) idle stage. The series of activities may be repeated for a plurality of digging cycles of the mining shovel 102. For example, the mining shovel 102 may perform a first excavate stage to initiate a first digging cycle, and then continue to a first swing full stage, a first dump stage, and a first swing empty stage to complete the first operating cycle, and then back to a second excavate stage to initiate a second operating cycle. In other examples, the operating cycle of the mining shovel 102 may include additional, fewer or alternative stages. For example, in certain situations, the operating cycle may include a bench cleaning subcycle with only (1) the excavate stage and (2) the dump stage. Excavate Stage

[0063] During the excavate stage, the operator will control the mining shovel 102 to excavate earthen material from the mine face 114, as shown in FIG. 3B. The excavate stage itself may include at least the following substages: (a) engage substage, (b) excavate substage, and (c) release substage.

[0064] During the engage substage, the handle 125 may be positioned in a lowered configuration (as shown in FIG. 1) such that the load container 154 may be positioned toward the mining area 114. To transition to the lowered configuration, the operator may control the mining shovel 102 to cause at least one of the frames and implement motorized spools 144 and 146 to extend the at least one cable 142 to lower the digging element 126, while causing the motorized pulley 140 to translate the mounting point between the boom 124 and the handle 125 towards the proximal end 138 of the handle 125.

[0065] During the excavate substage, the handle 125 transitions from the lowered configuration (also shown in FIG. 1) to a substantially horizontal configuration shown in Figure 3B (with a large angle relative to the y-axis and a small angle relative to the x-axis), which extends the digging element 126 into the mining area 114 to excavate the earthen material therefrom. To transition to the substantially horizontal configuration, the operator may control the mining shovel 102 to cause at least one of the frame and implement motorized spools 144 and 146 to retract the at least one cable 142 to raise the digging element 126, while causing the motorized pulley 140 to translate the mounting point between the boom 124 and the handle 125 towards the distal end 136 of the handle 125.

[0066] During the release substage, the handle 125 may be positioned in a raised configuration to maintain the load material 156 in the load container 154. To transition to the raised configuration, the operator may control the mining shovel 102 to cause at least one of the frame and implement motorized spools 144 and 146 to further retract the at least one cable 142 to further raise the digging element 126, while causing the motorized pulley 140 to further translate the mounting point between the boom 124 and the handle 125 towards the distal end 136 of the handle 125. Swing Full Stage

[0067] The swing full stage, as shown in FIGS. 3C and 3D, occurs after the excavate stage when the load 156 is within the load container 154. During the swing full stage, the operator may control the mining shovel 102 to rotate the load container 154 laterally away from the mine face 114 and towards a desired dump location. This may be done by rotating the frame 120 relative to the track 122 about the y-axis until the load container 154 containing the load 156 is positioned above the desired dump location. In the example shown in FIG. 3D, the desired dump location is a truck bed 171 of a haul truck 170. In such examples, the operator may also control additional components of the control mechanism 128 or the track 122, or an operator of the haul truck 170 may control the positioning of the haul truck 170, to fine-tune the relative positions of the load container 154 and the truck bed 171. In other examples, the desired dump location is another location in the open pit mine, such as a conveyor belt or another earthen material pile. In such examples, the operator may also control additional components of the control mechanism 128 or the track 122 to fine-tune the relative positions of the load container 154 and the desired dump location. Dump Stage

[0068] The dump stage, as shown in FIG. 3E, may occur after the swing full stage when the load 156 within the load container 154 may be positioned above the desired dump location. During the dump stage, the operator may control the mining shovel 102 to unload the load 156 from the load container 154 to the desired dump location. For example, the operator may control the door control mechanism 149 to open the bottom wall 162 of the load container 154. Additionally, or alternatively, the operator may cause at least one of the frame and implement motorized spools 144 and 146 to extend the at least one cable 142 to rotate the load container 154 about the z-axis relative to the handle 125 to a dump position to unload the load 156. Swing Empty Stage

[0069] The swing empty stage, as shown in FIG. 3F, may occur after the dump stage when the load 156 is unloaded and the load container 154 is emptied. During the swing empty stage, the operator may control the mining shovel 102 to rotate the now empty load container 154 laterally away from the desired dump location and towards another position proximate the mining area 114 to extract the earthen material therefrom in a subsequent excavate stage. The operator may initially control the door control mechanism 149 to close the bottom wall 162, and / or cause at least one of the frame and implement motorized spools 144 and 146 to retract the at least one cable 142 back from the dump position. Then, the operator may rotate the frame 120 relative to the track 122 about the y-axis until the load container 150 is positioned proximate the mining area 114 again. The operator may also control the track 122 to move the entire frame 120 to fine-tune the relative positions of the load container 126 and the mining area 114 for the next excavate stage. Idle stage

[0070] During each stage of the digging cycle, between the stages of the digging cycle or between a plurality of operating cycles, the operator may cause the mining shovel 102 to idle, as shown in FIG. 3 A, and not perform any specific activities which form the stages of the operating cycle described above, and / or be still. During the idle stage, the operator may be waiting for other equipment or personnel to be properly positioned. For example, during the dump stage, the operator may cause the mining shovel 102 to idle while waiting for the operator of the haul truck 170 to position the truck bed 171 below the load container 154. In another example, during the excavate stage, the operator may cause the mining shovel 102 to idle while waiting for personnel to reach a safe location. Machine Learning Framework

[0071] FIG. 5 illustrates a diagram of a machine learning framework 500 for bulk volume estimation and monitoring using machine learning models, in accordance with embodiments of the present disclosure. As described herein, embodiments use preprocessing techniques and feature extraction methods to produce relevant features extracted from image data to train machine learning models for bulk volume estimation and monitoring through scene understanding and operational phase detection. As shown, the machine learning framework 500 includes input data 502, a bulk volume estimation system 505, and an output 590. The bulk volume estimation system 505 includes a data preprocessor 510, a feature extractor 520, a scene understanding module 540, an operational phase detection module 550, a volume estimation module 560, and an aggregator 570. This combination of components enables the machine learning framework 500, through the bulk volume estimation system 505, to train and apply machine learning models to provide bulk volume estimations of materials excavated or held by a load container (e.g., shovel, bucket).

[0072] While previous volume estimation techniques utilize sensors to estimate bulk volume of materials, they merely rely on sensors placed on buckets that do not provide accurate estimations or may have rigid requirements in order to properly measure a bulk volume. Additionally, these methods are limited to the sensors being used by the heavy equipment. While previous other techniques also employed reference plane estimations, they merely fit a reference plane based on an exact bucket dimension and do not estimate pose or calculate a reference plane for each digging cycle based on key points of the buckets identified in images or via signal detection. In contrast, examples described herein utilize several advanced features extracted from images and / or mapped signals depicting a load container in operation, as described further down in more detail, to provide higher accuracy and adaptability when performing bulk volume estimations.

[0073] The input data 502 can include a sequence of images and / or a sequence of mapped signals representing the operational phases of a digging cycle performed by a digging element. The input data 502 can also include a sequence of images and / or a sequence of mapped signals representing a load container as material is being filled into the load container (e.g., a truck tray of a haul truck). The input data 502 can include stereoscopic images, red / green / blue (RGB) images, time-of-flight measurements, and / or radio wave or light measurements. Stereoscopic images generally refer to visual representations that create an illusion of depth by presenting slightly different perspectives. By presenting distinct perspectives, stereoscopic images can be used to measure depth and create a three-dimensional effect. As previously mentioned, stereoscopic images can be created using various techniques using a visual sensing system (e.g., visual sensing system 180), including dual-camera setups, machine vision sensors, image manipulation software, or specialized cameras that capture separate perspectives simultaneously.

[0074] The volume estimation system 505 is a component of the machine learning framework 500 configured to utilize datasets with relevant features extracted from the input data 502 to train and utilize machine learning models to provide volume estimations of material held by a load container. The volume estimation system 505 develops machine learning models used by the scene understanding module 530, the operational phase detection module 540, the feature segmentation module 550, and the volume estimation module 560. As such, the volume estimation system 505 provides volume estimations based on machine learning algorithms and techniques utilizing the relevant features from the input data 502 depicting a load container in operation.

[0075] The data preprocessor 510 is a component of the bulk volume estimation system 505 configured to preprocess the input data 502. Preprocessing transforms raw data into a usable format by removing anomalies such as noise, blue, occlusion, illumination variations, artifacts, and skewed or unrepresentative data. In some implementations, the data preprocessor 510 normalizes the images by scaling the pixel values of the input data 502 to a standardized range. For example, normalizing can be performed by subtracting the mean pixel value and dividing by the standard deviation or by rescaling the pixel values between 0 and 1. Normalization can help in reducing the impact of variations in lighting conditions and can enhance the convergence of the models during training.

[0076] In some implementations, the data preprocessor 510 applies denoising techniques the input data 502. Denoising techniques are used to reduce noise in images. Various algorithms such as Gaussian filtering, median filtering, and wavelet denoising can be employed to smooth out noisy pixels and improve image quality. Denoising can enhance the clarity of the images, making it easier for the models to extract relevant features.

[0077] In some implementations, the data preprocessor 510 applies image augmentation to the input data 502. Image augmentation techniques can artificially increase the size of a training dataset (e.g., the input data 502) by applying random transformations to the images. These transformations can include rotations, translations, scaling flipping, cropping, and changes in brightness and contrast. Augmentation can help in increasing the diversity of a training dataset and makes the models more robust to variations and anomalies.

[0078] In some implementations, the data preprocessor 510 applies image cropping and resizing techniques to the input data 502. Cropping and resizing images to a consistent size can help standardize the input dimensions for the models. This can be particularly useful when dealing with images of different sizes or aspect ratios and allows the volume estimation system 505 to be deployed on various heavy equipment using various image capturing systems. Cropping can also remove unwanted background or irrelevant parts of the image, focusing the attention on the relevant features.

[0079] In some implementations, the data preprocessor 510 performs feature extraction and feature selection on the input data 502 as part of the preprocessing process. Feature extraction involves computing a reduced set of values from a high-dimensional image capable of summarizing most of the information contained in the image. Feature extraction techniques develop a transformation of the input space onto the low-dimensional subspace that attempts to preserve the most relevant information. In feature selection, input dimensions that contain the most relevant information for solving a particular problem are selected. These methods aim to improve performance, such as estimated accuracy, visualization, and comprehensibility. An advantage of feature selection is that important information related to a single feature is not lost, but if a small set of features is required and original features are very diverse, there is chance of information being lost as some of the features must be omitted. On the other hand, with dimensionality reduction, also known as feature extraction, the size of the feature space can often be decreased without losing information about the original feature space.

[0080] In some implementations, the data preprocessor 510 applies feature extraction techniques to perform some transformation of original features to generate other features that are more significant. These feature extraction techniques include, but are not limited to Minimum Redundancy Maximum Relevance (“mRmR”), Relief, Conditional Mutual Information Maximization (“CMIM”), Correlation Coefficient, Between-Within Ratio (“BW-ratio”), Interact, Genetic Algorithms (“GA”), Support Vector Machine-Recursive Feature Elimination (“SVM-REF”), Principal Component Analysis (“PCA”), Non-Linear Principal Component Analysis, Independent Component Analysis, and Correlation based feature selection. These feature extraction techniques are useful for machine learning because they can reduce the complexity of input data and give a simple representation of data representing each variable in feature space as a linear combination of the original input variable.

[0081] In some implementations, the data preprocessor 510 performs depth processing on the sequences of input data 502. In some examples, the input data 502 may include associated depth values from an external source (e.g., an edge processor associated with a stereoscopic camera). In such examples, the data preprocessor 510 may not initiate any depth processing techniques or may only initiate a portion of the depth processing techniques. However, in some situations, the input data 502 may not have any associated depth values or may have associated depth values which are inaccurate. This may occur when image capture system imagers are partially or fully obscured by dust or debris during operation, have malfunctions due to extreme weather conditions, or dust and debris render the depth calculations inaccurate. In such examples, the data preprocessor 510 may initiate its own depth processing techniques to estimate depth values for each pixel of each image or each frame of the input data 502. The depth value is used in combination with the red, green and blue values of that pixel (combined “RGBD” values, [red, green, blue, depth] for example) in subsequent functions described below.

[0082] In some implementations, the data preprocessor 510 determines whether the images of the received sequential input data 502 (or frames of received video from the visual sensing system 180 includes 2D or 3D images, mapped images, or frames). For example, data preprocessor 510 may determine whether any depth values have been received and are associated with the input data 502. Alternatively, or additionally, data preprocessor 510 may determine whether one or more of a first image or frame or a second image or frame for a particular point in time is missing from the received sequential input data 502 or the received video. If one of the first image or frame or the second image or frame is missing, the data preprocessor 510 can determine that the data received for that particular point in time is two-dimensional data. Alternatively, or additionally, the data preprocessor 510 can determine whether the image capture system used to provide the input data 502 is a two-dimensional camera or a stereoscopic camera.

[0083] If the data preprocessor 510 determines that an initial image or frame is 2D data, the data preprocessor 510 generates depth values for that initial image or frame. For example, data preprocessor 510 can retrieve at least one sequential image or frame before and / or after the initial image or frame and identify at least one common feature in each image or frame. At least one sequential image or frame may include twenty sequential images, mapped images, or frames captured at time tl, t2, t3, and so on. In some examples, the sequential image or frame may include anywhere between 2 and 200 sequential images, mapped images, or frames captured at between 2 and 200 sequential points in time. Features of an image or a frame may be identified and described using a variety of different descriptors known to one or ordinary skill in the art, including using corner points of the images, mapped images, or frames, a scale-invariant feature transform (SIFT) detector based on reference images of an operating environment (e.g., the load container 126, the load 156, or the mine area 114), or a sped up robust features (SURF) detector for example. The data preprocessor 510 can track movement of the at least one common feature across the sequential images, mapped images, or frames to determine the depth of the pixels associated with the at least one common feature, based on estimating focal length relative to the visual sensing system 180 and radial distortion parameters for example.

[0084] Alternatively, or additionally, data preprocessor 510 can input the RGB values of the first image or frame, the second image or frame, and / or a third image or frame of the input data 502 into a depth estimator model. The depth estimator model can be trained with RGB values of two or more of: (a) first images or frames, (b) second images or frames, and (c) third images or frames, and corresponding camera parameters of a visual characterization system associated with the input data 502. The camera parameters, as shown in FIG. 4, include, but are not limited to, a first fixed distance between imagers 192 and 194, the second fixed distance 197 between the first and third imagers 192 and 198 and the third fixed distance 199 between the second and third imagers 194 and 198, as an example. The above RGB values and camera parameters are inputted into a neural network of the depth estimator model to iteratively generate and optimize coefficients which enable the neural network to estimate depth values based on an input of RGB values of one of: (a) the first image or frame captured by the first imager 192, (b) the second image or frame captured by the second imager 194 and (c) the third imager frame captured by the third imager 198.

[0085] In some example, the data preprocessor 510 can input a pair of monochrome images, with each image corresponding to the view from one of at least two cameras from the input data 502 into a depth estimator model configured as a neural network. The depth estimator model can be trained with a training dataset of monochrome images extracted from those images. Once trained, the depth estimator model can match the features between the two inputted images using techniques such as cost volumes, correlation layers, or other methods. The matching can assist the model in estimating the disparity between corresponding features in the two images. Once the depth estimator model has matched the features, it can calculate the disparity (i.e., the shift in position of the same feature when viewed from the two different viewpoints). Calculating the disparity can be performed by examining where the features align best in the two images, or it can involve further processing with additional network layers. Once the disparity is calculated, the depth estimator model can estimate the depth of each pixel or feature point based on the calculated disparity and the known geometry of the stereo camera setup. In some examples, the depth estimator model is configured as a StereoNet or Pyramid Stereo Matching Network (PSMNet) neural network capable of performing stereo depth estimation. In another example, a single imager may be color and the other two imagers may be monochrome and the values input to the neural network are the color image and the disparity between the monochrome images. In another example, the input into the neural network may be distance measurements taken from a LiDAR or radar to estimate depth values. In yet another example, a combination of the distance measurements and any combination of the images or disparity between them may be input into the neural network.

[0086] The data preprocessor 510 can store the determined depth value in association with the red, green and blue values (the “RGB” values, [red, green, blue] for example) of each pixel for a particular image or frame. For example, data preprocessor 510 can stack, or fuse or concatenate the depth value to the RRB values to generate “RGBD” values.

[0087] If data preprocessor 510 determines that the image or frame of the field-of-view includes 3D data, the data preprocessor 510 can generate depth values based on the 3D data. For example, the data preprocessor 510 can identify at least one common feature in each of the first image or frame captured by the first imager 192 (as shown in FIG. 4), the second image or frame captured by the second imager 194 or the third image or frame captured by the third imager 198 at a same point in time. Features of an image or a frame may be identified and described using a variety of different descriptors known to one or ordinary skill in the art, including the SIFT detector or the SURF detector, for example. For example, the data preprocessor 510 can calculate the relative displacement between at least one common feature in the first image versus the second image to determine the depth of the pixel associated with the at least one common feature, based on the calculated relative displacement and the fixed distance 195 between the first and second imagers 192 and 194. Additionally, or alternatively, the data preprocessor 510 can input the RGB values of both the first image or frame and the second image into the depth estimator model and then store the depth value as described above.

[0088] The feature extractor 520 may be a component of the volume estimation system 505 configured to extract relevant features from the input data 502. The feature extractor 520 can extract features using feature extraction techniques such as Scale-Invariant Feature Transform (SIFT), Speeded Up Robust Features (SURF), Oriented FAST and Rotated BRIEF (ORB), Histogram of Oriented Gradients (HOG), as well as machine learning techniques (e.g., convolutional neural networks (CNNs). The feature extraction techniques generally involve transforming each image of the input data 502 or each frame of a received video, or point cloud, into one or more feature maps operable to be used for subsequent key point identification, activity classification, feature segmentation, and bulk volume estimation described below. The feature extraction techniques thus form a backbone encoder for subsequent functions of the bulk volume estimation system 505.

[0089] In some implementations, the feature extractor 520 applies feature extraction techniques that derive features from a two-dimensional or three-dimensional point cloud (e.g., LiDAR data). These feature extraction techniques include geometric-based feature extraction, spin images, Point Feature Histograms (PFH) and Fast Point Feature Histograms (FPFH), Signature of Histograms of Orien Tations (SHOT), eigenvalue-based descriptors, and machine learning techniques. For example, machine learning techniques can utilize models such as PointNet and VoxNet take raw point cloud data and learn to extract complex features such as voxel-based features, point count, mean / max / min values, density, and the like. In another example, PFH and FPFH techniques can extract histogram-based features such as histogram counts. Histogram counts how many points fall into different bins. The bins could be defined based on intensity, height, or other properties. This gives a statistical summary of the distribution of those properties.

[0090] In some implementations, the feature extractor 520 applies feature extraction techniques that down sample the received image or frame to a smaller resolution height and a smaller resolution width, such as by using box down sampling or bilinear down sampling for example. For example, an image or a frame may be down sampled to a resolution height of 400 pixels and a resolution width of 640 pixels. However, it should be noted, the down-sampled resolution height and width may vary, and the resolutions described are for illustrative purposes only.

[0091] The feature extraction techniques are further configured to process the RGB values of each pixel for each image of the input data 502, or frame, using a feature extractor model configured to output one or more feature maps representing that image or frame. For example, each feature map may represent and report on a different feature or pattern in the image or frame, such as straight lines, curved lines, horizontal lines, vertical lines, shadow, light and may emphasize different components of image or frame, such as the teeth 150, the shrouds 152, and the lip 160, the load 156, or the mining area 114 (or earthen materials). Generally, the feature maps generated by the feature extractor model depend on convolutional layers described below and are not initially specified. In some implementations, the feature extractor model originally derives from and is adapted from an existing feature extraction model, such as U-Net, DenseNet, Segnet, Autoencoder, ResNet, and the like.

[0092] As an example, the feature extractor model of the feature extractor 520, can be based on the ResNet architecture and includes a first convolutional layer which operates directly on the RGB values of each pixel of the image or frame. The first convolutional layer may have kernel size of 3x3, stride of 2, filter size of 64 and may be followed by batch normalization (BN) and rectified linear activation function (ReLu) operations to generate a tensor shape of [IB, 200H, 320W, 64C], ‘B’ refers to batch size relating to the number of samples or data points processed simultaneously in a single forward or backward pass through a neural network. In this case, the batch size is 1 billion (IB), indicating that there are a billion samples or data points in the tensor. ‘H’ represents height referring to the vertical dimension or the number of rows in the tensor. In this case, the height is 200, meaning that the tensor has 200 rows. ‘W’ represents width relating to the horizontal dimension or the number of columns in the tensor. In this case, the width is 320, indicating that the tensor has 320 columns. ‘C’ represents channels relating to the number of channels or features in each element of the tensor. It is commonly used in computer vision tasks to represent different image channels, such as RGB channels. In this case, the tensor has 64 channels.

[0093] The BN operation generally functions to normalize the output tensor of the first convolutional layer before it is used as input for subsequent layers. The ReLu operation generates an activation function for the output tensor of the first convolutional layer and can reduce premature convergence of the feature extractor model. The output tensor of the first convolutional layer is inputted into a second convolutional layer which may have kernel size of 3x3, stride of 1, filter size of 64, and may be followed by corresponding BN and ReLu operations to generate a tensor shape of [IB, 200H, 320W, 64C], The output tensor of the second convolutional layer is inputted into a third convolutional layer which may have kernel size of 3x3, stride of 1, filter size of 128 and may be followed by corresponding BN and ReLu operations to generate a tensor shape of [IB, 200H, 320W, 128C], A maxpooling layer operates on the output tensor of the third convolutional layer and may have pool kernel size of 3x3, stride of 1 and pad of 1 and provides a flattened tensor to input into subsequent resnet blocks.

[0094] The feature extractor model can further includes a first resnet block and a second resnet block. The first resnet block 4may include fourth, fifth and sixth 2D convolutional layers. The fourth convolutional layer initially operates on the output tensor of the maxpooling layer and may have kernel size of 1x1, stride of 1, filter size of 64 and may be followed by corresponding BN and ReLU operations to generate a tensor of [IB, 100H, 160W, 64C], The output tensor of the fourth convolutional layer is inputted into the fifth convolutional layer which may have kernel size of 3x3, stride of 1, filter size of 64 and may be followed by corresponding BN and ReLU operations to generate a tensor of [IB, 100H, 160W and 64C], The sixth convolutional layer operates on both the output tensor of the maxpooling layer and the output tensor of the fifth convolutional layer, may have kernel size of 1x1, stride of 1, and filter size of 256 and may be followed by a corresponding BN operation to generate a tensor of [IB, 100H, 160W, 256C], The output tensor of the sixth convolutional layer based on the maxpooling layer may be summed with the output tensor of the sixth convolutional layer based on the fifth convolutional layer and may be followed by a ReLu operation. The summed output tensor of the sixth convolutional layer may be provided back to the fourth convolutional layer and the sixth convolutional layer to repeat the first resnet block. In the example shown, the first resnet block is repeated two more times. However, it should be noted in some examples, the first resnet block can be repeated fewer or a greater number of times based on a depth required for the feature extractor model.

[0095] The second resnet block operates on the output tensor of the final iteration of the first resnet block and may have a layer architecture substantially similar to the first resnet block except that one of the corresponding fourth and fifth convolutional layers of the second resnet block may have filter size of 128 and stride of 2 to generate tensors of [IB, 50H, 80W, 128C]; while the corresponding sixth convolutional layer of the second resnet block may have filter size of to generate a tensor of [IB, 50H, 80W, 512C], In the example shown, the second resnet block is repeated four times. However, in some examples, the second resnet block can be repeated fewer or a greater number of times based on the depth required. The output tensor of the final iteration of the second resnet block may be used for subsequent key point identification, activity classification, feature segmentation, and volume estimation functions described below.

[0096] The scene understanding module 530 may be a component of the volume estimation system 105 configured to identify and classify specific key points and landmarks in images and frames located in the input data 502. The key points and landmarks identified by the scene understanding module 530 can be used for subsequent activity classification, status analysis, and volume estimation described below.

[0097] The scene understanding module 530 is configured to utilize machine learning techniques that ingest samples (e.g., two dimensional images, a sequence of two dimensional images, or a three dimensional point cloud) and their associated features extracted by the feature extractor 520 and output key points and landmarks within the image data. As an example, the scene understanding model 530 can sample four sequential images, mapped images, or frames captured at four sequential points in time tl, t2, t3 and t4. Alternatively, or in addition, the scene understanding module 530 can sample between 2 and 200 sequential images, mapped images, or frames correspondingly captured at between 2 and 200 sequential points in time. The scene understanding module 530 can also input the features (generated by the feature extractor 520) corresponding to this subset of images, mapped images, or frames and output key point coordinates and key point labels for different key points in the subset of images, mapped images, or frames. As used herein, “key points” are computer-generated labels for at least one pixel of the image or frame which function to label components of a feature of the image or frame. The key point labels outputted by the scene understanding module 530 generally include components associated with the digging element 126, including, and without limitation, “tooth top,” “tooth bottom,” “shroud left,” “shroud right,” and “shroud top,” or any other component of a digging element 126 as described herein. These components can be seen as ‘landmarks’ of the digging element 126. The key point coordinates outputted by the scene understanding module 530 may include the x and y coordinates of at least one pixel corresponding to the identified key point in each image or frame. In some examples, the key point labels outputted by the scene understanding module 530 also include components associated with a truck tray of a haul truck, including, and without limitation, “truck tray top,” “truck tray bottom,” “tailgate,” “side walls,” “bulkhead,” “canopy,” “liners,” and “railing.”

[0098] In some implementations, the key point label “tooth top” may represent the center of the leading edge of the tip of each tooth of the teeth of the digging element 126, while the key point label “tooth bottom” may represent a center of a trailing edge of an intermediate adapter of each tooth. In some examples, the key point labels “tooth top” and “tooth bottom” may be divided into unique key point labels associated with an individual tooth of the teeth. In such examples, the key point labels may include “tooth top 1”, “tooth top 2”, “tooth top 3” up to “tooth top n,” and “tooth bottom 1”, “tooth bottom 2”, “tooth bottom 3” up to “tooth bottom n,” where n is the number of teeth making up the plurality of teeth.

[0099] As an example, a key point label “shroud left” may represent a left side of a leading edge of a shroud of the shroud, while the key point label “shroud right” may represent the right side of the leading edge of the shroud. The key point label “shroud top” may represent the center of the leading edge of the shroud, and in examples including the key point label “shroud bottom,” this key point label may represent a center of a trailing edge of the shroud. In some examples, the key point labels “shroud left,” “shroud right,” “shroud top,” and optionally “shroud bottom” may be divided into unique key point labels associated with individual shrouds of the shrouds. In such examples, the key point labels may include “shroud left 1”, “shroud left 2”, “shroud left 3” up to “shroud left m,”; “shroud right 1”, “shroud right 2”, “shroud right 3” up to “shroud right m”; “shroud top 1”, “shroud top 2”, “shroud top 3” up to “shroud top m” and “shroud bottom 1”, “shroud bottom 2”, “shroud bottom 3” up to “shroud bottom m,” where m is the number of shrouds making up the plurality of shrouds 152.

[00100] In some implementations, the scene understanding module 530 trains an associated machine learning model with the RGB values of each pixel of training images or frames, where the training images or frames include key point ground truth annotations such as key point labels “tooth top,” “tooth bottom,” “shroud left,” “shroud right,” “shroud top,” x and y coordinates of the key points, and optionally the visibility of the key point (such as with [key point label, x, y, z, visibility] for example). The above RGB values and key point ground truth annotations are inputted into a feature extraction model of the feature extractor 520 to generate one or more feature maps (output tensor of [IB, 50H, 80W, 512C]).

[00101] In some implementations, the feature maps are inputted into a neural network of the scene understanding module 530 to iteratively generate and optimize coefficients which enable the neural network to identify key point labels and generate corresponding key point coordinates for different subsets of images, mapped images, or frames based on an input including the feature maps of the subset of images, mapped images, or frames (e.g. based on an initial input of the RBG values of pixels of the subset of images or frames). The scene understanding model 530 can also be trained to generate the key point labels and key point coordinates based on alternative or additional inputs, such as the depth value of each pixel, time of flight measurement, or displacement of common features across the subset of images, mapped images, or frames. In such examples, the training images or frames may also be labeled with the depth values or displacement values.

[00102] In some examples, the scene understanding module 530 may originally derive from and be adapted from an existing key point identifier model such as KeypointNet, Keypoint RCNN, or PoseNet described in Bin XIAO et al., Simple Baselines for Human Pose Estimation and Tracking, arXiv: 1804.06208 [cs.CV] (2018), incorporated by reference herein.

[00103] The operational phase module 540 is a component of the volume estimation system 105 configured to identify and classify the operational phase of a load container detected in images, mapped images, and frames from the input data 502. The operational phases identified by the operational phase module 540 can be used for subsequent activity classification, status analysis, and volume estimation described below.

[00104] The operational phase module 540 is configured to estimate the operational phase of a load container during a digging cycle and identify a particular subset of images, mapped images, or frames reflecting the operational phase. As an example, the operational phase module 540 can retrieve a subset of images, mapped images, or frames captured over a period of time. These images can include sequential images, mapped images, or frames captured at twenty sequential points in time tl, t2, t3 [..] t20 from the input data 502. Alternatively, or additionally, samples can be anywhere between 2 and 200 sequential images, mapped images, or frames correspondingly captured at between 2 and 200 sequential points in time.

[00105] The operational phase module 540 is further configured to utilize machine learning techniques to output activity labels representing the different operational phases of the digging cycle to be associated with the subset of images, mapped images, or frames. For example, the operational phase module 540 can input the feature maps (generated by the feature extractor 530) corresponding to the subset of images, mapped images, or frames and output activity probabilities that the subset of images, image data, or frames corresponds to an activity label. In some examples, the operational phase module 540 can input the key point coordinates and the key point labels of the subset of images, mapped images, or frames and output activity labels. These activity labels include, without limitation, “background,” “excavate,” “swing full,” “dump truck,” “dump no truck,” “swing empty,” “idle,” “hauling, “maintenance,” and combinations and intermediaries of same, for example.

[00106] The activity probability outputted by the operational phase module 540 can also include a percentage for each of the activity labels, where the total percentage equals 1, such as [0, 0, 0.95, 0.02, 0.03, 0, 0] corresponding to the activity labels [“background,” “excavate,” “swing empty,” “dump truck,” “dump no truck,” “idle,” “swing full”] for example. As previously discussed, examples of various operational phases are described in FIG. 3.

[00107] In some examples, a neural network of the operational phase module 540 may be trained with the input data 502, where the training images, mapped images, or frames are further ground truth annotated with at least the activity labels “background,” “excavate,” “swing full,” “dump truck,” “dump no truck,” “swing empty,” and “idle.” In examples including the activity labels “engage substage,” “excavate substage,” “release substage,” and “maintenance,” the training images, mapped images, or frames may also be ground truth annotated with these activity labels as well. The above RGB values and activity ground truth annotations are then inputted into the feature extractor 520 to generate one or more feature maps (an output tensor [IB, 50H, 80W, 512C]). The feature maps or a sequence of the feature maps is inputted into a neural network of the operational phase module 540 to iteratively generate and optimize coefficients which enable the neural network to generate corresponding activity probabilities of the activity labels based on an input, including the feature maps of the subset of images, mapped images, and frames (e.g. based on an initial input of the RGB values of pixels of the subset of images or frames). As described above, in some examples, the key point coordinates and the key point labels of each image or frame may also be used by the operational phase module 540 to generate the activity probabilities of the activity labels. In other examples, the operational phase module 540 may also be trained to generate the activity probabilities of the activity labels based on alternative or additional inputs, such as the depth value of each pixel or displacement of common features across the subset of images, mapped images, or frames. In such examples, the training images, mapped images, or frames may also be labeled with the key point labels and the key point coordinates, the depth values and / or the displacement values. In some examples, the operational phase module 540 may originally derive from and be adapted from an existing spatial temporal activity classification model, such as the system described in Du TRAN et al., A Closer Look at Spatiotemporal Convolutions for Action Recognition, arXiv:1711.11248v3 [cs.CV] (2018), incorporated by reference herein.

[00108] The operational phase module 540 may be further configured to assign the subset of images, mapped images, or frames classified with an activity label to a particular operational phase. Based on a normal operational phase of the load container, each operational phase may include a plurality of subsets of images, mapped images, or frames labeled with “excavate,” followed by a plurality of subsets of images, mapped images, or frames labeled with “swing full,” followed by a plurality of subsets of images, mapped images, or frames labeled with “dump truck” or “dump no truck” and finally followed by a plurality of subsets of images, mapped images, or frames labeled with “swing empty,” before another operational phase may be initiated with another plurality of subsets of images, mapped images, or frames labeled with “excavate.” Thus, operational phase module 540 can identify each operational phase by determining transition points to “excavate,” where all images, mapped images, and frames falling before a transition point to “excavate” may be assigned to a previous operating cycle, and all images, mapped images, and frames falling after the transition point may be assigned to a current operating cycle.

[00109] The feature segmentation module 550 may be a component of the volume estimation system 505 configured to identify and classify specific features in each image, frame, or sequence of images or frames to generate feature segmentations operable for use in subsequent status analysis functions as well as determining a three dimensional profile of the material held by a load container. The feature segmentation module 550 is configured to input the feature maps corresponding to the subset of images, mapped images, or frames from the input data 502 into a machine learning model and have the machine learning model output feature labels for pixels of that subset of images, mapped images, or frames. In some examples, the machine learning model of feature segmentation module 550 can input the key point coordinates and the key point labels of the subset of images, mapped images, or frames (generated by the scene understanding module 530). The feature labels output by the feature segmentation module 550 may include, without limitation, “teeth,” “shroud,” and “load.” In some examples, the feature labels may also include “mining area.” A plurality of pixels having a common feature label are generally segmentations associated with that feature, such as a “teeth” segmentation or a “load” segmentation for example.

[00110] The feature label “teeth” (or the “teeth” segmentation) may generally be associated with pixels corresponding to portions of the teeth of the load container visible in the image or frame. In some examples, the feature label “teeth” may be divided into unique feature labels associated with the individual tooth of the teeth 150, such as “tooth 1”, “tooth 2”, “tooth 3” [...] up to “tooth n,” where n is the number of teeth making up the plurality of teeth.

[00111] The feature label “shroud” (or the “shroud” segmentation) may generally be associated with pixels corresponding to portions of the shrouds of the load container visible in the image or frame. Similar to the feature label “teeth,” in some examples, the feature label “shroud” may be divided into unique feature labels associated with individual shrouds of the shrouds 152, such as “shroud 1”, “shroud 2”, “shroud 3” [...] up to “shroud m”, where m is the number of shrouds making up the plurality of shrouds.

[00112] The feature label “load” (or the “load segmentation”) may generally be associated with pixels corresponding to portions of a load of materials within a load container of a load container that are visible in the image or frame. An example of a “load” segmentation is shown in FIG. 8. The feature label “mining area” (or the “mining area segmentation”) may generally be associated with pixels corresponding to portions of a mining area that are visible in the image or frame. The load segmentation may take into account the various different types or fragmentations of the load as shown.

[00113] In some examples, the feature segmentation module 550 implements a neural network. The neural network can be trained with the RGB values of each pixel of the input data 502, where pixels of the training images, mapped images, or frames are further feature ground truth annotations with at least the feature labels “teeth,” “shroud,” and “load,” (and optionally “mining area”). The above RGB values are inputted into the feature extractor 520 to generate one or more feature maps (output tensors of [IB, 50H, 80W, 512C]). The feature maps are then inputted into the neural network of the feature segmentation module 550 to iteratively generate and optimize coefficients which enable the neural network to generate corresponding feature labels for different subsets of images, mapped images, or frames based on an input, including the feature maps of the subset of images, mapped images, or frames (e.g. based on an initial input of the RGB values of pixels of the subset of images or frames). In some examples, the feature segmentation module 550 may also generate the feature labels based on alternative or additional inputs, such as the depth value of each pixel or displacement of common features across the subset of images, mapped images, or frames. In such examples, the training images, mapped images, or frames may also be ground truth annotated with the key point labels and the key point coordinates, the depth values or the displacement values. In some examples, the feature segmentation module 550 may originally derive from and be adapted from an existing spatial-temporal feature segmentation model, such as the system described in Huan FU et al., Deep Ordinal Regression Network for Monocular Depth Estimation, arXiv:1806.02446vl [cs.CV] (2018), which is incorporated herein by reference.

[00114] The volume estimation module 560 is a component of the bulk volume estimation system 505 configured to estimate the bulk volume of a load of material held by a load container. The bulk volume, or heap volume, can refer to the volume of material that is held by a load container above the edge of the load container. As such, the bulk volume can be a measurement of the excess material held by a load container during an operational phase of digging cycle.. Considering t factors, such as, the shape of the bucket, the angle at which it is positioned, and the characteristics of the material being handled factor into the overall calculation of the bulk volume, the volume estimation module 560 may be configured to retrieve the load segmentation of the material (via the feature segmentation module 550) while the load container is in a swing full state (via the operation phase module 540) and determine the three dimensional profile of the material by applying depth processing mechanics to the region of interest of the material determined by the load segmentation. Once profiled from the feature segmentation module 550, the volume estimation module 560 can fit a reference plane to key points detected by the scene understanding module 530 and estimate the volume of the profiled shape in relation the pose, or normal vector, of the reference plane.

[00115] In some examples, the volume estimation module 560 includes a computer model (e.g., a conditional probability model) for estimating key points, or landmarks, on the load container that have may been occluded by the material being held. As such, the conditional probability model can estimate the probability of certain features, objects, or classes being present in an image and conditioned on the image data during swing empty states. The volume estimation module 560 can implement the conditional probability model in a variety of ways. These include, but are not limited to, Naive Bayes classifier, Bayesian networks (belief networks), Hidden Markov Models (HMMs), Conditional Random Fields (CRFs), Markov Random Fields (MRFs), Generalized Linear Models (GLMs), decision trees, random forests, Gradient Boosting Machines (GBMs), neural networks, Recurrent Neural Networks (RNNs), and Gaussian processes.

[00116] As an example, the conditional probability model can be based on an HMM. HMMs are statistical models used to represent systems that change over time according to a Markov process, where the states of the process are not directly observable (e.g., states of a load container with occluded key points or landmarks). Observations can be generated at each state according to some probability distribution that depends on the hidden state.

[00117] In some implementations, the volume estimation module 560 trains the HMM using a gradient descent method. Using training data taken from the input data 502 and pre-processed by the data preprocessor 510, the HMM can input the training images to train the HMM to, based on observed key points, output probabilities of key points, or landmarks, occluded by material (e.g., occluded during the “swing full” or “swing empty” state). During training, a sequence of images representing a “swing empty” state may be considered a state sequence, and another sequence of images representing a swing full state can be an observation sequence. In another example, during training, a sequence of images representing a “swing full” state may be considered a state sequence, and another sequence of images representing a “swing empty” state can be an observation sequence. This example may be useful in determining carryback of material during an empty state. This can also be used in aggregating the calculation of fill factor for the next load container load. According to these parameters, the HMM can determine the probability of a certain key point, or landmark, present at a certain point during the swing full state. The HMM may predict the location of unobserved key points based on observed key points and known key points from a CAD model, for example. Inputs to the HMM include landmark points (discrete points), which learns transitions from the points as the points change as the load container moves and the load shifts and changes. The HMM learns the appearance of the load container and loads over time.

[00118] The landmarks, such as the tooth, lip shroud, wing shrouds, and heel shrouds, can be identified and marked in both sequences of images (e.g., the swing empty, swing full, and the like), representing the different operational phases. The HMM can consist of two parts. The first part can be a Markov chain, and the second part can be a random process. As such, the HMM can be described using the following Equation 1: A = (n,S,T) Equation 1 Where A represents the HMM, tt, and S represent the Markov chain, and T represents the random process. Given the initial model A = (n,S, T) and a sequence of observations (e.g., the sequence of images of a load container in a swing empty state) represented as L = (Z1; l2,..., lb), the HMM can calculate the probability P(L|A). The probability can be seen as a matching problem between the HMM and the sequence of observations and can be solved using a forward-backward algorithm.

[00119] Additionally, during training, the HMM provides parameter estimation. As such, assuming A = (n,S, T) and L = (Z1; l2, —,lb\ use a maximum likelihood estimation method to calculate the HMM parameters A = (ji, S, T) so that the probability of generating certain observation sequences P(L|A) (e.g., occluded landmarks in the swing full state) reaches the maximum. This can be accomplished using the Baum-Welch algorithm. The Baum-Welch algorithm, also known as the forward-backward algorithm or the expectation-maximization (EM) algorithm for HMMs, is an iterative procedure used to estimate the parameters of an HMM from a set of observed data.

[00120] For predictions, given A = (n, S, T) and L = l2,..., 1^, the HMM finds a state sequence represented as K = klt k2, ...,kb when the conditional probability C of the observation sequence is maximized. The state sequence can be realized by optimizing the Viterbi algorithm. The Viterbi algorithm is a dynamic programming algorithm used to find the most likely sequence of hidden states in a HMM, given an observed sequence of data. In the HMM, give a sequence of observations L = (Z1; l2,..., 1^, and each observation may be generated by a hidden state. However, the hidden states themselves are not directly observable (e.g., occluded). The Viterbi algorithm aims to find the sequence of hidden states that maximizes the probability of generating the observed sequence.

[00121] In some implementations, the HMM continuously learns the model parameters from a set of training images taken of a load container during a swing empty state, as illustrated in FIG. 9. As such, the model parameters are continuously refined. During training, the training images can be pre-processed by the data preprocessor 510 and labeled with the corresponding key points or landmarks. These landmarks can also define the hidden states of the HMM as well as the observed symbols. These symbols can capture relevant information for the analysis task. The Baum-Welch algorithm can be applied, also known as the EM, as described above, to estimate the model parameters iteratively. This can involve multiple iterations of the EM steps. In the E-step, the forward-backward algorithm is used to compute the posterior probabilities of the hidden states given the observed symbols for each training image. In the M-step, the HMM parameters are updated based on the computed posterior probabilities.

[00122] During training, the E-step and the M-step are repeated until convergence criteria are met. In some examples, the convergence can be determined based on the change in loglikelihood or the parameters themselves. After convergence, the estimated model parameters represented the learned HMM for image analysis of the sequence of images representing a digging element in a swing-full state.

[00123] Refinement of the HMM can occur during each swing-empty state a digging element is in or set by some other pre-defined metric, such as after every ten operational phase cycles. The refinement can modify the set of hidden states or observed symbols or incorporate additional training data (e.g., an additional sequence of images taken of the digging element). This iterative process can help improve the accuracy and robustness of the HMM.

[00124] The volume estimation module 560 may be further configured to fit a reference plane using observed landmarks, CAD models and predicted landmarks that may be occluded during a swing full state of a digging element. Fitting a reference plane to a set of key points can involve the best-fitting plane that minimizes the overall distances between the points and the plane. The volume estimation module 560 can apply several techniques to fit the reference plane to the point. These techniques include, but are not limited to, least squares fitting, Principal Component Analysis (PCA), Random Sample Consensus (RANSAC), Hough Transform, geometric methods, convex hull, and the like.

[00125] As an example, the volume estimation module 560 uses the least squares fitting method to fit a reference plane to the key points observed and predicted from the images. Once the key points are observed and predicted, the volume estimation module 560 can compute the centroid of the reference plane. Calculation of the centroid includes calculating the centroid of the point cloud by taking the average of the x, y, and z coordinates of all the key points. Once the centroid is calculated, the volume estimation module 560 can calculate the covariance matrix of the point cloud. The covariance matrix can describe the distribution of the points around the centroid and provide information about the orientation of the plane. The volume estimation module 560 can further perform eigenvalue decomposition on the covariance matrix of the point cloud. This will yield the eigenvalues and eigenvectors of the matrix. The eigenvector can correspond to the smallest eigenvalue that represents the normal vector (e.g., the pose of the load container) of the plane. Once the normal vector (i.e., the pose) is determined, the volume estimation module 560 can define the reference plane equation using the centroid as the reference point. The resulting plane equation can represent the best-fit plane that minimizes the distance between the point and the plane in the least squares sense.

[00126] In some implementations, the volume estimation module 560 calculates the reference plane of the load container during a swing empty state and a swing full state. Using the knowledge gained while calculating the swing empty reference plane, specific characteristics of those planes, such as constraints on the orientation or known relationships between key points, the volume estimation module 560 can incorporate that information into the fitting process to improve the accuracy of the fitting reference plane.

[00127] The volume estimation module 560 may be further configured to calculate the bulk volume of the material held by a load container. As mentioned, feature segmentation model 550 is configured to retrieve the load segmentation of the material while the load container is in a swing full state and determine a three-dimensional profile of the material by applying depth processing mechanics to the region of interest of the material determined by the load segmentation. Once profiled, the volume estimation module 560 calculates a reference plane from a point cloud of key points detected by the scene understanding module 530, as well as key points predicted by the conditional probability model to estimate the volume of the profiled shape. The conditional model also predicts the normal vector of the plane. The volume estimation module 560 can then estimate the bulk volume of the material based on the three-dimensional profile of the material in relation to the reference plane.

[00128] In some implementations, the volume module 560 calculates the bulk volume of the material held by the load container using a discrete volume summation methodology. Discrete volume summation is a method used to estimate the volume of an object or a region by dividing it into small, discrete elements and summing their individual volumes. This approach is particularly useful when dealing with irregular shapes or when the exact mathematical representation of the object is not available.

[00129] As an example, using the discrete volume summation technique, the volume estimation module 560 can convert the profiled surface into a triangulated mesh representation. This can involve subdividing the profiled surface into a set of triangles. For each triangle in the triangulated mesh, the volume estimation module 560 can form tetrahedra by connecting its vertices with the point on the reference plane. The volume estimation module 560 can calculate the volume of each tetrahedron using the signed volume formula. The signed volume formula accounts for whether the tetrahedron is above or below the reference plane. The volume estimation module 560 can then sum the volumes of all the tetrahedra to obtain the total volume of the profiled material. When accounting for volume above and below a reference plane, the reference plane is often considered as the base or zero level. Volumes above this reference plane and volumes below this plane are calculated separately. Each discrete element of volume above the reference plane can contribute positively to the total volume. Each discrete volume can be calculated as the area of the section multiplied by the thickness of the layer, and all these are summed up to give the total volume above the reference plane. Each discrete element of volume below the reference plane contributes negatively to the total volume or can be considered as a separate negative volume. Similar to the volume above the plane, the volume for each section is calculated and summed up. By summing the volume of the tetrahedra, the volume estimation module 560 effectively approximates the bulk volume of the profiled material within the defined reference plane.

[00130] It should be noted that the reference plane can be viewed as a cutting plane to divide the material into two distinct regions. As such, the volume estimation module 560 uses methods like discrete volume summation or mesh-based volume calculation techniques to calculate the volumes of the separate regions in relation to the reference plane. Additionally, the normal vector, or pose, of the reference plane can help determine the orientation of the reference plane relative to the material.

[00131] The aggregator 570 may be a component of the bulk volume estimation system 505 configured to aggregate different metrics associated with different stages of the operational phases to identify trends over multiple operational phases or trends associated with a common stage (e.g., excavate stages, swing full stages, dump stages, swing empty stages, etc.) of multiple operational phases. In this regard, a “common” stage of multiple operating cycles refers to the same stage over multiple operational phases. For example, one operational phase may include one excavate stage and one swing full stage, and the excavate stages of a plurality of operational phases may be common excavate stages, while the swing full stages of multiple operating cycles may be common swing full stages. One example of an aggregated metric may be the duration of different common stages over multiple operating cycles, such as an average or median bulk volume of material in a swing full stage, for example. Another example aggregated metric may be an aggregated bulk volume calculation per operator. The aggregated metrics can be correlated with an identity of an operator, the composition of the material, and a composition of a particular mining area.

[00132] It may be noted that FIG. 5 is intended to depict the major representative components of a machine learning framework 500 and a volume estimation system 505. In some examples, however, individual components may have greater or lesser complexity than as represented in FIG. 5, components other than or in addition to those shown in FIG. 5 may be present, and the number, type, and configuration of such components may vary.

[00133] FIG. 6 illustrates an example heavy equipment monitoring system 600, in accordance with examples of the present disclosure. As discussed, the techniques described herein provide a bulk volume estimation of material held by a digging element during a swing full operational phase of heavy equipment. Accordingly, in some examples, heavy equipment monitoring tools (e.g., cameras) train machine learning models (e.g., neural networks, conditional probability models) associated with the bulk volume estimation system 505 to generate a bulk volume estimations report 630 that can include information such as bulk volume estimations of material for each operational phase, bulk volume estimations of material per haul truck 634, and bulk volume estimations of material per operator 632, digging element efficiency reports 636, and the like. As shown in FIG. 6, optionally, the bulk volume estimation system 505 may be implemented as part of the heavy equipment monitoring system 600. Alternatively, the bulk volume estimation system 505 may be implemented in a separate system that provides at least the output 590 of the bulk volume estimation system 505 to the heavy equipment monitoring system 600 for evaluation and aggregation.

[00134] Once the output 590, representing the bulk volume estimations of a load container, has been obtained by the heavy equipment monitoring system 600, the output can be evaluated by a user (e.g., supervisor, analyst). In some implementations, user input (e.g., operator information, haul truck specifications) can be received by a monitoring manager 604 of the heavy equipment monitoring system 600. The user input 602 can include any information regarding operations relating to the heavy equipment utilizing the load container, including the operator information, the type of material being excavated, the components attached to the load container, the haul truck capacity, and the like. The monitoring manager 604 can generate the bulk volume estimations report 630, including diagnostics known in the art, and may include the output 590 produced by the bulk volume estimation system 505 in combination with evaluating the user input 602.

[00135] In some implementations, the resulting bulk volume estimation output 590 generated by the bulk volume estimation system 505 may output to a different or downstream monitoring system for evaluation and responsive action. The responsive action may take any known or later developed form, including output, to a field supervisor, such as via a user interface, when the bulk volume estimations indicate lower material yields. The user interface may comprise user interface elements for drilling down into the details of the notification, including identifying the specific operators, load containers, and material being excavated. In this way, the supervisor may identify if adjustments need to be made to the components of the excavation equipment. In some examples, user interface elements may be provided to allow the supervisor to provide input to indicate the correctness / incorrectness of the bulk volume estimations of the material such that this information may be used to update training datasets for updating the models used by the bulk volume estimation system 505 at a future time.

[00136] FIG. 7 illustrates a schematic diagram of a bulk volume estimation system 700 (e.g., “the bulk volume estimation system 505” described above), in accordance with embodiments of the present disclosure. As shown, the bulk volume estimation system 700 includes, but may be not limited to, a user interface manager 702, a training manager 704, a data preprocessor 705, a machine learning component 706, and a storage component 708. The machine learning component 706 includes a feature extractor module 710 (which can be the same as, or substantially similar to, the feature extractor 520 of FIG. 5), a scene understanding module 711 (which can be the same as, or substantially similar to, the scene understanding module 530 of FIG. 5), an operational phase module 712 (which can be the same as, or substantially similar to, the operational phase module 540 of FIG. 5), a feature segmentation module 713 (which can be the same as, or substantially similar to, the feature segmentation module 550 of FIG. 5), and a volume estimation module 714 (which can be the same as, or substantially similar to, the volume estimation module 560 of FIG. 5). The storage component 708 includes input data 718, bulk volume estimations 720, and estimation reports 722.

[00137] As illustrated in FIG. 7, the bulk volume estimation system 700 includes a user interface manager 702. For example, the user interface manager 702 allows users to input data 718 to the bulk volume estimation system 700. In some examples, the user interface manager 702 provides a user interface through which the user can upload the input data 718 representing the operational phases of a load container, as discussed above. Alternatively, or additionally, the user interface may enable the user to download the input data 718 from a local or remote storage location (e.g., by providing an address (e.g., a URL or other endpoint) associated with an input image source). In some examples, the user interface manager 702 also enables the user to provide specific input images relating to an operational phase or load container to be evaluated and analyzed. Additionally, the user interface manager 702 allows users to request the bulk volume estimation system 700 to provide specific bulk volume estimation of material in a load container. For example, a user can request a single bulk volume estimation after every digging cycle. In some examples, the user interface manager 702 enables the user to edit the input data 718. For example, the user can perform pre-processing techniques such as labeling and operational phase identification. Alternatively, the input data 718 can be evaluated in a separate heavy equipment monitoring system separate from bulk volume estimation system 700, as discussed above.

[00138] As illustrated in FIG. 7, the bulk volume estimation system 700 also includes a training manager 704. The training manager 704 can teach, guide, tune, and / or train one or more machine learning models, neural networks, hidden Markov models, and the like. In particular, the training manager 704 can train a machine learning model based on a plurality of training data (e.g., input data 718). As discussed, the input images 478 includes sequences of images captured from a heavy equipment operating a load container during the operational phases of the load container. More specifically, the training manager 704 can access, identify, generate, create, and / or determine training input and utilize the training input to train and fine-tune the machine learning models described above. For instance, the training manager 704 can train the neural networks of the feature extractor module 710, the scene understanding module 710, the operational phase module 712, and the feature segmentation module 713. The training manager 704 can also train the HMM of the volume estimation module 714, as well as provide evaluation metrics, as discussed above. As discussed, the models are trained specifically for a specific purpose in some examples, but other configurations are possible, such as more than one purpose may be trained by a single neural network. For example, a neural network may be trained (or the existing neural network may be retrained) to identify and classify specific key points and landmarks in images and frames, and another neural network may be trained to identify and classify the operational phase of a load container detected in images and frames, and another neural network may be trained to identify and classify specific features in each image and to generate feature segmentations operable for use in subsequent status analysis functions.

[00139] As illustrated in FIG. 7, the bulk volume estimation system 700 also includes a data preprocessor 705 (e.g., “the data preprocessor 710” described above). As discussed, the data preprocessor 705 implements data preprocessing techniques to generate a training dataset for training machine learning models by the training manager 704. The data preprocessor 705 can generate a plurality of relevant features from the input data 718 to create training datasets for the machine learning models described above, respectively. The resulting training datasets are provided to the training manager 704 for training, as discussed.

[00140] As illustrated in FIG. 7, the bulk volume estimation system 700 also includes a machine learning component 706. The machine learning component 706 may host a plurality of machine learning models or other machine learning models, such as the neural networks of the feature extractor module 710, the scene understanding module 711, the operational phase module 712, and the feature segmentation module 713. The machine learning component 706 can also include the HMM of the volume estimation module 714. The machine learning component 706 may include an execution environment, libraries, and / or any other data needed to execute the machine learning models. In some examples, the machine learning component 706 may be associated with dedicated software and / or hardware resources to execute the machine learning models.

[00141] Although depicted in FIG. 7 as being hosted by a machine learning component 706, in various embodiments, the machine learning models may be hosted in multiple machine learning components and / or as part of different components. For example, a scene understanding manager can host the neural network of the scene understanding module 710. Similarly, an operational phase manager can host the neural network of the operational phase module 711. In various embodiments, the scene understanding manager and the operational phase manager can each include their own machine learning component, or other host environments, in which the respective machine learning models execute.

[00142] As illustrated in FIG. 7, the bulk volume estimation system 700 also includes the storage component 708. The storage component 708 maintains data for the bulk volume estimation system 700. The storage component 708 can maintain data of any type, size, or kind as necessary to perform the functions of the bulk volume estimation system 700. The storage component 708, as shown in FIG. 7, includes the input data 718. The input data 718 can include a sequences of images associated with various operational phases of a load container, as discussed in additional detail above. In particular, in one or more embodiments, the input data 718 include sequences of operational phase images utilized by the training manager 704 to train the plurality of machine learning models to generate bulk volume estimation of material held by a load container.

[00143] Each of the components 702-708 of the bulk volume estimation system 700 and their corresponding elements (as shown in FIG. 7) may be in communication with one another using any suitable communication technologies. It will be recognized that although components 702-708 and their corresponding elements are shown to be separate in FIG. 7, any of components 702-708 and their corresponding elements may be combined into fewer components, such as into a single facility or module, divided into more components, or configured into different components as may serve a particular example.

[00144] The components 702-708 and their corresponding elements can comprise software, hardware, or both. For example, components 702-708 and their corresponding elements can comprise one or more instructions stored on a computer-readable storage medium and executable by processors of one or more computing devices. When executed by the one or more processors, the computer-executable instructions of the bulk volume estimation system 700 can cause a client device and / or a server device to perform the methods described herein. Alternatively, the components 702-708 and their corresponding elements can comprise hardware, such as a special purpose processing device to perform a certain function or group of functions. Additionally, the components 702-708 and their corresponding elements can comprise a combination of computerexecutable instructions and hardware.

[00145] Furthermore, the components 702-708 of the bulk volume estimation system 700 may, for example, be implemented as one or more stand-alone applications, as one or more modules of an application, as one or more plug-ins, as one or more library functions or functions that may be called by other applications, and / or as a cloud-computing model. Thus, the components 702-708 of the bulk volume estimation system 700 may be implemented as a standalone application, such as a desktop or mobile application, or on an edge device associated with heavy equipment (e.g., a processor circuit onboard the heavy equipment). Furthermore, components 702-708 of the bulk volume estimation system 700 may be implemented as one or more web-based applications hosted on a remote server.

[00146] FIGS. 5-7, the corresponding text, and the examples provide a number of different systems and devices that enable bulk volume estimation using relevant features extracted from images, allowing for continuous volume estimations of material during the operational phases of a digging cycle. In addition to the foregoing, embodiments can also be described in terms of flowcharts comprising acts and steps to accomplish a particular result. For example, FIG. 10 illustrates a flowchart of an exemplary method in accordance with one or more embodiments. The method described in relation to FIG. 10 may be performed with fewer or more steps / acts, or the steps / acts may be performed in differing orders. Additionally, the steps / acts described herein may be repeated or performed in parallel with one another or in parallel with different instances of the same or similar steps / acts. Example Flow Diagram

[00147] With reference now to FIG. 10, a flow diagram is provided illustrating a method. Each block of the method 1000 and any other methods described herein comprise a computing process performed using any combination of hardware, firmware, and / or software. For instance, in some examples, various functions are carried out by a processor executing instructions stored in memory. In some cases, the methods are embodied as computer-usable instructions stored on computer storage media. In some implementations, the methods are provided by a standalone application, a service or hosted service (standalone or in combination with another hosted service), or a plug-in to another product, to name a few.

[00148] FIG. 10 illustrates a flowchart 1000 of a series of acts in a method of estimating a bulk volume of a material being held by a load container during a digging cycle, in accordance with embodiments of the present disclosure. In one or more embodiments, the method 1000 may be performed in a digital medium environment that includes the volume estimation system 700. The method 1000 may be intended to be illustrative of one or more methods in accordance with the present disclosure and may be not intended to limit potential examples. Alternative examples can include additional, fewer, or different steps than those articulated in FIG. 10.

[00149] As illustrated in FIG. 10, the method 1000 includes a step 1010 of collecting image data depicting a load container in operation during a digging cycle. Similar to the input images discussed above, image data can include stereoscopic images, mapped images, frames, and / or RGB images. As previously mentioned, stereoscopic images can be created using various techniques using a visual sensing system (e.g., visual sensing system 180), including dual-camera setups, machine vision sensors, image manipulation software, or specialized cameras that capture separate perspectives simultaneously.

[00150] The method 1000 also includes a step 1020 of determining an operational phase of the load container based on the input images. Through the use of an operational phase module 712, the method 1000 can utilize a machine learning model to analyze the input images and output activity labels representing the different operational phases of the digging cycle associated with a subset of the images. As discussed, the activity labels can include labels such “background,” “excavate,” “swing full,” “dump truck,” “dump no truck,” “swing empty,” “idle,” “hauling,” “maintenance,” and combinations and intermediaries of same to represent the various operational phases of the load container during a digging cycle.

[00151] For instance, the activity label “swing full” may be associated with frames or images corresponding to the load container holding material excavated from a mining area. The activity label “swing empty” may be associated with frames or images corresponding to when the load container is not holding material (e.g., upon initial use, and prior to excavation, after releasing material, etc.).

[00152] Upon detecting an operational phase state (for example, the “swing full” operational phase state), the method 1000 further includes a step 1030 of identifying key points of the load container from the input images. Utilizing a scene understanding module 711, the method 1000 can utilize a machine learning model to analyze the subset of input images depicting the load container in the detected operational state. From that analysis, the machine learning model can output key points identified in the subset of images, mapped images, or frames. As explained, key points are computer-generated labels for at least one pixel of the image or frame which function to label components of a feature of the image or frame. The key point labels outputted by the scene understanding module 711 can include components associated with the load container 126, including a “tooth top,” “tooth bottom,” “shroud left,” “shroud right,” and “shroud top,” or any other component of a load container 126 as described herein. The key point can also include coordinates output by the scene understanding module 711 that include the x, y, and z coordinates of at least one pixel corresponding to the identified key point in each image or frame.

[00153] The method 1000 may also include a step 1040 of estimating occluded key points of the load container from the images. For example, during a “swing full” state, the load container is holding material which may occlude certain landmarks associated with the load container, such as the teeth or lip shrouds. Using the volume estimation module 714, a computer model such as a conditional probability model (e.g., an HMM) can be used to estimate the occluded key points. Using at least one key point identified by the scene understanding module 711, the conditional probability model can estimate the probability of certain features such as the occluded key points being present in the image data and conditioned on the image data during swing empty states. In this instance, the conditional probability model is used to estimate the occluded key points.

[00154] Applying training data, the conditional probability model can be trained to output probabilities of key points or landmarks occluded by material during the “swing full” state or other operational states. For instance, let a sequence of images representing a “swing empty” state be a state sequence, and another sequence of images representing a “swing full” state be an observation sequence. According to these parameters, the conditional probability model can determine the probability of a certain key point or landmark present at a certain point during the swing full state.

[00155] The method 1000 further includes a step 1050 of to calculate a three dimensional profile of the material. As mentioned, feature segmentation module 713 is configured to retrieve the load segmentation of the material while the load container is in the “swing full” state and determine a three-dimensional profile of the material by applying depth processing mechanics to the region of interest of the material determined by the load segmentation. As an example, the feature segmentation module 713 can utilize a machine learning model (e.g., a neural network) that inputs the key point coordinates and the key point labels of the subset of images, mapped images, or frames (generated by the scene understanding module 711). The feature labels outputted by the feature segmentation module 713 may include a plurality of pixels having a common feature label known as segmentations associated with that feature, such as a “teeth” segmentation 713a or a “shroud” segmentation 713b or a “load” segmentation 713c for example.

[00156] In some examples, the input images may include associated depth values from an external source (e.g., an edge processor associated with a stereoscopic camera). In such examples, the associated depth values can be applied to the load segmentation indicating a region of interest of the material. The resulting combination provides the three-dimensional profile of the material. In some examples, a data preprocessor can analyze the images and apply depth processing techniques, as described herein to generate the resulting three-dimensional profile of the material.

[00157] The method 1000 also includes a step 1060 of fitting a model with the key points and the occluded key points to the three dimensional profile of the material to determine pose. In some embodiments, this includes fitting a reference plane to the load container based on a normal vector associated with the key points and occluded key points. The volume estimation module 714 can calculate a reference plane from a point cloud of the key points detected by the scene understanding module 711, as well as key points predicted by the conditional probability model. From the point cloud, a normal vector (i.e., the pose of the load container) can be calculated and then used to fit a reference plane to the load container. The volume estimation module 714 can utilize various algorithms and techniques to fit the reference plane and can vary depending on the software or library being used. These methods and techniques include, but are not limited to, Random Sample Consensus (RANSAC), principal component analysis (PCA), and robust regression techniques.

[00158] For example, the volume estimation module 714 can use the Random Sample Consensus (RANSAC) algorithm to fit a plant to a 3D point cloud where it iteratively selects a subset of points, fits a plane using the least squares method, and then evaluates the quality of the fit by counting the number of points that lie within a certain threshold distance from the plane. The algorithm may be repeated for a fixed number of iterations, selecting the best fitting plan, which maximizes the number of inliers. The module may use least squares method to fit a plane to the subset of 3D points, where it estimates an affine transformation in the form of A.X + b where A is an nxl vector of constants and b is a scalar constant. A and b define the parameters of a hyper plane that best fits the 3D points, which minimizes the sum of squared errors between the point coordinates and their projections on the plane. Therefore, the gradients of the objective function are zero at the optimal point. Standard numerical algorithms can be used to find the arguments, which yield the zero gradients.

[00159] The method 1000 further includes a step 1070 of estimating the bulk volume of the material based on the three-dimensional profile of the material in relation to the reference plane. As an example, the volume estimation module 714 can calculate the bulk volume of the material held by the load container using a discrete volume summation methodology. Discrete volume summation, also known as discrete volume integration or voxelization, is a technique used to calculate the volume of a three-dimensional object or region by discretizing it into a grid of small volumetric elements called voxels. By discretizing the object into voxels and summing the volumes of the "inside" voxels, discrete volume summation provides an approximation of the object's volume. Example Computing Environment

[00160] FIG. 11 illustrates a schematic diagram of an exemplary computing environment 1100 in which the bulk volume estimation system 505, 700 can operate in accordance with one or more embodiments of the present disclosure. In one or more embodiments, the computing environment 1100 includes a service provider 1102 which may include one or more servers 1104 connected to a plurality of client devices 1106A - 1106C (e.g., heavy equipment) via one or more networks 1108. The client devices 1106A - 1106C, the one or more networks 1108, the service provider 1102, and the one or more servers 1104 may communicate with each other or other components using any communication platforms and technologies suitable for transporting data 47 and / or communication signals, including any known communication technologies, devices, media, and protocols supportive of remote data communications, examples of which will be described in more detail below with respect to FIG. 12.

[00161] Although FIG. 11 illustrates a particular arrangement of the client devices 1106A-1106C, the one or more networks 1108, the service provider 1102, and the one or more servers 1104, various additional arrangements are possible. For example, the client devices 1106A-1106C may directly communicate with the one or more servers 1104, bypassing the network 1108. Or alternatively, the client devices 1106A-1106C may directly communicate with each other. The service provider 1102 may be a public cloud service provider which owns and operates its own infrastructure in one or more data centers and provides this infrastructure to customers and end users on demand to host applications on the one or more servers 1104. The servers may include one or more hardware servers (e.g., hosts), each with its own computing resources (e.g., processors, memory, disk space, networking bandwidth, etc.), which may be securely divided between multiple customers, each of which hosts their own applications on the one or more servers.

[00162] In some examples, the service provider may be a private cloud provider who maintains cloud infrastructure for a single organization. The one or more servers 1104 may similarly include one or more hardware servers, each with its own computing resources, which are divided among applications hosted by the one or more servers for use by members of the organization or their customers.

[00163] Similarly, although the computing environment 1100 of FIG. 11 is depicted as having various components, the computing environment 1100 may have additional or alternative components. For example, the environment 1100 can be implemented on a single computing device with the bulk volume estimation system 700. In particular, the bulk volume estimation system 700 may be implemented in whole or in part on the client device 802A.

[00164] As illustrated in FIG. 11, the environment 1100 may include client devices 1106A - 1106C. The client devices 1106A - 1106C may comprise any computing device. For example, client devices 1106A - 1106C may comprise one or more personal computers, laptop computers, mobile devices, mobile phones, tablets, special purpose computers, TVs, or other computing devices, including computing devices described below with regard to FIG. 11. Although three client devices are shown in FIG. 11, it will be appreciated that client devices 1106A - 1106C may comprise any number of client devices (greater or smaller than shown).

[00165] Moreover, as illustrated in FIG. 11, the client devices 1106A-1106C and the one or more servers 1104 may communicate via one or more networks 1108. The one or more networks 1108 may represent a single network or a collection of networks (such as the Internet, a corporate Intranet, a virtual private network (VPN), a local area network (LAN), a wireless local network (WLAN), a cellular network, a wide area network (WAN), a metropolitan area network (MAN), or a combination of two or more such networks. Thus, the one or more networks 1108 may be any suitable network over which the client devices 1106A-1106N may access service provider 1102 and server 1104, or vice versa. The one or more networks 1108 will be discussed in more detail below with regard to FIG. 12.

[00166] In addition, the environment 1100 may also include one or more servers 1104. The one or more servers 1104 may generate, store, receive, and transmit any type of data, including input data 502 bulk volume estimation outputs 590, bulk volume estimation reports 630, or other information. For example, a server 1104 may receive data from a client device, such as the client device 1106A, and send the data to another client device, such as the client device 1102B and / or 1102C. The server 1104 can also transmit electronic messages between one or more users of the environment 1100. In one example embodiment, the server 1104 is a data server. The server 1104 can also comprise a communication server or a web-hosting server. Additional details regarding the server 1104 will be discussed below with respect to FIG. 12.

[00167] As mentioned, in one or more embodiments, the one or more servers 1104 can include or implement at least a portion of the bulk volume estimation system 400. In particular, the bulk volume estimation system 400 can comprise an application running on the one or more servers 1104, or a portion of the bulk volume estimation system 505, 700 can be downloaded from the one or more servers 1104. For example, the bulk volume estimation system 505, 700 can include a web hosting application that allows the client devices 1106A - 1106C to interact with content hosted at the one or more servers 1104. To illustrate, in one or more examples of the environment 1100, one or more client devices 1106A - 1106C can be visual sensing systems having sensors that communicate with server 1104.

[00168] Upon receiving the request, the one or more servers 1104 can automatically perform the methods and processes described above. The one or more servers 1104 can provide all or portions of the bulk volume estimations to the client device 1106A for display to the user. The one or more servers 1104 can also host a diagnostics application used to provide diagnostics to a subject.

[00169] As just described, the bulk volume estimation system 505, 700 may be implemented in whole, or in part, by the individual elements 1102-1108 of the computing environment 1100. It will be appreciated that although certain components of the bulk volume estimation system 505, 700 are described in the previous examples with regard to particular elements of the computing environment 1100, various alternative implementations are possible. For instance, in one or more examples, the bulk volume estimation system 505, 700 is implemented on any of the client devices 1106A-C. Similarly, in one or more examples, the bulk volume estimation system 505, 700 may be implemented on the one or more servers 1104. Moreover, different components and functions of the bulk volume estimation system 505, 700 may be implemented separately among client devices 1106A - 1106C, the one or more servers 1104, and the network 1108.

[00170] Examples of the present disclosure may comprise or utilize a special purpose or general-purpose computer including computer hardware, such as, for example, one or more processors and system memory, as discussed in greater detail below. Examples within the scope of the present disclosure also include physical and other computer readable media for carrying or storing computer-executable instructions and / or data structures. In particular, one or more of the processes described herein may be implemented at least in part as instructions embodied in a non-transitory computer-readable medium and executable by one or more computing devices (e.g., any of the media content access devices described herein). In general, a processor (e.g., a microprocessor) receives instructions, from a non - transitory computer-readable medium, (e.g., a memory, etc.), and executes those instructions, thereby performing one or more processes, including one or more of the processes described herein.

[00171] Computer-readable media can be any available media that can be accessed by a general purpose or special purpose computer system. Computer-readable media that store computer-executable instructions are non-transitory computer-readable storage media (devices). Computer-readable media that carry computer-executable instructions are transmission media. Thus, by way of example, and not limitation, examples of the disclosure can comprise at least two distinctly different kinds of computer-readable media: non-transitory computer-readable storage media (devices) and transmission media.

[00172] Non-transitory computer-readable storage media (devices) includes RAM, ROM, EEPROM, CD-ROM, solid state drives (SSDs) (e.g., based on RAM), Flash memory, phasechange memory (PCM), other types of memory, other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium which can be used to store desired 50 program code means in the form of computer-executable instructions or data structures and which can be accessed by a general purpose or special purpose computer.

[00173] A “network” is defined as one or more data links that enable the transport of electronic data between computer systems and / or modules and / or other electronic devices. When information is transferred or provided over a network or another communications connection (either hardwired, wireless, or a combination of hardwired or wireless) to a computer, the computer properly views the connection as a transmission medium. Transmission media can include a network and / or data links that can be used to carry desired program code means in the form of computer-executable instructions or data structures and which can be accessed by a general-purpose or special-purpose computer. Combinations of the above should also be included within the scope of computer-readable media.

[00174] Further, upon reaching various computer system components, program code means in the form of computer-executable instructions or data structures can be transferred automatically from transmission media to non-transitory computer-readable storage media (devices) (or vice versa). For example, computer-executable instructions or data structures received over a network or data link can be buffered in RAM within a network interface module (e.g., a “NIC”), and then eventually transferred to computer system RAM and / or to less volatile computer storage media (devices) at a computer system. Thus, it should be understood that non - transitory computer-readable storage media (devices) can be included in computer system components that also (or even primarily) utilize transmission media.

[00175] Computer-executable instructions comprise, for example, instructions and data which, when executed at a processor, cause a general-purpose computer, special-purpose computer, or special-purpose processing device to perform a certain function or group of functions. In some examples, computer-executable instructions are executed on a general-purpose computer to turn the general-purpose computer into a special-purpose computer implementing elements of the disclosure. The computer-executable instructions may be, for example, binaries, intermediate format instructions such as assembly language, or even source code. Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the described features or acts described above. Rather, the described features and acts are disclosed as example forms of implementing the claims.

[00176] Those skilled in the art will appreciate that the disclosure may be practiced in network computing environments with many types of computer system configurations, including 51 personal computers, desktop computers, laptop computers, message processors, handheld devices, multiprocessor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, mobile telephones, PDAs, tablets, pagers, routers, switches, and the like. The disclosure may also be practiced in distributed system environments where local and remote computer systems, which are linked (either by hardwired data links, wireless data links, or by a combination of hardwired and wireless data links) through a network, both perform tasks. In a distributed system environment, program modules may be located in both local and remote memory storage devices.

[00177] Examples of the present disclosure can also be implemented in cloud computing environments. In this description, “cloud computing” is defined as a model for enabling on-demand network access to a shared pool of configurable computing resources. For example, cloud computing can be employed in the marketplace to offer ubiquitous and convenient on-demand access to the shared pool of configurable computing resources. The shared pool of configurable computing resources can be rapidly provisioned via virtualization and released with low management effort or service provider interaction, and then scaled accordingly.

[00178] A cloud-computing model can be composed of various characteristics such as, for example, on-demand self-service, broad network access, resource pooling, rapid elasticity, measured service, and so forth. A cloud-computing model can also expose various service models, such as, for example, Software as a Service (“SaaS”), Platform as a Service (“PaaS”), and Infrastructure as a Service (“laaS”). A cloud computing model can also be deployed using different deployment models such as private cloud, community cloud, public cloud, hybrid cloud, and so forth. In this description and in the claims, a “cloud-computing environment” is an environment in which cloud computing is employed. Example Operating Environment

[00179] Having described an overview of examples of the present technology, an example operating environment in which examples of the present technology may be implemented is described in order to provide a general context for various aspects of the present technology. Referring now to FIG. 12, in particular, an exemplary operating environment for implementing examples of the present technology is shown and designated generally as computing device 1200. Computing device 1200 is but one example of a suitable computing environment and is not intended to suggest any limitation as to the scope of use or functionality of the technology. Neither should computing device 1200 be interpreted as having any dependency or requirement relating to any one or combination of components illustrated.

[00180] The technology of the present disclosure may be described in the general context of computer code or machine-useable instructions, including computer-executable instructions such as program modules, being executed by a computer or other machines, such as a personal data assistant or other handheld devices. Generally, program modules, including routines, programs, objects, components, data structures, etc., refer to code that performs particular tasks or implement particular abstract data types. The technology may be practiced in a variety of system configurations, including hand-held devices, consumer electronics, general-purpose computers, more specialty computing devices, etc. The technology may also be practiced in distributed computing environments where tasks are performed by remote-processing devices that are linked through a communications network.

[00181] With reference to FIG. 12, computing device 1200 includes bus 1202 that directly or indirectly couples the following devices: memory 1204, one or more processors 1206, one or more presentation components 1216, input / output ports 1210, input / output components 1212, and illustrative power supply 1214. Bus 1202 represents what may be one or more buses (such as an address bus, data bus, or combination thereof). Although the various blocks of FIG. 12 are shown with lines for the sake of clarity, in reality, delineating various components is not so clear, and metaphorically, the lines would more accurately be grey and fuzzy. For example, one may consider a presentation component, such as a display device, or an I / O component. Also, processors have memory. We recognize that such is the nature of the art and reiterate that the diagram of FIG. 12 merely illustrates an example computing device that can be used in connection with one or more examples of the present technology. A distinction is not made between such categories as “workstation,” “server,” “laptop,” “hand-held device,” etc., as all are contemplated within the scope of FIG. 12 and reference to “computing device.”

[00182] Computing device 1200 typically includes a variety of computer-readable media. Computer-readable media can be any available media that can be accessed by computing device 1200 and includes both volatile and nonvolatile media, removable and non-removable media. By way of example, and not limitation, computer-readable media may comprise computer storage media and communication media.

[00183] Computer storage media include volatile and nonvolatile, removable, and nonremovable media implemented in any method or technology for storage of information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium which can be used to store the desired information and which can be accessed by computing device 800. Computer storage media excludes signals per se.

[00184] Communication media typically embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism and includes any information delivery media. The term “modulated data signal” means a signal that has one or more of its characteristics set or changed in such a manner as to encode information in the signal. By way of example, and not limitation, communication media includes wired media such as a wired network or direct-wired connection, and wireless media such as acoustic, RF, infrared, and other wireless media. Combinations of any of the above should also be included within the scope of computer-readable media.

[00185] Memory 1204 includes computer storage media in the form of volatile or nonvolatile memory. The memory may be removable, non-removable, or a combination thereof. Examples of hardware devices include solid-state memory, hard drives, optical-disc drives, etc. Computing device 1200 includes one or more processors that read data from various entities, such as memory 1204 or VO components 1212. Presentation component(s) 1208 presents data indications to a user or other device. Examples of presentation components include a display device, speaker, printing component, vibrating component, etc.

[00186] VO ports 1210 allow computing device 1200 to be logically coupled to other devices, including VO components 1212, some of which may be built in. Illustrative components include a microphonejoystick, game pad, satellite dish, scanner, printer, wireless device, etc.

[00187] Having identified various components in the present disclosure, it should be understood that any number of components and arrangements may be employed to achieve the desired functionality within the scope of the present disclosure. For example, the components in the examples depicted in the figures are shown with lines for the sake of conceptual clarity. Other arrangements of these and other components may also be implemented. For example, although some components are depicted as single components, many of the elements described herein may be implemented as discrete or distributed components or in conjunction with other components, and in any suitable combination and location. Some elements may be omitted altogether. Moreover, various functions described herein as being performed by one or more entities may be carried out by hardware, firmware, and / or software, as described below. For instance, various functions may be carried out by a processor executing instructions stored in memory. As such, other arrangements and elements (e.g., machines, interfaces, functions, orders, and groupings of functions, etc.) can be used in addition to or instead of those shown.

[00188] The subject matter of the present disclosure is described with specificity herein to meet statutory requirements. However, the description itself is not intended to limit the scope of this patent. Rather, the inventor has contemplated that the claimed subject matter might also be embodied in other ways, to include different steps or combinations of steps similar to the ones described in this document, in conjunction with other present or future technologies. Moreover, although the terms “step” and / or “block” may be used herein to connote different elements of methods employed, the terms should not be interpreted as implying any particular order among or between various steps herein disclosed unless and except when the order of individual steps is explicitly described. For purposes of this disclosure, words such as “a” and “an,” unless otherwise indicated to the contrary, include the plural as well as the singular. Thus, for example, the requirement of “a feature” is satisfied where one or more features are present.

[00189] The present disclosure has been described in relation to particular examples, which are intended in all respects to be illustrative rather than restrictive. Alternative examples will become apparent to those of ordinary skill in the art to which the present disclosure pertains without departing from its scope.

[00190] From the foregoing, it will be seen that this disclosure is one well adapted to attain all the ends and objects set forth above, together with other advantages which are obvious and inherent to the system and method. It will be understood that certain features and subcombinations are of utility and may be employed without reference to other features and subcombinations. This is contemplated by and is within the scope of the claims.

Claims

1. A method for bulk volume estimation of material held by a load container, the method comprising:collecting image data of a load container during a digging cycle;identifying key points of the load container from the image data representing the load container byinputting the image data into a neural network trained to classify the key points of the load container;analyzing the image data by the neural network; andpredicting, via the neural network, the key points observable in the image data;estimating occluded key points of the load container using the image data representing the load container and the key points;calculating a three dimensional profile of the material held by the load container;fitting a model with the key points and the occluded key points to the three dimensional profile of the material to determine pose; andestimating a bulk volume of the material based on the three dimensional profile of the material and the pose.

2. The method of claim 1, further comprising:determining, prior to identifying the key points, an operational phase of the load container based on the image data, wherein the operational phase indicates the load container is in a swing full state and contains the material.

3. The method of claim 2, wherein determining the operational phase of the load container comprises:inputting the image data into a neural network trained to classify operational phases of the load container;analyzing the image data by the neural network; andestimating, via the neural network, the operational phase indicating the load container is in the swing full state.02 05 244. The method of claim 3, wherein the operational phases include the swing full state, a swing empty state, and an idle state.

5. The method of any preceding claim, wherein estimating the occluded key points comprises:inputting the image data into a neural network trained to identify the occluded key points, wherein the occluded key points are occluded by the material being held by the load container;analyzing the image data by the neural network;estimating, via the neural network, the occluded key points associated with the load container.

6. The method of claim 5, further comprising:collecting additional image data of the load container during the digging cycle;determining an operational phase of the load container based on the additionalimage data, wherein the operational phase indicates the load container is ina swing empty state;analyzing the additional image data by the neural network; andrefining the neural network based on the key points estimated during the swingempty state of the load container.

7. The method of any preceding claim, further comprising:determining the digging cycle of the load container has performed a plurality of operational phases;aggregating the bulk volume of the material with previous bulk volume estimations of the material; andreporting the aggregated bulk volume of the material.

8. The method of claim 6, wherein the previous bulk volume estimations of the material are associated with an operator of the load container.02 05 249. The method of claim 8, wherein estimating the bulk volume is performed using a discrete volume summation methodology or a continuous load profile prior model.

10. The method of any preceding claim, further comprising fitting a reference plane to the load container based on a normal vector associated with the key points and the occluded key points, wherein the normal vector represents the pose of the load container.

11. The method of claim 10, wherein estimating the bulk volume of the material comprises:retrieving a load segmentation of the material held by the load container;determining a three-dimensional contour of the material using the load segmentation as a region of interest; andestimating the bulk volume of the material based on the three-dimensional contour of the material positioned above and below the reference plane.

12. The method of any preceding claim, wherein the key points of the load container include ground engaging teeth, wing shrouds, and lip shrouds associated with the load container.

13. The method of any preceding claim, wherein the load container is an excavator bucket.

14. The method of any preceding claim, wherein the load container is a mining shovel.

15. A system for bulk volume estimation of material held by a load container, the system comprising:at least one processor; andone or more computer storage media storing computer executable instructions that, when executed by the at least one processor, cause the at least one processor to perform operations comprising:collecting image data of a load container;02 05 24identifying key points of the load container from the image data by inputting the image data into a neural network trained to classify the key points of the load container;analyzing the image data by the neural network; andpredicting, via the neural network, the key points observable in the image data;estimating occluded key points of the load container from the image data representing the load container having material and the key points;calculating a three dimensional profile of the material held by the load container;fitting a model with the key points and the occluded key points to the three dimensional model of the material to determine pose; and estimating a bulk volume of the material based on the three dimensional profile of the material and the pose.

16. The system of claim 15, wherein the load container is a truck tray of a haul truck.

17. The system of any of claims 15 to 16, further comprising fitting a reference plane to the load container based on a normal vector associated with the key points and the occluded key points, wherein the normal vector represents the pose of the load container.

18. The system of claim 17, wherein estimating the bulk volume of the material comprises:retrieving a load segmentation of the material held by the load container;determining the three dimensional profile of the material using the load segmentation as a region of interest;estimating the bulk volume of the material based on the three dimensional profile of the material positioned above and below the reference plane.

Citation Information

Patent Citations

  • An excavator bucket material volume and weight measuring system

    CN109948189A

  • Measurement of bulk density of the payload in a dragline bucket

    US20120136542A1

  • System and method for determining the material loading condition of a bucket of a material moving machine

    US20180239849A1

  • Container angle sensing using vision sensor for feedback loop control

    US20200040555A1