Intelligent monitoring system for mineral loading process

The GET Smart system employs AI and neural networks to accurately detect missing GETs on mining equipment, reducing downtime and improving efficiency by minimizing false alarms and physical inspections.

JP2024537640A5Pending Publication Date: 2025-12-11JEBI S A C +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024515449
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2022-09-12
Publication Date
2025-12-11

AI Technical Summary

Technical Problem

Existing systems for detecting the loss of ground-engaging tools (GETs) on mining equipment suffer from high false alarm rates, leading to unnecessary work stoppages and operator fatigue, while human inspections are impractical due to the size and hazardous conditions of excavators.

Method used

The GET Smart system uses AI modeling and neural networks to process data from multidimensional sensors, constructing enriched tensors and applying statistical and AI techniques to accurately detect missing GETs with a low false positive rate, notifying operators and remote users via in-cab monitors and mobile applications.

Benefits of technology

The system efficiently identifies missing GETs, reduces downtime, and provides metrics like wear level and mineral volume, enhancing mining operation efficiency and safety by minimizing false alarms and physical inspections.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The "GET Smart" system uses AI modeling and neural network techniques to efficiently identify lost wear part events and provide other useful metrics during excavator operation to improve efficiency and reduce downtime. The system uses multi-dimensional sensors to monitor the integrity of the Ground Engaging Tool (GET), determine areas of interest, and generate and process enriched tensors via an embedded system with a combination of CPU and TPU. The system also determines the wear level of the GET, the volume of mineral per shovel bucket, and the average grain size.
Need to check novelty before this filing date? Find Prior Art

Description

[Background technology]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority from earlier Peruvian Application No. 001494-2021 / DIN, filed September 10, 2021, the entire contents of which are incorporated herein by reference.

[0002] The present invention relates generally to systems and methods for monitoring mineral loading on mining equipment.

[0003] Heavy machinery such as excavators are routinely used in mineral and soil mining. Such machines are equipped with shovels or buckets to rapidly move loose ore into waiting vehicles for downstream processing. The operating implements that engage the loose rock include one or more ground-engaging tools (GETs) designed to be sacrificed at a specific stage of wear. Because these parts have high hardness, loss of a part can damage downstream equipment such as crushers and conveyor belts. While such events are rare, they can result in significant downtime and safety hazards. Therefore, it is important for mining operations to detect the loss of a wear part as close to the loss event as possible.

[0004] Various techniques for detecting a GET loss event have been contemplated in the prior art, such as capturing successive images of the operating implement and measuring pixel intensity values ​​to determine whether a subset of pixels corresponds to a wear part. Another technique embeds an RFID module within the GET to establish its location.

[0005] However, because actual lost GET events occur on average only once per year per excavator, many of these systems suffer from unacceptable levels of false alarms, resulting in unnecessary work stoppages or operator fatigue, ignoring the alarms ("boy cried wolf syndrome"), or requiring very expensive specialized GETs (such as RFID implementations).

[0006] Additionally, using humans for frequent inspections is impractical due to the sheer size of the excavators, which poses a risk to nearby personnel, and they must operate during the day, at night, in extremely hot or cold conditions, or in adverse weather conditions. DISCLOSURE OF THE INVENTION [Problem to be solved by the invention]

[0007] The invention, called the "GET Smart" system, uses AI modeling and neural network techniques to efficiently identify missing wear part events and provide other useful metrics during excavator operation to improve efficiency and reduce downtime. [Means for solving the problem]

[0008] In one embodiment, the present invention monitors the integrity of ground engaging tools (GETs) that have the potential to cause catastrophic damage within mining operations downstream of the earth moving step.

[0009] The present invention uses various multidimensional sensors to collect information about the excavator, its components, and its surroundings. All information is then structured into a unique data structure called an enriched tensor and processed in real time using an embedded system equipped with a CPU and a tensor processing unit (TPU). The data is processed through statistical and artificial intelligence (AI) techniques, including neural networks and visual transformers.

[0010] The invention uses several parallel and independent processing techniques to generate separate (or discrete) results, which are individually evaluated by a custom algorithm to accurately detect missing GETs within an acceptable false positive rate. When a "true positive" event is identified, the invention notifies the operator via an in-cab monitor and remote users via cloud and mobile applications.

[0011] Other metrics related to the earth moving process are calculated by the present invention. These include detecting the wear level of the GET, the volume of mineral per shovel bucket, and the average grain size. [Brief explanation of the drawings]

[0012] [Figure 1A] FIG. 1A shows the hardware assembly. [Figure 1B] FIG. 1B shows the hardware assembly attached to the top of the drilling rig. [Figure 1C] Figure 1C shows the shovel and wear parts. [Figure 1D] FIG. 1D shows the shovel and wear parts in a removed / exploded view. [Figure 1E] FIG. 1E shows the operator's cabin. [Figure 1F] FIG. 1F shows the data acquisition unit. [Figure 2A] FIG. 2A shows an exemplary camera view of an operating environment. [Figure 2B] FIG. 2B is a flow diagram illustrating creating an enriched tensor data structure. [Figure 2C] FIG. 2C shows an exemplary region of interest. [Figure 2D]FIG. 2D is a graphical representation of an exemplary frame of an enriched tensor. [Figure 2E] Figure 2E is a graphical representation of the point cloud. [Figure 2F] FIG. 2F is a graphical representation of a depth frame. [Figure 2G] Figure 2G is a flow diagram showing the artificial intelligence module. Table 1 is a summary chart of the training dataset. [Figure 3A] FIG. 3A illustrates the training process. [Figure 4A] FIG. 4A is a flow diagram showing data flow from an enriched tensor. [Figure 4B] FIG. 4B shows the labels and confidence values. [Figure 4C] FIG. 4C is a flow diagram illustrating the GET smart manager. [Figure 4D] FIG. 4D shows a depth map for the region of interest. [Figure 4E] Figure 4E shows the functionality of the wear detection app. [Figure 5A] Figure 5A shows object recognition in the volumetric analysis app. [Figure 5B] FIG. 5B shows the calculations performed by the Volume Analysis app. [Figure 5C] Figure 5C shows object recognition in the particle size analysis app. [Figure 5D] FIG. 5D shows the calculations performed by the particle size analysis app. DETAILED DESCRIPTION OF THE INVENTION

[0013] <Implementation environment, hardware, and user interface> Figure 1A is an exemplary embodiment of the hardware assemblies required for the system's function. Tower assembly 100 includes mast 101, data collection unit 102, and processing enclosure 103. The data collection unit includes sensory equipment used to collect visual, depth, and inertial data, as well as equipment used to illuminate the work area and maintain the sensors in good operational condition. The mast includes a base for installation in high-vibration and harsh environments and includes slots for engaging various hardware components and a protected location for routing electrical cables. A pair of headlamps 104 provides the illumination necessary for visual detection and operator assistance.

[0014] The processing enclosure 103 essentially comprises a computer on which the software portion of the present invention runs, and a power management unit.

[0015] The tower assembly 100 may be placed on an excavator, as shown in FIG. 1B, where the data collection unit is optimally located at a high point within the equipment to minimize the effects of dust or flying debris. The excavator includes a shovel 110 with exemplary ground-engaging tools / wear parts 111 mounted at a location that approximately defines the rim of the shovel. These parts are designed to wear out to prevent damage to the shovel or the excavator itself. The excavator arm 115 moves independently of the tower assembly, and an operator cabin 120 contains the user interface for the system, described in more detail below. A person is provided to indicate approximate scale.

[0016] Figure 1C is an exemplary view of the shovel and wear parts from the perspective of a data collection unit, and Figure 1D shows an exploded view of the wear parts 111. They are attached to the shovel via raw attachment points 112 that are not designed to wear. Whenever an attachment point is visible, it means that one or more wear parts have been removed. Any part, individually or in combination, may be removed at any time during a mining operation.

[0017] 1E is an exemplary diagram of the interior of the operator cabin. The GET smart system user interface is represented by a single monitor module 121 located in a prominent location for the equipment operator. In a preferred embodiment, this feature may be implemented using visual and audible alarms, and a mobile app that notifies off-site personnel is expressly contemplated.

[0018] 1F shows the data collection unit 102 alone. A sensor enclosure 131 houses various sensors, including at least one or more stereoscopic cameras and LIDAR sensors. Other video cameras known in the art may also be used to collect visual data. These cameras are protected from the elements via a camera window 132. In an exemplary embodiment, these sensors may be an Intel RealSense D455 active stereoscopic camera and a Continental High Resolution 3D Flash LiDAR.

[0019] The environmental sensors, which provide at least inertial data for tracking the movement of the excavator, include accelerometers, gyroscopes, and magnetometers. Additionally, temperature, barometric pressure, humidity, noise, and light emission data may be collected. In an exemplary embodiment, the sensors may be Bosch XDKs. The sensors are mounted within the tower assembly.

[0020] The sensor enclosure is protected by an air blade system 140, which provides conditioned airflow to keep the data collection unit in good operational condition. The system works by directing airflow through the top of the enclosure and downward past the camera window. This prevents dust and moisture buildup on the window and deflects flying debris that could damage the window and sensor. The conditioned air also keeps the sensor at an optimal operating temperature. This system reduces maintenance needs and the frequency with which personnel are required to access physical components.

[0021] <Enriched tensor and AI module> At the heart of the GET smart system are data structures known as enriched tensors, which store relevant portions of the sensor data captured by the data collection units, and software algorithms known as AI modules that operate on the enriched tensors.

[0022] FIG. 2A is an exemplary graphical representation of information viewed by a data collection unit associated with the system. Within the frame, a camera captures a typical operating environment with an excavator shovel in front, and most, if not all, of the GET 111 are visible. The overall terrain 200 may have different colors or textures depending on whether it is dry or wet, whether the sun is shining or cloudy, and whether it is daytime or nighttime. Features such as snowdrifts 201, puddles 202, buildings 203, and workers 204 are all visible, while rain or condensation droplets 205 are in close proximity to the camera lens, thus obscuring the view. Several clouds 206 may be present in the sky 207. While these features are easily discernible by the human eye, each presents its own challenges for machine perception, and the focus should be on detecting missing GETs and other desirable features of the mining process.

[0023] As shown in Figure 2B, the GET smart system begins the recognition process by collecting raw data 211 in different formats, both structured and unstructured, using a variety of sensors 210. An inertial measurement unit (IMU) generates information related to the specific force, angular velocity, and orientation of a mass over time, a stereo camera provides depth data within the camera's field of view, a LIDAR captures and builds a point cloud of distances for each point, and an imager captures image data in RGB format.

[0024] After capturing the raw data, a region of interest (ROI) is calculated by the tensor module 212. This calculation allows the system to efficiently identify the most relevant data for inclusion in the enriched tensor. The ROI is the smallest subset of data most likely to contain information related to the object of interest.

[0025] FIG. 2C shows an exemplary ROI. It can be represented by a specific region of the visible screen where an object of interest, such as an excavator shovel, and thus a wear part or attachment point, should appear at any given moment. It is defined by the region containing two horizontal lines 220 and 221 and two vertical lines 222 and 223, where the shovel bucket with the GET in view should appear most of the time. Defining such a region immediately limits the system's visual processing to a portion of the total captured data, significantly reducing processing time. It may also be represented in 3D space, as discussed further below.

[0026] The ROI may be determined via presets or solely via IMU data. The inertial data collected by the sensors can accurately determine the state of the excavator, whether the shovel is digging or hauling the ground, and therefore can optimally determine the state when the shovel is pointing up (and the GET is most visible). Therefore, the ROI may also be time-based, in which case the system simply ignores data collected when it knows the shovel is pointing down and the GET is not visible. By limiting the data, it reduces the chance of generating false positives.

[0027] In other embodiments, the horizon defining the ROI may be dynamic and generated by taking the center point of the detected GET and applying a minimum and maximum distance to the center point, which is then used as an axis to determine the area where the GET is likely to be found. This technique may prevent occasional edge cases where the GET drifts slightly outside the preset area, resulting in a false detection. The boundaries of the ROI vary from implementation to implementation and are customizable based on parameters such as the size of the excavator shovel and the location of the sensor. The ROI also need not necessarily be a geometric portion of the entire visible screen, but may be limited to a rectangle or cuboid surrounding the object of interest, such as the GET or its "raw" attachment point after it is detached.

[0028] Returning to Figure 2B, the tensor module normalizes the collected data within the region of interest and uses this data to construct enriched tensors.

[0029] FIG. 2D is a graphical representation of an exemplary frame of an enriched tensor, a data structure that includes IMU values ​​215, point cloud 216, red, green, and blue (RGB) images 217, 218, and 219, and depth frame 220. This data structure may be implemented via a multidimensional matrix or any data structure known in the visual processing arts. Each frame of the enriched tensor represents a specific moment in time, captured and stored at a frame rate that varies from 15 to 90 frames per second. Thus, the enriched tensor is a time series of inertial, point cloud, image, and depth data for any given region of interest and is the basis for further processing.

[0030] Figure 2E is a graphical representation of the point cloud 216 portion of the enriched tensor frame. Because this is a 3D representation, the shovel arm 115 appears near the front (as shown) (up), with the bucket adjacent but somewhat further back. The GET is expected to be in the position shown in the drawing. Details such as the ground, landscape, etc. are in the background.

[0031] The region of interest is represented by a cuboid 230, which is also the total volume of the point cloud; only the portion of the point cloud within the cuboid is present in the enriched tensor; the portion outside is largely irrelevant and left unprocessed (or, alternatively, the region of interest is represented by a cuboid 230, which is also the total volume of the point cloud; only the portion of the point cloud within the cuboid is present in the enriched tensor; the portion outside is largely irrelevant and left unprocessed).

[0032] 2F is a graphical representation of the depth frame 220 (or depth map) portion of the enriched tensor frame. The excavator 110 is closest to the camera, with other captured features further back. Environmental features 221 may be close to the camera's distance, but are typically excluded from consideration because they are outside the region of interest. Stereo cameras can easily exceed the granularity of the information represented by a line drawing, so this situation is rarely encountered.

[0033] FIG. 2G is a flowchart of the AI ​​module 250, which comprises multiple neural networks (NNs) adapted to process individual data flows separately. The enriched tensor is separated into its component streams and processed separately. Every neural network outputs prediction data that may include at least the labels of identified relevant objects and a confidence level for such predictions. In a preferred embodiment, the hardware used is a tensor processing unit (TPU) running Google's TensorFlow framework.

[0034] The AI ​​module includes a two-dimensional convolutional neural network (2D-CNN) configured to process the 2D image portion of the enriched tensor. The output is fed to a classification model, which makes a refined prediction. Both of these outputs are retained for later use. A weighting process assigns a higher score when the outputs match and a lower score when the outputs do not match. In a preferred embodiment, the 2D-CNN can be a single-shot detector (SSD) with a ResNet-18 neural network as its backbone. The classification model can be a dense neural network (or deep neural network / DNN), which is a CNN with a depth of 18 to 100 layers.

[0035] The AI ​​module further comprises a three-dimensional convolutional neural network (3D-CNN) configured to process the point cloud portion of the enriched tensor. This NN is similar to a 2D-CNN, except that it is trained to process 3D data. In an exemplary embodiment, this component may be PointPillars.

[0036] The AI ​​module further includes additional computations used to process the depth data to obtain the distance to the object of interest, and a recurrent neural network (RNN) for processing the IMU (inertial measurement unit) data. RNNs are the preferred embodiment because they are adept at processing time-series data (inertial data is structured as time-series data). Specifically, a long-short-term memory (LSTM) neural network is preferred because it handles long-term dependencies better than a typical RNN and does not "forget" excavator states that occur only intermittently. Each generates its own output.

[0037] Finally, the AI ​​module includes a foundational model that processes the entire enriched tensor, regardless of the individual data streams. In a preferred embodiment, the foundational model is a Vision Transformer (ViT) rather than a neural network. In machine learning, a transformer is composed of multiple self-attention layers that are familiar with general learning methods that can be applied to various data modalities. In the GET smart system, it is used to process the entire enriched tensor as a whole and output prediction data in the same way.

[0038] Training and Data Augmentation Neural networks or deep learning models require training before they can be used in an application environment. The goal of model training is to build the best mathematical representation of the relationship between detected data features and labels, and inure that relationship to the NN. In GET smart systems, training typically involves providing a dataset containing information about pre-labeled objects of interest. The training set can be thought of as a set of "right answers" that the NN must become familiar with.

[0039] The training dataset is split into at least three subsets: the training subset is used to improve the performance of the NN through supervised learning, the validation subset is used as a quiz to demonstrate proficiency in the task, and the test subset is used to obtain a set of proficiency metrics for the NN.

[0040] Table 1 is a quantitative summary chart of the relevant features of the training dataset. In one embodiment, the dataset used contains 37,104 training images and 4,106 test images. The relevant features (presence or absence of wear parts) are labeled so that the NN can recognize them. For example, in the dataset, tooth objects are labeled 92,430 times. Objects labeled "raw" are images of damaged GETs or raw attachment points 112 (see FIG. 1C) as seen after the GET has been removed. Table 1. Summary chart of the training dataset TIFF2023037344000001.tif130169

[0041] The number of features and images in the table is an example implementation for fully training a 2D-CNN neural network. Each type of model requires its own training dataset. However, if the dataset is ideal and as diverse as possible, it is estimated that at least 10,000 images are required, while less ideal datasets may require more than 40,000 images. For 3D-CNNs that process point clouds, at least 10,000 point clouds are required.

[0042] 3A shows an exemplary training process for a neural network. In a preferred embodiment, training is performed on batches of 64 images at a time. After processing each batch 3410, the NN calculates the most appropriate weights and determines an error value 3421 representing its accuracy in determining the features by performing an intersection 3420 (XNOR) between the predicted labels 3412 and the test labels 3413 generated by the NN. Once a full epoch (all 580 batches of 64 images) has been presented, a weight correction is performed that takes into account the previous error values ​​generated, and the process is repeated.

[0043] To reach the desired accuracy, 80-120 epochs are used per training run, with 10-20 training runs. This process can take hours or even days. In the preferred embodiment, the training process is performed on the NVIDIA TAO framework. Training is automated via a pipeline. The pipeline runs on dedicated hardware (dedicated to training) and is implemented in Python 3.

[0044] The procedure for training a 3D-CNN is identical to that for a 2D-CNN, except that the intersection of the test image labels and the predicted labels is in three dimensions. Each model is trained in its own context; for example, when training an RNN, only inertial data is used, as the RNN only processes data obtained from an IMU.

[0045] The training dataset is constructed from all combinations of inputs and sensors collected from similar devices and environments. Similar to actual operation, raw data is collected from visual cameras, LIDAR, environmental sensors, IMUs, and radars. However, the GET smart system generates additional training data by manipulating and transforming the initial dataset through a process called data augmentation.

[0046] The goal of data augmentation is to generate sufficiently diverse training data so that the trained AI model is more robust than what can typically be achieved through experiments or simulations. Since actual GET loss events are rare, it is not possible to collect all of the data sets required in a production environment. Also, GET loss events are not captured under sufficiently variable conditions to allow for the desired training diversity.

[0047] Some of the techniques explicitly contemplated are performing mathematical operations on images, such as presenting a mirror image, tilting an image up to 15 degrees in any direction, zooming in or out, or adjusting the contrast so that the image becomes brighter or darker.

[0048] Other techniques include applying the above operations to 3D point clouds. In a preferred embodiment, the point cloud of a particular object of interest is manipulated such that all data points are moved slightly further or closer to a calculated point representing the center of the object, generating a point cloud of an object that is slightly larger or smaller than the actual object.

[0049] In another embodiment, in a process known as a digital twin, a digital representation of an excavator and a shovel, including synthetic data, can be generated.

[0050] <GET Loss Detection and Wear Detection Application> One of the primary tasks of a GET smart system is to detect GET loss events as they occur, or as close to the event as possible, so that wear parts can be located and removed. This should be achieved while keeping the frequency of false alarms below operator tolerance levels so that mining operations are not unduly disrupted.

[0051] Another task of the system is to detect the wear level of any particular GET. Since these parts are sacrificial, wear is an integral part of their life cycle. Accurately predicting GETs approaching the end of their service life indicates the need for preventative maintenance and may avoid a lost GET event altogether.

[0052] Figure 4A shows the flow diagram and modules used in any of these tasks. Data from the enriched tensor is pulled into a series of queues for processing. If the stream is selected for recording for later use, e.g., training, the data is sent to a recording queue and a recording module for persistence to storage.

[0053] The data selected for processing is queued in the AI ​​module 250, which generates separate (or discrete) outputs from each model, as described above. These outputs include at least predicted labels and associated confidence levels for each label. The outputs include at least 3D-FM (three-dimensional foundation model); 3D-CNN; 2D-FM (two-dimensional foundation model); 2D-CNN; 2D classifier; depth; and RNN (inertia) outputs.

[0054] FIG. 4B is a graphical representation of labels generated by an exemplary module for processing visual data. In any given frame of the enriched tensor, the model recognizes multiple objects of interest, such as GET 111. Each recognized object is surrounded by a bounding box 410, 412, or 413. Each object has an associated label 411, which details the type of recognized object and is associated with the model's confidence level. Depending on the conditions, not all GETs can always be recognized with a high confidence level. For example, the GET within bounding box 412 displays a confidence level of only 72% because it is partially obscured by material 415.

[0055] In the figure, the GET within bounding box 413 is missing or damaged. The AI ​​module determines this and assigns it a "raw" label and confidence level for further processing downstream.

[0056] For ease of reading, not all of the bounding boxes or labels are shown herein. Note that although the illustrations are two-dimensional, three-dimensional results are produced and mathematically manipulated where the bounding box is a cube.

[0057] Figure 4C details the GET Smart Manager 401, which weights / tiebreaks the decisions of each AI model and determines whether an alert is warranted. A set of custom thresholds 402 helps make this decision. These thresholds can vary from implementation to implementation based on site, environment, or operator tolerance for error. If all thresholds are not met, the system determines there is a false event and does not sound an alert. However, once a consensus is reached, the system reports a true event, sends a notification, and flags the event for recording.

[0058] In one embodiment, the manager evaluates the output from the 2D classifier and generates a set of labels and confidence levels for each element detected within the pre-defined bounding box. The output from the 3D-CNN is evaluated similarly, except that only objects that overlap with the 2D bounding box are evaluated. If the model further agrees with the label, the object is considered valid.

[0059] More generally, if multiple models report a missing GET with high confidence, but the RNN reports that the shovel is in a location where the GET should not be visible, it is likely a false positive of an incorrect particle resembling the missing GET, and no alert will be sounded. If multiple models report a missing GET in an area that is too close or too far from the region of interest based on depth data, the manager may determine that the detection is not valid.

[0060] 4D details an embodiment in which the region of interest (ROI) is selected not from a preset but from recognizing the object of interest as described in FIG. 4B. It is clear that many objects similar to the GET may be mistakenly picked up by the AI ​​model. Therefore, depth map data needs to be taken into account to determine the relevant object of interest.

[0061] In this representation of the depth map, at least some of the GETs are shown selected in their ROIs. For each recognized object, the center of mass is calculated and mapped to a corresponding location on the depth map. Only those located within a precise distance (an exemplary distance could be between 7 and 8 meters) are selected for further processing.

[0062] Figure 4E illustrates the process for making the wear detection app work. The GET smart system generates a bounding box 420 containing each individual GET and determines a GET shape 421 that outlines the boundaries of the GET image. From this shape, a distance 423 to the center 422 is calculated, from which a polygonal approximation is determined, providing at least the major and minor axes of the shape.

[0063] The system also determines a GET point cloud 424 from the LIDAR data that outlines the GET, after which a polygonal approximation 425 is calculated.

[0064] This information is then calculated to determine physical parameters such as the area, mass, measurements, and volume of the GET. Now that these parameters are known for the new GET, the wear level for each individual GET can be calculated.

[0065] <Volume analysis and particle size analysis app> The GET Smart System may leverage its AI model to automate certain other tasks without the need to physically manipulate or inspect the collected minerals. The system performs a volumetric analysis of a given shovel load to determine the volume and weight of the material scooped. The system may also provide an estimate of the material's particle size. As is evident from the drawing, many different tasks can be accomplished with just one set of sensor data inputs.

[0066] Referring to Figure 5A, the volumetric analysis app first defines a region of interest delineated by a horizontal line 430, where the area of ​​the entire shovel load's surface 431 is likely to be found. Due to the nature of the minerals collected, this surface is likely to be uneven. From this region, a surface map of all points on the uneven surface can be constructed from at least the point cloud data.

[0067] 5B shows a surface map superimposed on the shape of a shovel bucket. The dimensions 432 of the shovel bucket are known constants. Each point on the surface map corresponds vertically to a known location on the bottom, so the depth and volume of the shovel load can be calculated. If the density of the material is known, the weight of the shovel load can also be determined.

[0068] These metrics are available via a user interface and allow, for example, an operator to quantify the load placed on a dump truck to ensure the load is loaded properly and efficiently. Underloading (or underloading) a dump truck can cause costly inefficiencies during the ground moving process, while overloading (or overloading) a dump truck can damage the truck.

[0069] With regard to granular applications, a typical environment in which an excavator operates is scooping material after a mine has already been blasted with explosives. Particle size analysis or particle size measurement is important for coordinating and optimizing both the blasting process and downstream processing.

[0070] 5C, the grain size analysis app begins by defining a region of interest delineated by line 440. Surface analysis involves delineating grain boundaries 441 on multiple grains present on the surface of the collected mineral.

[0071] Figure 5D shows multiple bounding boxes 442 defined around the identified particle boundaries. Similar to the wear detection app, the distance to the center of each particle is calculated, and a polygonal approximation is determined, providing at least the major and minor axes of the shape. The system also determines multiple particle points 443 that outline the particle, and then a polygonal approximation 444 is calculated.

[0072] This information is then calculated to determine physical parameters such as area and volume per particle. These samples at the surface represent the entire contents of the shovel as they are randomly picked up, and therefore there is no need to analyze sub-surface components.

[0073] The dimensional measurements are then converted into useful metrics that inform the mine blasting process and processing plant. These metrics are reported via the cloud to the mine blasting operators, allowing them to set mineral processing parameters and provide feedback to blasting operations to detect under- or over-blasting.

[0074] The advantage of these applications in the GET smart system is that less physical contact and downtime is required to inspect or measure the removed material, thus increasing efficiency and safety.

[0075] All publications and patent applications cited in this specification are herein incorporated by reference to the same extent as if each individual publication or patent application was specifically and individually indicated to be incorporated by reference.

[0076] While the invention has been described with reference to exemplary embodiments, those skilled in the art will recognize that various changes can be made and equivalents can be substituted for elements thereof without departing from the scope of the invention. In addition, many modifications can be made to adapt a particular situation or material to the teachings without departing from the essential scope. Therefore, it is not intended that the invention be limited to the particular embodiment disclosed as the best mode contemplated for carrying out this invention, but rather, the invention is intended to include all embodiments falling within the scope of the appended claims.

[0077] The following is the invention as originally described in the present application. <Claim 1> 1. An AI based monitoring system for use in detecting the condition of an excavator during mineral loading in a mining operation, comprising: one or more sensors; an enriched tensor data structure; An artificial intelligence module; a weighting mechanism; and one or more outputs. <Claim 2> The sensor further comprises: one or more color imaging cameras; one or more LIDAR sensors; one or more stereoscopic cameras; and one or more inertial measurement units. <Claim 3> The artificial intelligence module further comprises: one or more neural networks adapted to detect objects of interest; one or more foundation models adapted to detect objects of interest; the one or more neural networks are configured to process only color images; the one or more neural networks are configured to process only the point cloud; the one or more neural networks are configured to process only the depth maps; the one or more neural networks are configured to process only inertial data; the one or more foundation models are configured to process the color imagery, point clouds, depth maps, and inertial data collectively; The system of claim 1 , wherein the one or more neural networks and underlying models each return results in the form of a predicted object label and a confidence level. <Claim 4> The neural network further comprises: one or more convolutional neural networks; one or more dense neural networks; and one or more recurrent neural networks. <Claim 5> The base model further comprises: The system of claim 3 including one or more vision transformers. <Claim 6> The enriched tensor data structure further comprises: Color image data; Point cloud data, Depth data and and inertial data. <Claim 7> 1. A method for detecting a state of a shovel during loading of mineral in a mining operation, comprising: defining one or more regions of interest; constructing an enriched tensor data structure; processing the enriched tensor using an artificial intelligence module; processing one or more results of said artificial intelligence module using weightings; generating an alert based on the results of said weighting mechanism. <Claim 8> The step of defining one or more regions of interest may further include: Defining a rectangular region within a field of view of a camera; defining an area bounded by two horizontal lines within the field of view of the camera; defining an area bounded by two perpendicular lines within the field of view of the camera; defining a depth-based region of a particular region within the field of view of the camera; Defining time segments; defining a region based on object detection by determining a center point of the object and applying minimum and maximum distances to the center point; 8. The method of claim 7, further comprising one or more of: defining a region enclosed by the detected objects. <Claim 9> The step of constructing the enriched tensor data structure further comprises: collecting raw data from at least an inertial measurement unit, a stereoscopic camera, a color imaging camera, and a LIDAR; defining a region of interest based on preset inputs or analyzing the collected raw data; and creating a data structure containing normalized raw data from within the region of interest. <Claim 10> 1. A method of training an artificial intelligence module for determining a state of an excavator during mineral loading in a mining operation, comprising: introducing a batch of training images into the neural network; receiving a predicted label set; applying an intersection of a test label set to the predicted labels; generating a set of error values; and adjusting the weights of the neural network by applying the error value. <Claim 11> 1. A method of analyzing sensory data to determine the condition of a shovel during mineral loading in a mining operation, comprising: defining one or more regions of interest containing data relating to one or more wear parts; defining one or more regions of interest containing data relating to one or more attachment points on the excavator; constructing one or more data structures based on the collected sensor data within the region of interest; generating a set of labels and confidence intervals for each data structure; configuring a weighting mechanism with a set of values; For each data structure, weighting a set of labels and confidence intervals; generating an alert based on the result of said weighting. <Claim 12> 1. A method of analyzing sensor data to determine a condition of an excavator during mineral loading in a mining operation, comprising: defining one or more regions of interest containing data relating to one or more wear parts; constructing one or more data structures based on the collected sensor data within the region of interest; generating a polygonal geometric representation of each wear part; generating a point cloud for each wear part; and calculating a wear level based on said geometric representation of the worn parts and known values ​​of unworn parts. <Claim 13> The method of claim 12 , wherein the polygonal geometric representation includes a contour, a centroid, and one or more axes. <Claim 14> 1. A method of analyzing sensor data to determine a condition of an excavator during mineral loading in a mining operation, comprising: defining one or more areas of interest containing data about several minerals within the shovel; constructing one or more data structures based on the collected sensor data within the region of interest; generating a surface map of the mineral contours; and calculating the volume and weight of the mineral based on the surface map and known values ​​of the geometry of the shovel. <Claim 15> 1. A method of analyzing sensor data to determine a condition of an excavator during mineral loading in a mining operation, comprising: defining one or more regions of interest containing data about a plurality of mineral particles resting on a surface within the shovel; constructing one or more data structures based on the collected sensor data within the region of interest; generating a polygonal geometric representation of each particle; generating a point cloud for each particle; and calculating a size of each particle based on the geometric representation and the point cloud. <Claim 16> The method of claim 15 , wherein the polygonal geometric representation includes a contour, a centroid, and one or more axes.

Claims

1. 1. An AI based monitoring system for use in detecting the condition of a shovel during mineral loading in a mining operation, comprising: one or more sensors; an enriched tensor data structure; An artificial intelligence module; a weighting mechanism; and one or more outputs.

2. The sensor further comprises: one or more color imaging cameras; one or more LIDAR sensors; one or more stereoscopic cameras; and one or more inertial measurement units.

3. The artificial intelligence module further comprises: a plurality of neural networks adapted to detect objects of interest; one or more foundation models adapted to detect objects of interest; the one or more neural networks are configured to process only color images; the one or more neural networks are configured to process only the point cloud; the one or more neural networks are configured to process only the depth maps; the one or more neural networks are configured to process only inertial data; the one or more foundation models are configured to process the color imagery, point clouds, depth maps, and inertial data collectively; The system of claim 1 , wherein the one or more neural networks and underlying models each return results in the form of a predicted object label and a confidence level.

4. The neural network further comprises: one or more convolutional neural networks; one or more dense neural networks; and one or more recurrent neural networks.

5. The base model further comprises: The system of claim 3 including one or more vision transformers.

6. The enriched tensor data structure further comprises: color image data captured by a color image camera; Point cloud data captured by a LIDAR sensor; depth data captured by a stereo camera; and and inertial data captured by an inertial measurement unit.

7. 1. A method for detecting a state of a shovel during loading of mineral in a mining operation, comprising: defining one or more regions of interest; constructing an enriched tensor data structure; processing the enriched tensor using an artificial intelligence module; processing one or more results of said artificial intelligence module using weightings; generating an alert based on the results of said weighting mechanism.

8. The step of defining one or more regions of interest may further include: Defining a rectangular region within a field of view of a camera; defining an area bounded by two horizontal lines within the field of view of the camera; defining an area bounded by two perpendicular lines within the field of view of the camera; defining a depth-based region of a particular region within the field of view of the camera; Defining time segments; defining a region based on object detection by determining a center point of the object and applying minimum and maximum distances to the center point; The method of claim 7 , further comprising one or more of: defining a region enclosed by the detected objects.

9. The step of constructing the enriched tensor data structure further comprises: collecting raw data from at least an inertial measurement unit, a stereo camera, a color imaging camera, and a LIDAR; defining a region of interest based on preset inputs or analyzing the collected raw data; and creating a data structure containing normalized raw data from within the region of interest.

10. 1. A method of training an artificial intelligence module for determining a state of an excavator during mineral loading in a mining operation, comprising: introducing a batch of data-augmented training images into a neural network; receiving a predicted label set; applying an intersection of a test label set to the predicted labels; generating a set of error values; and adjusting the weights of the neural network by applying the error value.

11. 1. A method of analyzing sensor data to determine a condition of an excavator during mineral loading in a mining operation, comprising: defining one or more regions of interest containing data relating to one or more wear parts; defining one or more regions of interest containing data relating to one or more attachment points on the excavator; constructing one or more data structures based on the collected sensor data within the region of interest; generating a set of labels and confidence intervals for each data structure; configuring a weighting mechanism with a set of values; For each data structure, weighting a set of labels and confidence intervals; generating an alert based on the result of said weighting.

12. 1. A method of analyzing sensor data to determine a condition of an excavator during mineral loading in a mining operation, comprising: defining one or more regions of interest containing data relating to one or more wear parts; constructing one or more data structures based on collected sensor data within the region of interest, including LIDAR data; generating a polygonal geometric representation of each wear part; generating a point cloud for each wear part; and calculating a wear level based on said geometric representation of the worn parts and known values ​​of unworn parts.

13. The method of claim 12 , wherein the polygonal geometric representation includes a contour, a centroid, and one or more axes.

14. 1. A method of analyzing sensor data to determine a condition of an excavator during mineral loading in a mining operation, comprising: defining one or more areas of interest containing data about several minerals within the shovel; constructing one or more data structures based on the collected sensor data within the region of interest; generating a surface map of the mineral contours; and calculating the volume and weight of the mineral based on the surface map and known values ​​of the geometry of the shovel.

15. 1. A method of analyzing sensor data to determine a condition of an excavator during mineral loading in a mining operation, comprising: defining one or more regions of interest containing data regarding a plurality of mineral particles resting on a surface within the shovel; constructing one or more data structures based on the collected sensor data within the region of interest; generating a polygonal geometric representation of each particle; generating a point cloud for each particle; and calculating a size of each particle based on the geometric representation and the point cloud.

16. The method of claim 15 , wherein the polygonal geometric representation includes a contour, a centroid, and one or more axes.

17. The system of claim 3 , wherein the weighting mechanism further comprises a set of custom thresholds to arbitrate results returned by the plurality of one or more neural networks.

18. Data augmentation also Mirroring the training images; tilting the training image horizontally by up to 15 degrees, Zooming in or out of the training images, and adjusting the contrast of the training images; The method of claim 10, comprising one or more of:

19. The system of claim 1 , wherein the system is located in proximity to the sensor.