Mining truck difficult case identification model training method, difficult case identification method and device

By generating and expanding a set of difficult example samples and training the Transformer model, the problem of insufficient difficult example recognition in the mining truck autonomous driving system was solved, and the model's recognition ability and safety were improved.

CN120766058APending Publication Date: 2025-10-10ZHONGKE YUNGU TECH
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510882518.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

Existing technologies have difficulty effectively identifying and collecting difficult examples in mining trucks, resulting in insufficient training data for autonomous driving systems and affecting the model's recognition capabilities.

Method used

The first sample set is generated by acquiring collected data, real difficult examples are screened out and expanded to generate simulated difficult examples, a Transformer model is constructed for training, and a difficult example recognition model is generated.

Benefits of technology

The accuracy of difficult-example recognition and multi-target perception capabilities have been improved, enhancing the safety of autonomous driving for mining trucks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120766058A_ABST
    Figure CN120766058A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a mining truck difficulty identification model training method, a difficulty identification method and equipment. The method comprises the following steps that collection data are obtained, a first sample set is generated according to the collection data, and the first sample set comprises normal cases and real difficult cases; performing difficult case expansion processing on the real difficult case to generate a simulation difficult case; summarizing the real hard cases and the simulation hard cases into hard cases, and summarizing the hard cases and the normal cases into a second sample set; and training the initial model by using the second sample set, and marking the trained initial model as a difficult case identification model. Therefore, according to the method and the device, the existing collected data can be utilized to simulate and generate more difficult cases, and the difficult cases in the original sample set are replaced or added, so that the expansion of the difficult cases is realized, and the quality of the sample data is improved. And subsequent model training can improve difficult case identification and multi-target perception capabilities, and is helpful for popularization of mine card automatic driving.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of automatic control technology, and in particular relates to a mining truck difficult case identification model training method, a mining truck difficult case identification method, a computer device and a computer-readable storage medium. Background Art

[0002] Autonomous mining trucks (abbreviated as mining trucks) typically achieve environmental perception through data acquisition equipment such as lidar, millimeter-wave radar, and cameras. By retrofitting autonomous driving hardware and software packages onto mining trucks, they function as autonomous driving data collection vehicles. Currently, mainstream autonomous driving systems rely on models to perceive the surrounding environment, requiring extensive training data to enhance the system's perception and control capabilities. However, in the real world, due to differences in sensor characteristics and the impact of multiple targets, the system cannot fully identify the objects detected by the sensors. This results in unknown and difficult scenarios that the algorithm cannot recognize, leading to problems. These unknown and difficult scenarios that the algorithm cannot recognize are called difficult examples, or simply hard cases. Identifying hard cases is crucial to the autonomous driving training process. Effectively identifying hard cases allows the autonomous driving system to avoid corresponding risks and ensure autonomous driving safety.

[0003] However, difficult examples often occur in fragmented and extreme scenarios. Therefore, collecting data for these scenarios has become a key issue in training autonomous driving algorithms. Currently, difficult data samples for mining trucks are collected primarily by having dedicated data collectors drive the trucks repeatedly along the road section where the problem occurred, manually turning on the data collection switch on the trucks to collect data. Furthermore, when the autonomous driving system encounters an emergency brake signal (or similar specific rules), it automatically collects data for the short period before and after the emergency stop. This collection process can miss difficult examples, resulting in insufficient data for model training. Therefore, how to obtain more difficult examples and train models is a pressing technical challenge for those skilled in the art.

[0004] The preceding description is intended to provide general background information and does not necessarily constitute prior art. Summary of the Invention

[0005] Based on this, it is necessary to address the above problems and propose a mining truck difficult example recognition model training method, a mining truck difficult example recognition method, a computer device and a computer-readable storage medium, which can effectively simulate more difficult examples for training.

[0006] The present application solves the technical problem by adopting the following technical solutions: The present application provides a method for training a difficult example recognition model for a mining truck, the method comprising the following steps: acquiring collected data, generating a first sample set based on the collected data, the first sample set comprising normal examples and real difficult examples; performing difficult example expansion processing on the real difficult examples to generate simulated difficult examples; aggregating the real difficult examples and the simulated difficult examples into difficult examples, and aggregating the difficult examples and the normal examples into a second sample set; using the second sample set to train an initial model, and marking the trained initial model as a difficult example recognition model.

[0007] In an optional embodiment of the present application, a first sample set is generated based on the collected data, including: obtaining labels in the collected data, and distinguishing the collected data into normal examples and real hard examples based on the labels; and / or, screening suspicious scene data in the collected data, marking the collected data containing suspicious scene data as real hard examples, and marking the remaining collected data as normal examples.

[0008] In an optional embodiment of the present application, the collected data includes multiple frames of point cloud data and image data; screening suspicious scene data in the collected data includes: determining the time matching relationship between each frame of point cloud data and image data; determining the spatial matching relationship between the point cloud data and the image data through preset calibration parameters; determining the matching relationship between the image data and the point cloud data based on the time matching relationship and the spatial matching relationship; based on the matching relationship, processing each frame of image data using an edge detection and segmentation algorithm to determine at least one geometric target; performing cluster analysis on each frame of point cloud data to determine at least one three-dimensional object; calculating the overlap between each geometric target and the three-dimensional object in each frame; when the overlap meets the suspicious condition, the collected data of this frame is screened as suspicious scene data.

[0009] In an optional embodiment of the present application, a difficulty case expansion process is performed on a real difficult case to generate a simulated difficult case, including: determining the difficulty case type of the real difficult case, the difficulty case type including at least one of a rain and snow difficult case and a dust and fog difficult case; performing a first sub-expansion process on the rain and snow difficult case to obtain a plurality of simulated rain and snow difficult cases; and / or performing a second sub-expansion process on the dust and fog difficult case to obtain a plurality of simulated dust and fog difficult cases; and summarizing all simulated rain and snow difficult cases and all simulated dust and fog difficult cases as simulated difficult cases.

[0010] In an optional embodiment of the present application, a first sub-expansion process is performed on the rain and snow difficult examples to obtain multiple simulated rain and snow difficult examples, including: determining the rain and snow intensity value of each rain and snow difficult example, screening the rain and snow difficult example with the smallest rain and snow intensity value and marking it as the first sample difficult example; obtaining the point cloud data in the first sample difficult example and marking it as the first sample point cloud; determining the rain and snow interference point cloud in the first sample point cloud and marking the rain and snow interference point cloud as the first unit point cloud; adjusting the point cloud data of each rain and snow difficult example according to the first unit point cloud, and marking the adjusted rain and snow difficult example as a simulated rain and snow difficult example.

[0011] In an optional embodiment of the present application, the second sub-expansion processing is performed on the dust fog difficult case to obtain a plurality of simulated dust fog difficult cases, including: obtaining the acquisition data corresponding to each dust fog difficult case, marked as dust fog data; determining the modeling parameters of each dust fog difficult case according to the dust fog data; determining the interference calculation formula according to the modeling parameters; fitting the dust fog data to generate a plurality of simulated acquisition data by using the interference calculation formula; and generating at least one simulated dust fog difficult case according to the simulated acquisition data.

[0012] In an optional embodiment of the present application, the method further comprises: constructing an initial model, the initial model adopting a Transformer model architecture; training the initial model by using the second sample set, and marking the trained initial model as a difficult case recognition model, including: inputting the normal cases and the difficult cases in the second sample set into the initial model for training according to a preset ratio; determining whether the initial model of the current round meets the training completion condition according to the training result; the training result is obtained by the initial model after each round of training; if the training completion condition is not met, adjusting the model parameters of the initial model of the current round according to the training result, and entering the next round of training; or, if the training completion condition is met, it is determined that the initial model of the current round is trained, the initial model of the current round is marked as the difficult case recognition model and output.

[0013] The present application also provides a mine truck difficult case recognition method, the method comprising: obtaining and deploying a difficult case recognition model, the difficult case recognition model being trained by the method provided in the foregoing; inputting the collected detection data into the difficult case recognition model for processing, and obtaining the processing result output by the difficult case recognition model, the processing result being used to represent whether the mine truck is currently in a difficult scene corresponding to a difficult case.

[0014] The present application also provides a computer device comprising a processor and a memory: the processor is used to execute a computer program stored in the memory to implement the method as described in the foregoing.

[0015] The present application also provides a computer readable storage medium storing a computer program, when the computer program is executed by a processor, the method as described in the foregoing is implemented.

[0016] By adopting the embodiments of the present application, the following beneficial effects are achieved: The present application can simulate more difficult cases by using the existing acquisition data, replace or increase the difficult cases in the original sample set, expand the difficult cases, and thus improve the quality of the sample data. The subsequent model training can improve the ability of recognizing difficult cases and multi-target perception, which is helpful for the popularization of mine truck automatic driving.

[0017] The above description is only an overview of the technical solution of this application. In order to more clearly understand the technical means of this application, which can be implemented in accordance with the contents of the description, and to make the above and other purposes, features and advantages of this application more obvious and easy to understand, the following preferred embodiments are specifically described in detail with reference to the accompanying drawings. It should be understood that the above general description and the detailed description below are only exemplary and explanatory and do not limit this application. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0019] Figure 1 A flowchart of a method for training a difficult-example recognition model for mining trucks is provided in one embodiment.

[0020] Figure 2 A flowchart of a method for identifying difficult examples of mining trucks is provided in one embodiment.

[0021] Figure 3 The present invention is a schematic block diagram of the structure of a computer device provided by an embodiment. DETAILED DESCRIPTION

[0022] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0023] The autonomous driving applied to mining trucks requires not only the installation of autonomous driving software and hardware kits, but also the use of autonomous driving data collection vehicles to collect data for training. The autonomous driving kit package generally includes IMU (Inertial Measurement Unit), perception cameras, lidar, GPS and other sensors, intelligent driving domain controller, and autonomous driving system. The current collection of difficult data samples for mining trucks is mainly done by having a dedicated collector drive the vehicle multiple times on the road section where the problem occurs after a fault occurs, manually turning on the data collection switch on the vehicle, and collecting data. In addition, when the mining truck autonomous driving system encounters an emergency braking signal (specific rules), it automatically collects data for a period of time before and after the emergency stop. Therefore, it will lead to the problem of missing difficult examples, and the lack of difficult examples will directly affect the quality of autonomous driving training. For this reason, how to correctly identify difficult examples is an urgent problem to be solved. Based on this, this application proposes a method for training a difficult example recognition model for mining trucks. By expanding the existing collected data, simulating and generating more difficult examples, and training, the recognition accuracy of difficult examples can be improved. In order to clearly describe the method provided in this embodiment, please refer to Figure 1 , including steps S110~S130.

[0024] Step S110: Acquire collected data, and generate a first sample set based on the collected data. The first sample set includes normal examples and real difficult examples.

[0025] In one embodiment, the collected data can be acquired by controlling a collection device. The collection device is equipped with various sensors, including but not limited to lidar, perception cameras, millimeter-wave radar, infrared thermal imaging cameras, and intelligent driving domain controllers. The collected data is collectively referred to as collected data and is specifically divided into image data and point cloud data. Image data can include visible light images captured by perception cameras and point cloud data captured by infrared thermal imaging cameras. Data collected by lidar and millimeter-wave radar is point cloud data. Regarding the specific data acquisition process, the collection device can be controlled to travel to a designated area, such as a specific working area of ​​a mining truck, along the mining truck's operating route. Data from each road section, under different weather conditions, and at different times along this route is collected as collected data. Infrared cameras are used for data collection. Infrared cameras provide thermal imaging, meaning that the image data includes infrared images. Medium-wave infrared cameras have excellent anti-interference capabilities and can adapt to interference factors in different environments, such as rain, snow, dust, and haze. By comparing whether objects appear different under different sensor conditions, it can more accurately detect environments with significant sensor influence.

[0026] In one embodiment, step S110: generating a first sample set based on the collected data includes: obtaining labels in the collected data, and distinguishing the collected data into normal examples and real hard examples based on the labels; and / or, screening suspicious scene data in the collected data, marking the collected data containing suspicious scene data as real hard examples, and marking the remaining collected data as normal examples.

[0027] In one embodiment, the first sample set includes normal cases and real difficult cases. It is understandable that the collected data, including point cloud data and image data, are usually continuous data used to characterize the environment in which the collection equipment is located. Depending on the environment, there may be normal cases and difficult cases. The normal case refers to the environmental conditions under which the mining truck is in normal working condition. The difficult cases are specifically for situations that will affect the operation of the mining truck, which may include but are not limited to: dust and fog blocking the line of sight, extreme weather such as rain and snow, multiple obstacles, dynamic obstacles, etc. Therefore, the collected data segments obtained when the mining truck is in normal working condition are called normal cases; conversely, the collected data segments obtained when the mining truck is in abnormal condition are called difficult cases. It is also understandable that under normal circumstances, the mining truck is far more likely to be in normal cases than in difficult cases.

[0028] Obtain labels from the collected data and, based on the labels, distinguish the collected data into normal examples and true hard examples. These labels may be manually pre-labeled after the collected data is obtained. These labels are used to distinguish the types of a particular collected data segment, specifically, normal examples and true hard examples. Specifically, the normal examples and true hard examples in the first sample set may be the collected data itself or feature data extracted from the collected data. This is not a limitation.

[0029] In addition to manually pre-labeling the collected data, existing AI recognition tags can also be used, or specific algorithms can be used for verification. If a segment of the collected data matches the conditions of a difficult scenario, it can be labeled as a difficult example, while the rest are labeled as normal examples.

[0030] In one embodiment, the collected data includes multiple frames of point cloud data and image data; screening suspicious scene data in the collected data includes: determining a time matching relationship between each frame of point cloud data and image data; determining a spatial matching relationship between the point cloud data and the image data through preset calibration parameters; determining a matching relationship between the image data and the point cloud data based on the time matching relationship and the spatial matching relationship; processing each frame of image data using an edge detection and segmentation algorithm based on the matching relationship to determine at least one geometric target; performing cluster analysis on each frame of point cloud data to determine at least one three-dimensional object; calculating the overlap between each geometric target and the three-dimensional object in each frame; when the overlap meets a suspicious condition, the collected data of this frame is screened as suspicious scene data.

[0031] In one embodiment, to filter out the data collected for difficult scenes, the point cloud data and image data in the collected data can be calibrated first to unify the spatiotemporal relationship between the two. Both the point cloud data and the image data are collected continuously, with the minimum unit being one frame, which can be 30 frames or 60 frames. The specific number of frames depends on the acquisition frequency of the acquisition device. In addition, the number of frames per unit time for the point cloud data and the image data can also be different. To this end, it is necessary to synchronize the temporal relationship between the point cloud data and the image data. Specifically, both the point cloud data and the image data carry timestamps. By synchronizing the point cloud data and the image data with the same timestamp, the temporal matching relationship between the two can be determined regardless of whether the frame numbers are the same.

[0032] In addition, it is understandable that the radar that collects point cloud data and the camera that collects image data are definitely not installed in the same spatial position on the collection equipment, and there must be spatial differences between the two. For the same object, there must be differences in their respective construction coordinate systems. In order to eliminate this difference, preset calibration parameters can be obtained to determine the spatial matching relationship between the point cloud data and the image data. The calibration parameters are used to represent the spatial relationship between the radar and the camera in the real space. The calibration parameters can be used to unify the coordinate systems of the point cloud data and the image data and determine them in the same spatial coordinate system. By summarizing the time matching relationship and the space matching relationship, the spatiotemporal relationship between the point cloud data and the image data is completed, that is, the matching relationship between the two is determined.

[0033] Based on the matching relationships, the image data and point cloud data are processed separately. For image data, an edge detection and segmentation algorithm is used to segment the image by identifying and extracting edge information. Specifically, the Canny operator is used to detect edges in the image through a series of steps, effectively identifying geometric objects in the image and separating them from the background. After segmentation, each identified geometric object can be classified and labeled, for example, to determine whether it is a pedestrian, an obstacle, another mobile mining truck, or other transport equipment.

[0034] For point cloud data, the data can be rasterized and the height of each grid cell calculated. If the height difference is less than a preset threshold, it is considered a road grid; otherwise, it is considered an object grid. Road grids are ignored. For all object grids, the density-based DBSCAN (Density-Based Spatial Clustering of Applications with Noise) clustering algorithm can be used to calculate the 3D structure and outline of all objects in the point cloud from the object grids. Each 3D object is labeled using this 3D structure and outline.

[0035] For the process of calculating the overlap, the set of geometric targets determined by each frame of image data can be marked as A; based on the matching relationship, the 3D data of the 3D object is directly transformed into a 2D projection to obtain the 2D geometric object outline, and the geometry of each 2D geometric object outline is marked as B. Assuming A is used as the reference, find the matching target in B. Taking the kth geometric object in A as an example, first calculate A k The geometric center of B, and then calculate the distance between the center of each geometric object in A k The distance from the center of the object to determine the nearest geometric object Bi to Ak. Calculate A k and B i The overlap between A and B is calculated, for example, by calculating the intersection over union (IOU). The overlap is used to determine the matching between A and B. Assuming there are S2 geometric objects in A, any geometric objects with an overlap above a preset threshold (for example, an IOU above 80%) are marked as matched. Assuming there are S1 matched objects, the number of unmatched objects is S2 - S1. The proportion of unmatched objects in each frame can be determined using the formula (S2 - S1) / S2. If the proportion of unmatched objects in a frame exceeds the preset upper threshold, the suspicious condition is determined to be met, and the captured data for this frame is filtered as suspicious scene data.

[0036] For suspicious scene data in the collected data, the suspicious scene data can be processed as real difficult examples according to the processing method of obtaining real difficult examples based on labels as mentioned above; for other data in the collected data, they can be processed as normal examples according to the appointment method. The real difficult examples and normal examples are summarized to obtain the first sample set. It is worth noting that marking suspicious scene data as difficult examples does not mean identifying difficult examples. It only distinguishes between situations that may be difficult examples. This processing flow is to screen out situations that may be difficult examples in order to facilitate expansion and training. The recognition accuracy cannot be compared with manual labeling or the recognition model obtained by training later. This processing process is only a preliminary screening method for expansion. Moreover, the two are not mutually exclusive. The two processes of determining the first sample set, through label processing and through suspicious scene data processing, can be executed separately and then summarized.

[0037] Step S120: performing difficult example expansion processing on real difficult examples to generate simulated difficult examples; summarizing the real difficult examples and simulated difficult examples into difficult examples, and summarizing the difficult examples and normal examples into a second sample set.

[0038] In an embodiment, the difficult case expansion processing is performed on the real difficult case to generate the simulated difficult case, including: determining a difficult case type of the real difficult case, the difficult case type including at least one of a rain and snow difficult case and a dust and fog difficult case; performing a first sub-expansion processing on the rain and snow difficult case to obtain a plurality of simulated rain and snow difficult cases; and / or performing a second sub-expansion processing on the dust and fog difficult case to obtain a plurality of simulated dust and fog difficult cases; and aggregating all the simulated rain and snow difficult cases and all the simulated dust and fog difficult cases as the simulated difficult case.

[0039] In an embodiment, as described above, the difficult case can include a plurality of cases in reality, and can be classified into the following categories: 1. extreme / complex weather category; 2. complex terrain / landscape category, such as muddy, potholes, and light-reflecting ore accumulation area; 3. dynamic interference category, such as other moving mining trucks, transportation equipment, and personnel walking; 4. sensor abnormality or interference category, i.e., the case where the detection equipment randomly fails or is blocked and cannot work normally; and 5. abnormal driving behavior category (artificial or unexpected behavior), such as accidental touch and error control after manual takeover. For each category, a corresponding simulation method can be adopted for expansion. In the preferred embodiment of the present application, the extreme / complex weather category difficult case is expanded, wherein the extreme / complex weather category difficult case includes the following cases. Rain and snow weather, which is characterized by visual blurring, lens water droplet interference, ground reflection leading to misidentification, scene occlusion, contour blurring, and color change. Dust and fog, i.e., dust and haze, which greatly shortens the visible distance and makes it difficult for the camera / laser radar to identify the target. Light interference, such as strong light, backlight, and low light at night, which can cause visual overexposure or contour blurring of image acquisition and loss of target edge features. Among them, the light interference category mainly interferes with image data, and the simulation method is relatively simple, which will not be described in detail. The rain and snow difficult case and the dust and fog difficult case are mainly expanded.

[0040] As described above, there are various types of difficult cases, and the present application prioritizes the expansion of the rain and snow difficult case and the dust and fog difficult case. Therefore, the difficult cases need to be classified to determine the difficult case type of the real difficult case, so as to screen out the rain and snow difficult case and the dust and fog difficult case in the extreme / complex weather category and process them to achieve simulation and expansion. Other difficult cases can be processed according to the corresponding simulation generation method, which can refer to the simulation method of the rain and snow difficult case and the dust and fog difficult case, and will not be described in detail.

[0041] In an embodiment, the first sub-expansion processing is performed on the rain and snow difficult instance to obtain a plurality of simulated rain and snow difficult instances, including: determining the rain and snow intensity value of each rain and snow difficult instance, screening the rain and snow difficult instance with the smallest rain and snow intensity value as the first sample difficult instance; obtaining the point cloud data in the first sample difficult instance, and marking it as the first sample point cloud; determining the rain and snow interference point cloud in the first sample point cloud, and marking the rain and snow interference point cloud as the first unit point cloud; and adjusting the point cloud data of each rain and snow difficult instance according to the first unit point cloud, and marking the adjusted rain and snow difficult instance as the simulated rain and snow difficult instance.

[0042] In an embodiment, the expansion of the rain and snow difficult instance is implemented by performing the first sub-expansion processing. The rain and snow difficult instance has the following characteristics: the falling track of rain and snow is affected by gravity and air resistance, the speed is related to the diameter of raindrops or snowflakes (Stokes law), and it is represented as an inclined short line in the image (when the wind speed is not zero). The spatial distribution: the intensity of precipitation (mm / h) determines the density of raindrops or snowflakes, and large-particle raindrops are more sparse. Based on this, the rain and snow difficult instance can be simulated according to the intensity of precipitation, that is, the rain and snow intensity value. Specifically, the rain and snow intensity value of each rain and snow difficult instance can be directly obtained from the meteorological bureau or calculated by processing the collected data with a specific algorithm. The specific intensity value can be set in a stepped level, for example, high, medium and low, or it can be a stepless numerical design, for example, any numerical value between 1 and 10, which is not limited. For the case of obtaining the rain and snow intensity value from the meteorological bureau, not only the numerical value can be obtained, but also the rain and snow range and influence change in the collection time corresponding to the collected data, so that the simulation is better.

[0043] After determining the rain and snow intensity value, the rain and snow difficult instance with the smallest rain and snow intensity value is screened and marked as the first sample difficult instance, and the point cloud data in the first sample difficult instance is obtained and marked as the first sample point cloud. It can be known that rain and snow will mainly interfere with the point cloud data as an entity particle in the falling process. Therefore, if simulation is needed, the influence of the rain and snow particles on the point cloud data needs to be learned and understood, that is, the rain and snow interference point cloud is screened from the point cloud data. It can be understood that the higher the rain and snow intensity value, the stronger the interference with the point cloud data, and the accuracy of the obtained point cloud data and the rain and snow interference point cloud in the point cloud data will be reduced. Therefore, in order to learn the interference of the rain and snow particles on the point cloud data well, in a preferred embodiment, the point cloud data with the lowest rain and snow intensity value is used as a learning sample for simulation.

[0044] The rain and snow interference point cloud is filtered out from the point cloud data, and the core characteristics of the rain and snow particles are counted, including but not limited to particle size spatial distribution, velocity distribution, spatial density, distance mapping distribution, etc. For the particle size spatial distribution, the Marshall-Palmer distribution can be fitted according to the statistical rain and snow particle size spatial distribution, and the rain and snow particle distribution function coefficient is obtained. The rain and snow particle distribution function coefficient is used to describe that under a specified rain and snow intensity value, the number concentration of rain and snow particles with a particle size of D per unit volume. In the radar field of view, a set of initial positions (x, y, z) of rain and snow particles is generated according to the three-dimensional Poisson distribution, the density decreases with the vertical height (the attenuation rate is fitted according to the measured collection data statistics), and a continuous rain and snow point cloud frame is simulated according to the velocity distribution corresponding to the rain and snow intensity value. According to the simplified Mie scattering formula, the reflection intensity and the distance and particle size attenuation coefficient σ are obtained based on the measured mapping distribution fitting. According to the particle size density, the rain and snow particle size, the distance between the rain and snow particles and the radar, and the laser attenuation coefficient, the reflection intensity (x, y, z, r) of each point is obtained, 5% outliers and Gaussian noise with the same signal-to-noise ratio as the original sample are added to obtain the rain and snow interference point cloud under the minimum rain and snow intensity value.

[0045] The rain and snow interference point cloud is marked as the first unit point cloud, and through the first unit point cloud, the point cloud data of each rain and snow difficult example can be adjusted. The specific process can be that the original rain and snow difficult example point cloud data is filtered out of noise, and the rain and snow interference point cloud is also screened out and replaced by a preset multiple of the first unit point cloud; or a preset multiple of the first unit point cloud is added or filtered out based on the original rain and snow interference point cloud. Thus, based on the original point cloud data, more point cloud data is simulated and expanded, and each expanded point cloud data can generate a simulated rain and snow difficult example. For each simulated rain and snow difficult example, a corresponding label can be added, for example, it is distinguished from the real difficult example, and the simulated rain and snow intensity value and the obstacle information contained therein are labeled, etc., for subsequent training.

[0046] It is worth noting that although the rain and snow particles are processed for rain and snow, they can actually be simulated separately, and the corresponding processing process is basically processed in the manner described in the foregoing, that is, different intensity values and rain and snow particle point cloud data are simulated separately. That is, for rain conditions, only raindrop particles are simulated; for snow conditions, only snow particles can be simulated; and for sleet weather, rain and snow can be included at the same time, but for sleet conditions, the intensity of each needs to be adjusted and controlled to ensure that the finally simulated simulated difficult example meets the real situation. In addition, for particle simulation, the difference between snowflakes and raindrops can also be considered to simulate more accurate point cloud data.

[0047] In one embodiment, a second sub-expansion process is performed on the dust and fog difficulty case to obtain multiple simulated dust and fog difficulty cases, including: obtaining collected data corresponding to each dust and fog difficulty case and marking it as dust and fog data; determining modeling parameters of each dust and fog difficulty case based on the dust and fog data; determining an interference calculation formula based on the modeling parameters; using the interference calculation formula, fitting the dust and fog data to generate multiple simulated collected data; and generating at least one simulated dust and fog difficulty case based on the simulated collected data.

[0048] In one embodiment, for the simulation of the dust and fog difficulty case, it is necessary to perform a second sub-extension for processing. The dust and fog difficulty case is aimed at the real situation, such as dust, haze, etc., and has the following characteristics: it is affected by wind speed, terrain roughness, and vehicle movement, and the particle concentration decays with distance (in line with Mie scattering theory). Dynamic characteristics: the particles are dense in the near field (around the mine car), sparse in the far field, and diffuse over time (Gaussian diffusion model). The simulation of the dust and fog difficulty case will mainly interfere with the image data, and will also affect the radar detection range of the point cloud data. To this end, it is necessary to obtain the collected data corresponding to each dust and fog difficulty case and mark it as dust and fog data. By analyzing the dust and fog data, the modeling parameters of each dust and fog difficulty case are determined. The modeling parameters include but are not limited to dust and fog particle size distribution, vertical attenuation coefficient of concentration, horizontal distribution of dust and fog, etc. According to the modeling parameters, the interference calculation formula is determined. The interference calculation formula is used to characterize the four-dimensional factors of particle size-laser wavelength-distance-angle, which simplifies the description of the scattering intensity of laser irradiation to non-spherical particles. The calculation formula can be referred to as follows: (1) In the above formula, is the laser scattering intensity; is the size of the particles; Laser wavelength dependence function, reflecting the influence of laser wavelength; is the distance between the observation point and the particle; is the angle distribution function, which represents the radar azimuth and polar angle Therefore, it is necessary to model the dust and mist particles within the radar field of view. In a preferred embodiment, the dust and mist particles can be non-spherical (such as ellipsoids or polyhedrons), with a wide range of particle sizes (micrometers to millimeters), and distributed in accordance with the lognormal distribution or the Rosin-Rammler distribution. Randomly generate particle equivalent diameters , and then through the aspect ratio (major axis / minor axis) are further modeled as non-spherical particles. Using interference calculation formulas, the dust cloud data is fitted to generate multiple simulated data sets. The fitting process can be performed as described above by setting a dust cloud intensity value to adjust the density of dust cloud particles, thereby reversely calculating the data collected at that dust cloud intensity value. The dust cloud intensity value can be superimposed on or removed from existing dust cloud examples, meaning that simulated data is superimposed on or removed from the collected data to generate more examples. Alternatively, a replacement method can be used to filter out noise from the collected data containing dust cloud examples, replacing the original dust point cloud with simulated data at a preset dust cloud level to generate examples at a specified dust cloud intensity value. The specific simulation method is not limited; as long as more simulated data can be fitted, and at least one simulated dust cloud example is generated based on the simulated data, it will suffice.

[0049] Furthermore, for both the first and second sub-expansion processes, the fitting and superposition process can simulate the attenuation of lidar parameters. This takes into account two real-world scenarios: 1. Power attenuation; 2. Sensor noise increase. For the former, all of the aforementioned data can be adjusted by adjusting the attenuation coefficient of laser reflection intensity with distance (for example, increasing the attenuation by 20%) to simulate data generated when laser power attenuates. For the latter, the signal-to-noise ratio distribution in the original acquired data can be statistically analyzed. By increasing the preset mean and variance (for example, by 20%), a new noise distribution is formed, replacing all noise in the data with Gaussian noise that conforms to the new distribution. This results in more realistic simulation data, which can be used to generate difficult simulation examples.

[0050] After the above processing, more simulated difficult examples are obtained in the case of real difficult examples. The real difficult examples and simulated difficult examples are collectively referred to as difficult examples. The difficult examples and normal examples are summarized to obtain a second sample set with a larger number of samples.

[0051] Step S130: using the second sample set to train the initial model, and marking the trained initial model as a difficult example recognition model.

[0052] In one embodiment, the method also includes: constructing an initial model, the initial model adopts a Transformer model architecture; using the second sample set to train the initial model, and marking the trained initial model as a difficult example recognition model, including: inputting normal examples and difficult examples in the second sample set into the initial model for training according to a preset ratio; judging whether the initial model of the current round meets the training completion conditions based on the training results; the training results are output by the initial model after each round of training; if the training completion conditions are not met, adjusting the model parameters of the initial model of the current round according to the training results, and entering the next round of training; or, if the training completion conditions are met, determining that the initial model of the current round has completed training, marking the initial model of the current round as a difficult example recognition model and outputting it.

[0053] In an embodiment, for the initial model and the hard example identification model output after training is completed, they are the same model architecture, the difference is only that the latter is the model after training iteration and parameter adjustment. The model adopts a Transformer model architecture, including an input embedding layer, a Transformer encoder layer, and an output layer. The normal examples and hard examples in the second sample set are input into the initial model for training according to a preset ratio. The ratio can be 7:3 for normal examples:hard examples.

[0054] After each round, it is judged whether the initial model of the current round meets the training completion condition according to the training result output after each round of training. The initial model outputs the training result of the current round based on the loss function value, recognition accuracy, convergence degree index, etc. in the current training round. The training result can specifically include the recognition number of each object of the initial model for normal examples and hard examples, recognition accuracy, etc. According to the training result, the system judges whether the current initial model meets the training completion condition. The training completion condition can include but is not limited to: the model loss value converges to a preset threshold, the verification set accuracy reaches a preset upper limit of precision, the training round reaches a maximum iteration number, or the recognition recall rate of the hard example is stable, etc. The indicators can be arbitrarily set according to the actual demand, and the present application does not make specific limitations. If it is judged that the training completion condition is not met, the parameters of the initial model are updated based on the training result of the current round. The update method can include gradient descent, momentum update, adaptive learning rate adjustment, and other optimization strategies, so that the model better fits the training data in the next round of training. If it is judged that the training completion condition is met, it is determined that the initial model of the current round has the expected hard example identification capability, and it is marked as a "hard example identification model". The model can be subsequently deployed on a mining truck to perform tasks such as hard example discrimination, classification optimization, sample playback, etc. on actual samples.

[0055] For the trained hard example identification model, the present application also proposes a mining truck hard example identification method to realize the application of the model. For a clear description of the method, please refer to Figure 2 , including steps S210-S220.

[0056] Step S210: Obtain and deploy the hard example identification model.

[0057] In an embodiment, the hard example identification model is trained by the method provided in the foregoing, and the specific training process is described in the foregoing, which will not be repeated here.

[0058] Step S220: input the collected detection data into the hard example identification model for processing, and obtain the processing result output by the hard example identification model. The processing result is used to represent whether the mining truck is currently in a difficult scene corresponding to a hard example.

[0059] In one embodiment, the detection data is a superordinate of the collected data, and in addition to the image data and point cloud data in the collected data, it also includes driving data, environmental data, etc. Driving data is specifically the operation data of the mining truck at various locations on the road. After pre-processing such as noise filtering, the average acceleration, average speed, and heading angle of the mining truck at each location are calculated according to a preset period, and the speed (3D), acceleration (3D), and heading angle of the mining truck in the difficult case are divided by the average value to obtain the speed deviation. s , acceleration offset l , heading angle deviation k , the speed deviation s , acceleration offset l and heading angle deviation k As driving data, environmental data includes weather station data in the mining area, mining area map road topology, etc.

[0060] In one embodiment, step S220: the step of inputting the collected detection data into the difficult example recognition model for processing also includes preprocessing the detection data, that is, preprocessing the collected data, driving data and environmental data respectively, to achieve normalization of the motion parameters and historical mean values ​​of the three, so that the model is easier to process.

[0061] The main processing step for collected data is point cloud data, which includes the following steps: 1. Perform single-frame point cloud pre-processing, including processing the point cloud data in the collected data, removing outliers, and filtering the height (i.e. retaining the point cloud above the ground). 2. Use lightweight PointNet (delete the deep network, retain the core feature extraction), input a single-frame point cloud, and output frame-level features after processing. f 1 (256-dimensional matrix), the continuous point cloud is transformed into f Secondly, extract the dynamic targets in the continuous point cloud frames. The first step is preprocessing, filtering out independent noise points through Kalman filtering. The second step is point cloud registration: Based on the IMU data, obtain the pose transformation matrix between adjacent frames ( R , T ). t Frame point cloud P t After the pose transformation matrix and sensor calibration parameters, it can be converted into a global frame. The third step is static background recognition. If the same target appears in more than half of the frames in all point cloud frames, it is considered a static point and added to the static background. The fourth step is dynamic point detection. Dynamic points are identified by the displacement of adjacent frame point clouds: frame by frame comparison, search k 1<Movement Speed< k 2's proximity point ensures regular dynamic targets are obtained, k 1 and k2 is the preset speed upper and lower limits. Further filter the non-uniform motion points and calculate the average motion vector of the dynamic point neighborhood v If the motion vector of a single point is v The angle exceeds the preset upper limit θ (e.g. 30°), it is determined as an abnormal point and removed; further filtering is done for clusters with too low voxel density and too low frame ratio (e.g., most rain, dust). The fifth step is clustering to obtain dynamic targets. Using the DBSCAN clustering algorithm, cluster the center area of ​​each dynamic point to obtain the outline of the dynamic target. Assuming that the final calculation results in the dynamic target point cloud f 2. After setting the dynamic and static weights (e.g. dynamic weight 1.5, static weight 1), the dynamic frame and the static frame are spliced ​​together to obtain the final point cloud sequence. f The dynamic target point cloud is filtered out from the point cloud data through inter-frame difference and point cloud registration algorithms. The noise is further filtered through spatial distribution sparsity check and inconsistent motion check. The point cloud is then added to the original target point cloud as a matrix according to the set weights and input into the Transformer model to enhance the model's ability to learn the influence of dynamic targets.

[0062] Secondly, for driving data processing, the speed deviation s , acceleration offset l and heading angle deviation k Combined driving vector cl (6 dimensions), after convolution upscaling (128 dimensions), the time series driving data is transformed into sl driving sequence.

[0063] As for the processing of environmental data, it mainly processes the meteorological data of the mining area, specifically normalizing the weather station data to obtain the time series weather .

[0064] Finally, the point cloud sequence f , driving sequence sl , time series weather and map road topology are spliced ​​together and input into the input embedding layer of the difficult example recognition model for processing to obtain the processing result. The processing result is used to characterize whether the mining truck is currently in the difficult scene corresponding to the difficult example, and to mark the collected data corresponding to the scene. By normalizing the motion parameters and historical mean values ​​and embedding them into the model training, it is easier for the model to learn the impact of motion parameter changes on perception. By embedding dynamic target frames, the model's ability to learn the impact of multiple dynamic targets on the final perception with fewer data samples is enhanced. Furthermore, the difficult examples and corresponding collected data obtained by the processing results can be used as training sets in reverse to train the model again, and gradually improve the recognition accuracy of the model.

[0065] Therefore, the mining truck difficult case identification model training method provided in this application can utilize existing collected data to simulate and generate more difficult cases, replace or add difficult cases in the original sample set, and achieve the expansion of difficult cases, thereby improving the quality of sample data and overcoming the problems of sample omission and small sample size in the prior art. This allows the subsequent training of other autonomous driving models to improve the ability to identify difficult cases and perceive multiple targets, which will contribute to the popularization of mining truck autonomous driving. For the mining truck difficult case identification method, the difficult case identification model obtained by the mining truck difficult case identification model training method is used to input the collected data obtained by the mining truck into the difficult case identification model for processing. Before input processing, the detection data will undergo a series of preprocessing. The comparison of the collected data, driving data, and environmental data in the detection data will be normalized and then embedded in the model processing, making it easier for the model to obtain more accurate processing results. At the same time, the obtained out-of-column records can be used as a reverse training set to train the model again, gradually improving the recognition accuracy of the model. This overcomes the problem of insufficient application of difficult case identification models and difficulty in recognizing difficult case scenarios in the prior art. This enables the model to accurately predict difficult scenarios and improve the safety of intelligent driving.

[0066] Figure 3 FIG1 shows an internal structure diagram of a computer device in an embodiment. The computer device can be a terminal or a server. Figure 3 As shown, the computer device includes a processor, a memory and a network interface connected via a system bus. Among them, the memory includes a non-volatile storage medium and an internal memory. The non-volatile storage medium of the computer device stores an operating system and may also store a computer program. When the computer program is executed by the processor, the processor may implement a mining truck difficult case identification model training method and / or a mining truck difficult case identification method. The internal memory may also store a computer program. When the computer program is executed by the processor, the processor may implement a mining truck difficult case identification model training method and / or a mining truck difficult case identification method. Those skilled in the art will understand that Figure 3 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0067] In one embodiment, the present application further proposes a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the processor executes the steps of the method described in any of the aforementioned embodiments.

[0068] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0069] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0070] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be determined by the appended claims.

Claims

1. A mining truck difficult case recognition model training method, characterized in that: The method comprises the following steps: Acquire collected data, and generate a first sample set based on the collected data, wherein the first sample set includes normal examples and real difficult examples; Performing a difficult example expansion process on the real difficult example to generate a simulated difficult example; Aggregating the real difficult examples and the simulated difficult examples into difficult examples, and aggregating the difficult examples and the normal examples into a second sample set; The second sample set is used to train an initial model, and the trained initial model is marked as a difficult example recognition model.

2. The mining truck difficult example recognition model training method according to claim 1, characterized in that: Generating a first sample set according to the collected data includes: Obtaining labels in the collected data, and distinguishing the collected data into the normal examples and the real difficult examples according to the labels; and / or, The suspicious scene data in the collected data is screened, the collected data containing the suspicious scene data is marked as the real difficult example, and the remaining collected data is marked as normal examples.

3. The mining truck difficult case recognition model training method according to claim 2, characterized in that: The collected data includes multiple frames of point cloud data and image data; The screening of suspicious scene data in the collected data includes: Determining a temporal matching relationship between the point cloud data and the image data for each frame; determining a spatial matching relationship between the point cloud data and the image data using preset calibration parameters; Determining a matching relationship between the image data and the point cloud data according to the time matching relationship and the space matching relationship; Based on the matching relationship, each frame of the image data is processed using an edge detection and segmentation algorithm to determine at least one geometric target; each frame of the point cloud data is clustered and analyzed to determine at least one three-dimensional object; Calculating the degree of overlap between each of the geometric targets and the three-dimensional object in each frame; When the overlap degree meets the suspicious condition, the collected data of this frame is filtered as the suspicious scene data.

4. The mining truck difficult case recognition model training method according to claim 1, characterized in that: The performing a difficult example expansion process on the real difficult example to generate a simulated difficult example includes: Determining a difficult example type of the real difficult example, where the difficult example type includes at least one of a rain and snow difficult example and a dust and fog difficult example; Performing a first sub-expansion process on the rain and snow difficulty case to obtain a plurality of simulated rain and snow difficulty cases; and / or performing a second sub-expansion process on the dust and fog difficulty case to obtain a plurality of simulated dust and fog difficulty cases; All of the simulated rain and snow difficult cases and all of the simulated dust and fog difficult cases are summarized as the simulated difficult cases.

5. The mining truck difficult case recognition model training method according to claim 4, characterized in that: The performing a first sub-expansion process on the rain and snow difficult examples to obtain a plurality of simulated rain and snow difficult examples includes: Determining the rain and snow intensity value of each of the difficult rain and snow examples, screening the difficult rain and snow example with the smallest rain and snow intensity value and marking it as a first sample difficult example; Obtaining point cloud data in the first sample difficult example and marking it as a first sample point cloud; determining a rain and snow interference point cloud in the first sample point cloud and marking the rain and snow interference point cloud as a first unit point cloud; The point cloud data of each of the rain and snow difficult examples is adjusted according to the first unit point cloud, and the adjusted rain and snow difficult examples are marked as the simulated rain and snow difficult examples.

6. The mining truck difficult case recognition model training method according to claim 4, characterized in that: The performing a second sub-expansion process on the dust fog difficult example to obtain a plurality of simulated dust fog difficult examples includes: Acquire collected data corresponding to each of the dust and fog difficult cases and mark them as dust and fog data; Determining modeling parameters for each of the dust and fog difficulty cases according to the dust and fog data; Determining an interference calculation formula according to the modeling parameters; Using the interference calculation formula, fitting the dust and fog data to generate a plurality of simulated collected data; At least one simulated dust and fog instance is generated based on the simulated collected data.

7. The mining truck difficult example recognition model training method according to claim 1, characterized in that: The method further comprises: Constructing the initial model, wherein the initial model adopts a Transformer model architecture; The step of training the initial model using the second sample set and marking the trained initial model as a difficult example recognition model includes: Inputting the normal examples and the difficult examples in the second sample set into the initial model according to a preset ratio for training; Determine whether the initial model of the current round meets the training completion condition according to the training result; the training result is obtained by outputting the initial model after each round of training; If the training completion condition is not met, the model parameters of the initial model of the current round are adjusted according to the training results, and the next round of training is started; or If the training completion condition is met, the initial model of the current round is deemed to have completed training, and the initial model of the current round is marked as the difficult example recognition model and output.

8. A method for identifying difficult examples of mining trucks, characterized in that: The method comprises: Obtaining and deploying a difficult example recognition model, wherein the difficult example recognition model is trained by the method of any one of claims 1 to 7; The collected detection data is input into the difficult example recognition model for processing, and a processing result output by the difficult example recognition model is obtained. The processing result is used to characterize whether the mining truck is currently in a difficult scenario corresponding to the difficult example.

9. A computer device, characterized in that: including processor and memory; The processor is configured to execute the computer program stored in the memory to implement the method according to any one of claims 1 to 7, or the method according to claim 8.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 or the method according to claim 8 is implemented.

Citation Information

Patent Citations

  • Multi-scene automatic point cloud augmentation method for automatic driving

    CN111881029A

  • Target detection network training method and device and image augmentation method and device

    CN113743434A

  • Model training method and device, difficult case identification method and device, equipment, storage medium and program

    CN115359308A

  • Target detection model training method and device, target detection method and device, equipment and medium

    CN116665170A

  • Training sample augmentation method based on three-dimensional model simulation

    CN118781273A