Visual identification neural network structure and method for nonstandard driving of carry-scraper driver

By installing high-resolution cameras and convolutional neural networks on the shovel machine for image classification, identifying and monitoring the irregular driving behavior of shovel machine drivers, the problems of identification accuracy and resource limitation in the prior art are solved, and safety and management efficiency are improved.

CN120220123APending Publication Date: 2025-06-27SHANDONG GOLD GROUP +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510316867.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art is difficult to accurately identify and monitor irregular driving behaviors of shovel drivers in complex environments, and traditional methods are limited in application on resource-constrained devices.

Method used

A visual recognition neural network method for drivers’ irregular driving is adopted. By installing a high-resolution camera on the truck, collecting and preprocessing data, using convolutional neural networks to classify images, identifying irregular driving behaviors, and improving the robustness and generalization capabilities of the model through multi-branch parallel structure and data augmentation technology.

Benefits of technology

It realizes accurate classification and real-time monitoring of complex behaviors, automatically identify and warn of irregular behaviors, improves operational safety, reduces accident risks, and optimizes the operation management of construction machinery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220123A_ABST
    Figure CN120220123A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of engineering machinery monitoring, and particularly relates to a visual identification neural network structure and method for nonstandard driving of a carry-scraper driver. The visual identification neural network method for nonstandard driving of the carry-scraper driver is characterized by comprising the following steps: (1) installing a high-resolution camera on a carry-scraper, and clearly recording conditions in a cockpit and conditions of a road ahead; (2) collecting a data set when a carry-scraper driver drives; (3) preprocessing the collected data set; and (4) sending the preprocessed images to a classification model, and classifying the images into non-standard driving and standard driving categories of the carry-scraper driver. The method has the beneficial effects that accurate classification and real-time monitoring of complex behaviors are realized, irregular behaviors are automatically identified and early warned, a driver is reminded to correct the behaviors, the operation safety is improved, and the accident risk is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of construction machinery monitoring, and particularly relates to a visual recognition neural network structure and method for irregular driving of a scraper driver. Background Art

[0002] A scraper is an efficient construction machinery, mainly used in scenarios such as earth excavation, transportation, and unloading operations. The operation of a scraper is difficult and involves high-risk environments. Irregular driving behaviors of the driver are the main causes of accidents. The irregular driving behaviors of a scraper driver include improper postures, illegal operations, etc. These irregular behaviors may all increase the accident risk, threaten the lives and property safety of surrounding personnel, and affect work efficiency and equipment life.

[0003] Traditional driving behavior monitoring mainly relies on manual supervision or simple sensor detection, but these methods have limitations. For example, the real-time performance is insufficient. Manual supervision cannot achieve all-weather real-time monitoring and is prone to missing key details. Sensor-based detection methods are difficult to comprehensively capture complex driving behaviors, especially subtle movement changes, and the accuracy is limited. Moreover, traditional methods are difficult to adapt to complex operating environments, such as interference factors like light changes and equipment vibrations.

[0004] In recent years, researchers have conducted extensive research on the visual recognition neural network for irregular driving of a scraper driver. Traditional neural networks perform poorly in dealing with different light, angle, and shape changes, which may lead to a decrease in the recognition accuracy of irregular driving of a scraper in complex environments. The performance of a neural network largely depends on the quality and quantity of training data. If the training data is insufficient or of low quality, it may lead to overfitting or poor generalization ability of the model. Deep learning models usually require a large amount of computing resources for training and inference, which may limit their application on resource-constrained devices. Summary of the Invention

[0005] The present invention aims to make up for the deficiencies of the prior art and provides a visual recognition neural network structure and method for irregular driving of a scraper driver.

[0006] The present invention is realized through the following technical solutions: A visual recognition neural network method for irregular driving of a scraper driver, characterized by comprising the following steps: (1), Install a high-resolution camera on the scraper to clearly record the situation inside the cab and the road conditions ahead; (2), Collect a data set of the scraper driver during driving; (3), Preprocess the collected data set; (4) Send the preprocessed image to the classification model to classify the image into categories of improper driving and proper driving of the scraper driver.

[0007] Preferably, in step (3), during preprocessing, data augmentation is first performed on the data.

[0008] Preferably, in step (3), during the data augmentation stage, first process the image data with median blur, and then use the Sobel filter to calculate the gradient of the image to obtain the edge information in the image.

[0009] Preferably, in step (3), based on the edge information obtained from the Sobel filter, use a sharpening mask to further enhance the details and edges of the image, and then apply the CLAHE technique to enhance the contrast of the image.

[0010] A neural network structure used in the vision recognition neural network method for improper driving of a scraper driver, characterized in that: it includes three branches. The first branch is sequentially connected to the max pooling layer 3, the max pooling layer 4, and the convolutional block 1. The second branch is sequentially connected to the max pooling layer 1, the convolutional layer 2, the batch normalization layer 2, the ReLU activation function layer 2, the max pooling layer 5, and the convolutional block 2. The third branch is sequentially connected to three convolutional blocks. The three branches are connected to the fusion layer through the dropout layer. The fusion layer is sequentially connected to the convolutional block 7, the convolutional block 5, and the convolutional block 6. The convolutional block 6 is sequentially connected to the output layer through the fully connected layer and the SoftMax layer.

[0011] Preferably, each convolutional block in the third branch includes a convolutional layer, a batch normalization layer, a ReLU activation function layer, and a max pooling layer.

[0012] Preferably, the convolutional block 7, the convolutional block 5, and the convolutional block 6 all include a convolutional layer, a batch normalization layer, a ReLU activation function layer, and a max pooling layer.

[0013] The beneficial effects of the present invention are: using convolutional neural network and computer vision technologies to extract high-dimensional features from the video stream, realizing accurate classification and real-time monitoring of complex behaviors, automatically identifying and warning of improper behaviors, reminding the driver to correct the behaviors, improving the operation safety, reducing the accident risk, and optimizing the operation management of construction machinery. It provides technical support for the further improvement of the intelligent transportation system and helps to achieve more intelligent and automated traffic management. It can not only be applied to construction machinery such as scrapers, but also be extended to other types of vehicles and equipment, and has a wide application prospect. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] The present invention will be further described below with reference to the accompanying drawings.

[0015] Appendix Figure 1This is the workflow diagram of the present invention; Appendix Figure 2 This is the flowchart for preprocessing the dataset of the present invention; Appendix Figure 3 This is the structural diagram of the visual recognition neural network of the present invention. Specific implementation manner

[0016] The accompanying drawings show a specific embodiment of the present invention. This embodiment includes the following steps: Step 1: Install a high-resolution camera on the scraper to ensure that the situation inside the cab and the road conditions ahead can be clearly recorded. The process of the camera collecting data on the irregular driving of the scraper driver involves multiple steps, mainly including data planning, equipment installation, data collection, and data processing, etc. Step 2: Preprocess the dataset collected during the driving of the scraper driver. When performing data augmentation, certain preprocessing techniques are used, such as operations like resizing, scaling, median filtering, Sobel filtering, sharpening mask, and grayscaling. Resizing ensures that the data can be correctly input into the network, and normalizing the image data to a specific range can make the gradient descent algorithm converge faster. Data augmentation increases the diversity of the training data, enabling the model not to overly rely on limited original samples during the training process, thereby reducing the risk of overfitting. In a deep neural network, inappropriate data scales may lead to problems such as vanishing or exploding gradients during the backpropagation process, and normalizing the data helps to maintain the gradients within a reasonable range. Step 3: Send the preprocessed images to the proposed classification model to classify the images into categories of irregular and regular driving of the scraper driver. To identify irregular driving behaviors of the scraper driver at different scales, the proposed classification model architecture adopts three parallel streams, each designed to process the input at different resolutions. This multi-scale representation enhances the model's ability to capture and analyze the irregular driving of the scraper driver at different granularity levels.

[0017] In the stage of collecting and processing the original dataset, first clarify which driving behaviors are considered irregular, such as poking the head outwards, leaving the seat casually, etc., and set clear judgment criteria and thresholds for each irregular behavior. Annotate the collected data, remove invalid or incorrect data records, ensure the data quality, and mark the types of irregular behaviors and the time points when they occur.

[0018] Since this dataset consists of images of various sizes, during the augmentation stage, they have been scaled to a fixed size of 128×128. When augmenting the driving images of the scraper driver, the preprocessing methods used include rotation, scaling, translation in height and width, and shearing, etc. To retain the most features in the images, the settings for rotation, scaling, translation, and shearing, etc. are adjusted for different categories of samples in the training and testing folders to ensure the consistency of the input.

[0019] In the data augmentation stage, first, median blur is applied to the original image for preliminary processing. Its main function is to effectively remove random noises such as salt-and-pepper noise in the image. When a forklift driver is operating, the noises from components such as the engine and transmission are superimposed on the working environment noise, making the noise environment more complex. While smoothing the image, this method can better preserve the edge and detail information of the image, providing a relatively clean and stable image basis for subsequent processing steps and reducing the interference of noise on subsequent processing. The basic idea of median blur is to replace the value of a pixel with the median value within its pixel neighborhood. A 3×3 neighborhood window is set, centered on the current pixel, with a region of 1 pixel above, below, to the left, and to the right, totaling 9 pixels.

[0020] After median blur processing, the Sobel filter is used to calculate the gradient of the image, thereby highlighting the edge information in the image. This helps the CNN better capture the structural features of the image. The core principle of the Sobel filter is based on the gradient calculation of image intensity. It determines the position of the edge by calculating the degree of gray change of each pixel point in the image. For a two-dimensional image, its gradient can be decomposed into horizontal and vertical components. The Sobel filter uses two 3×3 convolutional kernels to approximately calculate the gradients in these two directions respectively.

[0021] Specifically, first, for the RGB image of the forklift driver during driving collected, it can be converted into a grayscale image using the grayscale conversion formula as follows:

[0022] where R, G, and B represent the intensity values of the red, green, and blue components respectively; the coefficients 0.299, 0.587, and 0.114 are set according to the perception sensitivity of the human eye to these three colors. By performing grayscale conversion through the weighted average method, important information in the forklift driver's driving picture can be retained, while reducing computational complexity, which helps to simplify the data and highlight the key features of the image.

[0023] Convolution operations are performed on the grayscale image using the Sobel convolution kernels in the horizontal and vertical directions respectively. For each pixel point (x, y) in the image, its horizontal gradient Gx(x, y) and vertical gradient Gy(x, y) are calculated through the following convolution formulas:

[0024]

[0025] Where, I(i,j) represents the pixel value at position (i,j). Gx and Gy are the Sobel operator matrices in the horizontal and vertical directions respectively. m and n are the indices for traversing the Sobel operator matrix, and (i+m,j+n) is the local region in the image corresponding to the Sobel operator matrix. By multiplying the corresponding elements of the local region of the image with the Sobel operator matrix and summing them up, the gradient value at this point can be obtained.

[0026] Based on the edge information obtained from the Sobel filter, a sharpening mask is used to further enhance the details and edges of the image. The Laplacian of Gaussian (LoG) operator is an edge detection operator that combines the effects of Gaussian smoothing and Laplacian sharpening. Its convolution operation mainly includes the following key steps: First, the input image is convolved with a two-dimensional Gaussian function to reduce noise and smooth the image. The Gaussian function formula is:

[0027] where σ is the standard deviation, which controls the degree of smoothing.

[0028] Then, the sum of the second-order derivatives of the Gaussian-smoothed image is calculated, which is the Laplace transform. The discrete form of the Laplacian operator is usually expressed as:

[0029] The Gaussian function is combined with the Laplacian operator to form the LoG operator:

[0030] It is applied to the original image through convolution operations. After completing the above processing, the CLAHE technique is applied to enhance the contrast of the image. CLAHE can adaptively adjust the contrast according to the local regions of the image, avoiding problems such as over-enhancement or noise amplification that may occur in global histogram equalization, enabling the CNN to extract richer features. The input image is divided into several non-overlapping 8×8 pixel small blocks, and the size of these small blocks is fixed. The purpose of overlapping is to achieve a better transition at the boundaries of the small blocks and avoid obvious block effects. For each small block, the distribution of its gray values is statistically analyzed to obtain the histogram of this small block.

[0031] For each sub-block, the number of pixels at each gray level is statistically analyzed, that is, the frequency of occurrence of each gray value within the sub-block is calculated to obtain the histogram of this sub-block. This step can be achieved by traversing all the pixel points within the sub-block and counting according to the pixel values. The cumulative distribution function (CDF) is calculated based on the histogram of the sub-block. The CDF represents the sum of the probabilities of pixels with gray values less than or equal to a specific value. The CDF is obtained by performing an accumulation operation on the histogram, that is

[0032] Among them, hist(i) is the number of pixels with a gray value of i.

[0033] Set a contrast limit threshold of 4.0. If the number of pixels at certain gray levels in a sub-block exceeds this threshold, then these excess pixel numbers are evenly distributed to other gray levels to achieve histogram clipping and redistribution, so as to avoid excessive enhancement of local contrast.

[0034] Perform equalization processing on the histogram of the sub-block after clipping and redistribution, so that the pixel value distribution within the sub-block is more uniform, thereby enhancing the local contrast of the image. The mapped gray value is shown as follows:

[0035] Among them, L is the total number of gray levels, and T(i) is the equalized gray value 1.

[0036] As Figure 2 shown, it shows the deep architecture for image recognition, which can extract the non-standard driving features of the scraper driver at different scales. The input of the network is an image with dimensions h×w, where "h" and "w" represent height and width respectively. The output of the network includes class probability prediction, indicating whether the analyzed image contains non-standard driving of the scraper driver.

[0037] The input image is analyzed at three different scales: h×w, h / 2×w / 2, and h / 4×w / 4. This multi-resolution analysis is crucial for capturing the global context and details of various non-standard driving of the scraper driver. The parallel branches of the architecture are customized to gradually reduce the size of the input picture, which is consistent with the reduction in the number of convolutional layers connected to each branch. This design ensures that the fusion layer receives feature maps of equal size.

[0038] The proposed network architecture consists of three branches. Among them, the first branch is sequentially connected to the max pooling layer 3, the max pooling layer 4, and the convolutional block 1. Inside the convolutional block 1, a convolutional layer, a batch normalization layer, a ReLU activation function layer, and a max pooling layer are sequentially connected.

[0039] In the max pooling operation 3, the pooling window is set to 64×64, and its core function is to perform downsampling on the input feature map. Specifically, after selecting a local area on the feature map, the maximum value in this area is taken as the output result to achieve the purpose of reducing the size of the feature map.

[0040] Subsequently, the max pooling operation is performed again. This time, the pooling window is adjusted to 32×32. This step can further reduce the resolution of the feature map, while continuously extracting key features, further reducing the data volume and computational volume.

[0041] After that, the convolutional layer in Convolution Block 1 performs a convolution operation on the input image using a 3×3 filter, outputting 8 feature maps to extract the initial features of the image. Then, batch normalization is performed on the output of the convolutional layer, which helps to stabilize the network training process, accelerate the convergence speed, and improve the generalization performance of the model. Finally, the ReLU activation function is applied to the result after batch normalization to introduce non-linearity and enable the network to learn more complex feature representations. The mathematical expression of the ReLU activation function is:

[0042] If the input x is greater than 0, the output is x; if the input x is less than or equal to 0, the output is 0. For the max pooling layer in Convolution Block 1, a filter with a stride of 8 and a size of 16×16 is used to perform max pooling on the feature map after ReLU activation. This operation can effectively reduce the resolution of the feature map, reduce the computational amount, and ensure that the main feature information is retained.

[0043] The processing flow of the second branch is successively a max pooling layer 1, a convolutional layer 2, a batch normalization layer 2, a ReLU activation function layer 2, a max pooling layer 5, and a Convolution Block 2. At the beginning, a max pooling layer with a pooling size of 64×64 is added, which is used to reduce the spatial dimension of the feature map and helps to achieve translational invariance while retaining key information. Inside Convolution Block 2, there are successively connected a convolutional layer, a batch normalization layer, a ReLU activation function layer, and a max pooling layer.

[0044] The filter size used in Convolutional Layer 2 is 3×3, and its number of output channels is 8; the pooling window size of Max Pooling Layer 5 is 32×32, and the number of output channels is also 8, which helps to reduce the computational amount and the number of model parameters.

[0045] Finally, through Convolution Block 2, the convolutional layer in this convolution block performs a convolution operation on the feature map after the previous max pooling layer using a 3×3 convolution kernel to further extract more complex features. The output feature map has 64 channels, enabling the model to capture more subtle feature patterns. Moreover, its max pooling layer has a pooling window size of 16×16 and an output channel number of 64, which can further reduce the resolution of the feature map while ensuring that important feature information is retained, making the extracted features more robust and representative.

[0046] The processing flow of the third branch goes through three convolutional blocks in sequence, and each convolutional block includes a convolutional layer, a batch normalization layer, a ReLU activation function, and a max pooling layer. In the first convolutional block, the filter size of the convolutional layer is 3×3, its output channels are 8, the corresponding pooling window size is set to 64×64, and the output channels are also 8. Then, in the second convolutional block, the convolutional layer uses a 3×3 convolutional kernel to perform convolution on the feature map after the first max pooling to further extract more complex features, and the output feature map has 32 channels. The second pooling window size of this convolutional block is 32×32, and the output channels are 32. Finally, in the last convolutional block, the kernel size is 3×3, the output channels are 64, and the pooling window size of the third max pooling is 16×16, and the output channels are 64.

[0047] After the last max pooling of the three branches, a dropout layer is added to this neural network. It will set the outputs of some neurons to 0 with a probability of 0.5 during the training process, so that the network uses different subsets of neurons for learning in each training iteration. This reduces the mutual dependence between neurons, prevents the network from relying too much on certain specific neurons or feature combinations, thus effectively reducing the risk of overfitting and improving the generalization ability of the model. Dropout has different calculation methods in the training and testing phases. In the training phase, a binary mask is generated, and a random value is generated for each neuron. If the random value is less than the dropout probability p, the output of this neuron is set to 0 (i.e., discarded); if the random value is greater than or equal to p, the output of this neuron remains unchanged. Expressed by the formula:

[0048] where M represents the mask value of the i-th neuron, taking values 0 or 1, and

[0049]

[0050] Forward propagation calculation: Multiply the input vector x by the mask m to obtain the output vector after Dropout processing . Since some neurons are discarded and the actual number of neurons participating in the calculation becomes smaller, the result needs to be scaled to ensure the consistency of the output. The specific calculation formula is:

[0051] In the testing phase, neurons are no longer randomly discarded, but all neurons are used, and the output is scaled to ensure that the expected value of the prediction is the same as that during training. The specific calculation formula is:

[0052] Among them, y test is the output in the test phase, y is the original output of the network without Dropout processing, and p is the dropout probability during training.

[0053] By randomly dropping neurons, the model will learn multiple different feature representations and patterns during training, which makes the model more robust to small changes or noise in the input data. Due to the existence of the dropout layer, the network will continuously adjust its learning strategy during training, prompting neurons to learn more representative and general features, avoiding falling into local optima, helping to accelerate the convergence speed of the model, improve training efficiency, and reduce training time.

[0054] The multiple feature maps output by the foregoing network layer are concatenated in the channel dimension, and the obtained feature map after concatenation has a spatial specification of 16×16 and 136 channels. Thereafter, this fusion layer goes through three convolutional blocks, and each convolutional block includes a convolutional layer, a batch normalization layer, a ReLU activation function, and a max pooling layer.

[0055] In convolutional block 7, its convolutional layer uses a 3×3 convolutional kernel to perform convolution operations on the concatenated feature map, extracts different feature patterns by sliding the convolutional kernel on the feature map, and outputs a feature map with 128 channels. The max pooling operation in this convolutional block 7 has a pooling window size of 8×8 and an output channel number of 128, which reduces the size of the feature map by selecting the maximum value in the local area of the feature map.

[0056] Convolutional block 5 then uses a 3×3 convolutional kernel again to perform convolution operations on the feature map processed by the previous max pooling layer, further extracts more complex features, and then outputs a feature map with 256 channels. The pooling window size of max pooling in this convolutional block 5 is 4×4, and the output channel number is 256, which can further reduce the resolution of the feature map.

[0057] In the last convolutional block 6, the filter size of the convolutional layer is 3×3, and it outputs a feature map with 256 channels, which further enhances the richness and expressiveness of the features. The pooling window size in convolutional block 6 is 2×2, and the output channel number is 256, thus providing a more refined feature input for the subsequent fully connected layer.

[0058] Subsequently, the features obtained after all the previous convolution and pooling operations are flattened into a one-dimensional vector, and then a fully connected layer is used to handle the classification task. The fully connected layer synthesizes the previously extracted features to serve the final decision-making. A SoftMax layer is connected after the fully connected layer. The SoftMax function can transform the output of the fully connected layer into a probability distribution for each class, ensuring that the output values are in the range of 0 to 1 and the sum of the probabilities for all classes is 1. The overall network architecture gradually extracts the features of the image through multiple layers of convolution, batch normalization, ReLU activation, and pooling operations, and determines the final classification result relying on the fully connected layer and the SoftMax layer, outputting the class to which the image belongs.

Claims

1. A neural network method for visually identifying irregular driving of a scraper driver, characterized by: The following steps are involved: (1) Install a high-resolution camera on the LHD to clearly record the situation inside the cockpit and the road ahead; (2) Collecting data sets of the LHD drivers while driving; (3) Preprocess the collected data set; (4) Send the preprocessed images to the classification model to classify the images into the categories of irregular driving and regular driving of the scraper driver.

2. The neural network method for visually identifying irregular driving of a scraper driver according to claim 1 is characterized by: In step (3), the data is first enhanced during preprocessing.

3. The neural network method for visually identifying irregular driving of a scraper driver according to claim 2 is characterized in that: In step (3), in the data enhancement processing stage, the image data is first processed using median blur, and then the gradient of the image is calculated using the Sobel filter to obtain the edge information in the image.

4. The neural network method for visually identifying irregular driving of a scraper driver according to claim 3 is characterized in that: In step (3), based on the edge information obtained by the Sobel filter, an unsharp mask is used to further enhance the details and edges of the image, and then the CLAHE technique is applied to enhance the contrast of the image.

5. A neural network structure used in the neural network method for visually identifying irregular driving of a scraper driver as claimed in any one of claims 1 to 4, characterized in that: It includes three branches. The first branch is connected to the maximum pooling layer 3, the maximum pooling layer 4 and the convolution block 1 in sequence. The second branch is connected to the maximum pooling layer 1, the convolution layer 2, the batch normalization layer 2, the ReLU activation function layer 2, the maximum pooling layer 5 and the convolution block 2 in sequence. The third branch is connected to the three convolution blocks in sequence. The three branches are connected to the fusion layer through the dropout layer. The fusion layer is connected to the convolution block 7, the convolution block 5 and the convolution block 6 in sequence. The convolution block 6 is connected to the output layer through the fully connected layer and the SoftMax layer in sequence.

6. The neural network structure according to claim 5, characterized in that: Each convolution block in the third branch includes a convolution layer, a batch normalization layer, a ReLU activation function layer, and a maximum pooling layer.

7. The neural network structure according to claim 5, characterized in that: convolution Block 7, convolution block 5 and convolution block 6 all include a convolution layer, a batch normalization layer, a ReLU activation function layer and a maximum pooling layer.