Stock yard mining efficiency analysis method based on multi-mode excavator intelligent monitoring

Through the intelligent monitoring method of multimodal excavator, combined with the preprocessing and feature fusion of video, audio and kinematic data, the problem of excavator status identification in material field mining is solved, and efficient and accurate mining efficiency analysis and safety monitoring is achieved, which is suitable for water conservancy and hydropower engineering construction.

CN120258219APending Publication Date: 2025-07-04HUANENG LANCANG RIVER HYDROPOWER CO LTD +2
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510342369.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The prior art is difficult to fully and accurately reflect the real-time operating status of the excavator in a complex material field mining environment, and cannot meet the needs of efficient mining and safety monitoring. In addition, multimodal data fusion technology has the challenges of information extraction and correlation capture in the operation efficiency analysis of excavator.

Method used

The intelligent monitoring method of multimodal excavator is adopted to obtain video, audio and kinematic data, preprocess and extract single-modal features, and use cross-modal attention, self-attention and multi-head self-attention mechanisms to perform multi-level feature fusion, train machine learning models to identify excavator activity status, calculate productivity and make future efficiency predictions.

Benefits of technology

It realizes accurate identification and classification of excavator activity status, improves the accuracy and safety of material field mining efficiency, provides real-time prediction support for production efficiency, and optimizes resource allocation and construction management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120258219A_ABST
    Figure CN120258219A_ABST
Patent Text Reader

Abstract

The invention discloses a stock ground mining efficiency analysis method based on multi-mode excavator intelligent monitoring. The method comprises the steps that multiple single-mode original features corresponding to preprocessed excavator monitoring data are extracted; calculating correlation among modals of the plurality of single-modal original features to generate a preliminary fusion feature, and fusing a self-attention mechanism and a multi-head self-attention mechanism to perform multi-level feature fusion to obtain a multi-modal fusion feature; training a machine learning model by using the multi-modal fusion features to obtain an optimal machine learning recognition model so as to output an excavator activity state recognition result; and calculating the action time, the average cycle time and the productivity of the excavator in different activity states, and predicting the future production efficiency according to the calculated productivity. According to the invention, accurate identification and classification of the activity state of the excavator can be realized, the identification accuracy of the activity state of the excavator is improved, and powerful support is provided for real-time prediction and optimization of the mining efficiency of a stock yard.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of water conservancy and hydropower engineering construction, and particularly relates to a method for analyzing the quarrying efficiency based on multi-modal intelligent monitoring of excavators. Background Art

[0002] Quarrying is a key link in the construction process of water conservancy projects, which is directly related to the construction progress, production cost and resource utilization efficiency of the project. As the main equipment for quarrying, the working efficiency and operating status of excavators have a significant impact on the overall efficiency of quarrying operations. Traditional methods for analyzing quarrying efficiency usually rely on manual inspections, single-sensor data or basic mechanical monitoring. Although they can provide some preliminary data, it is difficult to comprehensively and accurately reflect the real-time operating status of excavators in complex working environments. These methods not only have problems such as lagging information collection and single data processing, but also cannot meet the actual needs of efficient quarrying and safety monitoring.

[0003] Traditional excavation simulation detection methods, such as Xia Dayong et al. established an excavation simulation model using system simulation methods for the spillway excavation project of Tian Shengqiao First-class Power Station; Song Wenshuai et al. proposed a construction simulation parameter prediction method based on a width learning system and elastic net quantile regression, providing more accurate inputs for the simulation analysis of dam construction progress; Zhang Yan proposed an excavation construction simulation system considering the filling requirements of the upper reservoir dam body, optimizing the planning of excavation material transportation. Although these studies have deeply explored excavation simulation, most of them do not fully utilize multi-modal data and rarely involve the research on the excavation efficiency of excavators.

[0004] In recent years, with the rapid development of Internet of Things, cloud computing and big data technologies, the accurate identification of the activity status of construction machinery based on multi-modal data has become a research hotspot. Zhang Jun et al. used methods such as low-pass filters and Mel spectrograms to realize the real-time collection and preprocessing of multi-modal data of rockfill dam construction machinery, and proposed a deep learning model for rockfill dams for automatically extracting multi-modal data features. However, few studies have used multi-modal data to analyze the excavation efficiency of excavators. In the analysis of excavator operation efficiency, an intelligent monitoring system that integrates video monitoring, audio detection and kinematic data can monitor the equipment operation status from multiple angles and achieve accurate identification of quarrying efficiency. Among them, video data can provide visual information on the external operation of the excavator, helping to monitor the mechanical posture, working position and operation process; audio data reveals the working load and engine status of the excavator; kinematic data reflects the working dynamics and operation efficiency of the machinery by real-time monitoring of parameters such as position, speed and acceleration. Combining with artificial intelligence technology, the intelligent monitoring system can perform accurate analysis in complex operation environments and further improve the operation efficiency of excavators.

[0005] Although multi-modal data fusion technology has shown great potential, it still faces many challenges in practical applications. First of all, how to extract valuable information from massive and complex data and transform it into effective decision support remains a bottleneck in the development of the technology. In the complex quarrying environment, the operation efficiency is affected by various factors, such as equipment performance, operation methods, operator skills, and geological conditions. How to comprehensively evaluate these factors is still a key problem to be solved urgently. Secondly, due to the dynamic and complex nature of excavator operations, it is difficult to capture the correlation between different modal data, which poses challenges to data fusion and efficiency analysis. Finally, how to optimize the recognition model after multi-modal feature fusion to ensure the real-time and accuracy of quarrying efficiency analysis is an important direction for the development of technology in this field. Summary of the Invention

[0006] The present invention aims to solve at least one of the technical problems in the related art to some extent.

[0007] The present invention provides a method for analyzing quarrying efficiency based on multi-modal intelligent monitoring of excavators, which not only improves the accuracy of quarrying efficiency but also provides strong support for the evaluation of the production efficiency and safety of construction machinery.

[0008] Another object of the present invention is to provide a system for analyzing quarrying efficiency based on multi-modal intelligent monitoring of excavators.

[0009] To achieve the above object, on the one hand, the present invention provides a method for analyzing quarrying efficiency based on multi-modal intelligent monitoring of excavators, including:

[0010] Obtaining preprocessed excavator monitoring data; wherein, the excavator monitoring data includes video data, audio data, and kinematic data;

[0011] Extracting multiple single-modal original features corresponding to the preprocessed excavator monitoring data;

[0012] Using a cross-modal attention strategy to calculate the correlation between different modalities of multiple single-modal original features to generate preliminary fusion features, and integrating the preliminary fusion features into a self-attention mechanism and a multi-head self-attention mechanism to perform multi-level feature fusion to obtain multi-modal fusion features;

[0013] Training a machine learning model with the multi-modal fusion features to obtain an optimal machine learning recognition model to output the recognition result of the excavator activity state;

[0014] Calculating the action time, average cycle time, and productivity of the excavator in different activity states based on the recognition result of the excavator activity state, and predicting the future production efficiency according to the calculated productivity.

[0015] The method for analyzing the quarrying efficiency based on multimodal excavator intelligent monitoring according to the embodiments of the present invention may further have the following additional technical features:

[0016] In an embodiment of the present invention, preprocessing the excavator monitoring data includes:

[0017] Using the Gaussian filtering method to remove the noise of the video data, and performing size normalization processing on the video image to obtain a preprocessed video frame sequence;

[0018] Reading the original speech signal from the audio file, performing noise reduction using the spectral subtraction method, and converting the sound wave signal into a Mel spectrogram through short-time Fourier transform;

[0019] For the kinematic data, including acceleration, angular velocity, and position, a low-pass filter is used to remove the high-frequency noise of the kinematic data, retain the low-frequency signal, and use mean filtering to perform mean processing on the low-frequency signal with a sliding window to obtain preprocessed kinematic data;

[0020] Using a time synchronization algorithm to align the multimodal data to a unified time axis.

[0021] In an embodiment of the present invention, extracting multiple single-modal original features corresponding to the preprocessed excavator monitoring data includes:

[0022] Inputting the preprocessed video frame sequence into the S3D model based on separable spatial and temporal convolutions to output a high-dimensional feature vector;

[0023] Inputting the Mel spectrogram into the VGGish model to extract high-level depth features from the audio data, and processing the Mel spectrogram through a deep convolutional network to obtain the time-domain and frequency-domain features of the audio;

[0024] Inputting the preprocessed kinematic data into the Conformer model to capture temporal features; and,

[0025] Performing unified standardization and dimensionality reduction processing on the multiple single-modal original features, including:

[0026]

[0027] In an embodiment of the present invention, using a cross-modal attention strategy to calculate the correlation between each modality of the multiple single-modal original features to generate a preliminary fusion feature, and integrating the preliminary fusion feature into the self-attention mechanism and the multi-head self-attention mechanism to perform multi-level feature fusion to obtain a multi-modal fusion feature, including:

[0028] Self-attention mechanism:

[0029]

[0030] Among them, Q represents the query matrix, K represents the key matrix, V represents the value matrix, and d k is the dimension of the key vector;

[0031] Cross-modal attention mechanism:

[0032] CrossModalAttention(X1,X2) = Concat(Attention 12 ,Attention 21 ) (3)

[0033] Among them, Attention 12 , Attention 21 are the attention of X1 to X2 and the attention of X2 to X1, respectively;

[0034] Multi-head self-attention mechanism. The query, key, and value vectors are linearly transformed through multiple weight matrices W i Q , W i K , W i V for training, respectively, to generate multiple attention heads; each attention head independently calculates the attention weights and generates the corresponding attention output, and then the outputs of all attention heads are concatenated and linearly transformed through a final weight matrix W O for training to form the final multi-head self-attention output:

[0035] MultiHead(Q,K,V) = Concat(head1,head2,...,head n )W O (4)

[0036] Among them, each head i = Attention(QW i Q ,KW i K ,VW i V ),W i Q ,W i K ,W i V and W O are weight matrices for training.

[0037] In an embodiment of the present invention, training a machine learning model using multi-modal fusion features to obtain an optimal machine learning recognition model includes:

[0038] Divide the multi-modal fusion features into a training set and a validation set;

[0039] Input the training set into the machine learning model SVM, and adjust the model parameters to minimize the loss function through the optimization algorithms of backpropagation and gradient descent;

[0040] After the machine learning model is trained, use the validation set to evaluate the performance of the model, and adjust the model structure, hyperparameters, and training strategy according to the evaluation results to find the optimal machine learning recognition model; among them, the performance evaluation metrics include accuracy, precision, recall, F1 score, and confusion matrix;

[0041] Output the activity state of the excavator based on the optimal machine learning recognition model; among them, the activity state of the excavator includes multiple states such as excavation, transportation, loading, and shutdown.

[0042] In an embodiment of the present invention, calculate the action time, average cycle time, and productivity of the excavator in different activity states based on the excavator activity state recognition result, and predict the future production efficiency according to the calculated productivity, including:

[0043] The time of each action is calculated by the formula:

[0044]

[0045] where, AT is the action time, EF and SF are the end frame and start frame respectively, and FR is the frame rate;

[0046] Add up the time of each cyclic action to obtain the average time of the whole cycle:

[0047] ACT = ADG + AH + ADP + AS (6)

[0048] where, ADG is the excavation time, AH is the transportation time, ADP is the unloading time, and AS is the swing time.

[0049] The productivity calculation formula is as follows:

[0050]

[0051] where, P is the productivity, ACT is the average cycle time, and ABP is the average bucket load; calculated by ABP = HBC × BEF, where HBC is the full bucket capacity and BFF is the bucket filling factor;

[0052] Based on the productivity, predict the future productivity of the excavator through the deep learning model LSTM, and make production adjustments according to the prediction results.

[0053] To achieve the above object, on the other hand, the present invention provides a system for analyzing the production efficiency of a quarry based on multi-modal intelligent monitoring of excavators, including:

[0054] A monitoring data acquisition module for acquiring pre-processed excavator monitoring data; wherein, the excavator monitoring data includes video data, audio data, and kinematic data.

[0055] A feature extraction module for extracting a plurality of single-modal original features corresponding to the pre-processed excavator monitoring data.

[0056] A feature fusion module for calculating the correlation between modalities of the plurality of single-modal original features using a cross-modal attention strategy to generate preliminary fusion features, and integrating the preliminary fusion features into a self-attention mechanism and a multi-head self-attention mechanism for multi-level feature fusion to obtain multi-modal fusion features.

[0057] A model training module for training a machine learning model using the multi-modal fusion features to obtain an optimal machine learning recognition model for outputting the recognition result of the excavator activity state.

[0058] A production efficiency prediction module for calculating the action time, average cycle time, and productivity of the excavator in different activity states based on the excavator activity state recognition result, and predicting the future production efficiency according to the calculated productivity.

[0059] For the method and system for analyzing the production efficiency of a quarry based on multi-modal intelligent monitoring of excavators according to the embodiments of the present invention, through accurate recognition of the activity state, project managers can monitor the working state of equipment in real time, timely discover potential faults or safety hazards, thereby optimizing resource allocation, improving construction efficiency, and reducing project risks. In addition, the intelligent monitoring system of this method can automatically predict the production efficiency of quarrying, helping project decision-makers adjust the operation plan and strategy in real time, and further promoting the process of project intelligentization and digitization.

[0060] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] The above and / or additional aspects and advantages of the present invention will become apparent and be readily understood from the following description of the embodiments in conjunction with the accompanying drawings, wherein:

[0062] Figure 1 is an example diagram of the production efficiency of a quarry based on intelligent monitoring of excavators according to an embodiment of the present invention;

[0063] Figure 2 is a flowchart of a method for analyzing the production efficiency of a quarry based on multi-modal intelligent monitoring of excavators according to an embodiment of the present invention;

[0064] Figure 3 is a multi-modal feature extraction and fusion flowchart according to an embodiment of the present invention;

[0065] Figure 4 is a structural diagram of a yard mining efficiency analysis system based on multi-modal excavator intelligent monitoring according to an embodiment of the present invention. Specific embodiments

[0066] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other. The present invention will be described in detail below with reference to the drawings and in combination with the embodiments.

[0067] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present invention.

[0068] The method and system for analyzing the yard mining efficiency based on multi-modal excavator intelligent monitoring according to an embodiment of the present invention will be described below with reference to the drawings.

[0069] The purpose of the present invention is to address the defect in the existing yard mining efficiency analysis method that lacks comprehensive consideration of the complex associations among the excavator's activity status, operating environment, and multi-modal data features. A method for analyzing the yard mining efficiency based on multi-modal excavator intelligent monitoring is proposed. Through the preprocessing and feature extraction of multi-modal data (video, audio, kinematic data), combined with cross-modal attention mechanism, self-attention mechanism, and multi-head self-attention mechanism, a multi-level feature fusion model is established to achieve accurate identification and classification of the excavator's activity status. Through the application of the feature fusion method, the dynamic operation characteristics of the excavator can be captured more comprehensively and accurately, providing a reliable basis for yard mining efficiency analysis and safety status judgment; through the optimization of the training set and validation set, the efficiency and robustness of the recognition model can be ensured, providing strong support for the real-time prediction and evaluation of mining operations. An example diagram of yard mining based on excavator intelligent monitoring is as Figure 1 shown.

[0070] Figure 2 is a flowchart of the method for analyzing the yard mining efficiency based on multi-modal excavator intelligent monitoring according to an embodiment of the present invention, as Figure 2 shown, and the method includes:

[0071] S1, obtaining the preprocessed excavator monitoring data; wherein, the excavator monitoring data includes video data, audio data, and kinematic data.

[0072] By deploying a multi-modal intelligent monitoring system for quarry excavation machinery, including cameras, microphones, and inertial measurement units (IMUs), visual, audio, and kinematic data are obtained.

[0073] Among them, visual data: Video data during the operation process is obtained in real time through cameras installed on the excavator and in the quarry.

[0074] Among them, audio data: The working audio of the excavator is collected by installing microphones to capture the engine running sound and noise characteristics during the operation process.

[0075] Among them, kinematic data: Dynamic data of the excavator is collected using sensors such as accelerometers and gyroscopes to monitor the speed, acceleration, and operation angle information of the excavator in real time.

[0076] Furthermore, the visual, audio, and kinematic data obtained from the above multi-modal intelligent monitoring system are preprocessed to ensure that the information from different data sources can be effectively fused and provide accurate input for subsequent analysis.

[0077] Specifically, the main objective of the present invention is to deploy a multi-modal intelligent monitoring system for quarry excavation machinery, including cameras, microphones, and inertial measurement units (IMUs), so as to obtain visual, audio, and kinematic data. And the obtained visual, audio, and kinematic data are preprocessed to ensure that the information from different data sources provides accurate input for subsequent analysis.

[0078] First, for the acquisition of video data, perspective information involved in the operation process is obtained in real time through high-resolution cameras installed at different positions on the excavator and in the quarry. For the acquisition of audio data, the running sound of the excavator engine, mechanical noise generated during the operation process, and sound signals from the external environment are collected through highly sensitive microphones installed in the working area of the excavator. For kinematic data, sensors such as accelerometers and gyroscopes are installed at key parts of the excavator to collect dynamic data of the excavator in real time. These sensors can monitor motion parameters such as the speed, acceleration, tilt angle, and direction change of the excavator during the operation process.

[0079] Subsequently, the obtained visual, audio, and kinematic data are systematically preprocessed to extract effective information that is of great value for efficiency analysis and status recognition.

[0080] Among them, for video data, the Gaussian filtering method is first used to remove noise in the video and improve the video quality, and then the video is decomposed into individual frames. In addition, since the sizes of video frame images often vary. Therefore, it is necessary to perform size normalization processing on the images to ensure that the sizes of all frame images are the same, so as to be input into subsequent deep learning models for analysis.

[0081] Among them, for the audio data, the original voice signal, i.e., the sound wave, is first read from the audio file, and spectral subtraction is used to reduce the noise of the recorded sound signal, removing background noise and interference and improving the clarity of the audio signal. At the same time, since it is difficult to describe the frequency variation law of the sound wave, the sound wave is converted into a Mel spectrogram through short-time Fourier transform (STFT).

[0082] Among them, the kinematic data mainly includes acceleration, angular velocity, and position. For the kinematic data, a low-pass filter is first used to remove high-frequency noise and retain low-frequency signals. Then, mean filtering is used as a denoising technique. By performing mean processing on the data with a sliding window, the random fluctuations in the data are smoothed, thereby eliminating abnormal interference within a short period of time and ensuring that the signal is more stable.

[0083] Considering that the kinematic, audio, and video data come from different sensors and there may be differences in timestamps, a time synchronization algorithm is adopted to align the data of each modality to a unified time axis, ensuring the time consistency of multi-source data and laying a foundation for feature fusion and comprehensive analysis.

[0084] S2. Extract multiple single-modal original features corresponding to the preprocessed excavator monitoring data.

[0085] S3. Use the cross-modal attention strategy to calculate the correlation between each modality of multiple single-modal original features to generate preliminary fusion features, and incorporate the preliminary fusion features into the self-attention mechanism and the multi-head self-attention mechanism to perform multi-level feature fusion to obtain multi-modal fusion features.

[0086] Specifically, on the basis of data preprocessing, a deep learning model is used to extract features from each modality of data. Visual data: Use a 3D convolutional neural network (S3D) to extract spatio-temporal features in video frames and capture the movement trajectory and operation details of the excavator. Audio data: Use the VGGish model to extract audio features. By analyzing the spectral features of the audio signal, the working sound and abnormal noise of the excavator are identified. Kinematic data: Use the Conformer model to process kinematic sensor data and extract the dynamic change features of the excavator. Through the cross-modal attention mechanism, the self-attention mechanism, and the multi-head self-attention mechanism, the correlation between video, audio, and kinematic data is captured for multi-level feature fusion. Specifically in implementation: The cross-modal attention mechanism is adopted to enhance the correlation between each modality through a shared attention mechanism. The self-attention mechanism is used to capture the feature relationship within a single modality, and the multi-head self-attention mechanism is used to effectively fuse the features of different modalities. Weighted fusion is used to combine the features from different modalities, and after fusion, further processing is performed through a fully connected layer to obtain multi-modal fusion features. The multi-modal feature extraction and fusion of the present invention are as Figure 3 shown.

[0087] In one embodiment of the present invention, in step S2, the core task is to extract unimodal key features from the preprocessed visual, audio, and kinematic data, providing basic support for subsequent multimodal feature fusion and intelligent analysis.

[0088] Among them, for video data, the S3D (Separated 3D Convolutional Neural Network) model that separates spatial and temporal convolutions is adopted. By decoupling the convolutional operations of spatial and temporal features, it efficiently captures the spatio-temporal features in the excavator operation scenario. In the video feature extraction process, the preprocessed video frame sequence is used as the input. After passing through the S3D network, a high-dimensional feature vector is generated, comprehensively representing the spatio-temporal dynamics of the video. This model can not only accurately describe the excavator's movement trajectory but also identify the key visual information that dynamically changes in a complex operation environment and generate a high-dimensional spatio-temporal feature vector, providing rich visual features for the analysis of the equipment operating state.

[0089] Among them, for audio data, the VGGish model is used to extract deep features from the noise-reduced audio signal. Specifically, first, the preprocessed Mel spectrogram is input into the VGGish model. The VGGish model is used to extract high-level deep features from the audio data. The Mel spectrogram is processed through a deep convolutional network to capture the time-domain and frequency-domain features of the audio.

[0090] Among them, for kinematic data, the Conformer (Convolution-Transformer) model is adopted to extract the features of kinematic data. First, the preprocessed kinematic data is input into the Conformer model, and the Conformer model captures the temporal features.

[0091] After extracting the unimodal features of the above visual, audio, and kinematic data, all features need to be subjected to unified standardization and dimensionality reduction processing. The standardization processing eliminates the dimension difference, ensuring the consistency of different modal data features on the same scale. The implementation method is shown in Equation (1); the dimensionality reduction processing (principal component analysis, PCA) removes redundant information, reduces the feature dimension, improves the calculation efficiency, and at the same time retains the key features that are most representative for subsequent analysis.

[0092]

[0093] In one embodiment of the present invention, in step S3, it aims to effectively capture and fuse the correlations between different modal features through cross-modal attention mechanisms, self-attention mechanisms, and multi-head self-attention mechanisms, thereby generating multimodal fusion features with rich information expression capabilities.

[0094] Among them, the role of the self-attention mechanism is to capture the dependencies between different features within a single modality, thereby enhancing the expressive ability of important features. The implementation method is shown in Equation (2):

[0095]

[0096] Among them, Q represents the query matrix, K represents the key matrix, V represents the value matrix, and d k is the dimension of the key vector.

[0097] Among them, the role of the cross-modal attention mechanism is to strengthen the mutual correlation between the features of each modality by capturing the dependencies between different modalities

[0098] CrossModalAttention(X1,X2)=Concat(Attention 12 ,Attention 21 ) (3)

[0099] Among them, Attention 12 , Attention 21 are the attention of X1 to X2 and the attention of X2 to X1, respectively.

[0100] Among them, the multi-head self-attention mechanism realizes parallel computing of different attention distributions by introducing multiple attention heads (the number of heads is h), so as to capture multiple relationship patterns. The concept is to linearly transform the query, key, and value vectors respectively through multiple trainable weight matrices W i Q 、W i K 、W i V to generate multiple attention heads. Each attention head independently calculates the attention weights and generates the corresponding attention output. Subsequently, the outputs of all attention heads are concatenated (Concat) and linearly transformed through a final trainable weight matrix W O to form the final multi-head self-attention output. The implementation method is shown in Equation (4).

[0101] MultiHead(Q,K,V)=Concat(head1,head2,...,head n )W O (4)

[0102] Among them, each head i =Attention(QW i Q ,KW i K ,VWi V ),W i Q 、W i K 、W i V and W O are trainable weight matrices.

[0103] Furthermore, the feature fusion method adopts a multi-level feature fusion network. By constructing multi-layer attention modules, features of different modalities are fused in each layer, and higher-level semantic information is gradually extracted. In terms of the fusion strategy, cross-modal attention mechanism, self-attention mechanism and multi-head attention mechanism are simultaneously applied in each layer to weight-fuse the features and generate comprehensive feature representations.

[0104] Specifically, the specific process of the feature fusion method includes using the single-modal visual, audio and kinematic features extracted in step S2 as the initial input, calculating the correlation between modalities by the cross-modal attention mechanism to generate preliminary fusion features; then applying the self-attention mechanism and multi-head self-attention mechanism to the preliminary fusion features to further enhance the feature expression. Through this multi-level and comprehensive feature fusion method, not only can multi-modal information be effectively integrated, but also the recognition accuracy and robustness of the model can be improved.

[0105] S4. Use the multi-modal fusion features to train a machine learning model to obtain an optimal machine learning recognition model for outputting the recognition result of the excavator activity state.

[0106] It can be understood that in the steps of the present invention, the core goal is to achieve accurate classification and recognition of the excavator activity state through the training and optimization of the machine learning model, so as to provide a solid foundation for the analysis of the excavator production efficiency. The model is trained using the training data, and the model parameters are adjusted to improve the accuracy. It is tested on the validation set to evaluate the generalization ability of the model, and the model hyperparameters are optimized to ensure the robustness and real-time performance of the recognition model. An efficient classifier that can accurately identify the states such as excavation / transportation / shutdown of the excavator is constructed to provide support for safe production and production efficiency prediction. Potential abnormal states during the excavator operation are predicted, and fault warnings are provided in a timely manner to ensure operation safety. The specific operation process is as follows:

[0107] First, use the multi-modal fusion features obtained in step S3 as the input, and divide them into a training set and a validation set. The training set is used for model training. Through optimization algorithms such as backpropagation and gradient descent, the model parameters are continuously adjusted to minimize the loss function and improve the model's fitting ability to the data. The validation set is used to evaluate the performance of the model during the training process to help determine whether the model is overfitting or underfitting. Subsequently, the machine learning model SVM is selected for model training.

[0108] After the machine learning model is trained, a validation set is used to evaluate the performance of the model. The performance evaluation metrics include accuracy, precision, recall, F1-score, and confusion matrix, comprehensively measuring the performance of the model in classification and recognition tasks. According to the evaluation results of the validation set, the model structure, hyperparameters, and training strategy are adjusted to find the model with the best performance on the validation set.

[0109] Through the systematic training and validation process, the optimal recognition model with the best performance on the validation set is determined. This optimal recognition model has high classification accuracy and recognition efficiency, and can effectively distinguish different excavator activity states, such as excavation, transportation, loading, shutdown, etc. Combining the excavator activity state recognition results, the system can monitor the safety state of the machine, timely detect abnormal operations or potential risks, send warning signals, prevent accidents, and ensure the safety of the construction site. At the same time, through the excavator activity recognition results, the system can collect and record various operation data in real time, and further use it for the analysis and optimization of mechanical production efficiency. The optimal recognition model not only improves the overall operation efficiency, but also effectively reduces the cost and error of manual monitoring, and has broad application prospects and practical value.

[0110] S5. Based on the excavator activity state recognition results, calculate the excavator action time, average cycle time, and productivity under different activity states, and predict the future production efficiency according to the calculated productivity.

[0111] It can be understood that in the steps of the present invention, the purpose is to calculate the excavator action time, average cycle time (ACT), and productivity under different activity states based on the optimal recognition model trained in step S4, and predict the future production efficiency according to the calculated productivity.

[0112] Specifically, based on the optimal recognition model, the activity state of the excavator is efficiently recognized, the productivity is calculated according to the recognition results, and then the future production efficiency is predicted based on the productivity. Calculate the excavator action time and average cycle time (ACT) under different activity states. Use the calculated excavator action time and average cycle time (ACT) to calculate the excavator productivity. Predict the mining efficiency according to the calculated excavator productivity.

[0113] Among them, the action time is calculated based on the recognition results in the optimal recognition model. For each frame of visual data, the current operation action of the excavator is recognized. Each action (such as excavation, transportation, etc.) lasts for a certain period of time, and the productivity can be evaluated by the duration of each action. The time of each action is calculated by the formula:

[0114]

[0115] Among them, AT is the action time, EF and SF are the end frame and start frame respectively, and FR is the frame rate.

[0116] Among them, the average cycle time (ACT) is calculated. Add the time of each cyclic action (such as excavation time, transportation time, etc.) to obtain the average time of the whole cycle:

[0117] ACT = ADG + AH + ADP + AS (6)

[0118] Among them, ADG is the excavation time, AH is the transportation time, ADP is the unloading time, and AS is the swing time.

[0119] Among them, the productivity is calculated. Productivity reflects the working efficiency of the excavator, and the formula is as follows:

[0120]

[0121] Among them, P is the productivity (unit: cubic meters per hour), ACT is the average cycle time (unit: seconds), and ABP is the average bucket load (unit: cubic meters). It is calculated by ABP = HBC × BEF, where HBC is the full capacity of the bucket and BFF is the bucket filling factor.

[0122] Based on the productivity calculated above, the future productivity of the excavator is predicted through the deep learning model LSTM. Through the prediction results, managers can adjust the operation arrangement or dispatch equipment in advance to avoid the impact of low efficiency on the overall operation progress.

[0123] In summary, in the present invention, first, the visual, audio, and kinematic data are preprocessed, and the single-modal features are extracted by using the S3D, VGGish, and Conformer models respectively. Subsequently, cross-modal attention, self-attention, and multi-head self-attention mechanisms are adopted for multi-level feature fusion to generate high-quality multi-modal fusion feature vectors. Based on the above multi-modal fusion feature vectors, the deep learning model is trained and optimized to accurately classify and identify the activity states of construction machinery. Finally, the productivity is calculated by using the recognition results, and the future production efficiency is predicted based on the calculated productivity. This method significantly improves the accuracy and real-time performance of mechanical activity state recognition, is applicable to various types of excavators and operating environments, has broad application prospects and practical value, and provides an efficient, accurate, and reliable technical solution for the analysis of quarrying efficiency and the judgment of safety status.

[0124] The beneficial effects of the present invention are:

[0125] 1) Through the multi-modal data fusion technology, an efficient quarrying efficiency analysis model is established, which can comprehensively and accurately identify the activity states of excavators, providing a more intuitive basis for the real-time evaluation of quarrying efficiency.

[0126] 2) Combines cross-modal attention mechanism, self-attention mechanism and multi-head self-attention mechanism, which can effectively capture the complex correlations between different modal data and improve the performance of the model under multi-dimensional data fusion.

[0127] 3) Through the multi-level feature fusion method, the effective information of each modal data is maximally retained, improving the accuracy and robustness of the recognition of the excavator's activity state.

[0128] 4) By training and optimizing the multi-modal fusion features, a support vector machine (SVM) model is used to achieve efficient and accurate classification and recognition of the activity states of construction machinery, providing strong data support for the analysis of mechanical production efficiency and the assessment of safety states.

[0129] 5) Combining the action time, average cycle time (ACT) and productivity calculated by the excavator activity recognition model, accurate prediction of future production efficiency is achieved, further helping managers optimize operation arrangements and equipment scheduling and improving overall production efficiency.

[0130] According to the method for analyzing the quarrying efficiency based on multi-modal excavator intelligent monitoring according to the embodiments of the present invention, through the application of deep learning technology, not only the recognition accuracy of the excavator's activity state is improved, but also strong support is provided for the real-time prediction and optimization of the quarrying efficiency, thus providing a scientific and accurate decision-making basis for the construction management of water conservancy and hydropower projects.

[0131] To implement the above embodiments, as Figure 4 shown, in this embodiment, a system 10 for analyzing the quarrying efficiency based on multi-modal excavator intelligent monitoring is further provided, including:

[0132] A monitoring data acquisition module 100 for acquiring preprocessed excavator monitoring data; wherein, the excavator monitoring data includes video data, audio data and kinematic data;

[0133] A feature extraction module 200 for extracting a plurality of single-modal original features corresponding to the preprocessed excavator monitoring data;

[0134] A feature fusion module 300 for calculating the correlations between modalities of a plurality of single-modal original features by using a cross-modal attention strategy to generate preliminary fusion features, and integrating the preliminary fusion features into a self-attention mechanism and a multi-head self-attention mechanism to perform multi-level feature fusion to obtain multi-modal fusion features;

[0135] A model training module 400 for training a machine learning model by using the multi-modal fusion features to obtain an optimal machine learning recognition model to output the recognition result of the excavator's activity state;

[0136] The production efficiency prediction module 500 is used to calculate the action time, average cycle time, and productivity of the excavator in different activity states based on the excavator activity state recognition result, and predict the future production efficiency according to the calculated productivity.

[0137] Furthermore, before the monitoring data acquisition module 100, there is also a preprocessing module, which is used for:

[0138] Using the Gaussian filtering method to remove the noise of the video data, and performing size normalization processing on the video image to obtain the preprocessed video frame sequence;

[0139] Reading the original speech signal from the audio file, performing noise reduction using the spectral subtraction method, and converting the acoustic wave signal into a Mel spectrogram through the short-time Fourier transform;

[0140] The kinematic data, including acceleration, angular velocity, and position, uses a low-pass filter to remove the high-frequency noise of the kinematic data, retain the low-frequency signal, and use mean filtering to perform mean processing on the low-frequency signal with a sliding window to obtain the preprocessed kinematic data;

[0141] Using the time synchronization algorithm to align the multi-modal data to a unified time axis.

[0142] Furthermore, the feature extraction module 200 is also used for:

[0143] Inputting the preprocessed video frame sequence into the S3D model based on separable spatial and temporal convolutions, and outputting a high-dimensional feature vector;

[0144] Inputting the Mel spectrogram into the VGGish model, extracting high-level depth features from the audio data, and processing the Mel spectrogram through a deep convolutional network to obtain the time-domain and frequency-domain features of the audio;

[0145] Inputting the preprocessed kinematic data into the Conformer model to capture the temporal features; and,

[0146] Performing unified standardization and dimensionality reduction processing on multiple single-modal original features, including:

[0147]

[0148] Furthermore, the feature fusion module is also used for:

[0149] Self-attention mechanism:

[0150]

[0151] Among them, Q represents the query matrix, K represents the key matrix, V represents the value matrix, and d k is the dimension of the key vector;

[0152] Cross-modal attention mechanism:

[0153] CrossModalAttention(X1,X2) = Concat(Attention 12 , Attention 21 ) (3)

[0154] where Attention 12 , Attention 21 are the attention of X1 to X2 and the attention of X2 to X1, respectively;

[0155] Multi-head self-attention mechanism, which linearly transforms the query, key, and value vectors through multiple weight matrices W i Q , W i K , W i V for training to generate multiple attention heads; each attention head independently calculates the attention weights and generates the corresponding attention output, and then the outputs of all attention heads are concatenated and linearly transformed through a final weight matrix W O for training to form the final multi-head self-attention output:

[0156] MultiHead(Q,K,V) = Concat(head1,head2,...,head n )W O (4)

[0157] where each head i = Attention(QW i Q ,KW i K ,VW i V ),W i Q ,W i K ,W i V and W O are weight matrices for training.

[0158] The system for analyzing the quarrying efficiency based on multi-modal intelligent monitoring of excavators according to the embodiments of the present invention, and the method for analyzing the quarrying efficiency based on multi-modal intelligent monitoring of excavators. Through the application of deep learning technology, not only the recognition accuracy of the activity state of the excavator is improved, but also strong support is provided for the real-time prediction and optimization of the quarrying efficiency, thereby providing a scientific and accurate decision-making basis for the construction management of water conservancy and hydropower projects.

[0159] In the description of this specification, the descriptions referring to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms are not necessarily directed to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0160] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise specifically defined.

Claims

1. A method for analyzing the quarrying efficiency based on multi-modal intelligent monitoring of excavators, characterized in that, Including: Obtain the preprocessed excavator monitoring data; wherein, the excavator monitoring data includes video data, audio data, and kinematic data; Extract multiple single-modal original features corresponding to the preprocessed excavator monitoring data; Use a cross-modal attention strategy to calculate the correlation between modalities of multiple single-modal original features to generate a preliminary fusion feature, and integrate the preliminary fusion feature into the self-attention mechanism and the multi-head self-attention mechanism to perform multi-level feature fusion to obtain a multi-modal fusion feature; Use the multi-modal fusion feature to train a machine learning model to obtain an optimal machine learning recognition model to output the recognition result of the excavator activity state; Calculate the action time, average cycle time, and productivity of the excavator in different activity states based on the excavator activity state recognition result, and predict the future production efficiency according to the calculated productivity.

2. The method according to claim 1, characterized in that, Preprocess the excavator monitoring data, including: Use the Gaussian filtering method to remove the noise of the video data, and perform size normalization processing on the video image to obtain the preprocessed video frame sequence; Read the original speech signal from the audio file, perform noise reduction using the spectral subtraction method, and convert the sound wave signal into a Mel spectrogram through the short-time Fourier transform; The kinematic data, including acceleration, angular velocity, and position, uses a low-pass filter to remove the high-frequency noise of the kinematic data, retain the low-frequency signal, and use mean filtering to perform mean processing of the sliding window on the low-frequency signal to obtain the preprocessed kinematic data; Use the time synchronization algorithm to align the data of each modality to a unified time axis.

3. The method according to claim 2, wherein Extract multiple single-modal original features corresponding to the preprocessed excavator monitoring data, including: Input the preprocessed video frame sequence into the S3D model based on separable space and time convolution to output a high-dimensional feature vector; Input the Mel spectrogram into the VGGish model, extract high-level deep features from the audio data, and process the Mel spectrogram through a deep convolutional network to obtain the time domain and frequency domain features of the audio; Input the preprocessed kinematic data into the Conformer model to capture the temporal features; and Perform unified standardization and dimensionality reduction processing on multiple single-modal original features, including:

4. The method according to claim 1, characterized in that, Use a cross-modal attention strategy to calculate the correlation between modalities of multiple single-modal original features to generate a preliminary fusion feature, and integrate the preliminary fusion feature into the self-attention mechanism and the multi-head self-attention mechanism to perform multi-level feature fusion to obtain a multi-modal fusion feature, including: Self-attention mechanism: Among them, Q represents the query matrix, K represents the key matrix, V represents the value matrix, and d k is the dimension of the key vector; Cross-modal attention mechanism: CrossModalAttention(X1,X2) = Concat(Attention 12 , Attention 21 ) (3) Among them, Attention 12 , Attention 21 are the attention of X1 to X2 and the attention of X2 to X1 respectively; The multi-head self-attention mechanism linearly transforms the query, key, and value vectors respectively through multiple weight matrices W for training i Q , W i K , W i V to generate multiple attention heads; each attention head independently calculates the attention weights and generates the corresponding attention output, and then the outputs of all attention heads are concatenated and linearly transformed through a final weight matrix W for training O to form the final multi-head self-attention output: MultiHead(Q,K,V)=Concat(head1,head2,...,head n )W O (4) Among them, each head i = Attention(QW i Q , KW i K , VW i V ), W i Q 、W i K 、W i V and W O are weight matrices for training.

5. The method according to claim 1, wherein Use the multi-modal fusion feature to train a machine learning model to obtain an optimal machine learning recognition model, including: Divide the multi-modal fusion feature into a training set and a validation set; Input the training set into the machine learning model SVM, and use the optimization algorithms of backpropagation and gradient descent to adjust the model parameters to minimize the loss function; After the machine learning model is trained, use the validation set to evaluate the performance of the model, and adjust the model structure, hyperparameters, and training strategy according to the evaluation results to find the optimal machine learning recognition model; wherein, the performance evaluation indicators include accuracy, precision, recall, F1 score, and confusion matrix; Output the activity status of the excavator based on the optimal machine learning recognition model; among them, the activity status of the excavator includes multiple of digging, transporting, loading, and shutdown statuses.

6. The method according to claim 1, wherein Calculate the action time, average cycle time, and productivity of the excavator in different activity statuses based on the recognition result of the excavator activity status, and predict the future production efficiency based on the calculated productivity, including: The time of each action is calculated by the formula: Among them, AT is the action time, EF and SF are the end frame and start frame respectively, and FR is the frame rate; Add up the time of each cyclic action to obtain the average time of the whole cycle: ACT = ADG + AH + ADP + AS (6) Among them, ADG is the digging time, AH is the transportation time, ADP is the unloading time, and AS is the swing time. The productivity calculation formula is as follows: Among them, P is the productivity, ACT is the average cycle time, and ABP is the average bucket load; it is calculated by ABP = HBC × BEF, where HBC is the full bucket capacity and BFF is the bucket filling factor; Based on the productivity, predict the future productivity of the excavator through the deep learning model LSTM, and make production adjustments based on the prediction results.

7. A system for analyzing the quarrying efficiency based on multi-modal intelligent monitoring of excavators, characterized in that, Including: A monitoring data acquisition module for acquiring preprocessed excavator monitoring data; among them, the excavator monitoring data includes video data, audio data, and kinematic data; A feature extraction module for extracting multiple single-modal raw features corresponding to the preprocessed excavator monitoring data; A feature fusion module for calculating the correlation between each modality of multiple single-modal raw features using a cross-modal attention strategy to generate preliminary fusion features, and integrating the preliminary fusion features into a self-attention mechanism and a multi-head self-attention mechanism to perform multi-level feature fusion to obtain multi-modal fusion features; A model training module for training a machine learning model using multi-modal fusion features to obtain an optimal machine learning recognition model to output the recognition result of the excavator activity status; A production efficiency prediction module for calculating the action time, average cycle time, and productivity of the excavator in different activity statuses based on the recognition result of the excavator activity status, and predicting the future production efficiency based on the calculated productivity.

8. The system according to claim 7, wherein Before the monitoring data acquisition module, there is also a preprocessing module for: Using the Gaussian filtering method to remove the noise of the video data and perform size normalization processing on the video image to obtain a preprocessed video frame sequence; Read the original speech signal from the audio file, perform noise reduction using spectral subtraction, and convert the sound wave signal into a Mel spectrogram through short-time Fourier transform; The kinematic data includes acceleration, angular velocity, and position. Use a low-pass filter to remove the high-frequency noise of the kinematic data, retain the low-frequency signal, and use mean filtering to perform mean processing on the low-frequency signal with a sliding window to obtain preprocessed kinematic data; Use a time synchronization algorithm to align the data of each modality to a unified time axis.

9. The system according to claim 8, wherein The feature extraction module is also used for: Input the preprocessed video frame sequence into the S3D model based on separable spatial and temporal convolutions to output a high-dimensional feature vector; Input the Mel spectrogram into the VGGish model to extract high-level deep features from the audio data, and process the Mel spectrogram through a deep convolutional network to obtain the time-domain and frequency-domain features of the audio; Input the preprocessed kinematic data into the Conformer model to capture temporal features; And, Perform unified standardization and dimensionality reduction processing on multiple single-modal raw features, including:

10. The system according to claim 7, wherein The feature fusion module is also used for: Self-attention mechanism: Among them, Q represents the query matrix, K represents the key matrix, V represents the value matrix, and d k is the dimension of the key vector; Cross-modal attention mechanism: CrossModalAttention(X1,X2) = Concat(Attention 12 , Attention 21 ) (3) Among them, Attention 12 , Attention 21 are the attention of X1 to X2 and the attention of X2 to X1 respectively; The multi-head self-attention mechanism linearly transforms the query, key, and value vectors respectively through multiple weight matrices W for training i Q , W i K , W i V to generate multiple attention heads; each attention head independently calculates the attention weights and generates the corresponding attention output, and then the outputs of all attention heads are concatenated and linearly transformed through a final weight matrix W for training O to form the final multi-head self-attention output: MultiHead(Q,K,V)=Concat(head1,head2,...,head n )W O (4) Among them, each head i = Attention(QW i Q ,KW i K ,VW i V ),W i Q 、W i K 、W i V and W O are weight matrices for training.

Citation Information

Cited By

  • Power distribution network fault detection method and system based on multi-source data fusion

    CN121093165A