Method and system for detecting large blocks based on multi-sensor fusion

By integrating vibration and sound sensors onto the excavator and combining Fourier transform and Clip model feature matrix fusion technology, the problem of the excavator's inability to detect large blocks was solved, enabling efficient excavation.

CN115906000BActive Publication Date: 2026-01-30HUANENG YIMIN COAL POWER CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202211467044.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-22
Publication Date
2026-01-30
Estimated Expiration
2042-11-22

AI Technical Summary

Technical Problem

The existing excavators lack effective detection capabilities for excavating large rocks, which causes delays in the excavation process when encountering large rocks, affecting the progress of the project.

Method used

A multi-sensor fusion approach is adopted, which collects signals through vibration and sound sensors, and extracts time-domain and frequency-domain feature information of vibration and sound signals by combining Fourier transform and Clip model. Feature matrix fusion is then performed by orthographic projection nonlinear reweighting to achieve real-time and accurate large-block detection.

Benefits of technology

It enables real-time and accurate detection of large excavated blocks, ensuring the smooth progress of excavation work and improving detection accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115906000B_ABST
    Figure CN115906000B_ABST
Patent Text Reader

Abstract

This application relates to the field of excavator technology, specifically disclosing a method and system for detecting large excavated blocks based on multi-sensor fusion. It collects vibration and sound detection signals from the excavator during the excavation process using vibration and sound sensors, then performs Fourier transforms on each to extract frequency domain features. Next, a Clip model incorporating a sequence encoder and an image encoder is used to extract fused feature information of the time-domain and frequency-domain features of the vibration and sound signals during the excavation process. Preferably, considering that the vibration and sound feature matrices obtained through the Clip model may have opposite correlation directions at corresponding positions, a fully orthographic projection nonlinear reweighting method is used to fuse the vibration and sound feature matrices. This approach enables real-time and accurate detection of large excavated blocks, thereby ensuring the smooth progress of the excavation work.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of excavator technology, and more specifically, to a method and system for detecting large excavated blocks based on multi-sensor fusion. Background Technology

[0002] With the development of society and economy, urban construction is accelerating, and excavators are indispensable in urban construction. Excavators are earthmoving machines that use their buckets to dig materials above or below the machine's bearing surface and load them into transport vehicles or unload them into stockpiles. In recent years, the development of construction machinery has been relatively rapid, and excavators have become one of the most important pieces of construction machinery.

[0003] During excavation, excavators encounter numerous rocks and clods of varying sizes in mountainous terrain, making them difficult to distinguish. Current excavators lack the capability to detect large clods, which hinders excavation, slows down the process, and ultimately reduces the overall construction progress. Currently, when excavators unearth large rocks and clods, they need to be broken down into smaller pieces for further processing. Therefore, the ability to detect large clods during excavation is crucial.

[0004] Therefore, an optimized scheme for detecting large blocks is desired. Summary of the Invention

[0005] To address the aforementioned technical problems, this application is proposed. Embodiments of this application provide a method and system for detecting large excavated blocks based on multi-sensor fusion. This method collects vibration and sound detection signals from the excavator during the excavation process using vibration and sound sensors. Fourier transforms are then performed on each signal to extract frequency domain features. Next, a Clip model incorporating a sequence encoder and an image encoder is used to extract fused feature information of the time-domain and frequency-domain features of the vibration and sound signals during the excavation process. Preferably, considering that the vibration and sound feature matrices obtained through the Clip model may have opposite correlation directions at corresponding positions, a fully orthographic projection nonlinear reweighting method is used to fuse the vibration and sound feature matrices. This approach enables real-time and accurate detection of large excavated blocks, thereby ensuring the smooth progress of the excavation work.

[0006] According to one aspect of this application, a method for detecting large blocks in mining based on multi-sensor fusion is provided, comprising:

[0007] Acquire vibration and sound detection signals collected by vibration and sound sensors during the excavation process;

[0008] The vibration detection signal and the sound detection signal are subjected to Fourier transform to obtain multiple vibration frequency domain statistical feature values ​​and multiple sound frequency domain statistical feature values;

[0009] The vibration feature matrix is ​​obtained by passing the multiple vibration frequency domain statistical feature values ​​and the waveform of the vibration detection signal through a first Clip model that includes a sequence encoder and an image encoder.

[0010] The multiple audio frequency domain statistical feature values ​​and the waveform of the sound detection signal are passed through a second Clip model containing a sequence encoder and an image encoder to obtain a sound feature matrix;

[0011] The vibration feature matrix and the sound feature matrix are fused to obtain a classification feature matrix; and

[0012] The classification feature matrix is ​​passed through a classifier to obtain a classification result, which is used to indicate whether a large block has been discovered.

[0013] According to another aspect of this application, a large-block detection system based on multi-sensor fusion is provided, comprising:

[0014] The signal acquisition module is used to acquire vibration detection signals and sound detection signals collected by vibration sensors and sound sensors during the excavation process;

[0015] The frequency domain feature extraction module is used to perform Fourier transform on the vibration detection signal and the sound detection signal respectively to obtain multiple vibration frequency domain statistical feature values ​​and multiple sound frequency domain statistical feature values;

[0016] The vibration signal encoding module is used to obtain a vibration feature matrix by passing the multiple vibration frequency domain statistical feature values ​​and the waveform of the vibration detection signal through a first Clip model containing a sequence encoder and an image encoder;

[0017] The sound signal encoding module is used to obtain a sound feature matrix by passing the multiple audio frequency domain statistical feature values ​​and the waveform of the sound detection signal through a second Clip model containing a sequence encoder and an image encoder;

[0018] A fusion module is used to fuse the vibration feature matrix and the sound feature matrix to obtain a classification feature matrix; and

[0019] The classification module is used to pass the classification feature matrix through a classifier to obtain a classification result, which is used to indicate whether a large block has been discovered.

[0020] According to another aspect of this application, an electronic device is provided, comprising: a processor; and a memory storing computer program instructions, which, when executed by the processor, cause the processor to perform the multi-sensor fusion-based large-block detection method as described above.

[0021] According to another aspect of this application, a computer-readable medium is provided having computer program instructions stored thereon, which, when executed by a processor, cause the processor to perform the multi-sensor fusion-based large block detection method as described above.

[0022] Compared with existing technologies, this application provides a method and system for detecting large excavated blocks based on multi-sensor fusion. It collects vibration and sound detection signals from the excavator during the excavation process using vibration and sound sensors, then performs Fourier transforms on each to extract frequency domain features. Next, a Clip model containing a sequence encoder and an image encoder is used to extract fused feature information of the time-domain and frequency-domain features of the vibration and sound signals during the excavation process. Preferably, considering that the vibration and sound feature matrices obtained through the Clip model may have opposite correlation directions at corresponding positions, a fully orthographic projection nonlinear reweighting method is used to fuse the vibration and sound feature matrices. This approach enables real-time and accurate detection of large excavated blocks, thereby ensuring the smooth progress of the excavation work. Attached Figure Description

[0023] The above and other objects, features, and advantages of this application will become more apparent from the more detailed description of the embodiments of this application in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of this application and form part of the specification. They are used together with the embodiments of this application to explain this application and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.

[0024] Figure 1 The illustration shows an application scenario of a multi-sensor fusion-based method and system for detecting large blocks in mining, according to an embodiment of this application.

[0025] Figure 2 The illustration shows a flowchart of a method for detecting large blocks based on multi-sensor fusion according to an embodiment of this application.

[0026] Figure 3 The figure illustrates a schematic diagram of the system architecture of a large-block detection method based on multi-sensor fusion according to an embodiment of this application.

[0027] Figure 4The illustration shows a flowchart of a method and system for detecting large blocks of earth based on multi-sensor fusion according to an embodiment of this application, in which the vibration feature matrix is ​​obtained by passing the multiple vibration frequency domain statistical feature values ​​and the waveform of the vibration detection signal through a first Clip model including a sequence encoder and an image encoder.

[0028] Figure 5 The illustration shows a flowchart of a method and system for detecting large blocks in mining based on multi-sensor fusion according to an embodiment of this application, in which the multiple vibration frequency domain statistical feature values ​​are input into the sequence encoder of the first Clip model to obtain a vibration frequency statistical feature vector.

[0029] Figure 6 The figure shows a block diagram of a large-block detection system based on multi-sensor fusion according to an embodiment of this application.

[0030] Figure 7 The figure shows a block diagram of a vibration signal encoding module in a large-block excavation detection system based on multi-sensor fusion according to an embodiment of this application.

[0031] Figure 8 The figure shows a block diagram of a vibration frequency encoding unit in a large-block excavation detection system based on multi-sensor fusion according to an embodiment of this application.

[0032] Figure 9 A block diagram of an electronic device according to an embodiment of this application is illustrated. Detailed Implementation

[0033] Hereinafter, exemplary embodiments according to this application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this application, and not all embodiments of this application. It should be understood that this application is not limited to the exemplary embodiments described herein.

[0034] Scene Overview

[0035] As mentioned above, with the development of society and economy, urban construction is accelerating, and excavators are indispensable in urban construction. Excavators are earthmoving machines that use their buckets to dig materials above or below the machine's bearing surface and load them into transport vehicles or unload them into stockpiles. In recent years, the development of construction machinery has been relatively rapid, and excavators have become one of the most important construction machines.

[0036] During excavation, excavators encounter numerous rocks and clods of varying sizes in mountainous terrain, making them difficult to separate. Existing excavators lack the capability to detect large clods, which hinders excavation, slows down the process, and ultimately reduces the overall construction progress. Currently, when excavators encounter large rocks and clods, they need to be broken down into smaller pieces for further processing. Therefore, detecting large clods during excavation is crucial. Thus, an optimized solution for detecting large clods during excavation is desired.

[0037] Currently, deep learning and neural networks have been widely applied in fields such as computer vision, natural language processing, and speech signal processing. Furthermore, deep learning and neural networks have demonstrated near-human or even surpassed human-level performance in areas such as image classification, object detection, semantic segmentation, and text translation.

[0038] In recent years, the development of deep learning and neural networks has provided new ideas and solutions for intelligent detection of large blocks.

[0039] Correspondingly, when excavators are detecting large blocks, some existing solutions rely on the vibrations generated during the excavation process. This is because the vibration signals when large blocks are encountered differ from those of normal blocks. While this method can effectively identify abnormal vibrations within soil and rocks, its accuracy is lower in complex environments due to the varying structures within these blocks. Furthermore, considering the periodic changes in the sound signals emitted by excavators during normal excavation, fusing the characteristics of both—using sound signal features to enhance the representation of vibration signal features—would significantly improve the accuracy of large block detection.

[0040] Based on this, the technical solution of this application employs deep learning-based artificial intelligence detection technology to extract fused feature information of the time-domain and frequency-domain characteristics of vibration and sound signals during the excavation process. This information is used to construct a multi-sensor fusion-based large-block excavation detection scheme for intelligent detection of large blocks excavated by the excavator. This enables real-time and accurate detection of large blocks, thereby ensuring the smooth progress of the excavation work.

[0041] Specifically, in the technical solution of this application, firstly, vibration detection signals and sound detection signals are collected during the excavation process using vibration sensors and sound sensors, respectively. Next, considering that using the time-domain characteristics of the vibration and sound detection signals to detect large excavated blocks can contain significant environmental interference information, which can severely impact the detection results, the frequency-domain statistical characteristics of the detection signals are further combined to improve detection accuracy. That is, specifically, Fourier transforms are performed on the vibration and sound detection signals respectively to obtain multiple vibration frequency-domain statistical feature values ​​and multiple sound frequency-domain statistical feature values.

[0042] Then, for vibration feature extraction of the vibration detection signal, a first Clip model, including a sequence encoder and an image encoder, is used to process the multiple vibration frequency domain statistical feature values ​​and the waveform of the vibration detection signal to obtain a vibration feature matrix. Specifically, the sequence encoder of the first Clip model performs feature mining on the multiple vibration frequency domain statistical feature values ​​of the vibration detection signal to extract the multi-scale implicit feature distribution information of the frequency domain statistical features of the vibration detection signal; and the image encoder of the first Clip model performs feature mining on the waveform of the vibration detection signal to extract the temporal implicit feature information of the vibration detection signal; then, based on the multi-scale implicit feature information of the frequency domain statistical feature values ​​of the vibration detection signal, image attribute encoding optimization is performed on the temporal implicit features of the vibration detection signal waveform to obtain the vibration feature matrix. In this way, the obtained vibration feature matrix not only contains the frequency domain feature content of the vibration detection signal but also reflects the changing characteristics of the frequency domain content over time, improving the accuracy of large-block detection.

[0043] Similarly, for the extraction of sound features from the sound detection signal, considering that the periodic features of the sound detection signal and the periodic features of the vibration detection signal have similar regularity, the Clip model is also used for sound signal encoding in the technical solution of this application. Specifically, the multiple audio frequency domain statistical feature values ​​and the waveform of the sound detection signal are processed by a second Clip model containing a sequence encoder and an image encoder to obtain a sound feature matrix. Then, based on the multi-scale implicit features of the frequency domain statistical feature values ​​of the sound detection signal, image attribute encoding optimization is performed on the time domain implicit features of the sound detection signal waveform to obtain the sound feature matrix.

[0044] Furthermore, by fusing the feature information from the vibration feature matrix and the sound feature matrix, and then classifying them using a classifier, a classification result indicating whether a large chunk has been excavated can be obtained. This allows for intelligent detection of large chunks during the excavation process.

[0045] Specifically, in the technical solution of this application, since both the vibration feature matrix and the sound feature matrix are obtained by fusing the parametric semantics of frequency domain statistical feature values ​​and the image semantics of the corresponding waveform using the CLIP model, the feature value at each position of the vibration feature matrix and the sound feature matrix represents the bitwise correlation value between the parametric semantics of the frequency domain statistical feature values ​​and the image semantics of the corresponding waveform. Therefore, in the vibration feature matrix and the sound feature matrix, there are cases where the correlation directions represented by corresponding positions are opposite. This leads to a negative correlation between corresponding positions in the vibration feature matrix and the sound feature matrix, thereby affecting the bitwise fusion effect of the vibration feature matrix and the sound feature matrix.

[0046] Therefore, the applicant of this application uses a fully orthographic projection nonlinear reweighting method to fuse the vibration feature matrix and the sound feature matrix, as follows:

[0047]

[0048] Among them, M1, M2 and M c These are the vibration feature matrix, the sound feature matrix, and the classification feature matrix, respectively, where ReLU(·) represents the ReLU activation function. This indicates matrix multiplication, where the division between the numerator and denominator matrices is a positional division of the matrix eigenvalues. exp(·) represents matrix exponentiation, which means calculating the natural exponential function value raised to the power of the eigenvalues ​​at each position in the matrix.

[0049] Here, the orthographic projection nonlinear reweighting uses the ReLU function to ensure the projection is entirely positive to avoid aggregating negatively correlated information. Simultaneously, a nonlinear reweighting mechanism is introduced to aggregate the eigenvalue distribution of the feature matrix, allowing the intrinsic structure of the feature matrix to penalize long-distance connections and strengthen local coupling. This achieves a synergistic effect between the vibration feature matrix and the sound feature matrix in the high-dimensional feature space corresponding to the orthographic projection reweighting spatial feature transform, improving the fusion effect of the vibration feature matrix and the sound feature matrix. This enables real-time and accurate detection of large sections being excavated by the excavator during the excavation process, thus ensuring the smooth progress of the excavation work.

[0050] Based on this, this application provides a method for detecting large blocks during excavation based on multi-sensor fusion, comprising: acquiring vibration detection signals and sound detection signals collected by vibration sensors and sound sensors during the excavation process; performing Fourier transforms on the vibration detection signals and the sound detection signals respectively to obtain multiple vibration frequency domain statistical feature values ​​and multiple sound frequency domain statistical feature values; passing the waveforms of the multiple vibration frequency domain statistical feature values ​​and the vibration detection signals through a first Clip model including a sequence encoder and an image encoder to obtain a vibration feature matrix; passing the waveforms of the multiple sound frequency domain statistical feature values ​​and the sound detection signals through a second Clip model including a sequence encoder and an image encoder to obtain a sound feature matrix; fusing the vibration feature matrix and the sound feature matrix to obtain a classification feature matrix; and passing the classification feature matrix through a classifier to obtain a classification result, wherein the classification result is used to indicate whether a large block has been excavated.

[0051] Figure 1 The illustration shows an application scenario of a multi-sensor fusion-based method and system for detecting large blocks in mining, according to an embodiment of this application. For example... Figure 1 As shown, in this application scenario, by deploying on an excavator (e.g., such as...) Figure 1 The vibration sensor on C) as shown (e.g., such as Figure 1 The V shown is shown) and the sound sensor (e.g., such as Figure 1 The T shown in the diagram collects vibration and sound detection signals from the excavator during the excavation process. Then, the collected vibration and sound detection signals are input to a server deployed with a large-block excavation detection algorithm based on multi-sensor fusion (e.g., Figure 1 As shown in S), the server is able to use the multi-sensor fusion-based large block detection algorithm to process the vibration detection signal and the sound detection signal to generate a classification result indicating whether a large block has been discovered.

[0052] After introducing the basic principles of this application, various non-limiting embodiments of this application will be described in detail below with reference to the accompanying drawings.

[0053] Exemplary methods

[0054] Figure 2 The illustration shows a flowchart of a large-block detection method based on multi-sensor fusion according to an embodiment of this application. Figure 2As shown, the method for detecting large blocks in excavation based on multi-sensor fusion according to an embodiment of this application includes: S110, acquiring vibration detection signals and sound detection signals collected by vibration sensors and sound sensors during the excavation process; S120, performing Fourier transforms on the vibration detection signals and the sound detection signals respectively to obtain multiple vibration frequency domain statistical feature values ​​and multiple sound frequency domain statistical feature values; S130, passing the waveforms of the multiple vibration frequency domain statistical feature values ​​and the vibration detection signals through a first Clip model including a sequence encoder and an image encoder to obtain a vibration feature matrix; S140, passing the waveforms of the multiple sound frequency domain statistical feature values ​​and the sound detection signals through a second Clip model including a sequence encoder and an image encoder to obtain a sound feature matrix; S150, fusing the vibration feature matrix and the sound feature matrix to obtain a classification feature matrix; and S160, passing the classification feature matrix through a classifier to obtain a classification result, the classification result being used to indicate whether a large block has been excavated.

[0055] Figure 3 The illustration shows a schematic diagram of the system architecture of a large-block detection method based on multi-sensor fusion according to an embodiment of this application. Figure 3 As shown in the system architecture of the multi-sensor fusion-based method for detecting large blocks during excavation, according to an embodiment of this application, firstly, vibration and sound detection signals collected by vibration and sound sensors during the excavation process are acquired. Then, Fourier transforms are performed on the vibration and sound detection signals respectively to obtain multiple vibration frequency domain statistical feature values ​​and multiple sound frequency domain statistical feature values. Next, the waveforms of the multiple vibration frequency domain statistical feature values ​​and the vibration detection signals are passed through a first Clip model containing a sequence encoder and an image encoder to obtain a vibration feature matrix. Simultaneously, the waveforms of the multiple sound frequency domain statistical feature values ​​and the sound detection signals are passed through a second Clip model containing a sequence encoder and an image encoder to obtain a sound feature matrix. Finally, the vibration feature matrix and the sound feature matrix are fused to obtain a classification feature matrix, and the classification feature matrix is ​​passed through a classifier to obtain a classification result, which indicates whether a large block has been excavated.

[0056] In step S110 of this embodiment, vibration detection signals and sound detection signals collected by vibration sensors and sound sensors during the excavation process are acquired. As mentioned above, during the excavation process, existing excavators encounter numerous rocks and soil clumps of varying sizes in the excavated soil and mountainous areas, making them difficult to distinguish. Furthermore, existing excavators lack large-clump detection capabilities. When large rocks and soil clumps are encountered, it hinders the excavation work, delays the process, and slows down the overall construction progress. Currently, when excavators encounter large rocks and soil clumps, they need to be broken down into smaller pieces for further processing. Therefore, large-clump detection during the excavation process is crucial. Thus, an optimized large-clump detection solution is desired.

[0057] Specifically, when excavators are detecting large blocks, some existing solutions rely on the vibrations generated during the excavation process. This is because the vibration signal when a large block is encountered differs from the normal signal. While this method can effectively identify abnormal vibration information within soil and rocks, its accuracy is lower in complex environments due to the varying structures within these blocks. Furthermore, considering the periodic changes in the sound signal generated by the excavator during normal excavation, fusing the characteristics of both sound and vibration signals to enhance the representation of vibration signal features would significantly improve the accuracy of large block detection.

[0058] Based on this, the technical solution of this application employs deep learning-based artificial intelligence detection technology to extract fused feature information of the time-domain and frequency-domain characteristics of vibration and sound signals during the excavation process. This information is used to construct a multi-sensor fusion-based large-block excavation detection scheme for intelligent detection of large blocks excavated by the excavator. This enables real-time and accurate detection of large blocks, thereby ensuring the smooth progress of the excavation work.

[0059] Specifically, in the technical solution of this application, the vibration sensor and the sound sensor deployed on the excavator collect the vibration detection signal and sound detection signal of the excavator during the excavation process.

[0060] In step S120 of this embodiment, the vibration detection signal and the sound detection signal are subjected to Fourier transforms respectively to obtain multiple vibration frequency domain statistical feature values ​​and multiple sound frequency domain statistical feature values. It should be understood that, considering the use of the time domain features of the vibration detection signal and the sound detection signal to detect large excavations, the time domain features contain a significant amount of environmental interference information, which can severely affect the detection results. Therefore, the frequency domain statistical features of the detection signals are further combined to improve the accuracy of the detection. Specifically, the vibration detection signal and the sound detection signal are subjected to Fourier transforms respectively to obtain multiple vibration frequency domain statistical feature values ​​and multiple sound frequency domain statistical feature values.

[0061] In step S130 of this embodiment, the plurality of vibration frequency domain statistical feature values ​​and the waveform of the vibration detection signal are processed by a first Clip model comprising a sequence encoder and an image encoder to obtain a vibration feature matrix. It should be understood that, considering that the plurality of vibration frequency domain statistical feature values ​​and the waveform of the vibration detection signal belong to different modalities, the Clip model has significant advantages in data fusion across different modalities. Specifically, the first Clip model comprises a sequence encoder and an image encoder. The sequence encoder of the first Clip model performs feature mining on the plurality of vibration frequency domain statistical feature values ​​of the vibration detection signal to extract the frequency domain implicit feature information of the frequency domain statistical features of the vibration detection signal. Furthermore, the image encoder of the first Clip model performs feature mining on the waveform of the vibration detection signal to extract the time domain implicit feature information of the vibration detection signal. Furthermore, considering that the multiple vibration frequency domain statistical feature values ​​have different periodic distributions across different time spans, in order to extract the multi-scale implicit feature information of the vibration detection signal's frequency domain statistical feature values, the frequency domain statistical feature values ​​of the vibration detection signal are convolved at different scales and then fused. Then, based on the multi-scale implicit feature information of the vibration detection signal's frequency domain statistical feature values, image attribute encoding optimization is performed on the temporal implicit features of the vibration detection signal waveform to obtain the vibration feature matrix. In this way, the obtained vibration feature matrix not only contains the frequency domain feature content of the vibration detection signal but also reflects the changing characteristics of the frequency domain content over time, improving the accuracy of large-block detection.

[0062] Figure 4 The illustration shows a flowchart of a method and system for detecting large blocks in excavation based on multi-sensor fusion according to an embodiment of this application. The flowchart describes how multiple vibration frequency domain statistical feature values ​​and waveforms of the vibration detection signals are processed through a first Clip model including a sequence encoder and an image encoder to obtain a vibration feature matrix. Figure 4As shown, in a specific embodiment of this application, the step of obtaining a vibration feature matrix by passing the plurality of vibration frequency domain statistical feature values ​​and the waveform of the vibration detection signal through a first Clip model including a sequence encoder and an image encoder includes: S210, inputting the plurality of vibration frequency domain statistical feature values ​​into the sequence encoder of the first Clip model to obtain a vibration frequency statistical feature vector; S220, inputting the waveform of the vibration detection signal into the image encoder of the first Clip model to obtain a vibration waveform feature vector; and S230, performing image attribute encoding optimization on the vibration waveform feature vector based on the vibration frequency statistical feature vector to obtain the vibration feature matrix.

[0063] Figure 5 The illustration shows a flowchart of a method and system for detecting large blocks in excavation based on multi-sensor fusion according to an embodiment of this application, in which the multiple vibration frequency domain statistical feature values ​​are input into the sequence encoder of the first Clip model to obtain a vibration frequency statistical feature vector. Figure 5 As shown, in a specific embodiment of this application, the step of inputting the plurality of vibration frequency domain statistical feature values ​​into the sequence encoder of the first Clip model to obtain a vibration frequency statistical feature vector includes: S310, arranging the plurality of vibration frequency domain statistical feature values ​​into a vibration frequency domain statistical input vector; S320, using the first convolutional layer of the sequence encoder of the first Clip model to perform one-dimensional convolutional encoding on the vibration frequency domain statistical input vector at a first scale to obtain a first-scale vibration frequency domain statistical feature vector; S330, using the second convolutional layer of the sequence encoder of the first Clip model to perform one-dimensional convolutional encoding on the vibration frequency domain statistical input vector at a second scale to obtain a second-scale vibration frequency domain statistical feature vector; and S340, concatenating the first-scale vibration frequency domain statistical feature vector and the second-scale vibration frequency domain statistical feature vector to obtain the vibration frequency statistical feature vector.

[0064] In a specific embodiment of this application, the step of inputting the waveform of the vibration detection signal into the image encoder of the first Clip model to obtain the vibration waveform feature vector includes: using each layer of the image encoder of the first Clip model to perform convolution processing, pooling processing and nonlinear activation processing on the input data in the forward propagation of the layer, so that the vibration waveform feature vector is output by the last layer of the image encoder of the first Clip model.

[0065] In a specific embodiment of this application, the step of optimizing the vibration waveform feature vector by image attribute encoding based on the vibration frequency statistical feature vector to obtain the vibration feature matrix includes: optimizing the vibration waveform feature vector by image attribute encoding based on the vibration frequency statistical feature vector using the following formula to obtain the vibration feature matrix;

[0066] The formula is as follows:

[0067]

[0068] Where M is the vibration feature matrix, V1 is the vibration frequency statistical feature vector, and V2 is the vibration waveform feature vector. That is, by multiplying the transpose of the vibration frequency statistical feature vector with the vibration waveform feature vector, the multi-scale implicit feature information of the frequency domain statistical feature values ​​of the vibration detection signal is mapped to the high-dimensional space where the vibration waveform feature vector resides. This corrects the time-domain implicit features of the vibration detection signal waveform to obtain the vibration feature matrix, which includes both the frequency domain implicit features and the time-domain implicit features.

[0069] In step S140 of this embodiment, the plurality of audio frequency domain statistical feature values ​​and the waveform of the sound detection signal are processed by a second Clip model including a sequence encoder and an image encoder to obtain a sound feature matrix. Similarly, for the sound feature extraction of the sound detection signal, considering that the periodic feature information of the sound detection signal and the periodic feature of the vibration detection signal have similar regularity, the Clip model is also used for sound signal encoding in the technical solution of this application. That is, specifically, the plurality of audio frequency domain statistical feature values ​​and the waveform of the sound detection signal are processed by a second Clip model including a sequence encoder and an image encoder to obtain a sound feature matrix, and then the time domain implicit features of the sound detection signal waveform are optimized by image attribute encoding based on the multi-scale implicit features of the frequency domain statistical feature values ​​of the sound detection signal to obtain the sound feature matrix.

[0070] Furthermore, by fusing the feature information from the vibration feature matrix and the sound feature matrix, and then performing classification processing through a classifier, a classification result indicating whether a large block has been discovered can be obtained. For example, the vibration feature matrix and the sound feature matrix can be fused bit by bit.

[0071] Specifically, in the technical solution of this application, since both the vibration feature matrix and the sound feature matrix are obtained by fusing the parametric semantics of frequency domain statistical eigenvalues ​​and the image semantics of the corresponding waveforms using the CLIP model, the eigenvalue at each position of the vibration feature matrix and the sound feature matrix represents the bitwise correlation value between the parametric semantics of the frequency domain statistical eigenvalues ​​and the image semantics of the corresponding waveforms. Consequently, in the vibration feature matrix and the sound feature matrix, there are cases where the correlation directions represented by corresponding positions are opposite, which leads to a negative correlation between corresponding positions in the vibration feature matrix and the sound feature matrix, thus affecting the bitwise fusion effect of the vibration feature matrix and the sound feature matrix. Therefore, the applicant of this application uses a fully orthographic projection nonlinear reweighting method to fuse the vibration feature matrix and the sound feature matrix.

[0072] In step S150 of this embodiment, the vibration feature matrix and the sound feature matrix are fused to obtain a classification feature matrix.

[0073] In a specific embodiment of this application, fusing the vibration feature matrix and the sound feature matrix to obtain a classification feature matrix includes: fusing the vibration feature matrix and the sound feature matrix using the following formula to obtain the classification feature matrix;

[0074] The formula is as follows:

[0075]

[0076] Among them, M1, M2 and M c These are the vibration feature matrix, the sound feature matrix, and the classification feature matrix, respectively, where ReLU(·) represents the ReLU activation function. This indicates matrix multiplication, where the division between the numerator and denominator matrices is a positional division of the matrix eigenvalues. exp(·) represents matrix exponentiation, which means calculating the natural exponential function value raised to the power of the eigenvalues ​​at each position in the matrix.

[0077] Here, the orthographic projection nonlinear reweighting uses the ReLU function to ensure the projection is entirely positive to avoid aggregating negatively correlated information. Simultaneously, a nonlinear reweighting mechanism is introduced to aggregate the eigenvalue distribution of the feature matrix, allowing the intrinsic structure of the feature matrix to penalize long-distance connections and strengthen local coupling. This achieves a synergistic effect between the vibration feature matrix and the sound feature matrix in the high-dimensional feature space corresponding to the orthographic projection reweighting spatial feature transform, improving the fusion effect of the vibration feature matrix and the sound feature matrix. This enables real-time and accurate detection of large sections being excavated by the excavator during the excavation process, thus ensuring the smooth progress of the excavation work.

[0078] In step S160 of this embodiment, the classification feature matrix is ​​processed by a classifier to obtain a classification result, which indicates whether a large block has been excavated. Specifically, the classification feature matrix, which includes the time-frequency domain latent features of the vibration detection signal and the time-frequency domain latent features of the sound detection signal, is processed by a classifier to obtain a classification result indicating whether a large block has been excavated. This allows for real-time and accurate detection of large blocks, ensuring the smooth progress of the excavation work.

[0079] In a specific embodiment of this application, the step of passing the classification feature matrix through a classifier to obtain a classification result includes: projecting the classification feature matrix into a classification feature vector; using the fully connected layer of the classifier to perform fully connected encoding on the classification feature vector to obtain an encoded classification feature vector; and passing the encoded classification feature vector through the Softmax classification function of the classifier to obtain the classification result.

[0080] Specifically, the classification feature matrix is ​​first projected into a one-dimensional classification feature vector. Then, considering that each position in the classification feature matrix contains rich feature information, the fully connected layer of the classifier is used to fully encode the classification feature vector to obtain an encoded classification feature vector. Next, the Softmax function value of the encoded classification feature vector is calculated, that is, the probability value of the encoded classification feature vector belonging to each classification label. In this embodiment, the classification labels include "large block found" (first label) and "no large block found" (second label). Finally, the label corresponding to the larger probability value is taken as the classification result.

[0081] In summary, the multi-sensor fusion-based method for detecting large excavated blocks according to the embodiments of this application collects vibration and sound detection signals from the excavator during the excavation process using vibration and sound sensors. Fourier transforms are then performed on each signal to extract frequency domain features. Next, a Clip model containing a sequence encoder and an image encoder is used to extract the fused feature information of the time-domain and frequency-domain features of the vibration and sound signals during the excavation process. Preferably, considering that the vibration and sound feature matrices obtained through the Clip model may have opposite correlation directions at corresponding positions, a fully orthographic projection nonlinear reweighting method is used to fuse the vibration and sound feature matrices. This approach enables real-time and accurate detection of large excavated blocks, thereby ensuring the smooth progress of the excavation work.

[0082] Exemplary System

[0083] Figure 6 The diagram illustrates a block diagram of a large-block detection system based on multi-sensor fusion according to an embodiment of this application. Figure 6 As shown, the multi-sensor fusion-based large-block excavation detection system 100 according to an embodiment of this application includes: a signal acquisition module 110, used to acquire vibration detection signals and sound detection signals collected by vibration sensors and sound sensors during the excavation process; a frequency domain feature extraction module 120, used to perform Fourier transform on the vibration detection signals and the sound detection signals respectively to obtain multiple vibration frequency domain statistical feature values ​​and multiple sound frequency domain statistical feature values; a vibration signal encoding module 130, used to pass the waveforms of the multiple vibration frequency domain statistical feature values ​​and the vibration detection signals through a first Clip model including a sequence encoder and an image encoder to obtain a vibration feature matrix; a sound signal encoding module 140, used to pass the waveforms of the multiple sound frequency domain statistical feature values ​​and the sound detection signals through a second Clip model including a sequence encoder and an image encoder to obtain a sound feature matrix; a fusion module 150, used to fuse the vibration feature matrix and the sound feature matrix to obtain a classification feature matrix; and a classification module 160, used to pass the classification feature matrix through a classifier to obtain a classification result, the classification result being used to indicate whether a large block has been excavated.

[0084] Figure 7 The diagram illustrates a block diagram of a vibration signal encoding module in a large-block excavation detection system based on multi-sensor fusion according to an embodiment of this application. Figure 7As shown, in a specific embodiment of this application, the vibration signal encoding module 130 includes: a vibration frequency encoding unit 131, used to input the plurality of vibration frequency domain statistical feature values ​​into the sequence encoder of the first Clip model to obtain a vibration frequency statistical feature vector; a vibration waveform encoding unit 132, used to input the waveform of the vibration detection signal into the image encoder of the first Clip model to obtain a vibration waveform feature vector; and an optimization unit 133, used to perform image attribute encoding optimization on the vibration waveform feature vector based on the vibration frequency statistical feature vector to obtain the vibration feature matrix.

[0085] Figure 8 The diagram illustrates a block diagram of a vibration frequency encoding unit in a large-block excavation detection system based on multi-sensor fusion according to an embodiment of this application. Figure 8 As shown, in a specific embodiment of this application, the vibration frequency encoding unit 131 includes: an arrangement subunit 1311, used to arrange the plurality of vibration frequency domain statistical feature values ​​into a vibration frequency domain statistical input vector; a first scale encoding subunit 1312, used to perform one-dimensional convolutional encoding of the vibration frequency domain statistical input vector at a first scale using the first convolutional layer of the sequence encoder of the first Clip model to obtain a first-scale vibration frequency domain statistical feature vector; a second scale encoding subunit 1313, used to perform one-dimensional convolutional encoding of the vibration frequency domain statistical input vector at a second scale using the second convolutional layer of the sequence encoder of the first Clip model to obtain a second-scale vibration frequency domain statistical feature vector; and a concatenation subunit 1314, used to concatenate the first-scale vibration frequency domain statistical feature vector and the second-scale vibration frequency domain statistical feature vector to obtain the vibration frequency statistical feature vector.

[0086] In a specific embodiment of this application, the vibration waveform encoding unit is further configured to: use each layer of the image encoder of the first Clip model to perform convolution processing, pooling processing and nonlinear activation processing on the input data in the forward propagation of the layer so that the vibration waveform feature vector is output by the last layer of the image encoder of the first Clip model.

[0087] In a specific embodiment of this application, the optimization unit is further configured to: perform image attribute encoding optimization on the vibration waveform feature vector based on the vibration frequency statistical feature vector using the following formula to obtain the vibration feature matrix;

[0088] The formula is as follows:

[0089]

[0090] Where M is the vibration feature matrix, V1 is the vibration frequency statistical feature vector, and V2 is the vibration waveform feature vector.

[0091] In a specific embodiment of this application, the fusion module is further configured to: fuse the vibration feature matrix and the sound feature matrix using the following formula to obtain the classification feature matrix;

[0092] The formula is as follows:

[0093]

[0094] Among them, M1, M2 and M c These are the vibration feature matrix, the sound feature matrix, and the classification feature matrix, respectively, where ReLU(·) represents the ReLU activation function. This indicates matrix multiplication, where the division between the numerator and denominator matrices is a positional division of the matrix eigenvalues. exp(·) represents matrix exponentiation, which means calculating the natural exponential function value raised to the power of the eigenvalues ​​at each position in the matrix.

[0095] In a specific embodiment of this application, the classification module includes: a projection unit for projecting the classification feature matrix into a classification feature vector; a fully connected encoding unit for performing fully connected encoding on the classification feature vector using the fully connected layer of the classifier to obtain an encoded classification feature vector; and a classification unit for passing the encoded classification feature vector through the Softmax classification function of the classifier to obtain the classification result.

[0096] Here, those skilled in the art will understand that the specific functions and operations of each unit and module in the above-described multi-sensor fusion-based large-block detection system have been referenced above. Figures 1 to 5 The method for detecting large blocks based on multi-sensor fusion has been described in detail in the previous section, and therefore, its repeated description will be omitted.

[0097] As described above, the multi-sensor fusion-based large-block excavation detection system 100 according to the embodiments of this application can be implemented in various terminal devices, such as servers for multi-sensor fusion-based large-block excavation detection systems. In one example, the multi-sensor fusion-based large-block excavation detection system 100 can be integrated into the terminal device as a software module and / or a hardware module. For example, the multi-sensor fusion-based large-block excavation detection system 100 can be a software module in the operating system of the terminal device, or it can be an application developed for the terminal device; of course, the multi-sensor fusion-based large-block excavation detection system 100 can also be one of many hardware modules of the terminal device.

[0098] Alternatively, in another example, the multi-sensor fusion-based excavation large block detection system 100 and the terminal device can also be separate devices, and the multi-sensor fusion-based excavation large block detection system 100 can be connected to the terminal device via wired and / or wireless networks, and transmit interactive information in accordance with an agreed data format.

[0099] Exemplary electronic devices

[0100] Below, for reference Figure 9 This describes an electronic device according to embodiments of the present application.

[0101] Figure 9 A block diagram of an electronic device according to an embodiment of this application is illustrated.

[0102] like Figure 9 As shown, the electronic device 10 includes one or more processors 11 and memory 12.

[0103] The processor 11 may be a central processing unit (CPU) or other form of processing unit with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 10 to perform desired functions.

[0104] The memory 12 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 11 may execute the program instructions to implement the multi-sensor fusion-based large-block detection and / or other desired functions described in the various embodiments of this application above. The computer-readable storage medium may also store various contents such as vibration detection signals and sound detection signals collected by vibration sensors and sound sensors.

[0105] In one example, the electronic device 10 may also include an input device 13 and an output device 14, which are interconnected via a bus system and / or other forms of connection mechanism (not shown).

[0106] The input device 13 may include, for example, a keyboard, a mouse, etc.

[0107] The output device 14 can output various information to the outside, including classification results. The output device 14 may include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.

[0108] Exemplary computer program products and computer-readable storage media

[0109] In addition to the methods and devices described above, embodiments of this application may also be computer program products, which include computer program instructions that, when executed by a processor, cause the processor to perform the steps of the multi-sensor fusion-based large block detection method according to various embodiments of this application as described in the "Exemplary Methods" section above.

[0110] The computer program product can be written in any combination of one or more programming languages ​​to perform the operations of the embodiments of this application. The programming languages ​​include object-oriented programming languages ​​such as Java and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0111] Furthermore, embodiments of this application may also be computer-readable storage media storing computer program instructions thereon, which, when executed by a processor, cause the processor to perform the steps of the multi-sensor fusion-based large block detection method according to various embodiments of this application described in the "Exemplary Methods" section above.

[0112] The computer-readable storage medium may be any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may, for example, include, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0113] The basic principles of this application have been described above with reference to specific embodiments. However, it should be noted that the advantages, benefits, and effects mentioned in this application are merely examples and not limitations, and should not be considered as essential features of each embodiment of this application. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the application to the necessity of employing the aforementioned specific details for implementation.

[0114] The block diagrams of devices, apparatuses, devices, and systems involved in this application are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, devices, and systems can be connected, arranged, and configured in any manner. Words such as “comprising,” “including,” “having,” etc., are open-ended terms meaning “including but not limited to,” and are used interchangeably with them. The terms “or” and “and” as used herein refer to the terms “and / or,” and are used interchangeably with them unless the context clearly indicates otherwise. The term “such as” as used herein refers to the phrase “such as but not limited to,” and is used interchangeably with it.

[0115] It should also be noted that in the apparatus, equipment, and methods of this application, the components or steps can be disassembled and / or recombined. These disassemblies and / or recombinations should be considered as equivalent solutions of this application.

[0116] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this application. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of this application. Therefore, this application is not intended to be limited to the aspects shown herein, but rather to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0117] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of this application to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations thereof.

Claims

1. A multi-sensor fusion-based large block detection method, characterized in that, The method comprises: acquiring a vibration detection signal and a sound detection signal collected by a vibration sensor and a sound sensor during the excavation process; performing Fourier transform on the vibration detection signal and the sound detection signal respectively to obtain a plurality of vibration frequency domain statistical characteristic values and a plurality of sound frequency domain statistical characteristic values; inputting the plurality of vibration frequency domain statistical characteristic values into a sequence encoder of a first Clip model to obtain a vibration frequency statistical characteristic vector; inputting the waveform graph of the vibration detection signal into an image encoder of the first Clip model to obtain a vibration waveform characteristic vector; and performing image attribute coding optimization on the vibration waveform characteristic vector based on the vibration frequency statistical characteristic vector to obtain the vibration feature matrix. The method comprises: arranging the plurality of vibration frequency domain statistical characteristic values into a vibration frequency domain statistical input vector; using a first convolution layer of the sequence encoder of the first Clip model to perform one-dimensional convolution coding of a first scale on the vibration frequency domain statistical input vector to obtain a first scale vibration frequency domain statistical characteristic vector; using a second convolution layer of the sequence encoder of the first Clip model to perform one-dimensional convolution coding of a second scale on the vibration frequency domain statistical input vector to obtain a second scale vibration frequency domain statistical characteristic vector; and concatenating the first scale vibration frequency domain statistical characteristic vector and the second scale vibration frequency domain statistical characteristic vector to obtain the vibration frequency statistical characteristic vector. wherein M1, M2and M c are the vibration feature matrix, the sound feature matrix and the classification feature matrix respectively, ReLU(·) represents a ReLU activation function, denotes matrix multiplication, and the division between the numerator matrix and the denominator matrix is the positional division of the eigenvalues of the matrix, exp(·) represents the exponential operation of the matrix, and the exponential operation of the matrix represents the calculation of the natural exponential function value with the eigenvalue of each position in the matrix as the power.

2. The multi-sensor fusion based detection method of the boulder according to claim 1, wherein, The method comprises: using each layer of the image encoder of the first Clip model to perform convolution processing, pooling processing, and nonlinear activation processing on the input data in the forward transmission of the layer respectively to output the vibration waveform characteristic vector from the last layer of the image encoder of the first Clip model. The method comprises: using each layer of the image encoder of the first Clip model to perform convolution processing, pooling processing, and nonlinear activation processing on the input data in the forward transmission of the layer respectively to output the vibration waveform characteristic vector from the last layer of the image encoder of the first Clip model. 3.The multi-sensor fusion based large block detection method of claim 2, wherein, The method comprises: using the first Clip model to fuse the vibration feature matrix and the sound feature matrix to obtain a classification feature matrix; and using a classifier to obtain a classification result from the classification feature matrix, wherein the classification result is used to indicate whether a large block is excavated. The method comprises: 4.The multi-sensor fusion based shovel bulk detection method of claim 3, wherein, using the first Clip model to fuse the vibration feature matrix and the sound feature matrix to obtain a classification feature matrix; and using a classifier to obtain a classification result from the classification feature matrix, wherein the classification result is used to indicate whether a large block is excavated. The method comprises: using the first Clip model to fuse the vibration feature matrix and the sound feature matrix to obtain a classification feature matrix; and using a classifier to obtain a classification result from the classification feature matrix, wherein the classification result is used to indicate whether a large block is excavated. The method comprises: using the first Clip model to fuse the vibration feature matrix and the sound feature matrix to obtain a classification feature matrix; and using a classifier to obtain a classification result from the classification feature matrix, wherein the classification result is used to indicate whether a large block is excavated. The method comprises: using the first Clip model to fuse the vibration feature matrix and the sound feature matrix to obtain a classification feature matrix; and using a classifier to obtain a classification result from the classification feature matrix, wherein the classification result is used to indicate whether a large block is excavated. The method comprises: using the first Clip model to fuse the vibration feature matrix and the sound feature matrix to obtain a classification feature matrix; and using a classifier to obtain a classification result from the classification feature matrix, wherein the classification result is used to indicate whether a large block is excavated. The method comprises: using the first Clip model to fuse the vibration feature matrix and the sound feature matrix to obtain a classification feature matrix; and using a classifier to obtain a classification result from the classification feature matrix, wherein the classification result is used to indicate whether a large block is excavated. The method comprises: using the first Clip model to fuse the vibration feature matrix and the sound feature matrix to obtain a classification feature matrix; and using a classifier to obtain a classification result from the classification feature matrix, wherein the classification result is used to indicate whether a large block is excavated. The method comprises: using the first Clip model to fuse the vibration feature matrix and the sound feature matrix to obtain a classification feature matrix; and using a classifier to obtain a classification result from the classification feature matrix, wherein the classification result is used to indicate whether a large block is excavated. The method comprises: using the first Clip model to fuse the vibration feature matrix and the sound feature matrix to obtain a classification feature matrix; and using a classifier to obtain a classification result from the classification feature matrix, wherein the classification result is used to indicate whether a large block is excavated. The method comprises: using the first Clip model to fuse the vibration feature matrix and the sound feature matrix to obtain a classification feature matrix; and using a classifier to obtain a classification result from the classification feature matrix, wherein the classification result is used to indicate whether a large block is excavated. The method comprises: using the first Clip model to fuse the vibration feature matrix and the sound feature matrix to obtain a classification feature matrix; and using a classifier to obtain a classification result from the classification feature matrix, wherein the classification result is used to indicate whether a large block is excavated. The method comprises: using the first Clip model to fuse the vibration feature matrix and the sound feature matrix to obtain a classification feature matrix; and using a classifier to obtain a classification result from the classification feature matrix, wherein the classification result is used to indicate whether a large block is excavated. The method comprises: using the first Clip model to fuse the vibration feature matrix and the sound feature matrix to obtain a classification feature matrix; and using a classifier to obtain a classification result from the classification feature matrix, wherein the classification result is used to indicate whether a large block is excavated. The method comprises: using the first Clip model to fuse the vibration feature matrix and the sound feature matrix to obtain a classification feature matrix; and using a classifier to obtain a classification result from the classification feature matrix, wherein the classification result is used to indicate whether a large block is excavated. The method comprises: using the first Clip model to fuse the vibration feature matrix and the sound feature matrix to obtain a classification feature matrix; and using a classifier to obtain a classification result from the classification feature matrix, wherein the classification result is used to indicate whether a large block is excavated. The method comprises: using the first Clip model to fuse the vibration feature matrix and the sound feature matrix to obtain a classification feature matrix; and using a classifier to obtain a classification result from the classification feature matrix, wherein the classification result is used to indicate whether a large block is excavated. The method comprises: using the first Clip model to fuse the vibration feature matrix and the sound feature matrix to obtain a classification feature matrix; and using a classifier to obtain a classification result from the classification feature matrix, wherein the classification result is used to indicate whether a large block is excavated. The method comprises: using the first Clip model to fuse the vibration feature matrix and the sound feature matrix to obtain a classification feature matrix; and using a classifier to obtain a classification result from the classification feature matrix, wherein the classification result is used to indicate whether a large block is excavated. The method comprises: using the first Clip model to fuse the vibration feature matrix and the sound feature matrix to obtain a classification feature matrix; and using a classifier to obtain a classification result from the classification feature matrix, wherein the classification result is used to indicate whether a large block is excavated. The method comprises: using the first Clip model to fuse the vibration feature matrix and the sound feature matrix to obtain a classification feature matrix; and using a classifier to obtain a classification result from the classification feature matrix, wherein the classification result is used to indicate whether a large block is excavated. The method comprises: using the first Clip model to fuse the vibration feature matrix and the sound feature matrix to obtain a classification feature matrix; and using a classifier to obtain a classification result from the classification feature matrix, wherein the classification result is used to indicate whether a large block is excavated. The method comprises: using the first Clip model to fuse the vibration feature matrix and the sound feature matrix to obtain a classification feature matrix; and using a classifier to obtain a classification result from the classification feature matrix, wherein the classification result is used to indicate whether a large block is excavated. The method comprises: using the first Clip model to fuse the vibration feature matrix and the sound feature matrix to obtain a classification feature matrix; and using a classifier to obtain a classification result from the classification feature matrix, wherein the classification result is used to indicate whether a large block is excavated. The method comprises: using the first Clip model to fuse the vibration feature matrix and the sound feature matrix to obtain a classification feature matrix; and using a classifier to obtain a classification result from the classification feature matrix, wherein the classification result is used to indicate whether a large block is excavated. The method comprises: using the first Clip model to fuse the vibration feature matrix and the sound feature matrix to obtain a classification feature matrix; and using a classifier to obtain a classification result from the classification feature matrix, wherein the classification result is used to indicate whether a large block is excavated. The method comprises: using the first Clip model to fuse the vibration feature matrix and the sound feature matrix to obtain a classification feature matrix; and using a classifier to obtain a classification result from the classification feature matrix, wherein the classification result is used to indicate whether a large block is excavated. The method comprises: using the first Clip model to fuse the vibration feature matrix and the sound feature matrix to obtain a classification feature matrix; and using a classifier to obtain a classification result 5.The multi-sensor fusion based large block detection method of claim 4, wherein, The image attribute coding optimization is performed on the vibration waveform feature vector based on the vibration frequency statistical feature vector to obtain the vibration feature matrix, and the image attribute coding optimization includes: The image attribute coding optimization is performed on the vibration waveform feature vector based on the vibration frequency statistical feature vector to obtain the vibration feature matrix, and the image attribute coding optimization includes: The formula is: The formula is: 6.The multi-sensor fusion based large block detection method of claim 5, wherein, Wherein, M is the vibration feature matrix, V1 is the vibration frequency statistical feature vector, and V2 is the vibration waveform feature vector. The classification feature matrix is projected into a classification feature vector, and the classification feature vector is fully connected and coded using a full connection layer of the classifier to obtain a coded classification feature vector. The coded classification feature vector is input into a Softmax classification function of the classifier to obtain the classification result. The signal acquisition module is configured to acquire vibration detection signals and sound detection signals collected by vibration sensors and sound sensors during the mining process.

7. A multi-sensor fusion based boulder detection system for excavators, characterized in that, The frequency domain feature extraction module is configured to perform Fourier transform on the vibration detection signals and the sound detection signals respectively to obtain a plurality of vibration frequency domain statistical feature values and a plurality of sound frequency domain statistical feature values. The vibration signal coding module is configured to input the plurality of vibration frequency domain statistical feature values and a waveform graph of the vibration detection signals into a first Clip model including a sequence encoder and an image encoder to obtain a vibration feature matrix. The sound signal coding module is configured to input the plurality of sound frequency domain statistical feature values and a waveform graph of the sound detection signals into a second Clip model including a sequence encoder and an image encoder to obtain a sound feature matrix. The fusion module is configured to fuse the vibration feature matrix and the sound feature matrix to obtain a classification feature matrix. The classification module is configured to input the classification feature matrix into a classifier to obtain a classification result, and the classification result is used to indicate whether a large block is mined. The fusion module is further configured to: The fusion module is further configured to: The vibration signal coding module includes: The vibration frequency coding unit is configured to input the plurality of vibration frequency domain statistical feature values into a sequence encoder of the first Clip model to obtain a vibration frequency statistical feature vector. The vibration waveform coding unit is configured to input the waveform graph of the vibration detection signals into an image encoder of the first Clip model to obtain a vibration waveform feature vector. The optimization unit is configured to perform image attribute coding optimization on the vibration waveform feature vector based on the vibration frequency statistical feature vector to obtain the vibration feature matrix. wherein M1, M2and M c are the vibration feature matrix, the sound feature matrix and the classification feature matrix respectively, ReLU(·) represents a ReLU activation function, denotes matrix multiplication, and the division between the numerator matrix and the denominator matrix is the positional division of the eigenvalues of the matrix, exp(·) represents the exponential operation of the matrix, and the exponential operation of the matrix represents the calculation of the natural exponential function value with the eigenvalue of each position in the matrix as the power.

8. The multi-sensor fusion based detection system for detection of overburden material of claim 7, wherein, ​ ​ ​ ​

Citation Information

Patent Citations

  • Video target behavior anomaly detection method and system based on multi-modal feature fusion

    CN114782882A

  • Intelligent fault diagnosis system of servo motor and diagnosis method thereof

    CN115235612A

  • Soil property estimation method, trained model generation method, soil property estimation device, and program

    JP2022108568A