Underground engineering construction machinery activity cross-modal depth identification system, method and equipment

By adopting a cross-modal depth recognition system in the underground engineering construction environment, combining single-modal feature extraction, cross-modal attention mechanism and multi-head self-attention mechanism, the problem of insufficient accuracy in modeling of single-modal data is solved, and a higher accuracy of construction machinery activity recognition is achieved.

CN120216866APending Publication Date: 2025-06-27TIANJIN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510224144.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

When the prior art recognizes construction machinery activities in underground engineering construction environments, the accuracy and robustness of single-modal data modeling are insufficient, and the complex interaction relationships between multimodal data are not fully utilized, resulting in low recognition accuracy.

Method used

A cross-modal depth recognition system is proposed. Through the combination of data acquisition module, data preprocessing module and construction machinery activity state recognition model, the single-modal feature extraction module, the cross-modal attention mechanism module and the multi-head self-attention mechanism module are used to capture and fuse multi-modal features in video, audio and kinematic data.

Benefits of technology

The accuracy of the recognition of construction machinery activities is significantly improved, and the long-term dependence relationship between different modes is established through the multi-head self-attention mechanism, and the feature information of multi-modal data is fully explored, which improves the model's ability to identify key features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216866A_ABST
    Figure CN120216866A_ABST
Patent Text Reader

Abstract

The invention discloses an underground engineering construction machinery activity cross-modal depth identification system. The system comprises a data acquisition module, a data preprocessing module and a construction machinery activity state identification model which are connected in sequence. The construction machinery activity state recognition model comprises a single-mode feature extraction module, a cross-mode attention mechanism module and a multi-head self-attention mechanism module; the data acquisition module acquires kinematics, audio and video data of the construction machinery; the data preprocessing module is used for preprocessing the collected data; the single-mode feature extraction module is used for extracting single-mode features of kinematics, audio and video data; the cross-modal attention mechanism module is used for capturing correlation among multi-modal features, and the multi-head self-attention mechanism module is used for fusing and classifying the multi-modal features, performing multi-modal fusion on the input features, and outputting a construction machinery activity state classification result through a Softmax function. The method can be applied to the fields of multi-modal data processing, feature recognition and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of construction machinery activity recognition in water conservancy and hydropower projects, and particularly relates to a cross-modal depth recognition system, method and device for construction machinery activities in underground projects. Background Technique

[0002] At present, the effective recognition of the activity status of construction machinery can provide strong support for the analysis of mechanical production efficiency and the judgment of safety status. In the underground cavern group, the lighting conditions are poor, the dust concentration is high, and the mechanical sounds in the enclosed space are mixed and noisy. The mechanical activity recognition method based on single-modal data such as video and kinematics has insufficient recognition robustness due to the lack of supplementation and corroboration of other modal data. Multi-modal data contains richer modal information and can reflect the characteristics of the same activity of construction machinery from different perspectives. Introducing it into the mechanical activity recognition task helps to improve the recognition accuracy. However, most current studies conduct activity recognition based on single-modal data. Even in a few studies that consider the information of two modalities, they only fuse different modal data at the data level or the decision level, ignoring the rich associations between modalities of multi-modal data, resulting in waste of data resources and low recognition accuracy. Therefore, there is an urgent need to propose a method for recognizing construction machinery activities in underground projects that can consider the complex interaction relationships between multi-modal data.

[0003] Sensor-based non-visual technology recognition and vision-based technology recognition are the two major mainstream methods for recognizing the activity status of construction machinery. Sensor-based non-visual technology uses inertial measurement units (IMUs), global positioning systems (GPS), or wireless radio frequency devices such as wireless local area networks (WLANs), ultra-wideband (UWB), and radio frequency identification (RFID) to locate equipment components or their relative postures. However, these methods have their own limitations in the underground construction environment. For example, in tunnel construction, due to the problem of its own error accumulation, the inertial measurement unit (IMU) is difficult to be applied for a long time without regular calibration; GPS and wireless radio frequency devices are difficult to locate equipment components or their relative postures due to the lack of signals; poor lighting conditions and smoke and dust caused by blasting also make it difficult for vision-based methods to capture effective images.

[0004] In recent years, audio signals have also become a solution for analyzing construction activities because mechanical equipment generates specific sound patterns when performing different tasks. However, in underground projects, multiple construction machinery operate simultaneously in a closed environment, making the acoustic signals mixed and greatly interfered by noise, resulting in poor performance of the acoustic-based activity recognition method in the underground working environment. Therefore, single-modal data modeling is prone to insufficient accuracy and robustness of the model in a relatively complex construction environment. Summary of the Invention

[0005] The present invention provides an underground engineering construction machinery active cross-modal depth recognition system, method and device to solve the technical problems existing in the known technology.

[0006] The technical solution adopted by the present invention to solve the technical problems existing in the known technology is as follows:

[0007] An underground engineering construction machinery active cross-modal depth recognition system, which includes a data acquisition module, a data preprocessing module and a construction machinery activity state recognition model connected in sequence; the construction machinery activity state recognition model includes a single-modal feature extraction module, a cross-modal attention mechanism module and a multi-head self-attention mechanism module; the data acquisition module is used to collect video, audio and kinematic data of the construction machinery reflecting the activity state of the construction machinery; the data preprocessing module is used to preprocess the data collected by the data acquisition module to extract effective information; the single-modal feature extraction module is used to extract the single-modal original features of the video, audio and kinematic data; the cross-modal attention mechanism module is used to capture the correlation between various single-modal original features, its input is the single-modal original features extracted by the single-modal feature extraction module, and its output is the cross-modal interaction features used to simulate the interaction between different modalities; the multi-head self-attention mechanism module is used for multi-modal feature fusion and classification, its input is the single-modal original features extracted by the single-modal feature extraction module and the cross-modal interaction features output by the cross-modal attention mechanism module, it performs multi-modal fusion on the input features, and integrates the interaction information inside each modality feature and between various modality features through a fully connected layer network, and it outputs the classification result of the construction machinery activity state through an activation function.

[0008] Furthermore, the data preprocessing module includes an audio preprocessing sub-module, a video preprocessing sub-module and a kinematic preprocessing sub-module; the audio preprocessing sub-module performs the following preprocessing on the collected audio data: converting the audio data into time-frequency data through short-time Fourier transform, using a Mel filter bank to convert the frequency in the time-frequency data into Mel-scale frequency, and further performing discrete cosine transform on the logarithmic energy of the Mel spectrum of each frame to remove the correlation between the energies of different filters and extract Mel-frequency cepstral coefficients; arranging the Mel spectra of each frame in chronological order to obtain a complete Mel spectrogram;

[0009] The video preprocessing sub-module randomly flips, randomly crops and adjusts the color brightness of the extracted frames in the video clip to simulate different perspectives and distances;

[0010] The kinematic preprocessing sub-module is provided with a first-order low-pass filter, the first-order low-pass filter calculates the gravity acceleration component, and the kinematic preprocessing sub-module subtracts the gravity acceleration component from the raw data collected by the data acquisition module.

[0011] Furthermore, the video preprocessing sub-module includes an OpenCV unit, which uniformly samples 20 to 100 frames of pictures from each collected video clip.

[0012] Furthermore, the unimodal feature extraction module includes a video extraction sub-module, an audio extraction sub-module, and a kinematics extraction sub-module; the video extraction sub-module uses S3D to extract the original features of video data; the audio extraction sub-module uses VGGish to extract the original features of audio data, and the kinematics extraction sub-module uses a Conformer neural network to extract the original features of kinematics data.

[0013] Furthermore, the cross-modal attention mechanism module includes first to third cross-modal attention mechanism sub-modules; each of the first to third cross-modal attention mechanism sub-modules includes a one-dimensional temporal convolutional layer, a position embedding unit, and multiple cross-modal attention units. The one-dimensional temporal convolutional layer and the position embedding unit are connected in sequence, and the output end of the position embedding unit is connected in parallel to multiple cross-modal attention units;

[0014] The one-dimensional temporal convolutional layers of the first to third cross-modal attention mechanism sub-modules respectively correspond to the unimodal original features of video data, audio data, and kinematics data, map the feature dimensions of different modalities to a fixed-dimensional space, and output features

[0015] The position embedding unit is to make the feature sequence output by the one-dimensional temporal convolutional layer carry temporal information. According to the following formula: Add position embedding:

[0016]

[0017] In the formula:

[0018] is the low-level position-aware feature corresponding to video, audio, and kinematics modality data;

[0019] is the output feature of the one-dimensional temporal convolutional layer;

[0020] d is the position embedding dimension;

[0021] d {P,V,A} is the feature dimension corresponding to video, audio, and kinematics modality data;

[0022] Y {P,V,A} is the number of time steps corresponding to video, audio, and kinematics modality data;

[0023] PE(T {P,V,A} ,d) is the position embedding, which is used to add temporal information to each time step in the feature sequence;

[0024] PE() represents the fixed embedding function for calculating the index of each position;

[0025] Each cross-modal attention unit of the first to third cross-modal attention mechanism sub-modules is calculated using a cross-modal algorithm, mapping the data of one modality to the data of another modality for processing, and finally outputting: six pairs of cross-modal interaction features of kinematics and audio, kinematics and video, audio and kinematics, audio and video, video and kinematics, and video and audio.

[0026] Further, each cross-modal attention unit is composed of multiple layers of cascaded cross-modal attention sub-units. Let one cross-modal attention unit be cross-modal attention unit U, and let j be the layer number of the cross-modal attention sub-units in cross-modal attention unit U; let j = 1, …, D; cross-modal attention unit U performs feed-forward calculations layer by layer according to j = 1, …, D;

[0027] Cross-modal attention unit U maps the kinematic modality data to video modality data according to the following formula to obtain the cross-modal data interaction feature between video and kinematics:

[0028]

[0029] In the formula:

[0030] is the initial input feature;

[0031] is the initial feature of the kinematic modality data;

[0032] is the initial feature of the video modality data;

[0033] is the intermediate feature of the j-th layer of cross-modal attention sub-units;

[0034] is the output feature of the (j - 1)-th layer of cross-modal attention sub-units;

[0035] LN() represents the layer normalization function;

[0036] represents the multi-head function of the j-th layer of cross-modal attention sub-units;

[0037] represents the position feed-forward function of the j-th layer of cross-modal attention sub-units with θ as the parameter;

[0038] V→P represents the transfer of visual data to kinematic data.

[0039] Furthermore, an self-attention mechanism module is also connected between the cross-modal attention mechanism module and the multi-head self-attention mechanism module. The self-attention mechanism module is used to obtain the correlation and contribution degree distribution within each pair of cross-modal interaction features. The self-attention mechanism module inputs each pair of cross-modal interaction features and outputs multi-modal mixed features reflecting the correlation and contribution degree distribution within each pair of cross-modal interaction features.

[0040] The present invention also provides a method for cross-modal depth recognition of underground engineering construction machinery activities using the above-mentioned underground engineering construction machinery activity cross-modal depth recognition system. The method includes the following steps:

[0041] Step 1, construct a recognition model for the activity state of construction machinery, and connect the single-modal feature extraction module, the cross-modal attention mechanism module, and the multi-head self-attention mechanism module in the recognition model in sequence; at the same time, connect the output end of the single-modal feature extraction module to the input end of the multi-head self-attention mechanism module;

[0042] Step 2, enable the data acquisition module to collect sample data of kinematics, audio, and video reflecting the activity state of construction machinery; enable the data preprocessing module to preprocess the sample data to extract effective information; compile the preprocessed sample data into a sample set for training the recognition model of the activity state of construction machinery;

[0043] Step 3, divide the sample set into a training set, a validation set, and a test set; use the training set to train the recognition model, determine the optimal parameters of the recognition model according to the best performance of the validation set; use the test set to evaluate the final model performance of the recognition model;

[0044] Step 4, enable the data acquisition module to collect real-time data of kinematics, audio, and video of the activity state of construction machinery at the construction site; enable the data preprocessing module to preprocess the real-time data; input the preprocessed real-time data into the trained recognition model of the activity state of construction machinery, and the recognition model outputs the classification result of the activity state of construction machinery in real time.

[0045] Furthermore, Step 3 includes the following method steps:

[0046] Divide 60% of the sample set into a training set, 20% into a validation set, and 20% into a test set; input the data into the recognition model of the activity state of construction machinery in batches, and each batch is used to update the weights. Use the Adam optimization algorithm to update the weights of the model and reduce the value of the loss function; control the step size of weight update by dynamically adjusting the learning rate, and at the same time add a regularization term to prevent the model from overfitting; determine the final optimal recognition model based on the best performance of the validation set.

[0047] The present invention also provides a device for the method of identifying the activity cross-modal depth of underground engineering construction machinery, including a memory and a processor, where the memory is used to store computer programs; the processor is used to execute the computer programs and, when executing the computer programs, implement the steps of the method of identifying the activity cross-modal depth of underground engineering construction machinery as described above.

[0048] The advantages and positive effects of the present invention are:

[0049] An activity cross-modal depth identification system and method for underground engineering construction machinery proposed by the present invention make full use of the advantage that the attention mechanism can capture complex and dynamic interdependencies between modalities, can more deeply and flexibly understand and process complex data relationships, and significantly improve the accuracy of construction machinery activity identification.

[0050] The present invention establishes long-term dependency relationships between different modalities, fully excavates the characteristic information of multi-modal data, analyzes the internal correlations existing between three single-modal features and three fused-modal features through the multi-head self-attention mechanism, calculates the attention weights they contribute to the final activity identification, and redistributes them according to the attention weights to obtain the multi-modal fusion feature Z, thereby enhancing the model's ability to identify key features.

[0051] The present invention is not only applicable to specific construction machinery activity identification tasks, but also can be widely applied to fields such as multi-modal data processing and feature recognition, and has strong generalizability. Description of the Drawings

[0052] Figure 1 It is a schematic structural diagram of an activity cross-modal depth identification system for underground engineering construction machinery of the present invention.

[0053] In the figure:

[0054] Acc represents the accuracy rate.

[0055] F1 represents the F1 score, which represents the harmonic mean of the precision rate and the recall rate.

[0056] X A represents the original feature of the audio single modality.

[0057] X V represents the original feature of the visual single modality.

[0058] X P represents the original feature of the kinematics single modality.

[0059] Z V is the video feature after cross-modal feature interaction processing.

[0060] Z A is the audio feature after cross-modal feature interaction processing.

[0061] Z P The kinematic features after cross-modal feature interaction processing.

[0062] Denotes the concatenation operation.

[0063] V→P means visual data is passed to kinematic data.

[0064] V→A means visual data is passed to audio data.

[0065] P→A means kinematic data is passed to audio data.

[0066] P→V means kinematic data is passed to visual data.

[0067] A→V means audio data is passed to visual data.

[0068] A→P means audio data is passed to kinematic data.

[0069] Z P→A Denotes the cross-modal interaction features of kinematics and audio.

[0070] Z P→V Denotes the cross-modal interaction features of kinematics and video.

[0071] Z A→P Denotes the cross-modal interaction features of audio and kinematics.

[0072] Z A→V Denotes the cross-modal interaction features of audio and video.

[0073] Z V→P Denotes the cross-modal interaction features of video and kinematics.

[0074] Z V→A Denotes the cross-modal interaction features of video and audio. Detailed implementation manner

[0075] The present invention will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.

[0076] The Chinese interpretations of the following English words, phrases and abbreviations are as follows:

[0077] Acc: Accuracy rate.

[0078] F1: F1 score, the harmonic mean of precision and recall.

[0079] OpenCV: Open Source Computer Vision Library, an open-source computer vision and machine learning software library.

[0080] Softmax: A mathematical function that converts a vector of raw scores into probabilities and is mostly used in classification tasks.

[0081] S3D: A 3D convolutional neural network for video understanding.

[0082] VGGish: A neural network for audio feature extraction, adjusted based on the VGG network architecture.

[0083] Conformer: A deep learning model that combines Transformer and CNN.

[0084] Please refer to Figure 1 , an underground engineering construction machinery activity cross-modal depth recognition system, which includes a data acquisition module, a data preprocessing module, and a construction machinery activity state recognition model connected in sequence; the construction machinery activity state recognition model includes a single-modal feature extraction module, a cross-modal attention mechanism module, and a multi-head self-attention mechanism module; the data acquisition module is used to collect video, audio, and kinematic data of the construction machinery reflecting the activity state of the construction machinery; the data preprocessing module is used to preprocess the data collected by the data acquisition module to extract effective information; the single-modal feature extraction module is used to extract single-modal raw features of the video, audio, and kinematic data; the cross-modal attention mechanism module is used to capture the correlation between various single-modal raw features, its input is the single-modal raw features extracted by the single-modal feature extraction module, and its output is the cross-modal interaction features used to simulate the interaction between different modalities; the multi-head self-attention mechanism module is used for multi-modal feature fusion and classification, its input is the single-modal raw features extracted by the single-modal feature extraction module and the cross-modal interaction features output by the cross-modal attention mechanism module, it performs multi-modal fusion on the input features, and integrates the interaction information inside each modality feature and between modality features through a fully connected layer network, and outputs the classification result of the construction machinery activity state through an activation function.

[0085] Preferably, the data preprocessing module may include an audio preprocessing sub-module, a video preprocessing sub-module, and a kinematic preprocessing sub-module; the audio preprocessing sub-module can perform the following preprocessing on the collected audio data: convert the audio data into time-frequency data through short-time Fourier transform, use a Mel filter bank to convert the frequencies in the time-frequency data into Mel-scale frequencies, and further perform discrete cosine transform on the logarithmic energy of the Mel spectrum of each frame to remove the correlation between the energies of different filters and extract Mel frequency cepstral coefficients; arrange the Mel spectrum of each frame in chronological order to obtain a complete Mel spectrogram.

[0086] The video preprocessing sub-module can perform random flipping, random cropping, and color and brightness adjustment on the extracted frames in the video clip to simulate different perspectives and distances.

[0087] The kinematic preprocessing sub-module can be equipped with a first-order low-pass filter. The first-order low-pass filter calculates the gravity acceleration component, and the kinematic preprocessing sub-module subtracts the gravity acceleration component from the raw data collected by the data acquisition module.

[0088] Preferably, the video preprocessing sub-module can include an OpenCV unit, and the OpenCV unit uniformly samples 20 to 100 frame pictures from each collected video clip.

[0089] Preferably, the unimodal feature extraction module can include a video extraction sub-module, an audio extraction sub-module, and a kinematic extraction sub-module; the video extraction sub-module can use S3D to extract the original features of video data; the audio extraction sub-module can use VGGish to extract the original features of audio data, and the kinematic extraction sub-module can use a Conformer neural network to extract the original features of kinematic data.

[0090] Preferably, the cross-modal attention mechanism module can include first to third cross-modal attention mechanism sub-modules; each of the first to third cross-modal attention mechanism sub-modules can include a one-dimensional temporal convolutional layer, a position embedding unit, and multiple cross-modal attention units. The one-dimensional temporal convolutional layer and the position embedding unit can be connected in sequence, and the output end of the position embedding unit is connected in parallel with multiple cross-modal attention units.

[0091] The one-dimensional temporal convolutional layers of the first to third cross-modal attention mechanism sub-modules can respectively correspond to the unimodal original features of the input video data, audio data, and kinematic data, map the feature dimensions of different modalities to a fixed-dimensional space, and output features

[0092] The position embedding unit is used to make the feature sequence output by the one-dimensional temporal convolutional layer carry temporal information, and can be as follows according to the formula Add position embedding:

[0093]

[0094] In the formula:

[0095] is the low-level position perception feature corresponding to the video, audio, and kinematic modality data;

[0096] is the output feature of the one-dimensional temporal convolutional layer;

[0097] d is the position embedding dimension;

[0098] d{P,V,A} is the feature dimension corresponding to video, audio, and kinematic modal data;

[0099] Y {P,V,A} is the number of time steps corresponding to video, audio, and kinematic modal data;

[0100] PE(T {P,V,A} , d) is the positional embedding, which is used to add temporal information to each time step in the feature sequence;

[0101] PE() represents a fixed embedding function for calculating each position index;

[0102] Each cross-modal attention unit of the first to third cross-modal attention mechanism sub-modules can be calculated using a cross-modal algorithm, mapping the data of one modality to the data of another modality for processing, and finally outputting: six pairs of cross-modal interaction features between kinematics and audio, kinematics and video, audio and kinematics, audio and video, video and kinematics, and video and audio.

[0103] Preferably, each cross-modal attention unit can be composed of multiple layers of cascaded cross-modal attention sub-units. One cross-modal attention unit can be set as cross-modal attention unit U, and let j be the layer number of the cross-modal attention sub-units in cross-modal attention unit U; let j = 1, …, D; cross-modal attention unit U performs feed-forward calculation layer by layer according to j = 1, …, D.

[0104] Cross-modal attention unit U can map kinematic modal data to video modal data according to the following formula to obtain the cross-modal data interaction feature between video and kinematics:

[0105]

[0106] In the formula:

[0107] is the initial input feature;

[0108] is the initial feature of kinematic modal data;

[0109] is the initial feature of video modal data;

[0110] is the intermediate feature of the j-th layer of cross-modal attention sub-units;

[0111] is the output feature of the (j - 1)-th layer of cross-modal attention sub-units;

[0112] LN() represents the layer normalization function;

[0113] Denote the multi - head function of the cross - modal attention sub - unit at the j - th layer;

[0114] Denote the position - feed - forward function of the cross - modal attention sub - unit at the j - th layer with parameter θ;

[0115] V→P represents the transfer of visual data to kinematic data.

[0116] Preferably, an self - attention mechanism module can also be connected between the cross - modal attention mechanism module and the multi - head self - attention mechanism module. The self - attention mechanism module can be used to obtain the correlation and contribution degree distribution within each pair of cross - modal interaction features; the self - attention mechanism module inputs each pair of cross - modal interaction features and outputs multi - modal mixed features reflecting the correlation and contribution degree distribution within each pair of cross - modal interaction features.

[0117] The present invention also provides a method for cross - modal depth recognition of construction machinery activities using the above - mentioned cross - modal depth recognition system for underground engineering construction machinery. The method includes the following steps:

[0118] Step 1, construct a recognition model for construction machinery activity states, and connect the single - modal feature extraction module, cross - modal attention mechanism module, and multi - head self - attention mechanism module in the recognition model in sequence; at the same time, connect the output end of the single - modal feature extraction module to the input end of the multi - head self - attention mechanism module.

[0119] Step 2, enable the data acquisition module to collect sample data of kinematics, audio, and video reflecting the construction machinery activity states; enable the data pre - processing module to pre - process the sample data to extract effective information; compile the pre - processed sample data into a sample set for training the recognition model of construction machinery activity states.

[0120] Step 3, divide the sample set into a training set, a validation set, and a test set; use the training set to train the recognition model, determine the optimal parameters of the recognition model according to the best performance on the validation set; use the test set to evaluate the final model performance of the recognition model.

[0121] Step 4, enable the data acquisition module to collect real - time data of kinematics, audio, and video of construction machinery activity states at the construction site; enable the data pre - processing module to pre - process the real - time data; input the pre - processed real - time data into the trained recognition model of construction machinery activity states, and the recognition model outputs the classification result of construction machinery activity states in real time.

[0122] Preferably, step 3 may include the following method steps:

[0123] Divide 60% of the sample set into the training set, 20% into the validation set, and 20% into the test set; the data is input into the construction machinery activity state recognition model in batches, and each batch is used to update the weights. Use the Adam optimization algorithm to update the weights of the model and reduce the value of the loss function; by dynamically adjusting the learning rate, control the step size of weight update, and at the same time add a regularization term to prevent the model from overfitting; based on the optimal performance of the validation set, determine the final optimal recognition model.

[0124] The present invention also provides an apparatus embodiment of the cross-modal depth recognition method for underground engineering construction machinery activities, including a memory and a processor. The memory is used to store a computer program; the processor is used to execute the computer program and implement the steps of the cross-modal depth recognition method for underground engineering construction machinery activities as described above when executing the computer program.

[0125] The following further illustrates the working process and working principle of the present invention with a preferred embodiment of the present invention:

[0126] A cross-modal depth recognition system for underground engineering construction machinery activities, the system includes a data acquisition module, a data preprocessing module and a construction machinery activity state recognition model connected in sequence; the construction machinery activity state recognition model includes a single-modal feature extraction module, a cross-modal attention mechanism module, a self-attention mechanism module and a multi-head self-attention mechanism module.

[0127] The data acquisition module is used to collect videos, audios and kinematic data of construction machinery reflecting the activity state of construction machinery.

[0128] The data preprocessing module is used to preprocess the data collected by the data acquisition module to extract effective information.

[0129] The single-modal feature extraction module is used to extract single-modal original features of videos, audios and kinematic data.

[0130] The cross-modal attention mechanism module is used to capture the correlation between various single-modal original features. Its input is the single-modal original features extracted by the single-modal feature extraction module, and its output is cross-modal interaction features used to simulate the interaction between different modalities.

[0131] The self-attention mechanism module is used to obtain the correlation and contribution degree distribution inside each pair of cross-modal interaction features; the self-attention mechanism module inputs each pair of cross-modal interaction features and outputs multi-modal mixed features reflecting the correlation and contribution degree distribution inside each pair of cross-modal interaction features.

[0132] The multi-head self-attention mechanism module is used for multi-modal feature fusion and classification. It takes the single-modal raw features extracted by the single-modal feature extraction module and the cross-modal interaction features output by the cross-modal attention mechanism module as inputs. It performs multi-modal fusion on the input features, integrates the interaction information within each modal feature and between different modal features through a fully-connected layer network, and outputs the classification result of the construction machinery activity state through the Softmax function.

[0133] A method for cross-modal depth recognition of construction machinery activities in underground engineering using the above-mentioned cross-modal depth recognition system for construction machinery activities in underground engineering, the method comprising the following steps:

[0134] Step 1, construct a recognition model for construction machinery activity state, and connect the single-modal feature extraction module, the cross-modal attention mechanism module, the self-attention mechanism module, and the multi-head self-attention mechanism module in the recognition model in sequence; at the same time, connect the output end of the single-modal feature extraction module to the input end of the multi-head self-attention mechanism module.

[0135] Step 2, enable the data acquisition module to collect sample data of kinematics, audio, and video reflecting the construction machinery activity state; enable the data preprocessing module to preprocess the sample data to extract effective information; compile the preprocessed sample data into a sample set for training the recognition model of construction machinery activity state.

[0136] Step 3, divide the sample set into a training set, a validation set, and a test set; use the training set to train the recognition model, determine the optimal parameters of the recognition model according to the best performance of the validation set; evaluate the final model performance of the recognition model using the test set.

[0137] Step 4, enable the data acquisition module to collect real-time data of kinematics, audio, and video of the construction machinery activity state at the construction site; enable the data preprocessing module to preprocess the real-time data; input the preprocessed real-time data into the trained recognition model of construction machinery activity state, and the recognition model outputs the classification result of the construction machinery activity state in real time.

[0138] In step 2, the method for enabling the data preprocessing module to preprocess the sample data to extract effective information includes:

[0139] The audio preprocessing sub-module preprocesses the collected audio data as follows:

[0140] First, the audio signal is subjected to short-time Fourier transform (STFT) to obtain the time-frequency representation of the audio, and its formula is:

[0141]

[0142] In the formula:

[0143] n is the sample serial number within the window;

[0144] N is the number of samples; it corresponds to the window length;

[0145] m is the serial number of the time frame;

[0146] H is the window step size; it corresponds to the time offset;

[0147] k is the frequency;

[0148] i is the imaginary unit;

[0149] S(m, k) is the result of the short-time Fourier transform at the m-th time and frequency k;

[0150] x(m + mH) is the original data;

[0151] W(n) is the window function.

[0152] Secondly, a Mel filter bank is used to process the spectral data, converting the frequency to the frequency on the Mel scale. The processed spectrogram is closer to the auditory perception of the human ear, making audio analysis and processing more effective perceptually. The commonly used Hertz (Hz)-Mel (mel) transformation formula is:

[0153]

[0154] In the formula:

[0155] e is the frequency corresponding to the Mel scale;

[0156] f is the frequency value of the original audio signal, in Hertz (Hz).

[0157] Finally, the logarithmic energy of the Mel spectrum for each frame is taken for discrete cosine transform (DCT) to remove the correlation between the energies of different filters and extract the Mel-frequency cepstral coefficients (MFCC). After the above steps of processing, the Mel spectrum representation of each frame will be obtained, and arranging the Mel spectra of each frame in chronological order will result in a complete Mel spectrogram.

[0158] The video preprocessing sub-module performs the following preprocessing on the collected video data:

[0159] Using OpenCV, 50 frame pictures are uniformly sampled from each video clip. The original input size of the video frame is 455×256; to improve the generalization ability of the model, the extracted frames are randomly flipped, randomly cropped, and color brightness adjusted to simulate different perspectives and distances, thereby increasing sample diversity.

[0160] The kinematics preprocessing sub-module performs the following preprocessing on the collected kinematics data:

[0161] Since the measured value of the acceleration sensor is affected by both dynamic and static (i.e., gravitational) accelerations simultaneously, it is necessary to eliminate the gravitational component from the signal. The gravitational acceleration component g (with an initial value of 0) is calculated using first-order low-pass filtering, and this component is subtracted from the original sensor data;

[0162] g(t) = (1 - α) × g(t - 1) + α × s(t);

[0163] a(t) = s(t) - g(t);

[0164]

[0165] Where:

[0166] g(t) is the gravitational value at time t;

[0167] g(t - 1) is the gravitational value at time t - 1;

[0168] s(t) is the original data collected at time t;

[0169] a(t) is the acceleration after removing gravity at time t;

[0170] α is the filtering coefficient, used to control the cut-off point of the filter;

[0171] T s is the sampling time interval of the original data;

[0172] f c is the cut-off frequency of the low-pass filter, with a value range between 0.1 and 0.5 Hz.

[0173] α is the filtering coefficient, used to control the cut-off point of the filter, and the calculation formula for α is

[0174] f c is the cut-off frequency of the low-pass filter, generally taking values between 0.1 and 0.5 Hz; since the original acceleration sampling frequency is 100 Hz, the calculated value of α is 0.03.

[0175] The single-modal feature extraction module includes a video extraction sub-module, an audio extraction sub-module, and a kinematics extraction sub-module; the video extraction sub-module uses S3D to extract the original features of video data; the audio extraction sub-module uses VGGish to extract the original features of audio data, and the kinematics extraction sub-module uses a Conformer neural network to extract the original features of kinematic data.

[0176] The method by which the video extraction sub-module uses S3D to extract the original features of video data includes the following method steps:

[0177] Use a pre-trained S3D (Separable 3D Convolution) model to obtain the unimodal feature vectors of video data. The input of the model is the normalized video frames. The S3D model first performs 2D spatial convolution and then 1D temporal convolution, which can reduce the computational complexity and the number of parameters of the model, while maintaining the ability to capture the spatial and temporal dimensional information in the video data.

[0178] The method for the audio extraction sub-module to use VGGish to extract the original features of audio data includes the following method steps:

[0179] Use a pre-trained VGGish model to obtain the unimodal feature vectors of audio data. The audio data is input into the VGGish model in the form of a mel spectrogram. The model processes the data through a multi-layer convolutional network and gradually extracts the high-level features of the audio data.

[0180] The method for the kinematics extraction sub-module to use a Conformer neural network to extract the original features of kinematic data includes the following method steps:

[0181] Use a Conformer model to extract kinematic modal features. The Conformer model learns the low-level local features of the entire one-dimensional temporal and spatial convolutional layers through a convolutional module, and the self-attention module extracts the global correlations in the local temporal features, which can effectively extract the features of the high-dimensional data (triaxial acceleration and triaxial angular velocity) that changes with time in kinematics.

[0182] The cross-modal attention mechanism module, the self-attention mechanism module, and the multi-head self-attention mechanism module are connected in sequence to capture the correlations between features, establish a multi-level feature fusion method, and obtain multi-modal fusion features.

[0183] The specific working method and working principle of the cross-modal attention mechanism module are as follows:

[0184] The cross-modal attention mechanism module includes the first to third cross-modal attention mechanism sub-modules; the first to third cross-modal attention mechanism sub-modules all include a one-dimensional temporal convolutional layer, a position embedding unit, and multiple cross-modal attention units. The one-dimensional temporal convolutional layer and the position embedding unit are connected in sequence, and the output end of the position embedding unit is connected in parallel to multiple cross-modal attention units.

[0185] Let the input data of the recognition model of the construction machinery activity state include kinematic data, video data, and audio data; let the kinematic data be P, the video data be V, and the audio data be A; {P, V, A} represents the data set containing kinematic data, video data, and audio data; X {P,V,A} represents the input data containing kinematic data, video data, and audio data; d {P,V,A} correspondingly represents the input data X{P,V,A} dimensions

[0186] Use to represent the input feature sequences and their dimensions from the three modalities of kinematics, video, and audio.

[0187] First, preprocess the input of each modality through a one-dimensional temporal convolutional layer to map the feature dimensions of different modalities to d {P,V,A} , providing stable and effective input features for subsequent cross-modal interactions:

[0188]

[0189] In the formula:

[0190] is the new feature after convolutional processing for X {P,V,A} ;

[0191] k {P,V,A} is the convolutional kernel size for the modalities {P, V, A};

[0192] Conv1D() represents the one-dimensional temporal convolution function;

[0193] To enable the sequence to carry temporal information, add positional embedding (PE) to :

[0194]

[0195] Among them, calculate the fixed embedding of each position index.

[0196] In the formula:

[0197] are the low-level position-aware features corresponding to the video, audio, and kinematic modality data;

[0198] are the output features of the one-dimensional temporal convolutional layer;

[0199] d is the positional embedding dimension;

[0200] d {P,V,A} is the feature dimension corresponding to the video, audio, and kinematic modality data;

[0201] Y {P,V,A} is the number of time steps corresponding to the video, audio, and kinematic modality data;

[0202] PE(T {P,V,A} , d) is the positional embedding, used to add temporal information to each time step in the feature sequence;

[0203] PE() represents a fixed embedding function that calculates the index of each position;

[0204] The cross-modal attention unit enables one modality to receive information from another modality, and fixes the dimensions of the first to third cross-modal attention mechanism sub-modules to d d ; Each cross-modal attention unit is composed of multiple cascaded cross-modal attention sub-units. Let one of the cross-modal attention units be the cross-modal attention unit U. Let the cross-modal attention unit U transfer visual information (V) to kinematic signals (P), denoted as V→P.

[0205] Let j be the layer number of the cross-modal attention sub-units in the cross-modal attention unit U; Let j = 1, …, D; The cross-modal attention unit U performs feed-forward calculations layer by layer according to j = 1, …, D.

[0206] The cross-modal attention unit U maps the kinematic modality data to video modality data according to the following formula to obtain the cross-modal data interaction feature between video and kinematics:

[0207]

[0208] In the formula:

[0209] is the initial input feature;

[0210] is the initial feature of the kinematic modality data;

[0211] is the initial feature of the video modality data;

[0212] is the intermediate feature of the j-th layer cross-modal attention sub-unit;

[0213] is the output feature of the (j-1)-th layer cross-modal attention sub-unit;

[0214] LN() represents the layer normalization function;

[0215] represents the multi-head function of the j-th layer cross-modal attention sub-unit;

[0216] represents the position feed-forward function of the j-th layer cross-modal attention sub-unit with θ as the parameter;

[0217] V→P represents the transfer of visual data to kinematic data.

[0218] The six pairs of cross-modal interaction features between kinematics and audio, kinematics and video, audio and kinematics, audio and video, video and kinematics, and video and audio are respectively denoted as ZP→A , Z P→V , Z A→P , Z A→V , Z V→P , Z V→A .

[0219] The specific working method and working principle of the self-attention mechanism module are as follows:

[0220] Each pair of cross-modal interaction features is used to simulate the interaction between different modalities. For different target modalities, in order to further adopt the self-attention mechanism module to obtain the internal correlation and contribution distribution within each interaction modality, the self-attention mechanism module is applied respectively, which is specifically expressed as:

[0221]

[0222] In the formula:

[0223] Z V is the video feature after cross-modal feature interaction processing;

[0224] Z A is the audio feature after cross-modal feature interaction processing;

[0225] Z P is the kinematic feature after cross-modal feature interaction processing;

[0226] represents the concatenation operation;

[0227] attention() represents the self-attention mechanism module function;

[0228] Through the above operations, a multi-modal mixed feature Z that is fully aware of adjacent modality information is obtained V , Z A , Z P .

[0229] The specific working method and working principle of the multi-head self-attention mechanism module are as follows:

[0230] To establish long-term dependence relationships between different modalities and fully mine the feature information of multi-modal data, through the multi-head self-attention mechanism, analyze the internal correlation between three single-modal features and three fused-modal features, calculate the attention weights they make for the final activity recognition, and re-distribute according to the attention weights to obtain a multi-modal fusion feature Z containing rich feature information, thereby enhancing the model's ability to identify key features.

[0231] The multi-head self-attention mechanism module can capture different feature relationships in parallel from multiple subspaces, capture different features in different subspaces, and understand information more comprehensively.

[0232] A multi-level feature fusion sub-module is set up. The multi-level feature fusion sub-module is used to splice multiple single-modal features and multiple fused-modal features. It inputs the single-modal original features extracted by each single-modal feature extraction module and the cross-modal interaction features output by the cross-modal attention mechanism module, and performs splicing according to the following formula:

[0233] M = Concat(Z V , Z A , Z P , X V , X A , X P );

[0234] In the formula:

[0235] X A represents the audio single-modal original feature;

[0236] X V represents the visual single-modal original feature;

[0237] X P represents the kinematic single-modal original feature;

[0238] Concat() represents the splicing operation.

[0239] For each head, a linear transformation is performed:

[0240] The attention value output of each head: head m = Attention(Q m , K m , V m );

[0241] Splice the outputs of all heads and perform a linear transformation to obtain the final multi-level fusion feature vector Z:

[0242] Z = MultiHeadOutput = Concat(head1, head2,... head h )W O ;

[0243] In the formula:

[0244] X is the input matrix, that is, the spliced multi-modal feature matrix;

[0245] Z is the final multi-level fusion feature vector;

[0246] Q m is the query vector of the m-th attention head;

[0247] Km is the key vector for the m-th attention head;

[0248] V m is the value vector for the m-th attention head;

[0249] is the query linear transformation weight matrix for the m-th attention head;

[0250] is the key linear transformation weight matrix for the m-th attention head;

[0251] is the value linear transformation weight matrix for the m-th attention head;

[0252] W O is the weight matrix for the output linear transformation;

[0253] head m is the output of the m-th attention head;

[0254] W O is used to map the concatenated result to the desired output dimension and is learned during the model training process.

[0255] Input the final multi-level fusion feature vector into the fully connected layer network to integrate the interaction information within each modality and between modalities, and output the activity recognition classification result through the Softmax function.

[0256] Divide the sample set into a training set, a validation set, and a test set; use the training set to train the recognition model, determine the optimal parameters of the recognition model according to the best performance on the validation set; use the test set to evaluate the final model performance of the recognition model.

[0257] Divide 60% of the dataset into the training set, 20% into the validation set, and 20% into the test set. The data is input into the network in batches, and each batch is used to update the weights. Use the Adam optimization algorithm to update the weights of the model and reduce the value of the loss function. Control the step size of weight update by dynamically adjusting the learning rate, and add a regularization term to prevent the model from overfitting. Determine the final optimal recognition model based on the best performance on the validation set.

[0258] The above-mentioned functional modules such as the data acquisition module, data preprocessing module, construction machinery activity state recognition model, single-modal feature extraction module, cross-modal attention mechanism module, multi-head self-attention mechanism module, OpenCV unit, S3D, VGGish, Conformer neural network, etc. can all adopt the applicable functional modules in the prior art, or adopt the applicable functional modules and software in the prior art and be constructed by conventional technical means.

[0259] The embodiments described above are only used to illustrate the technical idea and features of the present invention, and the purpose is to enable those skilled in the art to understand the content of the present invention and implement it accordingly. The patent scope of the present invention cannot be limited only by these embodiments. That is, any equivalent changes or modifications made in accordance with the spirit disclosed by the present invention still fall within the patent scope of the present invention.

Claims

1. A cross-modal depth recognition system for underground construction machinery activities, characterized in that: The system includes a data acquisition module, a data preprocessing module and a construction machinery activity state recognition model connected in sequence; the construction machinery activity state recognition model includes a unimodal feature extraction module, a cross-modal attention mechanism module and a multi-head self-attention mechanism module; the data acquisition module is used to collect video, audio and kinematic data of the construction machinery reflecting the activity state of the construction machinery; the data preprocessing module is used to preprocess the data collected by the data acquisition module to extract effective information; the unimodal feature extraction module is used to extract unimodal original features of video, audio and kinematic data; The cross-modal attention mechanism module is used to capture the correlation between various unimodal original features. Its input is the unimodal original features extracted by the unimodal feature extraction module, and its output is used to simulate the cross-modal interaction features of the interaction between different modalities; the multi-head self-attention mechanism module is used for multimodal feature fusion and classification. Its input is the unimodal original features extracted by the unimodal feature extraction module and the cross-modal interaction features output by the cross-modal attention mechanism module. It performs multimodal fusion on the input features and integrates the interaction information within each modal feature and between each modal feature through a fully connected layer network. It outputs the classification result of the activity state of construction machinery through an activation function.

2. The underground engineering construction machinery activity cross-modal depth recognition system according to claim 1 is characterized in that: The data preprocessing module includes an audio preprocessing submodule, a video preprocessing submodule and a kinematic preprocessing submodule; The audio preprocessing submodule performs the following preprocessing on the collected audio data: convert the audio data into time-frequency data through short-time Fourier transform, convert the frequency in the time-frequency data into Mel-scale frequency using Mel filter bank, and further perform discrete cosine transform on the logarithmic energy of the Mel spectrum of each frame to remove the correlation between different filter energies and extract Mel-frequency cepstrum coefficients; arrange the Mel spectrum of each frame in chronological order to obtain a complete Mel spectrum graph; The video preprocessing submodule randomly flips, randomly crops, and adjusts the color brightness of the extracted frames in the video clips to simulate different viewing angles and distances; The kinematic preprocessing submodule is provided with a first-order low-pass filter, which calculates the gravity acceleration component. The kinematic preprocessing submodule subtracts the gravity acceleration component from the original data collected by the data acquisition module.

3. The underground engineering construction machinery activity cross-modal depth recognition system according to claim 2 is characterized in that: The video preprocessing submodule includes an OpenCV unit, which uniformly samples 20 to 100 frames of images for each captured video clip.

4. The underground engineering construction machinery activity cross-modal depth recognition system according to claim 1 is characterized in that: The single-modal feature extraction module includes a video extraction submodule, an audio extraction submodule, and a kinematics extraction submodule; the video extraction submodule uses S3D to extract the original features of the video data; The audio extraction submodule uses VGGish to extract the original features of audio data, and the kinematics extraction submodule uses the Conformer neural network to extract the original features of kinematics data.

5. The underground engineering construction machinery activity cross-modal depth recognition system according to claim 1 is characterized in that: The cross-modal attention mechanism module includes the first to third cross-modal attention mechanism sub-modules; The first to third cross-modal attention mechanism sub-modules each include a one-dimensional temporal convolution layer, a position embedding unit, and multiple cross-modal attention units. The one-dimensional temporal convolution layer and the position embedding unit are connected in sequence, and the output end of the position embedding unit is connected to multiple cross-modal attention units in parallel. The one-dimensional temporal convolutional layers of the first to third cross-modal attention mechanism submodules correspond to the unimodal original features of the input video data, audio data, and kinematic data, respectively, and map the feature dimensions of different modalities to a fixed-dimensional space, outputting the features. In order to make the feature sequence output by the one-dimensional time convolution layer carry time information, the position embedding unit is as follows: Add location embedding: Where: Low-level position-aware features corresponding to video, audio, and kinematic modality data; is the output feature of the one-dimensional temporal convolution layer; d is the position embedding dimension; d {P,V,A} is the feature dimension corresponding to video, audio, and kinematic modal data; T {P,V,A} is the number of time steps corresponding to the video, audio, and kinematic modal data; PE(T {P,V,A} ,d) is position embedding, which is used to add time information to each time step in the feature sequence; PE() represents a fixed embedding function that calculates each position index; Each cross-modal attention unit of the first to third cross-modal attention mechanism sub-modules adopts a cross-modal algorithm for calculation, maps the data of one modality to the data of another modality for processing, and finally outputs six pairs of cross-modal interaction features: kinematics and audio, kinematics and video, audio and kinematics, audio and video, video and kinematics, and video and audio.

6. The underground engineering construction machinery activity cross-modal depth recognition system according to claim 5 is characterized in that: Each cross-modal attention unit is composed of multiple layers of cross-modal attention sub-units connected in series. Let one of the cross-modal attention units be a cross-modal attention unit U, and let j be the layer number of the cross-modal attention sub-unit in the cross-modal attention unit U; Let j = 1, ..., D; the cross-modal attention unit U performs feed-forward calculation layer by layer according to j = 1, ..., D; The cross-modal attention unit U maps the kinematic modality data to the video modality data according to the following formula to obtain the cross-modal data interaction features between video and kinematics: Where: is the initial input feature; is the initial feature of the kinematic modal data; is the initial feature of the video modality data; is the intermediate feature of the j-th layer cross-modal attention sub-unit; is the output feature of the j-1th layer cross-modal attention sub-unit; LN() represents the layer normalization function; Represents the multi-head function of the j-th layer cross-modal attention sub-unit; represents the position feedforward function of the j-th layer cross-modal attention subunit with θ as parameter; V→P means that visual data is transferred to kinematic data.

7. The underground engineering construction machinery activity cross-modal depth recognition system according to claim 1 is characterized in that: A self-attention mechanism module is also connected between the cross-modal attention mechanism module and the multi-head self-attention mechanism module. The self-attention mechanism module is used to obtain the correlation and contribution distribution within each pair of cross-modal interaction features; the self-attention mechanism module inputs each pair of cross-modal interaction features, and outputs a multimodal mixed feature that reflects the correlation and contribution distribution within each pair of cross-modal interaction features.

8. A method for cross-modal depth recognition of underground construction machinery activities using the cross-modal depth recognition system for underground construction machinery activities according to any one of claims 1 to 7, characterized in that: The method comprises the following steps: Step 1: construct a construction machinery activity state recognition model, connect the unimodal feature extraction module, the cross-modal attention mechanism module and the multi-head self-attention mechanism module in the recognition model in sequence; and connect the output end of the unimodal feature extraction module to the input end of the multi-head self-attention mechanism module; Step 2, enabling the data acquisition module to collect kinematic, audio, and video sample data reflecting the activity state of the construction machinery; The data preprocessing module preprocesses the sample data to extract effective information; the preprocessed sample data is compiled into a sample set for training a construction machinery activity state recognition model; Step 3, dividing the sample set into a training set, a validation set and a test set; using the training set to train the recognition model, determining the optimal parameters of the recognition model based on the optimal performance of the validation set; using the test set to evaluate the final model performance of the recognition model; Step 4, enabling the data acquisition module to collect real-time kinematic, audio, and video data of the activity status of the construction machinery at the construction site; The data preprocessing module is used to preprocess the real-time data; the preprocessed real-time data is input into the trained construction machinery activity state recognition model, and the recognition model outputs the construction machinery activity state classification result in real time.

9. The method for cross-modal depth identification of underground construction machinery activities according to claim 8 is characterized in that: Step 3 includes the following method steps: 60% of the sample set is divided into training set, 20% is divided into validation set, and 20% is divided into test set; the data is input into the construction machinery activity state recognition model in batches, and each batch is used to update the weights. The Adam optimization algorithm is used to update the weights of the model and reduce the loss function value; the learning rate is dynamically adjusted to control the step size of the weight update, and a regularization term is added to prevent the model from overfitting; the final optimal recognition model is determined based on the optimal performance of the validation set.

10. A device for a method for cross-modal depth recognition of underground construction machinery activities, comprising a memory and a processor, characterized in that: The memory is used to store computer programs; the processor is used to execute the computer program and implement the steps of the cross-modal depth identification method for underground engineering construction machinery activities as described in claim 8 when executing the computer program.