Traffic target detection and recognition method based on acoustic vibration time-frequency characteristics and cross attention fusion mechanism

By using acoustic vibration time-frequency features and a cross-attention fusion mechanism, the stability problem of traffic target detection in complex environments is solved, and efficient traffic target recognition is achieved.

CN119992040BActive Publication Date: 2025-12-19NANJING UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411257787.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-09
Publication Date
2025-12-19
Estimated Expiration
2044-09-09

AI Technical Summary

Technical Problem

Existing traffic target detection technologies are susceptible to environmental interference, have high algorithm complexity, require large hardware devices, and are unstable under complex working conditions due to the single detection method.

Method used

The method employs acoustic vibration time-frequency features and cross-attention fusion mechanism. It uses acoustic vibration sensor to collect signals for filtering and noise reduction, variational mode decomposition, and extracts Mel spectrogram and wavelet transform time-frequency features. The CNN-transformer model is then used for feature fusion and recognition.

Benefits of technology

It improves the signal-to-noise ratio, deeply mines potential time-frequency features, overcomes the limitations of single detection equipment, and achieves stable identification of traffic targets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992040B_ABST
    Figure CN119992040B_ABST
Patent Text Reader

Abstract

The application provides a traffic target detection and recognition method based on sound-vibration time-frequency characteristics and cross-attention fusion mechanism. First, sound-vibration signals of different traffic targets are collected by a sound-vibration sensor, and normalized least mean square (NLMS) noise filtering preprocessing is performed. Second, the preprocessed sound-vibration signals are subjected to variational mode decomposition (VMD), and scale spectrum segmentation method and summation fuzzy entropy minimum value method are adopted to decompose the sound-vibration signals into multiple intrinsic mode functions (IMFs). Third, Mel spectrograms are extracted from the sound signal IMFs, and wavelet transform time-frequency diagrams are extracted from the vibration signal IMFs, and the results are subjected to CNN convolution pooling, and transformer encoder is further used to extract sound-vibration signal characteristics of different traffic targets. Finally, cross-attention mechanism is used for coding, sound signal characteristics and vibration signal characteristics are fused into new characteristics, and Softmax function and Dropout function are used for normalization and overfitting prevention. The application has the advantages of low algorithm complexity, strong real-time performance and low cost, and solves the traffic target detection problem under extreme climate, weather, light and other scenes.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of intelligent traffic detection, and specifically relates to a traffic target detection and recognition method based on sound-vibration time-frequency characteristics and a cross-attention fusion mechanism. BACKGROUND

[0002] The development of intelligent traffic information has given rise to an urgent need for traffic information detection, and traffic information detection has gradually upgraded from traditional coil, geomagnetic and other methods to detection methods represented by video and radar. However, video and radar detection have some inherent defects, video is easily disturbed by external climate, weather, light and other environments, radar is easily disturbed by electromagnetic interference and has a high cost, and both have large data volume, high algorithm complexity and high performance requirements for hardware devices. In special extreme scenarios, the detection requirements cannot be met.

[0003] At present, the related research on traffic target recognition mostly adopts machine learning and deep learning algorithms, sound and vibration signals have rich information, and the main features include short-time zero-crossing rate, mel-frequency cepstral coefficient MFCC, linear prediction cepstral coefficient LPCC, mel spectrogram, wavelet transform time-frequency diagram and the like. By extracting features of sound and vibration signals for different traffic targets, a multi-feature fusion feature library is constructed, a deep learning neural network method is adopted, and a related model is constructed for recognition. At present, the research and application of sound and vibration signal fusion for detecting traffic targets are relatively few. SUMMARY

[0004] The application provides a traffic target detection and recognition method based on sound-vibration time-frequency characteristics and a cross-attention fusion mechanism.

[0005] The technical scheme for achieving the object of the application is as follows: a traffic target detection and recognition method based on sound-vibration time-frequency characteristics and a cross-attention fusion mechanism, and the specific steps are as follows:

[0006] Step 1: filtering and denoising pretreatment is performed on the sound-vibration signals of different traffic targets collected;

[0007] Step 2: the sound-vibration signals of different traffic targets after pretreatment are subjected to variational mode decomposition, a scale spectrum segmentation method is adopted, the optimal modal parameter value in the VMD decomposition process of the sound-vibration signals is obtained, and the optimal penalty factor is obtained by using the minimum sum fuzzy entropy, and finally the sound-vibration signals are decomposed into a plurality of intrinsic mode functions IMF;

[0008] Step 3: feature extraction is performed on the independent sound-vibration signals IMF after VMD decomposition, and the extracted features include a mel spectrogram and a wavelet transform time-frequency diagram;

[0009] Step 4: input the sound signal Mel spectrogram and the vibration signal wavelet transform time-frequency diagram into the CNN-transformer model based on the cross attention fusion mechanism for training, and obtain the CNN-transformer model based on the cross attention fusion mechanism with optimal training parameters;

[0010] Step 5: real-time acquisition of sound and vibration signals of the traffic target, processing of the collected signals according to steps 1-3, obtaining of the sound signal Mel spectrogram and the vibration signal wavelet transform time-frequency diagram of the real-time traffic target, input of the cross attention fusion mechanism CNN-transformer model in step 4, utilization of the new features F of the traffic target after fusion, and obtaining of the recognition result of the traffic target according to the optimal training parameter model.

[0011] Compared with the prior art, the present application has the following advantages: (1) the multi-sensor fusion composite detection means overcomes the instability of single sound and vibration detection equipment under complex working conditions; (2) the VMD algorithm can improve the signal-to-noise ratio of the sound and vibration signals and deeply mine the real and effective potential time-frequency features; (3) the cross attention fusion mechanism constructs the internal connection between the sound features and the vibration features, and realizes the joint coding and information exchange of the two.

[0012] The present application will be further described in detail below with reference to the accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS

[0013] Figure 1 A flowchart of a traffic target detection and recognition method based on sound and vibration time-frequency features and a cross attention fusion mechanism is provided for the embodiments of the present application.

[0014] Figure 2 A traffic target sound and vibration signal time-frequency feature extraction schematic diagram is provided for the embodiments of the present application.

[0015] Figure 3 A sound and vibration time-frequency feature cross attention fusion mechanism algorithm flowchart is provided for the embodiments of the present application. DETAILED DESCRIPTION

[0016] A traffic target detection and recognition method based on acoustic vibration time-frequency characteristics and cross-attention fusion mechanism, first, an acoustic vibration sensor is used to collect acoustic vibration signals of different traffic targets (people, motor vehicles, non-motor vehicles), and preprocessing is performed. Secondly, the preprocessed acoustic vibration signals are subjected to variational mode decomposition (VMD), and scale spectrum segmentation method and summation fuzzy entropy minimum method are used to decompose into multiple intrinsic mode functions (IMF). Thirdly, the IMF of the sound signal is extracted Mel spectrogram, the vibration signal IMF is extracted wavelet transform time-frequency diagram, the results are subjected to CNN convolution pooling to obtain the input of the transformer, the acoustic vibration signal features of different traffic targets are extracted by using the transformer encoder layer. Finally, the cross-attention mechanism is used for coding, the sound signal features and vibration signal features are fused into new features, and the Softmax function and Dropout function are used for normalization and overfitting prevention processing. According to the optimal parameters, the test set is input to obtain the classification accuracy of traffic targets (people, motor vehicles, non-motor vehicles). The steps include:

[0017] Step 1: acoustic vibration signal filtering and noise reduction. Set the sampling rate of the sound and vibration sensor to be the same, that is, Fs, and after the acoustic vibration sensor collects the traffic target data, align the time stamp and truncate the effective signal of the same length. Let the traffic target sound signal be x a (t), and the traffic target vibration signal be x v (t), use the NLMS adaptive filtering and noise reduction algorithm to remove the direct current component of the acoustic vibration signal to facilitate subsequent data processing.

[0018] Specifically, the NLMS adaptive filtering and noise reduction process is:

[0019] NLMS is a normalized LMS filtering algorithm, considering the weight vector obtained in the nth iteration ω(n), and the weight obtained in the (n+1)th iteration ω(n+1), the NLMS filtering design is converted into an optimization problem of weight under the constraint condition; give the input vector x(t) of the filter order or tap number L at time t and the expected response d(t), so that the Euclidean norm of the weight iteration increment is minimized, and the optimization criterion can be expressed as:

[0020] ω H (n+1)x(t)=d(t)

[0021] min J=min||ω(n+1)-ω(n)|| P

[0022] Where (·) H represents taking the conjugate transpose, J is the optimization criterion function, ω(n) is the weight of the NLMS filter, and p represents the norm power term of the optimization criterion function.

[0023] The NLMS filtering algorithm denoising to a sound signal x a The specific process of (t) can be expressed as follows:

[0024] (1) Determine the initial condition, ω a (0) = 0 or give the initial filter weight ω a (0) according to the actual condition.

[0025] (2) Calculate the sound signal output response:

[0026]

[0027] (3) Calculate the error according to the expected response d a (t) and the output response y a (t):

[0028] e a (t) = y a (t) - d a (t)

[0029] (4) Filter weight update:

[0030]

[0031] Where, ω a (n) is the weight of the filter when processing the sound signal, e a (t) is the error signal at time t, x a (t) is the input vector of the sound signal at time t, ||x a (t) || 2 represents the norm of the input signal, y a (t) is the output response at time t, d a (t) is the expected response at time t, (·) * represents the complex conjugate, μ is the step factor, and 0 < μ < 2, α is the correction value, and 0 ≤ α ≤ 1.

[0032] Stop updating when the filter weight ω a (n) satisfies the minimum optimization criterion function mmnJ, and output y a (t).

[0033] The NLMS filtering algorithm denoising to a sound signal x v (t) can be expressed as follows:

[0034] (1) Determine the initial condition, ω v (0) = 0 or give the initial filter weight ω v (0) according to the actual condition.

[0035] (2) Calculate the vibration signal output response:

[0036]

[0037] (3) Calculate the error according to the desired response d v (t) and the output response y v (t) :

[0038] e v (t) = y v (t) - d v (t)

[0039] (4) Update the filter weight:

[0040]

[0041] where ω v (n) is the weight of the filter when processing the vibration signal, e v (t) is the error signal at time t, x v (t) is the input vector of the vibration signal at time t, ||x v (t) || 2 represents the norm of the input signal, y v (t) is the output response at time t, d v (t) is the desired response at time t, (·)* represents complex conjugate, μ is the step factor, and 0 < μ < 2, α is the correction value, and 0 ≤ α ≤ 1.

[0042] Stop updating when the filter weight ω v (n) satisfies the minimum optimization criterion function min J, and output y v (t).

[0043] The sound signal x a (t) is filtered and denoised by the above formula to output y a (t), and the vibration signal x v (t) is filtered and denoised by the above formula to output y v (t).

[0044] Step 2: Variational Modal Decomposition (VMD) decomposes the sound-vibration signal in step 1 into multiple Intrinsic Mode Functions (IMF). Before VMD decomposition, the optimal modal parameter values K1 and K2 and the optimal penalty factors α1 and α2 in the VMD decomposition process of the sound-vibration signal need to be determined.

[0045] 2.1 Determine the optimal modal number K using the scale spectrum segmentation method:

[0046]

[0047] Where Y(f) is the discrete time series of the input traffic target sound and vibration signal, y(t) is the input traffic target sound and vibration signal, and f represents the signal frequency.

[0048]

[0049] Where L(f, s) is the discrete scale spectrum of Y(f), g(f, s) is the sampling Gaussian kernel function, s is the scale parameter, and S is the window size of the truncated filter.

[0050] All the segmentation boundary points can be found in the scale spectrum by detecting each local extreme value of L(f, s) and searching for the local minimum value on both sides, and they are divided into continuous K wave bands, thus obtaining the modal parameters required by VMD.

[0051] 2.2 Determine the optimal penalty factor a by using SFE (sum fuzzy entropy):

[0052] (1) Set the maximum value of the penalty factor a max , the minimum value a min and the step size L. The maximum value is usually set to 6000, and the minimum value and the step size are usually set to 1% of the decomposition data amount.

[0053] (2) Calculate the modal component fuzzy entropy H(c k (t))

[0054]

[0055] H(c k (t)) = -∫p(c k (t))log(p(c k (t)))dt

[0056] Where F s is the sampling rate of the sound and vibration sensor, N is the length of each component of the IMF component {c k (t)} after decomposition of the sound and vibration signal, p(c k (t)) is the probability of the modal component c k (t) at time point t, and H(c k (t)) is the fuzzy entropy of the modal component c k (t).

[0057] (3) Calculate the sum fuzzy entropy SFE

[0058]

[0059] Where SFE is the sum fuzzy entropy, w k is the modal component c kThe weight of (t), i.e. the center frequency of each modal component.

[0060] (4) Let α = α min , find the SFE of VMD decomposition of K modal components when the penalty factor is α, when α ≤ α max , let α = α + L, and sequentially find the corresponding SFE; when α > α max , end the loop, and the abscissa corresponding to the minimum value of SFE is the optimal value of the penalty factor α.

[0061] 2.3 VMD variational modal decomposition:

[0062] (1) Obtain the analysis signal of y(t) through Hilbert transform, and calculate its one-sided spectrum. By multiplying the operator , the center band of y(t) is modulated to the corresponding baseband:

[0063]

[0064] (2) The modal component uses its own estimated center frequency to perform frequency mixing through exponential correction, and gradually adjusts the center frequency to the corresponding bandwidth.

[0065] (3) Introduce a Gaussian smoothed demodulation signal to estimate the bandwidth of the modal component band, and obtain the constraint expression of the variational model through the square root of the norm gradient, which is:

[0066]

[0067] In the formula: is the partial derivative with respect to t, δ(t) is the impulse function, K is the optimal modal decomposition number, {c k (t)} is the decomposed IMF component, {c k (t)} = {c1, c2, c3, …, c k}, {w k} is the center frequency of each modal component, {w k} = {w1, w2, w3, …, w k}, y(t) is the input traffic target sound vibration signal.

[0068] (4) Introduce the penalty factor α and the Lagrange multiplier λ(t), and further obtain the Lagrange function:

[0069]

[0070] Among them, λ(t) is the constraint enhancement.

[0071] Perform time-frequency domain conversion on the Lagrange function to obtain the corresponding extreme solution, and further obtain the modal component and center frequency expression:

[0072]

[0073] wherein, is the power spectral centroid of is the modal component formed by VMD in the iteration process, is the power spectral centroid of are the Fourier transforms of the modal component, the original signal and the Lagrangian operator, respectively.

[0074] Finally, by alternating direction multiplication, the saddle point of the Lagrangian function is constantly iterated and updated {u k}, {w k}, λ, so as to obtain the optimal solution of the variational model constraint expression, and finally obtain the K modal components of the input signal y(t). Finally, the input sound signal y a (t) is converted into the output signal c ak (t), which is represented as A k (t) hereafter; the input vibration signal y v (t) is converted into the output signal c vk (t), which is represented as V k (t) hereafter.

[0075] Step 3: Time-frequency feature extraction is performed on the independent sound IMF after VMD decomposition. The filtered and denoised sound and vibration signals are framed and windowed. The Mel spectrogram is extracted for the i-th frame of the i-th IMF of the sound, and the wavelet transform time-frequency diagram is extracted for the vibration signal extracted from the i-th IMF of the i-th frame of the sound.

[0076] 3.1 Framing and windowing

[0077] The length of each component of the IMF component {c k (t)} after decomposition of the sound and vibration signals is set as N, the sampling frequency is F s , the length of each frame is taken as W len , the frame shift is inc, and the overlapping part between adjacent two frames is overlap = W len -inc.

[0078] For the sound and vibration signals with a sample length of N, framing is performed as follows:

[0079] C n (t) = (N-w len ) / inc+1

[0080] Wherein, N is the length of each component of the IMF component {c k (t)} after decomposition of the sound and vibration signals, and W lenFor each frame length, inc is the frame shift, and overlap is the overlap between two adjacent frames.

[0081] After the frame division, the windowing operation is needed to get the frame signal. In this experiment, the Hanning window is selected. The result of applying the window function to each frame signal is as follows

[0082]

[0083] where {c k (t)} is the i-th frame signal of the k-th IMF component after decomposition, C k,i (t) is the i-th frame signal of the k-th IMF component after windowing, and N is the length of each component of the IMF component {c k (t)} after decomposition of the acoustic vibration signal, and n is the sample point index in the window.

[0084] This step finally converts the input sound signal A k (t) into the output signal A k ′(t); and converts the input vibration signal V k (t) into the output signal V k ′(t).

[0085] 3.2 Extracting the Mel spectrogram for the sound signal

[0086] Performing short-time Fourier transform on each frame signal A k ′(t) to obtain the frequency spectrum representation, and the formula is as follows:

[0087]

[0088] where A″(p, m) represents the complex value of the p-th frequency point of the m-th frame, and N is the window length of each frame.

[0089] For each frame Fourier transform result, calculate its power spectrum:

[0090] P(p, m) = |A″(p, m)| 2

[0091] where P(p, m) is the power spectrum density of each frame.

[0092] Filter the power spectrum through a set of Mel filters, and the response of the filter can be represented as:

[0093]

[0094] where f(m) is the center frequency of the m-th Mel filter, and k is the index on the frequency axis.

[0095] E m (n) = P(p, m) · Hm (k)

[0096] M m (n) = log[E m (n)]

[0097] Among them, M m (n) are the Mel spectrum coefficients after logarithmic transformation, E m (n) is the energy of the nth frame signal after passing through the mth filter.

[0098] With n as the horizontal axis, the eigenvector M of the Mel spectrum coefficients after logarithmic transformation m (n) is plotted on the vertical axis to show the energy distribution of the sound signal in different frequency bands, resulting in a Mel spectrogram. The Mel spectrograms of K1 sound signals are then represented as follows:

[0099] 3.3 Extracting wavelet transform time-frequency diagram from vibration signal

[0100]

[0101] Where ψ(t) is the mother wavelet function, ψ a,b (t) is the wavelet function after translation and scaling transformation, where a is the scaling factor (controlling the wavelet scaling) and b is the translation factor (controlling the position of the wavelet on the time axis).

[0102]

[0103] Among them, W x (a, b) are the wavelet coefficients obtained from continuous wavelet transform, V k ′(t) is the input vibration signal, It is ψ a,b The complex conjugate of (t).

[0104] For a given scale *a*, its corresponding feature scale is defined as: W is calculated at multiple scales and different time locations b. x (a, b), finally plotted with the translation factor b as the abscissa and the feature scale. W is the ordinate. x (a, b) are points The corresponding values ​​are represented by color intensities, and an RGB time-frequency diagram is plotted. The result is the wavelet transform time-frequency diagram of K2 vibration signals, represented as follows:

[0105] Step 4: Input the Mel spectrogram of the sound signal and the wavelet transform time-frequency graph of the vibration signal into the CNN-transformer model with cross-attention fusion mechanism for training to obtain the optimal training parameters.

[0106] The CNN-transformer model of the cross-attention fusion mechanism comprises: obtaining a sound signal feature map and a vibration signal feature map through CNN convolution pooling; then, the sound signal feature map and the vibration signal feature map are respectively converted into feature labels by using an encoder layer to extract sound signal features and vibration signal features of different traffic targets; finally, a cross-attention mechanism is used for encoding, and the sound signal features and the vibration signal features are fused into new features, and a Softmax function and a Dropout function are used for normalization and overfitting prevention processing.

[0107] 4.1 CNN convolution pooling

[0108] The K1 Mel spectrograms of sound signals and the K2 wavelet transform time-frequency graphs of vibration signals are respectively subjected to CNN convolution pooling by using a CNN network to obtain the input of the encoder. In the convolution layer, the convolution kernel performs convolution on the output of the previous layer, and a nonlinear activation function is used to construct the output features. In the pooling layer, a large matrix is converted into a small matrix through data downsampling.

[0109] Specifically, the convolution process of the convolution layer is as follows:

[0110]

[0111] wherein, represents the weight of the kernel in the i-th filter in the l-th layer, wherein, represents the bias of the kernel in the i-th filter in the l-th layer, x l (j) represents the j-th local region in the l-th layer, represents the input of the j-th neuron in the l+1-th layer.

[0112] Specifically, the nonlinear activation function is as follows:

[0113] After the convolution operation, the numerical output of each convolution is subjected to nonlinear conversion by using an activation function. In this paper, a ReLU function is used as the activation function, and the formula of the function is as follows:

[0114]

[0115] wherein, represents the output value of the convolution operation, represents the activation value of .

[0116] Specifically, the pooling process of the pooling layer is as follows:

[0117]

[0118] wherein, denotes the value of the t-th neuron in the i-th feature in the l-th layer after activation, (j-1)W P +1≤t≤W P , W P denotes the width of the pooling region, denotes the value of the i-th neuron in the l+1-th layer.

[0119] The sound signal Mel spectrogram and vibration signal wavelet transform time-frequency diagram of the traffic target are convolved and pooled to output a feature map X n ∈R C*H*W ,

[0120] where C represents the number of channels, H represents the image height, and W represents the image width

[0121] 4.2 The transformer encoder extracts the sound-vibration features.

[0122] The core structure of the transformer encoder is composed of a multi-head self-attention mechanism and a feedforward neural network. The self-attention mechanism captures global feature dependencies and improves the effectiveness of feature representation. The output results of K1 sound signal Mel spectrograms and K2 vibration signal wavelet transform time-frequency diagrams after CNN convolution and pooling are input into the transformer encoder, and the CNN output feature map is converted into a feature tag to extract sound-vibration signal features of different traffic targets (people, motor vehicles, and non-motor vehicles).

[0123] The feature map X n ∈R C*H*W is divided into blocks of the same size:

[0124] X∈RN L (a L *a L *C)

[0125] where (a L *a L ) represents the resolution of each block, is the length of the input sequence, and C represents the number of channels.

[0126] The blocks are projected into a fixed dimension space in the entire transformer encoder layer using a linear layer, spatial information is encoded, and a one-dimensional learnable position embedding is added in the block.

[0127] f0=[x 1 A;x 2 A;…x N A]+A position

[0128] where A ∈ RN a*a*C*M is the block-wise embedding projection; A position ∈ R N*M is a one-dimensional learnable position embedding.

[0129] Thus the i-th layer output can be represented as:

[0130] f′ i = MHA(VN(f i-1 ))+f i-1 , i = 1…V

[0131] f i = MLP(VN(f′ i ))+f′ i , i = 1…V

[0132] where f i and f′ i are the output of the i-th layer; VN(.) is layer normalization; MHA is V-layer multi-head self-attention; MLP is a multi-layer perceptron consisting of 2 linear layers activated by GRLU; i is an identifier; V is the number of layers of the transformer encoder.

[0133] The final output results are sound feature F audio and vibration feature F vibro .

[0134] 4.3 Cross-attention mechanism feature fusion

[0135] The cross-attention mechanism is used for encoding to realize feature information exchange, build the internal connection between sound features and vibration features, realize joint encoding of the two, and output classification results through a fully connected layer and Softmax. After cross-attention encoding, sound feature F audio and vibration feature F vibro are fused into new feature F. Sound feature F audio and vibration feature F vibro are mapped to two feature spaces through a linear layer, where q (Query) represents a query vector, k (Key) represents a key vector, and v (Value) represents a value vector. The results of the sound feature and the vibration feature after mapping are represented as q a , k a , v a and q v , k v , v v .

[0136] q a , k a , v a = Linear(F audio )

[0137] q v , k v , v v = Linear(F vibro )

[0138] The cross-attention score score is calculated by q a , k a , v a and q v , k v , v v , and normalized and over-fitting prevention processing is performed by using a Softmax function and a Dropout function:

[0139]

[0140] score' = Softmax(score)

[0141] score'' = Dropout(score')

[0142] F av = score'' · v a

[0143] Wherein, score is the cross-attention score, λ c is a custom constant, score', score'' and F av are intermediate quantities.

[0144] F av is processed through a full connection layer and normalization to obtain F:

[0145] F = Normalization(ω1F av + ω2)

[0146] Wherein, ω1 and ω2 are parameters to be learned.

[0147] For the learning parameters ω1 and ω2 in the model fused feature F = Normalization(ω1F av + ω2), when performing parameter iteration, first define the loss function L(ω) of the model, wherein ω is the learning parameter, and the specific learning process is as follows:

[0148] (1) Set the initial parameter values ω1(0), ω2(0) and define the learning rate α F .

[0149] (2) In each iteration process, the gradient of the loss function to the parameter is calculated and the parameter is updated:

[0150]

[0151]

[0152] where n represents the number of iterations, denotes the gradient of the loss function with respect to the parameters, and α F denotes the learning rate.

[0153] (3) Using the adaptive learning rate algorithm Adam optimizer, further improve the parameter optimization efficiency.

[0154]

[0155]

[0156]

[0157]

[0158]

[0159]

[0160] where m n and v n respectively represent the moving average of the first and second moments of the gradient of the parameter ω1, m n ' and v n ' respectively represent the moving average of the first and second moments of the gradient of the parameter ω2, β1 and β2 are hyperparameters, and ξ is a small constant to avoid the denominator being zero.

[0161] (4) By cross-validation method, evaluate different parameter combinations. For each parameter combination, record the classification accuracy of the model on the test set, select the best performance parameter combination as the final optimal parameter, and input the test set according to the optimal parameter to finally obtain the classification accuracy of the traffic target (person, motor vehicle, non-motor vehicle).

[0162] Step 5: Real-time acquisition of traffic target sound and vibration signals, processing the collected signals according to steps 1-3 to obtain real-time traffic target sound signal Mel spectrogram and vibration signal wavelet transform time-frequency diagram, inputting the cross-attention fusion mechanism CNN-transformer model in step 4, using the new traffic target fusion feature F, and obtaining the recognition result of the traffic target according to the optimal training parameter model.

[0163] The main advantages of detecting traffic targets by fusing acoustic vibration signals are as follows: (1) acoustic vibration signals have a long propagation wavelength and strong diffraction ability, and are not sensitive to small obstructions such as road debris; (2) acoustic and vibration detection devices have simple sensor structures, small size, light weight, and strong mobility, and can work all day, low power consumption, and for a long time; (3) acoustic vibration signal data has a small storage space and rich features, providing conditions for multi-feature fusion recognition.

[0164] Embodiments

[0165] The method divides traffic targets into three categories: people, motor vehicles, and non-motor vehicles. The sampling rate of the sound and vibration sensors is set to 10240Hz, and the road section of the sensor layout point should meet the following conditions:

[0166] (1) Select all-weather traffic saturation less than 0.6;

[0167] (2) The straight-line distance from the subway and high-speed rail line is less than 500m;

[0168] (3) No large factory or construction road within 1km;

[0169] Collect the sound and vibration signals of the traffic targets at the same time and save the original data.

[0170] Step 1: acoustic vibration signal filtering and noise reduction. After the acoustic vibration sensor collects the traffic target data, align the time stamp and truncate the effective signal of the same length. Let the traffic target sound signal be x a (t), and the traffic target vibration signal be x v (t), use the NLMS adaptive filtering and noise reduction algorithm to remove the direct current component of the acoustic vibration signal to facilitate subsequent data processing.

[0171] Specifically, the NLMS adaptive filtering and noise reduction process is as follows:

[0172] NLMS is a normalized LMS filtering algorithm. The weight vector obtained by the nth iteration is ω(n), and the weight obtained by the (n+1)th iteration is ω(n+1). The NLMS filtering design is converted into an optimization problem of the weight under the constraint condition; the input vector x(t) of the filter order or tap number L at time t and the expected response d(t) are given, so that the Euclidean norm of the weight iteration increment is minimized, and the optimization criterion can be expressed as:

[0173] ω H (n+1)x(t)=d(t)

[0174] minJ=min||ω(n+1)-ω(n)|| P

[0175] Where (·) Hwherein J represents an optimization criterion function, ω (n) represents a weight of the NLMS filter, and p represents a norm power term of the optimization criterion function.

[0176] The NLMS filtering algorithm reduces noise of a sound signal x a The specific process of the sound signal x

[0177] (1) determining an initial condition, ω a (0) = 0 or giving an initial filter weight ω a (0) according to an actual condition. (2) calculating a sound signal output response:

[0178]

[0179] (3) calculating an error according to an expected response d a (t) and the output response y a (t):

[0180] e a (t) = y a (t) - d a (t)

[0181] (4) updating the filter weight:

[0182]

[0183] wherein ω a (n) represents a weight of the filter during sound signal processing, e a (t) represents an error signal at time t, x a (t) represents an input vector of the sound signal at time t, ||x a (t) || 2 represents a norm of the input signal, y a (t) represents an output response at time t, d a (t) represents an expected response at time t, (·) * represents a complex conjugate, μ represents a step factor, and 0 < μ < 2, and α represents a correction value, and 0 ≤ α ≤ 1.

[0184] The updating is stopped when the filter weight ω a (n) satisfies a minimum optimization criterion function min J, and y a (t) is outputted.

[0185] The NLMS filtering algorithm reduces noise of a vibration signal x v (t) can be represented as follows:

[0186] (1) determining an initial condition, ω v (0) = 0 or giving an initial filter weight ωv (0).

[0187] (2) Calculate the vibration signal output response:

[0188]

[0189] (3) Calculate the error according to the desired response d v (t) and the output response y v (t) :

[0190] e v (t) = y v (t) - d v (t)

[0191] (4) Filter weight update:

[0192]

[0193] where ω v (n) is the weight of the filter when processing the vibration signal, e v (t) is the error signal at time t, x v (t) is the input vector of the vibration signal at time t, ||x v (t) || 2 represents the norm of the input signal, y v (t) is the output response at time t, d v (t) is the desired response at time t, (·)* represents complex conjugate, μ is the step factor, and 0 < μ < 2, α is the correction value, and 0 ≤ α ≤ 1.

[0194] When the filter weight ω v (n) satisfies the minimum optimization criterion function minJ, stop updating, and output y v (t).

[0195] The sound signal x a (t) is filtered and denoised by the above formula to output y a (t), and the vibration signal x v (t) is filtered and denoised by the above formula to output y v (t)

[0196] Step 2: Variational Modal Decomposition (VMD) decomposes the sound-vibration signal in step 1 into multiple Intrinsic Mode Functions (IMF). Before VMD decomposition, the optimal modal parameter values K1 and K2 and the optimal penalty factors α1 and α2 in the VMD decomposition process of the sound-vibration signal need to be determined.

[0197] 2.1 Determine the optimal modal number K using the scale spectrum segmentation method:

[0198]

[0199] where Y(f) is the discrete time series of the input traffic target vibro-acoustic signal, y(t) is the input traffic target vibro-acoustic signal, and f represents the signal frequency.

[0200]

[0201] where L(f, s) is the discrete scale spectrum of Y(f), g(f, s) is the sampling Gaussian kernel function, s is the scale parameter, and S is the window size of the truncated filter.

[0202] All the segmentation boundary points can be found in the scale spectrum by detecting each local extreme value of L(f, s) and searching for the local minimum values on both sides, and they are divided into continuous K wave bands, thus obtaining the modal parameters required by VMD.

[0203] 2.2 Determine the optimal penalty factor a by using SFE (sum fuzzy entropy):

[0204] (1) Set the maximum value of the penalty factor a max , the minimum value a min and the step size L. The maximum value is set to 6000, and the minimum value and the step size are set to 1% of the decomposition data amount.

[0205] (2) Calculate the modal component fuzzy entropy H(c k (t))

[0206]

[0207] H(c k (t)) = -∫p(c k (t))log(p(c k (t)))dt

[0208] where F s is the sampling rate of the sound and vibration sensor, N is the length of each component of the decomposed IMF component {c k (t)} of the vibro-acoustic signal, p(c k (t)) is the probability of the modal component c k (t) at time point t, and H(c k (t)) is the fuzzy entropy of the modal component c k (t).

[0209] (3) Calculate the sum fuzzy entropy SFE

[0210]

[0211] where SFE is the sum fuzzy entropy, w kis the modal component c k (t) the weight of each modal component, i.e. the center frequency.

[0212] (4) Let α = α min , find the SFE of VMD decomposition of K modal components when the penalty factor is α, when α≤α max , let α = α + L, and sequentially find the corresponding SFE; when α > α max , end the loop, and the horizontal coordinate corresponding to the minimum value of SFE is the optimal value of the penalty factor α.

[0213] 2.3 VMD variational modal decomposition:

[0214] (1) Get the analytic signal of y(t) through Hilbert transform, and calculate its one-sided spectrum. By multiplying the operator , the center band of y(t) is modulated to the corresponding baseband:

[0215]

[0216] (2) The modal component uses its own estimated center frequency to perform frequency mixing through exponential correction, and gradually adjusts the center frequency to the corresponding bandwidth.

[0217] (3) Introduce a Gaussian smoothed demodulation signal to estimate the bandwidth of the modal component frequency band, and get the constraint expression of the variational model through the square root of the norm gradient, which is:

[0218]

[0219] In the formula: is the partial derivative of t, δ(t) is the impulse function, K is the optimal modal decomposition number, {c k (t)} is the decomposed IMF component, {c k (t)} = {c1, C2, C3, …, c k}, {w k} is the center frequency of each modal component, {w k} = {w1, w2, w3, …, w k}, y(t) is the input traffic target sound vibration signal.

[0220] (4) Introduce the penalty factor α and the Lagrange multiplier λ(t), and then get the Lagrange function:

[0221]

[0222] Among them, λ(t) is the constraint enhancement.

[0223] The time-frequency domain conversion is performed on the Lagrange function to obtain the corresponding extreme value solution, and then the modal component and center frequency expression are obtained:

[0224]

[0225] wherein, is the Fourier transform of , is the modal component formed by VMD in the iteration process, is the power spectrum center of , are the Fourier transforms of the modal component, the original signal and the Lagrange operator respectively.

[0226] Finally, the saddle point of the Lagrange function is obtained by alternately multiplying and iteratively updating {u k}, {w k} and λ, so as to obtain the optimal solution of the variational model constraint expression, and finally obtain the K modal components of the input signal y(t). Finally, the input sound signal y a (t) is converted into the output signal c ak (t), which is represented as A k (t) hereafter; and the input vibration signal y v (t) is converted into the output signal c vk (t), which is represented as V k (t) hereafter.

[0227] Step 3: Time-frequency feature extraction is performed on the independent sound IMF after VMD decomposition. The filtered and denoised sound and vibration signals are framed and windowed. The Mel spectrogram is extracted for the i-th frame of the i-th IMF of the sound, and the wavelet transform time-frequency diagram is extracted for the vibration signal extracted from the i-th IMF of the i-th frame of the sound.

[0228] 3.1 Framing and windowing

[0229] The length of each component of the IMF component {c k (t)} obtained by decomposing the sound and vibration signals is set as N, the sampling frequency is F s = 10240 Hz, the frame length of each frame is w len = 128, the frame shift is inc = 64, and the overlapping part between adjacent two frames is overlap = w len -inc = 64.

[0230] Frame the sound and vibration signals with a sample length of N:

[0231] C n (t) = (N-W len ) / inc + 1

[0232] where {c k (t)} is the kth IMF component of the decomposed signal, W len is the frame length, inc is the frame shift, and overlap is the overlap length between two adjacent frames.

[0233] After the frame division, the windowing operation is needed. In this experiment, the Hanning window is used. The result of applying the window function to each frame is shown in the following figure.

[0234]

[0235] where {c k (t)} is the kth IMF component of the decomposed signal, W k,i (t) is the kth IMF component of the decomposed signal after windowing, and N is the length of each component of the decomposed signal {c k (t)}.

[0236] This step finally converts the input sound signal A k (t) into the output signal A k ′(t) and converts the input vibration signal V k (t) into the output signal V k ′(t).

[0237] 3.2 Extracting the Mel-spectrogram of the sound signal

[0238] The short-time Fourier transform is applied to each frame of the signal A k ′(t) to obtain the frequency representation, which is shown in the following equation:

[0239]

[0240] where A″(p, m) is the complex value of the pth frequency bin of the mth frame, and N is the window length of each frame.

[0241] The power spectrum of each frame is calculated as follows:

[0242] P(p, m) = |A″(p, m)| 2

[0243] where P(p, m) is the power spectral density of each frame.

[0244] The power spectrum is filtered by a set of Mel filters. The response of the filter can be represented as:

[0245]

[0246] where f(m) is the center frequency of the mth Mel filter, and k is the index on the frequency axis.

[0247] E m (n) = P(p, m) · H m (k)

[0248] M m (n) = log[E m (n)]

[0249] where M m (n) is the Mel-spectrum coefficient after log-transformation, and E m (n) is the energy of the nth frame signal after passing through the mth filter.

[0250] Taking n as the horizontal axis, and the feature vector M m (n) of the Mel-spectrum coefficient after log-transformation as the vertical axis, the energy distribution of the sound signal in different frequency segments is plotted, and the Mel-spectrum diagram is obtained. The Mel-spectrum diagram representation of K1 sound signals is

[0251] 3.3 Extracting wavelet transform time-frequency diagram for vibration signal

[0252]

[0253] where ψ(t) is the mother wavelet function, and ψ a,b (t) is the wavelet function after translation and scaling transformation, a is the scale factor (controls the scaling of the wavelet), and b is the translation factor (controls the position of the wavelet on the time axis)

[0254]

[0255] where W x (a, b) is the wavelet coefficient obtained by continuous wavelet transform, V k '(t) is the input vibration signal, is the complex conjugate of ψ a,b (t).

[0256] For a given scale a, the corresponding characteristic scale is defined as: W x (a, b) is calculated at multiple scales and different time positions b, and finally the horizontal coordinate is the translation factor b, the vertical coordinate is the characteristic scale , and the value of W x (a, b) is the point corresponding to the value, which is represented by the color depth, and the RGB time-frequency diagram is drawn, and the wavelet transform time-frequency diagram representation of K2 vibration signals is obtained as

[0257] Step 4: input the sound signal Mel spectrogram and the vibration signal wavelet transform time-frequency diagram into the CNN-transformer model of the cross attention fusion mechanism for training to obtain the optimal training parameters.

[0258] The CNN-transformer model of the cross attention fusion mechanism comprises: obtaining sound signal feature maps and vibration signal feature maps through CNN convolution pooling; then, using an encoder layer to convert the sound signal feature maps and the vibration signal feature maps into feature labels respectively, to extract sound signal features and vibration signals of different traffic targets; finally, using a cross attention mechanism to encode, fusing the sound signal features and the vibration signal features into new features, and using a Softmax function and a Dropout function to perform normalization and prevent overfitting.

[0259] 4.1 CNN convolution pooling

[0260] The K1 sound signal Mel spectrograms and the K2 vibration signal wavelet transform time-frequency diagrams are respectively subjected to CNN convolution pooling by using a CNN network to obtain the input of the encoder. In the convolution layer, the convolution kernel performs convolution on the output of the previous layer, and a nonlinear activation function is used to construct the output features. In the pooling layer, a large matrix is converted into a small matrix through data downsampling.

[0261] Specifically, the convolution process of the convolution layer is as follows:

[0262]

[0263] wherein, represents the weight of the kernel in the i-th filter in the l-th layer, wherein, represents the bias of the kernel in the i-th filter in the l-th layer, x l (j) represents the j-th local region in the l-th layer, represents the input of the j-th neuron in the l+1-th layer.

[0264] Specifically, the nonlinear activation function is as follows:

[0265] After the convolution operation, the numerical output of each convolution is subjected to nonlinear conversion by using an activation function. In this paper, the ReLU function is used as the activation function, and the formula of the function is as follows:

[0266]

[0267] wherein, represents the output value of the convolution operation, represents the activation value of

[0268] ​In particular, the pooling process of the pooling layer is:

[0269]

[0270] wherein, represents the value of the tth neuron in the ith feature in the lth layer after activation, (j-1)W P +1≤t≤W P , W P represents the width of the pooling region, represents the value of the ith neuron in the l+1th layer.

[0271] The sound signal Mel spectrogram and the vibration signal wavelet transform time-frequency diagram of the traffic target are convolved and pooled to output a feature map X n ∈R C*H*W ,

[0272] wherein, C represents the number of channels, H represents the image height, and W represents the image width

[0273] 4.2 The transformer encoder extracts the sound-vibration features.

[0274] The core structure of the transformer encoder is composed of a multi-head self-attention mechanism and a feedforward neural network. The self-attention mechanism captures global feature dependency relationships and improves the effectiveness of feature representation. The output results of K1 sound signal Mel spectrograms and K2 vibration signal wavelet transform time-frequency diagrams after CNN convolution and pooling are input into the transformer encoder, and the CNN output feature map is converted into a feature tag to extract sound-vibration signal features of different traffic targets (people, motor vehicles, and non-motor vehicles).

[0275] The feature map X n ∈R C*H*W is divided into blocks of the same size:

[0276] X∈RN L (a L *a L *C)

[0277] wherein, (a L *a L ) represents the resolution of each block, is the length of the input sequence, and C represents the number of channels;

[0278] A linear layer is used to project the blocks into a fixed dimension space in the entire transformer encoder layer, to encode the spatial information. A one-dimensional learnable position embedding is added in the block.

[0279] f0 = [x 1 A; x 2 A; ...x N A]+A position

[0280] Where A∈RN a*a*C*M For block embedding projection; A position ∈R N*M It is a one-dimensional learnable position embedding.

[0281] Therefore, the output of the i-th layer can be represented as:

[0282] f i =MHA(VN(f i-1 ))+f i-1 , i = 1…V

[0283] f i =MLP(VN(f i ))+f i , i = 1…V

[0284] Among them, f i and f′ i is the output of the i-th layer; VN(.) is the layer normalization; MHA is the multi-head self-attention of the V-layer; MLP is a multilayer perceptron, consisting of two linear layers with GRLU activation function; i is the identifier; V is the layer number of the transformer encoder.

[0285] The final output results are the sound features F audio and vibration characteristics F vibro .

[0286] 4.3 Feature Fusion of Cross-Attention Mechanism

[0287] Encoding is performed using a cross-attention mechanism to achieve feature information exchange, constructing an internal connection between sound features and vibration features, and realizing joint encoding of the two. The classification result is then output through a fully connected layer and a softmax layer. Sound feature F audio and vibration characteristics F vibro After cross-attention encoding, the features are fused into a new feature F. Sound feature F audio and vibration characteristics F vi b ro The vectors are mapped to two feature spaces via a linear layer, where q(Query) represents the query vector, k(Key) represents the key vector, and v(Value) represents the value vector. The results of the mapping of sound features and vibration features are represented as q... a k a v a and q v k v v v.

[0288] q a , k a , v a = Linear(F audio )

[0289] q v , k v , v v = Linear(F vibro )

[0290] The cross-attention score score is calculated by q a , k a , v a and q v , k v , v v , normalized by the Softmax function and the Dropout function, and overfitting is prevented:

[0291]

[0292] score' = Softmax(score)

[0293] score'' = Dropout(score')

[0294] F av = score''·v a

[0295] wherein score is the cross-attention score, λ c is a self-defined constant, score', score'' and F av are intermediate quantities.

[0296] F av is subjected to a full connection layer and normalization processing to obtain F:

[0297] F = Normalization(ω1F av + ω2)

[0298] wherein ω1 and ω2 are parameters to be learned.

[0299] For the learning parameters ω1 and ω2 in the model fused feature F = Normalization(ω1F av + ω2), when performing parameter iteration, first define the loss function L(ω) of the model, wherein ω is the learning parameter, and the specific learning process is as follows:

[0300] (1) Set the initial parameter values ω1(0), ω2(0) and define the learning rate αF .

[0301] (2) In each iteration process, the gradient of the loss function to the parameter is calculated And update the parameters:

[0302]

[0303]

[0304] Where n represents the number of iterations, The gradient of the loss function to the parameter, α F Indicates the learning rate.

[0305] (3) Using the adaptive learning rate algorithm Adam optimizer, further improve the parameter optimization efficiency.

[0306]

[0307]

[0308]

[0309]

[0310]

[0311]

[0312] Where m n And v n Respectively represent the moving average of the first and second moments of the gradient of the parameter ω1, m n ' and v n ' Respectively represent the moving average of the first and second moments of the gradient of the parameter ω2, β1 and β2 are hyperparameters, and ξ is a small constant to avoid zero denominator.

[0313] (4) Through cross-validation method, different parameter combinations are evaluated. For each parameter combination, record the classification accuracy of the model on the test set, select the best performance parameter combination as the final optimal parameter, and input the test set according to the optimal parameter to finally obtain the classification accuracy of the traffic target (person, motor vehicle, non-motor vehicle).

[0314] Step 5: Real-time acquisition of traffic target sound and vibration signals, processing the collected signals according to steps 1-3 to obtain real-time traffic target sound signal Mel spectrum and vibration signal wavelet transform time-frequency diagram, inputting the CNN-transformer model of cross-attention fusion mechanism in step 4, using the new traffic target fusion feature F, and obtaining the recognition result of the traffic target according to the optimal training parameter model.

Claims

1. A traffic target detection and recognition method based on acoustic vibration time-frequency features and cross-attention fusion mechanism, characterized in that, The specific steps are: Step 1: filtering and denoising preprocessing of the collected sound and vibration signals of different traffic targets; Step 2: variational mode decomposition of the preprocessed sound and vibration signals of different traffic targets, using scale spectrum segmentation method to obtain the optimal modal parameter value in the VMD decomposition process of the sound and vibration signals, and using the sum fuzzy entropy minimum value to obtain the optimal penalty factor, and finally decomposing the sound and vibration signals into multiple intrinsic mode functions (IMF); Step 3: feature extraction of the independent sound and vibration signals IMF after VMD decomposition, including Mel spectrogram and wavelet transform time-frequency diagram; Step 4: input the sound signal Mel spectrogram and vibration signal wavelet transform time-frequency diagram into the CNN-transformer model based on cross attention fusion mechanism for training, and obtain the CNN-transformer model based on cross attention fusion mechanism with optimal training parameters; Step 5: real-time acquisition of the sound and vibration signals of the traffic target, processing the collected signals according to steps 1-3 to obtain the sound signal Mel spectrogram and vibration signal wavelet transform time-frequency diagram of the real-time traffic target, inputting the cross attention fusion mechanism CNN-transformer model in step 4, using the new feature F of the traffic target fusion, and obtaining the recognition result of the traffic target according to the optimal training parameter model.

2. The traffic target detection and recognition method based on acoustic vibration time-frequency features and cross-attention fusion mechanism according to claim 1, characterized in that, The specific method for filtering and denoising preprocessing of the collected sound and vibration signals of different traffic targets is: Set the sampling rate of the sound and vibration sensors to be the same, align the time stamps of the sound and vibration sensors, and truncate the same length of effective signals; Utilizing NLMS filtering algorithm to reduce noise of sound signal x a (t) performing filtering and noise reduction preprocessing, and the specific process is: (1) Given initial filter weights ω a (0); (2) Calculate the output response of the sound signal: (3) Desired response d a (t) and output response y a (t) Compute error: e a (t) = y a (t) - d a (t) (4) Filter weight update: ω (t) = n (t) e (t), (1) a (n) is the weight of the filter in the sound signal processing, e a (t) is the error signal at time t, x a (t) is the input vector of the sound signal at time t, ||x a (t) || 2 represents the norm of the input signal, y a (t) is the sound signal after filtering and noise reduction at time t, d a (t) is the expected response at time t, (·) * represents the complex conjugate, μ is the step factor, taken 0<μ<2, α is the correction value, taken 0≤α≤1; When the filter weight ω a (n) stops updating and outputs the filtered and noise-reduced sound signal y a (t); The NLMS filtering algorithm is used for noise reduction on the vibration signal x v (t) filtering and noise reduction preprocessing, the specific process is: (1) Determine initial conditions, ω v (0) = 0 or give initial filter weights ω v (0) according to actual conditions; (2) Calculate the output response of the vibration signal: (3) Desired response d v (t) and output response y v (t) Compute error: e v (t) = y v (t) - d v (t) (4) Filter weight update: Where, ω v (n) represents the filter weights during vibration signal processing, e v (t) is the error signal at time t, x v (t) is the input vector of the vibration signal at time t, ||x v (t)|| 2 The norm of the input signal, y v (t) represents the filtered and noise-reduced vibration signal at time t, and d v (t) represents the expected response at time t, (·)* denotes complex conjugate, μ is the step size factor, which takes 0 < μ < 2, and α is the correction value, which takes 0 ≤ α ≤ 1; When the filter weights ω v (n) stop updating and output y v (t).

3. The traffic target detection and recognition method based on acoustic vibration time-frequency features and cross-attention fusion mechanism according to claim 2, characterized in that, The minimum optimization criterion function minJ is specifically: ω H (n + 1) x(t) = d(t) minJ = min || ω(n+1) - ω(n) || P wherein (·) H denotes the conjugate transpose, J is an optimization criterion function, ω(n) is a weight of the NLMS filter, p denotes a norm power term of the optimization criterion function, x(t) is x v (t) or x a (t), and d(t) is d a (t) or d v (t).

4. The traffic target detection and recognition method based on acoustic vibration time-frequency features and cross-attention fusion mechanism according to claim 1, characterized in that, The specific process of step 2 for variational mode decomposition is: Step 2.1: determine the optimal modal number K using scale spectrum segmentation method; Step 2.2: determine the optimal penalty factor a using sum fuzzy entropy, the specific method is: (1) Set the maximum value of the penalty factor a max , the minimum value a min , and the step size L; (2) calculating the modal component fuzzy entropy H(c k (t)) H(c k (t)) = -∫p(c k (t))log(p(c k (t)))dt where F s is the sampling rate of the sound and vibration sensors, N is the number of components of the decomposed sound and vibration signal {c k (t)} and p(c k (t)) is the probability of the modal component c k (t) at time point t, H(c k (t)) is the fuzzy entropy of the modal component c k (t). (3) Calculate the sum fuzzy entropy SFE where SFE is the sum fuzzy entropy, w k is the weight of the modal component c k (t), i.e. the center frequency of each modal component. (4) Let α = α min , find the SFE of VMD decomposition of K modal components when the penalty factor is α, when α ≤ α max , let α = α + L, and sequentially find the corresponding SFE; when α > α max , end the loop, and the abscissa corresponding to the minimum value of SFE is the optimal value of the penalty factor α; Step 2.3: according to the determined optimal penalty factor and optimal modal number, perform variational mode decomposition on the preprocessed sound and vibration signals of different traffic targets, which specifically includes: (1) The pre-processed traffic target sound vibration signal y(t) is analyzed by Hilbert transform to obtain the analysis signal, and the one-sided spectrum of the analysis signal is calculated. By multiplying the operator , the center band of y(t) is modulated to the corresponding base band: (2) The modal component uses its own estimated center frequency to perform frequency mixing through exponential correction, and gradually adjusts the center frequency to the corresponding bandwidth; (3) Introduce a Gaussian smoothing demodulation signal to estimate the bandwidth of the modal component frequency band, and obtain the constraint expression of the variational model through the square root of the norm gradient, which is: wherein is the partial derivative of t, δ(t) is an impulse function, K is the optimal modal decomposition number, {c k (t)} is the decomposed IMF component, {c k (t)} = {c1, c2, c3,..., c k}, {w k} is the center frequency of each modal component, {w k} = {w1, w2, w3,..., w k}; (4) Introduce the penalty factor a and the Lagrange multiplier operator lambda(t), and then obtain the Lagrange function: Where lambda(t) is the constraint enhancement; Perform time-frequency domain conversion on the Lagrange function to obtain the corresponding extreme value solution, and then obtain the modal component and center frequency expression: wherein is the Fourier transform of is a modal component formed by the VMD during the iteration process, is the power spectral centroid of are the Fourier transforms of the modal component, the original signal and the Lagrangian operator, respectively; By alternating direction multiplication, the {u k}, {w k}, λ are iteratively updated to the saddle point of the Lagrangian function, so as to obtain the optimal solution of the variational model constraint expression, and finally obtain the K modal components of the input signal y(t). Finally, the input sound signal y a (t) is converted into the output signal c ak (t), represented as A k (t); the input vibration signal y v (t) is converted into the output signal c vk (t), represented as V k (t).

5. The traffic target detection and recognition method based on acoustic vibration time-frequency features and cross-attention fusion mechanism according to claim 4, characterized in that, The specific process of determining the optimal modal number K using scale spectrum segmentation method is: Obtain the discrete time series of the sound and vibration signals of the traffic target after filtering and denoising preprocessing, which is: Wherein, Y(f) is a discrete time series of the traffic target sound vibration signal, y(t) is the preprocessed traffic target sound vibration signal, and f represents the signal frequency; The discrete scale spectrum of the discrete time series of the traffic target sound vibration signal is determined, specifically as follows: Wherein, L(f, s) is the discrete scale spectrum of Y(f), g(f, s) is a sampling Gaussian kernel function, s is a scale parameter, and S is the window size of the truncated filter; The segmentation boundary points are determined by detecting each local extreme value of L(f, s) and searching for the local minimum values on both sides, all segmentation boundary points are found in the scale spectrum, and they are divided into K continuous wave bands, so as to obtain the modal parameter K required by VMD.

6. The traffic target detection and recognition method based on acoustic vibration time-frequency features and cross-attention fusion mechanism according to claim 1, characterized in that, The specific method for feature extraction of the independent sound vibration signal IMF after VMD decomposition is as follows: The decomposed IMF components of the acoustic vibration signal {c k Each component of the long t) is set as N, and the sampling frequency is F s The frame length of each frame is w len The frame shift is inc, and the overlapping part between adjacent two frames is overlap = w len -inc; Frame the sound vibration signal with a sample length of N: C n (t) = (N - w len ) / inc + 1 wherein N is the number of components of the decomposed acoustic vibration signal {c k (t)}, w len is the length of each component, L is the frame length of each frame, inc is the frame shift, and overlap is the overlap between adjacent frames. After taking out a frame of signal and performing windowing operation, the window function is applied to each frame of signal as follows: wherein {c k (t)} is the i-th frame signal of the k-th IMF component after decomposition, C k,i (t) is the i-th frame signal of the k-th IMF component after decomposition with windowing, n is the index of sample points in the window, The Mel spectrogram of the framed and windowed sound signal is extracted, specifically as follows: for each frame of the windowed sound signal A k '(t) performing a short-time Fourier transform to obtain a spectral representation, according to the formula Wherein, A''(p, m) represents the complex value of the pth frequency point of the mth frame, and N is the window length of each frame; The power spectrum of each frame of Fourier transform result is calculated as follows: P(p,m) = |A"(p,m| 2 Wherein, P(p, m) is the power spectral density of each frame; The power spectrum is filtered through a set of Mel filters, and the response of the filter is represented as follows: Wherein, f(m) is the center frequency of the mth Mel filter, and k is the index on the frequency axis; E m (n) = P(p, m) · H m (k) M m (n) = log [E m (n)] where M m (n) is the log transformed Mel spectral coefficient, E m (n) is the energy of the nth frame of signal after passing through the mth filter. with n as the horizontal axis, the feature vector M of the log-transformed Mel-spectral coefficients m (n) as the vertical axis, the energy distribution of the sound signal over different frequency bins is plotted, resulting in a Mel-spectrogram, resulting in K1 Mel-spectrogram representations of the sound signal The wavelet transform time-frequency diagram of the framed and windowed vibration signal is extracted where ψ(t) is the mother wavelet function, ψ a,b (t) is the wavelet function after translation and scaling transformation, a is the scale factor, and b is the translation factor. The wavelet coefficient is obtained by continuous wavelet transform: where W x (a,b) are the wavelet coefficients obtained from the continuous wavelet transform, V k (t) is the input vibration signal, is the complex conjugate of ψ a,b (t). For a given scale a, its corresponding characteristic scale is defined as: W is computed at multiple scales and different time positions b x (a, b), finally with the translation factor b as the horizontal coordinate and the characteristic scale as the vertical coordinate, W x (a, b) as points corresponding values, the RGB time-frequency graph is drawn, and the result is that the wavelet transform time-frequency graph of K2 vibration signals is represented as 7. The traffic target detection and recognition method based on acoustic vibration time-frequency features and cross-attention fusion mechanism according to claim 1, characterized in that, The processing process of the CNN-transformer model based on the cross-attention fusion mechanism for the sound signal Mel spectrogram and the vibration signal wavelet transform time-frequency diagram is as follows: K1 Mel spectrograms of sound signals and K2 wavelet transform time-frequency diagrams of vibration signals are respectively subjected to CNN convolution and pooling by using the CNN network; The sound-vibration features are extracted by using a transformer encoder, a core structure of the transformer encoder is composed of a multi-head self-attention mechanism and a feedforward neural network, the global feature dependency is captured through the self-attention mechanism, the output results of K1 Mel spectrograms of sound signals and K2 wavelet transform time-frequency diagrams of vibration signals after CNN convolution and pooling are input into the transformer encoder, the feature map X n ∈R C*H*W are converted into feature labels, and the sound-vibration signal features of different traffic targets are extracted. The cross-attention mechanism is used for encoding to realize feature information exchange, build the internal relationship between sound features and vibration features, realize joint encoding of the two, and output the classification result through the full connection layer and Softmax.

8. The traffic target detection and recognition method based on acoustic vibration time-frequency features and cross-attention fusion mechanism according to claim 7, characterized in that, The specific process of extracting sound vibration features by using the transformer encoder is as follows: Feature map X n ∈ R C*H*W are divided into blocks of equal size: X ∈ RN L (a L *a L *C) wherein (a L *a L ) represents the resolution of each block, is the length of the input sequence, and C represents the number of channels. A linear layer is used to project the block into a fixed dimension space of the entire transformer encoder layer, and a one-dimensional learnable position embedding is added in the block to encode the spatial information: f0 = [x 1 A; x 2 a;... x N A] + A position where A ∈ RN a*a*C*M is the block embedding projection; A position ∈ R N*M is a one-dimensional learnable positional embedding; Therefore, the output of the i th layer of the transformer encoder is represented as follows: f′ i = MHA(VN(f i-1 ))+f i-1 ,i = 1...V f i = MLP(VN(f′ i ))+f′ i ,i=1…V where f i and f' i are the outputs of the i-th layer; VN(.) is the layer normalization; MHA is the V-layer multi-head self-attention; MLP is a multi-layer perceptron consisting of 2 linear layers activated by GRLU activation function; i is an identifier; V is the number of layers of the transformer encoder; The final output results are sound features F audio and vibration features F vibro respectively.

9. The traffic target detection and recognition method based on acoustic vibration time-frequency features and cross-attention fusion mechanism according to claim 7, characterized in that, The specific process of realizing feature information exchange by using the cross-attention mechanism for encoding to obtain the fused features is as follows: sound feature F audio and vibration feature F vibro The linear layer mapped to two feature spaces, sound feature and vibration feature, the results of the mapping are denoted as q a ,k a ,v a and q v ,k v ,v v : q a ,k a ,v a = Linear(F audio ) q v ,k v ,v v = Linear(F vibro ) Wherein, q represents a query vector, k represents a key vector, v represents a value vector, and the subscripts a and v represent sound features and vibration features respectively; Cross attention scores score are calculated by q a ,k a ,v a and q v ,k v ,v v for feature fusion, and normalized and overfitting prevention processing are performed by using a Softmax function and a Dropout function. score' = Softmax(score) score'' = Dropout(score') F av = score" · v a wherein score is the cross-attention score, λ c is a custom constant, score', score" and F av are intermediate quantities, F av After passing through the fully connected layer and normalization, F is obtained: F = Normalization(ω1F + ω2) av + ω2) Wherein, ω1 and ω2 are parameters to be learned.

Citation Information

Patent Citations

  • Vehicle feature recognition algorithm based on real-time coding

    CN102982802A

  • Power transformer fault voiceprint diagnosis method based on VMD-JS divergence

    CN116884432A