Security early warning monitoring method, device and system for rail transit station
Through high sensitivity microphone array and improved ECAPA-TDNN model, combined with multimodal data enhancement and dynamic attention statistics pooling, the problems of high missed detection rate and response delay in rail transit station security monitoring are solved, real-time early warning and efficient emergency response in complex scenarios are achieved.
Patent Information
- Application Number
- CN202510482919.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-07-04
AI Technical Summary
There are problems of high missed detection rate, delayed response and environmental interference in the security monitoring of rail transit stations. The existing voiceprint recognition technology lacks the discrimination of features and poor robustness in complex scenarios.
A high-sensitivity microphone array is used to collect multi-channel audio stream signals, combined with segmented noise reduction, standardization and multi-modal data enhancement processing, and feature mapping and classification are used to use the improved ECAPA-TDNN model to perform feature mapping and classification. Through multi-scale convolution, channel attention mechanism and dynamic attention statistics pooling, feature expression and classification are optimized by combining angle interval loss function.
It significantly improves the expression ability and robustness of voiceprint characteristics, realizes real-time detection and early warning of abnormal sounds, and improves the security monitoring efficiency and emergency response capabilities of rail transit stations.
Smart Images

Figure CN120260610A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical fields of voiceprint recognition and intelligent security, and in particular, to a security warning monitoring method, device, and system for rail transit stations. Background Art
[0002] As a public area with a large flow of people, the security monitoring of rail transit stations relies on traditional video surveillance and manual inspections, and there are the following problems: (1) High missed detection rate: Video surveillance cannot cover abnormal events (such as sudden screams, item collisions, etc.) in complex acoustic environments; (2) Response delay: Manual inspections are difficult to handle emergencies in a timely manner, resulting in a lag in emergency response; (3) Environmental interference: Station background noise (such as broadcasts, conversations of the crowd) seriously interferes with the extraction of sound features.
[0003] Based on the above problems, voiceprint recognition technology (such as the GMM model based on MFCC) is usually used for station monitoring at present, but there are still defects such as insufficient feature discriminability and poor robustness in complex scenarios. Summary of the Invention
[0004] The purpose of the present application is to provide a security warning monitoring method, device, and system for rail transit stations, which can improve the classification ability of the audio feature discrimination ability of the station in complex scenarios, and thus improve the accuracy of security warning for the station.
[0005] In a first aspect, the present application provides a security warning monitoring method for rail transit stations, and the method includes: collecting multi-channel audio stream signals in the station through a high-sensitivity microphone array; performing first preprocessing and second preprocessing on the multi-channel audio stream signals to obtain Mel spectrogram features; the first preprocessing includes: segmental noise reduction and normalization processing; the second preprocessing includes: multi-modal data augmentation processing; inputting the Mel spectrogram features into a preset voiceprint recognition model, and performing feature mapping and classification on the Mel spectrogram features through the voiceprint recognition model to output a classification feature result corresponding to the Mel spectrogram features; wherein, the voiceprint recognition model includes: an improved ECAPA-TDNN model based on multi-scale convolution, channel attention mechanism, dynamic attention statistical pooling, and angular margin loss function; performing threshold determination based on the classification feature result to determine whether to perform an alarm emergency response.
[0006] Further, the steps of performing the first preprocessing and the second preprocessing on the multi-channel audio stream signal to obtain the Mel spectral features include: performing segmented noise reduction and normalization processing on the multi-channel audio stream signal to obtain initial spectral features; performing multi-modal data augmentation processing on the initial spectral features to obtain Mel spectral features; wherein, the multi-modal data augmentation processing includes: additive noise injection processing, speed perturbation processing, and spectral masking processing; the additive noise injection processing includes: superimposing background noise on the clean audio according to the signal-to-noise ratio; the speed perturbation processing includes: changing the audio duration by resampling to generate speed perturbation samples; the spectral masking processing includes: randomly masking frequency bands and time domain regions on the initial spectral features.
[0007] Further, the above improved ECAPA-TDNN model includes: a multi-scale convolution module, a channel attention mechanism module, a dynamic attention statistical pooling module, and a linear compression module; the steps of inputting the Mel spectral features into a preset speaker recognition model, performing feature mapping and classification on the Mel spectral features through the speaker recognition model, and outputting the classification feature results corresponding to the Mel spectral features include: dividing the Mel spectral features into multiple sub-features through the multi-scale convolution module, performing progressive sub-feature accumulation and convolution on the multiple sub-features to obtain a convolutional feature map; performing global average pooling on the convolutional feature map through the channel attention mechanism module, and processing it through two linear transformations and an activation function to obtain channel weights; calculating the attention weights in the time dimension based on the feature vectors affected by the channel weights through the dynamic attention statistical pooling module, and aggregating the mean and standard deviation features based on the attention weights; linearly compressing the mean and standard deviation features through the linear compression module to obtain the classification feature results corresponding to the Mel spectral features.
[0008] Further, the steps of performing global average pooling on the convolutional feature map through the channel attention mechanism module, and processing it through two linear transformations and an activation function to obtain channel weights include: dynamically adjusting the channel weights w through the Squeeze-and-Excitation mechanism:
[0009] w = σ(W2 · ReLU(W1 · GAP(H out )));
[0010] wherein, GAP is the global average pooling process, and H out represents the convolutional feature map; W1 and W2 are the weights of the fully connected layers for linear transformation; ReLU is the activation function for introducing non-linearity; σ is the Softmax activation function for mapping the output to a specific range;
[0011] The steps of calculating the attention weights in the time dimension based on the feature vectors under the influence of channel weights through the dynamic attention statistical pooling module and aggregating the mean and standard deviation features based on the attention weights include: calculating the attention weights α in the time dimension according to the following formula t , the mean μ and the standard deviation feature σ:
[0012] α t = Softmax(v T tanh(Wh t + b));
[0013]
[0014] where v, W, and b are learnable parameters for calculating the attention weights; tanh is an activation function for non-linear transformation; h t represents the feature vector of the model at time step t; Softmax represents the activation function;
[0015] The steps of linearly compressing the mean and standard deviation features through the linear compression module to obtain the classification feature results corresponding to the Mel spectrum features include: determining the classification feature results z corresponding to the Mel spectrum features according to the following formula:
[0016] z = W liner ·z pool + b liner ; z pool = Concat(μ, σ); z ∈ R D ;
[0017] where W liner ∈ R D×2C , b liner ∈ R D are learnable parameters.
[0018] Furthermore, the above steps of making a threshold determination based on the classification feature results to determine whether to initiate an alarm emergency response include: calculating the cosine similarity scores between the classification feature results and the classification weights corresponding to multiple preset categories; taking the preset category corresponding to the highest cosine similarity score as the target category corresponding to the classification feature results; determining whether the highest cosine similarity score exceeds the corresponding preset threshold; if so, issuing a warning in the warning method corresponding to the target category.
[0019] Furthermore, the above preset threshold is determined by the ROC curve of the test set and satisfies the following formula:
[0020] τ = argmax(TPR(τ) - FPR(τ))
[0021] where TPR(τ) is the true positive rate and FPR(τ) is the false positive rate.
[0022] Further, during the training process of the above improved ECAPA-TDNN model, the cosine annealing strategy is adopted to dynamically adjust the learning rate, and the angular margin loss function is used for model convergence;
[0023] The learning rate formula is as follows:
[0024]
[0025] where η t is the learning rate at the training step t, η min and η max are the minimum and maximum values of the learning rate respectively, and T is the total number of training steps or the cycle length;
[0026] The angular margin loss function formula is as follows:
[0027]
[0028] where θ yi is the included angle between the sample and the true class weight w i ; s is the feature scaling factor used to amplify the inter-class difference; the angular margin m is used to increase the inter-class difference; cos(θ yi +m) means adding an interval m to the included angle of the true class to make the classification boundary more strict; represents the sum of the scores of other classes.
[0029] Further, the above improved ECAPA-TDNN model adopts multi-GPU parallel training, and the AllReduce gradient synchronization strategy is used to ensure consistent update of the model parameters on each GPU.
[0030] In a second aspect, the present application also provides a security warning monitoring device for a rail transit station. The device includes: an audio acquisition module for collecting multi-channel audio stream signals in the station through a high-sensitivity microphone array; a preprocessing module for segmenting, denoising, and normalizing the multi-channel audio stream signals and extracting Mel spectrum features; a model prediction module for inputting the Mel spectrum features into a preset voiceprint recognition model, performing feature mapping and classification on the Mel spectrum features through the voiceprint recognition model, and outputting a classification feature result corresponding to the Mel spectrum features; wherein, the voiceprint recognition model includes: an improved ECAPA-TDNN model based on multi-scale convolution, channel attention mechanism, dynamic attention statistical pooling, and angular margin loss function; a decision warning module for making a threshold determination based on the classification feature result to determine whether to perform an alarm emergency response.
[0031] In a third aspect, the present application also provides a security warning monitoring system for rail transit stations. The system includes a high-sensitivity microphone array and a server connected by communication. The server is used to execute the method described in the first aspect.
[0032] In the security warning monitoring method, device and system for rail transit stations provided by the present application, a multi-channel audio stream in the station is collected through a high-sensitivity microphone array, and after being processed by a preprocessing module for audio segmentation, noise reduction, normalization, and extraction of Mel spectrum features; the improved ECAPA-TDNN model significantly improves the expression ability and robustness of voiceprint features through multi-scale convolution and channel attention mechanism, combined with multi-modal data augmentation strategies (such as noise injection, speed perturbation, spectrum masking) and dynamic attention statistical pooling; the introduction of the angular interval loss function further enhances the intra-class compactness and inter-class separability of voiceprint features; real-time detection and warning of abnormal sounds (such as screams, explosions) are realized through a threshold determination and alarm mechanism, effectively solving problems such as high false negative rate, response delay and environmental interference of traditional video monitoring, and improving the security monitoring efficiency and emergency response ability of rail transit stations. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] In order to more clearly illustrate the specific embodiments of the present application or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0034] Figure 1 It is a flowchart of a security warning monitoring method for a rail transit station provided by an embodiment of the present application;
[0035] Figure 2 It is a schematic flowchart of a model prediction process provided by an embodiment of the present application;
[0036] Figure 3 It is a structural block diagram of a security warning monitoring device for a rail transit station provided by an embodiment of the present application;
[0037] Figure 4 It is a structural block diagram of a security warning monitoring system for a rail transit station provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0038] The technical solution of the present application will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.
[0039] Existing voiceprint recognition technologies (such as the GMM model based on MFCC) have defects such as insufficient feature discriminability and poor robustness in complex scenarios. The ECAPA-TDNN model significantly improves the expression ability of voiceprint features by introducing multi-scale convolution and channel attention mechanisms, but its application in the rail transit scenario is not yet mature.
[0040] Based on this, the embodiments of the present application provide a security warning and monitoring method, device and system for rail transit stations, and solve the above problems through the following innovative designs:
[0041] Multi-modal data augmentation strategy: Combining noise injection, speed perturbation and spectral masking to enhance the adaptability of the model to complex acoustic environments.
[0042] Dynamic attention statistical pooling: Optimize the feature aggregation method to improve the ability to capture key segments of abnormal sounds.
[0043] Angular interval loss function: By introducing an angular penalty term, enhance the intra-class compactness and inter-class separability of voiceprint features.
[0044] For the convenience of understanding this embodiment, first, a security warning and monitoring method for rail transit stations disclosed in the embodiments of the present application will be introduced in detail.
[0045] Figure 1 The following is a flowchart of a security warning and monitoring method for rail transit stations provided in the embodiments of the present application. The method specifically includes the following steps:
[0046] Step S102, collect multi-channel audio stream signals in the station through a high-sensitivity microphone array;
[0047] Step S104, perform first preprocessing and second preprocessing on the multi-channel audio stream signals to obtain Mel spectrum features; the first preprocessing includes: segmental noise reduction and normalization processing; the second preprocessing includes: multi-modal data augmentation processing;
[0048] The segmental noise reduction and normalization processing includes: dividing the original multi-channel audio stream signal into segments with a fixed duration, and then performing normalization processing. This process helps to eliminate the dimensionality impact in the data, so that different data sets can be compared and analyzed on the same scale.
[0049] Multimodal data augmentation processing may include: additive noise injection processing, speed perturbation processing, and spectral masking processing; the additive noise injection processing includes: superimposing background noise on the clean audio according to the signal-to-noise ratio; the speed perturbation processing includes: changing the audio duration by resampling to generate speed perturbation samples; the spectral masking processing includes: randomly masking frequency bands and time-domain regions on the initial spectral features.
[0050] After segmented noise reduction, normalization processing, and multimodal data augmentation processing, Mel spectral features are obtained; the Mel spectral features are a method of representing audio signals on the Mel scale and are mainly used in fields such as speech processing and audio recognition.
[0051] Step S106, input the Mel spectral features into a preset voiceprint recognition model, perform feature mapping and classification on the Mel spectral features through the voiceprint recognition model, and output the classification feature results corresponding to the Mel spectral features; among them, the voiceprint recognition model includes: an improved ECAPA-TDNN model based on multi-scale convolution, channel attention mechanism, dynamic attention statistical pooling, and angular margin loss function;
[0052] The training and application process of the improved ECAPA-TDNN model will be described in detail later.
[0053] Step S108, perform threshold determination based on the classification feature results to determine whether to trigger an alarm emergency response.
[0054] In specific implementation, it can be determined whether to trigger an emergency response by calculating the similarity between the classification feature results and multiple preset classifications and comparing with the corresponding thresholds.
[0055] In the rail transit station security warning and monitoring method provided by the embodiments of the present application, multi-channel audio streams in the station are collected through a high-sensitivity microphone array, and after being processed by a preprocessing module for audio segmentation, noise reduction, normalization, and Mel spectral features are extracted; the improved ECAPA-TDNN model significantly improves the expression ability and robustness of voiceprint features through multi-scale convolution and channel attention mechanism, combined with multimodal data augmentation strategies (such as noise injection, speed perturbation, spectral masking) and dynamic attention statistical pooling; the introduction of the angular margin loss function further enhances the intra-class compactness and inter-class separability of voiceprint features; through threshold determination and alarm mechanism, real-time detection and warning of abnormal sounds (such as screams, explosions) are realized, effectively solving problems such as high missed detection rate, response delay, and environmental interference of traditional video monitoring, and improving the security monitoring efficiency and emergency response ability of rail transit stations.
[0056] Another security warning and monitoring method for rail transit stations is also provided in an embodiment of the present application, which is implemented on the basis of the above embodiment; the feature extraction process, model prediction process and emergency response process are mainly described in this embodiment.
[0057] The above step S104, the step of performing the first preprocessing and the second preprocessing on the multi-channel audio stream signal to obtain the Mel spectrum feature, specifically includes:
[0058] (1) Segment the multi-channel audio stream signal for noise reduction and normalization processing to obtain the initial spectrum feature;
[0059] The original audio signal x(t) is segmented into segments with a fixed duration T and normalized. This process helps to eliminate the influence of dimensions in the data, enabling different data sets to be compared and analyzed on the same scale. The normalization processing formula is as follows:
[0060] μ = E[x(t)];
[0061] Among them, is the signal after normalization, x(t) is the original signal, μ is the mean of the signal, and σ is the standard deviation of the signal.
[0062] (2) Perform multi-modal data augmentation processing on the initial spectrum feature to obtain the Mel spectrum feature; among them, the multi-modal data augmentation processing includes: additive noise injection processing, speed perturbation processing and spectrum masking processing;
[0063] 1) The additive noise injection processing includes: superimposing the background noise n(t) on the clean audio according to the signal-to-noise ratio SNR;
[0064] x noisy (t) = x(t) + α·n(t); α = 10 -SNr / 20 ;
[0065] Among them, x noisy (t) is the audio signal after adding noise. The higher the signal-to-noise ratio, the smaller the scaling factor α, which means the smaller the influence of the noise. Conversely, the lower the signal-to-noise ratio, the larger the scaling factor α, and the greater the influence of the noise.
[0066] 2) The speed perturbation processing includes: changing the audio duration by resampling to generate the speed perturbation sample x speed (t);
[0067] 3) The spectrum masking processing includes: randomly masking the frequency band and time domain region on the initial spectrum feature to enhance the robustness of the model.
[0068] Furthermore, the above improved ECAPA-TDNN model includes: a multi-scale convolution module, a channel attention mechanism module, a dynamic attention statistical pooling module, and a linear compression module;
[0069] In step S106 above, the step of inputting the Mel spectrogram features into a preset voiceprint recognition model, performing feature mapping and classification on the Mel spectrogram features through the voiceprint recognition model, and outputting the classification feature result corresponding to the Mel spectrogram features specifically includes the following steps. See Figure 2 the process schematic diagram shown:
[0070] (1) The multi-scale convolution module divides the Mel spectrogram features into multiple sub-features, performs progressive sub-feature accumulation and convolution on the multiple sub-features, and obtains a convolution feature map;
[0071] The multi-scale convolution module divides the input feature, that is, the Mel spectrogram feature H∈R c×t into S sub-features {H1,...,H S}, where the dimension of each sub-feature map is H∈R (c / s)×t , performs progressive sub-feature accumulation and convolution, and obtains a convolution feature map H out as follows:
[0072] H out =Concat(Conv1D(H1),Conv1D(H1+H2),...,Conv1D(H1+
[0073] H2+...+H s ))。
[0074] (2) The channel attention mechanism module performs global average pooling on the convolution feature map, and processes it through two linear transformations and an activation function to obtain channel weights; the specific process is as follows:
[0075] Dynamically adjust the channel weights w through the Squeeze-and-Excitation mechanism:
[0076] w=σ(W2·ReLU(W1·GAP(H out )));
[0077] Among them, GAP is global average pooling processing, H oit represents the convolution feature map; W1 and W2 are the weights of the fully connected layers for linear transformation; ReLU is the activation function, ReLU=max(0,x) ; is used to introduce non-linearity; σ is the Softmax activation function, which is used to map the output to a specific range.
[0078] (3) Based on the feature vectors under the influence of channel weights through the dynamic attention statistical pooling module, calculate the attention weights in the time dimension, and aggregate the mean and standard deviation features based on the attention weights;
[0079] Calculate the attention weight α in the time dimension according to the following formula t , the mean μ and the standard deviation feature σ:
[0080] α t = Softmax(v T tanh(Wh t + b));
[0081]
[0082] where v, W, and b are learnable parameters for calculating the attention weights; tanh is an activation function for non-linear transformation; h t represents the feature vector of the model at time step t; Softmax represents the activation function.
[0083] (4) Linearly compress the mean and standard deviation features through the linear compression module to obtain the classification feature results corresponding to the Mel spectrogram features.
[0084] Determine the classification feature result z corresponding to the Mel spectrogram feature according to the following formula:
[0085] z = W liner · z pool + b liner ; z pool = Concat(μ, σ); z ∈ R D ;
[0086] where W liner ∈ R D×2C , b liner ∈ R D are learnable parameters; z pool represents the two-dimensional feature generated by the dynamic attention statistical pooling; R D represents the one-dimensional feature space.
[0087] Furthermore, the above step S108, the step of determining whether to perform an alarm emergency response based on the classification feature results, specifically includes:
[0088] (1) Calculate the cosine similarity scores between the classification feature results and the classification weights corresponding to multiple preset categories;
[0089] After preprocessing the input audio to generate the Mel spectrogram feature X ∈ R F×TAfter that (where F is the frequency dimension and T is the time dimension), one-dimensional features z ∈ R are extracted from the Mel spectrogram features through an improved ECAPA-TDNN model D , where D is the dimension of the features; then the similarity scores with preset categories (such as normal, scream, explosion) are calculated through the following formula:
[0090]
[0091] where score i represents the cosine similarity between the extracted feature z and the weight vector w i of the preset category i; z·w i represents the dot product of the feature vector z and the weight vector w i ; ‖z‖ and ‖w i ‖ are the norms of the feature vector z and the weight vector w i respectively.
[0092] (2) Take the preset category corresponding to the highest cosine similarity score as the target category corresponding to the classification feature result;
[0093] (3) Determine whether the highest cosine similarity score exceeds the corresponding preset threshold; the above preset threshold is determined by the ROC curve of the test set and satisfies the following formula:
[0094] τ = argmax(TPR(τ) - FPR(τ))
[0095] where TPR(τ) represents the true positive rate at the threshold τ, that is, the proportion of correctly detected positive examples. FPR(τ) represents the false positive rate at the threshold τ, that is, the proportion of incorrectly detected negative examples. By plotting TPR and FPR at different thresholds, the optimal threshold τ is selected to balance the detection accuracy and false alarm rate.
[0096] (4) If so, give an alarm in the warning method corresponding to the target category. The text form of the warning is, for example: "Ambient sound", "There is a scream", "There is an explosion".
[0097] The training and optimization process of this model is introduced below:
[0098] The data collected at the station is divided into a training set and a test set according to 7:3, and the model is trained. During the training process of the improved ECAPA-TDNN model, the cosine annealing strategy is adopted to dynamically adjust the learning rate, and the formula is:
[0099]
[0100] where η t is the learning rate at the training step t, η min and η maxare the minimum and maximum values of the learning rate, respectively, and T is the total number of training steps or the cycle length. The cosine function is used to smoothly adjust the learning rate.
[0101] This method uses a larger learning rate at the beginning of training to converge quickly, and then uses a smaller learning rate for fine-tuning at the later stage of training. By periodically adjusting the learning rate, it can help the model jump out of local optima and find a better global optimum solution.
[0102] During the training process, an angular margin loss function is also used for model convergence, that is, an angular margin m is introduced into the Softmax loss to force similar samples to be more compact in the feature space. The angular margin loss function has the following formula:
[0103]
[0104] where θ yi is the angle between the sample and the true class weight w i ; s is the feature scaling factor used to amplify the difference between classes; the angular margin m is used to increase the difference between classes; cos(θ yi +m) means adding a margin m to the angle of the true class, making the classification boundary more strict; represents the sum of the scores of other classes.
[0105] The angular margin loss function minimizes this loss during training by taking the negative logarithm, so that the score of the true class is as high as possible and the scores of other classes are as low as possible.
[0106] In addition, the above improved ECAPA-TDNN model can also adopt multi-GPU parallel training, and ensure the consistency of model parameter updates on each GPU through the AllReduce gradient synchronization strategy. This can improve the model training speed and accuracy.
[0107] The security warning and monitoring method for rail transit stations provided by the embodiments of the present application collects multi-channel audio streams in the station through a high-sensitivity microphone array, performs audio segmentation, noise reduction, and normalization through a preprocessing module, and extracts Mel spectrum features; then uses an improved ECAPA-TDNN model to extract and classify the Mel spectrum features to obtain a classification feature result; this model significantly improves the expression ability and robustness of voiceprint features through multi-scale convolution and channel attention mechanism, combined with multi-modal data augmentation strategies (such as noise injection, speed perturbation, spectral masking) and dynamic attention statistical pooling; introducing an angular margin loss function further enhances the intra-class compactness and inter-class separability of voiceprint features; finally, real-time detection and warning of abnormal sounds (such as screams, explosions) are realized through a threshold determination and alarm mechanism, effectively solving the problems of high false negative rate, response delay, and environmental interference of traditional video monitoring, and improving the security monitoring efficiency and emergency response ability of rail transit stations.
[0108] Based on the above method embodiments, the embodiments of the present application also provide a security warning and monitoring device for rail transit stations. Refer to Figure 3 As shown, the device includes: an audio acquisition module 302, configured to collect multi-channel audio stream signals in the station through a high-sensitivity microphone array; a preprocessing module 304: configured to perform segmentation, noise reduction, and normalization processing on the multi-channel audio stream signals, and extract Mel spectrum features; a model prediction module 306, configured to input the Mel spectrum features into a preset voiceprint recognition model, perform feature mapping and classification on the Mel spectrum features through the voiceprint recognition model, and output a classification feature result corresponding to the Mel spectrum features; wherein, the voiceprint recognition model includes: an improved ECAPA-TDNN model based on multi-scale convolution, channel attention mechanism, dynamic attention statistical pooling, and angular margin loss function; a decision-making and warning module 308, configured to perform threshold determination based on the classification feature result to determine whether to perform an alarm emergency response.
[0109] Further, the above preprocessing module 304: is configured to perform segmentation noise reduction and normalization processing on the multi-channel audio stream signals to obtain initial spectrum features; perform multi-modal data augmentation processing on the initial spectrum features to obtain Mel spectrum features; wherein, the multi-modal data augmentation processing includes: additive noise injection processing, speed perturbation processing, and spectral masking processing; the additive noise injection processing includes: superimposing background noise on the clean audio according to the signal-to-noise ratio; the speed perturbation processing includes: changing the audio duration through resampling to generate speed perturbation samples; the spectral masking processing includes: randomly masking frequency bands and time domain regions on the initial spectrum features.
[0110] Furthermore, the above improved ECAPA-TDNN model includes: a multi-scale convolution module, a channel attention mechanism module, a dynamic attention statistical pooling module, and a linear compression module; a model prediction module 306, which is used to divide the Mel spectrogram features into multiple sub-features through the multi-scale convolution module, perform progressive sub-feature accumulation and convolution on the multiple sub-features to obtain a convolution feature map; perform global average pooling on the convolution feature map through the channel attention mechanism module, and process it through two linear transformations and an activation function to obtain channel weights; calculate the attention weights in the time dimension based on the feature vectors affected by the channel weights through the dynamic attention statistical pooling module, and aggregate the mean and standard deviation features based on the attention weights; linearly compress the mean and standard deviation features through the linear compression module to obtain the classification feature results corresponding to the Mel spectrogram features.
[0111] Furthermore, the above model prediction module 306 is used to dynamically adjust the channel weight w through the Squeeze-and-Excitation mechanism:
[0112] w = σ(W2 · ReLU(W1 · GAP(H out )));
[0113] where GAP is the global average pooling process, and H out represents the convolution feature map; W1 and W2 are the weights of the fully connected layers for linear transformation; ReLU is the activation function for introducing non-linearity; σ is the Softmax activation function for mapping the output to a specific range;
[0114] The model prediction module 306 is used to calculate the attention weight α in the time dimension according to the following formula t , the mean μ and the standard deviation feature σ:
[0115] α t = Softmax(v T tanh(Wh t + b));
[0116]
[0117] where v, W, and b are learnable parameters for calculating the attention weights; tanh is the activation function for non-linear transformation; h t represents the feature vector of the model at time step t; Softmax represents the activation function;
[0118] The model prediction module 306 is used to determine the classification feature result z corresponding to the Mel spectrogram feature according to the following formula:
[0119] z = W liner · zpool +b liner ; z pool = Concat(μ, σ); z ∈ R D ;
[0120] Among them, W liner ∈ R D×2C , b liner ∈ R D is a learnable parameter.
[0121] Furthermore, the above decision warning module 308 is used to calculate the cosine similarity score between the classification feature result and the classification weights corresponding to multiple preset categories; take the preset category corresponding to the highest cosine similarity score as the target category corresponding to the classification feature result; judge whether the highest cosine similarity score exceeds the corresponding preset threshold; if so, issue a warning in the warning method corresponding to the target category.
[0122] Furthermore, the above preset threshold is determined by the ROC curve of the test set and satisfies the following formula:
[0123] τ = argmax(TPR(τ) - FPR(τ))
[0124] Among them, TPR(τ) is the true positive rate and FPR(τ) is the false positive rate.
[0125] Furthermore, during the training process of the above improved ECAPA-TDNN model, the cosine annealing strategy is adopted to dynamically adjust the learning rate, and the angular margin loss function is used for model convergence; the learning rate formula is:
[0126]
[0127] Among them, η t is the learning rate at the training step t, η min and η max are respectively the minimum and maximum values of the learning rate, and T is the total number of training steps or the cycle length;
[0128] The formula of the angular margin loss function is as follows:
[0129]
[0130] Among them, θ yi is the included angle between the sample and the true class weight w i ; s is the feature scaling factor used to amplify the difference between classes; the angular margin m is used to increase the difference between classes; cos(θ yi + m) means adding an interval m to the included angle of the true class to make the classification boundary more strict; represents the sum of the scores of other classes.
[0131] Furthermore, the above improved ECAPA-TDNN model adopts multi-GPU parallel training and ensures the consistent update of model parameters on each GPU through the AllReduce gradient synchronization strategy.
[0132] The device provided in the embodiment of the present application has the same implementation principle and technical effects as those in the foregoing method embodiment. For the sake of brief description, for the parts not mentioned in the device embodiment, reference may be made to the corresponding content in the foregoing method embodiment.
[0133] Based on the foregoing method embodiment, the embodiment of the present application further provides a security warning monitoring system for a rail transit station. Refer to Figure 4 As shown, the system includes a high-sensitivity microphone array 42 and a server 44 connected by communication; a security warning monitoring device 442 for a rail transit station as described in the device embodiment is provided in the server 44 and is used to execute the method as described in the method embodiment.
[0134] The system provided in the embodiment of the present application has the same implementation principle and technical effects as those in the foregoing method embodiment. For the sake of brief description, for the parts not mentioned in the system embodiment, reference may be made to the corresponding content in the foregoing method embodiment.
[0135] In the description of the present application, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present application. In addition, the terms "first", "second", and "third" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance.
[0136] Finally, it should be noted that the above embodiments are only specific implementation manners of the present application, used to illustrate the technical solutions of the present application, rather than limiting them. The protection scope of the present application is not limited thereto. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed in the present application can still modify the technical solutions recorded in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A security warning and monitoring method for rail transit stations, characterized in that, The method includes: Collecting multi-channel audio stream signals in the station through a high-sensitivity microphone array; Performing first preprocessing and second preprocessing on the multi-channel audio stream signals to obtain Mel spectrogram features; the first preprocessing includes: segmental noise reduction and normalization processing; the second preprocessing includes: multi-modal data augmentation processing; Inputting the Mel spectrogram features into a preset speaker recognition model, and performing feature mapping and classification on the Mel spectrogram features through the speaker recognition model, and outputting a classification feature result corresponding to the Mel spectrogram features; wherein, the speaker recognition model includes: an improved ECAPA-TDNN model based on multi-scale convolution, channel attention mechanism, dynamic attention statistical pooling, and angular margin loss function; Based on the classification feature result, perform threshold determination to determine whether to perform an alarm emergency response.
2. The method according to claim 1, wherein The step of performing first preprocessing and second preprocessing on the multi-channel audio stream signals to obtain Mel spectrogram features includes: Performing segmental noise reduction and normalization processing on the multi-channel audio stream signals to obtain initial spectrogram features; Performing multi-modal data augmentation processing on the initial spectrogram features to obtain Mel spectrogram features; Wherein, the multi-modal data augmentation processing includes: additive noise injection processing, speed perturbation processing, and spectral masking processing; the additive noise injection processing includes: superimposing background noise on the clean audio according to the signal-to-noise ratio; the speed perturbation processing includes: changing the audio duration by resampling to generate speed perturbation samples; the spectral masking processing includes: randomly masking frequency bands and time domain regions on the initial spectrogram features.
3. The method according to claim 1, wherein The improved ECAPA-TDNN model includes: a multi-scale convolution module, a channel attention mechanism module, a dynamic attention statistical pooling module, and a linear compression module; the step of inputting the Mel spectrogram features into a preset speaker recognition model, and performing feature mapping and classification on the Mel spectrogram features through the speaker recognition model, and outputting a classification feature result corresponding to the Mel spectrogram features includes: Dividing the Mel spectrogram features into multiple sub-features through the multi-scale convolution module, and performing progressive sub-feature accumulation and convolution on the multiple sub-features to obtain a convolution feature map; Performing global average pooling on the convolution feature map through the channel attention mechanism module, and processing it through two linear transformations and an activation function to obtain channel weights; Calculating attention weights in the time dimension based on the feature vectors affected by the channel weights through the dynamic attention statistical pooling module, and aggregating mean and standard deviation features based on the attention weights; Linearly compressing the mean and standard deviation features through the linear compression module to obtain a classification feature result corresponding to the Mel spectrogram features.
4. The method according to claim 3, wherein The step of performing global average pooling on the convolution feature map through the channel attention mechanism module, and processing it through two linear transformations and an activation function to obtain channel weights includes: dynamically adjusting the channel weights w through the Squeeze-and-Excitation mechanism: w = σ(W2·ReLU(W1·GAP(H out ))); Among them, GAP is the global average pooling process, and H out represents the convolutional feature map; W1 and W2 are the weights of the fully connected layers respectively, which are used for linear transformation; ReLU is the activation function, which is used to introduce non-linearity; σ is the Softmax activation function, which is used to map the output to a specific range; The step of calculating the attention weight in the time dimension based on the feature vector under the influence of the channel weight by means of the dynamic attention statistical pooling module, and aggregating the mean and standard deviation features based on the attention weight, includes: calculating the attention weight α in the time dimension according to the following formula t , the mean μ and the standard deviation feature σ: α t = Softmax(v T tanh(Wh t + b)); Among them, v, W, and b are learnable parameters used to calculate attention weights; tanh is an activation function used for non-linear transformation; h t represents the feature vector of the model at time step t; Softmax represents the activation function; The step of linearly compressing the mean and the standard deviation features through the linear compression module to obtain the classification feature result corresponding to the Mel spectrogram feature includes: determining the classification feature result z corresponding to the Mel spectrogram feature according to the following formula: z = W liner ·z pool + b liner ; z pool = Concat(μ, σ); z ∈ R D ; where, W liner ∈R D×2C , b luner ∈R D are learnable parameters.
5. The method according to claim 1, wherein The step of performing threshold determination based on the classification feature result to determine whether to perform an alarm emergency response includes: Calculating the cosine similarity score between the classification feature result and the classification weights corresponding to multiple preset categories; Taking the preset category corresponding to the highest cosine similarity score as the target category corresponding to the classification feature result; Judging whether the highest cosine similarity score exceeds the corresponding preset threshold; if so, giving an alarm in the warning manner corresponding to the target category.
6. The method according to claim 5, wherein The preset threshold is determined by the ROC curve of the test set and satisfies the following formula: τ=argmax(TPR(τ)-FPR(τ)) where TPR(τ) is the true positive rate and FPR(τ) is the false positive rate.
7. The method according to claim 1, characterized in that, During the training process of the improved ECAPA-TDNN model, the cosine annealing strategy is adopted to dynamically adjust the learning rate, and the angular margin loss function is adopted for model convergence; The learning rate formula is as follows: where η t is the learning rate at training step t, η min and η max are the minimum and maximum values of the learning rate respectively, and T is the total number of training steps or the cycle length; Angle interval loss function The formula is as follows: Among them, θ yi is the included angle between the sample and the true class weight w i ; s is the feature scaling factor used to amplify the difference between classes; the angular interval m is used to increase the difference between classes; cos(θ yi + m) means adding an interval m to the included angle of the true class, making the classification boundary more strict; represents the sum of the scores of other classes.
8. The method according to claim 1, characterized in that, The improved ECAPA-TDNN model adopts multi-GPU parallel training, and the AllReduce gradient synchronization strategy is used to ensure that the model parameters on each GPU are updated consistently.
9. An anti-theft warning and monitoring device for a rail transit station, characterized in that, The device includes: An audio acquisition module, configured to acquire multi-channel audio stream signals in the station through a high-sensitivity microphone array; A preprocessing module: configured to segment, denoise, and standardize the multi-channel audio stream signals, and extract Mel spectrogram features; A model prediction module, configured to input the Mel spectrogram features into a preset voiceprint recognition model, perform feature mapping and classification on the Mel spectrogram features through the voiceprint recognition model, and output the classification feature result corresponding to the Mel spectrogram features; wherein, the voiceprint recognition model includes: an improved ECAPA-TDNN model based on multi-scale convolution, channel attention mechanism, dynamic attention statistical pooling, and angular margin loss function; A decision warning module, configured to perform threshold determination based on the classification feature result to determine whether to perform an alarm emergency response.
10. A security warning and monitoring system for rail transit stations, characterized in that, The system includes: a high-sensitivity microphone array and a server connected by communication; the server is configured to execute the method according to any one of claims 1 to 8.