Gateway busy tone detection method, device, equipment and storage medium based on deep learning
Through multi-level feature extraction and pre-trained model detection based on deep learning methods, the problem of DSP busy tone detection being susceptible to interference is solved, the accuracy of busy tone detection and the stability of the FXO gateway are improved, ensuring the normal operation of the communication line.
Patent Information
- Application Number
- CN202510970563.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-15
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-07-15
AI Technical Summary
The existing DSP-based busy tone detection method is easily affected by interference signals, resulting in low accuracy of busy tone detection on FXO gateways, leading to idle communication lines and low utilization.
A deep learning-based method is used to obtain the original audio signal through the analog line interface, perform multi-level feature extraction to generate a timing feature matrix, and input the pre-trained deep learning model to output the busy tone probability value, and generate port control instructions according to the probability value.
The accuracy of busy tone detection is improved, the occurrence of line biting is reduced, and the stability and reliability of the FXO gateway are improved. It can accurately detect busy tones in complex environments and ensure the normal operation of communication lines.
Smart Images

Figure CN120526797B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of communication technologies, and in particular to a gateway busy tone detection method and apparatus, device, and storage medium based on deep learning. Background Art
[0002] In communications systems, FXO gateways are key devices that connect analog telephone lines to digital networks. Currently, many FXO gateways suffer from a serious line-biting problem, where ports remain abnormally busy, causing lines to become incorrectly idle and significantly impacting the effective utilization of communication lines.
[0003] The root of the problem lies in flaws in existing DSP-based busy tone detection methods. DSP busy tone detection is susceptible to various interference signals, such as voltage fluctuations and signal amplitude variations, resulting in inaccurate detection results and, in turn, wire biting failures. Existing detection methods are unable to meet the growing demand for communication stability and reliability.
[0004] The above content is only used to assist in understanding the technical solution of the present invention and does not constitute an admission that the above content is prior art. Summary of the Invention
[0005] The main purpose of the present invention is to provide a gateway busy tone detection method and apparatus, device and storage medium based on deep learning, aiming to solve the technical problem of low accuracy of FXO gateway busy tone detection.
[0006] To achieve the above object, the present invention provides a gateway busy tone detection method based on deep learning, which comprises the following steps:
[0007] Get the original audio signal through the analog line interface;
[0008] Performing multi-level feature extraction on the original audio signal to generate a time series feature matrix;
[0009] Input the time series feature matrix into a pre-trained deep learning model and output a busy tone probability value;
[0010] A port control instruction is generated according to a comparison result of the busy tone probability value and a preset threshold.
[0011] In one embodiment, the step of performing multi-level feature extraction on the original audio signal to generate a time series feature matrix includes:
[0012] Sampling the original audio signal at a preset sampling rate to generate a digital audio signal;
[0013] Performing dynamic noise suppression on the digital audio signal by using an adaptive Kalman filter to generate a noise reduction signal;
[0014] Performing frame processing on the noise reduction signal according to a preset frame length and frame shift to generate a frame sequence;
[0015] Extracting Mel-frequency cepstral coefficients of each frame in the frame sequence;
[0016] synchronously calculating the short-term energy of each frame in the frame sequence;
[0017] The Mel-frequency cepstral coefficients and the short-time energy are combined into a time series feature matrix.
[0018] In one embodiment, the step of inputting the time series feature matrix into a pre-trained deep learning model and outputting a busy tone probability value includes:
[0019] Performing frequency domain feature extraction on the temporal feature matrix through a multi-scale convolutional layer to generate a multi-scale feature map;
[0020] Performing dimension compression on the multi-scale feature map through a maximum pooling layer to generate a compressed feature map;
[0021] Inputting the compressed feature map into a bidirectional long short-term memory network for temporal modeling to generate a context feature vector;
[0022] The context feature vector is mapped into a busy tone probability value through a fully connected layer.
[0023] In one embodiment, the step of performing frequency domain feature extraction on the temporal feature matrix through a multi-scale convolutional layer to generate a multi-scale feature map includes:
[0024] Using a first convolution kernel to extract high-frequency components from the time series feature matrix to obtain high-frequency components;
[0025] Using a second convolution kernel to extract the intermediate frequency component of the time series feature matrix to obtain the intermediate frequency component;
[0026] Using a third convolution kernel to extract low-frequency components from the time series feature matrix to obtain low-frequency components;
[0027] Perform channel-wise splicing on the high-frequency component, the medium-frequency component, and the low-frequency component to generate a fusion feature map;
[0028] Batch normalization is performed on the fused feature map to obtain a multi-scale feature map.
[0029] In one embodiment, the step of generating a port control instruction according to a comparison result of the busy tone probability value and a preset threshold value includes:
[0030] When the busy tone probability value is greater than a preset threshold, generating a port occupation instruction;
[0031] When the busy tone probability value is not greater than a preset threshold, generating a port release instruction;
[0032] The port occupation instruction or the port release instruction is sent to the gateway control module, so that the gateway control module switches the physical port state according to the received instruction.
[0033] In one embodiment, before the step of inputting the time series feature matrix into the pre-trained deep learning model, the method further includes:
[0034] Collect labeled busy tone and non-busy tone sample data;
[0035] Performing voltage fluctuation simulation and phase shift enhancement on the sample data to generate an enhanced sample set;
[0036] Converting the enhanced sample set into a training feature matrix through a time-frequency transformation module;
[0037] The model is trained using a dynamic decay learning rate strategy, with the initial learning rate set to a preset value;
[0038] When the accuracy of the validation set does not improve for a preset number of consecutive rounds, the training is terminated and the model parameters are saved to obtain a pre-trained deep learning model.
[0039] In one embodiment, the step of performing voltage fluctuation simulation and phase offset enhancement on the sample data to generate an enhanced sample set includes:
[0040] Adding Gaussian noise of a preset amplitude range to the sample data to obtain a noisy sample;
[0041] Simulating a phase shift caused by line impedance mismatch on the noisy sample to obtain a phase shift sample;
[0042] generating a synthetic sample including voltage fluctuation for the phase shift sample;
[0043] The synthesized samples are subjected to time domain stretching and compression processing to obtain an enhanced sample set.
[0044] In addition, to achieve the above-mentioned purpose, the present invention also proposes a gateway busy tone detection device based on deep learning, the device comprising:
[0045] A signal acquisition module, used for acquiring original audio signals through an analog line interface;
[0046] A matrix generation module, configured to perform multi-level feature extraction on the original audio signal to generate a time series feature matrix;
[0047] A model calculation module, configured to input the time series feature matrix into a pre-trained deep learning model and output a busy tone probability value;
[0048] The control module is configured to generate a port control instruction according to a comparison result between the busy tone probability value and a preset threshold.
[0049] In addition, to achieve the above-mentioned purpose, the present invention also proposes a gateway busy tone detection device based on deep learning, which includes: a memory, a processor, and a gateway busy tone detection program based on deep learning stored in the memory and executable on the processor, wherein the gateway busy tone detection program based on deep learning is configured to implement the steps of the gateway busy tone detection method based on deep learning as described above.
[0050] In addition, to achieve the above-mentioned purpose, the present invention also proposes a storage medium, on which a gateway busy tone detection program based on deep learning is stored. When the gateway busy tone detection program based on deep learning is executed by a processor, the steps of the gateway busy tone detection method based on deep learning as described above are implemented.
[0051] In addition, to achieve the above-mentioned purpose, the present application also provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the steps of the gateway busy tone detection method based on deep learning as described above.
[0052] One or more technical solutions proposed in this application have at least the following technical effects:
[0053] The system acquires raw audio signals through an analog line interface; performs multi-level feature extraction on the raw audio signals to generate a time series feature matrix; inputs the matrix into a pre-trained deep learning model to output a busy tone probability value; and generates port control instructions based on the comparison of the busy tone probability value with a preset threshold. The deep learning model can automatically learn complex audio features. Compared with traditional DSP-based detection methods, it has higher accuracy in busy tone detection in various environments, effectively reducing the occurrence of line bite and improving the stability and reliability of the FXO gateway. Through extensive data training, the model has strong adaptability to interference such as voltage fluctuations and signal amplitude changes, and can accurately detect busy tones in complex communication environments, ensuring the normal operation of communication lines. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0055] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0056] Figure 1 A flowchart of the first embodiment of the gateway busy tone detection method based on deep learning provided in this application;
[0057] Figure 2 This is an overall flow chart of the first embodiment of the gateway busy tone detection method based on deep learning of this application;
[0058] Figure 3 A flowchart of the second embodiment of the gateway busy tone detection method based on deep learning is provided in this application;
[0059] Figure 4 This is a schematic diagram of the module structure of a gateway busy tone detection device based on deep learning in an embodiment of the present application;
[0060] Figure 5 This is a schematic diagram of the device structure of the hardware operating environment involved in the gateway busy tone detection method based on deep learning in an embodiment of the present application.
[0061] The purpose, features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0062] It should be understood that the specific embodiments described herein are merely used to explain the technical solutions of the present application and are not intended to limit the present application.
[0063] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.
[0064] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, mobile phone, etc., or an electronic device capable of implementing the above functions, a deep learning-based gateway busy tone detection device, etc. The following uses the deep learning-based gateway busy tone detection device as an example to illustrate this embodiment and the following embodiments.
[0065] Based on this, the embodiment of the present application provides a gateway busy tone detection method based on deep learning, referring to Figure 1 , Figure 1 This is a flowchart of the first embodiment of the gateway busy tone detection method based on deep learning in this application.
[0066] In this embodiment, the gateway busy tone detection method based on deep learning includes steps S10 to S40:
[0067] Step S10, obtaining the original audio signal through the analog line interface;
[0068] It should be noted that this refers to the raw, unprocessed audio data received by the FXO gateway from traditional analog telephone lines through its analog line interface. This data includes busy tone signals, but due to factors such as line interference, direct use may affect the judgment, so subsequent processing is required.
[0069] like Figure 2 As shown, audio signals are acquired from the analog line interface of the FXO gateway. Preprocessing begins by digitizing the signal at an 8kHz sampling rate and then applying a Kalman or Wiener filter for noise removal. The audio signal is then framed (each frame is 20-30ms with overlapping). This preprocessed, real-time framed data is fed into a trained deep learning model, which outputs a busy / not-busy probability value. Based on this value (>0.5 for busy, ≤0.5 for not-busy), the FXO gateway port status is determined and controlled. A training process also occurs: data labeled as busy / not-busy is collected and, after the same preprocessing, divided into training, validation, and test sets. The model is trained on the training set and parameters are adjusted through backpropagation. Training is then stopped or continued based on performance improvement on the validation set. The process concludes with port status control.
[0070] Step S20, performing multi-level feature extraction on the original audio signal to generate a time series feature matrix;
[0071] It's important to note that feature extraction involves calculating parameters that represent the signal's characteristics from the original audio signal. Different features can reflect different aspects of the signal, such as frequency distribution and energy variation. Choosing the right features is crucial for busy tone detection, as they directly impact the accuracy of subsequent model recognition.
[0072] In its implementation, the audio signal is acquired from the analog line interface of the FXO gateway, sampled, and converted into a discrete digital signal at a specific sampling rate (e.g., 8kHz). The sampled digital audio signal is then denoised using filtering algorithms (e.g., Kalman filtering or Wiener filtering) to remove interference such as ambient noise and power supply noise. The audio signal is then divided into frames, each with a suitable duration (e.g., 20-30ms). During the framing process, an overlap-add algorithm is used to avoid information loss at frame boundaries.
[0073] In the signal processing module within the FXO gateway, a sampling circuit is configured to sample analog audio signals at an 8kHz sampling rate. At the software level, a Kalman filter algorithm is used to denoise the sampled digital audio signals, removing electromagnetic interference and background noise from the line. Framing parameters are configured to divide the audio signal into 25ms frames, with a 10ms overlap between frames to ensure signal continuity and integrity.
[0074] In a feasible implementation, step S20 includes steps A11 to A16:
[0075] A11: samples the original audio signal at a preset sampling rate to generate a digital audio signal;
[0076] It's important to note that sampling is the process of converting a continuous analog signal into a discrete digital signal. Because computers can only process digital signals, sampling is necessary to convert analog audio signals into digital form. The sampling rate determines the frequency resolution of the converted digital signal.
[0077] A12: Dynamically suppresses noise on digital audio signals through an adaptive Kalman filter to generate a noise-reduced signal.
[0078] It’s important to note that dynamic noise suppression uses an adaptive Kalman filter to estimate and eliminate noise components in audio signals in real time. This method automatically adjusts filter parameters based on signal changes, effectively improving signal quality in varying noise environments.
[0079] A13: Frame the noise reduction signal according to the preset frame length and frame shift to generate a frame sequence;
[0080] It should be noted that framing is the process of dividing a continuous audio signal into multiple shorter frames for short-term feature analysis. The choice of frame length and frame shift affects the time and frequency domain resolution of the features.
[0081] A14: Extract the Mel-frequency cepstral coefficients of each frame in the frame sequence;
[0082] It's important to note that Mel-Frequency Cepstral Coefficients (MFCCs) are a feature extraction method that mimics the human auditory system's sensitivity to sound frequency. They convert signals to the Mel-Frequency scale and then calculate the Cepstral Coefficients, effectively capturing important features in speech and audio signals.
[0083] A15: Synchronously calculate the short-term energy of each frame in the frame sequence;
[0084] It should be noted that short-term energy is the sum of the energy of each frame of the signal and is used to measure the change in signal strength. In busy tone detection, short-term energy can help distinguish between sound and silence intervals, thus assisting in judgment.
[0085] A16: Combine the Mel-frequency cepstral coefficients and short-time energy into a time series feature matrix.
[0086] It should be noted that MFCC and short-term energy, two different types of features, are arranged in chronological order and combined into a matrix. This can form a comprehensive feature representation that contains both frequency and energy information, providing richer input for subsequent deep learning models.
[0087] Step S30, inputting the time series feature matrix into a pre-trained deep learning model and outputting a busy tone probability value;
[0088] It should be noted that a pre-trained deep learning model refers to a model that has been trained with a large amount of labeled data and can recognize patterns in input features and output the probability of the corresponding category.
[0089] In a feasible implementation manner, step S30 includes steps A21 to A25:
[0090] A21: Collects labeled busy and non-busy tone sample data;
[0091] In the specific implementation, busy tone and non-busy tone audio samples in actual communication are collected and labeled with category labels to form a training data set.
[0092] A22: Perform voltage fluctuation simulation and phase shift enhancement on the sample data to generate an enhanced sample set;
[0093] It should be noted that by adding simulated voltage fluctuations and phase shifts to expand the training sample set, the model can learn more diverse features and improve its robustness in practical applications.
[0094] Furthermore, step A22 includes:
[0095] Add Gaussian noise of a preset amplitude range to the sample data to obtain a noisy sample;
[0096] Simulating the phase shift caused by line impedance mismatch for the noisy sample to obtain a phase shift sample;
[0097] generating a synthetic sample including voltage fluctuations for the phase-shifted sample;
[0098] The synthetic samples are stretched and compressed in the time domain to obtain an enhanced sample set.
[0099] It should be noted that Gaussian noise is a type of noise whose probability density function is a Gaussian function (normal distribution). Adding Gaussian noise can simulate random interference in a real environment. The preset amplitude range is the noise amplitude range set according to actual needs to ensure that the noise intensity is within a reasonable range. Noisy samples are samples after adding noise, which are used to improve the robustness of the model to noise. Line impedance mismatch refers to the impedance mismatch in the line, which may cause signal reflection and phase change. Phase offset refers to the phase offset of the signal, which may affect the quality and characteristics of the signal. Phase offset samples are samples after simulating phase offset, which are used to enhance the model's adaptability to line anomalies. Voltage fluctuation refers to the unstable change of voltage in the line, which may affect signal transmission. Synthetic samples are samples generated by combining phase offset and voltage fluctuation to simulate complex line conditions. Time domain stretching and compression processing is to change the length of the signal on the time axis to simulate different speaking speeds or line delays. The enhanced sample set is a set of samples that has undergone multiple processing and has greater diversity and robustness, which helps to improve the generalization ability of the model.
[0100] A23: Convert the enhanced sample set into a training feature matrix through the time-frequency transformation module;
[0101] It should be noted that time-frequency transformation converts the signal into the time-frequency domain so that both the time domain and frequency domain characteristics of the signal can be utilized by the model, thereby enhancing the model's ability to represent the signal.
[0102] A24: The model is trained using a dynamic decay learning rate strategy, with the initial learning rate set to a preset value.
[0103] It should be noted that gradually reducing the learning rate during the training process allows the model to converge quickly in the early stages of training, while fine-tuning the parameters in the later stages to avoid overshooting.
[0104] A25: When the validation set accuracy does not improve for a preset number of consecutive rounds, terminate the training and save the model parameters to obtain a pre-trained deep learning model.
[0105] It is important to note that an independent validation set is used to evaluate the model's performance to prevent overfitting. If the accuracy on the validation set does not improve after several rounds of training, training is stopped and the current model parameters are saved.
[0106] In the specific implementation, appropriate deep learning models are selected, such as convolutional neural networks (CNNs), recurrent neural networks (RNNs), and their variants, long short-term memory (LSTMs), and gated recurrent units (GRUs). Taking CNNs as an example, a network structure consisting of multiple convolutional, pooling, and fully connected layers is constructed. Convolutional layers extract local features of the audio signal, while pooling layers reduce data dimensionality and computational complexity. Fully connected layers integrate the previously extracted features and output the final classification result. A large amount of busy and non-busy audio data is collected and labeled to indicate whether it is busy or not. The preprocessed audio data is divided into training, validation, and test sets according to a specific ratio (e.g., 70% training, 15% validation, and 15% test). The deep learning model is trained using the training data. The backpropagation algorithm continuously adjusts the model parameters (such as convolution kernel weights and fully connected layer weights) to continuously reduce the model's loss function (such as the cross-entropy loss function) on the training set. During the training process, the model is validated using the validation set data to prevent overfitting. Training is stopped when the model's performance on the validation set no longer improves.
[0107] An LSTM model was selected for busy tone detection. A network architecture consisting of three LSTM layers and one fully connected layer was constructed. The number of neurons in the LSTM layers was set to 128, 64, and 32, respectively, and the number of neurons in the fully connected layer was set to 2 (corresponding to the busy and non-busy categories). A large amount of busy and non-busy audio data, totaling over 1,000 hours, was collected through web crawling and acquisition of data from actual communication lines. Professional audio annotation software was used to annotate the data, labeling each audio segment as busy or non-busy. The annotated data was divided into training, validation, and test sets with a 70%, 15%, and 15% split ratio. The model was built and trained using Python and a deep learning framework such as TensorFlow or PyTorch. During training, a learning rate of 0.001, a batch size of 64, and 100 epochs were set. The accuracy and loss of the model on the validation set were monitored in real time during training. Training was terminated when the validation set accuracy did not improve for five consecutive epochs.
[0108] Step S40: Generate a port control instruction according to a comparison result between the busy tone probability value and a preset threshold.
[0109] It should be noted that the busy tone probability output by the model is compared with a manually set threshold, and whether to generate a control instruction is determined based on the comparison result.
[0110] In a feasible implementation, step S40 includes steps A31 to A33:
[0111] A31: When the busy tone probability value is greater than the preset threshold, a port occupation instruction is generated;
[0112] In a specific implementation, when a busy tone is detected, an instruction is generated to notify the FXO gateway to occupy the corresponding port to prevent the line from being mistakenly judged as idle.
[0113] A32: When the busy tone probability value is not greater than the preset threshold, a port release instruction is generated;
[0114] In a specific implementation, when no busy tone is detected, an instruction is generated to notify the FXO gateway to release the corresponding port so that it can be used for other calls.
[0115] A33: Sending a port occupation instruction or a port release instruction to the gateway control module, so that the gateway control module switches the physical port state according to the received instruction.
[0116] It should be noted that, according to the control instructions, the actual working state of the physical port is changed to achieve effective management of the communication line.
[0117] In the implementation, preprocessed real-time audio data is fed into a trained deep learning model, which then outputs the probability of the audio data being busy or not busy. A suitable threshold (e.g., 0.5) is set. When the busy tone probability output by the model exceeds the threshold, the current audio signal is considered busy; otherwise, it is considered not busy. Based on this detection result, the true status of the FXO gateway port is accurately determined, avoiding line jamming and idle lines caused by false detection.
[0118] In the FXO gateway's real-time signal processing, preprocessed audio data is fed into a trained LSTM model in real time. The model outputs busy and non-busy tone probabilities, with a threshold of 0.5. When the busy tone probability output by the model is greater than 0.5, the port status is determined to be busy; otherwise, it is considered non-busy. Based on this determination, the FXO gateway port status is accurately controlled to prevent line jamming and idleness.
[0119] This embodiment provides a gateway busy tone detection method based on deep learning. The method obtains raw audio signals through an analog line interface, performs multi-level feature extraction on the raw audio signals, and generates a time series feature matrix. The time series feature matrix is input into a pre-trained deep learning model to output a busy tone probability value. Port control instructions are generated based on the comparison of the busy tone probability value with a preset threshold. The deep learning model can automatically learn complex audio features. Compared with traditional DSP-based detection methods, it has higher accuracy in busy tone detection in different environments, effectively reduces the occurrence of line jamming, and improves the stability and reliability of FXO gateways. Through extensive data training, the model has strong adaptability to interference such as voltage fluctuations and signal amplitude changes. It can accurately detect busy tones in complex communication environments and ensure the normal operation of communication lines. This solves the problem of poor accuracy of existing DSP busy tone detection methods and their susceptibility to interference, which can lead to FXO gateway line jamming and line idleness. The model improves the accuracy of busy tone detection and ensures the normal operation and efficient utilization of communication lines.
[0120] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar contents as those in the above embodiment 1 can be referred to the above introduction and will not be described in detail later. Figure 3 Step S30 includes steps S301 to S303:
[0121] Step S301, extracting frequency domain features from the temporal feature matrix through a multi-scale convolutional layer to generate a multi-scale feature map;
[0122] It should be noted that the multi-scale convolution layer uses convolution kernels of different sizes to capture features of different scales. Small convolution kernels can capture local details (such as high-frequency features), while large convolution kernels can capture broader patterns (such as low-frequency features).
[0123] Frequency domain feature extraction is to extract frequency-related features from the time series feature matrix. It helps the model understand the distribution of different frequency components in the audio signal.
[0124] Multi-scale feature maps contain maps of features at different scales, providing rich and multi-dimensional information for subsequent processing.
[0125] Step S302: compress the multi-scale feature map using a maximum pooling layer to generate a compressed feature map.
[0126] It should be noted that the max pooling layer is a layer used to reduce the spatial size of the feature map. It performs downsampling by taking the maximum value in a local area, reducing the amount of computation and improving the robustness of the model.
[0127] Dimensionality compression is to reduce the size of the feature map while retaining the most important feature information.
[0128] In a feasible implementation, step S302 includes steps A41 to A45:
[0129] A41: Use the first convolution kernel to extract the high-frequency components of the time series feature matrix to obtain the high-frequency components;
[0130] It should be noted that the first convolution kernel is a small-size convolution kernel designed to extract high-frequency features.
[0131] High-frequency component extraction is to capture the high-frequency part of the audio signal, such as high-frequency noise or high-frequency tones.
[0132] A42: Use the second convolution kernel to extract the intermediate frequency component of the time series feature matrix to obtain the intermediate frequency component;
[0133] It should be noted that the second convolution kernel is a medium-sized convolution kernel used to extract medium-frequency features.
[0134] Intermediate frequency component extraction is to capture the intermediate frequency part of the audio signal, such as the main frequency range of speech.
[0135] A43: Use the third convolution kernel to extract the low-frequency component of the time series feature matrix to obtain the low-frequency component;
[0136] It should be noted that the third convolution kernel is a large-size convolution kernel, which is used to extract low-frequency features.
[0137] Low-frequency component extraction is to capture the low-frequency part of the audio signal, such as background sound or low-frequency vibration.
[0138] A44: Concatenate the high-frequency component, the mid-frequency component, and the low-frequency component in the channel dimension to generate a fusion feature map;
[0139] It should be noted that channel dimension splicing is to splice the feature maps of different frequency components in the channel dimension to form a comprehensive feature map.
[0140] The fused feature map is a feature map that integrates high-, low- and medium-frequency information, providing a more comprehensive frequency feature representation.
[0141] A45: Perform batch normalization on the fused feature map to obtain a multi-scale feature map.
[0142] It should be noted that batch normalization is to normalize the fused feature map so that it has zero mean and unit variance, which accelerates training and improves model stability.
[0143] The multi-scale feature map is a normalized feature map that provides better input for subsequent time series modeling.
[0144] Step S303: Input the compressed feature map into a bidirectional long short-term memory network for temporal modeling to generate a context feature vector;
[0145] It should be noted that the BiLSTM is a neural network structure that can process sequential data and is good at capturing long-term dependencies in time series. The bidirectional structure can simultaneously utilize past and future contextual information.
[0146] Time series modeling is to model the time series information in the feature graph to capture the dynamic changes in the audio signal.
[0147] The context feature vector contains the feature vector of the time series context information and can provide a richer time series feature representation.
[0148] Step S304: Map the context feature vector to a busy tone probability value through a fully connected layer.
[0149] It should be noted that the fully connected layer is a layer in the neural network, in which each neuron is connected to all neurons in the previous layer and is used to map input features to the output space.
[0150] Mapping to a busy tone probability value is to convert the context feature vector into a probability value, which indicates the possibility that the current audio signal is a busy tone.
[0151] This embodiment provides a gateway busy tone detection method based on deep learning. A multi-scale convolutional layer extracts high-frequency, mid-frequency, and low-frequency features of an audio signal to generate a multi-scale feature map that comprehensively captures feature information at different frequencies. The fused feature map undergoes batch normalization and is then input into a bidirectional long short-term memory network for temporal modeling to generate a contextual feature vector. Finally, a fully connected layer maps the feature vector into a busy tone probability value, capturing the temporal characteristics and long-term dependencies of the audio signal and improving the accuracy of busy tone detection.
[0152] It should be noted that the above examples are only used to understand the present application and do not constitute a limitation on the gateway busy tone detection method based on deep learning in the present application. More simple transformations based on this technical concept are all within the scope of protection of the present application.
[0153] This application also provides a gateway busy tone detection device based on deep learning, please refer to Figure 4 , the gateway busy tone detection device based on deep learning includes:
[0154] The signal acquisition module 10 is used to obtain the original audio signal through the analog line interface;
[0155] The matrix generation module 20 is used to perform multi-level feature extraction on the original audio signal and generate a time series feature matrix;
[0156] The model calculation module 30 is used to input the time series feature matrix into the pre-trained deep learning model and output a busy tone probability value;
[0157] The control module 40 is configured to generate a port control instruction according to a comparison result between the busy tone probability value and a preset threshold.
[0158] The deep learning-based gateway busy tone detection device provided in this application utilizes the deep learning-based gateway busy tone detection method described in the aforementioned embodiments, addressing the technical issue of low FXO gateway busy tone detection accuracy. Compared to the prior art, the deep learning-based gateway busy tone detection device provided in this application offers the same beneficial effects as the deep learning-based gateway busy tone detection method described in the aforementioned embodiments. Other technical features of the deep learning-based gateway busy tone detection device are the same as those disclosed in the aforementioned embodiments and are not further elaborated here.
[0159] In one embodiment, the matrix generation module 20 is further configured to sample the original audio signal at a preset sampling rate to generate a digital audio signal;
[0160] Dynamic noise suppression is performed on digital audio signals through an adaptive Kalman filter to generate a noise reduction signal;
[0161] The noise reduction signal is divided into frames according to the preset frame length and frame shift to generate a frame sequence;
[0162] Extract the Mel-frequency cepstral coefficients of each frame in the frame sequence;
[0163] Synchronously calculate the short-term energy of each frame in the frame sequence;
[0164] The Mel-frequency cepstral coefficients and short-time energy are combined into a time series feature matrix.
[0165] In one embodiment, the model calculation module 30 is further configured to perform frequency domain feature extraction on the time series feature matrix through a multi-scale convolution layer to generate a multi-scale feature map;
[0166] The multi-scale feature map is dimensional compressed through the maximum pooling layer to generate a compressed feature map;
[0167] The compressed feature map is input into the bidirectional long short-term memory network for temporal modeling to generate a context feature vector;
[0168] The context feature vector is mapped to a busy tone probability value through a fully connected layer.
[0169] In one embodiment, the model calculation module 30 is further configured to extract high-frequency components from the time series feature matrix using a first convolution kernel to obtain high-frequency components;
[0170] The second convolution kernel is used to extract the intermediate frequency component of the time series feature matrix to obtain the intermediate frequency component;
[0171] The third convolution kernel is used to extract the low-frequency component of the time series feature matrix to obtain the low-frequency component;
[0172] The high-frequency component, the medium-frequency component and the low-frequency component are spliced in the channel dimension to generate a fusion feature map;
[0173] The fused feature maps are batch normalized to obtain multi-scale feature maps.
[0174] In one embodiment, the control module 40 is further configured to generate a port occupation instruction when the busy tone probability value is greater than a preset threshold;
[0175] When the busy tone probability value is not greater than a preset threshold, a port release instruction is generated;
[0176] The port occupation instruction or the port release instruction is sent to the gateway control module, so that the gateway control module switches the physical port state according to the received instruction.
[0177] In one embodiment, the model calculation module 30 is further configured to collect labeled busy tone and non-busy tone sample data;
[0178] Perform voltage fluctuation simulation and phase shift enhancement on sample data to generate an enhanced sample set;
[0179] The enhanced sample set is converted into a training feature matrix through the time-frequency transformation module;
[0180] The model is trained using a dynamic decay learning rate strategy, with the initial learning rate set to a preset value;
[0181] When the accuracy of the validation set does not improve for a preset number of consecutive rounds, the training is terminated and the model parameters are saved to obtain a pre-trained deep learning model.
[0182] In one embodiment, the model calculation module 30 is further configured to add Gaussian noise of a preset amplitude range to the sample data to obtain a noisy sample;
[0183] Simulating the phase shift caused by line impedance mismatch for the noisy sample to obtain a phase shift sample;
[0184] generating a synthetic sample including voltage fluctuations for the phase-shifted sample;
[0185] The synthetic samples are stretched and compressed in the time domain to obtain an enhanced sample set.
[0186] The present application provides a deep learning-based gateway busy tone detection device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the deep learning-based gateway busy tone detection method in the above-mentioned embodiment 1.
[0187] Reference below Figure 5 , which shows a schematic structural diagram of a deep learning-based gateway busy tone detection device suitable for implementing embodiments of the present application. The deep learning-based gateway busy tone detection device in the embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 5 The deep learning-based gateway busy tone detection device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0188] like Figure 5As shown, the deep learning-based gateway busy tone detection device may include a processing device 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes based on programs stored in ROM (Read Only Memory) 1002 or programs loaded from storage device 1003 into RAM (Random Access Memory) 1004. RAM 1004 also stores various programs and data required for the operation of the deep learning-based gateway busy tone detection device. Processing device 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, a magnetic tape, hard disk, etc.; and communication devices 1009. Communication devices 1009 can allow the deep learning-based gateway busy tone detection device to communicate with other devices wirelessly or by wire to exchange data. While the figure shows a deep learning-based gateway busy tone detection device with various systems, it should be understood that implementation or presence of all the illustrated systems is not required. More or fewer systems may alternatively be implemented or present.
[0189] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are executed.
[0190] The deep learning-based gateway busy tone detection device provided in this application utilizes the deep learning-based gateway busy tone detection method described in the aforementioned embodiment to address the technical issue of low FXO gateway busy tone detection accuracy. Compared to the prior art, the deep learning-based gateway busy tone detection device provided in this application achieves the same beneficial effects as the deep learning-based gateway busy tone detection method described in the aforementioned embodiment. Other technical features of this deep learning-based gateway busy tone detection device are the same as those disclosed in the aforementioned embodiment and are not further elaborated here.
[0191] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any one or more embodiments or examples in a suitable manner.
[0192] The above are only specific embodiments of the present application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
[0193] The present application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, and the computer-readable program instructions are used to execute the gateway busy tone detection method based on deep learning in the above embodiment.
[0194] The computer-readable storage medium provided herein may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems, or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more wires, a portable computer disk, a hard disk, RAM (Random Access Memory), ROM (Read Only Memory), EPROM (Erasable Programmable Read Only Memory or Flash memory), optical fiber, CD-ROM (CD-Read Only Memory), optical storage device, magnetic storage device, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including, but not limited to, wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0195] The above-mentioned computer-readable storage medium can be included in the gateway busy tone detection device based on deep learning; or it can exist independently without being assembled into the gateway busy tone detection device based on deep learning.
[0196] The computer-readable storage medium carries one or more programs. When the one or more programs are executed by a deep learning-based gateway busy tone detection device, the deep learning-based gateway busy tone detection device: obtains an original audio signal through an analog line interface; performs multi-level feature extraction on the original audio signal to generate a time series feature matrix; inputs the time series feature matrix into a pre-trained deep learning model to output a busy tone probability value; and generates a port control instruction based on a comparison result of the busy tone probability value and a preset threshold.
[0197] The computer program code for performing the operations of the present application may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a LAN (Local Area Network) or a WAN (Wide Area Network), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0198] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.
[0199] The modules described in the embodiments of the present application may be implemented in software or hardware, wherein the name of a module does not necessarily limit the unit itself.
[0200] The computer-readable storage medium provided in this application stores computer-readable program instructions (i.e., a computer program) for executing the aforementioned deep learning-based gateway busy tone detection method. This computer-readable storage medium can address the technical issue of low FXO gateway busy tone detection accuracy. Compared to the prior art, the beneficial effects of the computer-readable storage medium provided in this application are similar to those of the deep learning-based gateway busy tone detection method provided in the aforementioned embodiments, and are not further elaborated here.
[0201] The present application also provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of the above-mentioned deep learning-based gateway busy tone detection method.
[0202] The computer program product provided in this application can address the technical issue of low FXO gateway busy tone detection accuracy. Compared to the prior art, the computer program product provided in this application offers the same beneficial effects as the deep learning-based gateway busy tone detection method provided in the aforementioned embodiments, and will not be further elaborated here.
[0203] The above description is only part of the embodiments of the present application and does not limit the patent scope of the present application. All equivalent structural transformations made by using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect application in other related technical fields are included in the patent protection scope of the present application.
Claims
1. A gateway busy tone detection method based on deep learning, characterized in that: The method comprises: Get the original audio signal through the analog line interface; Performing multi-level feature extraction on the original audio signal to generate a time series feature matrix; Input the time series feature matrix into a pre-trained deep learning model and output a busy tone probability value; generating a port control instruction according to a comparison result of the busy tone probability value and a preset threshold; Before the step of inputting the time series feature matrix into the pre-trained deep learning model, the method further includes: Collect labeled busy tone and non-busy tone sample data; Performing voltage fluctuation simulation and phase shift enhancement on the sample data to generate an enhanced sample set; Converting the enhanced sample set into a training feature matrix through a time-frequency transformation module; The model is trained using a dynamic decay learning rate strategy, with the initial learning rate set to a preset value; When the accuracy of the validation set does not improve for a preset number of consecutive rounds, the training is terminated and the model parameters are saved to obtain a pre-trained deep learning model.
2. The method according to claim 1, wherein The step of performing multi-level feature extraction on the original audio signal to generate a time series feature matrix includes: Sampling the original audio signal at a preset sampling rate to generate a digital audio signal; Performing dynamic noise suppression on the digital audio signal by using an adaptive Kalman filter to generate a noise reduction signal; Performing frame processing on the noise reduction signal according to a preset frame length and frame shift to generate a frame sequence; Extracting Mel-frequency cepstral coefficients of each frame in the frame sequence; synchronously calculating the short-term energy of each frame in the frame sequence; The Mel-frequency cepstral coefficients and the short-time energy are combined into a time series feature matrix.
3. The method according to claim 1, wherein The step of inputting the time series feature matrix into a pre-trained deep learning model and outputting a busy tone probability value includes: Performing frequency domain feature extraction on the temporal feature matrix through a multi-scale convolutional layer to generate a multi-scale feature map; Performing dimension compression on the multi-scale feature map through a maximum pooling layer to generate a compressed feature map; Inputting the compressed feature map into a bidirectional long short-term memory network for temporal modeling to generate a context feature vector; The context feature vector is mapped into a busy tone probability value through a fully connected layer.
4. The method according to claim 3, wherein The step of extracting frequency domain features from the time series feature matrix through a multi-scale convolutional layer to generate a multi-scale feature map includes: Using a first convolution kernel to extract high-frequency components from the time series feature matrix to obtain high-frequency components; Using a second convolution kernel to extract the intermediate frequency component of the time series feature matrix to obtain the intermediate frequency component; Using a third convolution kernel to extract low-frequency components from the time series feature matrix to obtain low-frequency components; Perform channel-wise splicing on the high-frequency component, the medium-frequency component, and the low-frequency component to generate a fusion feature map; Batch normalization is performed on the fused feature map to obtain a multi-scale feature map.
5. The method according to claim 1, wherein The step of generating a port control instruction according to a comparison result of the busy tone probability value and a preset threshold value includes: When the busy tone probability value is greater than a preset threshold, generating a port occupation instruction; When the busy tone probability value is not greater than a preset threshold, generating a port release instruction; The port occupation instruction or the port release instruction is sent to the gateway control module, so that the gateway control module switches the physical port state according to the received instruction.
6. The method according to claim 5, wherein The step of performing voltage fluctuation simulation and phase offset enhancement on the sample data to generate an enhanced sample set includes: Adding Gaussian noise of a preset amplitude range to the sample data to obtain a noisy sample; Simulating a phase shift caused by line impedance mismatch on the noisy sample to obtain a phase shift sample; generating a synthetic sample including voltage fluctuation for the phase shift sample; The synthesized samples are subjected to time domain stretching and compression processing to obtain an enhanced sample set.
7. A gateway busy tone detection device based on deep learning, characterized in that: The device comprises: A signal acquisition module is used to obtain the original audio signal through an analog line interface; A matrix generation module, configured to perform multi-level feature extraction on the original audio signal to generate a time series feature matrix; A model calculation module, configured to input the time series feature matrix into a pre-trained deep learning model and output a busy tone probability value; The model calculation module is also used to collect labeled busy tone and non-busy tone sample data; Performing voltage fluctuation simulation and phase shift enhancement on the sample data to generate an enhanced sample set; Converting the enhanced sample set into a training feature matrix through a time-frequency transformation module; The model is trained using a dynamic decay learning rate strategy, with the initial learning rate set to a preset value; When the accuracy of the validation set does not improve for a preset number of consecutive rounds, the training is terminated and the model parameters are saved to obtain a pre-trained deep learning model; The control module is configured to generate a port control instruction according to a comparison result between the busy tone probability value and a preset threshold.
8. A gateway busy tone detection device based on deep learning, characterized in that: The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the gateway busy tone detection method based on deep learning according to any one of claims 1 to 6.
9. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by the processor, the steps of the gateway busy tone detection method based on deep learning are implemented as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Busy tone detecting method and apparatus
CN101272420A
Method for automatically detecting 'Busy' signal
CN1383315A