An Air Traffic Control Audio Coding Method and System Based on Dynamic Acoustic Masking
By constructing a perceptual saliency map using a personalized auditory feature extraction model and Hadamard product, and combining the basic codebook and tree-structured residual codebook for compression encoding, the problem of traditional models being unable to dynamically suppress background noise and adapt to individual auditory characteristics is solved, thereby improving the reliability and compression efficiency of air traffic control communication instructions.
Patent Information
- Application Number
- CN202511121490.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-12
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2045-08-12
AI Technical Summary
Traditional psychoacoustic models cannot dynamically suppress the masking effect of background noise on speech signals in air traffic control communications, and they are difficult to adapt to individual auditory characteristics in real time, which may lead to the risk of misinterpretation of instructions.
A personalized auditory feature extraction model is adopted. The time-frequency map is generated by complex short-time Fourier transform. The masking matrix is generated by combining physiological features and attention perception branches. The perceptual saliency map is constructed by Hadamard product. The code triples are generated by compression encoding through basic codebook and tree residual codebook.
It enables dynamic perception modeling of individual air traffic controllers' auditory characteristics, enhances the expressive power of key speech components, suppresses background noise interference, improves compression efficiency, and enhances the intelligibility and transmission reliability of aviation communications.
Smart Images

Figure CN120636422B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of audio coding technology, specifically relating to an air traffic control audio coding method and system based on dynamic acoustic masking. Background Technology
[0002] In the field of air traffic control communications, audio coding quality is crucial to flight safety and command transmission efficiency, placing high demands on coding real-time performance, frequency band sensitivity, and noise robustness. While traditional psychoacoustic models consider the characteristics of human hearing, they have significant limitations in aviation scenarios.
[0003] On the one hand, aviation communications face complex acoustic environments such as high-frequency electromagnetic interference and cabin noise, and traditional models cannot dynamically suppress the masking effect of background noise on speech signals. On the other hand, controllers are in a high-intensity auditory environment for a long time, and the differences in individual hearing in high-frequency response and noise tolerance make it difficult for traditional models to adapt to individual auditory characteristics in real time, which may lead to the risk of misjudgment of instructions. Summary of the Invention
[0004] This invention addresses the problems in the prior art by providing an air traffic control audio coding method and system based on dynamic acoustic masking. This solves the problem that traditional models in the prior art cannot dynamically suppress the masking effect of background noise on speech signals. It also solves the problem that traditional models in the prior art are difficult to adapt to individual auditory characteristics in real time, which may lead to the risk of misjudgment of instructions.
[0005] The technical solution adopted in this invention is as follows:
[0006] In a first aspect, this application provides an air traffic control audio coding method based on dynamic acoustic masking, the method comprising the following steps:
[0007] Step S1: Collect the audio signal from air traffic controllers, and process the audio signal to obtain an audio time-frequency diagram;
[0008] Step S2: Input the audio time-frequency map into the personalized auditory feature extraction model to extract features and generate the final masking matrix;
[0009] Step S3: Perform a Hadamard product between the final masking matrix and the audio time-frequency map to obtain the perceptual saliency map;
[0010] Step S4: Extract audio features based on the perceptual saliency map, compress and encode the audio features to obtain the corresponding compressed representation;
[0011] Step S5: Convert the compressed representation into encoded triples, which are then output as the final encoded result of the corresponding audio signal.
[0012] Furthermore, in step S1, the audio signal is subjected to complex short-time Fourier transform processing, and a time-frequency diagram is generated using a complex convolution kernel;
[0013] For input audio signal Time-frequency representation is generated using complex convolution kernels:
[0014]
[0015] in Represents time and frequency The amplitude of the audio signal at that location; It is the first There are learnable complex convolution kernels, where H is the conjugate transpose; Re and Im represent taking the real and imaginary parts, respectively.
[0016] Furthermore, in step S2, the personalized auditory feature extraction model includes a physiological feature branch and an attention perception branch;
[0017] The physiological feature branch extracts feature maps from the audio time-frequency graph;
[0018] The attention-aware branch extracts a dynamic weight matrix from the audio time-frequency graph;
[0019] The feature map and dynamic weight matrix are fused through band attention to generate the final masking matrix.
[0020] Furthermore, the physiological feature branch includes Mel filter banks and 1D convolutions;
[0021] Using the Mel filter bank Feature learning is performed on the time-frequency representation in the audio time-frequency graph;
[0022] The parameters of the Mel filter bank are initialized to a Bark scale distribution, which is adjusted through backpropagation.
[0023] The calculation formula using the Mel filter bank is:
[0024]
[0025] based on Following this are three layers of 1D convolutions and the Mish activation function to obtain the feature map. ;
[0026] The attention perception branch divides the spectrum into 24 critical bands and calculates the energy of each sub-band:
[0027]
[0028] This represents the set of frequencies contained within the b-th critical frequency band, where the value of b is related to the number of critical frequency bands and ranges from 1 to 24.
[0029] For all frequencies in the set Summing is performed to obtain the energy of the critical frequency band subband. After one gated loop unit, the initial dynamic weight matrix is obtained. The final dynamic weight matrix is obtained through two fully connected layers. .
[0030] Furthermore, regarding W att and A prediction masking matrix is generated through frequency band attention fusion, and the fusion method is as follows:
[0031]
[0032] in, It is an activation function. For learnable mixing coefficients, The personalized masking threshold representing the final generated value in time. and frequency The value at that location, To predict the masking matrix;
[0033] The loss function is designed as follows:
[0034]
[0035] Where T is the number of time frames and F is the number of frequency points; ;
[0036] By iteratively training by minimizing this loss function, the final masking matrix learned is obtained. .
[0037] Furthermore, in step S4, a basic codebook is constructed, and the perceptual saliency map is divided into blocks for feature extraction and clustering encoding to obtain the residual vector;
[0038] A tree-structured residual codebook is constructed based on residual vectors. The Euclidean distance between each residual vector and each path node is calculated according to the hierarchical structure. The path with the minimum distance is selected to generate a tree path index. Path matching is terminated when the residual energy is lower than a preset threshold, resulting in a compressed representation.
[0039] Furthermore, before constructing the tree-shaped residual codebook based on the residual vector, features that do not need to be encoded are filtered out. The squares of each dimension of the residual vector are added together to obtain a scalar value. When the scalar value is less than the threshold, the current residual vector is no longer matched and encoded in the subsequent tree-shaped codebook, and the feature is discarded.
[0040] Furthermore, in step S5, the encoded triplet consists of a base index, a path index, and a termination flag, which are used to represent the prototype codebook index, the hierarchical path index, and the dynamic quantization termination flag, respectively.
[0041] Secondly, this application provides an air traffic control audio coding system based on dynamic acoustic masking, the system comprising:
[0042] The audio preprocessing module is used to acquire the audio signals of air traffic controllers and perform complex short-time Fourier transform on the audio signals to obtain the audio time-frequency diagram.
[0043] The personalized masking modeling module is used to input the audio time-frequency map into the personalized auditory feature extraction model. The model includes a physiological feature branch and an attention perception branch, which are used to extract feature maps and dynamic weight matrices, respectively. Based on frequency band attention fusion, a predicted masking matrix is generated. The predicted masking matrix is trained and optimized by minimizing the loss function, and the final masking matrix is output.
[0044] The saliency map generation module is used to perform a Hadamard product between the final masking matrix and the audio time-frequency map to generate a perceptual saliency map;
[0045] The encoding processing module is used to extract audio features based on the perceptual saliency map, construct a basic codebook and a tree-structured residual codebook, encode using a hierarchical vector quantization method, and output a compressed representation;
[0046] The encoding output module is used to convert the compressed representation into an encoded triplet consisting of a base index, a path index, and a termination flag, as the final encoded result of the corresponding audio signal.
[0047] Furthermore, the personalized masking modeling module includes:
[0048] The physiological feature extraction submodule is used to extract features from the audio time-frequency graph. The submodule includes a Mel filter bank and a one-dimensional convolutional layer. The parameters of the Mel filter bank are initialized with a Bark scale distribution and are trained and optimized through backpropagation.
[0049] The attention perception submodule is used to divide the audio spectrum into several critical frequency bands, calculate the energy of each frequency band and generate a dynamic weight matrix. The dynamic weight matrix is output after being processed by a gated recurrent unit and a fully connected layer.
[0050] The fusion computation submodule is used to fuse the feature map and the dynamic weight matrix through the frequency band attention mechanism to generate a prediction mask matrix, and to perform training and optimization based on the set loss function to output the final mask matrix.
[0051] As can be seen from the above technical solutions, the advantages of the present invention are:
[0052] This invention constructs a personalized auditory feature extraction model, fusing physiological features with frequency band attention weights and incorporating a learnable masking threshold generation mechanism to achieve dynamic perceptual modeling tailored to the individual auditory characteristics of air traffic controllers. Compared to traditional fixed psychoacoustic models, this model can adaptively adjust the importance of different frequency information, effectively enhancing the expressive power of key speech components. By constructing a perceptual saliency map using the spectral Hadamard product, effective perceptual regions are highlighted, suppressing background noise interference. In the compression coding stage, by constructing a basic codebook and a tree-structured residual codebook, combined with a dynamic termination mechanism, low-perceptual-weight regions are skipped, significantly reducing redundant data and improving compression efficiency. The overall solution balances high-frequency band recognition accuracy, robustness in complex noise environments, and low-latency compression performance, effectively improving the intelligibility and transmission reliability of aviation communication commands, and is particularly suitable for aviation voice communication scenarios with high interference and significant individual differences. Attached Figure Description
[0053] To more clearly illustrate the technical solution of the present invention, the accompanying drawings used in the description will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0054] Figure 1 This is a flowchart illustrating the steps of the air traffic control audio coding method based on dynamic acoustic masking in this embodiment. Detailed Implementation
[0055] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0056] Please see Figure 1 As shown, this invention provides an air traffic control audio coding method based on dynamic acoustic masking, comprising the following steps:
[0057] Step S1: Collect the audio signal from air traffic controllers, and process the audio signal to obtain an audio time-frequency diagram;
[0058] Collecting audio signals from air traffic controllers Then, it is divided evenly, and the length of each segment is denoted as N;
[0059] In air traffic control audio coding, bit allocation strategies directly affect the intelligibility and noise resistance of air-to-ground instructions. Traditional fixed bit allocation modes cannot dynamically enhance key frequency bands for aviation communications, are easily affected by cabin noise or electromagnetic interference, and cause critical voice signals to be easily submerged; furthermore, they do not adapt to the long-term professional hearing characteristics of controllers, resulting in insufficient coding accuracy for high-frequency instructions.
[0060] Therefore, in this embodiment, a DPM-Net network is designed to achieve dynamic bit allocation based on human auditory perception characteristics. The input to this network is the raw audio signal. The output is a dynamic masking threshold matrix in the time-frequency domain. To address the problem that traditional acoustic models cannot adapt to air traffic control noise, an adjustable auditory feature extractor and an attention gating mechanism are proposed.
[0061] Each audio segment is processed by a complex short-time Fourier transform, and a time-frequency graph is generated using a complex convolution kernel;
[0062] Input audio signal The image is converted to a time-frequency graph using a complex convolution kernel STFT (kernel size of 1024 points, frame shift hop = 256), and its dimensions are as follows: For the input audio signal Time-frequency representation is generated using a complex convolution kernel STFT:
[0063]
[0064] in Represents time and frequency The intensity or amplitude of the audio signal. It is the first There are learnable complex convolution kernels (f=1,2,...,513), where H is the conjugate transpose. Re and Im represent extracting the real and imaginary parts, respectively, where the real part filter extracts the frequency... In-phase audio components, imaginary part filter extraction and frequency The audio components with orthogonal phase. This ultimately yields the time-frequency representation. .
[0065] Step S2: Construct a dynamic acoustic masking network. Input the audio time-frequency map into the dynamic acoustic masking network, extract individual auditory features and frequency band attention weights, and train and optimize based on a preset loss function. Then, fuse the two to generate a personalized masking threshold matrix.
[0066] Step S2: Input the audio time-frequency map into the personalized auditory feature extraction model to extract features and generate the final masking matrix;
[0067] Physiological feature branch: This branch includes a Mel filter bank (64 channels) and 1D convolution (kernel=5), yielding a size of The feature map. Specifically, firstly, the Mel filter bank is used. Feature learning is performed on the time-frequency representation. The parameters of the Mel filter bank are initialized with a Bark-scale distribution and can be adjusted through backpropagation. The Bark scale is a frequency scale based on the characteristics of human hearing, which better matches the human ear's perception of sound frequencies. The calculation formula using the Mel filter bank is:
[0068]
[0069] Then, based on Following this are three 1D convolutional layers (kernel=5, groups=64) and the Mish activation function, finally yielding the feature map. The size is .
[0070] Attention Perception Branch: The human ear's perception of sounds at different frequencies is not linear; the critical frequency band reflects the ear's ability to distinguish frequencies. Therefore, the frequency spectrum is divided into 24 critical frequency bands, and the energy of each sub-band is calculated:
[0071]
[0072] This represents the set of frequencies contained within the b-th critical band. The range of values for b is related to the number of critical bands; here, the spectrum is divided into 24 critical bands, so b ranges from 1 to 24. Each frequency within this set... Corresponding time-frequency representation The frequency dimension information in the set is obtained by analyzing the frequencies corresponding to all frequencies within this set. By summing the results, we can obtain the energy of the subband of the critical frequency band. Then, after passing through one gated recurrent unit (GRU), the initial dynamic weight matrix is obtained. Its size is Then, through two fully connected layers FC(256) and FC(513), the 24-dimensional weights are mapped to 513 dimensions to obtain the final dynamic weight matrix. .
[0073] In some embodiments, for W att and A prediction masking matrix is generated through frequency band attention fusion, and the fusion method is as follows:
[0074]
[0075] in, It is an activation function. For learnable mixing coefficients, The personalized masking threshold representing the final generated value in time. and frequency The value at that location, To predict the masking matrix, the size of the masking matrix is ;
[0076] The loss function is designed as follows:
[0077]
[0078] Where T is the number of time frames, i.e., N / 256, and F=513 is the number of frequency points; The second term is a sparsity (L1 regularization) constraint, which is applied to the predicted masking matrix. Applying L1 regularization encourages most values in the masking matrix to be 0, activating only at key frequencies and time points, thus simulating the locality of human auditory masking while reducing the risk of overfitting.
[0079] The above model is trained 200 times to obtain a well-learned masking matrix, denoted as . .
[0080] Step S3: Perform a Hadamard product between the final masking matrix and the audio time-frequency map to obtain the perceptual saliency map;
[0081] We design a hierarchical residual vector quantization network, HRVQ-Net, to achieve efficient and low-loss neural code compression. The input of this network is a masking matrix. With time-frequency graph The output is a compressed bitstream. To reduce codebook size and training complexity, a tree-structured residual codebook and an adaptive hierarchical quantization strategy are proposed.
[0082] Specifically, the original time-frequency map is multiplied by the masking matrix using the Hadamard product (i.e., pointwise multiplication) to generate a perceptual saliency map, namely:
[0083]
[0084] in , The dimension of A is ,Right now .
[0085] Step S4: Extract audio features based on the perceptual saliency map, and compress and encode the features using a hierarchical vector quantization method to obtain the compressed representation corresponding to each segment;
[0086] A basic codebook is constructed, and after dividing the perceptual saliency map into blocks, feature extraction and clustering encoding are performed to obtain the residual vector.
[0087] A tree-structured residual codebook is constructed based on residual vectors. The Euclidean distance between each residual vector and each path node is calculated according to the hierarchical structure. The path with the minimum distance is selected to generate a tree path index. Path matching is terminated when the residual energy is lower than a preset threshold, and a compressed representation is obtained.
[0088] The hierarchical encoding process is designed, where the first layer divides the perceptual saliency map into 16×16 sub-blocks, totaling 256. Considering that 513 is not divisible by 16, and N / 256 may not be divisible by 16, the last block is divided by overlapping with other blocks. Let the i-th sub-block be represented as... The feature representation is obtained after passing through a 256-dimensional fully connected layer. , i=1,...,256.
[0089] Based on feature h i A basic codebook C is constructed using clustering methods. Specifically, for N samples in the training set, features h are extracted using the methods described above. i Then, k-means clustering is performed: 64 features are randomly selected from N*256 features as initial cluster centers. The distance from each feature to these cluster centers is then calculated, and each feature is assigned to the nearest cluster center, resulting in 64 clusters. The mean value of each cluster is updated as its new center, and the process is iterated. This process is repeated 100 times to obtain stable cluster centers, which serve as the codebase. , This is the prototype vector. Calculate... With each prototype vector Euclidean distance Find the prototype vector index with the smallest distance. The residual vector is obtained. .
[0090] In the second layer, a tree-like hierarchical structure is used to construct the residual codebook based on the residual vector, with each layer having a splitting factor of 1. Depth is This forms a tree with 3 levels and 4 branches per level. There are 4^3 = 64 distinct paths from the root to a leaf node, and each path corresponds to a leaf node. Let the set of node vectors on the path from the root node to a leaf node be the node set. , ( Representing different paths, with values ranging from 1 to 64. Each node vector is 256-dimensional, therefore... It is a set of three 256-dimensional vectors. For the residual vector... Calculate its L2 distance with each path node vector. ,in, For residual vectors The k-th dimension element; The node vector of the kth layer of path l The k-th element represents the path. Each node at each level filters the residual vector; the smaller the cumulative distance, the better the overall match between the path's nodes and the residual vector. The path with the smallest distance is selected as the optimal path, and a hierarchical index is generated. The format is: n i =(a,b,c), where a represents the branch index from the root node to the intermediate layer, b represents the branch index from the intermediate layer to the child intermediate layer, and c represents the branch index from the child intermediate layer to the leaf layer, with values of 1, 2, 3, and 4.
[0091] In some embodiments, a dynamic depth control method is constructed to filter features that do not require encoding. Let the residual energy... That is, the residual vector The squares of each dimension are summed to obtain a scalar value. A larger value indicates more information in the residual; a smaller value indicates the residual is closer to noise or negligible details. The masking threshold is... Here it is defined as 0.01, when When this happens, the current residual vector is no longer matched or encoded using a tree-based codebook, and the feature is discarded directly.
[0092] Step S5: Convert the compressed representation into encoded triples, which are then used as the final encoded result for the corresponding audio segment and output.
[0093] The encoded triple consists of a base index, a path index, and a termination flag, which are used to represent the prototype codebook index, the hierarchical path index, and the dynamic quantization termination flag, respectively.
[0094] The final output is a triplet encoding. That is, (basic codebook index + tree path index + termination flag), where m i The prototype vector index in the basic codebook can be encoded using 6 bits, n i It is the path index of the tree-structured residual codebook (64 paths in total, which can be encoded with 6 bits). The termination flag is encoded with 1 bit to determine whether quantization is terminated early (1 indicates termination, 0 indicates no early termination). The basic codebook index and the tree path index record the mapping relationship of audio features in the codebook in a compact way, and the termination flag is used to dynamically control the quantization process and reduce redundant operations.
[0095] For a given input speech for testing, it is segmented into segments of length N. For each segment, its feature vector h is obtained according to the steps described above. i Then, the cluster centers (obtained from the training set) are calculated, and their residual vectors are obtained. Encoding is then performed to obtain the final encoded output. The above operations are performed synchronously on all segments to obtain the final speech code.
[0096] Using the above scheme, a segment of length N is encoded into 256 feature vectors h. i Thus, the input audio information is ultimately represented in a highly compressed form. This encoding method can effectively reduce the bit rate and reduce the space and bandwidth required for data storage and transmission.
[0097] In some embodiments, this application provides an air traffic control audio coding system based on dynamic acoustic masking, the system comprising:
[0098] The audio preprocessing module is used to acquire the audio signals of air traffic controllers and perform complex short-time Fourier transform on the audio signals to obtain the audio time-frequency diagram.
[0099] The personalized masking modeling module is used to input the audio time-frequency map into the personalized auditory feature extraction model. The model includes a physiological feature branch and an attention perception branch, which are used to extract feature maps and dynamic weight matrices, respectively. Based on frequency band attention fusion, a predicted masking matrix is generated. The predicted masking matrix is trained and optimized by minimizing the loss function, and the final masking matrix is output.
[0100] The saliency map generation module is used to perform a Hadamard product between the final masking matrix and the audio time-frequency map to generate a perceptual saliency map;
[0101] The encoding processing module is used to extract audio features based on the perceptual saliency map, construct a basic codebook and a tree-structured residual codebook, encode using a hierarchical vector quantization method, and output a compressed representation;
[0102] The encoding output module is used to convert the compressed representation into an encoded triplet consisting of a base index, a path index, and a termination flag, as the final encoded result of the corresponding audio signal.
[0103] In some embodiments, the personalized masking modeling module includes:
[0104] The physiological feature extraction submodule is used to extract features from the audio time-frequency graph. The submodule includes a Mel filter bank and a one-dimensional convolutional layer. The parameters of the Mel filter bank are initialized with a Bark scale distribution and are trained and optimized through backpropagation.
[0105] The attention perception submodule is used to divide the audio spectrum into several critical frequency bands, calculate the energy of each frequency band and generate a dynamic weight matrix. The dynamic weight matrix is output after being processed by a gated recurrent unit and a fully connected layer.
[0106] The fusion computation submodule is used to fuse the feature map and the dynamic weight matrix through the frequency band attention mechanism to generate a prediction mask matrix, and to perform training and optimization based on the set loss function to output the final mask matrix.
[0107] In some embodiments, this application provides a terminal, including:
[0108] Memory for storing air traffic control audio coding programs based on dynamic acoustic masking;
[0109] A processor is configured to implement the steps of the dynamic acoustic masking-based air traffic control audio coding method when executing the dynamic acoustic masking-based air traffic control audio coding system.
[0110] In some embodiments, this application provides a computer-readable storage medium that stores computer instructions. When a computer reads the computer instructions in the storage medium, the computer executes the air traffic control audio coding method based on dynamic acoustic masking.
[0111] It is understood that the systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer, which can be a personal computer, a laptop computer, a personal digital assistant, a tablet computer, a wearable device, or any combination of these devices.
[0112] In a typical configuration, a computer includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0113] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0114] Computer-readable media, including both permanent and non-permanent, removable and non-removable media, can store information using any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, disk storage, quantum memory, graphene-based storage media or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined in this embodiment, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0115] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0116] It should be understood that although the terms first, second, third, etc., may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first information may also be referred to as second information without departing from the scope of one or more embodiments of this specification, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "in response to a determination," or "when," or "in the event of a determination."
[0117] The above description is merely a preferred embodiment of one or more embodiments of this specification and is not intended to limit the scope of one or more embodiments of this specification. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of one or more embodiments of this specification should be included within the protection scope of one or more embodiments of this specification.
Claims
1. An air traffic control audio coding method based on dynamic acoustic masking, characterized in that, Includes the following steps: Step S1: Collect the audio signal from air traffic controllers, and process the audio signal to obtain an audio time-frequency diagram; Step S2: Input the audio time-frequency map into the personalized auditory feature extraction model to extract features and generate the final masking matrix; Personalized auditory feature extraction models include physiological feature branches and attention perception branches; The physiological feature branch extracts feature maps from the audio time-frequency graph; The attention-aware branch extracts a dynamic weight matrix from the audio time-frequency graph; The feature map and dynamic weight matrix are fused through frequency band attention to generate the final masking matrix; The physiological feature branch includes Mel filter banks and 1D convolutions; Using the Mel filter bank Feature learning is performed on the time-frequency representation in the audio time-frequency graph; The parameters of the Mel filter bank are initialized to a Bark scale distribution, which is adjusted through backpropagation. The calculation formula using the Mel filter bank is: based on Following this are three layers of 1D convolutions and the Mish activation function to obtain the feature map. ; The attention perception branch divides the spectrum into 24 critical bands and calculates the energy of each sub-band: This represents the set of frequencies contained within the b-th critical frequency band, where the value of b is related to the number of critical frequency bands and ranges from 1 to 24. For all frequencies in the set Summing is performed to obtain the energy of the critical frequency band subband. ; After one gated loop unit, the initial dynamic weight matrix is obtained. The final dynamic weight matrix is obtained through two fully connected layers. ; Step S3: Perform a Hadamard product between the final masking matrix and the audio time-frequency map to obtain the perceptual saliency map; Step S4: Extract audio features based on the perceptual saliency map, compress and encode the audio features to obtain the corresponding compressed representation; Step S5: Convert the compressed representation into encoded triples, which are then output as the final encoded result of the corresponding audio signal. The encoded triple consists of a base index, a path index, and a termination flag, which are used to represent the prototype codebook index, the hierarchical path index, and the dynamic quantization termination flag, respectively.
2. The air traffic control audio coding method based on dynamic acoustic masking according to claim 1, characterized in that, In step S1, the audio signal is subjected to complex short-time Fourier transform processing, and a time-frequency diagram is generated using a complex convolution kernel; For input audio signal Time-frequency representation is generated using complex convolution kernels: in Represents time and frequency The amplitude of the audio signal at that location; It is the first There are learnable complex convolution kernels, where H is the conjugate transpose; Re and Im represent taking the real and imaginary parts, respectively.
3. The air traffic control audio coding method based on dynamic acoustic masking according to claim 1, characterized in that, For W att and A prediction masking matrix is generated through frequency band attention fusion, and the fusion method is as follows: in, It is an activation function. For learnable mixing coefficients, The personalized masking threshold representing the final generated value in time. and frequency The value at that location, To predict the masking matrix; The loss function is designed as follows: Where T is the number of time frames and F is the number of frequency points; ; By iteratively training by minimizing this loss function, the final masking matrix learned is obtained. .
4. The air traffic control audio coding method based on dynamic acoustic masking according to claim 3, characterized in that, In step S4, a basic codebook is constructed, and the perceptual saliency map is divided into blocks for feature extraction and clustering encoding to obtain the residual vector. A tree-structured residual codebook is constructed based on residual vectors. The Euclidean distance between each residual vector and each path node is calculated according to the hierarchical structure. The path with the minimum distance is selected to generate a tree path index. Path matching is terminated when the residual energy is lower than a preset threshold, resulting in a compressed representation.
5. The air traffic control audio coding method based on dynamic acoustic masking according to claim 4, characterized in that, Before constructing a tree-structured residual codebook based on the residual vector, features that do not need to be encoded are filtered out. The squares of each dimension of the residual vector are summed to obtain a scalar value. When the scalar value is less than the threshold, the current residual vector is no longer matched and encoded in the subsequent tree-structured codebook, and the feature is discarded.
6. An air traffic control audio coding system based on dynamic acoustic masking, characterized in that, The system includes: The audio preprocessing module is used to acquire the audio signals of air traffic controllers and perform complex short-time Fourier transform on the audio signals to obtain the audio time-frequency diagram. The personalized masking modeling module is used to input the audio time-frequency map into the personalized auditory feature extraction model. The model includes a physiological feature branch and an attention perception branch, which are used to extract feature maps and dynamic weight matrices, respectively. Based on frequency band attention fusion, a predicted masking matrix is generated. The predicted masking matrix is trained and optimized by minimizing the loss function, and the final masking matrix is output. The personalized masking modeling module includes: The physiological feature extraction submodule is used to extract features from the audio time-frequency graph. The submodule includes a Mel filter bank and a one-dimensional convolutional layer. The parameters of the Mel filter bank are initialized with a Bark scale distribution and are trained and optimized through backpropagation. The attention perception submodule is used to divide the audio spectrum into several critical frequency bands, calculate the energy of each frequency band and generate a dynamic weight matrix. The dynamic weight matrix is output after being processed by a gated recurrent unit and a fully connected layer. The fusion computation submodule is used to fuse the feature map and the dynamic weight matrix through the frequency band attention mechanism to generate the prediction mask matrix, and to perform training and optimization based on the set loss function to output the final mask matrix. The saliency map generation module is used to perform a Hadamard product between the final masking matrix and the audio time-frequency map to generate a perceptual saliency map; The encoding processing module is used to extract audio features based on the perceptual saliency map, construct a basic codebook and a tree-structured residual codebook, encode using a hierarchical vector quantization method, and output a compressed representation; The encoding output module is used to convert the compressed representation into an encoded triplet consisting of a base index, a path index, and a termination flag, as the final encoded result of the corresponding audio signal.
Citation Information
Patent Citations
Sound encoder and sound encoding method
CN101044552A
Bit rate allocation method and bit rate allocation device in digital audio coding
CN106653035A