Audio signal selection method and apparatus, related device, and signal reception system
By extracting and decoding the acoustic features of audio signals, the problem of inaccurate audio signal quality assessment caused by manual listening is solved, and accurate audio signal selection based on multi-dimensional factors is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HEFEI IFLY DIGITAL TECH CO LTD
- Filing Date
- 2023-07-05
- Publication Date
- 2026-07-21
AI Technical Summary
Existing audio signal optimization schemes rely on human listening, resulting in inaccurate quality assessments and varying effects from person to person. They also fail to fully consider influencing factors such as microphone cavity structure, signal resampling, echo, noise, reverberation, gain, and encoding/decoding.
By extracting acoustic features of candidate audio signals, including multi-dimensional information such as microphone cavity structure, signal sampling rate, echo, noise, reverberation, gain, and encoding/decoding, the target audio signal is decoded and restored using a self-supervised learning model or acoustic encoder. Based on the acoustic features, the quality is determined and audio signals that meet the set conditions are selected.
It achieves accuracy and consistency in audio signal quality assessment, eliminates the influence of human cognitive differences, and can comprehensively consider multi-dimensional factors to select the audio signal with the best quality.
Smart Images

Figure CN116682461B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of signal quality assessment technology, and more specifically, to an audio signal selection method, apparatus, related equipment, and signal receiving system. Background Technology
[0002] In some communication scenarios, audio signal optimization is often involved, which means selecting the audio signal with the best quality from several audio signals for subsequent use.
[0003] Taking shortwave and ultra-shortwave wireless communication scenarios as an example, multiple signal receivers located at different positions can be used to collect signals from the same signal source. Due to variations in the distance and direction between the signal source and the receivers, as well as changes in the electromagnetic environment such as weather, surrounding terrain, interference sources, and multipath effects during signal reception, the audio signals received by different receivers are complex and varied, with inconsistent quality. Therefore, it is necessary to select the highest quality audio signal from the audio signals received by multiple receivers for subsequent use.
[0004] Existing audio signal optimization schemes typically rely on experienced listening personnel to select the best signal. However, human listening is susceptible to individual differences in the listener's perception, leading to inconsistent optimization results. Furthermore, many factors contribute to variations in audio signal quality, such as microphone cavity structure, signal resampling, echo, noise, reverberation, gain, encoding / decoding, and distortion caused by packet loss. Relying solely on human listening cannot comprehensively consider all these influencing factors, resulting in inaccurate audio signal quality assessments. Summary of the Invention
[0005] In view of the above problems, this application is made to provide an audio signal selection method, apparatus, related equipment, and signal receiving system, so as to improve the accuracy of audio signal quality assessment and thus select the audio signal with the best quality. The specific solution is as follows:
[0006] Firstly, an audio signal selection method is provided, including:
[0007] Obtain an audio signal set, wherein the audio signal set contains at least one candidate audio signal;
[0008] The acoustic features of each candidate audio signal are extracted, and the acoustic features are such that the target audio signal can be decoded based on them. The target audio signal is close to or equivalent to the candidate audio signal.
[0009] The quality of each candidate audio signal is determined based on its acoustic characteristics.
[0010] Based on the quality of each candidate audio signal, the candidate audio signal that meets the set quality conditions is selected as the audio signal selection result.
[0011] Secondly, a signal receiving system is provided, including: a plurality of signal receivers and processing terminals;
[0012] Each of the aforementioned signal receivers acquires audio signals from the same signal source and sends them to the processing terminal to form an audio signal set;
[0013] The processing terminal uses the aforementioned audio signal selection method to select audio signals that meet the set quality conditions from the set of audio signals.
[0014] Thirdly, an audio signal selection device is provided, comprising:
[0015] A signal set acquisition unit is used to acquire an audio signal set, wherein the audio signal set contains at least one candidate audio signal;
[0016] An acoustic feature extraction unit is used to extract acoustic features of each candidate audio signal. The acoustic features are those that can be used to decode the target audio signal, and the target audio signal is close to or equivalent to the candidate audio signal.
[0017] A quality determination unit is configured to determine the quality of each candidate audio signal based on the acoustic characteristics of each candidate audio signal.
[0018] The quality selection unit is used to select candidate audio signals that meet the set quality conditions by referring to the quality of each candidate audio signal, and use them as the audio signal selection result.
[0019] Fourthly, an audio signal selection device is provided, including: a memory and a processor;
[0020] The memory is used to store programs;
[0021] The processor is used to execute the program to implement each step of the aforementioned audio signal selection method.
[0022] Fifthly, a storage medium is provided on which a computer program is stored, which, when executed by a processor, implements the various steps of the audio signal selection method as described above.
[0023] By employing the above technical solution, this application acquires each candidate audio signal and extracts the acoustic features of each candidate audio signal. These acoustic features are used as a basis for decoding the acoustic features of the target audio signal, wherein the target audio signal is close to or equivalent to the candidate audio signal. Since the acoustic features extracted from the candidate audio signals can be decoded to reconstruct a target audio signal that is close to or equivalent to the candidate audio signal, the acoustic features extracted from the candidate audio signals contain intrinsic information of various dimensions of the candidate audio signals, such as: microphone cavity structure, signal sampling rate, echo, noise, reverberation, gain, encoding / decoding, packet loss, and other multi-dimensional information. Only in this way can the original candidate audio signal be decoded and reconstructed based on the rich intrinsic information of each dimension. It is evident that the intrinsic information of each dimension contained in the extracted acoustic features necessarily includes influencing factors affecting the quality of the audio signal in various dimensions. Based on this, the quality of the candidate audio signal can be determined based on the acoustic features, and by referring to the quality of each candidate audio signal, the candidate audio signal that meets the set quality conditions can be selected as the final selection result. This application overcomes the problem of inconsistent selection results due to human factors in existing methods of selecting the best signal through manual listening. Furthermore, by extracting acoustic features that contain intrinsic information of candidate audio signals in various dimensions, it can comprehensively consider all factors affecting the quality of audio signals. Based on this, the quality of the candidate audio signals determined is more accurate and is not affected by differences in human perception. Attached Figure Description
[0024] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the scope of this application. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:
[0025] Figure 1 This is a schematic flowchart of an audio signal selection method provided in an embodiment of this application;
[0026] Figure 2 This is a schematic diagram of a signal receiving system structure provided in an embodiment of this application;
[0027] Figure 3 This example illustrates a flowchart for selecting the highest quality audio signal from multiple audio signals acquired by multiple signal receivers.
[0028] Figure 4 This is a schematic diagram of an audio signal selection device provided in an embodiment of this application;
[0029] Figure 5 This is a schematic diagram of the structure of the audio signal selection device provided in the embodiments of this application. Detailed Implementation
[0030] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0031] This application provides an audio signal selection scheme applicable to various tasks involving selecting the highest quality audio signal from multiple candidate audio signals. For example, in scenarios where multiple signal receivers simultaneously acquire audio signals from a signal source, the method described in this application can be used to determine the signal quality of each candidate audio signal received by each receiver, and then, based on the determined signal quality, select the highest quality audio signal for subsequent processing.
[0032] The proposed solution can be implemented based on a terminal with data processing capabilities, such as a mobile phone, computer, server, or cloud platform.
[0033] Next, combined Figure 1 The audio signal selection method of this application may include the following steps:
[0034] Step S100: Obtain an audio signal set, wherein the audio signal set contains at least one candidate audio signal.
[0035] Specifically, the candidate audio signals included in the audio signal set can be audio signals for subsequent selection. Depending on the applicable scenario, the number, source, and band of the candidate audio signals included in the audio signal set can also be different.
[0036] Taking a scenario where multiple different signal receivers are acquiring signals from the same signal source as an example, the audio signals acquired by each signal receiver can be used as candidate audio signals and added to the audio signal set. The signal receivers can receive signals from various frequency bands, such as shortwave, VHF, and longwave signals.
[0037] Step S110: Extract the acoustic features of each candidate audio signal. The acoustic features are those that can be used to decode the target audio signal. The target audio signal is close to or equivalent to the candidate audio signal.
[0038] Specifically, unlike traditional acoustic feature extraction, this application designs an acoustic feature extraction method that provides a feature extraction direction, namely, extracting acoustic features that can be used as a basis for decoding the target audio signal. Here, the target audio signal can be close to or equivalent to the candidate audio signal; that is, the closer the target audio signal is to the candidate audio signal, the better. The best case is that the target audio signal is equivalent to the candidate audio signal.
[0039] This setup allows the extracted acoustic features from the candidate audio signal to be decoded and reconstructed into a target audio signal that is close to or equivalent to the candidate audio signal. In other words, the extracted acoustic features contain intrinsic information from various dimensions of the candidate audio signal, such as microphone cavity structure, signal sampling rate, echo, noise, reverberation, gain, encoding / decoding, packet loss, and other multi-dimensional information. Only in this way can the original candidate audio signal be decoded and reconstructed based on this rich intrinsic information. Clearly, the intrinsic information contained in the extracted acoustic features must include various influencing factors affecting audio signal quality; that is, the acoustic features extracted in this step contain rich feature information for measuring audio signal quality.
[0040] Step S120: Determine the quality of each candidate audio signal based on the acoustic characteristics of each candidate audio signal.
[0041] Specifically, in the previous step, corresponding acoustic features were extracted for each candidate audio signal in the audio signal set. In this step, the quality of the candidate audio signal is determined by referring to the acoustic features of each candidate audio signal. Since acoustic features contain various influencing factors that affect the quality of audio signals, this step can obtain the quality of candidate audio signals more comprehensively and accurately based on acoustic features.
[0042] Step S130: Refer to the quality of each candidate audio signal and select the candidate audio signal that meets the set quality conditions as the audio signal selection result.
[0043] Specifically, quality conditions can be set according to task requirements. For example, the quality condition can be to select the candidate audio signal with the highest quality as the selection result, or to set a quality threshold and then select the candidate audio signal with a quality exceeding the quality threshold as the selection result, etc.
[0044] The speech signal selection method provided in this application acquires each candidate audio signal and extracts the acoustic features of each candidate audio signal. These acoustic features are used to decode the target audio signal, where the target audio signal is close to or equivalent to the candidate audio signal. Since the acoustic features extracted from the candidate audio signals can be decoded to reconstruct a target audio signal that is close to or equivalent to the candidate audio signal, the acoustic features extracted from the candidate audio signals contain intrinsic information of various dimensions of the candidate audio signals, such as microphone cavity structure, signal sampling rate, echo, noise, reverberation, gain, encoding / decoding, packet loss, and other multi-dimensional information. This allows for the decoding and reconstruction of the original candidate audio signal based on the rich intrinsic information of each dimension. It is evident that the intrinsic information of each dimension contained in the extracted acoustic features necessarily includes influencing factors affecting the quality of the audio signal. Based on this, the quality of the candidate audio signals can be determined based on these acoustic features, and by referring to the quality of each candidate audio signal, a candidate audio signal that meets the set quality conditions can be selected as the final selection result. This application overcomes the problem of inconsistent selection results due to human factors in existing methods of selecting the best signal through manual listening. Furthermore, by extracting acoustic features that contain intrinsic information of candidate audio signals in various dimensions, it can comprehensively consider all factors affecting the quality of audio signals. Based on this, the quality of the candidate audio signals determined is more accurate and is not affected by differences in human perception.
[0045] Furthermore, in certain scenarios, the propagation process of a signal source is easily affected by factors such as geographical environment, weather phenomena, and other electromagnetic devices, leading to frequent changes in the signal quality received by the signal receiver. This makes it impossible to perform quality assessment based on long-term signals. For example, traditional methods that assess audio signal quality by statistically analyzing the effective speech duration of a long audio signal are not applicable in these scenarios. However, the solution proposed in this application can perform quality assessment based on short-term audio signals, meaning that the solution is applicable to a wider range of scenarios.
[0046] In some embodiments of this application, taking a scenario where multiple different signal receivers acquire signals from the same signal source as an example, when the signal source is transmitting in an open space, signal receivers deployed at different locations can all receive the audio signal, forming an audio signal set. However, due to the different distances between the different signal receivers and the signal source, and the influence of multipath effects, the receiving clocks for the same content may be out of sync, resulting in a certain difference in the initial time of the audio signals received by different signal receivers.
[0047] To eliminate the impact of potential temporal differences between different candidate audio signals in the audio signal set on signal quality assessment and to ensure the consistency of content among the candidate audio signals, this embodiment can add the following processing before extracting the acoustic features of each candidate audio signal in the aforementioned step S110:
[0048] The candidate audio signals in the audio signal set are time-aligned. This ensures that step S110 extracts the acoustic features of each time-aligned candidate audio signal.
[0049] The process of time-aligning each candidate audio signal can be performed using the DTW (Dynamic Time Warping) algorithm or other optional time alignment algorithms.
[0050] During time alignment, one candidate audio signal can be selected as the reference from among the candidate audio signals, and the remaining candidate audio signals can be time aligned to obtain a set of time-aligned audio signals.
[0051] This embodiment performs time alignment on each candidate audio signal in the audio signal set before acoustic feature extraction, ensuring the consistency of content among the processed candidate audio signals. Based on this, the quality of the candidate audio signals obtained after acoustic feature extraction and quality assessment is more accurate.
[0052] Step S110 in the above embodiments of this application introduces an extraction approach for acoustic features of candidate audio signals, namely, extracting acoustic features that can be used to decode the target audio signal (the target audio signal is close to or equivalent to the candidate audio signal) as the direction of acoustic feature extraction.
[0053] Guided by this extraction approach, this application presents an optional implementation method, as follows:
[0054] In this embodiment, a self-supervised learning model can be pre-trained, which includes an acoustic encoder and a decoder. The acoustic encoder is used to extract acoustic features from the input training audio signal, and the decoder decodes the recovered audio signal based on the acoustic features extracted by the acoustic encoder. To achieve self-supervised learning, this application updates the network parameters of the acoustic encoder and decoder with the goal of making the recovered audio signal approximate the training audio signal.
[0055] After the model converges, a trained acoustic encoder is obtained, which can be used to extract acoustic features from candidate audio signals. Furthermore, since the acoustic encoder is trained jointly with the decoder, according to the above training objectives, it can be guaranteed that the acoustic features extracted by the acoustic encoder can be used to reconstruct a target audio signal that is close to or even the same as the candidate audio signal after decoding.
[0056] In addition to the above, this application may also adopt other optional implementation methods, such as:
[0057] A learning model consisting of an acoustic encoder and a task processing module can be pre-trained. This task processing module can be used to perform tasks such as speaker recognition. Taking the speaker recognition module as an example:
[0058] During training, the acoustic encoder extracts acoustic features from the input training audio signal. The speaker recognition processing module predicts the speaker's identity corresponding to the training audio signal based on the acoustic features extracted by the acoustic encoder. Further, a loss is calculated based on the predicted speaker identity and the speaker identity labels annotated in the training audio signal. Minimizing this loss is the training objective, and the network parameters of the acoustic encoder and the speaker recognition processing module are updated until training is complete. This results in a trained acoustic encoder used to extract acoustic features from candidate audio signals. Following this training method, the acoustic features extracted by the acoustic encoder can be used to identify the speaker corresponding to the input candidate audio signal. However, to identify the speaker in a candidate audio signal, it is essential to use some intrinsic features from the candidate audio signal; that is, the acoustic features contain some intrinsic features from the candidate audio signal. Based on this, decoding these acoustic features can also reconstruct a target audio signal that approximates the candidate audio signal.
[0059] Of course, the above are just two examples of implementation methods. Under the guidance of the extraction idea in step S110, other optional implementation methods can also be adopted, which will not be listed one by one in this application.
[0060] Furthermore, this application describes the process of extracting acoustic features of each candidate audio signal using a pre-trained acoustic encoder.
[0061] This application can use a pre-trained acoustic encoder to encode and model each candidate audio signal in different frequency bands, and obtain latent variables as the acoustic features of the candidate audio signal.
[0062] The latent variables contain intrinsic information of the candidate audio signal in each dimension.
[0063] This application uses an acoustic encoder to encode and model candidate audio signals across different frequency bands, achieving decoupling of intrinsic information across different dimensions of the frequency division scale. That is, different dimensions of latent variables correspond to different frequency bands, and intrinsic information of different dimensions can be represented by values of different dimensions in the latent variables. For example, intrinsic information of different dimensions such as microphone cavity structure, sampling rate, echo, noise, reverberation, encoding / decoding, and packet loss can be represented by values of different dimensions in the latent variables. An illustrative example is that the latent variables are represented as z = {z1, z2, ..., z...}. m}, where z1-z2 represent acoustic cavity structure information, z3-z5 represent sampling rate information, and z m This indicates packet loss information, etc.
[0064] Since latent variables implement frequency-decoupled encoding of intrinsic information in different dimensions of the candidate audio signal, using latent variables as acoustic features of the candidate audio signal allows us to obtain acoustic features after decoupling intrinsic information in different dimensions. Based on these acoustic features, the quality of the candidate audio signal can be accurately predicted.
[0065] Optionally, the acoustic encoder described above can be a variational autoencoder (VAE-encoder) or other types of encoders.
[0066] In some embodiments of this application, the process of determining the quality of each candidate audio signal based on the acoustic characteristics of each candidate audio signal in step S120 is further described.
[0067] In this embodiment of the application, an audio quality scoring model can be pre-trained. This model can be trained using the acoustic features of the training audio signal as training samples and the quality scores labeled on the training audio signal as sample labels.
[0068] Specifically, training audio signal data T = {t1, t2, ..., t} can be collected. n For each training audio signal t in the training audio signal data, its quality score is assigned. This quality score can be calibrated by experts or by other methods.
[0069] After quality scoring is calibrated, training data pairs D = {(t1,m1),(t2,m2),…,(t...} can be obtained. n ,m n )}, where m represents the quality score corresponding to the training audio signal.
[0070] By obtaining the acoustic features z of each training audio signal, the training data pair D can be obtained. T ={(z1,m1),(z2,m2),…,(z n ,m n )}.
[0071] Furthermore, based on the training data, D T Train the audio quality scoring model.
[0072] The audio quality scoring model can use a time-delay neural network structure (TDNN) or other types of network structures, such as CNN and RNN.
[0073] In some embodiments of this application, a signal receiving system is further provided, combined with Figure 2 As shown, the system may include several signal receivers 10 and processing terminals 20.
[0074] Each signal receiver 10 collects audio signals from the same signal source and sends them to the processing terminal 20 to form an audio signal set.
[0075] The processing terminal 20 can use the audio signal selection method of the aforementioned embodiments to select audio signals that meet the set quality conditions from the audio signal set for subsequent processing.
[0076] like Figure 3 It illustrates a flowchart of selecting the highest quality audio signal from multiple audio signals acquired by multiple signal receivers.
[0077] There are p signal receivers that collect audio signals.
[0078] After the audio signals collected by each signal receiver are clock-aligned, the time-synchronized audio signals corresponding to each signal receiver are obtained.
[0079] After each time synchronization, the audio signal is processed by an audio encoder to extract acoustic features, resulting in acoustic features (x1, x2, ..., x...). p ), where x p This represents the acoustic characteristics of the audio signal after the p-th time synchronization.
[0080] Each acoustic feature is input into a pre-trained audio quality scoring model to obtain the predicted quality score (s1, s2, ..., s) of the audio signal acquired by each signal receiver. p ), where s p This represents the quality score of the audio signal acquired by the p-th signal receiver. Finally, the highest-scoring s is selected. i , will s i The audio signal collected by the corresponding i-th signal receiver is used as the final selection result.
[0081] The audio signal selection device provided in the embodiments of this application is described below. The audio signal selection device described below can be referred to in correspondence with the audio signal selection method described above.
[0082] See Figure 4 , Figure 4 This is a schematic diagram of an audio signal selection device disclosed in an embodiment of this application.
[0083] like Figure 4 As shown, the device may include:
[0084] The signal set acquisition unit 11 is used to acquire an audio signal set, wherein the audio signal set contains at least one candidate audio signal;
[0085] The acoustic feature extraction unit 12 is used to extract the acoustic features of each candidate audio signal. The acoustic features are the acoustic features that can be used to decode the target audio signal, and the target audio signal is close to or equivalent to the candidate audio signal.
[0086] The quality determination unit 13 is used to determine the quality of each candidate audio signal based on the acoustic characteristics of each candidate audio signal.
[0087] The quality selection unit 14 is used to select candidate audio signals that meet the set quality conditions by referring to the quality of each candidate audio signal, and use them as the audio signal selection result.
[0088] Optionally, the candidate audio signals included in the audio signal set are audio signals acquired from the same signal source by different signal receivers. Based on this, the apparatus of this application may further include: a time alignment processing unit, used to perform time alignment processing on each candidate audio signal in the audio signal set before the acoustic feature extraction unit processes it.
[0089] Optionally, the process by which the acoustic feature extraction unit extracts the acoustic features of each candidate audio signal may include:
[0090] A pre-trained acoustic encoder is used to extract the acoustic features of each candidate audio signal;
[0091] The acoustic encoder and decoder are trained together. During the training process, the acoustic encoder extracts acoustic features from the training audio signal, and the decoder decodes the restored audio signal based on the acoustic features. The network parameters of the acoustic encoder and decoder are updated with the goal of making the restored audio signal approximate the training audio signal.
[0092] Optionally, the acoustic feature extraction unit described above employs a pre-trained acoustic encoder. The process of extracting acoustic features for each candidate audio signal may include:
[0093] A pre-trained acoustic encoder is used to encode and model each candidate audio signal in different frequency bands to obtain latent variables as the acoustic features of the candidate audio signal. The latent variables contain the intrinsic information of the candidate audio signal in each dimension.
[0094] Optionally, the acoustic encoder may be a variational autoencoder (VAE-encoder).
[0095] Optionally, the process by which the quality determination unit determines the quality of each candidate audio signal based on the acoustic characteristics of each candidate audio signal may include:
[0096] Each candidate audio signal is input into a pre-trained audio quality scoring model to obtain a quality score for each candidate audio signal output by the model.
[0097] The audio quality scoring model is trained using the acoustic features of the training audio signal as training samples and the quality scores of the training audio signal as sample labels.
[0098] Optionally, the above audio quality scoring model can adopt a time-delay neural network structure.
[0099] Optionally, the process by which the quality selection unit selects candidate audio signals that meet the set quality conditions as the audio signal selection result, with reference to the quality of each candidate audio signal, may include:
[0100] Among all candidate audio signals, the candidate audio signal with the highest quality is selected as the audio signal selection result;
[0101] or,
[0102] Among the candidate audio signals, the candidate audio signals whose quality exceeds the set quality threshold are selected as the audio signal selection results.
[0103] The audio signal selection device provided in this application embodiment can be applied to audio signal selection equipment, such as the processing terminal in the aforementioned signal receiving system. This processing terminal can be a mobile phone, computer, server, etc. Optionally, Figure 5 The hardware block diagram of the audio signal selection device is shown below. Figure 5 The hardware structure of the audio signal selection device may include: at least one processor 1, at least one communication interface 2, at least one memory 3, and at least one communication bus 4;
[0104] In this embodiment of the application, the number of processor 1, communication interface 2, memory 3, and communication bus 4 is at least one, and processor 1, communication interface 2, and memory 3 communicate with each other through communication bus 4;
[0105] Processor 1 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present invention.
[0106] Memory 3 may include high-speed RAM, and may also include non-volatile memory, such as at least one disk storage device;
[0107] The memory stores a program, which the processor can call. The program is used for:
[0108] Obtain an audio signal set, wherein the audio signal set contains at least one candidate audio signal;
[0109] The acoustic features of each candidate audio signal are extracted, and the acoustic features are such that the target audio signal can be decoded based on them. The target audio signal is close to or equivalent to the candidate audio signal.
[0110] The quality of each candidate audio signal is determined based on its acoustic characteristics.
[0111] Based on the quality of each candidate audio signal, the candidate audio signal that meets the set quality conditions is selected as the audio signal selection result.
[0112] Optionally, the refined and extended functions of the program can be found in the description above.
[0113] This application embodiment also provides a storage medium that can store a program suitable for execution by a processor, the program being used for:
[0114] Obtain an audio signal set, wherein the audio signal set contains at least one candidate audio signal;
[0115] The acoustic features of each candidate audio signal are extracted, and the acoustic features are such that the target audio signal can be decoded based on them. The target audio signal is close to or equivalent to the candidate audio signal.
[0116] The quality of each candidate audio signal is determined based on its acoustic characteristics.
[0117] Based on the quality of each candidate audio signal, the candidate audio signal that meets the set quality conditions is selected as the audio signal selection result.
[0118] Optionally, the refined and extended functions of the program can be found in the description above.
[0119] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0120] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The various embodiments can be combined as needed, and the same or similar parts can be referred to each other.
[0121] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for selecting audio signals, characterized in that, include: Obtain an audio signal set, wherein the audio signal set contains at least one candidate audio signal; The acoustic features of each candidate audio signal are extracted. These acoustic features are latent variables obtained by encoding and modeling the candidate audio signal in different frequency bands. They contain intrinsic information of the candidate audio signal in various dimensions. The values of different dimensions in the latent variables correspond to different frequency bands, and the intrinsic information in different dimensions is represented by the values of those dimensions. Each dimension of intrinsic information contains influencing factors affecting the quality of the audio signal. The acoustic features are such that the acoustic features of the target audio signal can be decoded based on them. The target audio signal is close to or equivalent to the candidate audio signal. The quality of each candidate audio signal is determined based on its acoustic characteristics. Based on the quality of each candidate audio signal, the candidate audio signal that meets the set quality conditions is selected as the audio signal selection result.
2. The method according to claim 1, characterized in that, The candidate audio signals included in the audio signal set are audio signals collected from the same signal source by different signal receivers; Prior to extracting the acoustic features of each candidate audio signal, the method further includes: Time alignment processing is performed on each candidate audio signal in the audio signal set.
3. The method according to claim 1, characterized in that, The extraction of acoustic features for each candidate audio signal includes: A pre-trained acoustic encoder is used to extract the acoustic features of each candidate audio signal; The acoustic encoder and decoder are trained together. During the training process, the acoustic encoder extracts acoustic features from the training audio signal, and the decoder decodes the restored audio signal based on the acoustic features. The network parameters of the acoustic encoder and decoder are updated with the goal of making the restored audio signal approximate the training audio signal.
4. The method according to claim 3, characterized in that, A pre-trained acoustic encoder is used to extract acoustic features for each candidate audio signal, including: A pre-trained acoustic encoder is used to encode and model each candidate audio signal in different frequency bands to obtain latent variables as the acoustic features of the candidate audio signal. The latent variables contain the intrinsic information of the candidate audio signal in each dimension.
5. The method according to claim 3, characterized in that, The acoustic encoder is a variational autoencoder (vae-encoder).
6. The method according to claim 1, characterized in that, Determining the quality of each candidate audio signal based on its acoustic features includes: Each candidate audio signal is input into a pre-trained audio quality scoring model to obtain a quality score for each candidate audio signal output by the model. The audio quality scoring model is trained using the acoustic features of the training audio signal as training samples and the quality scores of the training audio signal as sample labels.
7. The method according to claim 6, characterized in that, The audio quality scoring model uses a time-delay neural network structure.
8. The method according to any one of claims 1-7, characterized in that, The step of selecting candidate audio signals that meet the set quality conditions based on the quality of each candidate audio signal, and using these as the audio signal selection result, includes: Among all candidate audio signals, the candidate audio signal with the highest quality is selected as the audio signal selection result; or, Among the candidate audio signals, the candidate audio signals whose quality exceeds the set quality threshold are selected as the audio signal selection results.
9. A signal receiving system, characterized in that, include: Several signal receivers and processing terminals; Each of the aforementioned signal receivers acquires audio signals from the same signal source and sends them to the processing terminal to form an audio signal set; The processing terminal uses the audio signal selection method of any one of claims 1-8 to select an audio signal that meets the set quality conditions from the audio signal set.
10. An audio signal selection device, characterized in that, include: A signal set acquisition unit is used to acquire an audio signal set, wherein the audio signal set contains at least one candidate audio signal; An acoustic feature extraction unit is used to extract acoustic features of each candidate audio signal. The acoustic features are latent variables obtained by encoding and modeling the candidate audio signal in different frequency bands. These latent features contain intrinsic information of the candidate audio signal in various dimensions. Values of different dimensions in the latent variables correspond to different frequency bands, and the intrinsic information in different dimensions is represented by the values of those dimensions. Each dimension of intrinsic information contains influencing factors affecting the quality of the audio signal. The acoustic features are such that the acoustic features of the target audio signal can be decoded based on them. The target audio signal is close to or equivalent to the candidate audio signal. A quality determination unit is configured to determine the quality of each candidate audio signal based on the acoustic characteristics of each candidate audio signal. The quality selection unit is used to select candidate audio signals that meet the set quality conditions by referring to the quality of each candidate audio signal, and use them as the audio signal selection result.
11. An audio signal selection device, characterized in that, include: Memory and processor; The memory is used to store programs; The processor is configured to execute the program to implement each step of the audio signal selection method as described in any one of claims 1 to 8.
12. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements each step of the audio signal selection method as described in any one of claims 1 to 8.