Double-talk state detection method and device, equipment and storage medium
By combining multi-dimensional feature extraction and dimensionality reduction with a dual-talk status classification model, the reliability problem of traditional dual-talk status detection in complex acoustic environments is solved, achieving fast and accurate dual-talk status detection and improving user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZHUHAI MOJIE TECH CO LTD
- Filing Date
- 2025-12-31
- Publication Date
- 2026-05-01
AI Technical Summary
Traditional dual-talk status detection technology is not reliable enough in complex acoustic environments, leading to frequent misjudgments and affecting echo cancellation function and user experience.
By employing multi-dimensional feature extraction and dimensionality reduction techniques, combined with a trained dual-talk state classification model, audio signals are classified to obtain dimensionality-reduced features that comprehensively reflect the characteristics of the audio signals, thereby reducing misjudgments and omissions.
It significantly improves the speed, robustness, and reliability of dual-talk status detection, quickly and accurately identifies dual-talk status, prevents the false elimination of valid voice signals, and enhances the user experience.
Smart Images

Figure CN121963769A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of communication technology, and in particular to a dual-talk status detection method, apparatus, device, and storage medium. Background Technology
[0002] Dual-talk status detection is a key technical component of echo cancellation. Under normal circumstances, echo cancellation continuously works to remove speaker echoes. However, when two or more people are speaking simultaneously, the echo cancellation function may mistakenly interpret one party's speech signal as an echo. Therefore, dual-talk status detection technology pauses echo cancellation when dual-talk occurs to avoid erroneously eliminating valid speech signals.
[0003] Traditional duotalk status detection techniques typically rely on a single feature for judgment, such as the energy or frequency analysis of the audio signal. However, these techniques are often unreliable in complex acoustic environments. For example, in situations with high background noise, changing echo paths, or the presence of multiple sound sources, a single feature may fail to accurately reflect the actual duotalk status, leading to misjudgments. Such misjudgments not only affect echo cancellation functionality but may also degrade the user experience. Summary of the Invention
[0004] This invention provides a dual-talk status detection method, apparatus, device, and storage medium, aiming to solve the technical problem that traditional dual-talk status detection technology is not reliable enough in complex acoustic environments.
[0005] In a first aspect, embodiments of the present invention provide a dual-talk state detection method, comprising: Acquire the target audio signal; The target audio signal is subjected to multi-dimensional feature extraction processing to obtain the multi-dimensional audio features of the target audio signal; The multi-dimensional audio features are subjected to dimensionality reduction processing to obtain the dimensionality-reduced audio features of the target audio signal; By using a trained dual-talk state classification model, the reduced-dimensional audio features are subjected to dual-talk state classification processing to obtain the result of whether the target audio signal is in a dual-talk state.
[0006] Secondly, embodiments of the present invention also provide a dual-talk status detection device, comprising: The signal acquisition module is used to acquire the target audio signal; The feature extraction module is used to perform multi-dimensional feature extraction processing on the target audio signal to obtain multi-dimensional audio features of the target audio signal; The feature dimensionality reduction module is used to perform dimensionality reduction processing on the multi-dimensional audio features to obtain the dimensionality-reduced audio features of the target audio signal; The state classification module is used to perform dual-talk state classification processing on the dimensionality-reduced audio features using a trained dual-talk state classification model to obtain the result of whether the target audio signal is in a dual-talk state.
[0007] Thirdly, embodiments of the present invention also provide a computer device, the computer device including a processor, a memory, a computer program stored in the memory and executable by the processor, and a data bus for implementing connection communication between the processor and the memory, wherein when the computer program is executed by the processor, it implements the dual-talk state detection method as described in the first aspect.
[0008] Fourthly, embodiments of the present invention also provide a storage medium for computer-readable storage, the storage medium storing one or more programs, the one or more programs being executable by one or more processors to implement the dual-talk state detection method as described in the first aspect.
[0009] This invention provides a method, apparatus, device, and storage medium for dual-talk state detection. By extracting multi-dimensional features from the target audio signal, this invention obtains multi-dimensional audio features that comprehensively reflect the important characteristics of the target audio signal, effectively avoiding the failure of a single feature and reducing the possibility of misjudgment and missed judgment. By performing dimensionality reduction processing on the multi-dimensional audio features, irrelevant feature dimensions are effectively removed, while only the core discriminative information in the multi-dimensional audio features is retained. This allows for faster, more robust, and more accurate dual-talk state classification processing of the dimensionality-reduced audio features using a trained dual-talk state classification model, resulting in whether the target audio signal is in a dual-talk state. This significantly improves the speed, robustness, and reliability of dual-talk state detection. Furthermore, when the target audio signal is determined to be in a dual-talk state, the echo cancellation function is quickly paused to prevent the false cancellation of valid speech signals in the target audio signal, thereby improving the user experience. Attached Figure Description
[0010] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 This is a flowchart illustrating a dual-talk state detection method provided in an embodiment of the present invention; Figure 2 This is a flowchart illustrating another dual-talk state detection method provided in an embodiment of the present invention; Figure 3 yes Figure 1A schematic diagram of a specific implementation of step S104; Figure 4 This is a schematic block diagram of a dual-talk status detection device provided in an embodiment of the present invention; Figure 5 This is a schematic block diagram of the structure of a computer device provided in an embodiment of the present invention. Detailed Implementation
[0012] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0013] The flowchart shown in the attached diagram is for illustrative purposes only and does not necessarily include all content and operations / steps, nor does it necessarily have to be performed in the order described. For example, some operations / steps can be broken down, combined, or partially merged, so the actual execution order may change depending on the actual situation.
[0014] It should be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0015] This invention provides a dual-talk state detection method. By extracting multi-dimensional features from the target audio signal, multi-dimensional audio features that comprehensively reflect the important characteristics of the target audio signal can be obtained, effectively avoiding the failure of a single feature and reducing the possibility of misjudgment and missed judgment. By performing dimensionality reduction processing on the multi-dimensional audio features, dimensionality-reduced audio features are obtained, effectively removing irrelevant feature dimensions while retaining only the core discriminative information in the multi-dimensional audio features. This allows for faster, more robust, and more accurate dual-talk state classification processing of the dimensionality-reduced audio features using a trained dual-talk state classification model, resulting in the determination of whether the target audio signal is in a dual-talk state. This significantly improves the speed, robustness, and reliability of dual-talk state detection. Consequently, when the target audio signal is determined to be in a dual-talk state, the echo cancellation function is quickly paused to prevent the false cancellation of valid speech signals in the target audio signal, thereby improving the user experience.
[0016] The dual-talk status detection method provided in this application can be applied to a terminal, a server, or software running on either a terminal or a server. In some embodiments, the terminal can be a vehicle-mounted terminal, a smartphone, a tablet computer, a laptop computer, a desktop computer, etc.; the server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application implementing the dual-talk status detection method, etc., but is not limited to the above forms.
[0017] Please see Figure 1 , Figure 1 This is a flowchart illustrating a dual-talk state detection method provided in an embodiment of the present invention.
[0018] like Figure 1 As shown, the dual-talk state detection method includes steps S101 to S104.
[0019] S101: Acquire the target audio signal.
[0020] The dual-talk state detection method provided in this invention shifts from traditional single-feature judgment to multi-dimensional feature fusion, and introduces dimensionality reduction and artificial intelligence technologies to achieve effective dual-talk state detection, significantly improving detection accuracy and robustness in complex acoustic environments while maintaining low computational complexity.
[0021] The dual-talk state detection method provided in this embodiment of the invention mainly includes two stages: The first stage involves constructing training samples for training a pre-defined dual-talk state classification model, and then training the dual-talk state classification model based on the training samples to obtain a well-trained dual-talk state classification model with enhanced dual-talk detection performance.
[0022] The second stage involves extracting multi-dimensional audio features from the target audio signal to be detected, and then performing dimensionality reduction processing on the extracted multi-dimensional audio features to obtain dimensionality-reduced audio features that are easy for a well-trained dual-talk state classification model to understand and classify. Thus, the well-trained dual-talk state classification model can efficiently and accurately decide whether the target audio signal is in a dual-talk state based on the dimensionality-reduced audio features.
[0023] The first stage will be explained in detail below.
[0024] Please see Figure 2 Before step S101 in some embodiments, the following steps may be included, but are not limited to: S105: Construct training samples for training the pre-defined dual-talk state classification model; S106: Train the dual-talk state classification model based on the training samples to obtain the trained dual-talk state classification model.
[0025] For example, the preset two-lecture state classification model can be a Support Vector Machine (SVM) based on a Radial Basis Function (RBF) kernel, or it can be a binary classification model based on other function kernels (such as linear kernels, multinomial kernels), such as SVM, Logistic Regression, Decision Tree, Random Forest, or Gradient Boosting Trees. This embodiment of the invention does not limit this. For ease of understanding, this embodiment of the invention uses a lighter-weight SVM based on an RBF kernel as an example for illustration.
[0026] To enable the RBF kernel-based SVM to learn more robust and accurate bilingual state classification decisions, it can be trained using supervised training. Therefore, training samples for training the RBF kernel-based SVM are constructed, which include positive and negative samples.
[0027] In step S105 of some embodiments, a first dimensionality-reduced audio feature of the first audio signal in a two-talk scenario can be obtained; a two-talk state label can be attached to the first dimensionality-reduced audio feature; the first dimensionality-reduced audio feature and the two-talk state label can be used as positive samples; a second dimensionality-reduced audio feature of the second audio signal in a one-talk scenario can be obtained; a one-talk state label can be attached to the second dimensionality-reduced audio feature; the second dimensionality-reduced audio feature and the one-talk state label can be used as negative samples; the positive samples and negative samples constitute training samples.
[0028] The process of constructing positive samples includes: First, the first audio signal of the two-way conversation scenario is collected. The first audio signal is the microphone audio signal when the near-end speaker and the far-end speaker are active at the same time. That is, the first audio signal includes the near-end audio signal and the far-end audio signal. The far-end audio signal is the far-end echo from the speaker (the far-end echo of the speaker refers to the phenomenon that the sound of the far-end speaker is picked up again by the microphone after being played through the speaker).
[0029] Then, multi-dimensional feature extraction processing is performed on the first audio signal from the time domain, frequency domain, cepstral domain and cross-correlation domain to obtain the first multi-dimensional audio features of the first audio signal.
[0030] The first multi-dimensional audio features include the first time-domain features, the first frequency-domain features, the first cepstral domain features, and the first cross-correlation domain features: (a) The first time-domain feature includes the first zero-crossing rate and the first short-time energy.
[0031] Zero crossing rate (ZCR) reflects the frequency at which an audio signal crosses zero in the time domain; it reflects the degree of change in the audio signal. The formula for calculating zero crossing rate is as follows:
[0032] in, N Indicates the frame length (number of sampling points). Indicates the first n The signal amplitude at each sampling point.
[0033] Therefore, the first zero-crossing rate can be calculated using the formula for zero-crossing rate.
[0034] Short-Time Energy (STE) is used to characterize the energy intensity of an audio signal within a single frame; high energy typically indicates rapid changes in the audio signal. The formula for calculating short-time energy is as follows:
[0035] Therefore, the first short-time energy can be calculated using the formula for short-time energy.
[0036] (ii) The first frequency domain characteristics include the first energy entropy, the first spectral centroid, the first spectral roll-off point, and the first spectral flux.
[0037] Energy entropy (EE) reflects the uniformity of energy distribution in an audio signal's frequency band; the more uniform the distribution, the higher the entropy value. The formula for calculating energy entropy is as follows:
[0038] in, Y This indicates the number of sub-bands. The number of sub-bands applicable to the audio signal can be flexibly selected based on the sampling rate of the audio signal, such as between 8 and 32. L This indicates the length of each sub-band. L = N / Y ; Indicates the first y Energy of each frequency band.
[0039] Therefore, the first energy entropy can be calculated using the formula for calculating energy entropy.
[0040] The spectral centroid (SC) reflects the location of the center of gravity in the audio signal's spectrum and is commonly used to measure the perceived brightness of an audio signal. The formula for calculating the spectral centroid is as follows:
[0041] in, X[k] Represents the Discrete Fourier Transform (DFT) coefficients of an audio signal frame; k Indicates frequency index, k The corresponding frequency is , fs Indicates the sampling frequency.
[0042] Therefore, the first spectral centroid can be calculated using the formula for calculating the spectral centroid.
[0043] The spectral rolloff (SR) is the frequency at which the accumulated spectral energy reaches a specific proportion of the total energy.
[0044] in, This indicates the energy ratio threshold.
[0045] Therefore, the first spectral roll-off point can be calculated using the formula for calculating the spectral roll-off point.
[0046] Spectral flux (SF) measures the degree of spectral change between adjacent frames. The formula for calculating spectral flux is as follows:
[0047] in, Represents the Discrete Fourier Transform (DFT) coefficients of the current frame; This represents the DFT coefficients of the previous frame.
[0048] Therefore, the first spectral flux can be calculated using the formula for calculating spectral flux.
[0049] (iii) The first cepstral domain features include the first Mel frequency cepstral coefficients.
[0050] Mel-Frequency Cepstral Coefficients (MFCCs) consist of 12 coefficients and one energy term. The MFCCs can be obtained as follows: The audio signal is pre-emphasized to obtain a pre-emphasized audio signal. This pre-emphasized audio signal is then framed using a windowing method. Each frame is then subjected to a Fourier Transform (DFT) to obtain its spectrum. The spectrum is then weighted using a Mel filter bank to obtain a weighted spectrum. The logarithm of the weighted spectrum is taken to obtain the logarithmic energy. Finally, the logarithmic energy is subjected to a Discrete Cosine Transform (DCT) to obtain the MFCCs.
[0051] The formula for pre-emphasis processing is as follows: , This represents the pre-emphasis coefficient.
[0052] The formula for adding a window is: , Indicates Hanming window, .
[0053] The formula for DFT transformation is: .
[0054] Mel filter bank: , Indicates the first m The frequency response of a Mel filter (the frequency domain of the filter). Indicates the first m The center frequency of the Mel filter, Indicates the first m +1 center frequency of Mel filter.
[0055] The formula for calculating logarithmic energy is: , Indicates the first m Logarithmic energy corresponding to a Mel filter X[k] Indicates the first k Complex spectral values in the frequency domain.
[0056] The formula for DCT transformation is: , M Indicates the number of Mel filters. C[i] Indicates the first i MFCC coefficients.
[0057] Therefore, the first Mel frequency cepstral coefficient can be calculated using the above method.
[0058] (iv) The first cross-correlation domain features include the first normalized cross-correlation function value, the first energy ratio, and the first high-spectral divergence.
[0059] The normalized cross-correlation (NCC) function measures the similarity between the near-end and far-end audio signals. In two-way speech, the near-end mainly contains the far-end echo and is highly correlated with the far-end, resulting in a high NCC peak. When the near-end exclusively controls the speech, the correlation is low, and the NCC peak is low. Therefore, NCC is a core feature for two-way speech discrimination. The formula for calculating the normalized cross-correlation function is shown below:
[0060] in, y[n] This refers to near-end audio signals, which are audio signals directly collected from the vicinity of a sound source (such as a speaker at the near end) using a microphone; x[n] This refers to a distant audio signal, which is an audio signal captured by a microphone at a distance from the sound source, such as the audio signal of a distant speaker that is played through a speaker and then captured by the microphone again. This represents the time delay parameter.
[0061] Therefore, the first normalized cross-correlation function value can be calculated using the formula for calculating the normalized cross-correlation function value.
[0062] Energy ratio (ER) measures the intensity ratio of the near-end audio signal to the far-end audio signal and is a core feature for two-talk status determination. The formula for calculating energy ratio is shown below:
[0063] This represents a very small constant, used to prevent division by zero.
[0064] Therefore, the first energy ratio can be calculated using the formula for calculating the energy ratio.
[0065] High Spectral Divergence (HSD) measures the difference in the high-frequency components of a signal. The formula for calculating HSD is shown below:
[0066] These represent the lower and upper frequency indices of the high-frequency region, respectively. These represent the spectra of the near-end audio signal and the far-end audio signal, respectively.
[0067] Therefore, the first high-frequency divergence can be calculated using the formula for high-frequency divergence.
[0068] The first multi-dimensional audio feature includes 10 dimensions, covering two levels: the statistics of the first audio signal itself (time / frequency / cepstrum) and the near-far relationship (correlation / energy / high-frequency difference). They are highly complementary and can comprehensively characterize the near-far audio signal, the far-far audio signal, and the relationship between the two in the first audio signal.
[0069] Considering that the first multi-dimensional audio feature is a high-dimensional feature and has strong correlation, in order to reduce the training complexity of the RBF kernel-based SVM and avoid overfitting, the first multi-dimensional audio feature is dimensionality reduced to obtain the first dimensionality-reduced audio feature.
[0070] Considering that the 10 dimensions of the first multi-dimensional audio features may have different dimensions and numerical ranges, in order to avoid some features dominating the calculation, the first multi-dimensional audio features can be standardized to obtain standardized first multi-dimensional audio features. For example, standardization can be performed using the Z-score (zero mean, unit variance) method.
[0071] in, Indicates the first i The first audio signal j The value of each feature; Indicates the first j The mean of each feature, ; Indicates the first j The standard deviation of each feature , (Indicates sample index) B Indicates the total number of the first audio signals. (Indicates feature dimension index).
[0072] Next, the standardized first multi-dimensional audio features can be subjected to dimensionality reduction calculations. For example, Principal Component Analysis (PCA) can be used to perform dimensionality reduction calculations on the standardized first multi-dimensional audio features to obtain the first dimensionality-reduced audio features.
[0073] PCA is a statistical method used to transform high-dimensional data into low-dimensional data while preserving as much variation information as possible. PCA mainly uses orthogonal transformations to convert potentially correlated features into linearly uncorrelated principal components, thus achieving feature dimensionality reduction. The computational steps include covariance matrix calculation, eigenvalue decomposition, principal component selection, and PCA projection matrix construction.
[0074] The formula for calculating the covariance matrix is shown below:
[0075] in, Z This represents the normalized characteristic matrix of D×10. This represents a 10×10 covariance matrix.
[0076] The formula for eigenvalue decomposition is shown below:
[0077] in, V Represents the eigenvector matrix, V = [v 1 , v 2 , ..., v 10 ] , v i It is a 10-dimensional column vector. v i Represents the covariance matrix One of the characteristic directions, v i Each element in the normalized feature matrix Z represents the corresponding original feature in the normalized feature matrix Z. v i The importance of the feature direction it represents, but a single element does not represent an independent feature vector; Represents the eigenvalue matrix. This represents the characteristic values arranged in descending order.
[0078] The formula for principal component selection is shown below:
[0079] in, r This indicates the number of principal components selected; typically, the top ones with a cumulative variance contribution rate ≥ 85% are chosen. r Principal components, Indicates the first i Large eigenvalues.
[0080] Final choice r The principal components are used to construct a projection matrix, which is used to project the original feature vectors onto the dimensionality-reduced space. The formula for constructing the PCA projection matrix is as follows:
[0081] in, Indicates 10× r The projection matrix; v i Indicates the corresponding number i Eigenvectors with large eigenvalues; r This represents the dimension after dimensionality reduction, typically ranging from 2 to 3.
[0082] Therefore, the covariance matrix of the standardized first audio feature can be calculated to obtain the first covariance matrix of the standardized first audio feature; then, the first covariance matrix can be decomposed into eigenvalues to obtain the first eigenvalue matrix and its corresponding first eigenvector matrix; finally, based on the magnitude of the eigenvalues in the first eigenvalue matrix, the largest eigenvector matrix is selected. r The eigenvectors corresponding to each eigenvalue are used as principal components. The larger the eigenvalue, the higher the contribution rate of the corresponding principal component. Therefore, the standardized first audio feature is projected onto the selected principal component to obtain the first dimensionality-reduced audio feature.
[0083] This effectively reduces the dimensionality of the standardized first audio features while retaining as much original information as possible from them, facilitating subsequent model training.
[0084] Finally, the first dimensionality-reduced audio feature is labeled with a two-talk state label (e.g., labeled "+1"), and the first dimensionality-reduced audio feature and the two-talk state label are used as positive samples.
[0085] Similarly, the second audio signal of the single-speaker scenario is collected. The second audio signal is the microphone audio signal when the near-end speaker is present (the far-end speaker is not speaking). That is, the second audio signal only contains the near-end audio signal and does not contain any far-end echo from the speaker or the far-end speaker's voice.
[0086] Then, the second multi-dimensional audio features of the second audio signal are extracted. The second multi-dimensional audio features include second time-domain features, second frequency-domain features, second cepstral domain features, and second cross-correlation domain features. The second time-domain features include the second zero-crossing rate and the first short-time energy. The second frequency-domain features include the first energy entropy, the second spectral centroid, the second spectral roll-off point, and the second spectral flux. The second cepstral domain features include the second Mel-frequency cepstral coefficients. The second cross-correlation domain features include the second normalized cross-correlation function value, the second energy ratio, and the second high-spectral divergence. The extraction process is the same as the extraction process of the first multi-dimensional audio features described above, and will not be repeated here.
[0087] The second multi-dimensional audio features are subjected to dimensionality reduction processing to obtain the second dimensionality-reduced audio features of the second audio signal. The dimensionality reduction process is the same as that of the first multi-dimensional audio features, and will not be repeated here.
[0088] Finally, the second dimensionality-reduced audio feature is labeled with a single-lecture state label (e.g., labeled "-1"), and the first dimensionality-reduced audio feature and the single-lecture state label are used as negative samples.
[0089] Thus, positive and negative samples constitute the training samples for training the RBF kernel-based SVM.
[0090] Furthermore, the RBF kernel-based SVM is trained using the training samples to obtain a trained dual-state classification model.
[0091] In step S106 of some embodiments, positive and negative samples can be input into the dual-talk state classification model for training to calculate the decision function and obtain the trained dual-talk state classification model.
[0092] Among them, the RBF kernel-based SVM is an algorithm that achieves classification by constructing an optimal hyperplane, and its decision boundary depends on the support vectors.
[0093] Specifically, positive and negative samples can be input into an SVM based on an RBF kernel for training, and an optimization problem for the SVM can be constructed using the positive and negative samples. The goal of this optimization problem is to maximize the margin between the two classes, the two-talk state and the one-talk state, while correctly classifying the first and second dimensionality-reduced audio features. By solving the optimization problem, the Lagrange multipliers are calculated. Lagrange multipliers are used to evaluate the importance of each sample in the RBF kernel-based SVM. During training, non-zero values are multiplied by the lagrange multipliers. The corresponding samples serve as support vectors, which are the key data points constituting the decision boundary. The number of support vectors is counted. Based on the support vectors, their corresponding state labels and counts, and the Lagrange multipliers, the decision function is calculated. This decision function is a core component of the RBF kernel-based SVM, used to determine the category of the current input feature. The parameters of the RBF kernel function within the decision function are then solved. γ and bias terms b The trained dual-state classification model is obtained, and the formula for the decision function is shown below:
[0094] in, Indicates the number of support vectors; Represents the Lagrange multipliers; The status label indicates that +1 represents dual-talk mode and -1 represents single-talk mode. The RBF kernel function is defined as follows: ; This represents the features of the current input; Represents support vectors; The parameters represent the RBF kernel function; This represents the bias term, indicating the distance shifted in the final decision result.
[0095] A well-trained dual-talk state classification model can effectively classify between dual-talk and single-talk states.
[0096] Furthermore, a well-trained dual-talk state classification model can be updated and retrained by continuously inputting new training samples, thereby enhancing its adaptability and enabling it to maintain good performance in constantly changing audio environments.
[0097] The following is a detailed explanation of Phase Two.
[0098] For step S101, acquiring the target audio signal, which refers to the microphone audio signal to be detected, may involve the following situations: Scenario 1: Only near-end audio signals (the sound emitted by the near-end speaker) are included, and there are no far-end audio signals, indicating that there is no far-end echo from the speaker (the far-end echo of the speaker refers to the phenomenon that the sound of the far-end speaker is picked up again by the microphone after being played through the speaker). The second scenario is: only the far-end audio signal is included, meaning that the microphone receives the sound played by the speaker, but not the sound from the speaker at the near end; Case 3 is: It contains both near-end audio signals and far-end audio signals.
[0099] The core of dual-talk state detection of target audio signal is to determine whether the target audio signal has both near-end audio signal and far-end audio signal at the same time (if they exist at the same time, it is dual-talk state; otherwise, it is single-talk state).
[0100] S102: Perform multi-dimensional feature extraction processing on the target audio signal to obtain the multi-dimensional audio features of the target audio signal.
[0101] For step S102, multi-dimensional feature extraction processing is performed on the target audio signal from the time domain, frequency domain, cepstral domain and cross-correlation domain to obtain multi-dimensional audio features of the target audio signal. For the specific extraction process, please refer to the extraction process of the first multi-dimensional audio feature mentioned above, which will not be repeated here.
[0102] Multi-dimensional audio features can comprehensively reflect the important characteristics of the target audio signal, covering key aspects from basic volume to complex acoustics. By comprehensively extracting multi-dimensional audio features covering multiple domains and angles, the failure of a single feature can be effectively avoided, thereby enabling more accurate and robust dual-talk status detection of the target audio signal.
[0103] In some embodiments, the multi-dimensional audio features include at least two of the following: target time-domain features, target frequency-domain features, target cepstral domain features, and target cross-correlation domain features. The target time-domain features include the target zero-crossing rate and target short-time energy; the target frequency-domain features include the target energy entropy, target spectral centroid, target spectral roll-off point, and target spectral flux; the target cepstral domain features include the target Mel-frequency cepstral coefficients; and the target cross-correlation domain features include the target normalized cross-correlation function value, target energy ratio, and target high-spectral divergence. The multi-dimensional audio features can be expressed as follows:
[0104] in, Indicates the target zero-crossing rate; Indicates the short-term energy of the target; Represents the target energy entropy; Indicates the centroid of the target spectrum; Indicates the target spectral roll-off point; Indicates the target spectral flux; Represents the cepstral coefficients of the target Mel frequency; This represents the target normalized cross-correlation function value; Indicates the target energy ratio; The target is high spectral divergence.
[0105] S103: Perform dimensionality reduction processing on the multi-dimensional audio features to obtain the dimensionality-reduced audio features of the target audio signal.
[0106] For step S103, the multi-dimensional audio features are reduced in dimensionality, for example, to 2-3 dimensions, to obtain the dimensionality-reduced audio features of the target audio signal.
[0107] Dimensionality reduction places the audio features in a low-dimensional space, making them easier to linearly separate and reducing the computational complexity of subsequent classification. At the same time, it retains the main discriminative information in the original multi-dimensional audio features, improving the accuracy of subsequent classification.
[0108] In step S103 of some embodiments, the multi-dimensional audio features can be standardized to obtain standardized multi-dimensional audio features; principal component analysis is then used to perform dimensionality reduction calculation on the standardized multi-dimensional audio features to obtain dimensionality-reduced audio features.
[0109] The multi-dimensional audio features are standardized to obtain standardized multi-dimensional audio features. The standardization process is the same as that described above for the standardization of the first multi-dimensional audio features, and will not be repeated here. Principal component analysis is then used to perform dimensionality reduction calculations on the standardized multi-dimensional audio features to obtain dimensionality-reduced audio features. The dimensionality reduction calculation process is the same as that described above for the dimensionality reduction calculations of the first standardized multi-dimensional audio features, and will not be repeated here.
[0110] S104: Using a trained dual-talk state classification model, perform dual-talk state classification on the dimensionality-reduced audio features to obtain the result of whether the target audio signal is in a dual-talk state.
[0111] Finally, by using the trained dual-talk state classification model, the reduced-dimensional audio features are processed for dual-talk state classification to obtain the result of whether the target audio signal is in a dual-talk state, thereby improving the accuracy of dual-talk state classification and the interpretability of the classification results.
[0112] In some embodiments, please refer to Figure 3 Step S104 may include the following steps: S1041: Using the decision function of the trained dual-talk state classification model, perform binary classification decision calculation on the dimensionality-reduced audio features to obtain the decision value of the dimensionality-reduced audio features; S1042: If the decision value is greater than or equal to zero, determine that the target audio signal is in dual-talk mode; S1043: If the decision value is less than zero, determine that the target audio signal is in a single-talk state.
[0113] Specifically, the dimensionality-reduced audio features are used as the features of the current input. The input is fed into a trained dual-talk state classification model. The trained dual-talk state classification model substitutes the dimensionality-reduced audio features into its decision function to perform binary classification decision calculation, obtaining the decision value of the target audio signal, denoted as... Based on decision value The decision logic for determining whether the target audio signal is in dual-talk mode is as follows:
[0114] In other words, if the decision value is greater than or equal to zero, the target audio signal is determined to be in a dual-talk state; if the decision value is less than zero, the target audio signal is determined to be in a single-talk state.
[0115] The dual-talk state detection method provided in this embodiment of the invention can bring the following beneficial effects: On the one hand, it improves detection accuracy and robustness: by extracting multi-dimensional features from the target audio signal, it can comprehensively capture various characteristics of the target audio signal, effectively avoid the failure of a single feature, reduce the possibility of misjudgment and missed judgment, and thus improve the accuracy and robustness of dual-talk status detection.
[0116] On the other hand, it reduces computational complexity and resource consumption: dimensionality reduction can effectively remove irrelevant or highly correlated feature dimensions, reducing the computational burden. This not only speeds up subsequent classification but also reduces memory usage, making it suitable for applications in resource-constrained environments.
[0117] On the other hand, it enhances robustness: dimensionality reduction helps alleviate the "curse of dimensionality" problem, avoids overfitting, and retains the main discriminative information, making the trained dual-state classification model more robust under different acoustic environments, background noise or equipment conditions, and improving its anti-interference ability.
[0118] On the other hand, it improves real-time processing capabilities: By combining multi-dimensional feature extraction, multi-dimensional feature dimensionality reduction and artificial intelligence classification, the embodiments of the present invention make the dual-talk status detection process efficient and real-time, which can be widely used in real-time scenarios such as video conferencing and telephone communication, thereby improving the user experience.
[0119] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention. Where there is no conflict, the above embodiments and features can be combined with each other.
[0120] Please see Figure 4 , Figure 4 This is a schematic block diagram of the structure of a dual-talk status detection device provided in an embodiment of the present invention.
[0121] like Figure 4 As shown, the dual-talk status detection device 100 includes: Signal acquisition module 110 is used to acquire target audio signals; The feature extraction module 120 is used to perform multi-dimensional feature extraction processing on the target audio signal to obtain multi-dimensional audio features of the target audio signal; The feature dimensionality reduction module 130 is used to perform dimensionality reduction processing on the multi-dimensional audio features to obtain the dimensionality-reduced audio features of the target audio signal; The state classification module 140 is used to perform dual-talk state classification processing on the dimensionality-reduced audio features using a trained dual-talk state classification model to obtain the result of whether the target audio signal is in a dual-talk state.
[0122] In some embodiments, the multi-dimensional audio features include at least two of the following: target time-domain features, target frequency-domain features, target cepstral domain features, and target cross-correlation domain features. The target time-domain features include target zero-crossing rate and target short-time energy. The target frequency-domain features include target energy entropy, target spectral centroid, target spectral roll-off point, and target spectral flux. The target cepstral domain features include target Mel-frequency cepstral coefficients. The target cross-correlation domain features include target normalized cross-correlation function value, target energy ratio, and target high-spectral divergence.
[0123] In some embodiments, the feature dimensionality reduction module 130 is specifically used for: The multi-dimensional audio features are standardized to obtain standardized multi-dimensional audio features. Principal component analysis is used to perform dimensionality reduction calculations on the standardized multi-dimensional audio features to obtain the dimensionality-reduced audio features.
[0124] In some embodiments, the state classification module 140 is specifically used for: The decision value of the reduced-dimensional audio feature is obtained by performing binary classification decision calculation on the reduced-dimensional audio feature using the decision function of the trained dual-talk state classification model. If the decision value is greater than or equal to zero, the target audio signal is determined to be in dual-talk mode; If the decision value is less than zero, the target audio signal is determined to be in a single-talk state.
[0125] In some embodiments, the dual-talk state detection device further includes a model training module, which is used for: Construct training samples for training a pre-defined dual-state classification model; The dual-talk state classification model is trained based on the training samples to obtain the trained dual-talk state classification model.
[0126] In some embodiments, the model training module is specifically used for: Obtain the first dimension-reduced audio features of the first audio signal in a dual-talk scenario; Label the first dimensionality-reduced audio features with dual-talk status tags; The first dimensionality-reduced audio feature and the dual-talk state label are used as positive samples; Obtain the second dimension-reduced audio features of the second audio signal in a single-lecture scenario; Label the second dimensionality-reduced audio features with single-lecture status tags; The second dimensionality-reduced audio feature and the single-lecture state label are used as negative samples; The positive samples and the negative samples constitute the training samples.
[0127] In some embodiments, the model training module is further configured to: The positive and negative samples are input into the dual-talk state classification model for training to calculate the decision function and obtain the trained dual-talk state classification model.
[0128] It should be noted that those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the dual-talk status detection device described above can be referred to the corresponding process in the aforementioned dual-talk status detection method embodiments, and will not be repeated here.
[0129] Please see Figure 5 , Figure 5 This is a schematic block diagram of the structure of a computer device provided in an embodiment of the present invention.
[0130] like Figure 5 As shown, the computer device 200 includes a processor 201 and a memory 202, which are connected via a bus 203, such as an I2C (Inter-integrated Circuit) bus.
[0131] Specifically, processor 301 provides computing and control capabilities to support the operation of the entire computer device. Processor 201 can be a Central Processing Unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.
[0132] Specifically, the memory 202 can be a Flash chip, a read-only memory (ROM) disk, an optical disk, a USB flash drive, or a portable hard drive, etc.
[0133] Those skilled in the art will understand that Figure 5The structures shown are merely block diagrams of some structures related to the embodiments of the present invention, and do not constitute a limitation on the computer devices on which the embodiments of the present invention are applied. Specific computer devices may include more or fewer components than those shown in the figures, or combine certain components, or have different component arrangements.
[0134] The processor 201 is used to run a computer program stored in the memory 202, and implements any of the dual-talk state detection methods provided in the embodiments of the present invention when executing the computer program.
[0135] In one embodiment, the processor 201 is configured to run a computer program stored in a memory, and to perform the following steps when executing the computer program: Acquire the target audio signal; The target audio signal is subjected to multi-dimensional feature extraction processing to obtain the multi-dimensional audio features of the target audio signal; The multi-dimensional audio features are subjected to dimensionality reduction processing to obtain the dimensionality-reduced audio features of the target audio signal; By using a trained dual-talk state classification model, the reduced-dimensional audio features are subjected to dual-talk state classification processing to obtain the result of whether the target audio signal is in a dual-talk state.
[0136] In some embodiments, the multi-dimensional audio features include at least two of the following: target time-domain features, target frequency-domain features, target cepstral domain features, and target cross-correlation domain features. The target time-domain features include target zero-crossing rate and target short-time energy. The target frequency-domain features include target energy entropy, target spectral centroid, target spectral roll-off point, and target spectral flux. The target cepstral domain features include target Mel-frequency cepstral coefficients. The target cross-correlation domain features include target normalized cross-correlation function value, target energy ratio, and target high-spectral divergence.
[0137] In some embodiments, when the processor 201 performs dimensionality reduction processing on the multi-dimensional audio features to obtain the dimensionality-reduced audio features of the target audio signal, it is configured to: The multi-dimensional audio features are standardized to obtain standardized multi-dimensional audio features. Principal component analysis is used to perform dimensionality reduction calculations on the standardized multi-dimensional audio features to obtain the dimensionality-reduced audio features.
[0138] In some embodiments, when the processor 201 performs dual-talk state classification processing on the dimensionality-reduced audio features using a trained dual-talk state classification model to obtain a result indicating whether the target audio signal is in a dual-talk state, it is configured to: The decision value of the reduced-dimensional audio feature is obtained by performing binary classification decision calculation on the reduced-dimensional audio feature using the decision function of the trained dual-talk state classification model. If the decision value is greater than or equal to zero, the target audio signal is determined to be in dual-talk mode; If the decision value is less than zero, the target audio signal is determined to be in a single-talk state.
[0139] In some embodiments, before acquiring the target audio signal, the processor 201 is further configured to: Construct training samples for training a pre-defined dual-state classification model; The dual-talk state classification model is trained based on the training samples to obtain the trained dual-talk state classification model.
[0140] In some embodiments, when the processor 201 constructs training samples for training a preset dual-talk state classification model, it is configured to: Obtain the first dimension-reduced audio features of the first audio signal in a dual-talk scenario; Label the first dimensionality-reduced audio features with dual-talk status tags; The first dimensionality-reduced audio feature and the dual-talk state label are used as positive samples; Obtain the second dimension-reduced audio features of the second audio signal in a single-lecture scenario; Label the second dimensionality-reduced audio features with single-lecture status tags; The second dimensionality-reduced audio feature and the single-lecture state label are used as negative samples; The positive samples and the negative samples constitute the training samples.
[0141] In some embodiments, when the processor 201 trains the dual-talk state classification model based on the training samples to obtain the trained dual-talk state classification model, it is configured to: The positive and negative samples are input into the dual-talk state classification model for training to calculate the decision function and obtain the trained dual-talk state classification model.
[0142] It should be noted that those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the computer device described above can be referred to the corresponding process in the aforementioned dual-talk state detection method embodiment, and will not be repeated here.
[0143] This invention also provides a storage medium for computer-readable storage, wherein the storage medium stores one or more programs that can be executed by one or more processors to implement any of the dual-talk status detection methods provided in the specification of this invention.
[0144] The storage medium can be volatile or non-volatile. It can be an internal storage unit of the computer device described in the foregoing embodiments, such as the hard drive or memory of the computer device. Alternatively, it can be an external storage device of the computer device, such as a plug-in hard drive, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card.
[0145] Those skilled in the art will understand that all or some of the steps, systems, or apparatuses disclosed above, and their functional modules / units, can be implemented as software, firmware, hardware, or suitable combinations thereof. In hardware embodiments, the division between functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include computer storage media (or non-transitory media) and communication media (or transient media). As is known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by a computer. Furthermore, it is well known to those skilled in the art that communication media typically contain computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.
[0146] It should be understood that the term "and / or" as used in this specification and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes such combinations. It should be noted that, herein, the terms "comprising," "including," or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.
[0147] The sequence numbers of the above embodiments of the present invention are merely for descriptive purposes and do not represent the superiority or inferiority of the embodiments. The above descriptions are only specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A dual-talk status detection method, characterized in that, include: Acquire the target audio signal; The target audio signal is subjected to multi-dimensional feature extraction processing to obtain the multi-dimensional audio features of the target audio signal; The multi-dimensional audio features are subjected to dimensionality reduction processing to obtain the dimensionality-reduced audio features of the target audio signal; By using a trained dual-talk state classification model, the reduced-dimensional audio features are subjected to dual-talk state classification processing to obtain the result of whether the target audio signal is in a dual-talk state.
2. The dual-talk status detection method according to claim 1, characterized in that, The multi-dimensional audio features include at least two of the following: target time-domain features, target frequency-domain features, target cepstral domain features, and target cross-correlation domain features. The target time-domain features include the target zero-crossing rate and the target short-time energy. The target frequency-domain features include the target energy entropy, the target spectral centroid, the target spectral roll-off point, and the target spectral flux. The target cepstral domain features include the target Mel-frequency cepstral coefficients. The target cross-correlation domain features include the target normalized cross-correlation function value, the target energy ratio, and the target high-spectral divergence.
3. The dual-talk status detection method according to claim 1, characterized in that, The step of performing dimensionality reduction processing on the multi-dimensional audio features to obtain the dimensionality-reduced audio features of the target audio signal includes: The multi-dimensional audio features are standardized to obtain standardized multi-dimensional audio features. Principal component analysis is used to perform dimensionality reduction calculations on the standardized multi-dimensional audio features to obtain the dimensionality-reduced audio features.
4. The dual-talk status detection method according to claim 1, characterized in that, The step of classifying the dimensionality-reduced audio features using a trained dual-talk state classification model to obtain a result indicating whether the target audio signal is in a dual-talk state includes: The decision value of the reduced-dimensional audio feature is obtained by performing binary classification decision calculation on the reduced-dimensional audio feature using the decision function of the trained dual-talk state classification model. If the decision value is greater than or equal to zero, the target audio signal is determined to be in dual-talk mode; If the decision value is less than zero, the target audio signal is determined to be in a single-talk state.
5. The dual-talk status detection method according to claim 1, characterized in that, Before acquiring the target audio signal, the process also includes: Construct training samples for training a pre-defined dual-state classification model; The dual-talk state classification model is trained based on the training samples to obtain the trained dual-talk state classification model.
6. The dual-talk status detection method according to claim 5, characterized in that, The training samples used to construct the pre-defined dual-talk state classification model include: Obtain the first dimension-reduced audio features of the first audio signal in a dual-talk scenario; Label the first dimensionality-reduced audio features with dual-talk status tags; The first dimensionality-reduced audio feature and the dual-talk state label are used as positive samples; Obtain the second dimension-reduced audio features of the second audio signal in a single-lecture scenario; Label the second dimensionality-reduced audio features with single-lecture status tags; The second dimensionality-reduced audio feature and the single-lecture state label are used as negative samples; The positive samples and the negative samples constitute the training samples.
7. The dual-talk status detection method according to claim 6, characterized in that, The step of training the dual-talk state classification model based on the training samples to obtain the trained dual-talk state classification model includes: The positive and negative samples are input into the dual-talk state classification model for training to calculate the decision function and obtain the trained dual-talk state classification model.
8. A dual-talk status detection device, characterized in that, include: The signal acquisition module is used to acquire the target audio signal; The feature extraction module is used to perform multi-dimensional feature extraction processing on the target audio signal to obtain multi-dimensional audio features of the target audio signal; The feature dimensionality reduction module is used to perform dimensionality reduction processing on the multi-dimensional audio features to obtain the dimensionality-reduced audio features of the target audio signal; The state classification module is used to perform dual-talk state classification processing on the dimensionality-reduced audio features using a trained dual-talk state classification model to obtain the result of whether the target audio signal is in a dual-talk state.
9. A computer device, characterized in that, The computer device includes a processor, a memory, a computer program stored in the memory and executable by the processor, and a data bus for implementing communication between the processor and the memory, wherein when the computer program is executed by the processor, it implements the steps of the dual-talk state detection method as described in any one of claims 1 to 7.
10. A storage medium for computer-readable storage, characterized in that, The storage medium stores one or more programs, which can be executed by one or more processors to implement the steps of the dual-talk state detection method according to any one of claims 1 to 7.