Real-time communication audio equipment detection method based on multiple threads
By combining multi-threaded parallel detection with neural networks, the compatibility and efficiency issues of audio device detection in real-time communication software are solved, and comprehensive coverage of multiple sets of audio drivers and device configurations is achieved, ensuring accurate identification and rapid response in complex environments, and improving users' communication quality and experience.
Patent Information
- Application Number
- CN202510780808.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-09-23
AI Technical Summary
Audio device detection in existing real-time communication software relies on a single or limited system audio interface API, resulting in a complex and time-consuming detection process. This makes it difficult to meet the needs of quickly entering the communication link, and it is unable to fully adapt to the multiple sets of audio drivers and device configurations that have been constantly changing over the years. This leads to compatibility issues and the inability to directly feedback repair solutions based on test results, which seriously restricts the improvement of audio quality.
A multi-threaded parallel detection method is adopted. By starting multiple threads corresponding to the underlying audio acquisition interface API of different operating systems, combined with automatic adjustment of sampling rate, number of channels and sampling format, and introducing a neural network voice activity detection model, combined with time domain statistical indicators and frequency domain feature extraction, it can achieve comprehensive coverage and accurate identification of multiple sets of audio drivers and device configurations.
It significantly improves the compatibility and efficiency of detection, can accurately identify effective speech in complex noisy environments, dynamically adjust the detection time to balance speed and accuracy, ensure the accuracy and stability of the final detection results, and solve the problems of poor compatibility, long time consumption and inaccurate detection in traditional detection methods, thereby improving users' communication quality and experience.
Smart Images

Figure CN120690230A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of communication audio device detection, and in particular to a multi-threaded real-time communication audio device detection method. Background Art
[0002] With the rapid development of modern information technology, software applications based on real-time communication are becoming increasingly widespread, particularly in areas such as video conferencing and instant voice communication. This places increasingly high demands on the performance and stability of audio equipment. As a crucial component of real-time communication systems, the detection and optimization of audio equipment directly impacts the user's communication quality and user experience. Therefore, research on audio device detection technology in real-time communication environments has significant practical significance and broad application prospects.
[0003] In existing real-time communication software, the detection of audio devices usually relies on a single or limited system audio interface API, and the detection process is complex and time-consuming, making it difficult to meet the user's needs for quick access to the communication link. In addition, with the frequent upgrades of operating systems, mainstream platforms such as Windows and macOS have successively launched a variety of underlying audio acquisition and playback API solutions. However, due to the lack of synchronization between the software and the operating system and the underlying driver version, compatibility issues often occur, resulting in widespread abnormalities such as microphone silence, low volume, excessive background noise, voice stuttering and sound loss. To make matters more complicated, existing technologies are difficult to fully adapt to the multiple sets of audio drivers and device configurations that have been constantly changing over the years, and the detection results often cannot directly feedback the repair plan, which seriously restricts the improvement of the audio quality of real-time communication systems.
[0004] Therefore, in real-time communication audio equipment detection, how to achieve concurrent detection of multiple system underlying audio acquisition interfaces, improve detection speed and compatibility, and ensure accurate recognition of valid voice signals in noisy or equipment abnormal environments has become an urgent problem that needs to be solved. Summary of the Invention
[0005] The present application provides a multi-threaded real-time communication audio device detection method, apparatus, computer equipment and storage medium, aiming to solve the problem that the existing technology is difficult to fully adapt to the multiple sets of audio drivers and device configurations that have been constantly changing over the years, and the detection results often cannot directly feedback the repair solution, which seriously restricts the improvement of the audio quality of the real-time communication system.
[0006] A multi-threaded real-time communication audio device detection method, the method comprising:
[0007] When the device starts, multiple threads are started, each thread corresponds to a different operating system's underlying audio acquisition interface API;
[0008] Each thread executes the audio acquisition process in parallel, specifically including: acquiring the original audio signal, calculating the time domain short-time energy and short-time zero-crossing rate indicators, and automatically adjusting the sampling rate, number of sampling channels, and sampling format parameters if the indicators are lower than a preset threshold, and reopening the corresponding audio device for acquisition;
[0009] Resample the collected audio signals to unify the output sampling rate and channel format;
[0010] The neural network-based voice activity detection model uses frame windowing and short-time Fourier transform, combined with a 64-channel Gammatone filter bank to extract frequency domain features. These frequency domain features are normalized and then input into the neural network. The neural network outputs the probability value of speech presence in each frame and smoothes the probability sequence to obtain the voice activity determination result.
[0011] A weighted comprehensive score is calculated based on the time-domain statistical indicators and the neural network voice activity detection results to determine whether the current detection period has reached the preset voice activity probability threshold. If not, the detection period is extended and data collection and analysis continues.
[0012] Summarize the detection results of all threads and select the optimal audio detection solution based on the comprehensive score and relevant sampling parameters.
[0013] In the above solution, optionally, the operating system bottom-level audio acquisition interface APIs respectively initialized and called by the multiple threads include Windows audio architecture, DirectShow and Windows multimedia interface.
[0014] In the above scheme, optionally, the automatic adjustment of sampling parameters includes trying different sampling rates, mono and stereo modes and sampling formats in sequence, and repeatedly collecting audio data and calculating time domain features for each set of parameters until the audio validity judgment condition is met or the maximum number of adjustments is reached.
[0015] In the above scheme, optionally, the resampling process uses a digital signal processing algorithm to convert the audio signal into a unified sampling rate and channel format to ensure the consistency of the neural network input data.
[0016] In the above solution, optionally, the neural network voice activity detection model is implemented by the following steps:
[0017] Perform frame and window processing on the collected audio data;
[0018] Short-time Fourier transform is used to extract time-frequency features;
[0019] Frequency domain filtering is performed through a 64-channel Gammatone filter bank to obtain a 64-dimensional frequency feature vector;
[0020] After normalizing the frequency feature vector, it is input into the pre-trained neural network model and outputs the speech probability of each frame;
[0021] A median filter is applied to smooth the speech probability sequence to remove abnormal fluctuations.
[0022] In the above solution, optionally, the voice activity probability threshold is dynamically adjusted through remote configuration, and different thresholds can be used in different threads and detection stages to adapt to changing environments.
[0023] In the above scheme, optionally, the comprehensive score is based on preset weights, and the time domain short-time energy and zero-crossing rate statistics are weighted and synthesized with the neural network speech probability to evaluate the effectiveness of the current detection scheme.
[0024] In the above solution, optionally, the detection duration adopts a dynamic increasing strategy, with an initial duration of 2 seconds, which is gradually increased to 4 seconds, 5 seconds, until the maximum preset duration is reached or the judgment condition is met.
[0025] In the above solution, optionally, the multi-thread detection results are summarized and compared in real time by the scheduling processing module, and the optimal thread and corresponding sampling parameters are dynamically selected as the final detection result based on the comprehensive score.
[0026] In the above solution, optionally, the method further includes automatically generating audio device abnormality diagnosis information and corresponding repair parameter suggestions based on the final detection result for subsequent real-time communication system call and optimization.
[0027] Compared with the prior art, this application has at least the following beneficial effects:
[0028] This application, based on further analysis and research of existing technical issues, recognizes that existing technologies struggle to fully adapt to the multiple audio drivers and device configurations that have evolved over the years, and that detection results often fail to directly inform repair solutions, severely hindering improvements in audio quality in real-time communication systems. By launching multiple threads in parallel, corresponding to the underlying audio acquisition APIs of different operating systems, this approach achieves comprehensive coverage across multiple system environments and driver versions, significantly improving detection compatibility and efficiency. This method introduces a mechanism for automatically adjusting the sampling rate, number of channels, and sampling format, effectively addressing the issue of silence or abnormal noise caused by sampling parameter mismatches and improving the effectiveness of collected audio data. Furthermore, a comprehensive scoring strategy combining a neural network voice activity detection model with time-domain statistical indicators enables the system to accurately identify valid speech in complex noisy environments, avoiding the misjudgments and omissions associated with traditional time-domain threshold detection. A dynamically increasing detection duration strategy balances detection speed and accuracy, ensuring users can quickly obtain preliminary detection results while fully analyzing abnormal situations and improving detection robustness. Real-time aggregation and optimization of multi-threaded detection results further ensure the accuracy and stability of the final detection solution. The overall solution systematically solves key problems faced by real-time communication software in existing technologies, such as poor compatibility, long detection time, abnormal sound quality, and inaccurate detection. It greatly improves the efficiency and reliability of audio device anomaly detection and ensures users' communication quality and usage experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 A flowchart of a multi-threaded real-time communication audio device detection method provided by one embodiment of the present application;
[0030] Figure 2 A schematic diagram of the operation flow of a method for implementing multi-threaded real-time communication audio device detection provided by one embodiment of the present application. DETAILED DESCRIPTION
[0031] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0032] In one embodiment, Figure 1 As shown, a multi-threaded real-time communication audio device detection method is provided, comprising the following steps:
[0033] When the device starts, multiple threads are started, each thread corresponds to a different operating system's underlying audio acquisition interface API;
[0034] Each thread executes the audio acquisition process in parallel, specifically including: acquiring the original audio signal, calculating the time domain short-time energy and short-time zero-crossing rate indicators, and automatically adjusting the sampling rate, number of sampling channels, and sampling format parameters if the indicators are lower than a preset threshold, and reopening the corresponding audio device for acquisition;
[0035] Resample the collected audio signals to unify the output sampling rate and channel format;
[0036] The neural network-based voice activity detection model uses frame windowing and short-time Fourier transform, combined with a 64-channel Gammatone filter bank to extract frequency domain features. These frequency domain features are normalized and then input into the neural network. The neural network outputs the probability value of speech presence in each frame and smoothes the probability sequence to obtain the voice activity determination result.
[0037] A weighted comprehensive score is calculated based on the time-domain statistical indicators and the neural network voice activity detection results to determine whether the current detection period has reached the preset voice activity probability threshold. If not, the detection period is extended and data collection and analysis continues.
[0038] Summarize the detection results of all threads and select the optimal audio detection solution based on the comprehensive score and relevant sampling parameters.
[0039] In this embodiment, the operating system bottom audio acquisition interface APIs respectively initialized and called by the multiple threads include Windows audio architecture, DirectShow and Windows multimedia interface.
[0040] In this embodiment, the automatic adjustment of sampling parameters includes sequentially trying different sampling rates, mono and stereo modes, and sampling formats, and repeatedly collecting audio data and calculating time domain features for each set of parameters until the audio validity judgment condition is met or the maximum number of adjustments is reached.
[0041] In this embodiment, the resampling process uses a digital signal processing algorithm to convert the audio signal into a uniform sampling rate and channel format to ensure the consistency of the neural network input data.
[0042] In this embodiment, the neural network voice activity detection model is implemented by the following steps:
[0043] Perform frame and window processing on the collected audio data;
[0044] Short-time Fourier transform is used to extract time-frequency features;
[0045] Frequency domain filtering is performed through a 64-channel Gammatone filter bank to obtain a 64-dimensional frequency feature vector;
[0046] After normalizing the frequency feature vector, it is input into the pre-trained neural network model and outputs the speech probability of each frame;
[0047] A median filter is applied to smooth the speech probability sequence to remove abnormal fluctuations.
[0048] In this embodiment, the voice activity probability threshold is dynamically adjusted through remote configuration, and different thresholds can be used in different threads and detection stages to adapt to changing environments.
[0049] In this embodiment, the comprehensive score is based on preset weights, and combines the time domain short-time energy and zero-crossing rate statistics with the neural network speech probability to evaluate the effectiveness of the current detection scheme.
[0050] In this embodiment, the detection duration adopts a dynamic increasing strategy, with an initial duration of 2 seconds, which is gradually increased to 4 seconds and 5 seconds until the maximum preset duration is reached or the judgment condition is met.
[0051] In this embodiment, the multi-thread detection results are summarized and compared in real time by the scheduling processing module, and the optimal thread and corresponding sampling parameters are dynamically selected as the final detection result based on the comprehensive score.
[0052] In this embodiment, the method further includes automatically generating audio device abnormality diagnosis information and corresponding repair parameter suggestions based on the final detection result, for subsequent real-time communication system call and optimization.
[0053] The multi-threaded real-time communication audio device detection method described in this embodiment aims to solve the problems of poor compatibility, long detection time, and insufficient detection accuracy in audio device detection in existing real-time communication software.
[0054] First, during device startup, the system simultaneously launches multiple independent threads, each bound to a corresponding underlying operating system audio capture API, such as the Windows Audio Architecture API (WASAPI), DirectShow, and Windows Multimedia Interface (WinMM). By running multiple threads in parallel, multiple audio capture solutions can be tested simultaneously, significantly improving detection speed and ensuring compatibility across different system versions and drivers.
[0055] Within each thread, the audio acquisition process includes several key steps. The thread first collects the original audio signal through the corresponding interface, and then calculates two basic statistical indicators: short-time energy in the time domain and short-time zero-crossing rate. If the time domain indicator of the collected signal is lower than the preset effective threshold, it means that the collected audio signal is silent, the volume is too low, or there is abnormal noise. The system will automatically trigger the sampling parameter adjustment mechanism, recursively try different sampling rates (such as from 16kHz to 48kHz), number of channels (mono or stereo) and sampling formats (PCM format depth, etc.), and reinitialize the audio device for collection until a valid audio signal that meets the time domain indicators is collected or the maximum number of adjustments is reached.
[0056] After obtaining a valid sampling signal, the system resamples the collected audio data, uniformly adjusting it to the preset standard sampling rate and channel format to ensure consistent input data for subsequent processing modules. Resampling utilizes high-precision digital signal processing algorithms to ensure undistorted sound quality and signal characteristics.
[0057] After preprocessing, the audio signal enters the neural network voice activity detection (VAD) module. The VAD module's specific implementation process is as follows: the audio signal is framed and windowed, and time-frequency domain features are extracted using a short-time Fourier transform. Further frequency filtering is performed using a 64-channel gammatone filter bank to obtain a 64-dimensional frequency feature vector. After normalization, the feature vector is input into a pretrained deep neural network model, which outputs the speech probability judgment value for the corresponding frame. To improve judgment stability and accuracy, the judgment results are also smoothed using median filtering and other methods to remove short-term glitches and abnormal noise, ultimately resulting in a continuous and stable voice activity judgment sequence.
[0058] This embodiment calculates a comprehensive score based on the time domain short-time energy and zero-crossing rate indicators and the speech probability output by the neural network VAD, through preset weights for weighted fusion, to evaluate the audio quality and voice activity status of the current detection period. The score serves as the basis for determining whether to continue collecting. If it is lower than the dynamically set threshold, the detection time will increase (for example, from 2 seconds to 4 seconds, 5 seconds), and the thread continues to collect more audio data and repeat the above process until the maximum time is reached or the score reaches the qualified standard. The detection results of all threads are summarized in real time through a unified scheduling processing module, and the optimal audio detection scheme is finally selected based on the comprehensive score and sampling parameters. This scheme can not only respond quickly to device anomalies, but also has high compatibility and adapts to different system versions and drive environments.
[0059] The multi-threaded concurrent detection method proposed in this embodiment significantly addresses the technical bottlenecks of low efficiency and poor compatibility in traditional real-time communication software audio device detection. By concurrently calling multiple operating system low-level audio acquisition APIs through multiple threads, it effectively addresses device compatibility issues arising from the frequent operating system and driver version updates in recent years, avoiding the missed detection of audio device anomalies caused by single-interface detection. A mechanism for automatically recursively adjusting the sampling rate, number of channels, and sampling format enables the detection process to dynamically adapt to the hardware characteristics and driver configurations of various devices, greatly improving the effectiveness of collected audio data and avoiding audio silence, low volume, or abnormal noise caused by sampling parameter mismatches. This mechanism fills the gap in the existing technology, which lacks adaptive sampling parameter adjustment. The introduction of voice activity detection technology based on deep neural networks enables the system to accurately identify valid voice signals in complex noise environments and when the device is in an abnormal state. Combining high-dimensional frequency features extracted by a multi-channel gammatone filter with probabilistic smoothing ensures the robustness and accuracy of voice detection, effectively overcoming the misjudgment and missed detection issues of traditional detection methods based on simple time-domain thresholds in noisy environments. The dynamic detection duration increment strategy not only meets users' needs for rapid detection, but also ensures the stability and accuracy of test results. Users can obtain preliminary test results in a very short time, greatly improving the user experience. Furthermore, for complex or abnormal situations, the system can automatically extend the detection time to complete sufficient data collection and analysis, enhancing detection reliability. Finally, the real-time aggregation and optimization of multi-threaded detection results balances detection efficiency and quality, ensuring that the selected detection solution is the optimal solution for the current environment, effectively improving the anomaly detection and adaptability of real-time communication software and audio equipment.
[0060] In summary, this solution systematically solves key technical problems brought about by operating system updates, such as multi-API compatibility issues, sampling parameter mismatch issues, the contradiction between detection time and accuracy, and insufficient voice recognition accuracy in complex environments. It greatly improves the audio device detection efficiency and quality of real-time communication software, ensures the stability and experience of user audio communication, and has significant practical value and promotion prospects.
[0061] like Figure 2As shown, in one embodiment, a method for implementing multi-threaded real-time communication audio device detection is provided. Due to updates and upgrades to PC host systems such as Windows and macOS, the operating systems will successively launch a variety of underlying audio device acquisition and playback system API solutions. The latest solutions often bring many benefits. However, frequent system upgrades are accompanied by application compatibility issues and the risk of degraded sound quality. This phenomenon is mainly caused by the lack of optimization of application code, incompatibility between the operating system and the underlying device driver, bugs in the operating system itself, and even damage to the left and right audio channels of the device. Therefore, when operating system updates, driver updates, and program updates all have their own release timelines, end users face abnormalities such as no microphone sound, low microphone volume, excessive microphone background noise, audio stuttering, and audio dropout when using real-time communication software for video conferencing and instant voice communication. Existing real-time communication software generally has obvious defects in device detection: first, the functional units are independent and complex, and detection takes a long time, but users are accustomed to entering the communication link directly without detection; second, due to the complexity of the user system environment, it is difficult for the new software to adapt to all PC host system changes and underlying device driver compatibility in the past 10 years. Even if the device can detect the problem, there is a lack of solution, which makes it impossible to repair it, or it is not known which system solution is suitable.
[0062] The detection process of this embodiment is as follows:
[0063] Start multiple threads at boot time, where each thread corresponds to a different API backend solution; threads refer to the WASAPI driver, dshow driver, and WinMM driver in the flowchart below running in parallel.
[0064] Based on the language characteristics, each thread is selected to execute for 2-4-5s and gradually accumulate to a total time of 1 minute. After the collection is completed, the last scheduling process is performed to summarize the data and score it.
[0065] Each execution period is executed sequentially. Time-domain short-term energy and short-term zero-crossing rate statistics are calculated for the raw audio data. If the statistics are too low, the system API is recursively reconfigured with a new sampling rate, number of sampling channels, and sampling format, and device statistics are enabled again. The next step is audio resampling to prepare for output adaptation. The resampled audio is then passed through a neural network VAD speech detection algorithm to calculate the decision value (frame windowing -> STFT -> frequency filtering of each frame using a 64-channel gammatone filter bank, outputting a 64-dimensional vector -> feature normalization -> 64-dimensional features are input to the neural network -> the probability of speech in the current frame is predicted -> post-processing is performed by smoothing the output probability through the post-processing step ift (e.g., a median filter) to remove short glitches, and the decision value is output. The resulting VAD score is the result. This process is shown in the bottom row of the flowchart. Any commonly used neural network VAD solution can be used; this is not the subject of our protection. Our protection is to consider this process as using VAD with post-VAD smoothing (e.g., a median filter). The total decision results for the time period are calculated. The current value is calculated. If it is lower than the expected value, testing continues. The expected value can be configured remotely and can vary between modes. For example, a uniform value of 0.8 can be used initially. The optimal solution is selected based on the cumulative score of the low-weighted time domain results and the high-weighted VAD results. The most effective audio mode is determined. Each mode has detailed parameters, and the parameters to be used are confirmed when selecting. This can be seen in the node before the end node in the figure.
[0066] This embodiment uses a gradually increasing calculation time. If the sound is suitable, the comparison result can be obtained in as little as 2 seconds. The selected solution avoids fixed-duration device detection. The system adapts to multiple solutions such as WASAPI, WinMM, DirectSound, and CoreAudio, performing simultaneous multi-threaded detection to reduce detection time. The integrated time-domain signal threshold detection and neural network detection can handle primary sound and silence, voice and noise interference, and resampled data inversion.
[0067] This embodiment provides concurrency testing, supports testing multiple API solutions, and comprehensively assesses the startup time and data quality of each API. Test results are returned quickly, using aggressive testing to return results first and then gradually reducing the test duration to 10 seconds for stable testing, with adjustable policies. Test results incorporate statistical methods for quiet environments and neural network-based VAD detection to address noisy environments, resulting in faster, more reliable results.
[0068] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
Claims
1. A multi-threaded real-time communication audio device detection method, characterized in that: The method comprises: When the device starts, multiple threads are started, each thread corresponds to a different operating system's underlying audio acquisition interface API; Each thread executes the audio acquisition process in parallel, specifically including: acquiring the original audio signal, calculating the time domain short-time energy and short-time zero-crossing rate indicators, and automatically adjusting the sampling rate, number of sampling channels, and sampling format parameters if the indicators are lower than a preset threshold, and reopening the corresponding audio device for acquisition; Resample the collected audio signals to unify the output sampling rate and channel format; The neural network-based voice activity detection model uses frame windowing and short-time Fourier transform, combined with a 64-channel Gammatone filter bank to extract frequency domain features. These frequency domain features are normalized and then input into the neural network. The neural network outputs the probability value of speech presence in each frame and smoothes the probability sequence to obtain the voice activity determination result. A weighted comprehensive score is calculated based on the time-domain statistical indicators and the neural network voice activity detection results to determine whether the current detection period has reached the preset voice activity probability threshold. If not, the detection period is extended and data collection and analysis continues. Summarize the detection results of all threads and select the optimal audio detection solution based on the comprehensive score and relevant sampling parameters.
2. The method according to claim 1, characterized in that The operating system bottom audio acquisition interface APIs respectively initialized and called by the multiple threads include Windows audio architecture, DirectShow and Windows multimedia interface.
3. The method according to claim 1, characterized in that The automatic adjustment of sampling parameters includes sequentially trying different sampling rates, mono and stereo modes, and sampling formats, and repeatedly collecting audio data and calculating time domain features for each set of parameters until the audio validity determination condition is met or the maximum number of adjustments is reached.
4. The method according to claim 1, wherein The resampling process uses a digital signal processing algorithm to convert the audio signal into a uniform sampling rate and channel format to ensure the consistency of the neural network input data.
5. The method according to claim 1, wherein The neural network voice activity detection model is implemented by the following steps: Perform frame and window processing on the collected audio data; Short-time Fourier transform is used to extract time-frequency features; Frequency domain filtering is performed through a 64-channel Gammatone filter bank to obtain a 64-dimensional frequency feature vector; After normalizing the frequency feature vector, it is input into the pre-trained neural network model and outputs the speech probability of each frame; A median filter is applied to smooth the speech probability sequence to remove abnormal fluctuations.
6. The method according to claim 1, characterized in that The voice activity probability threshold is dynamically adjusted through remote configuration, and different thresholds can be used in different threads and detection stages to adapt to changing environments.
7. The method according to claim 1, characterized in that The comprehensive score is based on preset weights, and combines the time domain short-time energy and zero-crossing rate statistics with the neural network speech probability to evaluate the effectiveness of the current detection scheme.
8. The method according to claim 1, characterized in that The detection duration adopts a dynamic increasing strategy, with an initial duration of 2 seconds, and gradually increasing to 4 seconds and 5 seconds until the maximum preset duration is reached or the judgment condition is met.
9. The method according to claim 1, characterized in that The multi-thread detection results are summarized and compared in real time by the scheduling processing module, and the optimal thread and corresponding sampling parameters are dynamically selected as the final detection result based on the comprehensive score.
10. The method according to claim 1, characterized in that The method further includes automatically generating audio device abnormality diagnosis information and corresponding repair parameter suggestions based on the final detection result for subsequent real-time communication system call and optimization.
Citation Information
Cited By
Multi-level sound effect switching method and system
CN121151761A