Intelligent audio optimization method
By employing deep learning models for environmental assessment and dynamic processing strategies, the problems of unstable voice communication quality and low resource utilization efficiency were solved, achieving stable communication and reduced energy consumption in complex networks and noisy environments.
Patent Information
- Application Number
- CN202511627953.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-07
- Publication Date
- 2026-02-17
AI Technical Summary
Existing voice communication technologies struggle to maintain high-quality audio and efficient energy use in complex network environments and noisy background environments, and fixed bandwidth allocation methods lead to excessive energy consumption and resource waste.
By collecting speech signals and network state parameters, inputting them into a deep learning model for joint analysis, generating environmental assessment results, dynamically generating speech activity detection thresholds, and performing discontinuous transmission processing or comfort noise supplementation processing under different states, including discontinuous transmission processing and comfort noise supplementation processing.
Maintain stable communication quality in complex networks and noisy backgrounds, reduce device power consumption, and improve user experience and transmission efficiency.
Smart Images

Figure CN121545535A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the technical field of audio optimization, and in particular to an intelligent audio optimization method. Background Technology
[0002] Currently, voice communication has become an indispensable means of people's daily life and work. However, under complex network conditions and noisy background environments, traditional voice communication technology often struggles to balance high-quality audio effects with efficient energy utilization. Most existing audio processing methods rely on preset static noise reduction parameters and fixed bandwidth allocation strategies. While these methods can alleviate noise interference to some extent, they lack dynamic adaptability and sophisticated signal processing mechanisms. When network conditions fluctuate frequently or external noise levels change drastically, existing systems often struggle to adjust their processing strategies in a timely manner, leading to a significant deterioration in call quality. Furthermore, under conditions of limited network resources, fixed bandwidth allocation can result in excessive energy consumption and wasted transmission resources. Summary of the Invention
[0003] To address the issues of decreased call quality and insufficient energy efficiency in existing voice communication technologies under complex network environments and high background noise, this application provides an intelligent audio optimization method.
[0004] A smart audio optimization method, the smart audio optimization method comprising: The system collects voice signals and network state parameters during the current voice communication process, and inputs the voice signals and network state parameters into a deep learning model for joint analysis to generate corresponding environmental assessment results. Based on the environmental assessment results, a voice activity detection threshold is dynamically generated, and the voice activity detection threshold is used to detect the voice signal to determine whether the user is in a voice state. If the user is not in voice mode, the network status information is invoked to perform discontinuous transmission processing; If the user is in a voice state and there are voice gaps in the voice signal, the voice activity features in the environmental assessment results are called to perform comfort noise supplementation processing. The comfort noise supplementation processing is used to generate simulated background sound according to a preset noise library or user personalized configuration to be inserted into the voice signal. The output is the target communication audio stream generated after the discontinuous transmission processing or the comfort noise supplementation processing.
[0005] By adopting the above technical solution, voice signals and network status parameters are simultaneously collected during voice communication and input into a deep learning model to generate environmental assessment results. Based on these results, a voice activity detection threshold is dynamically generated. This enables accurate identification of the user's voice state in a real-time environment and triggers discontinuous transmission processing or comfort noise supplementation processing under different states. This achieves the technical effect of maintaining stable communication quality and reducing equipment energy consumption under complex network and noisy background conditions.
[0006] Preferably, in the step of inputting the speech signal and the network state parameters into a deep learning model for joint analysis to generate a corresponding environmental assessment result, the deep learning model includes at least a parsing sub-model, an extraction sub-model, and a correction sub-model, and the step includes: Based on the parsing sub-model, the network state parameters are parsed to generate corresponding state information and feature parsing strategies; Based on the extraction sub-model and the feature parsing strategy, feature extraction is performed on the speech signal to generate corresponding speech activity features; By integrating the voice activity features and the state information, a corresponding environmental assessment result is generated.
[0007] By adopting the above technical solution, and by introducing the parsing sub-model, extraction sub-model, and correction sub-model, the network state parameters and speech signals are parsed, extracted, and integrated to generate environmental assessment results, making the environmental assessment process more refined and hierarchical. This ensures that network information and speech features can be accurately coupled, providing a reliable basis for subsequent threshold generation and processing strategies.
[0008] Preferably, the step of parsing the network state parameters based on the parsing sub-model to generate corresponding state information and feature parsing strategies includes: Based on the parsing sub-model, the network state parameters are parsed to obtain the corresponding state indicators, which include bandwidth utilization indicators, latency fluctuation indicators and packet loss rate indicators. The weighted average of each of the aforementioned state indicators is processed to generate corresponding state information, and the corresponding network state score is determined based on the scoring mapping relationship of the state information. Determine whether the network state score is greater than a preset state threshold, generate a corresponding judgment result, and determine a corresponding feature parsing strategy based on the judgment result. The feature parsing strategy includes at least a network good parsing strategy and a network abnormal parsing strategy.
[0009] By adopting the above technical solution, when parsing network status parameters, status information is generated by weighted average of bandwidth usage index, latency fluctuation index and packet loss rate index, and network status score is obtained by mapping. Then, a good network parsing strategy or an abnormal network parsing strategy is determined based on the score, so that the system can flexibly select the processing path according to different network conditions and improve its adaptability to complex network environments.
[0010] Preferably, the step of extracting features from the speech signal based on the extraction sub-model and the feature parsing strategy to generate corresponding speech activity features includes: If the feature parsing strategy is a network-good parsing strategy, then the speech signal is directly extracted based on the extraction sub-model to generate corresponding speech activity features; If the feature parsing strategy is a network anomaly parsing strategy, then the speech signal is divided into confidence levels to filter out speech segments with target confidence levels, and features are extracted from the speech segments with target confidence levels to generate corresponding speech activity features.
[0011] By adopting the above technical solution, when extracting features from speech signals, if a good network parsing strategy is used, speech activity features are directly generated. However, under the abnormal network parsing strategy, confidence level division is performed and only speech segments with target confidence levels are extracted. This differentiated processing method can avoid low-confidence segments occupying bandwidth when network quality is insufficient, while retaining more speech details when network quality is good, thereby improving overall transmission efficiency and user experience.
[0012] Preferably, the step of dividing the speech signal into confidence levels to filter out speech segments with target confidence levels includes: According to the network anomaly analysis strategy, the shortening range of the preset time window is determined, and the preset time window is shortened based on the shortening range to generate a corresponding target time window; The speech signal is segmented according to the target time window to generate multiple segments; Confidence scores are assigned to each of the extracted segments, and the confidence score of each of the extracted segments is calculated. Each confidence score is compared with a preset confidence threshold, and the segments with confidence scores greater than or equal to the confidence threshold are selected as the target confidence speech segments.
[0013] By adopting the above technical solution, when performing confidence segmentation, the preset time window is shortened and the speech signal is divided into multiple segments. Then, the confidence score of each segment is calculated and compared with the preset threshold to select the speech segment with the target confidence score. This makes the speech feature segmentation granularity more refined, effectively avoids low-quality segments from entering subsequent processing, and improves the accuracy of feature extraction and transmission.
[0014] Preferably, the step of determining the shortening range of the preset time window according to the network anomaly analysis strategy includes: According to the network anomaly analysis strategy, the pre-stored computing resource limits are invoked, and the maximum amplitude is determined according to the amplitude mapping relationship of the computing resource limits; Based on the difference in the scoring thresholds included in the judgment result, a shortening range not greater than the maximum value of the range is determined.
[0015] By adopting the above technical solution, when determining the shortening range of the time window, the maximum value of the range is mapped by combining the network anomaly analysis strategy with the pre-stored computing resource limit, and then the final shortening range is determined according to the difference in the scoring threshold. This allows for dynamic adjustment of the time window while ensuring that the computing capacity is within acceptable limits, preventing excessive shortening that could lead to feature instability, and ensuring the discrimination accuracy under network anomaly conditions.
[0016] Preferably, the step of scoring the confidence of each of the extracted segments and calculating the confidence score of each of the extracted segments includes: Calculate the short-time energy and zero-crossing rate for each of the truncated segments to generate corresponding temporal feature values; Calculate the spectral entropy and dominant frequency distribution range for each of the extracted segments, and generate corresponding frequency domain feature values; The time-domain feature value and the frequency-domain feature value are weighted and summed according to a preset weight to generate the confidence score of the extracted segment.
[0017] By adopting the above technical solution, when scoring the confidence of the extracted segments, the time-domain feature value is obtained by calculating the short-time energy and zero-crossing rate, and the frequency-domain feature value is obtained by combining the spectral entropy and the main frequency distribution range. The confidence score is generated by weighted summation. This multi-dimensional feature fusion scoring method effectively improves the accuracy of confidence determination and makes the selected target segments more representative of effective speech.
[0018] Preferably, the step of dynamically generating a speech activity detection threshold based on the environmental assessment results includes: Based on the speech activity characteristics in the environmental assessment results, noise baseline parameters are extracted from the non-speech region of the speech signal. The noise baseline parameters include at least the short-time energy mean and the short-time energy variance. Based on the state information, a threshold adjustment weight factor is determined, and the threshold adjustment weight factor is mapped to the noise baseline parameter to generate a corresponding first detection threshold. The first detection threshold is subjected to an exponential moving average to obtain a smoothed threshold. An uplink threshold and a downlink threshold are generated by the smoothed threshold. Threshold boundary constraints are applied to the uplink threshold and the downlink threshold to limit the numerical range between the uplink threshold and the downlink threshold to between a preset minimum threshold value and a preset maximum threshold value, so as to generate a corresponding voice activity detection threshold.
[0019] By adopting the above technical solution, when dynamically generating the speech activity detection threshold, noise baseline parameters are extracted using speech activity features and the threshold adjustment factor is determined by combining state information. Then, the uplink and downlink thresholds are generated through exponential moving average and hysteresis control, and boundary constraints are applied to both to finally obtain a stable and reliable detection threshold. This avoids misjudgment caused by environmental noise fluctuations and improves the accuracy and robustness of speech state detection.
[0020] Preferably, the step of invoking the network status information to perform discontinuous transmission processing includes: The corresponding transmission rate is determined according to the state information and a preset rate mapping relationship. Perform the corresponding discontinuous transmission processing according to the transmission rate.
[0021] By adopting the above technical solution, when performing discontinuous transmission processing, the transmission rate is determined by matching the network status score with the preset rate mapping relationship, and the transmission process of voice signal is compressed accordingly. This enables the system to adaptively adjust the transmission bandwidth when the network status changes, which reduces unnecessary data transmission and ensures the stability and energy efficiency of the communication link.
[0022] Preferably, the step of calling the speech activity features in the environmental assessment results and performing comfort noise supplementation processing includes: When the speech activity features indicate the presence of speech gaps in the speech signal, the gap length and background noise level of the speech gaps are determined. Based on the gap length and background noise level, a corresponding noise sample is selected from a preset noise library or user-customized configuration, and the noise sample is adjusted in amplitude and trimmed in duration to generate a corresponding processed sample. The processed sample is identified as comfort noise and inserted into the speech gap to complete the comfort noise supplementation process.
[0023] By adopting the above technical solution, when performing comfort noise supplementation processing, when the speech activity feature detects a speech gap, a suitable noise sample is selected according to the gap length and background noise level, and the amplitude and duration are adjusted before being inserted into the speech gap. This prevents abrupt silence during the call, thereby ensuring the continuity and naturalness of the communication and improving the user's immersion and comfort.
[0024] In summary, this application includes at least one of the following beneficial technical effects: This application addresses the issues of unstable communication quality and low resource utilization efficiency by introducing environmental assessment and adaptive adjustment during voice communication. First, voice signals and network state parameters are collected and fed into a deep learning model for joint analysis, generating an environmental assessment result that comprehensively reflects the current network condition and voice characteristics. This allows subsequent processing to move beyond fixed parameters and possess dynamic adaptability. A dynamic voice activity detection threshold is generated based on the environmental assessment result, and this threshold is used to detect voice signals, thereby more accurately distinguishing between when the user is in a voice state and when not in a voice state. When a user is detected not in a voice state, the system invokes network state information to perform discontinuous transmission processing, compressing the voice signal transmission rate while maintaining the communication link, thus reducing bandwidth usage and energy consumption. When the user is in a voice state and there are speech gaps in the voice signal, the system further invokes the voice activity features from the environmental assessment result to perform comfort noise supplementation processing. By inserting simulated background noise from a preset noise library or user-customized settings into the speech gaps, the natural continuity and immersive experience of the communication process are ensured. Ultimately, the system outputs the target communication audio stream after discontinuous transmission processing or comfort noise supplementation processing. Through this adaptive optimization method across the network and voice dimensions, it effectively solves the problems of decreased call quality and energy waste in complex networks and noisy backgrounds, maintains stable voice communication quality and reduces device power consumption in variable environments. Attached Figure Description
[0025] Figure 1 This is a flowchart of an intelligent audio optimization method according to an embodiment of this application. Detailed Implementation
[0026] The present application will be further described in detail below with reference to the accompanying drawings.
[0027] In one embodiment, such as Figure 1 As shown, this application discloses an intelligent audio optimization method, which includes: S10. Collect the voice signal and network status parameters during the current voice communication process, and input the voice signal and network status parameters into the deep learning model for joint analysis to generate the corresponding environmental assessment results. The voice signal refers to analog or digital waveform data containing voice information collected by the user terminal. This data is acquired through a microphone array or a single pickup unit and converted from analog to digital to form a discrete signal stream that can be processed by the computing module. Network status parameters are operational indicators that characterize the current communication link status, including bandwidth utilization, transmission delay fluctuations, and packet loss rate. These parameters can be obtained through the real-time monitoring module in the communication protocol stack and are provided numerically for subsequent processing. The deep learning model is an analysis module composed of multiple neural network layers. It simultaneously receives both voice signals and network status parameters as input. The parsing sub-model analyzes the network status parameters to obtain status information and a preliminary parsing strategy. The extraction sub-model extracts multi-dimensional features from the voice signal to obtain voice activity features. Finally, the correction sub-model weightedly integrates the voice activity features and status information to generate the environmental assessment results.
[0028] First, the speech signal acquired by the acquisition module is segmented and windowed to form speech frame data with temporal continuity. Each frame contains several time-domain and frequency-domain features, such as short-time energy, zero-crossing rate, and spectral distribution. Simultaneously, the network status monitoring module outputs real-time bandwidth usage, latency fluctuation, and packet loss rate metrics as another type of input parameter. Subsequently, the parsing sub-model in the deep learning model performs normalization and weighting operations on these three metrics to generate unified state information and outputs a network status score based on a predefined scoring mapping relationship. Next, the extraction sub-model extracts features from the speech signal frames, using convolutional computation and time-frequency analysis to obtain speech activity features. Finally, the correction sub-model fuses the speech activity features with the state information, outputting an environmental assessment result that includes the current speech state, noise level, and network quality evaluation.
[0029] S20. Based on the environmental assessment results, dynamically generate a voice activity detection threshold, and use the voice activity detection threshold to detect the voice signal to determine whether the user is in a voice state. The voice activity detection threshold is a judgment boundary value used to distinguish between voice segments and non-voice segments in a voice signal. Its value determines the sensitivity and stability of the detection module when facing noise interference or signal fluctuations.
[0030] S30. If the user is not in voice mode, call the network status information to perform discontinuous transmission processing. Discontinuous transmission processing is used to reduce the transmission rate of audio data packets. It can dynamically generate voice activity detection thresholds in complex noise and network environments. It not only ensures the stability of the thresholds under static noise conditions, but also has the ability to adjust in real time according to network status and noise level, thereby effectively improving the accuracy of distinguishing between voice segments and non-voice segments.
[0031] S40. If the user is in a speech state and there are speech gaps in the speech signal, the system calls upon the speech activity features from the environmental assessment results to perform comfort noise supplementation processing. Comfort noise supplementation processing generates simulated background noise based on a preset noise library or user-defined configurations to be inserted into the speech signal. Comfort noise supplementation processing means that when the user is in a speech state but there are speech gaps in the speech signal, the system fills the blank intervals by inserting appropriate background noise to avoid the unnaturalness and abruptness caused by complete silence. Speech gaps are determined by the speech activity features from the environmental assessment results. Speech activity features include speech energy distribution, zero-crossing rate changes, and spectral fluctuations. These parameters can be used to identify whether speech has temporarily stopped. When a gap is detected in the speech signal, the system first calculates the length of the speech gap and obtains the current background noise level by combining it with the noise baseline parameters. Subsequently, the comfort noise generation module selects a suitable noise sample from the preset noise library or user-defined configurations based on the gap length and background noise level, such as breathing sounds, soft keyboard sounds, or gentle ambient sounds. Then, the system performs amplitude adjustment on the selected noise sample to make its intensity consistent with the current background noise, and performs duration trimming or splicing to match its length with the gap. Finally, the processed noise samples are inserted as comfort noise into the speech gaps, and after being spliced with the original speech signal, a continuous target communication audio stream is generated.
[0032] S50. Output the target communication audio stream generated after discontinuous transmission processing or comfort noise supplementation processing. The target communication audio stream refers to the output audio signal stream after discontinuous transmission processing or comfort noise supplementation processing during voice communication. This signal stream has different coding and continuity characteristics from the original speech signal, and can adapt to network conditions and meet the user's auditory needs. In this embodiment, by simultaneously acquiring speech signals and network state parameters during voice communication and inputting them into a deep learning model to generate environment evaluation results, and then dynamically generating a speech activity detection threshold based on these results, the system can accurately identify the user's speech state in a real-time environment and trigger discontinuous transmission processing or comfort noise supplementation processing in different states. This maintains stable communication quality and reduces device power consumption under complex network and noisy background conditions. After the speech signal is processed by the preceding steps, the system first buffers and reassembles the obtained audio frames to ensure that the blank frames discarded in the discontinuous transmission mode are correctly omitted, while the noise samples inserted in the comfort noise supplementation mode can be seamlessly spliced with the original speech frames. Subsequently, the encoding module performs compression encoding on the target audio stream according to the transmission control strategy. For example, it allocates different encoding rates to audio frames through adaptive bitrate control to minimize redundant information when bandwidth is limited, while maintaining high voice fidelity when bandwidth is sufficient. Furthermore, the system integrates the target audio stream into data packets that conform to the transmission protocol through packetization and synchronization mechanisms, ensuring the continuity of timestamps and frame sequences and avoiding playback interruptions or delays at the receiving end.
[0033] Furthermore, in the step of inputting speech signals and network state parameters into a deep learning model for joint analysis to generate corresponding environmental assessment results, the deep learning model includes at least a parsing sub-model, an extraction sub-model, and a correction sub-model. The steps include: S101. Based on the parsing sub-model, network state parameters are parsed to generate corresponding state information and feature parsing strategies. The parsing sub-model receives network state parameters such as bandwidth usage, latency fluctuation, and packet loss rate, and normalizes and weights them to obtain a unified set of state indicators. Subsequently, the sub-model generates state information based on the state indicator mapping, and simultaneously generates corresponding feature parsing strategies based on the scoring results to guide subsequent speech signal processing. For example, when the state information indicates high network latency, the feature parsing strategy tends to tighten the time window, thereby enhancing the ability to extract high-confidence speech segments.
[0034] S102. Based on the extraction sub-model and feature parsing strategy, features are extracted from the speech signal to generate corresponding speech activity features. Under the constraints of the feature parsing strategy, the extraction sub-model performs multi-dimensional feature extraction on the speech signal. Specifically, the system performs frame-by-frame windowing processing on the speech signal and calculates indicators such as short-time energy, zero-crossing rate, spectral entropy, and dominant frequency distribution in each frame to generate corresponding speech activity features. When the feature parsing strategy is a good network parsing strategy, the extraction sub-model directly extracts features from the complete speech signal; when the feature parsing strategy is an abnormal network parsing strategy, feature extraction is prioritized for the speech segment with the target confidence level to improve processing efficiency.
[0035] S103. Integrate speech activity features and state information to generate corresponding environmental assessment results. The modified sub-model integrates state information and speech activity features with weights to generate the final environmental assessment result. This environmental assessment result not only reflects the characteristics of the speech signal in the time and frequency domains but also includes a comprehensive evaluation of network quality, thus providing a reliable basis for subsequent speech activity detection threshold generation, discontinuous transmission processing, and comfort noise supplementation processing.
[0036] Furthermore, the steps of parsing the network state parameters based on the parsing sub-model to generate corresponding state information and feature parsing strategies include: S1011. Based on the analytical sub-model, the network state parameters are analyzed to obtain the corresponding state indicators, including bandwidth utilization, latency fluctuation, and packet loss rate. The modified sub-model weightedly integrates the state information output from S101 with the speech activity features output from S102 to generate the final environmental assessment result. This environmental assessment result not only reflects the characteristics of the speech signal in the time and frequency domains but also includes a comprehensive evaluation of network quality, thus providing a reliable basis for subsequent speech activity detection threshold generation, discontinuous transmission processing, and comfort noise supplementation processing.
[0037] S1012. A weighted average is applied to each status indicator to generate corresponding status information, and the corresponding network status score is determined based on the scoring mapping relationship of the status information. The system performs a weighted average of the above status indicators to obtain unified status information. The allocation of weighting factors can be adjusted according to the actual scenario. For example, in a video conferencing scenario, the weight of bandwidth usage can be increased, while in a weak network scenario, the weight of latency fluctuation and packet loss rate indicators can be increased. Subsequently, the parsing sub-model calculates the network status score based on the status information and the preset scoring mapping relationship. This score fluctuates between 0 and 1 and is used to quantify the overall quality of the current communication link.
[0038] S1013. Determine whether the network state score is greater than a preset state threshold, generate the corresponding judgment result, and determine the corresponding feature parsing strategy based on the judgment result. The feature parsing strategy includes at least a network good parsing strategy and a network abnormal parsing strategy. The system compares the network state score with the preset state threshold and generates the corresponding judgment result. When the network state score is higher than the preset threshold, it is determined to be a network good state. The parsing sub-model determines to adopt the network good parsing strategy, that is, to perform full feature extraction on the speech signal to retain more speech details. When the network state score is lower than or equal to the preset threshold, it is determined to be a network abnormal state. The parsing sub-model determines to adopt the network abnormal parsing strategy, that is, to divide the speech signal into confidence levels and perform feature extraction only on the speech segment with the target confidence level to reduce processing complexity and bandwidth consumption.
[0039] Furthermore, the step of extracting features from the speech signal and generating corresponding speech activity features based on the extraction sub-model and feature parsing strategy includes: S1021. If the feature parsing strategy is a network-optimized parsing strategy, then features are directly extracted from the speech signal based on the extraction sub-model to generate corresponding speech activity features. After the speech signal undergoes framing and windowing processing, the extraction sub-model calculates parameters such as short-time energy, zero-crossing rate, fundamental frequency distribution, and spectral entropy, and generates a complete set of speech activity features. Because the network is in good condition, the system can tolerate a high data transmission rate, thus eliminating the need for filtering or compressing the speech signal, thereby ensuring the integrity and fidelity of the speech feature information.
[0040] S1022. If the feature parsing strategy is a network anomaly parsing strategy, then the speech signal is divided into confidence levels to filter out speech segments with target confidence levels, and features are extracted from these target confidence speech segments to generate corresponding speech activity features. After the speech signal undergoes framing and windowing processing, the extraction sub-model calculates parameters such as short-time energy, zero-crossing rate, fundamental frequency distribution, and spectral entropy, and generates a complete set of speech activity features. Because the network is in good condition, the system can tolerate a high data transmission rate, therefore there is no need to filter or compress the speech signal, thus ensuring the integrity and fidelity of the speech feature information.
[0041] Furthermore, the step of dividing the speech signal into confidence levels to filter out speech segments with target confidence levels includes: S10221. Based on the network anomaly analysis strategy, determine the shortening range of the preset time window, and shorten the preset time window based on the shortening range to generate the corresponding target time window. The preset time window is the basic time unit for speech signal segmentation, and its size directly affects the precision of segmentation. When the network condition is poor, the system will shorten the time window to analyze speech features within a shorter time scale, thereby avoiding the mixing of weak speech segments with invalid data. The shortening range is determined based on the mapping relationship between the network state score and the computational resource limit in the state information to obtain the final target time window.
[0042] S10222: The speech signal is segmented according to the target time window to generate multiple truncated segments. Each truncated segment retains local time and frequency domain characteristics, providing a basic unit for subsequent confidence scoring. When segmenting the speech signal, it is first input into the segmentation module in the form of a linear time series. This module segments the speech signal segment by segment based on the length parameter of the target time window. Specifically, if the target time window is set to Tw milliseconds, the system starts from the beginning of the speech signal and sequentially extracts continuous samples of length Tw as a truncated segment. When the sample length is less than Tw, it can be padded with zero padding or mirror compensation to ensure that all segments have consistent lengths. The segmentation module uses a "frame shift + frame length" approach, where the target time window is used as the frame length, and the frame shift is set according to a preset overlap ratio. For example, when the target time window is 20ms and the frame shift is 10ms, there will be a 50% overlap between adjacent segments, thus avoiding the loss of speech boundary features. This approach ensures both the integrity of the segment's temporal coverage and improves the continuity and robustness of subsequent confidence scoring. The segmentation module uses a "frame shift + frame length" method, where the target time window is used as the frame length, and the frame shift is set according to a preset overlap ratio. For example, when the target time window is 20ms and the frame shift is 10ms, there will be a 50% overlap between adjacent segments, thus avoiding the loss of speech boundary features. This approach ensures both the integrity of the segment's temporal coverage and improves the continuity and robustness of subsequent confidence scoring.
[0043] S10223. Calculate the confidence score for each extracted segment by assigning a confidence score to it. Extract time-domain features such as short-time energy and zero-crossing rate from each segment, then calculate frequency-domain features such as spectral entropy and dominant frequency distribution range. Finally, perform a weighted calculation on these features according to preset weights to generate the corresponding confidence score. This confidence score is used to measure whether there are effective speech components in the segment and its reliability.
[0044] S10224. Compare each confidence score with a preset confidence threshold, and select segments with confidence scores greater than or equal to the confidence threshold as target confidence speech segments. Target confidence speech segments are considered to be the most representative part of the effective speech signal and will be prioritized for subsequent feature extraction processes.
[0045] Furthermore, the step of determining the shortening range of the preset time window based on the network anomaly analysis strategy includes: S102211. According to the network anomaly analysis strategy, the pre-stored computing resource limits are invoked, and the maximum amplitude is determined based on the amplitude mapping relationship of the computing resource limits. The purpose of introducing computing resource limits is to ensure that the computational complexity required by the system does not exceed the processing capacity of the processor or hardware platform when performing time window shortening operations. Since shortening the time window will result in more segmentation and a higher feature calculation frequency, unconstrained operation may lead to excessive processing latency or even loss of real-time performance. Therefore, the system needs to invoke the pre-stored computing resource limits before shortening the window. Computing resource limits are usually pre-set based on processor performance, allocable memory size, and real-time requirements, such as the maximum number of segments allowed to be processed per second or the maximum number of feature calculations allowed. The system invokes the computing resource limits based on the network anomaly analysis strategy and determines the maximum amplitude through the amplitude mapping relationship. The maximum amplitude refers to the maximum proportion by which the time window can be shortened without exceeding the computing resource limits. For example, when the resource limit allows a maximum of 200 segments to be processed per second, and the original window size corresponds to 100 segments / second, the maximum amplitude is 50%, indicating that the time window can be shortened by up to half. This ensures that shortening the operation will not overload the system.
[0046] S102212. Based on the score threshold difference included in the judgment result, a shortening range not exceeding the maximum amplitude is determined. The system determines the actual shortening range based on the score threshold difference in the judgment result. The score threshold difference reflects the degree of deviation between the current network state and the preset threshold. For example, the larger the difference between the network state score and the threshold, the worse the network condition, requiring more refined segmentation. The system maps the score threshold difference to a specific shortening ratio, while ensuring that this ratio does not exceed the aforementioned determined maximum amplitude. In this way, if the network state deteriorates slightly, the shortening range is small; if the network state deteriorates severely, the shortening range is close to the maximum value, but still within the range that computing resources can bear. This achieves dual control of performance constraints and network drive: it avoids computing resource overload caused by excessive shortening of the time window, and ensures more refined segmentation and confidence determination of the speech signal when the network state is abnormal, thereby improving the robustness and real-time performance of the system in complex environments.
[0047] Furthermore, the step of scoring the confidence level of each extracted segment and calculating the confidence score of each extracted segment includes: S102231. Calculate the short-time energy and zero-crossing rate for each truncated segment to generate corresponding time-domain feature values. First, the target truncated segment of the speech signal is represented as a discrete sampling point sequence, and windowing is applied to this sequence to reduce boundary effects. Subsequently, the system calculates the short-time energy value of the segment by squaring and summing the sampling points, which is used to characterize the overall energy of the speech signal within the segment. At the same time, the system counts the number of sign changes at the sampling points in the segment, i.e., the number of zero-crossing points where the positive and negative signs change, and normalizes it to obtain the zero-crossing rate, which is used to characterize the frequency components and noise characteristics of the speech signal in the segment. Through the above calculations, the generated short-time energy value and zero-crossing rate together serve as the time-domain feature values of the truncated segment, which can accurately reflect the energy distribution and coarse spectral characteristics of the speech segment in the time domain.
[0048] S102232. Calculate the spectral entropy and dominant frequency distribution range for each extracted segment, generating corresponding frequency domain feature values. First, the time-domain signal of the extracted segment is converted into a frequency-domain spectrum using a Fast Fourier Transform (FFT) to obtain the amplitude distribution of the segment at different frequencies. Subsequently, the system normalizes the spectral amplitude to form a probability distribution sequence, and calculates the Shannon entropy based on this sequence to obtain the spectral entropy value. This spectral entropy value reflects the uniformity of the spectral distribution; a higher value indicates a more dispersed signal energy distribution and a larger proportion of noise components, while a lower value indicates a higher energy concentration, making it more likely to be a valid speech signal. Simultaneously, the system searches for the dominant frequency with the highest energy concentration in the frequency-domain spectrum and statistically analyzes the energy coverage interval within a preset bandwidth above and below it, thereby determining the dominant frequency distribution range, which reflects the main frequency characteristics of the speech segment. The spectral entropy and dominant frequency distribution range obtained through the above calculations serve as the frequency domain feature values for the extracted segment. Combined with the time-domain feature values, this provides more comprehensive feature support for subsequent confidence score calculations.
[0049] S102233. The time-domain feature values and frequency-domain feature values are weighted and summed according to preset weights to generate a confidence score for the truncated segment. First, the time-domain signal of the truncated segment is converted into a frequency-domain spectrum through a Fast Fourier Transform to obtain the amplitude distribution of the segment at different frequencies. Subsequently, the system normalizes the spectral amplitude to form a probability distribution sequence, and calculates the Shannon entropy based on this sequence to obtain the spectral entropy value. This spectral entropy value reflects the uniformity of the spectral distribution; a higher value indicates a more dispersed signal energy distribution and a larger proportion of noise components, while a lower value indicates a higher energy concentration and a greater likelihood of a valid speech signal. At the same time, the system finds the dominant frequency with the highest energy concentration in the frequency-domain spectrum and statistically analyzes the energy coverage range within a preset bandwidth above and below it to determine the dominant frequency distribution range, which reflects the main frequency characteristics of the speech segment. The spectral entropy and dominant frequency distribution range obtained through the above calculations are used together as the frequency-domain feature values of the truncated segment. Combined with the time-domain feature values, this provides more comprehensive feature support for subsequent confidence score calculations.
[0050] Furthermore, the step of dynamically generating speech activity detection thresholds based on environmental assessment results includes: S201. Based on the speech activity characteristics in the environmental assessment results, extract noise baseline parameters from the non-speech region of the speech signal. The noise baseline parameters include at least the short-time energy mean and the short-time energy variance. S202. Determine the threshold adjustment weight factor based on the status information, and perform a mapping operation between the threshold adjustment weight factor and the noise baseline parameter to generate the corresponding first detection threshold. S203. Perform an exponential moving average on the first detection threshold to obtain a smoothed threshold. S204. Generate uplink and downlink thresholds by smoothing the threshold, apply threshold boundary constraints to the uplink and downlink thresholds, and limit the numerical range between the uplink and downlink thresholds to between the preset minimum and maximum threshold values to generate the corresponding speech activity detection thresholds.
[0051] In this embodiment, the system selects the non-speech region of the speech signal from the environmental assessment results and extracts the short-time energy mean and short-time energy variance of this region as noise baseline parameters to reflect the overall level and fluctuation of background noise. Subsequently, the network state score in the state information is input to the threshold adjustment module to generate a threshold adjustment weight factor. This factor is used to amplify or reduce the influence of the noise baseline parameters on the threshold. For example, when the network state score is low, the weight factor increases, making the threshold more conservative and improving the reliability of valid speech; while when the network state score is high, the weight factor decreases, making the threshold more sensitive to capture small speech segments. Next, the system performs a mapping operation between the noise baseline parameters and the threshold adjustment weight factor to obtain a first detection threshold, and then smooths the first detection threshold using an exponential moving average method to reduce misjudgments caused by instantaneous fluctuations. Further, the system generates an uplink threshold and a downlink threshold based on the smoothed threshold. The uplink threshold is used to determine the start point of the speech, and the downlink threshold is used to determine the end point of the speech. Finally, boundary constraints are applied to the generated uplink and downlink thresholds, limiting them to a preset minimum and maximum value, thus forming the final speech activity detection threshold. This system can dynamically generate speech activity detection thresholds in complex noise and network environments, ensuring the stability of the thresholds under static noise conditions while also possessing the ability to adjust in real time according to network status and noise levels, thereby effectively improving the accuracy of distinguishing between speech segments and non-speech segments.
[0052] Furthermore, the step of invoking network state information to perform discontinuous transmission processing includes: S301. Determine the corresponding transmission rate according to the preset rate mapping relationship based on the status information; S302. Perform the corresponding discontinuous transmission processing according to the transmission rate.
[0053] In this embodiment, discontinuous transmission processing is used to reduce the audio data packet transmission rate. When the voice activity detection threshold determines that the user is not in a voice state, the system calls the network status information and enters the transmission control module. First, the transmission control module matches the network status score with a preset rate mapping table: when the score is in the high range, full transmission is maintained; when the score is in the middle range, the transmission rate is compressed by a certain percentage; and when the score is below the set lower limit, discontinuous transmission mode is activated, sending control signaling or voice header data only when necessary. Subsequently, the system controls the voice signal encoding module, adjusting the quantization level and frame packing strategy so that only necessary audio feature information is retained during compressed transmission, such as maintaining the energy envelope of the voice while omitting detail frequency bands. Furthermore, in discontinuous transmission mode, the system can also use the voice activity detection result as a trigger condition, resuming transmission only when a new voice segment is detected, avoiding continuous bandwidth occupation during long periods of idle state. This achieves discontinuous transmission processing based on network status adaptation, effectively reducing bandwidth consumption and device power consumption caused by invalid data transmission while ensuring communication continuity. It is particularly suitable for scenarios with frequent network condition fluctuations or limited resources, thereby improving the overall system operating efficiency.
[0054] Furthermore, the step of invoking speech activity features from the environmental assessment results and performing comfort noise supplementation processing includes: S401. When speech activity features indicate the presence of speech gaps in the speech signal, determine the gap length and background noise level of the speech gaps. The gap length is calculated from the start-end time difference between adjacent speech segments, and the background noise level is calculated by extracting the short-time energy mean and variance in the non-speech interval, which is used to characterize the intensity and stability of environmental noise.
[0055] S402. Based on the gap length and background noise level, select a corresponding noise sample from a preset noise library or user-defined configuration, and adjust the amplitude and trim the duration of the noise sample to generate a processed sample. The preset noise library may include faint breathing sounds, keyboard sounds, wind sounds, or low-frequency ambient sounds, while user-defined configuration may include custom background sound effects. After selection, the system adjusts the amplitude of the noise sample to match the loudness level of the current background noise, and trims or splices the duration of the noise sample to match its length with the speech gap, thereby generating the processed sample.
[0056] S403. The processed sample is identified as comfort noise and inserted into the speech gaps to complete the comfort noise supplementation process. The system identifies the processed sample as comfort noise and inserts it into the speech gaps, seamlessly splicing it with the original speech signal to form a continuous audio output stream. In this way, the speech gaps are filled with background sound that conforms to natural communication habits, effectively avoiding the embarrassment or discomfort caused by prolonged silence.
[0057] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A method of intelligent audio optimization, the method comprising: The intelligent audio optimization method comprises the following steps: Collecting voice signals and network state parameters in a current voice communication process, and inputting the voice signals and the network state parameters into a deep learning model for joint analysis to generate corresponding environment evaluation results; Based on the environment evaluation results, a voice activity detection threshold is dynamically generated, and the voice activity detection threshold is used to detect the voice signals to determine whether the user is in a voice state; If the user is not in a voice state, the network state information is called to perform discontinuous transmission processing; If the user is in a voice state and there is a voice gap in the voice signals, the voice activity feature in the environment evaluation results is called to perform comfort noise supplement processing, which is used to generate simulated background sound from a preset noise library or user individual configuration to insert into the voice signals; Output the target communication audio stream generated after the discontinuous transmission processing or the comfort noise supplement processing.
2. The intelligent audio optimization method of claim 1, wherein, In the step of inputting the voice signals and the network state parameters into a deep learning model for joint analysis to generate corresponding environment evaluation results, the deep learning model at least comprises an analysis sub-model, an extraction sub-model and a correction sub-model, and the step comprises: Based on the analysis sub-model, the network state parameters are analyzed to generate corresponding state information and feature analysis strategies; Based on the extraction sub-model and the feature analysis strategies, the voice signals are feature-extracted to generate corresponding voice activity features; The voice activity features and the state information are integrated to generate corresponding environment evaluation results.
3. The intelligent audio optimization method of claim 2, wherein, In the step of analyzing the network state parameters based on the analysis sub-model to generate corresponding state information and feature analysis strategies, it comprises: Based on the analysis sub-model, the network state parameters are analyzed to obtain corresponding state indicators, including bandwidth occupation indicators, delay fluctuation indicators and packet loss rate indicators; Each state indicator is weighted and averaged to generate corresponding state information, and a network state score is determined based on a scoring mapping relationship of the state information; Determine whether the network state score is greater than a preset state threshold to generate a corresponding judgment result, and determine a corresponding feature analysis strategy according to the judgment result, the feature analysis strategy at least comprising a network good analysis strategy and a network abnormal analysis strategy.
4. The intelligent audio optimization method of claim 3, wherein, In the step of extracting the voice signals based on the extraction sub-model and the feature analysis strategies to generate corresponding voice activity features, it comprises: If the feature analysis strategy is a network good analysis strategy, the voice signals are directly feature-extracted based on the extraction sub-model to generate corresponding voice activity features; If the feature analysis strategy is a network abnormal analysis strategy, the voice signals are divided into confidence levels to filter out target confidence voice segments, and the target confidence voice segments are feature-extracted to generate corresponding voice activity features.
5. The intelligent audio optimization method of claim 4, wherein, The step of performing confidence division on the voice signal to screen out a target confidence voice segment comprises: According to the network anomaly analysis strategy, determine the shortening amplitude of the preset time window, and shorten the preset time window based on the shortening amplitude to generate a corresponding target time window; The voice signal is divided according to the target time window to generate a plurality of intercepted segments; Each of the confidence score of the intercepted segment is calculated to calculate the confidence score of each of the intercepted segment; Each of the confidence score is compared with the preset confidence threshold value, and the intercepted segment with the confidence score greater than or equal to the confidence threshold value is screened out as the target confidence voice segment.
6. The intelligent audio optimization method of claim 5, wherein, The step of determining the shortening amplitude of the preset time window according to the network anomaly analysis strategy comprises: According to the network anomaly analysis strategy, the amplitude maximum value is determined according to the amplitude mapping relationship of the pre-stored computing resource limit; According to the score threshold difference contained in the judgment result, the shortening amplitude not greater than the amplitude maximum value is determined by mapping.
7. The intelligent audio optimization method of claim 5, wherein, The step of calculating the confidence score of each of the intercepted segment to calculate the confidence score of each of the intercepted segment comprises: The short-time energy and zero-crossing rate of each of the intercepted segment are calculated to generate corresponding time domain characteristic values; The frequency spectrum entropy and the main frequency distribution range of each of the intercepted segment are calculated to generate corresponding frequency domain characteristic values; The time domain characteristic values and the frequency domain characteristic values are weighted and summed according to the preset weight to generate the confidence score of the intercepted segment.
8. The intelligent audio optimization method of claim 2, wherein, The step of dynamically generating a voice activity detection threshold value based on the environment evaluation result comprises: Based on the voice activity feature in the environment evaluation result, the noise baseline parameter is extracted from the non-voice interval of the voice signal, and the noise baseline parameter at least includes the short-time energy mean and the short-time energy variance; According to the state information, a threshold adjustment weight factor is determined, and the threshold adjustment weight factor is mapped with the noise baseline parameter to generate a corresponding first detection threshold value; The first detection threshold value is subjected to exponential moving average processing to obtain a smoothed threshold value; The smoothed threshold value generates an uplink threshold value and a downlink threshold value, and the uplink threshold value and the downlink threshold value are subjected to threshold boundary constraint, so that the numerical interval between the uplink threshold value and the downlink threshold value is limited between the preset threshold minimum value and the preset threshold maximum value, to generate a corresponding voice activity detection threshold value.
9. The intelligent audio optimization method of claim 2, wherein, The step of calling the network state information to perform discontinuous transmission processing comprises: According to the state information, a corresponding transmission rate is determined according to a preset rate mapping relationship; According to the transmission rate, corresponding discontinuous transmission processing is performed.
10. The intelligent audio optimization method of claim 1, wherein, The step of calling the voice activity feature in the environment evaluation result to perform comfort noise supplement processing comprises: When the voice activity feature indicates that there is a voice gap in the voice signal, the gap length of the voice gap and the background noise level are determined; Based on the gap length and the background noise level, a corresponding noise sample is selected from a preset noise library or a user personalized configuration, and the noise sample is amplitude adjusted and time length cropped to generate a corresponding processed sample; The processed sample is determined as the comfort noise and inserted into the speech gap to complete the comfort noise supplement processing.