High-performance voice processing method

By generating scene fingerprint information based on scene features and dynamically selecting acoustic compensation or semantic processing paths, the problem of interruption and breakage caused by lost voice data packets is solved, achieving efficient voice recovery and transmission, and improving the stability and efficiency of voice processing.

CN121545548APending Publication Date: 2026-02-17SHENZHEN SOUNDFIT TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511684924.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-17
Publication Date
2026-02-17

AI Technical Summary

Technical Problem

Existing voice processing technologies suffer from voice data packet loss due to unstable network links, bandwidth fluctuations, and sudden interference, resulting in audio playback interruptions, voice breaks, and content loss. They also exhibit low transmission efficiency and poor recovery performance.

Method used

By acquiring voice data and device information, scene features are determined to generate scene fingerprint information, acoustic compensation or semantic processing paths are dynamically selected, and processing strategies are adjusted according to the duration of packet loss. A lightweight multi-scene adaptation framework is adopted to achieve parallel or sequential execution of acoustic compensation and semantic processing, and a smooth transition is achieved by combining boundary alignment markers.

Benefits of technology

Ensuring voice clarity and continuity in complex communication environments improves the quality and stability of voice transmission, solves the problems of smooth recovery under short-term packet loss and semantic integrity under long-term packet loss, avoids voice breaks and abruptness, and improves transmission efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121545548A_ABST
    Figure CN121545548A_ABST
Patent Text Reader

Abstract

The invention relates to a high-performance voice processing method, and the method comprises the steps: obtaining voice data transmitted through a network, generating corresponding scene fingerprint information, determining a frame parameter of a lightweight multi-scene adaptation frame, and a corresponding processing parameter, and if the duration of a lost segment is smaller than a preset threshold value, determining that the segment is lost. If yes, generating a corresponding first compensation result according to the acoustic compensation parameter; if the duration is larger than or equal to a preset threshold value, whether the acoustic compensation process and the semantic processing process are performed in parallel or not is judged according to the frame parameters, if not, a corresponding second compensation result is generated, and if yes, a corresponding second compensation result is generated according to the acoustic compensation process and the semantic processing process; and outputting the corresponding target voice transmission data. The definition and the stability of the voice signal in a complex environment can be effectively improved, different compensation modes are flexibly switched or executed in parallel according to the packet loss duration and the resource state, and the effectiveness and the stability of a recovery mechanism are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of speech processing technology, and in particular to a high-performance speech processing method. Background Technology

[0002] Currently, voice processing technology is widely used in various communication devices and smart terminals, especially in wireless audio products and real-time voice call scenarios. However, due to factors such as unstable network links, bandwidth fluctuations, and sudden interference, voice data packets are often lost, directly leading to audio playback interruptions, voice breaks, or even content loss, thus seriously affecting the user's auditory experience. Existing technologies commonly employ packet loss recovery strategies based on redundancy coding, which add a certain amount of redundant information before data transmission. When packet loss occurs, recovery is performed using the redundant data. However, the increase in redundant data leads to a heavier transmission burden and reduces overall transmission efficiency. In the event of a sudden large number of packet losses, the redundant information itself may also be lost, thus significantly reducing the recovery effect. Summary of the Invention

[0003] To address the problems of low transmission efficiency and poor recovery performance in the face of sudden large-scale packet loss and complex environments in existing voice packet loss recovery methods, this application provides a high-performance voice processing method.

[0004] A high-performance speech processing method, the high-performance speech processing method comprising: Acquire voice data transmitted over the network, as well as device information from the device source; The scene features are determined based on the voice data and the device information. The corresponding scene fingerprint information is generated based on the scene features. The framework parameters of the lightweight multi-scene adaptation framework and the corresponding processing parameters are determined based on the scene fingerprint information. The processing parameters include acoustic compensation parameters and semantic processing parameters. Determine whether there are any missing segments in the voice data; if so, determine the duration of the missing segments. If the duration is less than a preset threshold, then the corresponding acoustic compensation process is performed on the lost segment according to the acoustic compensation parameters to generate the corresponding first compensation result; If the duration is greater than or equal to the preset threshold, then it is determined whether the acoustic compensation process and the semantic processing process are performed in parallel according to the framework parameters. If they cannot be performed in parallel, then the corresponding semantic processing process is performed on the lost segment according to the semantic processing parameters to generate the corresponding second compensation result. The semantic processing is used to perform context-based semantic prediction and synthesis recovery on the lost segment. If they can be performed in parallel, then the corresponding second compensation result is generated according to the acoustic compensation process and the semantic processing process. The first compensation result or the second compensation result is concatenated with the unlost segments in the speech data to output the corresponding target speech transmission data.

[0005] By adopting the above technical solution, the system can dynamically select acoustic compensation or semantic processing paths by acquiring voice data transmitted over the network and device information from the device source, and combining the packet loss detection results. This enables the system to adaptively adjust the processing strategy for different packet loss durations, thereby ensuring both voice clarity and continuity under short-term packet loss and semantic integrity and coherence under long-term packet loss, and achieving high-quality voice transmission even in complex communication environments.

[0006] Preferably, the step of determining scene features based on the voice data and the device information, and generating corresponding scene fingerprint information based on the scene features, includes: Identify and analyze environmental noise features and user speech rate features in the speech data; Based on the device information, environmental noise characteristics, and user speech rate characteristics, multi-dimensional scene characteristics are determined; The scene features are transformed into a corresponding discrete level set, which includes device level, environmental noise level and user speech rate level. The scene fingerprint information is generated based on the discrete level set. The scene fingerprint information is a data sequence of the discrete level set filled with a preset blank template to sequentially arrange the device level, the environmental noise level and the user speech rate level.

[0007] By adopting the above technical solution, the system identifies and analyzes the environmental noise features and user speech rate features of the voice data, and combines them with device information to form multi-dimensional scene features. Then, the scene features are transformed into a discrete level set and scene fingerprint information is generated. This enables the system to characterize different communication environments and user characteristics with a structured fingerprint sequence, thereby providing a unified feature input for subsequent processing and improving scene adaptability.

[0008] Preferably, the step of determining the framework parameters of the lightweight multi-scene adaptation framework based on the scene fingerprint information includes: Based on the device level in the scenario fingerprint information, extract the corresponding resource consumption index and latency tolerance index; Determine whether the resource consumption index is greater than the preset resource threshold for parallel operation of the acoustic compensation process and the semantic processing process. If it is greater, generate the corresponding first parallel permission parameter. The first parallel permission parameter is used to determine whether the acoustic compensation process and the semantic processing process are performed in parallel. Determine whether the latency tolerance index is greater than the preset latency threshold due to the excessive execution time of the semantic processing flow. If it is greater, generate the corresponding second parallel allowance parameter. The second parallel allowance parameter is used to determine the execution order of the acoustic compensation flow and the semantic processing flow when they are executed in parallel. The first parallelism allowance parameter and the second parallelism allowance parameter are integrated to generate the corresponding framework parameters.

[0009] By adopting the above technical solution, resource occupancy indicators and latency tolerance indicators are extracted from the device level in the scene fingerprint information, and parameters for judging parallelism and execution order are generated respectively. These parameters are then integrated into the framework parameters, enabling the system to flexibly adjust the parallel mode of acoustic compensation and semantic processing in resource-constrained or latency-sensitive scenarios. This ensures stable operation under diverse devices and large-scale concurrency conditions, thereby improving system stability and real-time performance.

[0010] Preferably, the step of determining the corresponding processing parameters based on the scene fingerprint information includes: Based on the first mapping relationship of the environmental noise level in the scene fingerprint information, a filter intensity parameter applied to the acoustic compensation process and a semantic weight adjustment amplitude parameter applied to the semantic processing process are generated. Based on the second mapping relationship of the user's speech rate level in the scene fingerprint information, a reduction parameter for the truncated window applied in the acoustic compensation process and a expansion parameter for the context window applied in the semantic processing process are generated. The filter intensity parameter and the reduction amplitude parameter are integrated into acoustic compensation parameters, and the prediction weight adjustment amplitude parameter and the expansion amplitude parameter are integrated into semantic processing parameters to generate corresponding processing parameters.

[0011] By adopting the above technical solution, the environmental noise level is mapped to the filter intensity parameter, and the user's speech rate level is mapped to the reduction of the capture window and the expansion of the context window, and then integrated into acoustic compensation parameters and semantic processing parameters, respectively. This enables the system to finely adjust the compensation process for different scene characteristics, thereby improving signal clarity in noisy environments, avoiding the loss of speech information when the speech rate is fast, and enhancing the flexibility and adaptability of compensation.

[0012] Preferably, the step of performing a corresponding acoustic compensation process on the lost segment according to the acoustic compensation parameters to generate a corresponding first compensation result includes: Based on the reduction parameter, the size of the first window of the truncating window is determined, and the lost segment is truncated based on the truncating window to generate multiple reference speech segments; Filtering is performed on the reference speech segment corresponding to the filter strength parameter to generate the corresponding filtered segment; Predictive modeling is performed on the filtered segment to generate the corresponding compensated speech segment; A first compensation result is generated based on the first boundary alignment identifier corresponding to the compensated speech segment marker. The first boundary alignment identifier is used to align the boundaries of the compensated speech segment and the unlost segment to complete the splicing.

[0013] By adopting the above technical solution, the size of the truncation window is determined by the reduction amplitude parameter to generate a reference speech segment. Then, the reference segment is filtered according to the filtering intensity parameter and predictive modeling is carried out. Finally, the boundary alignment mark is marked in the compensated speech segment and spliced ​​with the unlost segment, so that the lost segment can be repaired with natural waveform continuity, thereby effectively eliminating speech breaks and abruptness, and achieving fast and smooth recovery under short-term packet loss.

[0014] Preferably, the step of predictive modeling of the filtered segment to generate the corresponding compensated speech segment includes: Continuity features are extracted based on multiple filtered segments, and the continuity features include energy distribution and spectral morphology; Based on the continuity features, a prediction relationship is established, and the correlation between the previous filtered segment and the next filtered segment is used to estimate the lost segment point by point, generating the corresponding compensated speech segment.

[0015] By adopting the above technical solution, continuous features such as energy distribution and spectral morphology are extracted from multiple filtered segments, and a prediction relationship is established by utilizing the correlation between the segments before and after. The lost segments are estimated point by point to generate compensation segments, so that the recovery process can maintain the consistency of energy and spectrum with the context, thereby avoiding unnatural waveform abrupt changes and achieving high-fidelity speech compensation under short-term packet loss.

[0016] Preferably, the step of generating the corresponding second compensation result based on the acoustic compensation process and the semantic processing process includes: The execution order of the acoustic compensation process and the semantic processing process is determined based on the second parallel permission parameter in the framework parameters. If the semantic processing flow precedes the acoustic compensation flow, then the corresponding second compensation result is directly generated based on the semantic processing flow. If the acoustic compensation process precedes the semantic processing process, a corresponding first compensation result is generated according to the acoustic compensation process, and the first compensation result is determined as a temporary second compensation result. When the semantic processing process generates a new second compensation result, the new second compensation result adaptively replaces the temporary second compensation result based on the semantic weight adjustment amplitude parameter.

[0017] By adopting the above technical solution, the execution order of acoustic compensation and semantic processing is determined based on the framework parameters, and direct semantic results or prior acoustic results are generated under different orders. Then, adaptive replacement is performed after the semantic results are generated, so that the system can take into account both immediacy and accuracy, continuously optimize the output quality while ensuring uninterrupted speech, and maintain stable speech transmission under variable resource conditions.

[0018] Preferably, the step of adaptively replacing the temporary second compensation result with the new second compensation result includes: Based on the semantic weight adjustment magnitude parameter, the transition duration of the replacement transition time window is determined; Based on the transition duration, the temporary second compensation result is subjected to amplitude attenuation processing at the beginning of the replacement transition time window to obtain the attenuation result; The attenuation result and the new second compensation result are weighted and superimposed at the beginning of the replacement transition time window to generate the initial fusion result; Based on the transition control parameters, the new second compensation result is subjected to amplitude enhancement processing at the end of the replacement transition time window to obtain the enhanced result; The enhancement result and the temporary second compensation result are weighted and superimposed at the end of the replacement transition time window to generate the final fusion result; Based on the initial fusion result and the final fusion result, a transition fusion waveform is generated point by point using linear interpolation to form a continuous and smooth replacement segment; The replacement fragment is determined as the updated second compensation result, thereby completing the adaptive replacement of the temporary second compensation result by the new second compensation result.

[0019] By adopting the above technical solution, the replacement transition time window is determined by adjusting the amplitude parameter based on semantic weight. Within the window, amplitude attenuation processing is performed on temporary results and amplitude enhancement processing is performed on new semantic results. Then, weighted superposition and interpolation are used to generate replacement segments with smooth transitions, so that the final output can avoid abrupt replacement traces. This ensures that the speech is continuous and natural during long-term packet loss repair, improving the smoothness of recovery and the user's listening experience.

[0020] Preferably, the step of performing a corresponding semantic processing flow on the lost segment according to the semantic processing parameters to generate a corresponding second compensation result includes: The size of the second window of the context window is determined based on the amplification parameter, and the speech segments before and after the lost segment are extracted according to the context window to form a context reference segment; Feature extraction is performed on the context reference fragment to generate a corresponding semantic feature sequence; A semantic prediction model is established based on the semantic feature sequence, and the semantic prediction model is used to predict and complete the text content of the missing segment to generate predicted text. The predicted text is input into the speech synthesis model to generate the corresponding semantic synthesized segment; The semantically synthesized fragment is marked with a corresponding second boundary alignment identifier, and a corresponding second compensation result is generated. The boundary alignment identifier is used to smoothly splice the semantically synthesized fragment with the fragment that has not been lost.

[0021] By adopting the above technical solution, the size of the context window is determined by using the expansion amplitude parameter, and semantic feature sequences are extracted from the context fragments based on the window. A semantic prediction model is then established to complete the lost content, and a semantic synthesis fragment is generated through the speech synthesis model. Finally, the fragments are spliced ​​into the unlost fragments using boundary alignment markers. This ensures that the system can maintain the semantic coherence and content integrity even when there is long-term packet loss, thereby maintaining the effectiveness of dialogue understanding even when there is a large area of ​​speech loss and improving the accuracy of semantic recovery.

[0022] In summary, this application includes at least one of the following beneficial technical effects: This application utilizes scene fingerprinting to perform targeted scheduling of the compensation process. For example, it increases the filtering intensity in high-noise environments and narrows the acoustic compensation truncation window while expanding the context range of semantic prediction when the speech rate is fast. This effectively improves the clarity and stability of the speech signal in complex environments, solving the problem of insufficient anti-interference capability in existing technologies. Simultaneously, after detecting packet loss, this application dynamically selects the compensation path based on the duration of the lost segment. For short-term packet loss, it employs waveform prediction and splicing driven by acoustic compensation parameters; for long-term packet loss, it employs context prediction and speech synthesis recovery driven by semantic processing parameters. When resources permit, it implements parallel processing of the two compensation methods through framework parameters, dynamically adjusting the compensation path based on packet loss duration and resource status. This invention employs a method that allows for the switching or parallel execution of different compensation methods, thereby ensuring smooth and natural continuity under short-term packet loss and semantic integrity and coherence under long-term packet loss. This addresses the problem of poor recovery performance in existing technologies when there is a sudden large amount of packet loss. Furthermore, this application sets boundary alignment markers in the results generated by acoustic compensation and semantic processing, and performs smooth transitions during splicing. Natural segment connection is achieved through energy transition and semantic boundary control, thereby avoiding breaks, pops, or abruptness and ensuring a continuous and smooth auditory experience. Compared with existing methods based on redundant coding, this application relies entirely on intelligent prediction and adaptive compensation to achieve packet loss recovery, significantly improving the effectiveness and stability of the recovery mechanism, thus solving the problem of low transmission efficiency in existing technologies. Attached Figure Description

[0023] Figure 1 This is a flowchart of a high-performance speech processing method according to an embodiment of this application. Detailed Implementation

[0024] The present application will be further described in detail below with reference to the accompanying drawings.

[0025] In one embodiment, such as Figure 1 As shown, this application discloses a high-performance speech processing method, which includes: S10. Acquire voice data transmitted over the network and device information from the device source; voice data refers to the raw audio signal data obtained through network transmission. This data is usually carried in the form of packets and may be interrupted during transmission due to bandwidth jitter or link loss; device information refers to the operating status data extracted from the voice acquisition end or playback end device, including hardware processing power, available memory, bandwidth utilization and power consumption status. This information can reflect the resource conditions of the current device when performing compensation or predictive processing. S20. Determine scene features based on voice data and device information, generate corresponding scene fingerprint information based on scene features, and determine the framework parameters and corresponding processing parameters of the lightweight multi-scene adaptation framework based on scene fingerprint information. The processing parameters include acoustic compensation parameters and semantic processing parameters. Scene features refer to the attributes of the communication and usage environment jointly characterized by voice data and device information, including noise level, user speech rate, channel stability, and device operating level. It is an intermediate layer feature. Scene fingerprint information refers to the standardized and discretized feature sequence generated based on scene features. It can uniquely characterize the combined state of the environment at a certain moment, similar to a feature fingerprint map generated for different scenes, used for subsequent parameter mapping and decision-making. The lightweight multi-scene adaptation framework is a software and hardware framework that dynamically configures processing logic under different communication scenarios. This framework emphasizes on-demand loading and releasing to avoid occupying too many system resources, thereby ensuring smooth operation even on devices with low latency requirements or limited resources. Framework parameters refer to the set of control parameters mapped from scene fingerprint information, which determine different compensation flows. The parallelism, execution order, and resource scheduling of the process are equivalent to the framework's control instructions at a certain moment. Processing parameters refer to the specific parameters obtained from scene feature mapping that drive the actual compensation process, including acoustic compensation parameters and semantic processing parameters. Acoustic compensation parameters are parameters used to control waveform-level processing in packet loss repair, such as filter strength, window size, and energy constraints. They determine how to utilize continuity information to generate smooth transition compensation segments in short-term packet loss. Semantic processing parameters are parameters used to drive semantic prediction and synthesis recovery, including semantic weights, context window length, and language model invocation strength. They determine how to utilize contextual information to complete missing content in long-term packet loss. The acoustic compensation process refers to the short-term packet loss repair process based on acoustic compensation parameters, typically including reference segment extraction, filtering, predictive modeling, and splicing; it is a repair method based on signal continuity. The semantic processing process refers to the long-term packet loss repair process executed based on semantic processing parameters, typically including contextual semantic modeling, missing content prediction, and speech synthesis; it is a repair method based on semantic integrity. S30. Determine if there are any missing segments in the voice data. If so, determine the duration of the missing segments. S40. If the duration is less than the preset threshold, the corresponding acoustic compensation process is executed on the lost segment according to the acoustic compensation parameters to generate the corresponding first compensation result. The first compensation result refers to the repaired segment generated after the acoustic compensation process is executed, which ensures the smoothness and naturalness of the speech waveform in the short-term packet loss scenario. S50. If the duration is greater than or equal to a preset threshold, determine whether to perform the acoustic compensation process and the semantic processing process in parallel based on the framework parameters. If they cannot be performed in parallel, execute the corresponding semantic processing process on the lost segments according to the semantic processing parameters to generate the corresponding second compensation result. Semantic processing is used to perform context-based semantic prediction and synthetic recovery of the lost segments. If they can be performed in parallel, generate the corresponding second compensation result according to the acoustic compensation process and the semantic processing process. The second compensation result refers to the repaired segment generated after the semantic processing process is executed alone or in parallel with the acoustic compensation process, ensuring the integrity and coherence of semantic content in long-term packet loss scenarios. S60. The first compensation result or the second compensation result is concatenated with the unlost segments in the speech data to output the corresponding target speech transmission data; the target speech transmission data refers to the complete output speech data after compensation processing and concatenation with the unlost segments, which is used to transmit to the end user to ensure a continuous and natural auditory experience.

[0026] Specifically, in an online meeting, a user was speaking via mobile phone. Due to network fluctuations, a short-term loss of approximately 50 milliseconds occurred in the audio data. At this point, the system, based on scene fingerprint information, determined that the user was on a low-end device in a moderately noisy environment. Therefore, it generated corresponding acoustic compensation parameters. In the acoustic compensation process, the continuity of adjacent segments was used to generate a smooth transition compensation result, forming the first compensation result, which was successfully spliced ​​into the output speech stream. The user could hardly perceive the packet loss. However, when the network deteriorated further, resulting in a long-term loss of 600 milliseconds, the system adjusted the framework parameters based on scene fingerprint information and switched to the semantic processing process. Based on the context, it predicted the missing content of "Hello everyone, let's start the meeting" and generated synthesized speech, obtaining the second compensation result. This second compensation result was then spliced ​​with the original speech through boundary alignment, ensuring that the output target speech transmission data remained coherent and natural, avoiding semantic breaks and auditory discomfort.

[0027] Furthermore, the step of determining scene features based on voice data and device information, and generating corresponding scene fingerprint information based on scene features, includes: S2011. Identify and analyze environmental noise features and user speech rate features in speech data; environmental noise features refer to the background noise attributes extracted from speech data. By analyzing the energy distribution, spectral density, and noise interference patterns of audio signals, the strength and type of noise in the current environment can be quantified, providing a basis for subsequent noise level classification; user speech rate features refer to the average speech rate and syllable pronunciation density of speech segments detected in speech data. They are usually obtained by statistically analyzing the number of phonemes or words per unit time and are used to characterize the speaker's speech rhythm and content density.

[0028] S2012. Based on device information, environmental noise characteristics, and user speech rate characteristics, determine multi-dimensional scene characteristics; environmental noise characteristics refer to the background noise attributes extracted from speech data. By analyzing the energy distribution, spectral density, and noise interference patterns of audio signals, the strength and type of noise in the current environment can be quantified, providing a basis for subsequent noise level classification; user speech rate characteristics refer to the average speech rate and syllable pronunciation density of speech segments detected in speech data. They are usually obtained by statistically analyzing the number of phonemes or words per unit time and are used to characterize the speaker's speech rhythm and content density.

[0029] S2013. Transform the scene features into a corresponding discrete level set. The discrete level set includes device level, environmental noise level, and user speech rate level. The discrete level set refers to the standardization of the above multi-dimensional scene features into a gradeable numerical result through a preset threshold. For example, the device level can be divided into three ranges: low, medium, and high; the environmental noise level can be divided into quiet, medium noise, and high noise; and the user speech rate level can be divided into slow, normal, and fast. This transforms continuous features into a finite set, which is convenient for subsequent processing by the system.

[0030] S014. Generate corresponding scene fingerprint information based on the discrete level set. The scene fingerprint information is a data sequence of device level, environmental noise level, and user speech rate level, which is filled according to a preset blank template. The discrete level set refers to the standardized numerical results of the above multi-dimensional scene features through preset thresholds. For example, the device level can be divided into three intervals: low, medium, and high; the environmental noise level can be divided into quiet, medium noise, and high noise; and the user speech rate level can be divided into slow, normal, and fast. This transforms continuous features into a finite set, which is convenient for subsequent processing by the system.

[0031] Furthermore, the step of determining the framework parameters of the lightweight multi-scenario adaptation framework based on scenario fingerprint information includes: S2021. Based on the device level in the scene fingerprint information, extract the corresponding resource consumption indicators and latency tolerance indicators. Extract these indicators using a pre-established device performance mapping table. This table is pre-set during system initialization based on the hardware configuration and historical performance of different types of devices. For example, low-end devices typically allow only low CPU and memory usage during voice processing, so their resource consumption indicators are limited to a low threshold, while their latency tolerance indicators are relatively high to ensure that even slightly slower processing speeds do not cause system crashes. Conversely, high-end devices have strong computing power, and their resource consumption indicators can be set to a high threshold, allowing complex algorithms to run simultaneously, but their latency tolerance indicators are lower, requiring near-real-time completion of voice data processing. During system runtime, the system reads the device level field from the scene fingerprint information and directly searches and extracts the corresponding resource consumption indicators and latency tolerance indicators from the mapping table, thus providing a basis for subsequently determining the parallelism and execution order of the compensation process.

[0032] S2022. Determine whether the resource usage index is greater than the preset resource threshold for parallel execution of the acoustic compensation process and the semantic processing process. If it is greater, generate the corresponding first parallel permission parameter. The first parallel permission parameter is used to determine whether to perform the acoustic compensation process and the semantic processing process in parallel. Parallel execution of the two processes requires simultaneous invocation of multiple operation modules such as filtering, prediction, and semantic modeling. If the device's computing resources are insufficient, it will lead to excessive processing latency or even stuttering, which will affect the real-time performance and continuity of voice transmission. Therefore, the system evaluates whether the device has the ability to execute in parallel by comparing the resource usage index with the preset resource threshold. When the index is greater than the threshold, it means that the device performance is sufficient to support parallel processing. At this time, the generated first parallel permission parameter will be used as a control signal to trigger the parallel execution of acoustic compensation and semantic processing, thereby improving the accuracy and robustness of packet loss recovery as much as possible while ensuring system stability.

[0033] S2023. Determine whether the latency tolerance index is greater than the preset latency threshold due to the excessive execution time of the semantic processing flow. If it is greater, generate the corresponding second parallel allowance parameter. The second parallel allowance parameter is used to determine the execution order of the acoustic compensation flow and the semantic processing flow when they are executed in parallel. The semantic processing flow usually involves context modeling and speech synthesis, which takes longer than the acoustic compensation flow. If the device or application scenario has a low tolerance for latency, it is not suitable to use the semantic processing result as the main output, otherwise it will cause voice transmission stuttering and poor interaction. Therefore, it is necessary to compare the latency tolerance index with the preset latency threshold to determine whether the system can accept the time overhead brought by semantic processing. When the latency tolerance index is greater than the threshold, it means that the system has a low sensitivity to latency and can give priority to the semantic processing result and use the acoustic compensation as a transition result. When the latency tolerance index is insufficient, it is necessary to give priority to quickly output the acoustic compensation result and then update it to the semantic processing result when conditions permit. The second parallel allowance parameter generated in this way can be used as a scheduling basis to reasonably arrange the order of the two types of processes when they are executed in parallel, so as to ensure the real-time performance of voice transmission and take into account the semantic integrity of the final output. S2024. Integrate the first parallel permission parameter and the second parallel permission parameter to generate the corresponding framework parameters. The first parallel permission parameter and the second parallel permission parameter are logically integrated as the parallel feasibility judgment result and the execution order control signal, respectively. For example, the final judgment result of whether parallel execution is possible is obtained through logical AND operation. When parallelism is allowed, the second parallel permission parameter determines the execution order of the acoustic compensation process and the semantic processing process, thereby generating the corresponding framework parameters.

[0034] Furthermore, the step of determining the corresponding processing parameters based on the scene fingerprint information includes: S2031. Based on the first mapping relationship of environmental noise levels in the scene fingerprint information, generate filter intensity parameters for the acoustic compensation process and semantic weight adjustment amplitude parameters for the semantic processing process. S2032. Based on the second mapping relationship of user speech rate level in scene fingerprint information, generate the reduction amplitude parameter for the interception window applied in the acoustic compensation process and the expansion amplitude parameter for the context window applied in the semantic processing process. S2033 integrates the filter intensity parameter and the reduction amplitude parameter as acoustic compensation parameters, and integrates the prediction weight adjustment amplitude parameter and the expansion amplitude parameter as semantic processing parameters to generate corresponding processing parameters.

[0035] In this embodiment, firstly, a first mapping relationship is established between environmental noise levels and processing parameters. This transforms different noise levels into executable parameter instructions. For example, a lower filter intensity parameter is generated at low noise levels to preserve the naturalness of the speech, while a higher filter intensity parameter is generated at high noise levels to enhance anti-interference capabilities. Simultaneously, a semantic weight adjustment amplitude parameter is obtained based on the same noise level mapping. When noise is high, the weight ratio of semantic prediction results in the overall recovery is increased, thus compensating for insufficient acoustic information. Secondly, a second mapping relationship is established between user speech rate levels and processing parameters. This transforms speech rate features into a dynamic adjustment strategy for the time window. When the speech rate is slow, a smaller reduction amplitude parameter is generated to maintain the normal size of the interception window in the acoustic compensation process, ensuring the stability of waveform prediction. Conversely, when the speech rate is fast, a larger reduction amplitude parameter is generated to shorten the time required for compensation. A reference window is used to quickly track continuous phonemes. Simultaneously, a parameter for expanding the context window is generated based on the speech rate level. This increases the coverage of contextual speech segments at faster speech rates, allowing semantic prediction to acquire more complete contextual information and avoiding semantic omissions or sentence incoherence. Finally, the filter intensity parameter mapped from the noise level and the reduction parameter mapped from the speech rate level are integrated to form acoustic compensation parameters, specifically used to drive the waveform-level compensation process. Simultaneously, the semantic weight adjustment parameter and the context window expansion parameter are integrated to form semantic processing parameters, used to drive the semantic prediction and synthesis recovery process. The working principle of the entire integration process is to quantify the features of environmental noise and user speech rate into controllable numerical parameters, and drive different levels of compensation through parameter set combinations, thereby maintaining the clarity and coherence of speech recovery in different scenarios.

[0036] Furthermore, the step of performing the corresponding acoustic compensation process on the lost segment according to the acoustic compensation parameters to generate the corresponding first compensation result includes: S401. Based on the reduction amplitude parameter, determine the size of the first window of the truncating window, and truncate the lost segment based on the truncating window to generate multiple reference speech segments. Using the start and end boundaries of the lost segment as reference points, divide several continuous sampling intervals before and after them, and dynamically adjust the original window length through the reduction amplitude parameter. For example, in scenarios with a fast speech rate, the reduction amplitude parameter will reduce the window length to ensure that the captured reference segments can more sensitively reflect the instantaneous speech features. During truncating, the speech data on both sides of the lost segment is usually framed with a fixed frame length of 10ms to 20ms and a fixed frame shift of 5ms to 10ms to generate multiple short reference speech segments. These segments retain the continuous features of speech energy distribution, pitch period, and spectral envelope.

[0037] S402. Based on the filter intensity parameters, perform filtering on the corresponding reference speech segment to generate the corresponding filtered segment. First, convert the reference speech segment to its frequency domain representation and extract its spectral components using a Fast Fourier Transform. Then, determine the bandwidth and attenuation characteristics of the filter based on the filter intensity parameters. For example, at low noise levels, use a filter with a wider bandwidth and slower attenuation to preserve the detailed features of the speech signal as much as possible, while at high noise levels, use a filter with a narrower bandwidth and stronger attenuation to suppress background noise and interference signals to the maximum extent. The filtering process typically employs adaptive filtering or weighted spectral subtraction methods, allowing the filter coefficients to be dynamically updated with noise characteristics, ensuring optimal noise suppression in different environments. The filtered segment effectively improves the signal-to-noise ratio while preserving the main energy distribution and spectral envelope of the speech, providing cleaner and more stable input data for subsequent prediction modeling, thereby improving the accuracy and naturalness of lost segment reconstruction.

[0038] S403. Perform prediction modeling on the filtered segments to generate corresponding compensated speech segments; first, convert the filtered segments into feature sequences.

[0039] S404. Based on the first boundary alignment identifier corresponding to the compensated speech segment marker, generate the corresponding first compensation result. The first boundary alignment identifier is used to align the compensated speech segment and the unlost segment to complete the splicing. First, based on the end / start boundaries of the compensated speech segment and the unlost segment, calculate the transition interval for splicing. The length of the transition interval is determined by the system's time delay budget and the energy continuity requirements of the current scenario, for example, 10–30ms. Then, extract alignment features in both the frequency and time domains of the two segments simultaneously, including frame index, sampling timestamp, short-time energy, zero crossover rate, main frequency / pitch period, and phase trajectory, and generate a boundary alignment identifier based on this. The boundary alignment identifier records the splicing start and end frame numbers, crossfade window type and length, gain gradient curve, target energy reference, phase correction method, and maximum allowable phase deviation in the form of structured metadata. According to the boundary alignment identifier, first perform normalized energy matching and DC component correction on the two segments, then unwrap and drift correct the boundary phase of the compensated segment in the STFT domain to make its phase trajectory continuous with that of the unlost segment within the overlap band, and then return to In the temporal domain, an overlap-addition strategy is performed using a symmetrical crossfade: a fade-out envelope from 1 to 0 is applied to the unlost segments, and a fade-in envelope from 0 to 1 is applied to the compensated segments. The envelope can use Hann / Hamming or piecewise linear curves, and the shape of the envelope is specified by the window type in the identifier. The envelope is then weighted and summed by sample points within the transition region. To avoid pops and splicing clicks, energy slope limits and peak limits are added at both ends of the transition region. If necessary, a small time shift adjustment is made within a micro-window of 1–2 ms to align the peak and valley positions of the pitch period. After the overlap is completed, a smoothing filter of one frame width is added to each of the splicing results outside the transition region. Finally, the spliced ​​continuous waveform and the corresponding boundary alignment identifier are output together as the first compensation result. The boundary alignment identifier is written to the frame-level metadata in the buffer along with the result for subsequent possible replacement or rollback strategy calls. At the same time, the entire process is guaranteed to meet the predetermined end-to-end latency and computational load constraints.

[0040] Furthermore, the step of predictive modeling of the filtered segment to generate the corresponding compensated speech segment includes: S4031. Extract continuity features based on multiple filtered segments. The continuity features include energy distribution and spectral morphology. S4032. Establish a prediction relationship based on continuity features, and use the correlation between the previous filtered segment and the next filtered segment to estimate the lost segment point by point, and generate the corresponding compensated speech segment.

[0041] Specifically, for each segment, short-time energy distribution, spectral morphology parameters, and fundamental frequency and harmonic structure are calculated. The continuity of the speech signal is characterized by the energy curves and spectral envelope variation trends between adjacent segments, thus obtaining a set of feature sequences that reflect the smoothness in the time and frequency domains. Subsequently, a prediction relationship is established based on these continuity features, typically using a combination of linear interpolation and nonlinear modeling. For example, a transition curve is first constructed on the energy envelope using linear interpolation, and then a prediction model based on time-series modeling, such as an autoregressive model or a neural network prediction module, is used to estimate the spectral morphology point by point, thereby generating a compensation waveform that is consistent with the preceding and following segments in terms of energy, frequency, and phase. During the prediction process, the preceding filtered segment provides constraint information on the starting boundary of the lost segment, and the following filtered segment provides constraint information on the ending boundary of the lost segment. The model fills in the missing parts in the middle point by point through the correlation between the two ends, ensuring that the compensation segment is continuous in the time domain and matches the original speech characteristics in the frequency domain, ultimately obtaining a natural and smooth compensated speech segment.

[0042] Furthermore, the step of generating the corresponding second compensation result based on the acoustic compensation process and the semantic processing process includes: S5011. Determine the execution order of the acoustic compensation process and the semantic processing process according to the second parallel allowance parameter in the framework parameters; schedule the execution order of the two processes using the second parallel allowance parameter in the framework parameters. This parameter is usually determined by the device latency tolerance and the real-time requirements of the task. When the parameter value indicates that latency is allowed, the system prioritizes scheduling the semantic processing process; otherwise, it prioritizes scheduling the acoustic compensation process to ensure that speech recovery meets both real-time requirements and integrity.

[0043] S5012. If the semantic processing flow precedes the acoustic compensation flow, the corresponding second compensation result is directly generated based on the semantic processing flow. When the semantic processing flow precedes the acoustic compensation flow, the system directly calls the semantic processing module to complete context prediction and speech synthesis, generating an audio segment containing the semantic inference result. This segment is marked as the second compensation result and output after generation to ensure the logical integrity and semantic coherence of the speech content. In this mode, the acoustic compensation flow can be used as an alternative path to further refine the compensation waveform when necessary.

[0044] S5013. If the acoustic compensation process precedes the semantic processing process, a corresponding first compensation result is generated based on the acoustic compensation process, and this first compensation result is designated as a temporary second compensation result. When the semantic processing process generates a new second compensation result, the new second compensation result adaptively replaces the temporary second compensation result based on the semantic weight adjustment amplitude parameter. When the acoustic compensation process executes before the semantic processing process, the system first generates a first compensation result based on filtering and prediction modeling, and marks it as a temporary second compensation result. This temporary result has strong temporal continuity and low latency characteristics, and can quickly fill gaps in the speech data, avoiding obvious blanks or breaks during playback. Subsequently, when the semantic processing process completes context analysis and semantic synthesis and generates a new second compensation result, the system calls the semantic weight adjustment amplitude parameter to compare and evaluate the new result and the temporary result. Specifically, this includes comparing the differences between the two in terms of energy envelope, spectral consistency, and contextual semantic matching degree. Under the control of the weight adjustment parameter, the newly generated semantic synthesis segment gradually replaces the temporary compensation result. Through weighted fusion and amplitude gradual control within the transition window, adaptive replacement is achieved, thereby avoiding the popping, jumping, or semantic abruptness problems caused by direct switching.

[0045] In summary, low-latency output is ensured through initial acoustic compensation, and the overall semantic naturalness and contextual integrity are improved through subsequent semantic processing. The second parallel allowance parameter serves as the core scheduling basis, dynamically coordinating the execution order and replacement timing of the two processes, enabling the system to balance real-time performance and recovery quality in different scenarios.

[0046] Furthermore, the step of adaptively replacing the temporary second compensation result with the new second compensation result includes: S50131. Based on the semantic weight adjustment amplitude parameter, determine the transition duration of the replacement transition time window. The semantic weight adjustment amplitude parameter is used as an indicator to measure the proportion of the newly generated semantic result in the overall output. When the semantic weight adjustment amplitude parameter is large, it indicates that the new result has high credibility in terms of semantic completeness and contextual consistency. In this case, the system will allocate a shorter transition duration, allowing the new result to quickly replace the temporary result to improve the overall speech quality. Conversely, when the semantic weight adjustment amplitude parameter is small, it indicates that the new result has some uncertainty or limited connection with the preceding and following segments. To avoid abruptness during the switching process, the system will extend the transition duration, achieving smooth replacement through a longer fade-in / fade-out window. The value of the transition duration is usually calculated by a preset function mapping relationship. For example, using the semantic weight adjustment amplitude parameter as input, it is mapped to a millisecond-level transition window duration through a linear or piecewise function, thereby ensuring that the replacement process adaptively adjusts in the time dimension, satisfying both speech fluency and semantic recovery accuracy.

[0047] S50132. Based on the transition duration, amplitude attenuation processing is performed on the temporary second compensation result at the beginning of the replacement transition time window to obtain the attenuation result. First, the range of the attenuation function is determined according to the transition duration, and then the waveform amplitude of the temporary result is reduced point by point using the attenuation curve. Commonly used attenuation functions can be linear decreasing functions, exponential attenuation functions, or Hanning window functions, etc. Among them, linear decreasing is suitable for scenarios with high latency requirements, which can quickly reduce the amplitude, while exponential or window function attenuation can maintain high smoothness during the energy reduction process and reduce the abruptness of the sound. During the execution, the system will gradually reduce the amplitude of the sampling points of the temporary result in the transition area according to the weight assigned by the attenuation function, from 100% at the beginning to close to 0% at the end of the window, so as to ensure that there will be no excessive amplitude conflict when it is superimposed with the newly generated second compensation result. Through this attenuation processing, the energy of the temporary result can be released naturally, providing sufficient dynamic space for the fade-in of the new result, thereby avoiding popping or phase jump phenomena during splicing.

[0048] S50133. The attenuation result and the new second compensation result are weighted and superimposed at the beginning of the replacement transition time window to generate the initial fusion result. A fade-in function is applied to the newly generated second compensation result according to the transition duration, causing its amplitude to gradually increase from near zero at the beginning of the window. Simultaneously, the temporary second compensation result has already had its amplitude gradually reduced by the attenuation function, forming a pair of complementary weight distributions within the same time interval. During superposition, the system sums the two waveforms point-by-point according to the weighting coefficients. The weight of the attenuation result gradually decreases from 1 to 0, while the weight of the new compensation result gradually increases from 0 to 1. The weight change process can use a symmetrical linear curve, or a cosine or Hanning window function to improve the smoothness of the transition. Through this weighted superposition, the generated initial fusion result maintains continuity in energy and phase, avoiding waveform abrupt changes or auditory jumps caused by direct replacement, thus achieving a smooth and natural transition effect.

[0049] S50134. Based on the transition control parameters, the new second compensation result undergoes amplitude enhancement processing at the end of the replacement transition time window to obtain the enhanced result. The effective range of the enhancement function is set according to the transition duration. The newly generated second compensation result maintains a normal or slightly lower amplitude at the beginning of the window and gradually increases to the target amplitude level at the end of the window to ensure consistency with the energy baseline of subsequent unlost segments. The enhancement function can be designed using a linear increasing curve, an exponential growth curve, or the rising portion of a smooth window function (such as a Hanning window or a Blackman window). The choice should comprehensively consider enhancement speed and auditory smoothness. For example, linear functions are suitable for low-latency scenarios and can quickly complete amplitude enhancement, while smooth window functions can avoid abrupt energy jumps during enhancement. During execution, the system multiplies the sampling points of the new compensation result point by point by the weighting factor of the enhancement function, ensuring that its amplitude is at a lower level at the beginning of the window and gradually recovers to an energy state consistent with the context segments at the end of the window, thus achieving a smooth transition when superimposed with temporary results or unlost segments. Through this enhancement process, the new compensation result can naturally take over the speech stream at the end of the transition zone, ensuring auditory continuity and overall stability at the splicing point.

[0050] S50135. The enhanced result and the temporary second compensation result are weighted and superimposed at the end of the replacement transition time window to generate the final fusion result. First, an increasing weight is applied to the enhanced new result, gradually bringing its weight close to 1 at the end of the window. Simultaneously, a decreasing weight is applied to the temporary result, gradually bringing its weight close to 0 at the end of the window. The superposition process is completed point-by-point at the sampling point level. The system typically uses a linear weight curve to ensure a smooth transition in weight ratios; a cosine window or hyperbolic tangent function can also be used to achieve a smoother weight switching effect. Through this complementary weight allocation, the two results can transition naturally at the end of the window, avoiding sudden energy changes or phase discontinuities, thus generating the final fusion result. The final fusion result not only maintains energy balance with the preceding and following segments but also achieves a smooth transition in spectral and phase characteristics, allowing the newly generated compensation segment to seamlessly replace the temporary result and continuously splice with the unlost segments, ultimately ensuring a natural and smooth overall audio output in subjective listening experience.

[0051] S50136. Based on the initial and final fusion results, a transition fusion waveform is generated point-by-point using linear interpolation to form a continuous and smooth replacement segment. The initial and final fusion results are used as the two endpoints of the transition window. Within the transition window, the system calculates the interpolation weight for each sampling point sequentially, according to the interpolation formula: y(i)=ystart(i)×(1−α)+yend(i)×α; Where y(i) represents the amplitude of the transition fusion waveform at the i-th sampling point, ystart(i) represents the amplitude of the initial fusion result at the i-th sampling point, yend(i) represents the amplitude of the final fusion result at the i-th sampling point, and α is the normalized position parameter, which takes values ​​in the range [0,1] and increases linearly with the sampling point index.

[0052] In this way, a smooth transition of the waveform in the time dimension can be ensured, while avoiding abrupt changes in amplitude and phase. If further optimization of smoothness is required, the system can also introduce weighted corrections during the interpolation process, such as using piecewise linear interpolation or piecewise cosine interpolation, to make the waveform more natural in terms of energy and spectral changes. The final generated transition fusion waveform covers the entire transition window, seamlessly splicing with the preceding and following segments as well as the new compensation segments, thereby ensuring that the output speech sounds continuous, smooth, and without abruptness.

[0053] S50137. The replacement fragment is determined as the updated second compensation result, thereby completing the adaptive replacement of the temporary second compensation result by the new second compensation result.

[0054] Furthermore, the step of performing the corresponding semantic processing flow on the lost segment according to the semantic processing parameters to generate the corresponding second compensation result includes: S5021. The size of the second window of the context window is determined based on the amplification amplitude parameter, and the speech segments before and after the lost segment are extracted according to the context window to form a context reference segment. First, the start and end boundaries of the lost segment are used as center points, and context sampling intervals of a certain duration are divided before and after them. The interval length is controlled by the amplification amplitude parameter. When the user's speech rate is fast or the semantic context dependency is strong, the amplification amplitude parameter will be increased, thereby extending the extraction range to ensure that more semantic information before and after is included. Conversely, when the speech rate is slow or the semantic dependency is weak, the amplification amplitude parameter will be decreased to avoid introducing noise with too much redundant information. During extraction, a fixed frame length of 20ms and an overlapping frame shift of 10ms are used to frame the speech signals on both sides of the lost segment, and the resulting frame sequences are spliced ​​to form a context reference segment. This reference segment completely preserves the speech energy distribution, prosodic features, and semantic continuity before and after the lost segment, thereby providing reliable context input for the subsequent semantic prediction model.

[0055] S5022. Feature extraction is performed on the context reference segment to generate a corresponding semantic feature sequence. First, the reference segment is divided into several short-time windows of fixed frame length, such as a 20ms frame length and a 10ms frame shift. Each frame is pre-emphasized and windowed to enhance high-frequency information and reduce boundary effects. Then, the spectrum of each frame is obtained through Fast Fourier Transform, and phonetic features such as Mel frequency cepstral coefficients, formant parameters, and energy envelopes are extracted based on the spectral energy distribution. These features can well characterize the acoustic structure of speech. At the same time, time-domain features such as short-time energy and zero-crossing rate, as well as prosodic parameters such as speech rate and pause positions, can be combined to reflect the semantic rhythm characteristics of the context. Finally, the above acoustic features and prosodic features are combined in chronological order to form a semantic feature sequence, ensuring that it contains both the continuity information of the speech before and after the lost segment and multi-dimensional input features that can be processed by the semantic prediction model, thus providing a data foundation for subsequent prediction and completion.

[0056] S5023. A semantic prediction model is established based on semantic feature sequences, and this model is used to predict and complete the text content of missing segments, generating predicted text. First, the extracted semantic feature sequences are input into a context modeling network, typically composed of a bidirectional long short-term memory network, a Transformer encoder, or an improved version thereof, to capture the long-term dependencies between speech features before and after the missing segment. During network training, the model undergoes supervised learning using numerous labeled speech-text pairs, enabling it to establish a mapping relationship between acoustic features and semantic units. During runtime, the model combines the feature sequences of contextual reference segments to predict the word sequences or sub-word units corresponding to the missing segments at the semantic level. At the output layer, a language model probability distribution is used to constrain the selection, prioritizing semantically reasonable and context-coherent word combinations. The final generated predicted text not only conforms to the contextual relationships in semantic logic but also maintains a natural connection with the preceding and following speech in syntactic structure, thus providing accurate semantic content input for subsequent speech synthesis steps.

[0057] S5024. The predicted text is input into the speech synthesis model to generate corresponding semantic synthesized segments. First, the predicted text is normalized and phoneme-converted, mapping the predicted word sequences to corresponding phonemes or articulatory units. Then, these phoneme sequences are input into the speech synthesis model, preferably using an end-to-end neural network-based synthesis framework, such as the Tacotron series or FastSpeech model, to generate corresponding acoustic feature parameters, such as Mel spectrograms. During this process, the model combines the prosodic features of the context reference segments to adjust the intonation, pauses, and energy distribution of the predicted speech, ensuring that the synthesized segment remains consistent with the preceding and following speech at the prosody level. Next, a vocoder module, such as WaveNet, HiFi-GAN, or LPCNet, converts the acoustic feature parameters into playable temporal waveforms, forming the final semantic synthesized segment. The entire process ensures that the synthesized speech not only has clear and natural sound quality but also closely connects with the context in terms of semantics and prosody, thus seamlessly filling in missing speech segments.

[0058] S5025. Mark the semantically synthesized segment with the corresponding second boundary alignment identifier, and generate the corresponding second compensation result. The boundary alignment identifier is used to smoothly splice the semantically synthesized segment with the unlost segment. First, the boundary position is determined at the junction of the semantically synthesized segment and the unlost segment. This boundary is determined by the start and end timestamps of the predicted text combined with the actual alignment features of the context reference segment. Then, to ensure splicing smoothness, a boundary alignment identifier is generated at this boundary position. This identifier includes a time index, transition interval length, energy matching factor, and phase correction parameters, which are used to guide subsequent splicing processing. When performing the marking, the system performs amplitude normalization and fundamental frequency adjustment on the semantically synthesized segment to ensure that its energy envelope is consistent with the end or start of the unlost segment, and corrects potential phase drift in the spectral domain through a phase compensation algorithm. Finally, based on the boundary alignment identifier, the semantically synthesized segment and the unlost segment are smoothly connected within the splicing interval using cross-fade-in / fade-out or linear interpolation to avoid popping sounds or discontinuities, ultimately obtaining the second compensation result with the boundary alignment identifier, achieving natural recovery of lost speech.

[0059] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A high-performance speech processing method, characterized in that, The high-performance speech processing method includes: Acquire voice data transmitted over the network, as well as device information from the device source; The scene features are determined based on the voice data and the device information. The corresponding scene fingerprint information is generated based on the scene features. The framework parameters of the lightweight multi-scene adaptation framework and the corresponding processing parameters are determined based on the scene fingerprint information. The processing parameters include acoustic compensation parameters and semantic processing parameters. Determine whether there are any missing segments in the voice data; if so, determine the duration of the missing segments. If the duration is less than a preset threshold, then the corresponding acoustic compensation process is performed on the lost segment according to the acoustic compensation parameters to generate the corresponding first compensation result; If the duration is greater than or equal to the preset threshold, then it is determined whether the acoustic compensation process and the semantic processing process are performed in parallel according to the framework parameters. If they cannot be performed in parallel, then the corresponding semantic processing process is performed on the lost segment according to the semantic processing parameters to generate the corresponding second compensation result. The semantic processing is used to perform context-based semantic prediction and synthesis recovery on the lost segment. If they can be performed in parallel, then the corresponding second compensation result is generated according to the acoustic compensation process and the semantic processing process. The first compensation result or the second compensation result is concatenated with the unlost segments in the speech data to output the corresponding target speech transmission data.

2. The high-performance speech processing method according to claim 1, characterized in that, The step of determining scene features based on the voice data and the device information, and generating corresponding scene fingerprint information based on the scene features, includes: Identify and analyze environmental noise features and user speech rate features in the speech data; Based on the device information, environmental noise characteristics, and user speech rate characteristics, multi-dimensional scene characteristics are determined; The scene features are transformed into a corresponding discrete level set, which includes device level, environmental noise level and user speech rate level. The scene fingerprint information is generated based on the discrete level set. The scene fingerprint information is a data sequence of the discrete level set filled with a preset blank template to sequentially arrange the device level, the environmental noise level and the user speech rate level.

3. The high-performance speech processing method according to claim 2, characterized in that, The step of determining the framework parameters of the lightweight multi-scene adaptation framework based on the scene fingerprint information includes: Based on the device level in the scenario fingerprint information, extract the corresponding resource consumption index and latency tolerance index; Determine whether the resource consumption index is greater than the preset resource threshold for parallel operation of the acoustic compensation process and the semantic processing process. If it is greater, generate the corresponding first parallel permission parameter. The first parallel permission parameter is used to determine whether the acoustic compensation process and the semantic processing process are performed in parallel. Determine whether the latency tolerance index is greater than the preset latency threshold due to the excessive execution time of the semantic processing flow. If it is greater, generate the corresponding second parallel allowance parameter. The second parallel allowance parameter is used to determine the execution order of the acoustic compensation flow and the semantic processing flow when they are executed in parallel. The first parallelism allowance parameter and the second parallelism allowance parameter are integrated to generate the corresponding framework parameters.

4. The high-performance speech processing method according to claim 2, characterized in that, The step of determining the corresponding processing parameters based on the scene fingerprint information includes: Based on the first mapping relationship of the environmental noise level in the scene fingerprint information, a filter intensity parameter applied to the acoustic compensation process and a semantic weight adjustment amplitude parameter applied to the semantic processing process are generated. Based on the second mapping relationship of the user's speech rate level in the scene fingerprint information, a reduction parameter for the truncated window applied in the acoustic compensation process and a expansion parameter for the context window applied in the semantic processing process are generated. The filter intensity parameter and the reduction amplitude parameter are integrated into acoustic compensation parameters, and the prediction weight adjustment amplitude parameter and the expansion amplitude parameter are integrated into semantic processing parameters to generate corresponding processing parameters.

5. The high-performance speech processing method according to claim 4, characterized in that, The step of performing a corresponding acoustic compensation process on the lost segment according to the acoustic compensation parameters to generate a corresponding first compensation result includes: Based on the reduction parameter, the size of the first window of the truncating window is determined, and the lost segment is truncated based on the truncating window to generate multiple reference speech segments; Filtering is performed on the reference speech segment corresponding to the filter strength parameter to generate the corresponding filtered segment; Predictive modeling is performed on the filtered segment to generate the corresponding compensated speech segment; A first compensation result is generated based on the first boundary alignment identifier corresponding to the compensated speech segment marker. The first boundary alignment identifier is used to align the boundaries of the compensated speech segment and the unlost segment to complete the splicing.

6. The high-performance speech processing method according to claim 5, characterized in that, The step of predicting and modeling the filtered segment to generate the corresponding compensated speech segment includes: Continuity features are extracted based on multiple filtered segments, and the continuity features include energy distribution and spectral morphology; Based on the continuity features, a prediction relationship is established, and the correlation between the previous filtered segment and the next filtered segment is used to estimate the lost segment point by point, generating the corresponding compensated speech segment.

7. The high-performance speech processing method according to claim 5, characterized in that, The step of generating the corresponding second compensation result based on the acoustic compensation process and the semantic processing process includes: The execution order of the acoustic compensation process and the semantic processing process is determined based on the second parallel permission parameter in the framework parameters. If the semantic processing flow precedes the acoustic compensation flow, then the corresponding second compensation result is directly generated based on the semantic processing flow. If the acoustic compensation process precedes the semantic processing process, a corresponding first compensation result is generated according to the acoustic compensation process, and the first compensation result is determined as a temporary second compensation result. When the semantic processing process generates a new second compensation result, the new second compensation result adaptively replaces the temporary second compensation result based on the semantic weight adjustment amplitude parameter.

8. The high-performance speech processing method according to claim 7, characterized in that, The step of adaptively replacing the temporary second compensation result with the new second compensation result includes: Based on the semantic weight adjustment magnitude parameter, the transition duration of the replacement transition time window is determined; Based on the transition duration, the temporary second compensation result is subjected to amplitude attenuation processing at the beginning of the replacement transition time window to obtain the attenuation result; The attenuation result and the new second compensation result are weighted and superimposed at the beginning of the replacement transition time window to generate the initial fusion result; Based on the transition control parameters, the new second compensation result is subjected to amplitude enhancement processing at the end of the replacement transition time window to obtain the enhanced result; The enhancement result and the temporary second compensation result are weighted and superimposed at the end of the replacement transition time window to generate the final fusion result; Based on the initial fusion result and the final fusion result, a transition fusion waveform is generated point by point using linear interpolation to form a continuous and smooth replacement segment; The replacement fragment is determined as the updated second compensation result, thereby completing the adaptive replacement of the temporary second compensation result by the new second compensation result.

9. A high-performance speech processing method according to claim 4, characterized in that, The step of performing a corresponding semantic processing flow on the lost segment according to the semantic processing parameters to generate a corresponding second compensation result includes: The size of the second window of the context window is determined based on the amplification parameter, and the speech segments before and after the lost segment are extracted according to the context window to form a context reference segment; Feature extraction is performed on the context reference fragment to generate a corresponding semantic feature sequence; A semantic prediction model is established based on the semantic feature sequence, and the semantic prediction model is used to predict and complete the text content of the missing segment to generate predicted text. The predicted text is input into the speech synthesis model to generate the corresponding semantic synthesized segment; The semantically synthesized fragment is marked with a corresponding second boundary alignment identifier, and a corresponding second compensation result is generated. The boundary alignment identifier is used to smoothly splice the semantically synthesized fragment with the fragment that has not been lost.

Citation Information

Cited By

  • Digital human interaction method and system for predicting semantics and preemption scheduling, electronic equipment and storage medium

    CN121938347A

  • Method and system for predicting semantics and preemptively scheduling digital human interaction, electronic device and storage medium

    CN121938347B