Dynamic calculation voice separation method based on human voice feature extraction and related equipment thereof
By acquiring speech data for feature extraction and complexity scoring, and dynamically adjusting the exit point of the speech separation network, the problems of high computational resource consumption and insufficient separation accuracy in existing technologies are solved, achieving efficient speech separation and semantic understanding in multi-speaker and complex environments.
Patent Information
- Application Number
- CN202511396824.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-28
- Publication Date
- 2025-12-05
AI Technical Summary
Existing dynamic computational speech separation methods based on human voice feature extraction suffer from problems such as high computational resource consumption, insufficient separation accuracy in complex speech environments, and lack of adaptability to device operating status.
By acquiring voice data, feature extraction is performed and a complexity score is generated. Based on the score, the exit point is determined, voice separation processing is performed, and sound quality optimization and semantic reasoning are carried out. The exit point of the voice separation network is dynamically adjusted in combination with the device's operating status, and metadata is used for differentiated optimization and semantic reasoning.
It reduces computing resource consumption, improves speech separation accuracy and semantic understanding in multi-speaker and complex environments, and achieves efficient speech separation and understanding on low-computing-power devices.
Smart Images

Figure CN121075355A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of intelligent speech recognition, and in particular to a dynamic computing speech separation method and device based on human voice feature extraction, an electronic device and a storage medium thereof. BACKGROUND
[0002] With the development of artificial intelligence technology, speech interaction has gradually become an important application scenario of intelligent devices. In Internet of Things terminals, smart homes and vehicle systems, devices need to extract and understand target speech in complex acoustic environments, thereby supporting voice control, dialogue interaction and speech recognition. However, in real environments, there are often problems such as multiple speaker overlap, background noise interference and limited device computing power, which seriously affect the speech separation and recognition effect.
[0003] Most existing speech separation methods use neural network models of fixed depth to perform full inference on input speech data. Although such methods can achieve a certain degree of separation effect, they have two defects: on the one hand, in low complexity speech scenarios, a complete calculation process must be performed, resulting in waste of computing power and energy; on the other hand, in multi-speaker or strong noise environments, traditional methods lack a dynamic adjustment mechanism, and are prone to problems such as incomplete separation or loss of semantic information. In addition, existing methods cannot combine with device running states, so that real-time performance and accuracy cannot be considered in a lightweight hardware environment.
[0004] Therefore, the existing dynamic computing speech separation method based on human voice feature extraction has the problems of large consumption of computing resources, insufficient separation accuracy in complex speech environments, and lack of adaptation to device running states. SUMMARY
[0005] The present application provides a dynamic computing speech separation method based on human voice feature extraction to solve the problems of large consumption of computing resources, insufficient separation accuracy in complex speech environments, and lack of adaptation to device running states in the existing dynamic computing speech separation method based on human voice feature extraction.
[0006] In a first aspect, the present application provides a dynamic computing speech separation method based on human voice feature extraction, comprising the following steps: obtaining speech data; performing feature extraction on the speech data, and generating a complexity score according to the extracted feature data; based on the complexity score, determining an exit point, and performing speech separation processing at the exit point to obtain separated speech data and corresponding metadata; performing voice quality optimization and semantic reasoning processing based on the separated speech data and the corresponding metadata to obtain target speech data.
[0007] Optionally, the obtaining voice data comprises: acquiring an environmental voice signal through a preset microphone array; preprocessing the environmental voice signal to obtain voice data.
[0008] Optionally, the voice data is subjected to feature extraction, and a complexity score is generated according to the extracted feature data, comprising: extracting and processing the voice data through a preset hybrid network to obtain Mel spectrum features of each frame of audio data, thereby obtaining first audio feature data; modeling and processing inter-frame time sequence dependence of the first audio feature data based on an LSTM model to obtain second audio feature data; determining a voice overlap index and background sound effect complexity data in the second audio feature data; generating a complexity score based on the voice overlap index and the background sound effect complexity data.
[0009] Optionally, the complexity score is used to determine an exit point, comprising: obtaining a running load parameter and a temperature parameter of a current device; determining a target running state of the current device based on the complexity score, the running load parameter, and the temperature parameter; in a preset strategy table, determining a corresponding exit point in real time according to the target running state of the current device.
[0010] Optionally, the separated data comprises target voice data and non-target voice data, and the voice separation processing is performed at the exit point to obtain separated voice data and corresponding metadata, comprising: calling a corresponding voice separation network based on the exit point; performing hierarchical separation processing on the voice data through the corresponding voice separation network to obtain target voice data and non-target voice data; performing asymmetric encoding processing on the target voice data and the non-target voice data to obtain corresponding target encoding results and non-target encoding results, respectively; generating corresponding metadata based on the exit point, the target encoding results, and the non-target encoding results.
[0011] Optionally, the separated voice data and the corresponding metadata are subjected to voice quality optimization and semantic reasoning processing to obtain target voice data, comprising: analyzing exit point information and encoding parameters in the metadata, and performing voice quality optimization processing on the separated voice data to obtain optimized voice data; Performing semantic reasoning processing on the optimized voice data, and outputting the target voice data.
[0012] Optionally, after the quality optimization and semantic reasoning processing on the separated voice data and the corresponding metadata are performed to obtain the target voice data, the method further comprises: Compressing the target voice data by using a first encoding mode and compressing the non-target voice data by using a second encoding mode to obtain a differentiated compression result; Forming a transmission data packet based on the compression result and the metadata; Adaptively switching between short-distance communication and long-distance communication according to a current network signal strength, so that the transmission data packet is sent to the cloud in real time through short-distance communication or long-distance communication.
[0013] In a second aspect, the present application further provides a dynamic computing voice separation device based on human voice feature extraction, which comprises: A first acquisition module for acquiring voice data; A first extraction module for performing feature extraction on the voice data and generating a complexity score according to the extracted feature data; A first determination module for determining an exit point based on the complexity score and performing voice separation processing at the exit point to obtain separated voice data and corresponding metadata; A first processing module for performing quality optimization and semantic reasoning processing on the separated voice data and the corresponding metadata to obtain target voice data.
[0014] In a third aspect, the present application provides an electronic device, which comprises a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor implements the steps of the dynamic computing voice separation method based on human voice feature extraction provided by the present application when executing the computer program.
[0015] In a fourth aspect, the present application provides a computer readable storage medium, which stores a computer program, wherein the computer program is executed by a processor to implement the steps of the dynamic computing voice separation method based on human voice feature extraction provided by the present application.
[0016] This invention acquires speech data; extracts features from the speech data and generates a complexity score based on the extracted features; determines an exit point based on the complexity score, and performs speech separation processing at the exit point to obtain separated speech data and corresponding metadata; and performs sound quality optimization and semantic reasoning processing based on the separated speech data and corresponding metadata to obtain target speech data. By combining the complexity score with the device's operating status, the exit point of the speech separation network is dynamically determined, reducing computational resource consumption. Furthermore, metadata is used to perform differentiated optimization and semantic reasoning on the separation results, improving the accuracy of speech separation and semantic understanding in multi-speaker and complex environments. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a flowchart of a dynamic computational speech separation method based on human voice feature extraction provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of another dynamic computational speech separation device based on human voice feature extraction provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0020] like Figure 1 As shown, Figure 1 This is a flowchart of a dynamic computational speech separation method based on human voice feature extraction provided by an embodiment of the present invention. The dynamic computational speech separation method based on human voice feature extraction includes the following steps: 101. Obtain voice data.
[0021] In the embodiment of the present application, the dynamic calculation speech separation method based on human voice feature extraction can be applied to a dynamic calculation speech separation platform based on human voice feature extraction. The dynamic calculation speech separation platform based on human voice feature extraction has functions such as speech separation data processing, speech separation data transceiving, and speech separation data memory storage, and can be constructed based on a server or a server cluster. The server or the server cluster can be an electronic device with speech separation data processing capability.
[0022] The speech data can be a sound signal collected by a preset microphone array and preprocessed through echo cancellation (AEC), noise suppression (NS), automatic gain control (AGC), etc. The speech data can be a time-domain waveform sequence or a spectrogram converted into a frequency domain. Specifically, the speech data can be collected at a sampling rate of 16 kHz / 16 bit to ensure the input quality of subsequent feature extraction and separation processing.
[0023] 102. Feature extraction is performed on the speech data, and a complexity score is generated according to the extracted feature data.
[0024] In the embodiment of the present application, feature extraction can be performed on the speech data by converting the original speech data into a feature representation that can be used for calculation and modeling. Specifically, Mel spectrum calculation can be performed on each frame of audio (such as 25 ms frame length and 10 ms frame shift) to obtain time-frequency features, and then local pattern features can be extracted through a convolutional neural network (CNN), while a long short-term memory network (LSTM) is used to model the cross-frame temporal dependence to preserve the temporal dynamic information of the speech.
[0025] The feature data can be an intermediate calculation result obtained through the feature extraction process, including first audio feature data and second audio feature data. The first audio feature data mainly refers to the frame-level Mel spectrum features extracted by the CNN, which are used to capture the instantaneous spectral distribution; the second audio feature data is the temporal features modeled by the LSTM, which reflect the continuity and dependence of the speech in the time dimension. It can be understood that the feature data collectively provides the basis for the complexity score and subsequent separation.
[0026] In one possible embodiment, the dynamic calculation speech separation platform based on human voice feature extraction generates new data or indicators through algorithm operation, thereby calculating the corresponding complexity score. Specifically, the speech overlap index and the background sound effect complexity can be calculated according to the first audio feature data and the second audio feature data, and then they are combined to form the complexity score. For example, the O value is generated by detecting the time proportion of multiple speaker overlaps in the speech frame, and then the C value is generated by calculating the spectral entropy and the signal-to-noise ratio, and finally the complexity score S is synthesized.
[0027] The complexity score can be a quantitative indicator for measuring the difficulty of separating the current speech scene, which can generally be mapped to a level score between 1 and 5. The higher the score, the more complex the speech and the more difficult it is to separate, and a deeper model needs to be used for separation processing. It can be understood that the complexity score can be calculated in combination with the speech overlap index and the background sound effect complexity as an input for dynamic calculation scaling to determine whether a deep network needs to be run.
[0028] 103、Based on the complexity score, determine the exit point and perform speech separation processing at the exit point to obtain separated speech data and corresponding metadata.
[0029] In the embodiments of the present application, the exit point can be a position in the hierarchical speech separation network that allows early termination of calculation, and the selection of the exit point can be explained by Table 1 as follows:
[0030] The separation network can be divided into four layers (L1-L4). In different complexity scores and device states, the dynamic calculation speech separation platform based on human voice feature extraction dynamically selects the corresponding exit point, runs only to this layer and outputs the result, thereby reducing unnecessary calculation and reducing delay and energy consumption.
[0031] In the embodiments, after determining the exit point, the dynamic calculation speech separation platform based on human voice feature extraction calls the corresponding hierarchical separation network model to perform mask estimation or demixing operation on the input speech data, separates the target speech signal from the mixed speech, and simultaneously generates non-target speech signal. Specifically, the dynamic calculation speech separation platform based on human voice feature extraction stops calculation at the network depth corresponding to the exit point, and outputs the target / non-target speech as the separation result.
[0032] The separated speech data can be two-channel speech results obtained by speech separation processing, including target speech data and non-target speech data. The target speech data corresponds to the user's pronunciation or the main speaker's voice that needs to be extracted, and the non-target speech data includes but is not limited to environmental noise, other speaker's speech or background sound, and abnormal noise.
[0033] The metadata can be a special composite package that contains not only the separated speech data itself (usually the encoded target and non-target speech), but also descriptive information related to these speeches, such as exit point level, complexity score, encoding method, signal-to-noise ratio, timestamp, and channel number. Metadata is designed as a unified data carrier to guide cloud-based sound quality optimization and semantic reasoning.
[0034] 104, based on the separated speech data and the corresponding metadata, the speech quality optimization and semantic reasoning processing are performed to obtain target speech data.
[0035] In the embodiment of the present application, the separated speech data can be enhanced and repaired based on the parameter information contained in the metadata. Specifically, the speech defects caused by the shallow exit are repaired through the generative adversarial network (GAN), and the noise residue is removed by using the adaptive spectral subtraction method, so as to improve the intelligibility and naturalness of the target speech, and make it more suitable for subsequent semantic reasoning processing.
[0036] In another possible embodiment, after the above-mentioned dynamic calculation speech separation platform based on human voice feature extraction obtains the optimized speech data, the speech recognition and natural language understanding technology are used to analyze the speech content and extract the semantic information therein. For example, after performing automatic speech recognition (ASR) on the story type speech, logical completion is performed, and for the dialogue type speech, intent recognition and slot filling are performed, so as to generate a higher level semantic understanding result.
[0037] The above-mentioned target speech data can refer to the final output, that is, the target speaker speech obtained after separation, encoding, speech quality optimization and semantic reasoning, which has the characteristics of high intelligibility, complete semantics and can be used for recognition and interaction. The result can be returned to the user as an audio signal, or can be called by downstream applications as a structured semantic result.
[0038] In a possible embodiment, the above-mentioned dynamic calculation speech separation platform based on human voice feature extraction collects environmental speech signals through a microphone array at a preset sampling rate, performs echo cancellation, noise suppression, automatic gain control and other preprocessing operations to obtain speech data; Mel spectrum feature extraction is performed on the speech data, long short-term memory network modeling is used to model inter-frame time sequence dependence, speech overlap index and background sound effect complexity are generated, and complexity score is formed; the above-mentioned dynamic calculation speech separation platform based on human voice feature extraction determines the exit point in the preset strategy table in combination with the complexity score, the device running load and the temperature parameter; the exit point corresponding layered speech separation network is called, and the target speech and non-target speech are obtained by running to the specified depth; the above-mentioned dynamic calculation speech separation platform based on human voice feature extraction adopts a differential encoding method for the two-way speech data, packs the encoding result, exit point information, signal-to-noise ratio and other parameters to generate metadata; the metadata is transmitted to the cloud through the communication layer, the cloud analyzes the exit point and the encoding method, and performs defect repair based on the generative adversarial network and noise suppression by using the adaptive spectral subtraction method on the speech data; the above-mentioned dynamic calculation speech separation platform based on human voice feature extraction performs logical completion and intent recognition on the optimized speech, and outputs clear, natural and semantic complete target speech data.
[0039] Through the above method steps, the computational amount can be adaptively reduced, the energy consumption is reduced, and the speech separation and understanding effect in a multi-speaker and noisy environment is improved.
[0040] In the embodiment of the application, speech data is acquired, feature extraction is performed on the speech data, and a complexity score is generated according to the extracted feature data; based on the complexity score, an exit point is determined, and speech separation processing is performed at the exit point to obtain separated speech data and corresponding metadata; sound quality optimization and semantic reasoning processing are performed based on the separated speech data and the corresponding metadata to obtain target speech data. Through the combination of the complexity score and the device running state, the exit point of the speech separation network is dynamically determined, the computational resource consumption is reduced, and the separation result is differentially optimized and semantically reasoned using the metadata, thereby improving the speech separation precision and semantic understanding ability in a multi-speaker and complex environment.
[0041] Optionally, in the step of acquiring speech data, an environmental speech signal can also be acquired through a preset microphone array; the environmental speech signal is preprocessed to obtain the speech data.
[0042] In the embodiment of the application, the dynamic computing speech separation platform based on human voice feature extraction described above acquires an environmental speech signal through a preset microphone array, inputs the acquired original signal to an end-side processing unit, and sequentially performs preprocessing operations such as echo cancellation, noise suppression, and automatic gain control to obtain speech data with low noise interference and balanced signal amplitude. The microphone array described above can adopt a ring array or a linear array structure to realize directional sound pickup of a target speaker. The preprocessing described above can reduce the interference of background noise and device echo on the speech signal through filtering, speech activity detection, and dynamic gain adjustment, and ensure the quality of the output speech data.
[0043] Through the above method steps, spatial sound pickup using a microphone array can enhance the directivity of the target speech and highlight the target speech in a multi-speaker or noisy environment; in combination with preprocessing operations such as echo cancellation, noise suppression, and automatic gain control, the environmental noise and echo interference can be reduced, the intelligibility and stability of the speech signal can be improved, the input quality of subsequent feature extraction and speech separation is ensured, and a high speech intelligibility and recognition accuracy can still be maintained in a noisy environment, thereby improving the overall performance of speech separation and semantic understanding.
[0044] Optionally, in the step of feature extraction on the speech data and generating the complexity score according to the extracted feature data, further comprising: performing extraction processing on the speech data through a preset hybrid network to obtain Mel spectrum features of each frame of audio data, and obtaining first audio feature data; performing modeling processing on inter-frame time sequence dependence of the first audio feature data based on an LSTM model to obtain second audio feature data; determining a speech overlap index and background sound effect complexity data in the second audio feature data; and generating the complexity score based on the speech overlap index and the background sound effect complexity data.
[0045] In the embodiment of the application, the preset hybrid network mentioned above can be a deep learning structure for speech feature extraction, which is usually composed of a convolutional neural network (CNN) and a fully connected layer. In the training stage, the preset hybrid network can learn the mapping relationship between the spectrum pattern and the speech structure through a large amount of speech data to obtain fixed parameters. In actual operation, the input speech signal can be directly processed by the network for feature calculation. For example, a 25ms frame length and a 10ms frame shift of the speech frame are input into the hybrid network, the CNN layer can capture the local energy pattern of the spectrogram, such as the formant or noise interference, and the fully connected layer can integrate these patterns to form a feature representation for subsequent processing.
[0046] The Mel spectrum feature mentioned above can be feature representation data obtained by decomposing the speech signal into a spectrum through a short-time Fourier transform (STFT) and then mapping the frequency to a Mel scale conforming to the human ear perception characteristics through a Mel filter bank. Specifically, the Mel scale can be used to reflect the characteristics of the human ear in which the resolution is high at low frequencies and low at high frequencies. For example, a 16kHz sampling rate speech signal is mapped to a 40-dimensional Mel filter bank after being decomposed into a spectrum through STFT, forming a 40-dimensional Mel feature vector for describing the frequency spectrum energy distribution of the speech in the time frame.
[0047] The first audio feature data mentioned above refers to the frame-level Mel spectrum feature extracted by the preset hybrid network. It reflects the frequency energy distribution of the speech in each frame of time segment. For example, a 40-dimensional vector is obtained by passing a frame of speech through a Mel filter bank, which is the first audio feature data for subsequent time sequence modeling.
[0048] The LSTM model mentioned above can be an improved recurrent neural network with a forgetting gate, an input gate, and an output gate mechanism, which can effectively capture long-term dependencies and short-term dynamics in long sequence data. Unlike traditional RNN, LSTM can avoid the gradient disappearance problem and is more suitable for modeling continuous time sequence signals such as speech.
[0049] The inter-frame temporal dependency can refer to the dynamic connection between speech frames in the time dimension. For example, in the continuous pronunciation of "ba" and "da", the transition part of different frames contains the dynamic characteristics of the speech, and the LSTM model can capture this trend by analyzing the frame sequence. Since the speech signals of different speakers can overlap in time, the inter-frame temporal dependency plays a key role in separating multi-speaker speech.
[0050] The second audio feature data is a time sequence feature representation obtained by modeling the first audio feature data through the LSTM model. It not only contains the spectral distribution of a single frame, but also integrates the dynamic information across frames. For example, if the Mel feature sequence of the previous 10 frames is input into the LSTM, the output second audio feature data can describe the speech intensity variation, timbre continuity, or speaker characteristics within the time window.
[0051] The speech overlap index can refer to a quantitative indicator of the multi-speaker overlap in the speech. It can be determined by calculating the proportion of frames in which multiple speakers speak simultaneously in the detected speech segment. For example, in 10 seconds of speech data, if 2 seconds contain two people speaking simultaneously, the speech overlap index is 0.2. It can be understood that the larger the speech overlap index, the more difficult the separation.
[0052] The background sound complexity data can be a quantitative indicator of the complexity of environmental noise, which is usually calculated by parameters such as spectral entropy, signal-to-noise ratio (SNR), and energy distribution diversity. For example, in a coffee shop scene, there are multiple people talking, music, and machine sounds in the background, the noise components are complex, the spectral entropy value is high, the SNR is low, and the complexity data is therefore large. In a quiet office environment, the background noise is single, and the complexity data is low.
[0053] In one possible embodiment, when the dynamic speech separation platform based on human voice feature extraction extracts features from the speech data and generates a complexity score, the preprocessed speech data is input into a preset mixing network for analysis, the Mel spectrum features of each frame of audio are extracted to obtain first audio feature data, and then the first audio feature data is input into a long short-term memory network (LSTM model) to model the temporal dependency between speech frames and generate second audio feature data. Based on the second audio feature data, the dynamic speech separation platform based on human voice feature extraction determines whether multiple speakers are speaking simultaneously in the speech segment, thereby calculating the speech overlap index, and analyzes the energy distribution, spectral entropy, and signal-to-noise ratio of the background noise to obtain the background sound complexity data. Finally, the dynamic speech separation platform based on human voice feature extraction combines the speech overlap index and the background sound complexity data by weighting to generate a complexity score, which is used to quantify the complexity of the current speech scene.
[0054] Specifically, the above complexity score can be calculated by the following formula: Wherein, the voice overlap index (O): for dialogue scenes, measure the degree of overlap of multi-speaker voice, the calculation formula is:
[0055] Wherein, T is the total number of audio frames (frame length 20ms), S(t) is the number of detected speakers in the t frame, The indicator function (satisfies the condition 1, otherwise 0). O takes the value range of 0-1, O<0.2 is low overlap (single person dialogue), O>0.5 is high overlap (more than 3 people interact).
[0056] Background sound effect complexity (C): for story scenes, measure the interference degree of background sound effect and narrative voice, the calculation formula is:
[0057] Wherein, The signal-to-noise ratio of the target narrative voice, The overall signal-to-noise ratio of the mixed audio. C takes the value range of 0-1, C<0.3 is low complexity (slight background sound), C>0.6 is high complexity (multiple sound effects superimposed).
[0058] Complexity score generation: the "voice overlap index (O)" and "background sound effect complexity (C)" are weighted and fused to generate a complexity score of 1-5 points (1 point is the lowest, 5 points is the highest), the weight is adjusted according to the scene (dialogue scene O weight 0.7, C weight 0.3; story scene O weight 0.3, C weight 0.7).
[0059] Through the above method steps, the complexity of the voice scene can be dynamically quantified according to the spectral features and timing characteristics of the voice frame, not only can accurately reflect the influence of multi-speaker overlap and background noise interference on the separation difficulty, but also can provide quantifiable reference basis for subsequent exit point determination. The method can improve the rationality of exit point selection, reduce the waste of computing resources, reduce the device power consumption while ensuring the separation accuracy, enhance the stability and robustness of voice separation in complex environment.
[0060] Optionally, in the step of determining the exit point based on the complexity score, the running load parameter and the temperature parameter of the current device are also obtained; based on the complexity score, the running load parameter and the temperature parameter, the target running state of the current device is determined; in the preset strategy table, the corresponding exit point is determined in real time according to the target running state of the current device.
[0061] In the embodiments of the present application, the running load parameter can be quantitative data of the current device real-time operation capacity occupation, including but not limited to CPU usage, NPU (neural network processing unit) utilization or memory occupancy. For example, when the CPU usage exceeds 80%, it indicates that the device is in a high load state, and continuing to run the deep separation network may cause the delay to increase.
[0062] The temperature parameter is the monitoring result of the real-time temperature of the device hardware, which is usually collected by an integrated temperature sensor on the chip surface or core area. For example, when the device temperature exceeds 75℃, the frequency reduction protection mechanism may be triggered, so it is necessary to select a shallow exit point in a high temperature scenario to avoid overheating.
[0063] The target running state can be the device working level determined based on the complexity score and the device running parameter, which is usually divided into three categories: low load, normal load and high load. For example, when the complexity score is 4 (complex scenario) and the device temperature is 40℃ and the CPU occupancy is 30%, the target running state is "normal load"; if the complexity score is 4 but the CPU occupancy is 90% and the temperature is 80℃, the target running state is "high load".
[0064] The preset strategy table can be a mapping relationship table predefined by the dynamic calculation voice separation platform based on the human voice feature extraction, which is used to establish a corresponding relationship between the target running state and the corresponding exit point. For example, if the target running state is "low load", the exit point is set to L4 (deep separation network); if it is "normal load", the exit point is L2~L3; if it is "high load", the exit point is L1, so as to ensure that acceptable separation results can still be output in the case of high device pressure.
[0065] Specifically, the preset strategy table can be explained by Table 2, and the structure of Table 2 is as follows:
[0066] According to Table 2, the mapping rule of "complexity score-chip load-exit point" can be established to provide an interpretable basis for dynamic decision-making.
[0067] In a possible embodiment, in the process of determining the exit point based on the complexity score, the human voice feature extraction-based dynamic computing speech separation platform obtains the running load parameter and temperature parameter of the current device, calculates the real-time operation occupation of the device processor or acceleration unit and the current temperature level of the chip or device hardware, inputs the complexity score, the running load parameter and the temperature parameter into the judgment together, and determines the target running state of the device. Finally, the human voice feature extraction-based dynamic computing speech separation platform looks up the corresponding relationship in the preset strategy table, determines the exit point in real time according to the target running state, and ensures that the speech separation network can realize reasonable operation depth and resource allocation under different complexity scenarios.
[0068] Optionally, in the step of performing speech separation processing at the exit point to obtain separated speech data and corresponding metadata, the step further includes calling a corresponding speech separation network based on the exit point; performing hierarchical separation processing on the speech data through the corresponding speech separation network to obtain target speech data and non-target speech data; performing asymmetric encoding processing on the target speech data and the non-target speech data to obtain corresponding target encoding results and non-target encoding results, respectively; and generating corresponding metadata based on the exit point, the target encoding results and the non-target encoding results.
[0069] In the embodiment of the application, the separated data can include target speech data and non-target speech data.
[0070] The speech separation network can be a hierarchical neural network model matched with the exit point. Specifically, the recommended network layer corresponding to the exit point can be selected according to Table 1. For example, in the case of a four-layer separation network, if the exit point is L2, the output result of the separation network calculated only to the second layer is called to ensure that a deep network is used in a complex scenario and a shallow network is used in a simple scenario, so as to balance separation accuracy and computing efficiency.
[0071] In this embodiment, the speech signal can be separated step by step in the selected separation network through a hierarchical progressive manner. For example, the first layer roughly distinguishes target speech from background noise, the second layer further distinguishes multi-speaker signals, and the third layer optimizes the details. Through this layer-by-layer separation, target speech data and non-target speech data are finally obtained.
[0072] The asymmetric encoding processing can refer to using different encoding methods for different types of speech data. Generally, the target speech data can be encoded using a high-fidelity encoding method (such as Opus or high-bit rate PCM) to ensure speech clarity, and the non-target speech data can be encoded using a low-bit rate or high-compression encoding method (such as low-bit rate MP3) to reduce storage and bandwidth consumption. More specifically, to balance sound quality and storage, the target speech (such as the main speaker in a conversation or the narrator in a story) is encoded using 16-bit PCM, and the non-target audio (such as background noise or secondary speakers) is encoded using 8-bit μ-law, reducing storage usage by 40%.
[0073] The target encoding result can refer to the output file or data stream obtained after high-fidelity encoding processing of the target speech data, which has high restoration degree and low distortion characteristics, and can ensure the accuracy of subsequent optimization and semantic analysis.
[0074] The non-target encoding result can refer to the output obtained after low-bit rate or high-compression encoding processing of the non-target speech data, which is mainly used to preserve environmental sound information or as a background reference, and does not emphasize high quality.
[0075] In one possible embodiment, the dynamic computing speech separation platform based on human voice feature extraction integrates and packages the exit point information, target encoding result and non-target encoding result, and outputs structured metadata. For example, the generated metadata can include fields: exit point=L2, target encoding=high-fidelity stream, non-target encoding=low-bit rate stream, timestamp=1023ms, which are used to guide the execution of cloud audio quality optimization and semantic reasoning.
[0076] In another possible embodiment, the dynamic computing speech separation platform based on human voice feature extraction calls the corresponding speech separation network according to the determined exit point, performs hierarchical separation processing on the input speech data through the speech separation network to obtain target speech data and non-target speech data, and to optimize the subsequent transmission and storage efficiency, the dynamic computing speech separation platform based on human voice feature extraction performs asymmetric encoding processing on the target speech data and the non-target speech data, i.e., uses different encoding methods according to the importance of the two types of data, respectively generates corresponding target encoding result and non-target encoding result, and then generates corresponding metadata based on the exit point information, target encoding result and non-target encoding result.
[0077] By the above method steps, the adapted speech separation network can be called at the exit point, the hierarchical separation processing of the speech signal is realized, and the asymmetric coding is performed according to the importance of different categories of data, and finally the metadata containing the exit point and the coding result is generated. Not only the intelligibility and understandability of the target speech are ensured, but also the data amount of the non-target speech is reduced, the storage and transmission overheads are reduced, meanwhile, the generated metadata provides a standardized input for subsequent optimization and semantic processing, and the adaptability and processing efficiency of the above dynamic computing speech separation platform based on human voice feature extraction in low-power devices and complex speech environments are improved.
[0078] Optionally, in the step of performing voice quality optimization and semantic reasoning processing on the separated speech data and the corresponding metadata to obtain target speech data, the exit point information and the coding parameter in the metadata are analyzed, the separated speech data is subjected to voice quality optimization processing to obtain optimized speech data, and the optimized speech data is subjected to semantic reasoning processing to output target speech data.
[0079] In the embodiment of the application, the exit point information can refer to the depth at which the separation process is terminated in the hierarchical network corresponding to the exit point after the exit point is determined. For example, when the exit point is L2, it means that the separation network only runs to the second layer, and the result may have some residual noise; this information can guide the voice quality optimization module to use a stronger enhancement strategy.
[0080] The coding parameter can refer to the way and parameter setting used when encoding the target speech data and the non-target speech data in the separation stage, such as bit rate, sampling rate, encoding format, etc. By analyzing these parameters, the optimization module can select appropriate enhancement and decoding methods according to the data quality difference.
[0081] The voice quality optimization processing can refer to the process of post-processing the separated speech data to improve the speech clarity and naturalness. Specifically, it can include but is not limited to GAN-based spectrum repair, noise suppression based on spectral subtraction, and loudness equalization based on adaptive filtering, etc. For example, when the exit point is shallow and the separation result has errors, the signal in a specific frequency band can be enhanced to compensate for the defects.
[0082] Specifically, the repair can be performed by the following steps: GAN-based speech repair: an improved WaveNet-GAN model is used to repair the possible speech discontinuity (such as frame loss in high overlap scenarios) after device-side separation, the input is the device-side separated audio and the spectral features in the metadata, and the output is continuous speech, with a repair rate > 92%; Adaptive noise reduction: according to the noise level in the metadata, the spectral subtraction method is used to dynamically adjust the noise reduction strength, and the formula is:
[0083] wherein, is the time-frequency spectrum of the noisy speech, is the estimated noise spectrum, is the noise reduction coefficient (dynamically adjusted according to the noise level in the metadata, ranging from 1.0 to 1.5), ensuring that noise reduction does not damage the target speech.
[0084] The above-mentioned optimized speech data refers to the target speech signal after quality optimization processing. Compared with the separated speech data, it has higher clarity, and distortion and noise components are suppressed, so it is more suitable for semantic reasoning analysis.
[0085] The above-mentioned semantic reasoning processing can refer to the process of speech recognition and semantic understanding of the optimized speech data. For example, in the instruction scenario, the speech is transcribed into text and the intent and slot are extracted; in the dialogue scenario, the logical completion and context understanding of the speech content are performed, thereby generating semantic complete and structured target speech data.
[0086] Specifically, the semantic reasoning processing can be performed by the following steps: Story content enhancement: using a 1B parameter lightweight LLM, according to the story scene label (such as "fairy tale" "popular science") in the metadata, completing the narrative logic (such as repairing the story's plot discontinuity), generating coherent story audio; Dialogue semantic understanding: using a fine-tuned BERT model, identifying user dialogue intent (such as "query weather" "control device"), combining with the speaker identification in the metadata, ensuring the response is targeted; Multi-device synchronization calibration: through NTP protocol to realize the playback synchronization of multiple AI voice devices (such as multiple smart speakers in the home), time difference control in < 50ms, to avoid echo interference when multiple devices play.
[0087] Through the above method steps, after analyzing the exit point information and encoding parameters of the metadata, the separated speech data can be optimized in quality, making the results clearer and more natural, and then combined with semantic reasoning processing to output semantic complete target speech data. Not only does it improve the recognition accuracy of speech in a multi-speaker and noisy environment, but also ensures low latency and high robustness of the system, achieving collaborative optimization of speech separation and understanding.
[0088] Optionally, after the speech data and corresponding metadata are processed by voice quality optimization and semantic reasoning to obtain target speech data, the method further includes compressing the target speech data using a first encoding method and compressing the non-target speech data using a second encoding method to obtain a differentiated compression result; forming a transmission data packet based on the compression result and the metadata; and adaptively switching between short-distance communication and long-distance communication according to a current network signal strength, so that the transmission data packet is sent to the cloud in real time through short-distance communication or long-distance communication.
[0089] In the embodiment of the application, the first encoding method refers to a high-fidelity and low-compression-rate encoding method for target speech data, such as high-bit-rate Opus encoding or PCM encoding, to ensure that the target speech still maintains good clarity and integrity after transmission, and is suitable for core data that needs to be recognized and further processed.
[0090] The second encoding method refers to a low-bit-rate and high-compression-rate encoding method for non-target speech data, such as low-code-rate MP3 or ADPCM encoding, to reduce the data volume as much as possible and only retain environmental information or background reference signals without emphasizing high-quality output.
[0091] The differentiated compression result refers to a compressed file or data stream obtained by using different encoding strategies for target speech data and non-target speech data. The compression result corresponding to the target speech data has high clarity, and the compression result corresponding to the non-target speech data has small data volume, thereby achieving a balance between high quality and high efficiency.
[0092] The transmission data packet refers to a standardized data unit generated by integrating the differentiated compression result and metadata information. The data packet not only contains the encoding result of the speech signal, but also includes exit point information, encoding parameters, timestamps, etc., to ensure that the cloud can correctly parse and process.
[0093] In the embodiment, the best transmission path can be dynamically selected between short-distance communication (such as Wi-Fi, Bluetooth) and long-distance communication (such as 4G / 5G cellular network) according to the real-time detected network signal strength. For example, when the device is under indoor Wi-Fi signal coverage, short-distance communication is preferred; when leaving the Wi-Fi environment, the cellular network is automatically switched to, to ensure uninterrupted transmission.
[0094] The cloud can refer to a remote server or computing platform connected to the device, which has computing and storage capabilities. The data packet transmitted to the cloud can be further used for voice quality enhancement, semantic depth analysis or application service calling, such as executing a large-scale speech recognition model in the cloud to improve recognition accuracy.
[0095] In a possible embodiment, after the dynamic computing speech separation platform based on human voice feature extraction described above performs voice quality optimization and semantic reasoning processing on the separated speech data and the corresponding metadata to obtain target speech data, the target speech data is compressed using a first encoding method to ensure high clarity and intelligibility of the speech during transmission; the non-target speech data is compressed using a second encoding method to reduce storage and bandwidth occupation, thereby obtaining a differentiated compression result, and then a transmission data packet is generated according to the compression result and the metadata, and a transmission path is dynamically selected according to the current network signal strength to adaptively switch between short-distance communication and long-distance communication, so that the transmission data packet can be sent to the cloud in real time through short-distance communication modes such as Bluetooth and Wi-Fi, or long-distance communication modes such as cellular networks, to support subsequent storage, optimization or interaction processing.
[0096] Through the above method steps, the target speech and the non-target speech can be differentiated and compressed, which not only ensures the clarity of the key speech data, but also reduces the bandwidth consumption of the redundant data; at the same time, by generating a transmission data packet containing metadata and combining the network signal strength to realize adaptive switching between short-distance and long-distance communication, it is ensured that the speech data can be stably and real-time transmitted to the cloud. While improving the data transmission efficiency and system robustness, the availability and interaction experience of the target speech in a complex network environment are also ensured.
[0097] As shown in Figure 2 The embodiment of the application also provides a dynamic computing speech separation device 200 based on human voice feature extraction, which comprises: A first acquisition module 201 is configured to acquire speech data. A first extraction module 202 is configured to perform feature extraction on the speech data and generate a complexity score according to the extracted feature data. A first determination module 203 is configured to determine an exit point based on the complexity score, and perform speech separation processing at the exit point to obtain separated speech data and corresponding metadata. A first processing module 204 is configured to perform voice quality optimization and semantic reasoning processing on the separated speech data and the corresponding metadata to obtain target speech data.
[0098] Optionally, the first acquisition module 201 comprises: A first acquisition sub-module is configured to acquire environmental speech signals through a preset microphone array. A first acquisition sub-module is configured to acquire environmental speech signals through a preset microphone array.
[0099] Optionally, the first extraction module 202 comprises: a first extraction submodule configured to perform extraction processing on the voice data by using a preset mixing network to obtain Mel spectrum features of each frame of audio data, and obtain first audio feature data; a first processing submodule configured to perform modeling processing on inter-frame time sequence dependence of the first audio feature data based on an LSTM model, and obtain second audio feature data; a first determination submodule configured to determine a voice overlap index and background sound effect complexity data in the second audio feature data; a first generation submodule configured to generate a complexity score based on the voice overlap index and the background sound effect complexity data.
[0100] Optionally, the first determination module 203 includes: a second acquisition submodule configured to acquire a running load parameter and a temperature parameter of a current device; a second determination submodule configured to determine a target running state of the current device based on the complexity score, the running load parameter, and the temperature parameter; a third determination submodule configured to determine a corresponding exit point in real time according to the target running state of the current device in a preset strategy table.
[0101] Optionally, the first determination module 203 further includes: a first calling submodule configured to call a corresponding voice separation network based on the exit point; a second processing submodule configured to perform hierarchical separation processing on the voice data by using the corresponding voice separation network to obtain target voice data and non-target voice data; a third processing submodule configured to perform asymmetric coding processing on the target voice data and the non-target voice data to obtain corresponding target coding results and non-target coding results, respectively; a second generation submodule configured to generate corresponding metadata based on the exit point, the target coding results, and the non-target coding results.
[0102] Optionally, the first processing module 204 includes: a first analysis submodule configured to analyze exit point information and coding parameters in the metadata, perform sound quality optimization processing on the separated voice data, and obtain optimized voice data; a fourth processing submodule configured to perform semantic reasoning processing on the optimized voice data, and output the target voice data.
[0103] Optionally, the apparatus further includes: The compression module is configured to compress the target voice data in a first encoding mode and compress the non-target voice data in a second encoding mode to obtain a differentiated compression result. The transmission module is configured to form a transmission data packet based on the compression result and the metadata. The switching module is configured to adaptively switch between short-distance communication and long-distance communication according to a current network signal strength, so that the transmission data packet is transmitted to the cloud in real time through short-distance communication or long-distance communication.
[0104] As Figure 3 shown, the embodiment of the present application also provides an electronic device 300, comprising a processor, and the processor can execute any one of the dynamic calculation voice separation methods based on human voice feature extraction.
[0105] Specifically, the electronic device 300 comprises a processor 301 and a memory 302, and a computer program for executing the dynamic calculation voice separation method based on human voice feature extraction stored in the memory 302 and capable of running on the processor 301, wherein: The processor 301 runs the computer program of the dynamic calculation voice separation method based on human voice feature extraction stored in the memory 302, and executes the following steps: Obtain voice data; Feature extraction is performed on the voice data, and a complexity score is generated according to the extracted feature data; Based on the complexity score, a exit point is determined, and a voice separation process is performed at the exit point to obtain separated voice data and corresponding metadata; Based on the separated voice data and the corresponding metadata, voice quality optimization and semantic reasoning processing are performed to obtain target voice data.
[0106] Optionally, the processor 301 executes the voice data acquisition, comprising: An environmental voice signal is collected through a preset microphone array; The environmental voice signal is preprocessed to obtain voice data.
[0107] Optionally, the processor 301 executes the feature extraction on the voice data and generates a complexity score according to the extracted feature data, comprising: The voice data is extracted by a preset mixed network to obtain Mel spectrum features of each frame of audio data, and first audio feature data is obtained; The inter-frame time sequence dependence of the first audio feature data is modeled based on an LSTM model to obtain second audio feature data; The voice overlap index and the background sound effect complexity data in the second audio feature data are determined; generate a complexity score based on the voice overlap index and the background sound complexity data.
[0108] Optionally, the processor 301 performs the determining an exit point based on the complexity score, comprising: obtaining a running load parameter and a temperature parameter of the current device; determining a target running state of the current device based on the complexity score, the running load parameter and the temperature parameter; in a preset strategy table, determining a corresponding exit point in real time according to the target running state of the current device.
[0109] Optionally, the processor 301 performs the separating data into target voice data and non-target voice data, and performs voice separation processing at the exit point to obtain separated voice data and corresponding metadata, comprising: calling a corresponding voice separation network based on the exit point; performing hierarchical separation processing on the voice data through the corresponding voice separation network to obtain target voice data and non-target voice data; performing asymmetric encoding processing on the target voice data and the non-target voice data to obtain corresponding target encoding results and non-target encoding results, respectively; generating corresponding metadata based on the exit point, the target encoding results and the non-target encoding results.
[0110] Optionally, the processor 301 performs the performing voice quality optimization and semantic reasoning processing on the separated voice data and the corresponding metadata to obtain target voice data, comprising: analyzing exit point information and encoding parameters in the metadata to perform voice quality optimization processing on the separated voice data to obtain optimized voice data; performing semantic reasoning processing on the optimized voice data to output the target voice data.
[0111] Optionally, after the processor 301 performs the performing voice quality optimization and semantic reasoning processing on the separated voice data and the corresponding metadata to obtain target voice data, the method further comprises: compressing the target voice data using a first encoding method and compressing the non-target voice data using a second encoding method to obtain differential compression results; forming a transmission data packet based on the compression results and the metadata; performing adaptive switching between short-distance communication and long-distance communication according to a current network signal strength, so as to send the transmission data packet to the cloud in real time through short-distance communication or long-distance communication.
[0112] The computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement each process of the dynamic computing speech separation method based on human voice feature extraction or the application end dynamic computing speech separation method based on human voice feature extraction provided by the embodiments of the present application, and the same technical effects can be achieved. To avoid repetition, details are not described here.
[0113] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiments can be completed by a computer program instructing related hardware, and can be stored in a computer readable storage medium. When the program is executed, it can include the processes of the above-mentioned embodiments. The storage medium can be a magnetic disc, an optical disc, a read-only memory (ROM), or a random access memory (RAM), etc.
[0114] The above only discloses preferred embodiments of the present application, and of course cannot limit the scope of the right of the present application, so equivalent changes made according to the claims of the present application are still within the scope of the present application.
Claims
1. A dynamic computational speech separation method based on human voice feature extraction, characterized in that, include: Acquire voice data; Feature extraction is performed on the speech data, and a complexity score is generated based on the extracted feature data; Based on the complexity score, an exit point is determined, and speech separation processing is performed at the exit point to obtain separated speech data and corresponding metadata. The target speech data is obtained by performing sound quality optimization and semantic reasoning processing based on the separated speech data and corresponding metadata.
2. The dynamic computational speech separation method based on human voice feature extraction as described in claim 1, characterized in that, The acquisition of voice data includes: Environmental voice signals are collected using a preset microphone array; The environmental speech signal is preprocessed to obtain speech data.
3. The dynamic computational speech separation method based on human voice feature extraction as described in claim 1, characterized in that, The step of extracting features from the speech data and generating a complexity score based on the extracted features includes: The speech data is extracted and processed by a preset hybrid network to obtain the Mel spectrum features of each frame of audio data, thus obtaining the first audio feature data. The inter-frame temporal dependency of the first audio feature data is modeled and processed based on the LSTM model to obtain the second audio feature data. Determine the speech overlap index and background sound effect complexity data in the second audio feature data; A complexity score is generated based on the speech overlap index and background sound effect complexity data.
4. The dynamic computational speech separation method based on human voice feature extraction as described in claim 1, characterized in that, The process of determining the exit point based on the complexity score includes: Obtain the current operating load parameters and temperature parameters of the device; Based on the complexity score, operating load parameters, and temperature parameters, the target operating state of the current device is determined; In the preset strategy table, the corresponding exit point is determined in real time based on the current target operating state of the device.
5. The dynamic computational speech separation method based on human voice feature extraction as described in claim 1, characterized in that, The separated data includes target speech data and non-target speech data. The speech separation process is performed at the exit point to obtain the separated speech data and corresponding metadata, including: Based on the exit point, the corresponding speech separation network is invoked; The speech data is processed by the corresponding speech separation network to obtain target speech data and non-target speech data. The target speech data and non-target speech data are subjected to asymmetric coding processing to obtain the corresponding target coding results and non-target coding results, respectively. Based on the exit point, target encoding result, and non-target encoding result, corresponding metadata is generated.
6. The dynamic computational speech separation method based on human voice feature extraction as described in claim 1, characterized in that, The process of optimizing sound quality and performing semantic reasoning based on the separated speech data and corresponding metadata to obtain target speech data includes: The exit point information and encoding parameters in the metadata are parsed, and the separated speech data is subjected to sound quality optimization processing to obtain optimized speech data. Semantic reasoning processing is performed on the optimized speech data to output the target speech data.
7. The dynamic computational speech separation method based on human voice feature extraction as described in claim 6, characterized in that, After performing sound quality optimization and semantic reasoning processing on the separated speech data and corresponding metadata to obtain the target speech data, the method further includes: The target speech data is compressed using a first encoding method, and the non-target speech data is compressed using a second encoding method to obtain differentiated compression results; Based on the compression result and the metadata, a transmission data packet is formed; The system adaptively switches between short-distance and long-distance communication based on the current network signal strength, so that the data packets can be sent to the cloud in real time via either short-distance or long-distance communication.
8. A dynamic computational speech separation device based on human voice feature extraction, characterized in that, include: The first acquisition module is used to acquire voice data; The first extraction module is used to extract features from the speech data and generate a complexity score based on the extracted feature data. The first determining module is used to determine the exit point based on the complexity score, and perform speech separation processing at the exit point to obtain separated speech data and corresponding metadata; The first processing module is used to perform sound quality optimization and semantic reasoning processing based on the separated speech data and corresponding metadata to obtain the target speech data.
9. An electronic device, characterized in that, include: The memory, the processor, and the computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the dynamic computational speech separation method based on human voice feature extraction as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the dynamic computational speech separation method based on human voice feature extraction as described in any one of claims 1 to 7.