Voice processing methods, devices, terminal equipment and storage media
Patent Information
- Application Number
- CN202310689788.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-09
- Publication Date
- 2026-09-01
- Estimated Expiration
- 2043-06-09
AI Technical Summary
[0005]本申请的主要目的在于提供一种语音处理方法、装置、终端设备及存储介质,旨在解决或改善目前识别目标声音分区的方法容易导致音区泄露的问题
[0016]本申请实施例提出的语音处理方法、装置、终端设备及存储介质,通过获取若干个声音分区各自对应的待处理语音信号;基于预设的清晰度评估模型对所述待处理语音信号进行清晰度评估,得到对应的评估结果;基于所述评估结果,确定目标声音分区。基于本申请方案,采用清晰度评估模型可以对待处理语音信号进行清晰度评估,得到反映语音清晰度的评估结果,并进一步根据评估结果确定目标声音分区。如此可以摆脱对参考语音信号的依赖,并且清晰度评估模型能够适应环境噪声和说话人身姿改变等因素对待处理语音信号造成的动态影响,在此基础上能够准确地确定目标声音分区,有效降低了音区泄露的情况发生。
Smart Images

Figure CN116741161B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of speech processing technology, and in particular to a speech processing method, apparatus, terminal device and storage medium. Background Technology
[0002] Cars are becoming increasingly intelligent, and their voice interaction functions are becoming more and more sophisticated. Based on the seating arrangement, the interior of a car can be divided into several sound zones. In certain voice interaction scenarios, it is necessary to accurately identify the speaker's sound zone, i.e., the target sound zone.
[0003] Currently, the method for identifying target voice regions involves comparing the acquired speech signal to be processed with a reference speech signal and determining the target voice region based on the comparison result. However, environmental noise and changes in the speaker's posture can have a continuous and dynamic impact on the acquired speech signal to be processed, and the reference speech signal cannot dynamically adapt to these impacts. This can easily lead to the identified target voice region not being the speaker's voice region, a misidentification known as voice region leakage.
[0004] In summary, current methods for identifying target sound regions are prone to causing sound region leakage. Summary of the Invention
[0005] The main purpose of this application is to provide a speech processing method, apparatus, terminal device and storage medium, which aims to solve or improve the problem that current methods for identifying target voice partitions are prone to causing voice partition leakage.
[0006] To achieve the above objectives, this application provides a speech processing method, the speech processing method comprising: Acquire the speech signals to be processed corresponding to each of the several sound partitions; The clarity of the speech signal to be processed is evaluated based on a preset clarity evaluation model to obtain the corresponding evaluation results; Based on the evaluation results, the target sound zone is determined.
[0007] Optionally, the sharpness evaluation model includes a frame-by-frame convolutional model, a temporal model, and a pooling model. The step of evaluating the sharpness of the speech signal to be processed based on the preset sharpness evaluation model to obtain the corresponding evaluation result includes: The speech signal to be processed is spectrum segmented based on a preset window length to obtain several corresponding frames of spectrum. The pre-defined frame-by-frame convolution model is used to convolve the spectrum of the several frames one by one to obtain the first type of high-dimensional features corresponding to each of the several frames of spectrum. The first type of high-dimensional features are modeled in a time-dependent manner using a preset time-series model to obtain the second type of high-dimensional features corresponding to each of the several frame spectra. The second type of high-dimensional features corresponding to the spectrum of each of the several frames are aggregated using a preset pooling model to obtain aggregated features; The corresponding evaluation results are obtained based on the aggregation feature analysis.
[0008] Optionally, the step of obtaining the corresponding evaluation result based on the aggregation feature analysis includes: The corresponding evaluation score is obtained based on the aggregated feature analysis, wherein the type of the evaluation score includes at least one of MOS value, noise evaluation score, and human voice evaluation score.
[0009] Optionally, the step of obtaining the corresponding evaluation score based on the aggregation feature analysis includes: Based on the aggregated features and the preset speech quality scoring criteria, the corresponding evaluation score is obtained through analysis.
[0010] Optionally, the step of determining the target sound partition based on the evaluation results includes: Based on the evaluation score and the preset threshold filtering rules, the several sound partitions are filtered to determine at least one candidate sound partition; The target sound partition is determined by comparing the evaluation score corresponding to the candidate sound partition with the preset score comparison rules.
[0011] Optionally, before the step of evaluating the clarity of the speech signal to be processed based on a preset clarity evaluation model to obtain the corresponding evaluation result, the method further includes: The speech signal to be processed is subjected to speech activity detection to determine at least one sound zone with speech activity; The step of evaluating the clarity of the speech signal to be processed based on a preset clarity evaluation model to obtain the corresponding evaluation result includes: The clarity of the speech signal to be processed corresponding to the speech-active sound partition is evaluated based on the preset clarity evaluation model, and the corresponding evaluation results are obtained.
[0012] Optionally, after determining the target sound partition based on the evaluation result, the method further includes: The preset sound zone allocation strategy is adjusted according to the target sound zone to obtain the adjusted sound zone allocation strategy, wherein the adjusted sound zone allocation strategy is used to control the voice interaction task corresponding to the target sound zone.
[0013] This application also proposes a voice processing device, the voice processing device comprising: The acquisition module is used to acquire the audio signals to be processed corresponding to each of the several sound partitions. The evaluation module is used to evaluate the clarity of the speech signal to be processed based on a preset clarity evaluation model, and obtain the corresponding evaluation results. The determination module is used to determine the target sound partition based on the evaluation results.
[0014] This application also proposes a terminal device, which includes a memory, a processor, and a voice processing program stored in the memory and executable on the processor. When the voice processing program is executed by the processor, it implements the steps of the voice processing method described above.
[0015] This application also proposes a computer-readable storage medium storing a speech processing program, which, when executed by a processor, implements the steps of the speech processing method described above.
[0016] The speech processing method, apparatus, terminal device, and storage medium proposed in this application acquire speech signals to be processed corresponding to several sound zones; evaluate the speech signals to be processed based on a preset clarity evaluation model to obtain corresponding evaluation results; and determine the target sound zone based on the evaluation results. Based on the scheme of this application, the clarity evaluation model can evaluate the clarity of the speech signal to be processed, obtain evaluation results reflecting speech clarity, and further determine the target sound zone based on the evaluation results. This eliminates the dependence on reference speech signals, and the clarity evaluation model can adapt to the dynamic effects of environmental noise and changes in speaker posture on the speech signal to be processed. Based on this, the target sound zone can be accurately determined, effectively reducing the occurrence of sound zone leakage. Attached Figure Description
[0017] Figure 1 This is a schematic diagram of the functional modules of the terminal device to which the voice processing device of this application belongs; Figure 2 This is a schematic diagram of the first exemplary embodiment of the speech processing method of this application; Figure 3 This is a schematic diagram of the second exemplary embodiment of the speech processing method of this application; Figure 4 This is a schematic diagram of the clarity assessment model involved in the speech processing method of this application; Figure 5 This is a schematic diagram of the third exemplary embodiment of the speech processing method of this application; Figure 6This is a schematic diagram of the fourth exemplary embodiment of the speech processing method of this application; Figure 7 This is a schematic diagram of the fifth exemplary embodiment of the speech processing method of this application; Figure 8 This is a schematic diagram of the sixth exemplary embodiment of the speech processing method of this application; Figure 9 This is a schematic diagram of the seventh exemplary embodiment of the speech processing method of this application.
[0018] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0019] It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit this application.
[0020] The main solution of this application embodiment is as follows: acquiring the speech signals to be processed corresponding to several sound partitions; evaluating the speech signals to be processed based on a preset clarity evaluation model to obtain corresponding evaluation results; and determining the target sound partition based on the evaluation results. Based on this application solution, the clarity evaluation model can evaluate the clarity of the speech signals to be processed, obtain evaluation results reflecting speech clarity, and further determine the target sound partition based on the evaluation results. This eliminates the dependence on reference speech signals, and the clarity evaluation model can adapt to the dynamic effects of environmental noise and changes in speaker posture on the speech signals to be processed. Based on this, the target sound partition can be accurately determined, effectively reducing the occurrence of sound partition leakage.
[0021] Specifically, refer to Figure 1 , Figure 1 This is a functional module diagram of the terminal device to which the voice processing device of this application belongs. The voice processing device can be a device capable of voice processing, independent of the terminal device, and can be implemented on the terminal device in hardware or software form. The terminal device can be a smart mobile terminal with data processing capabilities, such as a mobile phone or tablet computer, or it can be a fixed terminal device or server with data processing capabilities.
[0022] In this embodiment, the terminal device to which the voice processing device belongs includes at least an output module 110, a processor 120, a memory 130, and a communication module 140.
[0023] The memory 130 stores the operating system and the voice processing program. The voice processing device can acquire several voice signals corresponding to each of the voice zones to be processed; evaluate the voice signals to be processed based on a preset clarity evaluation model, and obtain the corresponding evaluation results; based on the evaluation results, the voice zone information corresponding to the target voice zone is stored in the memory 130; the output module 110 may be a display screen, etc. The communication module 140 may include a WIFI module, a mobile communication module, and a Bluetooth module, etc., and communicates with external devices or servers through the communication module 140.
[0024] When the voice processing program in memory 130 is executed by the processor, it performs the following steps: Acquire the speech signals to be processed corresponding to each of the several sound partitions; The clarity of the speech signal to be processed is evaluated based on a preset clarity evaluation model to obtain the corresponding evaluation results; Based on the evaluation results, the target sound zone is determined.
[0025] Furthermore, when the voice processing program in memory 130 is executed by the processor, it also performs the following steps: The speech signal to be processed is spectrum segmented based on a preset window length to obtain several corresponding frames of spectrum. The pre-defined frame-by-frame convolution model is used to convolve the spectrum of the several frames one by one to obtain the first type of high-dimensional features corresponding to each of the several frames of spectrum. The first type of high-dimensional features are modeled in a time-dependent manner using a preset time-series model to obtain the second type of high-dimensional features corresponding to each of the several frame spectra. The second type of high-dimensional features corresponding to the spectrum of each of the several frames are aggregated using a preset pooling model to obtain aggregated features; The corresponding evaluation results are obtained based on the aggregation feature analysis.
[0026] Furthermore, when the voice processing program in memory 130 is executed by the processor, it also performs the following steps: The corresponding evaluation score is obtained based on the aggregated feature analysis, wherein the type of the evaluation score includes at least one of MOS value, noise evaluation score, and human voice evaluation score.
[0027] Furthermore, when the voice processing program in memory 130 is executed by the processor, it also performs the following steps: Based on the aggregated features and the preset speech quality scoring criteria, the corresponding evaluation score is obtained through analysis.
[0028] Furthermore, when the voice processing program in memory 130 is executed by the processor, it also performs the following steps: Based on the evaluation score and the preset threshold filtering rules, the several sound partitions are filtered to determine at least one candidate sound partition; The target sound partition is determined by comparing the evaluation score corresponding to the candidate sound partition with the preset score comparison rules.
[0029] Furthermore, when the voice processing program in memory 130 is executed by the processor, it also performs the following steps: The speech signal to be processed is subjected to speech activity detection to determine at least one sound zone with speech activity; The clarity of the speech signal to be processed corresponding to the speech-active sound partition is evaluated based on the preset clarity evaluation model, and the corresponding evaluation results are obtained.
[0030] Furthermore, when the voice processing program in memory 130 is executed by the processor, it also performs the following steps: The preset sound zone allocation strategy is adjusted according to the target sound zone to obtain the adjusted sound zone allocation strategy, wherein the adjusted sound zone allocation strategy is used to control the voice interaction task corresponding to the target sound zone.
[0031] This embodiment, through the above-described scheme, specifically acquires the speech signals to be processed corresponding to several sound zones; evaluates the speech signals to be processed based on a preset clarity assessment model to obtain corresponding evaluation results; and determines the target sound zone based on the evaluation results. In this embodiment, the clarity assessment model can evaluate the clarity of the speech signals to be processed, obtain evaluation results reflecting speech clarity, and further determine the target sound zone based on the evaluation results. This eliminates the dependence on reference speech signals, and the clarity assessment model can adapt to the dynamic effects of environmental noise and changes in speaker posture on the speech signals to be processed. Based on this, it can accurately determine the target sound zone, effectively reducing the occurrence of sound zone leakage.
[0032] Reference Figure 2 The first embodiment of the speech processing method of this application provides a flowchart, the speech processing method including: Step S10: Obtain the speech signals to be processed corresponding to each of the several sound partitions.
[0033] Specifically, the voice processing method involved in this embodiment can be applied to scenarios of multi-zone voice interaction, such as in a car's smart cockpit. Sound zoning refers to the partitioning of a whole sound system according to different areas. Taking a four- or five-seat smart cockpit system as an example, it can be divided into a driver's seat sound zone, a front passenger seat sound zone, a left rear sound zone, and a right rear sound zone. Furthermore, based on the sound zoning, each sound zone can independently support at least one function among sound acquisition, sound playback, and control of linked components.
[0034] To determine the speaker's location within a specific audio zone, it is necessary to acquire the corresponding audio signals for each of these zones. More specifically, a recording unit, such as a microphone unit, can be pre-set for each audio zone, thus allowing the acquisition of the corresponding audio signals for each zone.
[0035] Step S20: Based on a preset clarity assessment model, the clarity of the speech signal to be processed is assessed to obtain the corresponding assessment result.
[0036] Specifically, after acquiring the speech signals to be processed corresponding to several sound partitions, these signals are used as input to a clarity assessment model. The model evaluates the clarity of the speech signals to be processed, yielding assessment results for each sound partition. The clarity assessment model performs spectral segmentation, convolution, temporal dependency modeling, and feature aggregation on the speech signals to be processed, initially outputting aggregated features that characterize the clarity of the speech signals. Further quantification of these aggregated features yields the corresponding assessment results.
[0037] Understandably, the evaluation results can be evaluation scores, evaluation grades, or other data formats that can be used to characterize the clarity of the speech signal being processed.
[0038] Step S30: Based on the evaluation results, determine the target sound zone.
[0039] Specifically, the evaluation results reflect the clarity of the corresponding sound partition, and the sound partition with the best clarity can be determined as the target sound partition based on the evaluation results. For example, when the evaluation result is an evaluation score, the sound partition with the highest evaluation score can be determined as the target sound partition by comparing scores; similarly, when the evaluation result is an evaluation level, the sound partition with the highest evaluation level can be determined as the target sound partition by comparing levels. Likewise, if the evaluation result uses other data formats that can characterize the clarity of the speech signal to be processed, the sound partition with the best clarity can also be determined as the target sound partition based on the corresponding analysis method.
[0040] This embodiment, through the above-described scheme, specifically acquires the speech signals to be processed corresponding to several sound zones; evaluates the speech signals to be processed based on a preset clarity assessment model to obtain corresponding evaluation results; and determines the target sound zone based on the evaluation results. In this embodiment, the clarity assessment model can evaluate the clarity of the speech signals to be processed, obtain evaluation results reflecting speech clarity, and further determine the target sound zone based on the evaluation results. This eliminates the dependence on reference speech signals, and the clarity assessment model can adapt to the dynamic effects of environmental noise and changes in speaker posture on the speech signals to be processed. Based on this, it can accurately determine the target sound zone, effectively reducing the occurrence of sound zone leakage.
[0041] Furthermore, referring to Figure 3 The second embodiment of the speech processing method of this application provides a flowchart, based on the above. Figure 2 In the embodiment shown, step S20 involves evaluating the clarity of the speech signal to be processed based on a preset clarity evaluation model, and further refining the corresponding evaluation results, including: Step S201: The speech signal to be processed is spectral segmented based on a preset window length to obtain several corresponding frames of spectrum.
[0042] Specifically, the window length (i.e., window function) can be preset according to the actual spectrum analysis needs. The short-time Fourier transform (STFT) is used to divide the speech signal to be processed corresponding to several sound partitions into multiple frames. Then, each frame is multiplied by the window length to obtain the corresponding spectrum, that is, to obtain the corresponding spectrum of several frames.
[0043] Step S202: Perform frame-by-frame convolution on the spectrum of the several frames using a preset frame-by-frame convolution model to obtain the first type of high-dimensional features corresponding to each of the several frames of spectrum.
[0044] Specifically, such as Figure 4 As shown, Figure 4 This diagram illustrates the clarity assessment model involved in the speech processing method of this application. The clarity assessment model includes a frame-by-frame convolutional model, which is a model based on Convolutional Neural Networks (CNN). Several frames of spectrum are used as input to the frame-by-frame convolutional model. By performing frame-by-frame convolution on the several frames of spectrum, the model can obtain the first type of high-dimensional features corresponding to each frame of spectrum. It can be understood that the first type of high-dimensional features are high-dimensional features extracted from space.
[0045] Step S203: The first type of high-dimensional features are modeled on a time-dependent basis using a preset time-series model to obtain the second type of high-dimensional features corresponding to each of the several frame spectra.
[0046] Specifically, such as Figure 4 As shown, the sharpness assessment model includes a temporal model, which is a model based on Long Short-Term Memory (LSTM) networks. The second-type high-dimensional features corresponding to several frames of spectrum are used as input to the temporal model. The temporal model models the temporal dependencies of the first-type high-dimensional features, thus obtaining the second-type high-dimensional features corresponding to each frame of spectrum. It can be understood that the second-type high-dimensional features are high-dimensional features extracted over time.
[0047] Step S204: The second type of high-dimensional features corresponding to the spectrum of each of the several frames are aggregated using a preset pooling model to obtain aggregated features.
[0048] Specifically, such as Figure 4 As shown, the clarity assessment model includes a pooling model, which can reduce computational cost while preserving key features. By using the second-type high-dimensional features corresponding to several frames of spectrum as input to the pooling model, the pooling model aggregates these features to obtain aggregated features. It can be understood that aggregated features are also a type of acoustic feature characterizing speech clarity.
[0049] Step S205: Obtain the corresponding evaluation result based on the aggregation feature analysis.
[0050] Specifically, after obtaining the aggregated features, further quantification of these features yields the corresponding evaluation results. These evaluation results can be evaluation scores, evaluation grades, or other data formats that characterize the clarity of the speech signal being processed.
[0051] This embodiment employs the above-described scheme, specifically by performing spectral segmentation on the speech signal to be processed based on a preset window length to obtain several corresponding frames of spectrum; then, by performing frame-by-frame convolution on these frames of spectrum using a preset frame-by-frame convolution model, to obtain first-type high-dimensional features corresponding to each frame of spectrum; finally, by performing time-dependent modeling on these first-type high-dimensional features using a preset temporal model, to obtain second-type high-dimensional features corresponding to each frame of spectrum; and finally, by performing feature aggregation on these second-type high-dimensional features using a preset pooling model, to obtain aggregated features; and finally, by analyzing these aggregated features, a corresponding evaluation result is obtained. In this embodiment, by performing spectral segmentation, convolution, time-dependent modeling, and feature aggregation on the speech signal to be processed corresponding to several sound partitions, aggregated features of the speech signal to be processed corresponding to each sound partition can be obtained. Further quantization of the aggregated features can yield an evaluation result reflecting speech clarity. Thus, based on the evaluation result, the target sound partition can be accurately identified, effectively reducing the occurrence of sound region leakage.
[0052] Furthermore, referring to Figure 5 The third embodiment of the speech processing method in this application provides a flowchart, based on the above. Figure 3 In the illustrated embodiment, step S205, further refining the evaluation results obtained from the aggregation feature analysis, includes: Step S2051: Obtain the corresponding evaluation score based on the aggregated feature analysis, wherein the type of the evaluation score includes at least one of MOS value, noise evaluation score, and human voice evaluation score.
[0053] Specifically, this embodiment uses evaluation scores as the data form of the evaluation results. The types of evaluation scores include at least one of MOS (Mean Opinion Score), noise evaluation score, and human voice evaluation score. It is worth noting that the human voice evaluation score reflects the loudness of human voices.
[0054] Generally speaking, the MOS score is a necessary assessment score type. Based on the MOS score, at least one of the following assessment score types can be combined: noise assessment score and human voice assessment score. For example, two assessment score types can be used: MOS score and noise assessment score; or, MOS score and human voice assessment score; or, MOS score, noise assessment score, and human voice assessment score.
[0055] It is worth noting that the MOS value, noise assessment score, and human voice assessment score all have corresponding score ranges, such as 1 to 5 points. Generally, the higher the score, the better the corresponding assessment performance.
[0056] This embodiment, through the above-described scheme, specifically obtains the corresponding evaluation score based on the aggregated feature analysis. The evaluation score type includes at least one of MOS value, noise evaluation score, and human voice evaluation score. In this embodiment, selectively classifying MOS value, noise evaluation score, and human voice evaluation score into the evaluation score type allows for a more comprehensive assessment of the clarity of the speech signal to be processed, thereby accurately determining the target sound zone and effectively reducing the occurrence of sound zone leakage.
[0057] Furthermore, referring to Figure 6 The fourth embodiment of the speech processing method in this application provides a flowchart, based on the above. Figure 5 In the embodiment shown, step S2051, further refining the corresponding evaluation score obtained from the aggregation feature analysis, includes: Step S2052: Based on the aggregation features and the preset speech quality scoring criteria, the corresponding evaluation score is obtained through analysis.
[0058] Specifically, to ensure that the evaluation score more reasonably reflects the clarity of the speech signal being processed, this embodiment analyzes and obtains the corresponding evaluation score based on aggregation features and a preset speech quality scoring standard. The preset speech quality scoring standard can be a scoring standard based on prior knowledge. For example, the International Telecommunication Union's P.808 speech quality assessment standard (ITU-P.808) can be used as the preset speech quality scoring standard. The scope of the speech quality score can involve multiple dimensions such as MOS value, noise assessment score, and human voice assessment score.
[0059] This embodiment, through the above-described scheme, specifically analyzes and obtains the corresponding evaluation score based on the aggregated features and a preset speech quality scoring standard. This embodiment introduces a speech quality scoring standard into the evaluation process. This standard can be derived from prior knowledge such as international standards, thus enabling the evaluation score to more reasonably represent the clarity of the speech signal to be processed.
[0060] Furthermore, referring to Figure 7 The fifth embodiment of the speech processing method in this application provides a flowchart, based on the above. Figure 5 In the illustrated embodiment, step S30, based on the evaluation results, further refines the target sound partition, including: Step S301: Based on the evaluation score and the preset threshold filtering rules, filter the several sound partitions to determine at least one candidate sound partition.
[0061] Specifically, if the evaluation score for a certain sound partition is extremely poor, it can be determined that the speech signal to be processed for that sound partition is not clear enough and lacks the need for further processing. Therefore, several sound partitions can be screened based on their respective evaluation scores and a preset threshold filtering rule to determine at least one candidate sound partition. The preset threshold filtering rule refers to setting a threshold based on the type of evaluation score. The evaluation score of a certain type is compared with the corresponding threshold; if it is greater than (or greater than or equal to) the corresponding threshold, the evaluation score of that type is considered a valid evaluation score; if it is less than or equal to (or less than) the corresponding threshold, the evaluation score of that type is considered an invalid evaluation score.
[0062] For example, we can preset the threshold for the MOS value to 1.5, the threshold for the noise assessment score to 1, and the threshold for the human voice assessment score to 2. Then, when the MOS value is greater than or equal to 1.5, the noise assessment score is greater than or equal to 1, and the human voice assessment score is greater than or equal to 2, we can determine that the MOS value, noise assessment score, and human voice assessment score are all valid, and further determine the corresponding sound zones as candidate sound zones.
[0063] It is understandable that the candidate sound partitions refer to the sound partitions whose evaluation scores for the corresponding speech signals to be processed are all valid.
[0064] Step S302: Compare the scores of the candidate sound partitions with the preset score comparison rules to determine the target sound partition.
[0065] Specifically, after identifying at least one candidate sound partition, the scores can be further compared according to the evaluation scores corresponding to the candidate sound partitions and the preset score comparison rules. The candidate sound partition with the highest score is then identified as the target sound partition. For the score comparison process, the preferred evaluation score type is the MOS value, because the MOS value can better characterize the clarity of the speech signal to be processed in a candidate sound partition.
[0066] For example, if there are three candidate sound partitions, with the first candidate sound partition having a MOS value of 3, the second candidate sound partition having a MOS value of 2, and the third candidate sound partition having a MOS value of 1, then by comparing the scores, the first candidate sound partition with the highest MOS value can be selected as the target sound partition.
[0067] For example, if there is only one candidate sound partition, that candidate sound partition can be directly selected as the target sound partition.
[0068] This embodiment, through the above-described scheme, specifically filters the plurality of sound partitions according to the evaluation score and a preset threshold filtering rule to determine at least one candidate sound partition; then, it compares the evaluation score corresponding to the candidate sound partition with a preset score comparison rule to determine the target sound partition. In this embodiment, firstly, some sound partitions with lower clarity are filtered out based on the threshold filtering rule, resulting in at least one candidate sound partition. Further, the evaluation score corresponding to each candidate sound partition is compared with the preset score comparison rule, and the candidate sound partition with the higher evaluation score is determined as the target sound partition. In this way, the target sound partition can be accurately determined, effectively reducing the occurrence of sound leakage.
[0069] Furthermore, referring to Figure 8 The sixth embodiment of the speech processing method of this application provides a flowchart, based on the above. Figure 2 In the embodiment shown, before step S20, which involves evaluating the clarity of the speech signal to be processed based on a preset clarity evaluation model and obtaining the corresponding evaluation result, the method further includes: Step S001: Perform speech activity detection on the speech signal to be processed to determine at least one sound zone with speech activity.
[0070] Specifically, Voice activity detection (VAD) is a technique used in speech processing to detect the presence of speech signals. To eliminate interference from invalid speech signals, VAD can be performed on the speech signals to be processed corresponding to several voice regions, identifying at least one voice region with speech activity. It is understood that the speech signal to be processed corresponding to a voice region with speech activity contains the speaker's speech information.
[0071] Step S20: Based on a preset clarity assessment model, the clarity of the speech signal to be processed is assessed, and the corresponding assessment results are further refined, including: Step S206: Based on a preset clarity assessment model, the clarity of the speech signal to be processed corresponding to the speech-active sound partition is assessed to obtain the corresponding assessment result.
[0072] Specifically, after identifying at least one speech-active sound partition, the speech signal to be processed corresponding to the speech-active sound partition can be used as the input of the clarity assessment model. The clarity assessment model can then evaluate the clarity of the speech signal to be processed corresponding to the speech-active sound partition to obtain the corresponding evaluation result.
[0073] This embodiment, through the above-described scheme, specifically identifies at least one voice region with voice activity by performing voice activity detection on the voice signal to be processed. Then, based on a preset clarity assessment model, the clarity of the voice signal to be processed corresponding to the voice region with voice activity is assessed to obtain the corresponding assessment result. In this embodiment, firstly, voice activity detection is performed on the voice signals to be processed corresponding to several voice regions, eliminating some voice regions without voice activity, and identifying at least one voice region with voice activity. Then, based on the preset clarity assessment model, the clarity of the voice signal to be processed corresponding to the voice region with voice activity is assessed to obtain the corresponding assessment result. This reduces the computational burden on the clarity assessment model and improves the accuracy of the assessment results.
[0074] Furthermore, referring to Figure 9 The seventh embodiment of the speech processing method of this application provides a flowchart, based on the above. Figure 2 In the illustrated embodiment, after determining the target sound partition based on the evaluation results in step S30, the method further includes: Step S002: Adjust the preset sound zone allocation strategy according to the target sound zone to obtain the adjusted sound zone allocation strategy, wherein the adjusted sound zone allocation strategy is used to control the voice interaction task corresponding to the target sound zone.
[0075] Specifically, in this embodiment, after determining the target sound partition, the preset sound partition allocation strategy can be adjusted according to the target sound partition to obtain the adjusted sound partition allocation strategy. The adjusted sound partition allocation strategy can control the execution of voice interaction tasks corresponding to the target sound partition.
[0076] Taking an intelligent cockpit system as an example, the target sound zone is determined to be the driver's seat sound zone. The sound zone allocation strategy can then be adjusted based on this zone to obtain the adjusted strategy. Furthermore, if the speaker (in this case, the driver) says "open the window" through voice recognition, the adjusted sound zone allocation strategy can be used to determine that the window the speaker wants to open is the driver's side window, further controlling the corresponding linkage component to open the driver's side window. It can be understood that the aforementioned voice control task regarding opening the window is a voice interaction task.
[0077] This embodiment, through the above-described scheme, specifically adjusts a preset sound zone allocation strategy based on the target sound zone to obtain an adjusted sound zone allocation strategy. The adjusted sound zone allocation strategy is used to control the voice interaction task corresponding to the target sound zone. In this embodiment, based on the determined target sound zone, the sound zone allocation strategy can be further adjusted to control the corresponding voice interaction task. This can be applied to voice interaction scenarios such as intelligent cockpit systems in vehicles, effectively improving the user's voice interaction experience or driving experience.
[0078] Furthermore, embodiments of this application also propose a voice processing device, the voice processing device comprising: The acquisition module is used to acquire the audio signals to be processed corresponding to each of the several sound partitions. The evaluation module is used to evaluate the clarity of the speech signal to be processed based on a preset clarity evaluation model, and obtain the corresponding evaluation results. The determination module is used to determine the target sound partition based on the evaluation results.
[0079] The principle and implementation process of speech processing in this embodiment are explained in the above embodiments and will not be repeated here.
[0080] Furthermore, this application also proposes a terminal device, which includes a memory, a processor, and a voice processing program stored in the memory and executable on the processor. When the voice processing program is executed by the processor, it implements the steps of the voice processing method described above.
[0081] Since this voice processing program employs all the technical solutions of all the foregoing embodiments when executed by the processor, it has at least all the beneficial effects brought about by all the technical solutions of all the foregoing embodiments, which will not be elaborated here.
[0082] Furthermore, embodiments of this application also propose a computer-readable storage medium storing a speech processing program, which, when executed by a processor, implements the steps of the speech processing method described above.
[0083] Since this voice processing program employs all the technical solutions of all the foregoing embodiments when executed by the processor, it has at least all the beneficial effects brought about by all the technical solutions of all the foregoing embodiments, which will not be elaborated here.
[0084] Compared to existing technologies, the speech processing method, apparatus, terminal device, and storage medium proposed in this application acquire speech signals to be processed corresponding to several sound partitions; evaluate the speech signals to be processed based on a preset clarity evaluation model to obtain corresponding evaluation results; and determine the target sound partition based on the evaluation results. Based on this application's solution, the clarity evaluation model can evaluate the clarity of the speech signals to be processed, obtain evaluation results reflecting speech clarity, and further determine the target sound partition based on the evaluation results. This eliminates the dependence on reference speech signals, and the clarity evaluation model can adapt to the dynamic effects of environmental noise and changes in speaker posture on the speech signals to be processed. Based on this, the target sound partition can be accurately determined, effectively reducing the occurrence of sound partition leakage.
[0085] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.
[0086] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0087] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, controlled terminal, or network device, etc.) to execute the methods of each embodiment of this application.
[0088] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
Claims
1. A speech processing method, characterized in that, The speech processing method includes: Acquire the speech signals to be processed corresponding to each of the several sound partitions; The clarity of the speech signal to be processed is evaluated based on a preset clarity evaluation model to obtain the corresponding evaluation results; Based on the evaluation results, the target sound zone is determined; The preset sound zone allocation strategy is adjusted according to the target sound zone to obtain the adjusted sound zone allocation strategy, wherein the adjusted sound zone allocation strategy is used to control the voice interaction task corresponding to the target sound zone. The step of evaluating the clarity of the speech signal to be processed based on a preset clarity evaluation model to obtain the corresponding evaluation result includes: The speech signal to be processed is spectrum segmented based on a preset window length to obtain several corresponding frames of spectrum. The pre-defined frame-by-frame convolution model is used to convolve the spectrum of the several frames one by one to obtain the first type of high-dimensional features corresponding to each of the several frames of spectrum. The first type of high-dimensional features are modeled in a time-dependent manner using a preset time-series model to obtain the second type of high-dimensional features corresponding to each of the several frame spectra. The second type of high-dimensional features corresponding to the spectrum of each of the several frames are aggregated using a preset pooling model to obtain aggregated features; The corresponding evaluation results are obtained based on the aggregation feature analysis.
2. The speech processing method as described in claim 1, characterized in that, The step of obtaining the corresponding evaluation result based on the aggregation feature analysis includes: The corresponding evaluation score is obtained based on the aggregated feature analysis, wherein the type of the evaluation score includes at least one of MOS value, noise evaluation score, and human voice evaluation score.
3. The speech processing method as described in claim 2, characterized in that, The step of obtaining the corresponding evaluation score based on the aggregation feature analysis includes: Based on the aggregated features and the preset speech quality scoring criteria, the corresponding evaluation score is obtained through analysis.
4. The speech processing method as described in claim 2, characterized in that, The step of determining the target sound partition based on the evaluation results includes: Based on the evaluation score and the preset threshold filtering rules, the several sound partitions are filtered to determine at least one candidate sound partition; The target sound partition is determined by comparing the evaluation score corresponding to the candidate sound partition with the preset score comparison rules.
5. The speech processing method as described in claim 1, characterized in that, Before the step of evaluating the clarity of the speech signal to be processed based on a preset clarity evaluation model and obtaining the corresponding evaluation result, the method further includes: The speech signal to be processed is subjected to speech activity detection to determine at least one sound zone with speech activity; The step of evaluating the clarity of the speech signal to be processed based on a preset clarity evaluation model to obtain the corresponding evaluation result includes: The clarity of the speech signal to be processed corresponding to the speech-active sound partition is evaluated based on the preset clarity evaluation model, and the corresponding evaluation results are obtained.
6. A voice processing device, characterized in that, The voice processing device includes: The acquisition module is used to acquire the audio signals to be processed corresponding to each of the several sound partitions. The evaluation module is used to evaluate the clarity of the speech signal to be processed based on a preset clarity evaluation model, and obtain the corresponding evaluation results. The determination module is used to determine the target sound partition based on the evaluation results; The evaluation module is further configured to: The speech signal to be processed is spectrum segmented based on a preset window length to obtain several corresponding frames of spectrum. The pre-defined frame-by-frame convolution model is used to convolve the spectrum of the several frames one by one to obtain the first type of high-dimensional features corresponding to each of the several frames of spectrum. The first type of high-dimensional features are modeled in a time-dependent manner using a preset time-series model to obtain the second type of high-dimensional features corresponding to each of the several frame spectra. The second type of high-dimensional features corresponding to the spectrum of each of the several frames are aggregated using a preset pooling model to obtain aggregated features; The corresponding evaluation results are obtained based on the aggregation feature analysis. The voice processing device is further configured to: adjust a preset voice region allocation strategy according to the target voice region to obtain an adjusted voice region allocation strategy, wherein the adjusted voice region allocation strategy is used to control the voice interaction task corresponding to the target voice region.
7. A terminal device, characterized in that, The terminal device includes a memory, a processor, and a voice processing program stored in the memory and executable on the processor. When the voice processing program is executed by the processor, it implements the steps of the voice processing method as described in any one of claims 1-5.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a speech processing program, which, when executed by a processor, implements the steps of the speech processing method as described in any one of claims 1-5.
Citation Information
Patent Citations
Recording data identification method and device, and recording equipment
CN112509597A