Method, device and system for evaluating voice quality for telecommunication

By calculating and displaying voice quality in real time during remote communication, the problem of difficulty in providing real-time feedback on voice quality in existing technologies is solved, enabling accurate evaluation and feedback of voice quality and improving communication effectiveness.

CN115512720BActive Publication Date: 2025-12-23ZHONGKE YOUSHENG (SUZHOU) TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202211115357.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-14
Publication Date
2025-12-23
Estimated Expiration
2042-09-14

AI Technical Summary

Technical Problem

Existing technologies are difficult to apply in long-distance communication, wireless communication, and real-time, accurate feedback of voice quality, leading to communication problems.

Method used

By acquiring the signal of the first voice terminal in the target communication channel, and based on the voice quality evaluation device, the voice quality evaluation system, the voice quality evaluation device, the voice quality evaluation system, the voice technology, the voice quality evaluation device, the voice quality evaluation device, the voice quality evaluation device, the voice quality evaluation device, the voice quality evaluation device, the voice quality evaluation result, and the voice quality evaluation result are calculated and displayed in real time.

Benefits of technology

It enables real-time and accurate evaluation and feedback of voice quality in remote communication, thereby improving communication effectiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115512720B_ABST
    Figure CN115512720B_ABST
Patent Text Reader

Abstract

The application discloses a voice quality evaluation method, device and system for remote communication. The evaluation method comprises the following steps: obtaining a first target voice signal corresponding to a first signal source received by a first voice terminal of a target communication channel, wherein the first signal source corresponds to a second voice terminal in the target communication channel; obtaining a corresponding first voice quality evaluation result in real time based on the first target voice signal; and performing front-end display of the first voice quality evaluation result on the second voice terminal. The voice quality evaluation method for remote communication can test and feed back the voice quality received by a listening terminal in wireless remote communication in real time and accurately, thereby effectively improving the voice communication effect.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the evaluation technology of voice quality in remote communication, in particular to a voice quality evaluation method, device and system for remote communication. BACKGROUND

[0002] In remote communication, from the source (the sound emitting end), the voice signal passes through all the encoding and modulation processes, via the transmitter and the wireless channel, to the receiver containing all the signal processing functions, and finally ends at the sink (the sound receiving end). The voice quality is closely related to each step. Of course, in the process of wireless voice communication, the environment factors such as reverberation and noise of the place where the source is located, and the human factors such as the distance between the speaker and the microphone, all have an impact on the voice quality received by the sink. The source end cannot know the voice quality received by the other party in real time, which causes communication difficulties.

[0003] Therefore, we need to find a method to obtain real-time feedback of the voice quality in remote communication.

[0004] SUMMARY

[0005] The present application aims to provide a voice quality evaluation method, device and system for remote communication, which can evaluate the voice quality in voice communication in real time and accurately.

[0006] To achieve the above-mentioned application purposes, the present application proposes the following technical solutions:

[0007] In a first aspect, a voice quality evaluation method for remote communication is provided, and the evaluation method comprises:

[0008] Obtaining a first target voice signal corresponding to a first source received by a first voice terminal of a target communication channel, the first source corresponding to a second voice terminal in the target communication channel;

[0009] Real-time calculation of a corresponding first voice quality evaluation result based on the first target voice signal;

[0010] Front-end display of the first voice quality evaluation result at the second voice terminal.

[0011] In a preferred embodiment, the method further comprises:

[0012] Obtaining a second target voice signal corresponding to a first source received by a first voice terminal of a target communication channel;

[0013] Real-time calculation of a corresponding second voice quality evaluation result based on the second target voice signal;

[0014] The second speech quality evaluation result is displayed in front of the second speech terminal.

[0015] In a preferred embodiment, the speech quality evaluation result is calculated in real time based on the target speech signal, comprising:

[0016] The target speech signal is extracted to obtain at least one group of target speech signal features;

[0017] The at least one group of target speech signal features is inputted into a pre-trained speech quality model to obtain a corresponding speech quality evaluation result.

[0018] In a preferred embodiment, the first speech quality evaluation result includes, but is not limited to, one of the speech transmission index or the mean opinion score.

[0019] In a preferred embodiment, when the first speech quality evaluation result includes the speech transmission index, the first speech quality evaluation result is calculated in real time based on the first target speech signal, comprising:

[0020] The first target speech signal is extracted according to p different octave filter signal bands to obtain p groups of target features, p≥2;

[0021] The first target position corresponding to the first target speech signal is obtained based on the p groups of target features.

[0022] In a preferred embodiment, the first target speech signal is extracted according to p different octave filter signal bands to obtain p groups of target features, comprising:

[0023] The first target speech signal is filtered to obtain p groups of different octave filter signal bands, and any group of the octave filter signal bands includes n modulation frequencies f m , n≥1, m≥1;

[0024] The p groups of different octave filter signal bands are envelope extracted to obtain p groups of sub-band envelope features;

[0025] Any group of the sub-band envelope features corresponding to the octave filter signal band is inputted into a pre-trained speech quality model corresponding to the corresponding octave filter signal band to obtain p groups of reverberation time T;

[0026] The p groups of the reverberation time T are based on to obtain p groups of target features, and the target features are modulation transfer function values.

[0027] In a preferred embodiment, the step of extracting the envelope features of the p groups of different octave band filtered signal bands to obtain the envelope features of the p groups of sub-bands includes:

[0028] The envelope characteristics of the p groups of sub-bands are obtained by performing half-wave envelope detection on the different octave bands of the p groups of filtered signals.

[0029] In a preferred embodiment, the step of taking any set of octave-band filtered signal bands corresponding to the envelope features of the sub-bands as input, and obtaining p sets of reverberation times T through a pre-trained speech quality model corresponding to the corresponding octave-band filtered signal bands, includes:

[0030] Divide any of the octave band filtered signals into N consecutive speech segments of equal duration, where N≥2;

[0031] For any of the N speech segments included in any octave band filtered signal band, feature extraction is performed on any speech segment using a combination structure of one or more of the following: convolutional neural network, linear connection layer, activation layer, and normalization layer, to obtain a matrix of shape [P,Q], thereby obtaining the corresponding N speech segment features;

[0032] Any speech segment feature among the N obtained speech segment features is interacted through a combination of one or more of the following structures: Long Short-Term Memory module, multi-head / single-head attention module, linear connection layer, activation layer, and normalization layer, to obtain the corresponding speech segment interaction feature;

[0033] Based on the obtained N speech segment interaction features, N reverberation times T corresponding to the N speech segment interaction features are predicted by a linear regression layer or a classification layer, respectively. N ;

[0034] For N reverberation times T corresponding to any octave band filtered signal band N The average values ​​are taken separately to obtain the p-group reverberation time T corresponding to the respective octave band filtered signal bands.

[0035] In a preferred embodiment, obtaining the corresponding p groups of target features based on the reverberation times T of the p groups includes:

[0036] Based on any modulation frequency f m The value and corresponding reverberation time T are used to obtain any modulation frequency f of any octave band filtered signal. m The modulation transfer function value m k,fm .

[0037] In a preferred embodiment, obtaining a first speech quality evaluation result based on the p sets of target features corresponding to the first target location in the first target speech signal includes:

[0038] Based on any of the modulation transfer function values ​​m k,fm Obtain any modulation frequency f of the corresponding octave-band filtered signal band k. m Effective signal-to-noise ratio (SNR) effk,fm ;

[0039] Based on any of the aforementioned signal-to-noise ratios (SNR) effk,fm Obtain any modulation frequency f of the corresponding octave-band filtered signal band k. m Transmission index TI at the location k,fm ;

[0040] Calculate the n transmission indices TI for any octave-band filtered signal band k. k,fm The mean value is used to obtain the modulation transfer index M of the corresponding octave band filtered signal band k. k ;

[0041] Based on the modulation transfer index M of p octave band filtered signal bands k The first speech quality evaluation result corresponding to the first target position is calculated.

[0042] In a preferred embodiment, the evaluation method further includes pre-training p speech quality models corresponding to the p different octave band filtered signal bands, including:

[0043] Based on any speech signal sample in the speech signal sample set, obtain p different octave band filtered signal band sample sets. Each octave band filtered signal band sample set includes q modulation frequency samples and corresponding q impulse response samples. Each impulse response sample includes a reverberation time sample T0, where q ≥ 2.

[0044] Using the q modulation frequency samples as input and the corresponding q reverberation time samples T0 as output, p speech quality models corresponding to p octave band filtered signal bands are obtained by training based on a neural network.

[0045] In a preferred embodiment, the voice quality evaluation results are displayed on the front end, and the display methods include, but are not limited to:

[0046] The voice quality evaluation results are displayed on the interface in the form of numerical values ​​and dynamic movement signals; or...

[0047] The voice quality evaluation results are displayed on the interface as numerical values ​​and dynamic Wi-Fi signals; or...

[0048] The voice quality evaluation results are displayed on the interface in the form of numerical values ​​and a dynamic dashboard; or...

[0049] The voice quality evaluation result is displayed in the form of a numerical value and a dynamic attitude bar.

[0050] In a second aspect, a voice quality evaluation device for remote communication is provided, which comprises: an acquisition module configured to acquire a first target voice signal corresponding to a first source received by a first voice terminal of a target communication channel, the first source corresponding to a second voice terminal in the target communication channel;

[0051] a processing module configured to calculate a first voice quality evaluation result in real time based on the first target voice signal;

[0052] a display module configured to display the first voice quality evaluation result in front of the second voice terminal.

[0053] In a third aspect, a voice quality evaluation system for remote communication is provided, which comprises:

[0054] at least two voice terminals, including a first voice terminal and a second voice terminal, the target communication channel being formed between the first voice terminal and the second voice terminal, and the second voice terminal corresponding to a first source;

[0055] an intelligent device configured to acquire a first target voice signal corresponding to a first source received by a first voice terminal of a target communication channel, the first source corresponding to a second voice terminal in the target communication channel other than the first voice terminal, and perform the operation of any one of the first aspect in real time based on the first target voice signal to calculate a corresponding first voice quality evaluation result, and send the first voice quality evaluation result to the second voice terminal;

[0056] at least one display device, including a first display device arranged in the second voice terminal, and the first display device displays the first voice quality evaluation result in front of the second voice terminal.

[0057] In a fourth aspect, an electronic device is provided, which comprises:

[0058] one or more processors; and

[0059] a memory associated with the one or more processors, the memory being configured to store program instructions, which, when executed by the one or more processors, perform the operation of any one of the first aspect.

[0060] In a fifth aspect, a computer readable storage medium is provided, which stores a computer program, wherein the program, when executed by a processor, implements the method of any one of the first aspect.

[0061] Compared with the prior art, the present application has the following beneficial effects:

[0062] The present application provides a voice quality evaluation method, device and system for remote communication, the evaluation method comprising obtaining a first target voice signal corresponding to a first source received by a first voice terminal of a target communication channel, the first source corresponding to a second voice terminal in the target communication channel; obtaining a corresponding first voice quality evaluation result in real time based on the first target voice signal; and performing front-end display of the first voice quality evaluation result at the second voice terminal. The voice quality evaluation method for remote communication in the present application can realize real-time testing and real-time display of the voice quality received by the receiving end in a wireless communication environment, and the voice quality evaluation result is accurate and efficient. BRIEF DESCRIPTION OF DRAWINGS

[0063] Figure 1 is a flowchart of the voice quality evaluation method for remote communication in the present embodiment;

[0064] Figure 2 is a wireless remote communication structure and communication principle explanation;

[0065] Figure 3 is an exemplary remote communication structure diagram in the present embodiment;

[0066] Figure 4 is a schematic diagram of envelope extraction to obtain envelope boundary in the present embodiment;

[0067] Figure 5 is a half-wave envelope detection circuit diagram in the present embodiment;

[0068] Figure 6 is a structure schematic diagram of the neural network in the present embodiment;

[0069] Figure 7 is an STI score meaning explanation;

[0070] Figures 8a-8d is exemplary display content when the voice quality evaluation result is displayed on an interface. DETAILED DESCRIPTION

[0071] In order to make the purpose, technical scheme and advantages of the present application clearer, the technical scheme in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0072] In the description of the present application, it should be understood that the terms "first", "second" are only for the purpose of description, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of indicated technical features. Therefore, the features defined as "first", "second" can be explicitly or implicitly included one or more of the features. In the description of the present application, unless otherwise stated, the meaning of "a plurality of" is two or more.

[0073] In view of the current remote communication factors affecting the quality of voice, it is difficult to comprehensively test the quality of voice in real time, and the current situation, the present embodiment provides a simple and fast and can real-time feedback of voice quality evaluation method. In the following, the voice quality evaluation method, device and system for remote communication will be further described in detail.

[0074] Embodiment

[0075] As Figure 1 shown, the present embodiment provides a voice quality evaluation method for remote communication, which is suitable for evaluating the intelligibility of language during voice communication through remote communication. As Figure 2 shown, during the process of voice transmission through remote communication, the source is transmitted through channel after source coding and pulse modulation, and then reaches the receiving end through demodulation and decoding process, thereby completing the voice transmission. The voice quality evaluation method for remote communication provided by the present embodiment can evaluate the voice quality of the remote communication in real time and accurately, and perform real-time front-end display to improve the voice communication effect of remote communication.

[0076] The present embodiment includes but is not limited to two or more end-to-end communication scenarios. For the sake of description, the present embodiment will be further described in detail taking two end-to-end communication scenarios as an example. It can be understood that in voice communication, two voice terminals are simultaneously the sound emitting end and the receiving end.

[0077] As an example, as Figure 3 shown, the wireless communication channel required for voice quality evaluation is taken as the target communication channel, which includes the first voice terminal and the second voice terminal in the voice communication state. The first voice terminal or the second voice terminal in the present embodiment is the sound emitting end and the receiving end, the second voice terminal corresponds to the first source, and the first voice terminal corresponds to the second source. It can be understood that when the first source of the second voice terminal emits sound, the first target voice signal is transmitted to the first voice terminal, and the first voice quality evaluation result is obtained. Correspondingly, the same sound emits the second target voice signal in the local second voice terminal, and the second voice transmission evaluation result is obtained. Of course, the second source of the first terminal emits sound, which is similar to the first source of the second voice terminal.

[0078] The embodiment will take the first source sound of the second voice terminal as an example to further specifically describe the voice quality evaluation method for remote communication, which comprises the following steps:

[0079] S1, obtaining the first target voice signal corresponding to the first source received by the first voice terminal of the target communication channel. As described above, the first source corresponds to the second voice terminal in the target communication channel which is in a communication connection state with the first voice terminal.

[0080] In the specific implementation process, the background server or the cloud server can obtain the source signal data in the plurality of wireless channels in real time. The target voice signal sent by the target source (sound emitting end) in the target communication channel at the receiving end is obtained.

[0081] S2, obtaining the corresponding first voice quality evaluation result based on the first target voice signal in real time. Specifically, step S2 comprises:

[0082] S21, performing feature extraction on the target voice signal to obtain at least one group of target voice signals;

[0083] S22, taking at least one group of target voice signals as input, and obtaining the corresponding voice quality evaluation result through the pre-trained voice quality model.

[0084] It should be noted that the method used in the voice quality evaluation of the embodiment includes but is not limited to one of the speech transmission index (STI) or the mean opinion score (MOS). In order to facilitate the description, the embodiment uses STI to evaluate the voice quality.

[0085] In the specific implementation process, step S21 is specifically: performing feature extraction on the first target voice signal according to p different octave filter signal bands k to obtain p groups of target features, p≥2, 1≤k≤p.

[0086] Specifically, step S21 comprises:

[0087] S21a, filtering the first target voice signal to obtain p groups of different octave filter signal bands k, any group of octave filter signal bands k comprises n modulation frequencies f m , n≥1, m≥1.

[0088] It should be noted that human voice is usually divided into seven frequency bands, so p=7 is preferred in the embodiment, i.e. 1≤k≤7. Therefore, the target voice is filtered to obtain the center frequency f cS21a, the input is a set of p different octave filter signal bands k, each octave filter signal band k has an upper limit frequency f and a lower limit frequency f, and each octave filter signal band k has a corresponding speech quality model, and each speech quality model is pre-trained, and each speech quality model includes a data preprocessing module, a feature extraction module, a time interaction module, and a prediction module. u and a lower limit frequency f l respectively as shown in the following formulas (1) and (2):

[0089]

[0090]

[0091] S21b, envelope extraction is performed on the p different octave filter signal bands k respectively to obtain p sets of sub-band envelope features, and the envelope extraction result is an envelope boundary as shown in Figure 4 .

[0092] The envelope extraction algorithm is not limited in this embodiment, and preferably, p sets of different octave filter signal bands k are respectively subjected to half-wave envelope detection to obtain p sets of sub-band envelope features (as shown in Figure 5 ), and the differential equation is expressed as:

[0093]

[0094]

[0095] S21c, taking the octave filter signal band k corresponding to any set of sub-band envelope features as input, p sets of reverberation times T are obtained respectively through the pre-trained speech quality model corresponding to the corresponding octave filter signal band k.

[0096] Any speech quality model includes a data preprocessing module, a feature extraction module, a time interaction module, and a prediction module connected in sequence. Specifically:

[0097] The data preprocessing module: any octave filter signal band k is divided into continuous N speech slices (such as Figure 6 ) with a preset time length (such as x seconds), N≥2.

[0098] The feature extraction module: any speech slice in the N speech slices included in any octave filter signal band k is subjected to feature extraction by a combination structure of one or more of a convolutional neural network, a linear connection layer, an activation layer, and a normalization layer to obtain a matrix of shape [P, Q], and obtain corresponding N speech slice features, such as Figure 6 speech slice features.

[0099] The time interaction module: any of the obtained N speech segment features is interacted through a combination structure of one or more of a long short-term memory module, a multi-head / single-head attention module, a linear connection layer, an activation layer, and a normalization layer to obtain a corresponding speech segment interaction feature, such as Figure 6 The slice interaction feature.

[0100] The prediction module: based on the obtained N speech segment interaction features, a linear regression layer or a classification layer is used for prediction to obtain N reverberation times T N , respectively, corresponding to the N speech segment interaction features. Figure 6 The slice reverberation time T N . N T N is obtained by averaging to obtain the reverberation time T corresponding to the corresponding octave filter signal band k. The method of averaging includes but is not limited to any one of a simple average method, a weighted average method, or a harmonic average method.

[0101] To this end, before step S21c, the evaluation method further includes: Sa, p speech quality models corresponding to p different octave filter signal bands k are trained in advance, including:

[0102] Sa1, based on any speech signal sample in the speech signal sample set, a corresponding p set of different octave filter signal band sample sets are obtained, any set of octave filter signal band sample sets includes q modulation frequency samples and corresponding q impulse response samples, any impulse response sample includes a reverberation time sample T0, q≥2;

[0103] Sa2, taking q modulation frequency samples as input and corresponding q reverberation time samples T0 as output, p speech quality models corresponding to p octave filter signal bands are trained based on a neural network.

[0104] S21d, based on the p sets of reverberation times T, corresponding p sets of target features are obtained, and the target feature is a modulation transfer function value.

[0105] The modulation transfer function (Modulation Transfer Function, MTF) describes the degree of transmission of modulation m from a target object (sound source) to a receiving sensor, and is a function of modulation frequency f m m k,fm , and MTF determines the modulation reduction degree of the first target speech signal. Specifically, the range of modulation frequency f m is 0.63 to 12.5 Hz. Therefore, the MTF function value depends on the system environment characteristics and background noise. The MTF calculation process is shown in the following formula (5):

[0106]

[0107] S22 specifically is: obtaining the first speech quality evaluation result corresponding to the first target speech signal at the current position based on the p target characteristics.

[0108] This step S3 includes:

[0109] S22a, based on any of the modulation transfer function values m k,fm The effective signal-to-noise ratio SNR m of any modulation frequency f effk,fm of the corresponding octave filter signal band k is obtained; specifically, the effective signal-to-noise ratio SNR effk,fm is obtained by the following formula (6):

[0110]

[0111] S22b, based on any effective signal-to-noise ratio SNR effk,fm The transmission index TI m at the corresponding modulation frequency f k,fm of the octave filter signal band k is obtained; specifically, the transmission index TI k,fm is obtained by the following formula (7):

[0112]

[0113] S22c, the mean value of the n transmission indexes TI k,fm of any octave filter signal band k is calculated to obtain the modulation transfer index M k of the corresponding octave filter signal band k; the value range of the modulation transfer index M k is -15dB~+15dB. Specifically, the modulation transfer index M k is obtained by the following formula (8):

[0114]

[0115] S22d, based on the modulation transfer indexes M k of the p octave filter signal bands, the first speech quality evaluation result (STI) corresponding to the first target speech signal at the current position is calculated; specifically, the STI is obtained by the following formula (9):

[0116]

[0117] Wherein, α k represents the different gender weight factor of the octave filter signal band k;

[0118] β kdenotes the different gender redundancy factor between octave band k and octave band k+1 of the octave filtered signal;

[0119] M k denotes the modulation transfer index of the octave band of the octave filtered signal.

[0120] Preferably, after obtaining the STI score result, the current speech quality level is determined according to the STI level division standard as shown in Figure 7 The speech quality level can be bad, mediocre, excellent, etc. as shown in Figure 7

[0121] It should be noted that the STI method can distinguish between male and female voice signals, but in practice, in order to simplify the measurement process, only male voice is used to evaluate the speech transmission path. Table 1 gives the male STI weight factor a and the redundancy factor β as a function of the octave band.

[0122] Table 1

[0123]

[0124] S3, displaying the first speech quality evaluation result on the second voice terminal.

[0125] As shown in Figures 8a-8d The interface display method for front-end display of the speech quality evaluation result includes but is not limited to numerical and dynamic moving signal interface display of the speech quality evaluation result; or numerical and dynamic wifi signal interface display of the speech quality evaluation result; or numerical and dynamic dashboard interface display of the speech quality evaluation result; or numerical and dynamic bar interface display of the speech quality evaluation result.

[0126] The embodiment displays the first speech quality evaluation result in real time through the display screen integrated in the second voice terminal, or the display screen provided in the first voice terminal and connected thereto, or the display screen provided independently in the background. The display screen integrated in the second voice terminal can realize that the user knows the speech quality evaluation result, i.e. the language intelligibility evaluation result, of the voice transmitted from the second voice terminal to the first voice terminal through the target communication channel while speaking in the second voice terminal.

[0127] In the embodiment, the second voice terminal displays the STI score in the configured display interface to know the language intelligibility feedback of the first voice terminal to the second voice terminal, so that the user of the second voice terminal can adjust the influencing factors such as the distance between the mouth and the microphone according to the feedback in real time, to adjust the room reverberation, noise, etc. to improve the speech quality.

[0128] ​As preferred, the voice quality evaluation method for remote communication further comprises:

[0129] The second speech terminal receiving the second target speech signal corresponding to the first source on the target communication channel acquires the second target speech signal corresponding to the first source; the second target speech signal is used to calculate the corresponding second voice quality evaluation result in real time; and the second voice quality evaluation result is displayed in front of the second speech terminal.

[0130] Finally, the first speech quality evaluation result and the first speech quality evaluation result are used to comprehensively evaluate and feed back the transmission effect of the first source in the target communication channel.

[0131] The above, based on the second target speech signal, the specific technical scheme for calculating the corresponding second voice quality evaluation result in real time is described based on the first target speech signal.

[0132] In summary, the voice quality evaluation method for remote communication provided by the embodiment can realize real-time evaluation and real-time display of voice quality in a remote communication call scenario. Furthermore, the method obtains modulation transfer function values by dividing different octave band filter signals and finally obtains voice quality evaluation scores based on all modulation transfer function values. In combination with the application of machine learning, especially the voice quality model, the accuracy of voice quality evaluation is effectively improved. Moreover, the final voice quality evaluation can be obtained based on the target speech signal in the channel, without the need to evaluate the sound environment and channel parameters, effectively reducing the data processing amount and data processing speed, improving efficiency and effectively reducing costs.

[0133] Corresponding to the voice quality evaluation method for remote communication, the embodiment further provides a voice quality evaluation device for remote communication, which realizes the method through various functional modules. The voice quality evaluation device for remote communication comprises:

[0134] The acquisition module is configured to acquire a first target speech signal corresponding to a first source received by a first speech terminal on a target communication channel, wherein the first source corresponds to a second speech terminal in the target communication channel.

[0135] The processing module is configured to calculate a corresponding first voice quality evaluation result in real time based on the first target speech signal.

[0136] The display module is configured to display the first voice quality evaluation result in front of the second speech terminal.

[0137] The model training module is configured to pre-train p voice quality models corresponding to p different octave band filter signal bands, respectively.

[0138] As preferred, the acquisition module is further configured to acquire a second target voice signal corresponding to the first source and received by a second voice terminal of the target communication channel; the processing module is further configured to calculate a second voice quality evaluation result in real time based on the second target voice signal; and the display module is further configured to display the second voice quality evaluation result on the second voice terminal.

[0139] Further, the processing module comprises:

[0140] a feature extraction sub-module configured to extract features from the target voice signal to obtain at least one group of target voice signals, and specifically configured to extract features from the first target voice signal according to p different octave filter signal bands to obtain p groups of target features, where p≥2.

[0141] an evaluation sub-module configured to take the at least one group of target voice signals as input, and obtain a corresponding voice quality evaluation result through a pre-trained voice quality model, and specifically configured to take the p groups of target features as input, and obtain a first voice quality evaluation result corresponding to the first target voice signal at a current position through a pre-built voice quality evaluation model.

[0142] Specifically, the feature extraction sub-module comprises:

[0143] a first extraction unit configured to filter the first target voice signal to obtain p groups of different octave filter signal bands, and any group of the octave filter signal bands comprises n modulation frequencies f m , n≥1, m≥1;

[0144] a second extraction unit configured to extract envelopes from the p groups of different octave filter signal bands to obtain p groups of sub-band envelope features;

[0145] a third extraction unit configured to take any group of the octave filter signal bands corresponding to the sub-band envelope features as input, and obtain p groups of reverberation times T through a pre-trained voice quality model corresponding to the corresponding octave filter signal band;

[0146] a fourth extraction unit configured to obtain p groups of target features corresponding to the p groups of reverberation times T, respectively, and the target features are modulation transfer function values.

[0147] The second extraction unit is specifically configured to extract half-wave envelope from the p groups of different octave filter signal bands to obtain p groups of sub-band envelope features.

[0148] The third extraction unit is configured to:

[0149] divide any group of the octave filter signal bands into N voice segments with continuous and equal time lengths, where N≥2.

[0150] Any of the N speech segments included in any of the octave filter signal bands is subjected to feature extraction by a combination structure of one or more of convolutional neural network, linear connection layer, activation layer, normalization layer to obtain a matrix of [P, Q] shape, and N speech segment features are obtained accordingly;

[0151] Any of the N speech segment features obtained is interacted by a combination structure of one or more of long short-term memory module, multi / single attention module, linear connection layer, activation layer, normalization layer to obtain corresponding speech segment interaction features;

[0152] Based on the N speech segment interaction features obtained, linear regression layer or classification layer is used for prediction respectively to obtain N reverberation times T N corresponding to N speech segment interaction features respectively.

[0153] N reverberation times T N corresponding to any octave filter signal band are respectively averaged to obtain p groups of reverberation times T corresponding to the corresponding octave filter signal bands respectively.

[0154] The fourth extraction unit is specifically configured to: based on any modulation frequency f m and the corresponding reverberation time T, obtain the modulation transfer function value m m of any modulation frequency f k,fm of any group of octave filter signal bands.

[0155] The device further comprises a training module, and the training module comprises:

[0156] The sample acquisition unit obtains p groups of different octave filter signal band sample sets based on any speech signal sample in the speech signal sample set, and any group of the octave filter signal band sample set comprises q modulation frequency samples and corresponding q impulse response samples, any of the impulse response samples comprises a reverberation time sample T0, and q≥2.

[0157] The training unit is configured to take the q modulation frequency samples as input and the corresponding q reverberation time samples T0 as output, and train p speech quality models corresponding to p octave filter signal bands based on a neural network.

[0158] Further, the evaluation submodule comprises:

[0159] The first processing unit is configured to obtain an effective signal-to-noise ratio SNR of any modulation frequency f k,fm of the corresponding octave filter signal band k based on any modulation transfer function value m m .effk,fm ;

[0160] a second processing unit, configured to calculate a transmission index TI of any modulation frequency f at a corresponding octave filter signal band k based on any of the signal-to-noise ratios SNR effk,fm m k,fm ;

[0161] a third processing unit, configured to calculate a mean value of n transmission indexes TI of any octave filter signal band k to obtain a modulation transfer index M of the corresponding octave filter signal band k k,fm k ;

[0162] a fourth processing unit, configured to calculate a first speech quality evaluation result corresponding to the first target speech signal at a current position based on the modulation transfer indexes M of p octave filter signal bands k

[0163]

[0164] wherein, a k represents a different gender weight factor of the octave filter signal band k;

[0165] β k represents a different gender redundancy factor between the octave filter signal band k and the octave filter signal band k+1;

[0166] M k represents a modulation transfer index of the octave filter signal band.

[0167] It should be noted that the speech quality evaluation device for remote communication provided in the above embodiment is only used for example to divide the above-mentioned functional modules when performing the speech quality evaluation service. In actual application, the above-mentioned functions can be completed by different functional modules according to the needs, that is, the internal structure of the system is divided into different functional modules to complete all or part of the functions described above. In addition, the speech quality evaluation device for remote communication provided in the above embodiment and the speech quality evaluation method for remote communication belong to the same concept, that is, the device is based on the method, and the specific implementation process is described in detail in the method embodiment, which will not be repeated here.

[0168] In addition, the embodiment also provides a speech quality evaluation system for remote communication, which comprises:

[0169] ​​​​At least two voice terminals, the at least two voice terminals comprising a first voice terminal and a second voice terminal, a target communication channel being formed between the first voice terminal and the second voice terminal, the second voice terminal corresponding to a first signal source;

[0170] An intelligent device, the intelligent device obtaining a first target voice signal corresponding to the first signal source received by a first voice terminal of a target communication channel, the first signal source corresponding to a second voice terminal of the target communication channel other than the first voice terminal; performing real-time calculation to obtain a corresponding first voice quality evaluation result based on the first target voice signal according to any one of the voice quality evaluation methods, and sending the first voice quality evaluation result to the second voice terminal;

[0171] At least one display device, the at least one display device comprising a first display device arranged at the second voice terminal, the first display device performing front-end display of the first voice quality evaluation result.

[0172] In addition, the embodiment also provides an electronic device, comprising:

[0173] One or more processors; and

[0174] A memory associated with the one or more processors, the memory being configured to store program instructions, the program instructions being configured to perform operations as described in any one of the voice quality evaluation methods for remote communication when executed by the one or more processors.

[0175] The specific execution details and corresponding benefits of the voice quality evaluation method for remote communication performed by the program instructions are consistent with the description in the foregoing method, which will not be repeated here.

[0176] In addition, the embodiment also provides a computer-readable storage medium having a computer program stored thereon, wherein the program is executed by a processor to implement the method as described in any one of the voice quality evaluation methods for remote communication.

[0177] All the optional technical solutions described above can be combined in any way to form optional embodiments of the present application, that is, any multiple embodiments can be combined to meet the needs of different application scenarios, which are all within the protection scope of the present application, and will not be repeated here.

[0178] It should be noted that the above description is only a preferred embodiment of the present application, and is not intended to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. A method for voice quality assessment for telecommunication, characterized by, The evaluation method comprises: obtaining a first target voice signal corresponding to a first source received by a first voice terminal of a target communication channel, the first source corresponding to a second voice terminal in the target communication channel; real-time calculation of a corresponding first voice quality evaluation result based on the first target voice signal, real-time calculation of a corresponding voice quality evaluation result based on a target voice signal, comprising: performing feature extraction on the target voice signal to obtain at least one group of target voice signals; inputting the at least one group of target voice signals into a pre-trained voice quality model to obtain a corresponding voice quality evaluation result; when the first voice quality evaluation result includes a voice transmission index, the real-time calculation of the corresponding first voice quality evaluation result based on the first target voice signal comprises: performing feature extraction on the first target voice signal according to p different octave filter signal bands to obtain p groups of target features, p≥2; obtaining a first voice quality evaluation result corresponding to the first target voice signal based on the p groups of target features; The evaluation method further comprises pre-training p voice quality models corresponding to the p different octave filter signal bands, comprising: obtaining a corresponding p group of different octave filter signal band sample sets based on any voice signal sample in a voice signal sample set, any group of the octave filter signal band sample set comprising q modulation frequency samples and corresponding q impulse response samples, any of the impulse response samples comprising a reverberation time sample T0, q≥2; inputting the q modulation frequency samples into the corresponding q reverberation time samples T0 as output, and training p voice quality models corresponding to p octave filter signal bands based on a neural network; The first voice quality evaluation result is displayed in the front end of the second voice terminal.

2. The evaluation method according to claim 1, characterized by, The method further comprises: obtaining a second target voice signal corresponding to a first source received by a second voice terminal of the target communication channel; real-time calculation of a corresponding second voice quality evaluation result based on the second target voice signal; The second voice quality evaluation result is displayed in the front end of the second voice terminal.

3. The evaluation method according to claim 1, characterized by, The first voice quality evaluation result includes but is not limited to one of the voice transmission index or the mean opinion score.

4. The evaluation method according to claim 1, characterized by, The feature extraction on the first target voice signal according to p different octave filter signal bands to obtain p groups of target features comprises: The first target voice signal is filtered to obtain p groups of different octave filter signal bands, any group of the octave filter signal bands includes n modulation frequencies f m , n≥1, m≥1; performing envelope extraction on the p different octave filter signal bands to obtain p groups of sub-band envelope features; inputting the octave filter signal band corresponding to any group of the sub-band envelope features into a pre-trained voice quality model corresponding to the corresponding octave filter signal band to obtain p groups of reverberation time T; based on p groups of the reverberation time T, p groups of target features are obtained respectively, and the target features are modulation transfer function values.

5. The evaluation method according to claim 4, wherein The envelope extraction on the p different octave filter signal bands to obtain p groups of sub-band envelope features comprises: The p groups of different octave band filtered signal bands are respectively half-wave envelope detected to obtain p groups of sub-band envelope features.

6. The evaluation method according to claim 4, wherein The octave band filtered signal band corresponding to any one of the sub-band envelope features is input, and p groups of reverberation times T are respectively obtained through a pre-trained speech quality model corresponding to the corresponding octave band filtered signal band, including: Any octave band filtered signal band is divided into N speech segments with continuous and equal time lengths, N≥2; Any speech segment in the N speech segments included in any octave band filtered signal band is subjected to feature extraction by a combination structure of one or more of a convolutional neural network, a linear connection layer, an activation layer, and a normalization layer to obtain a [P, Q] shaped matrix, where P is a set of p, and Q is a set of q, to obtain corresponding N speech segment features; Any speech segment feature in the obtained N speech segment features is interacted through a combination structure of one or more of a long short-term memory module, a multi-head / single-head attention module, a linear connection layer, an activation layer, and a normalization layer to obtain corresponding speech segment interaction features; Based on the obtained N speech segment interaction features, N reverberation times T corresponding to the N speech segment interaction features are respectively obtained by linear regression layer or classification layer prediction N ; N reverberation times T corresponding to any of the octave filtered signal bands N averaging separately to obtain p sets of reverberation times T corresponding to the respective said octave filtered signal bands, respectively.

7. The evaluation method according to claim 4, wherein The p groups of the reverberation times T are respectively used to obtain corresponding p groups of target features, including: Based on any modulation frequency f m The value and corresponding reverberation time T are used to obtain any modulation frequency f of any octave band filtered signal. m The modulation transfer function value m k,fm .

8. The evaluation method of claim 7, characterized in that, The p groups of target features are used to obtain a first speech quality evaluation result corresponding to the first target speech signal at the first target position, including: based on any of the modulation transfer function values m k,fm acquiring any modulation frequency f m of the effective signal-to-noise ratio SNR eff k,fm ; based on any of the signal-to-noise ratios, SNRs eff k,fm acquiring any of the modulation frequencies, f, of the corresponding octave filtered signal band, k m at the transmission index, TI k,fm ; calculating the mean of the n transmission indices TI of any octave filtered signal band k k,fm to obtain the modulation transfer index M of the respective octave filtered signal band k k ; modulation transfer index M based on p said octave filtered signal bands k A first speech quality evaluation result corresponding to the first target speech signal is obtained by calculation for the first target position.

9. The evaluation method according to claim 1, wherein The speech quality evaluation result is front-end displayed, and the display mode includes but is not limited to: The speech quality evaluation result is displayed on the interface in the form of a numerical value and a dynamic moving signal; or, The speech quality evaluation result is displayed on the interface in the form of a numerical value and a dynamic wifi signal; Or, The speech quality evaluation result is displayed on the interface in the form of a numerical value and a dynamic dashboard; or, The speech quality evaluation result is displayed on the interface in the form of a numerical value and a dynamic bar.

10. Apparatus for speech quality assessment for telecommunication, characterized in that The evaluation method of any one of claims 1-9 is executed, and the device includes: an acquisition module configured to acquire a first target speech signal corresponding to a first source received by a first speech terminal of a target communication channel, the first source corresponding to a second speech terminal in the target communication channel; A processing module configured to calculate a corresponding first speech quality evaluation result in real time based on the first target speech signal; A display module configured to front-end display the first speech quality evaluation result on the second speech terminal.

11. A speech quality assessment system for telecommunication, characterized by The evaluation system includes: At least two speech terminals, including a first speech terminal and a second speech terminal, a target communication channel is formed between the first speech terminal and the second speech terminal, and the second speech terminal corresponds to a first source; An intelligent device acquires a first target voice signal corresponding to a first source received by a first voice terminal of a target communication channel, the first source corresponding to a second voice terminal other than the first voice terminal in the target communication channel; performs the evaluation method according to any one of claims 1-9 based on the first target voice signal to obtain a corresponding first voice quality evaluation result in real time, and sends the first voice quality evaluation result to the second voice terminal; At least one display device, including a first display device arranged on the second voice terminal, which displays the first voice quality evaluation result in front end.

12. An electronic device, comprising: Comprising: one or more processors; and a memory associated with the one or more processors, the memory for storing program instructions that, when read and executed by the one or more processors, perform the evaluation method according to any one of claims 1-9.

13. A computer-readable storage medium, characterized in that, A computer program is stored thereon, wherein the program is executed by a processor to implement the method according to any one of claims 1-9.

Citation Information

Patent Citations

  • Automatic switching between omnidirectional and directional microphone modes in a hearing aid

    CN101433098A

  • Method for testing intelligibility of speech transmission index

    CN102148033A

  • Voice quality evaluation method and device, equipment and medium

    CN111383657A

  • Sound amplifying system with classroom speech intelligibility measuring function

    CN111757235A

  • Voice quality evaluation device and method, medium and MOS scoring device

    CN112242151A