Voice quality evaluation method, device and system for local communication
By receiving and processing target voice signals, using pre-trained voice quality models for feature extraction and evaluation, the convenience and real-time problems of voice quality evaluation in the prior art are solved, and real-time voice quality feedback and adjustment to the local communication environment are achieved.
Patent Information
- Application Number
- CN202211115352.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-14
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2042-09-14
AI Technical Summary
The prior art requires professional equipment and complex calculations when performing voice quality evaluation, and cannot achieve convenient and real-time local communication voice quality feedback.
By receiving the target voice signal, the speech quality evaluation results are calculated in real time, and the pre-trained speech quality model is used for feature extraction and evaluation, including octave filtering, envelope extraction and neural network processing, providing real-time feedback of the speech quality index.
Real-time and convenient voice quality evaluation of any location in the local communication environment is achieved, and communication smoothness and feedback efficiency are improved.
Smart Images

Figure CN115512719B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of acoustic measurement technology, and in particular to a method, device, and system for evaluating voice quality for local communication. Background Art
[0002] Speech is the primary means of communication between people. In many cases, speech signals are attenuated by the signal path or transmission path between the speaker and listener, resulting in reduced speech quality at the listener's location. Classrooms, concert halls, conference rooms, and various venues equipped with public address systems, such as shopping malls, stadiums, airports, train stations, conference halls, and theaters, must ensure that speech is clearly heard. In emergencies, especially, clear information about escape routes and directions must be provided to those in danger. Therefore, reproducible and accurate methods are needed to verify speech intelligibility in these venues—in other words, to conduct speech quality evaluation.
[0003] In order to determine the degree of degradation of speech quality after passing through the transmission path, speech quality evaluation parameters are usually used for evaluation. By sending a specific test signal to the transmission path and then analyzing the received signal, the transmission quality of the transmission path is derived and expressed in a score. The larger the value, the better the speech quality (clarity or intelligibility) (see the attached figure). Figure 1 The main limitation of existing testing methods and technologies is that they require professionals to use special playback equipment, such as an artificial mouth or talkbox, to play special standard modulation test signals, or use professional equipment to collect room impulse responses for complex calculations. This is not only time-consuming and expensive, but also inaccessible to ordinary users. It is impossible to obtain real-time feedback or guidance on the voice quality of current local language communication.
[0004] Therefore, it is necessary to find a simple, convenient, and real-time feedback method for voice quality evaluation for local communication.
[0005] Application Contents
[0006] The purpose of this application is to provide a method, device and system for evaluating the voice quality of local communication, which can conveniently and accurately evaluate and display the voice quality in local communication in real time.
[0007] To achieve the above application objectives, this application proposes the following technical solutions:
[0008] In a first aspect, a method for evaluating speech quality is provided, the method comprising:
[0009] receiving a target voice signal, the target voice signal being formed by real-time voice received at a first target position and emitted by a target object at a second target position in the same space, the first target position being any position in the same space different from the second target position;
[0010] Obtaining a corresponding speech quality evaluation result in real-time calculation based on the target speech signal;
[0011] The voice quality evaluation results are displayed on the front end.
[0012] In a preferred embodiment, the step of obtaining a corresponding speech quality evaluation result by real-time calculation based on the target speech signal includes:
[0013] Performing feature extraction on the target speech signal to obtain at least one group of target speech signals;
[0014] The at least one set of target speech signals is used as input, and a corresponding speech quality evaluation result is obtained through a pre-trained speech quality model.
[0015] In a preferred embodiment, the speech quality evaluation result includes but is not limited to one of a speech transmission index and a mean opinion score.
[0016] In a preferred embodiment, when the speech quality evaluation result includes a speech transmission index, the step of obtaining the corresponding speech quality evaluation result by real-time calculation based on the target speech signal includes:
[0017] Performing feature extraction on the target speech signal according to p different octave filter signal bands to obtain corresponding p groups of target features, where p is greater than or equal to 2;
[0018] A speech quality evaluation result of the first target position corresponding to the target speech signal is obtained based on the p groups of target features.
[0019] In a preferred embodiment, the step of extracting features from the target speech signal according to p different octave filter signal bands to obtain corresponding p groups of target features includes:
[0020] The target speech signal is filtered to obtain p groups of different octave filter signal bands, and any group of the octave filter signal bands includes n modulation frequencies f m , n≥1, m≥1;
[0021] Performing envelope extraction on the p groups of different octave filter signal bands to obtain p groups of sub-band envelope features;
[0022] Taking any group of octave filter signal bands corresponding to the sub-band envelope features as input, respectively obtaining p groups of reverberation times T using a pre-trained speech quality model corresponding to the corresponding octave filter signal bands;
[0023] Based on the p groups of reverberation times T, corresponding p groups of target features are obtained respectively, where the target features are modulation transfer function values.
[0024] In a preferred embodiment, said extracting envelopes of the p groups of different octave filter signal bands to obtain p groups of sub-band envelope features includes:
[0025] The p groups of different octave-band filtered signal bands are respectively subjected to half-wave envelope detection to obtain p groups of sub-band envelope features.
[0026] In a preferred embodiment, the method of taking an octave filter signal band corresponding to any group of the sub-band envelope features as input and obtaining p groups of reverberation times T respectively by using a pre-trained speech quality model corresponding to the corresponding octave filter signal band comprises:
[0027] Divide any of the octave-band filtered signal bands into N consecutive speech segments of equal duration, where N is greater than or equal to 2;
[0028] Performing feature extraction on any of the N speech segments included in any of the octave-band filtered signal bands using a combination of one or more of a convolutional neural network, a linear connection layer, an activation layer, and a normalization layer to obtain a [P, Q]-shaped matrix, and obtaining corresponding N speech segment features;
[0029] Interacting any of the N obtained speech segment features through a combination of one or more structures selected from a long short-term memory module, a multi-head / single-head attention module, a linear connection layer, an activation layer, and a normalization layer to obtain a corresponding speech segment interaction feature;
[0030] Based on the obtained N speech segment interaction features, the linear regression layer or the classification layer is used to respectively predict and obtain N reverberation times T corresponding to the N speech segment interaction features. N ;
[0031] For N reverberation times T corresponding to any octave filter signal band N The p groups of reverberation times T corresponding to the corresponding octave filter signal bands are obtained by taking averages respectively.
[0032] In a preferred embodiment, obtaining corresponding p groups of target features based on the p groups of reverberation times T includes:
[0033] Based on any modulation frequency f mThe value and the corresponding reverberation time T are used to obtain any modulation frequency f of any set of octave filter signal bands. m The modulation transfer function value m k,fm .
[0034] In a preferred embodiment, obtaining a speech quality evaluation result corresponding to the target speech signal at the first target position based on the p groups of target features includes:
[0035] Based on any of the modulation transfer function values m k,fm Obtain any modulation frequency f corresponding to the octave filter signal band k m The effective signal-to-noise ratio (SNR) effk,fm ;
[0036] Based on any effective signal-to-noise ratio (SNR) effk,fm Obtain any modulation frequency f corresponding to the octave filter signal band k m Transmission index TI at k,fm ;
[0037] Calculate the n transmission indices TI for any octave-band filtered signal band k k,fm The mean of the octave filter signal band k is obtained by taking the modulation transfer index M of the corresponding octave filter signal band k. k ;
[0038] The modulation transfer index M based on the p octave filtered signal bands k A speech quality evaluation result of the target speech signal corresponding to the first target position is obtained by calculation.
[0039] In a preferred embodiment, the evaluation method further comprises pre-training p speech quality models corresponding to the p different octave filter signal bands, respectively, including:
[0040] Based on any speech signal sample in the speech signal sample set, obtaining corresponding p groups of different octave filter signal band sample sets, wherein any group of the octave filter signal band sample sets includes q modulation frequency samples and corresponding q impulse response samples, and any impulse response sample includes a reverberation time sample T0, q≥2;
[0041] With the q modulation frequency samples as input and the corresponding q reverberation time samples T0 as output, p speech quality models corresponding to p octave filter signal bands are obtained through training based on a neural network.
[0042] In a preferred embodiment, the voice quality evaluation result is displayed on the front end, and the display method includes but is not limited to:
[0043] Displaying the speech quality evaluation result on an interface in the form of numerical values and dynamic moving signals; or
[0044] Display the voice quality evaluation result on the interface in the form of numerical values and dynamic WiFi signals; or
[0045] Displaying the speech quality evaluation results in the form of numerical values and dynamic dashboards; or
[0046] The speech quality evaluation result is displayed on the interface in the form of a numerical value and a progress bar.
[0047] In a second aspect, a voice quality evaluation device for local communication is provided, characterized in that the device includes:
[0048] a receiving module, configured to receive a target voice signal, the target voice signal being formed by real-time voice emitted by a target object located at a second target position in the same space and received at a first target position, the first target position being any position in the same space different from the second target position;
[0049] A processing module, configured to calculate in real time based on the target speech signal to obtain a corresponding speech quality evaluation result;
[0050] The display module is used to display the voice quality evaluation result on the front end.
[0051] In a third aspect, a voice quality evaluation system for local communication is provided, the evaluation system comprising:
[0052] at least one voice receiving device;
[0053] At least one display device, the at least one display device being used to display the voice quality evaluation result on the front end;
[0054] An intelligent device is used to receive a target voice signal sent by the at least one voice receiving device, perform an operation as described in any one of the first aspects according to the target voice signal, calculate in real time to obtain a corresponding voice quality evaluation result, and send the voice quality evaluation result to the at least one display device for front-end display.
[0055] In a fourth aspect, an electronic device is provided, including:
[0056] one or more processors; and
[0057] A memory associated with the one or more processors, the memory being used to store program instructions, wherein the program instructions, when read and executed by the one or more processors, perform the operations described in any one of the first aspects.
[0058] In a fifth aspect, a computer-readable storage medium is provided, on which a computer program is stored, wherein when the program is executed by a processor, the method as described in any one of the first aspects is implemented.
[0059] Compared with the prior art, this application has the following beneficial effects:
[0060] The present application provides a voice quality evaluation method, device and system for local communication. The evaluation method includes receiving a target voice signal, where the target voice signal is formed by real-time voice emitted by a target object located at a second target position in the same space and received at a first target position, where the first target position is any position in the same space that is different from the second target position; obtaining a corresponding voice quality evaluation result in real time according to the target voice signal; and displaying the voice quality evaluation result on the front end. The voice quality evaluation method for local communication in the present application can realize voice quality evaluation for a known sound-making object at any position in the local communication environment. The present application is convenient, has real-time feedback and is universal, thereby helping the speaker to obtain voice quality feedback in a timely manner and make timely adjustments, thereby improving the communication fluency in local communication. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] Figure 1 It is an explanation of the meaning of the STI score;
[0062] Figure 2 is a flow chart of the voice quality evaluation method for local communication in this embodiment;
[0063] Figure 3 Schematic diagram of envelope extraction to obtain envelope boundaries in this embodiment;
[0064] Figure 4 1 is a circuit diagram of a half-wave envelope detector in this embodiment;
[0065] Figure 5 It is a schematic diagram of the neural network structure;
[0066] Figures 6a to 6d It is an exemplary display content when the voice quality evaluation result is displayed on the interface;
[0067] Figure 7 This is a system architecture diagram of a voice quality evaluation system for local communication. DETAILED DESCRIPTION
[0068] To make the purpose, technical solutions, and advantages of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0069] In the description of this application, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Therefore, a feature specified as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of this application, unless otherwise specified, "plurality" means two or more.
[0070] To address the current difficulties in achieving real-time measurement and feedback for speech intelligibility in local communication, this embodiment provides a speech quality assessment method with real-time feedback and accurate measurement results. The following describes the speech quality assessment method, device, and system for local communication in more detail, using specific embodiments.
[0071] Example
[0072] like Figure 2 As shown, this embodiment provides a method for evaluating speech quality for local communication, which is suitable for evaluating the intelligibility of language in local communication. In a local communication scenario, the position of the voice receiver is used as the first target position, the target object (speaker) is located at the second target position, and the first target position is any position in the same space that is different from the second target position. For example, in a classroom, the podium is the second target position, the teacher on the podium is the target object, and any other position other than the podium is used as the first target position. A voice transmission channel is formed between the podium and the determined first target position, and channel parameters such as room parameters will have a direct impact on the transmission quality of the voice transmission channel. It can be understood that in a determined communication scenario, each first target position to be measured corresponds to a uniquely determined voice transmission channel. This embodiment can realize real-time and accurate evaluation of voice quality at any position in a determined space, and perform real-time front-end display to further improve the local communication effect.
[0073] Specifically, the voice quality evaluation method for local communication in this embodiment includes the following steps:
[0074] S1. Receive a target voice signal. As described above, the target voice signal is formed by receiving real-time voice from a target object at a second target position in the same space at a first target position. The first target position is any position in the same space that is different from the second target position.
[0075] S2. Calculate the target speech signal in real time to obtain the corresponding speech quality evaluation result.
[0076] Typically, the above step S2 includes the following steps:
[0077] S21, extracting features from the target speech signal to obtain at least one group of target speech signals;
[0078] S22. Using at least one set of target speech signals as input, obtain a corresponding speech quality evaluation result using a pre-trained speech quality model. It should be noted that the speech quality evaluation result includes, but is not limited to, a Speech Transmission Index (STI) or a Mean Opinion Score (MOS).
[0079] For ease of description, the speech quality evaluation results in this embodiment are based on the Speech Transmission Index (STI) as an example, but are not limited to this. Generally, STI sends a specific test signal to the transmission path, analyzes the received signal, derives the speech transmission quality of the transmission path, and expresses it using a score between 0 and 1 (e.g. Figure 1 ).
[0080] Step S21 is specifically as follows: performing feature extraction on the target speech signal according to p different octave filter signal bands k to obtain corresponding p groups of target features, where p≥2, 1≤k≤p.
[0081] Furthermore, step S21 includes:
[0082] S21a, filtering the target speech signal to obtain p groups of different octave filter signal bands k, where any group of octave filter signal bands k includes n modulation frequencies f m , n≥1, m≥1.
[0083] It should be noted that human speech is usually divided into seven frequency bands, so in this embodiment, p=7 is preferred, that is, 1≤k≤7. Therefore, the center frequency f is obtained by filtering the target speech separately. c The octave filter signal bands k are 125Hz, 250Hz, 500Hz, 1kHz, 2kHz, 4kHz and 8kHz respectively, and the upper limit frequency f in each octave filter signal band k is u and the lower frequency f lAs shown in the following formulas (1) and (2):
[0084]
[0085]
[0086] S21b, respectively perform envelope extraction on p groups of different octave filter signal bands k to obtain p groups of sub-band envelope features. The envelope extraction results are as follows: Figure 3 The envelope boundaries are shown.
[0087] This embodiment does not limit the envelope extraction algorithm. It is preferred to perform half-wave envelope detection on p groups of different octave filter signal bands k to obtain p groups of sub-band envelope features (such as Figure 4 As shown), the expression is a differential equation as shown in the following formula (3) (4):
[0088]
[0089]
[0090] S21c. Taking the octave filter signal band k corresponding to any group of sub-band envelope features as input, obtain p groups of reverberation times T respectively through a pre-trained speech quality model corresponding to the corresponding octave filter signal band k.
[0091] Any speech quality model consists of a data preprocessing module, a feature extraction module, a time interaction module, and a prediction module. Specifically:
[0092] Data preprocessing module: divide any octave filter signal band k into N consecutive speech segments (such as Figure 5 Chinese phonetic sl ice), N≥2.
[0093] Feature extraction module: For any of the N speech segments included in any octave filter signal band k, feature extraction is performed through a combination of one or more structures including convolutional neural network, linear connection layer, activation layer, and normalization layer to obtain a matrix of [P, Q] shape, and obtain the corresponding N speech segment features, such as Figure 5 Chinese speech slice features.
[0094] Temporal interaction module: any of the N speech segment features obtained is interacted through a combination of one or more structures selected from the long short-term memory module, multi-head / single-head attention module, linear connection layer, activation layer, and normalization layer to obtain the corresponding speech segment interaction feature, such as Figure 5 Slice interaction features.
[0095] Prediction module: Based on the obtained N speech segment interaction features, the linear regression layer or classification layer is used to predict the N reverberation times T corresponding to the N speech segment interaction features. N ,like Figure 5 Slice reverberation time T N . For N T N The reverberation time T corresponding to the corresponding octave filter signal band k is obtained by averaging. The averaging method includes but is not limited to any one of a simple averaging method, a weighted averaging method, or a harmonic mean method.
[0096] To this end, before step S21c, the evaluation method further includes: Sa, pre-training p speech quality models corresponding to p different octave filter signal bands k, including:
[0097] Sa1. Based on any speech signal sample in the speech signal sample set, obtain corresponding p groups of different octave filter signal band sample sets, where any group of octave filter signal band sample sets includes q modulation frequency samples and corresponding q impulse response samples, and any impulse response sample includes a reverberation time sample T0, where q ≥ 2;
[0098] Sa2. Taking q modulation frequency samples as input and corresponding q reverberation time samples T0 as output, p speech quality models corresponding to p octave filter signal bands k are obtained by training based on a neural network.
[0099] S21d: Based on the p groups of reverberation times T, corresponding p groups of target features are obtained respectively, where the target features are modulation transfer function values.
[0100] The Modulation Transfer Function (MTF) describes the degree to which the modulation m is transmitted from the target object (sound source) to the receiving sensor and is the modulation frequency f m Function m k,fm , MTF determines the degree of modulation reduction of the target speech signal. Specifically, the modulation frequency f m The range is 0.63Hz to 12.5Hz. Therefore, the MTF function value depends on the system environment characteristics and background noise. The MTF calculation process is shown in the following formula (5):
[0101]
[0102] Step S22 specifically includes: obtaining a speech quality evaluation result of the first target position corresponding to the target speech signal based on the p groups of target features.
[0103] Specifically, step S22 includes:
[0104] S221, based on any modulation transfer function value m k,fm Get any modulation frequency f of the corresponding octave filter signal band k m The effective signal-to-noise ratio (SNR) effk,fm ; Specifically, the effective signal-to-noise ratio SNR effk ,fm is calculated by the following formula (6):
[0105]
[0106] S222, based on any effective signal-to-noise ratio SNR effk,fm Obtain any modulation frequency f corresponding to the octave filter signal band k m Transmission index TI at k,fm Specifically, the transmission index TI k,fm It is calculated by the following formula (7):
[0107]
[0108] S223, calculating the n transmission indices TI of any octave filtered signal band k k,fm The mean of the octave filter signal band k is obtained by taking the modulation transfer index M of the corresponding octave filter signal band k. k ; Modulation transfer index M k The value range of M is -15dB to +15dB. k It is calculated by the following formula (8):
[0109]
[0110] S224, modulation transfer index M based on p octave filter signal bands k The speech quality evaluation result (STI) of the target speech signal corresponding to the first target position is calculated; specifically, the STI is calculated by the following formula (9):
[0111]
[0112] Among them, α k represents the different gender weighting factors of the octave-band filtered signal band k;
[0113] β k represents the different gender redundancy factor between the octave filtered signal band k and the octave filtered signal band k+1;
[0114] M k Refers to the modulation transfer index of the octave-band filtered signal band.
[0115] It should be noted that the STI method can distinguish between male and female speech signals. However, in practice, to simplify the measurement process, only male speech is used to evaluate the speech transmission path. Table 1 shows the STI weighting factor α and redundancy factor β for male voices as a function of octave band.
[0116] Table 1
[0117]
[0118] S3. Display the voice quality evaluation result obtained in step S2 on the front end.
[0119] It is understandable that in this embodiment, there is at least one first target position, and the voice quality evaluation results of at least one first target position can be obtained at the same time, and the obtained voice quality evaluation results are displayed on the front end. When the second target object obtains the voice quality evaluation results in real time, it can adjust the volume, adjust the distance from the microphone, adjust the airflow conditions in the space (such as opening and closing doors / windows), etc. according to the feedback results, so that any first target position in the space can clearly hear the target object's words, thereby improving the current language communication effect.
[0120] like Figures 6a to 6d As shown, the interface display method used for the front-end display of the voice quality evaluation results includes but is not limited to displaying the voice quality evaluation results in the form of numerical values and dynamic mobile signals; or, displaying the voice quality evaluation results in the form of numerical values and dynamic wifi signals; or, displaying the voice quality evaluation results in the form of numerical values and dynamic dashboards; or, displaying the voice quality evaluation results in the form of numerical values and progress bars.
[0121] In summary, the voice quality evaluation method for local communication provided in this embodiment can realize voice quality evaluation of any target location in the local communication environment. Compared with the existing methods that require the use of professional equipment and standard methods to complete the evaluation, this application has the advantages of greater universality, convenience, and real-time feedback.
[0122] Furthermore, the speech quality evaluation method for local communication provided in this embodiment adopts a method of constructing speech quality models for different octaves to obtain reverberation times respectively when calculating the speech quality evaluation result, which has strong robustness and reproducibility.
[0123] Corresponding to the above-mentioned voice quality evaluation method, this embodiment further provides a voice quality evaluation device corresponding to the evaluation method, which implements the method through various functional modules. The voice quality evaluation device includes:
[0124] a receiving module, configured to receive a target voice signal, the target voice signal being formed by real-time voice emitted by a target object located at a second target position in the same space and received at a first target position, the first target position being any position in the same space different from the second target position;
[0125] A processing module, configured to calculate in real time based on the target speech signal to obtain a corresponding speech quality evaluation result;
[0126] A display module is used to display the voice quality evaluation results on the front end;
[0127] The model training module is used to pre-train p speech quality models corresponding to the p different octave filter signal bands.
[0128] The processing module includes:
[0129] A feature extraction unit is used to extract features of the target speech signal according to p different octave filter signal bands to obtain corresponding p groups of target features, where p is greater than or equal to 2;
[0130] An evaluation unit is configured to obtain a speech quality evaluation result of the first target position corresponding to the target speech signal based on the p groups of target features.
[0131] Furthermore, the feature extraction unit specifically includes:
[0132] The first processing subunit is used to filter the target speech signal to obtain p groups of different octave filter signal bands, and any group of the octave filter signal bands includes n modulation frequencies f m , n≥1, m≥1;
[0133] A second processing sub-unit is configured to perform envelope extraction on the p groups of different octave filter signal bands to obtain p groups of sub-band envelope features;
[0134] a third processing subunit, configured to take as input an octave filter signal band corresponding to any group of the sub-band envelope features, and obtain p groups of reverberation times T respectively through a pre-trained speech quality model corresponding to the corresponding octave filter signal band;
[0135] The fourth processing subunit is configured to obtain corresponding p groups of target features based on the p groups of reverberation times T, where the target features are modulation transfer function values.
[0136] The second processing sub-unit is specifically configured to perform half-wave envelope detection on the p groups of different octave-band filtered signal bands to obtain p groups of sub-band envelope features.
[0137] The third processing subunit is specifically configured to:
[0138] Divide any of the octave-band filtered signal bands into N consecutive speech segments of equal duration, where N is greater than or equal to 2;
[0139] Performing feature extraction on any of the N speech segments included in any of the octave-band filtered signal bands using a combination of one or more of a convolutional neural network, a linear connection layer, an activation layer, and a normalization layer to obtain a [P, Q]-shaped matrix, and obtaining corresponding N speech segment features;
[0140] Interacting any of the N obtained speech segment features through a combination of one or more structures selected from a long short-term memory module, a multi-head / single-head attention module, a linear connection layer, an activation layer, and a normalization layer to obtain a corresponding speech segment interaction feature;
[0141] Based on the obtained N speech segment interaction features, the linear regression layer or the classification layer is used to respectively predict and obtain N reverberation times T corresponding to the N speech segment interaction features. N ;
[0142] For N reverberation times T corresponding to any octave filter signal band N The p groups of reverberation times T corresponding to the corresponding octave filter signal bands are obtained by taking averages respectively.
[0143] The fourth processing subunit is specifically configured to process the signal based on any modulation frequency f m The value and the corresponding reverberation time T are used to obtain any modulation frequency f of any set of octave filter signal bands. m The modulation transfer function value m k,fm .
[0144] Furthermore, the evaluation unit specifically includes:
[0145] A fifth processing subunit is configured to process the modulation transfer function value m based on any one of the modulation transfer function values m. k,fm Obtain any modulation frequency f corresponding to the octave filter signal band k m The effective signal-to-noise ratio (SNR) effk,fm ;
[0146] The sixth processing subunit is configured to process the signal based on any of the effective signal-to-noise ratios SNR. effk,fm Obtain any modulation frequency f corresponding to the octave filter signal band k m Transmission index TI at k,fm ;
[0147] The seventh processing subunit is used to calculate the n transmission indices TI of any octave filtered signal band k k,fm The mean of the octave filter signal band k is obtained by taking the modulation transfer index M of the corresponding octave filter signal band k. k ;
[0148] An eighth processing subunit is configured to: k A speech quality evaluation result of the target speech signal corresponding to the first target position is obtained by calculation.
[0149] The eighth processing subunit specifically performs the calculation shown in the following formula (9) to obtain the speech quality evaluation result of the first target position corresponding to the target speech signal:
[0150]
[0151] Among them, α k represents the different gender weighting factors of the octave-band filtered signal band k;
[0152] β k represents the different gender redundancy factor between the octave filtered signal band k and the octave filtered signal band k+1;
[0153] M k Refers to the modulation transfer index of the octave-band filtered signal band.
[0154] The model training module is specifically used for:
[0155] Based on any speech signal sample in the speech signal sample set, obtaining corresponding p groups of different octave filter signal band sample sets, wherein any group of the octave filter signal band sample sets includes q modulation frequency samples and corresponding q impulse response samples, and any impulse response sample includes a reverberation time sample T0, q≥2;
[0156] With the q modulation frequency samples as input and the corresponding q reverberation time samples T0 as output, p speech quality models corresponding to p octave filter signal bands are obtained through training based on a neural network.
[0157] It should be noted that the voice quality assessment device provided in the above embodiment is only illustrated by the division of the above functional modules when performing voice quality assessment services. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the system can be divided into different functional modules to complete all or part of the functions described above. In addition, the voice quality assessment device provided in the above embodiment and the embodiment of the voice quality assessment method are based on the same concept, that is, the device is based on the method. The specific implementation process is detailed in the method embodiment and will not be repeated here.
[0158] As well as Figure 7 As shown, this embodiment further provides a voice quality evaluation system for local communication, the evaluation system comprising:
[0159] at least one voice receiving device;
[0160] At least one display device, the display device being used to display the voice quality evaluation result on the front end;
[0161] An intelligent device is used to receive a target voice signal sent by the at least one voice receiving device, perform a local voice quality evaluation method according to the target voice signal to obtain a corresponding voice quality evaluation result in real time, and send the voice quality evaluation result to the at least one display device for front-end display.
[0162] Furthermore, this embodiment further provides an electronic device, including:
[0163] one or more processors; and
[0164] A memory associated with the one or more processors, the memory being used to store program instructions, wherein when the program instructions are read and executed by the one or more processors, the program instructions perform the operations described in any one of the methods for evaluating voice quality for local communication.
[0165] Regarding the voice quality evaluation method performed by executing program instructions, the specific execution details and corresponding beneficial effects are consistent with the description of the aforementioned method and will not be repeated here.
[0166] Furthermore, this embodiment further provides a computer-readable storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the method described in any one of the methods for evaluating voice quality for local communication is implemented.
[0167] All of the above optional technical solutions can be combined in any way to form optional embodiments of the present application, that is, any multiple embodiments can be combined to meet the needs of different application scenarios. They are all within the scope of protection of the present application and will not be described in detail here.
[0168] It should be noted that the above is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should be included in the scope of protection of the present application.
Claims
1. A method for evaluating speech quality for local communication, characterized in that: The evaluation method includes: receiving a target voice signal, the target voice signal being formed by real-time voice received at a first target position and emitted by a target object at a second target position in the same space, the first target position being any position in the same space different from the second target position; Extracting features from the target speech signal includes: extracting features from the target speech signal according to p different octave filter signal bands to obtain corresponding p groups of target features, where p is greater than or equal to 2, specifically including: The target speech signal is filtered to obtain p groups of different octave filter signal bands, and any group of the octave filter signal bands includes n modulation frequencies f m , n≥1, m≥1; Performing envelope extraction on the p groups of different octave filter signal bands to obtain p groups of sub-band envelope features; Taking any group of octave filter signal bands corresponding to the sub-band envelope features as input, respectively obtaining p groups of reverberation times T using a pre-trained speech quality model corresponding to the corresponding octave filter signal bands; Based on the p groups of reverberation times T, corresponding p groups of target features are obtained, wherein the target features are modulation transfer function values; Taking the p groups of target features as input, a corresponding speech quality evaluation result is obtained through a pre-trained speech quality model; and the speech quality evaluation result is displayed on the front end.
2. The evaluation method according to claim 1, wherein The speech quality evaluation result includes but is not limited to one of a speech transmission index and a mean opinion score.
3. The evaluation method according to claim 1, wherein The step of extracting envelopes of the p groups of different octave filter signal bands to obtain p groups of sub-band envelope features comprises: The p groups of different octave-band filtered signal bands are respectively subjected to half-wave envelope detection to obtain p groups of sub-band envelope features.
4. The evaluation method according to claim 3, wherein: The method of taking an octave filter signal band corresponding to any group of the sub-band envelope features as input and obtaining p groups of reverberation times T respectively through a pre-trained speech quality model corresponding to the corresponding octave filter signal band comprises: Divide any of the octave-band filtered signal bands into N consecutive speech segments of equal duration, where N is greater than or equal to 2; Performing feature extraction on any of the N speech segments included in any of the octave-band filtered signal bands using a combination of one or more of a convolutional neural network, a linear connection layer, an activation layer, and a normalization layer to obtain a [P, Q]-shaped matrix, and obtaining corresponding N speech segment features; Interacting any of the N obtained speech segment features through a combination of one or more structures selected from a long short-term memory module, a multi-head / single-head attention module, a linear connection layer, an activation layer, and a normalization layer to obtain a corresponding speech segment interaction feature; Based on the obtained N speech segment interaction features, the linear regression layer or the classification layer is used to respectively predict and obtain N reverberation times T corresponding to the N speech segment interaction features. N ; For N reverberation times T corresponding to any octave filter signal band N The p groups of reverberation times T corresponding to the corresponding octave filter signal bands are obtained by taking averages respectively.
5. The evaluation method according to claim 3, wherein: The obtaining corresponding p groups of target features based on the p groups of reverberation times T respectively includes: Based on any modulation frequency f m The value and the corresponding reverberation time T are used to obtain any modulation frequency f of any set of octave filter signal bands. m The modulation transfer function value m k,fm .
6. The evaluation method according to claim 5, wherein: The obtaining, based on the p groups of target features, a speech quality evaluation result of the first target position corresponding to the target speech signal comprises: Based on any of the modulation transfer function values m k,fm Obtain any modulation frequency f corresponding to the octave filter signal band k m The effective signal-to-noise ratio (SNR) effk,fm ; Based on any effective signal-to-noise ratio (SNR) effk,fm Obtain any modulation frequency f corresponding to the octave filter signal band k m Transmission index TI at k,fm ; Calculate the n transmission indices TI for any octave-band filtered signal band k k,fm The mean of the octave filter signal band k is obtained by taking the modulation transfer index M of the corresponding octave filter signal band k. k ; The modulation transfer index M based on the p octave filtered signal bands k A speech quality evaluation result of the target speech signal corresponding to the first target position is obtained by calculation.
7. The evaluation method according to any one of claims 2 to 6, wherein: The evaluation method further includes pre-training p speech quality models corresponding to the p different octave filter signal bands, including: Based on any speech signal sample in the speech signal sample set, obtaining corresponding p groups of different octave filter signal band sample sets, wherein any group of the octave filter signal band sample sets includes q modulation frequency samples and corresponding q impulse response samples, and any impulse response sample includes a reverberation time sample T0, q≥2; With the q modulation frequency samples as input and the corresponding q reverberation time samples T0 as output, p speech quality models corresponding to p octave filter signal bands are obtained through training based on a neural network.
8. The evaluation method according to claim 1, wherein: The voice quality evaluation result is displayed on the front end, and the display method includes but is not limited to: Displaying the speech quality evaluation result on an interface in the form of numerical values and dynamic moving signals; or The voice quality evaluation result is displayed on the interface in the form of numerical values and dynamic WiFi signals; or, Displaying the speech quality evaluation results in the form of numerical values and dynamic dashboards; or The speech quality evaluation result is displayed on the interface in the form of a numerical value and a progress bar.
9. A voice quality evaluation device for local communication, characterized in that: The device comprises: a receiving module, configured to receive a target voice signal, the target voice signal being formed by real-time voice emitted by a target object located at a second target position in the same space and received at a first target position, the first target position being any position in the same space different from the second target position; A processing module is configured to extract features from the target speech signal, comprising: extracting features from the target speech signal according to p different octave filter signal bands to obtain corresponding p groups of target features, where p is greater than or equal to 2, specifically comprising: The target speech signal is filtered to obtain p groups of different octave filter signal bands, and any group of the octave filter signal bands includes n modulation frequencies f m , n≥1, m≥1; Performing envelope extraction on the p groups of different octave filter signal bands to obtain p groups of sub-band envelope features; Taking any group of octave filter signal bands corresponding to the sub-band envelope features as input, respectively obtaining p groups of reverberation times T using a pre-trained speech quality model corresponding to the corresponding octave filter signal bands; Based on the p groups of reverberation times T, corresponding p groups of target features are obtained, wherein the target features are modulation transfer function values; Taking the p groups of target features as input, obtaining corresponding speech quality evaluation results through a pre-trained speech quality model; The display module is used to display the voice quality evaluation result on the front end.
10. A voice quality evaluation system for local communication, characterized in that: The evaluation system includes: at least one voice receiving device; At least one display device, the display device being used to display the voice quality evaluation result on the front end; An intelligent device, wherein the intelligent device is used to receive a target voice signal sent by the at least one voice receiving device, execute the evaluation method according to any one of claims 1 to 8 in real time to calculate and obtain a corresponding voice quality evaluation result based on the target voice signal, and send the voice quality evaluation result to the at least one display device for front-end display.
11. An electronic device, characterized in that: include: one or more processors; as well as a memory associated with the one or more processors, the memory being configured to store program instructions, wherein the program instructions, when read and executed by the one or more processors, execute the evaluation method according to any one of claims 1 to 8; as well as A display associated with the one or more processors, the display being used to display in real time the speech quality evaluation results obtained after the one or more processors execute the program instructions.
12. A computer-readable storage medium, characterized in that A computer program is stored thereon, wherein when the program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Voice quality evaluation method and device and storage medium
CN114694685A
Speech Intelligibility Measurement and Open Space Noise Masking
US20150243297A1