Voice communication quality evaluation method and device, server and storage medium

By obtaining and splicing the embedded feature vectors of the evaluation voice fragments and clean voice fragments in audio and video communication, and combining the actual communication scene data set and attention mechanism training model, the accuracy of voice communication quality evaluation in audio and video communication scenarios is solved, and more accurate MOS evaluation is achieved.

WO2025161664A1PCT designated stage Publication Date: 2025-08-07DINGTALK (CHINA) INFORMATION TECH CO LTD

Patent Information

Application Number
PCT/CN2024/136210
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-29
Filing Date
2024-12-02
Publication Date
2025-08-07

AI Technical Summary

Technical Problem

In audio and video communication scenarios, it is difficult to obtain clean voice clips corresponding to the voice clip to be evaluated, resulting in the inability to accurately evaluate the voice communication quality.

Method used

By obtaining the embedded feature vectors of the target user's to be evaluated and the clean voice fragments, the voice communication quality evaluation model is called for scoring, and the model is trained using the data set and attention mechanism in the actual communication scenario to improve the evaluation accuracy.

Benefits of technology

The accurate evaluation of voice communication quality in audio and video communication scenarios is achieved, and the robustness and accuracy of the model are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024136210_07082025_PF_FP_ABST
    Figure CN2024136210_07082025_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of artificial intelligence, and provides a voice communication quality evaluation method and device, a server and a storage medium. The method comprises: processing a voice segment to be evaluated of a target user to obtain a first embedded feature vector; acquiring a second embedded feature vector corresponding to a clean voice segment of the target user; concatenating the first embedded feature vector and the second embedded feature vector to obtain a third embedded feature vector; and calling a voice communication quality evaluation model to process the third embedded feature vector to obtain the MOS of the voice segment to be evaluated. According to the present application, a first embedded feature vector corresponding to a voice segment to be evaluated of a target user and a second embedded feature vector corresponding to a clean voice segment of the target user are concatenated, and then a voice communication quality evaluation model is called to process the third embedded feature vector obtained by the concatenation, thereby implementing evaluation of the voice communication quality in audio and video communication scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Voice communication quality evaluation method, device, server and storage medium

[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on January 29, 2024, with application number 202410123614.2, and entitled “Voice Communication Quality Evaluation Method, Device, Server and Storage Medium”, all contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of artificial intelligence technology, and in particular to a method, device, server, and storage medium for evaluating voice communication quality. Background Art

[0003] The Mean Opinion Score (MOS) is used to quantify voice communication quality and can be obtained through subjective listening tests. MOS scores range from 1 to 5, with 5 indicating the best quality and 1 indicating the worst. With the development of 5G (fifth-generation mobile communication technology) networks and mobile devices, billions of audio and video communications are conducted daily around the world. To ensure the quality of these communications, it is necessary to evaluate the quality of voice communications in these scenarios.

[0004] Currently, when evaluating the quality of voice communication, a user's voice clip to be evaluated is obtained, and a clean voice clip of the user is obtained. The clean voice clip has the highest MOS, and its content, duration, etc. are exactly the same as those of the voice clip to be evaluated. Then, based on the clean voice clip, the voice clip to be evaluated is scored to obtain the MOS score of the voice clip to be evaluated.

[0005] However, in audio and video communication scenarios, it is difficult to obtain clean voice clips corresponding to the voice clips to be evaluated, making it impossible to evaluate the voice communication quality in audio and video communication scenarios. Therefore, it is urgent to provide a voice communication quality evaluation method to evaluate the voice communication quality in audio and video communication scenarios. Summary of the Invention

[0006] The embodiments of the present application provide a method, device, server, and storage medium for evaluating voice communication quality, which can evaluate the voice communication quality in audio and video communication scenarios. The technical solution is as follows:

[0007] In a first aspect, a method for evaluating voice communication quality is provided, the method comprising:

[0008] During the audio and video communication process of the target user, obtaining the target user's voice clip to be evaluated;

[0009] Processing the speech segment to be evaluated to obtain a first embedded feature vector, where the first embedded feature vector is used to characterize a voiceprint feature distribution of the speech segment to be evaluated in the frequency domain in a current audio and video communication scenario;

[0010] Obtaining a second embedded feature vector, where the second embedded feature vector is a feature vector obtained by processing a clean voice segment of the target user, and the second embedded feature vector is used to represent the voiceprint feature distribution of the clean voice segment in the frequency domain in an interference-free audio and video communication scenario, wherein the clean voice segment has the highest voice quality score (MOS);

[0011] concatenating the first embedded feature vector and the second embedded feature vector to obtain a third embedded feature vector;

[0012] A voice communication quality assessment model is called to process the third embedded feature vector to obtain the MOS of the voice segment to be evaluated. The voice communication quality assessment model is used to compare the voiceprint features corresponding to the clean voice segment in the third embedded feature vector with the voiceprint features corresponding to the voice segment to be evaluated, and score the voice segment to be evaluated based on the comparison result and the MOS of the clean voice segment.

[0013] In a second aspect, a device for evaluating voice communication quality is provided, the device comprising:

[0014] An acquisition module is configured to acquire a speech segment to be evaluated from a target user during audio and video communication with the target user;

[0015] a processing module configured to process the speech segment to be evaluated to obtain a first embedded feature vector, wherein the first embedded feature vector is used to represent the voiceprint feature distribution of the speech segment to be evaluated in the frequency domain in the current audio and video communication scenario;

[0016] The acquisition module is configured to acquire a second embedded feature vector, where the second embedded feature vector is a feature vector obtained by processing a clean voice segment of the target user, the second embedded feature vector being used to characterize a voiceprint feature distribution of the clean voice segment in a frequency domain in an interference-free audio and video communication scenario, and the clean voice segment has the highest voice quality score (MOS);

[0017] a concatenation module configured to concatenate the first embedded feature vector and the second embedded feature vector to obtain a third embedded feature vector;

[0018] The processing module is configured to call a voice communication quality assessment model to process the third embedded feature vector to obtain the MOS of the voice segment to be evaluated. The voice communication quality assessment model is used to compare the voiceprint features corresponding to the clean voice segment in the third embedded feature vector with the voiceprint features corresponding to the voice segment to be evaluated, and score the voice segment to be evaluated based on the comparison result and the MOS of the clean voice segment.

[0019] In a third aspect, a server is provided, comprising a processor and a memory; the memory stores at least one program code; the at least one program code is used to be called and executed by the processor to implement the voice communication quality evaluation method described in the first aspect.

[0020] In a fourth aspect, a computer-readable storage medium is provided, wherein at least one computer program is stored in the computer-readable storage medium, and when the at least one computer program is executed by a processor, the method for evaluating the voice communication quality described in the first aspect can be implemented.

[0021] In a fifth aspect, a computer program product is provided, which includes a computer program, and when the computer program is executed by a processor, it can implement the method for evaluating the voice communication quality described in the first aspect.

[0022] The beneficial effects of the technical solution provided by the embodiments of the present application are:

[0023] In the embodiment of the present application, during the audio and video communication process of the target user, an embedded feature vector representing the target user's voiceprint feature distribution is obtained. Since the voiceprint feature distribution is related to the voice communication quality but not to the voice content, the method provided by the embodiment of the present application can realize the evaluation of the voice communication quality in the audio and video communication scenario. To improve the quality of voice evaluation, the embodiment of the present application pre-stores a second embedded feature vector of the target user, which is used to represent the voiceprint feature distribution of the user's clean voice segment in the frequency domain in the undisturbed audio and video communication scenario. After obtaining the first embedded feature vector representing the voiceprint feature distribution of the target user's voice segment to be evaluated in the current audio and video communication scenario, the first embedded feature vector and the second embedded feature vector corresponding to the target user are concatenated to obtain a third embedded feature vector corresponding to the target user. The third embedded feature vector contains the voiceprint feature distribution of the target user in both the current audio and video communication scenario and the undisturbed audio and video communication scenario. Then, a pre-trained voice communication quality evaluation model is called to process the third embedded feature vector corresponding to the target user to obtain the MOS of the voice segment to be evaluated. Compared with the first embedded feature vector, the third embedded feature vector contains richer information. In actual application, the voice communication quality assessment model uses the second embedded feature vector corresponding to the clean voice segment as the prior feature, compares the voiceprint features corresponding to the clean voice segment in the third embedded feature vector with the voiceprint features corresponding to the voice segment to be evaluated, and then scores the voice segment to be evaluated based on the comparison results and the MOS of the clean voice segment, and the obtained MOS is more accurate. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0025] FIG1 shows an architecture diagram of an audio and video communication system used in an embodiment of the present application;

[0026] FIG2 shows a block diagram of a voice communication quality assessment system provided by an embodiment of the present application;

[0027] FIG3 is a flow chart of a method for training a voice communication quality assessment model according to an embodiment of the present application;

[0028] FIG4 shows a block diagram of a training process of a voice communication quality assessment model provided in an embodiment of the present application;

[0029] FIG5 is a flow chart of another method for training a voice communication quality assessment model according to an embodiment of the present application;

[0030] FIG6 is a flow chart of a method for evaluating voice communication quality provided in an embodiment of the present application;

[0031] FIG7 is a schematic structural diagram of a device for evaluating voice communication quality according to an embodiment of the present application;

[0032] FIG8 shows a structural block diagram of a server provided by an exemplary embodiment of the present application. DETAILED DESCRIPTION

[0033] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.

[0034] It should be understood that the terms "each," "plurality," and "any" used in the embodiments of this application include two or more, "each" refers to each of the corresponding plurality, and "any" refers to any one of the corresponding plurality. For example, if a plurality of words includes 10 words, "each" refers to each of the 10 words, and "any" refers to any one of the 10 words.

[0035] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and provide corresponding operation portals for users to choose to authorize or refuse. For example, the voice clips involved in this application (including the voice clips to be evaluated and the clean voice clips) are all obtained with full authorization.

[0036] Before executing the embodiments of the present application, the nouns involved in the embodiments of the present application are first explained.

[0037] MOS is a commonly used subjective evaluation method that requires users to rate audio quality, typically on a scale of 1 to 5. By averaging the scores of multiple users, a comprehensive evaluation score can be obtained, which can reflect the overall user experience of audio quality.

[0038] Intrusive SQA (intrusive voice communication quality assessment): A voice communication quality assessment method in which each speech segment to be evaluated requires a corresponding clean speech segment.

[0039] Non-intrusive SQA (non-intrusive SQA): A method for assessing the quality of voice communication that requires only the speech segment to be evaluated, without requiring the corresponding clean speech segment.

[0040] Speaker Embedding: Through the speaker recognition neural network, embedded information related to a specific speaker is extracted and used as voiceprint feature information.

[0041] Artificial intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve results. In other words, AI is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also involves studying the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.

[0042] With the research and advancement of artificial intelligence technology, it has been studied and applied in many fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, unmanned driving, autonomous driving, robotics, smart healthcare, and smart customer service. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role. The solutions provided in the embodiments of this application involve technologies such as natural language processing in artificial intelligence, which are specifically explained through the subsequent embodiments. Natural language processing is an important field in the fields of computer science and artificial intelligence. It studies various theories and methods that enable effective communication between humans and computers using natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, research in this field involves natural language, that is, the language people use in daily life, and is closely related to the study of linguistics. Natural language processing technologies generally include text processing, semantic understanding, machine translation, robot question answering, knowledge graphs, and other technologies.

[0043] With the development of internet technology, various online communication applications (including instant messaging and video conferencing) have been developed. Audio and video communication based on these applications is not only low-cost but also convenient and fast, and has gradually become the primary means of communication. During audio and video communication, voice quality can be measured using Mean Score (MOS). MOS can be obtained through subjective listening tests conducted according to ITU-T P.800 (a subjective voice communication quality evaluation standard). Although ITU-T P.808 provides guidance for crowdsourced MOS testing, subjective testing remains laborious and time-consuming. To overcome these issues, objective methods such as Perceptual Evaluation of Voice Communication Quality (PESQ) and Perceptual Objective Listening Quality Analysis (POLQA) can be used for evaluation. These objective methods can, to a certain extent, replace subjective listening tests. However, these methods are intrusive voice communication quality assessment methods, requiring a clean audio clip with identical content and duration for each audio clip to be evaluated. Except in laboratory scenarios, obtaining a clean audio clip for each audio clip to be evaluated is not possible in most audio and video communication scenarios. Therefore, intrusive voice communication quality assessment methods cannot assess voice communication quality in audio communication scenarios.

[0044] With the development of artificial intelligence technology, non-intrusive voice communication quality assessment methods have been proposed. These methods use a dataset to train a neural network model. Based on the trained neural network model, the voice clips to be assessed are processed to assess the voice communication quality in audio communication scenarios. Although non-intrusive voice communication quality assessment methods can assess voice communication quality in audio communication scenarios without requiring the acquisition of corresponding clean voice clips, the datasets used to train the neural network models are typically not derived from actual communication scenarios and lack generalizability. This results in a lack of accuracy and robustness in the trained neural network models, and accordingly, inaccurate assessment results based on these neural network models. In summary, neither intrusive nor non-intrusive voice communication quality assessment methods can accurately assess voice communication quality in audio and video communication scenarios.

[0045] In order to solve the above problems, an embodiment of the present application proposes a method for evaluating the quality of voice communication, which concatenates a first embedded feature vector corresponding to the voice segment to be evaluated and a second embedded feature vector corresponding to the user's clean voice segment to obtain a third embedded feature vector, and then calls a voice communication quality evaluation model to process the third embedded feature vector to obtain the MOS of the voice segment to be evaluated. The embodiment of the present application uses voice data in actual communication scenarios as a data set, and the data set has generalization properties. In addition, an attention mechanism is used in the training process, so that the model can learn the voiceprint features of clean voice segments, thereby being able to compare the voiceprint features of clean voice segments with the voiceprint features of the voice segment to be evaluated, so that the trained model has accuracy and robustness. Therefore, the model can accurately evaluate the MOS of the voice segment to be evaluated.

[0046] Please refer to Figure 1, which shows an architecture diagram of the audio and video communication system adopted in an embodiment of the present application. The audio and video communication system includes a terminal 101, a terminal 102 and a server 103, etc. (It should be noted here that the number of terminals included in the audio and video communication system is at least two, and Figure 1 only exemplarily shows two terminals in the audio and video communication process, and is not used to limit the number of terminals in the system).

[0047] Among them, terminal 101 and terminal 102 can be devices used by a first user and a second user for audio and video communication, and an online communication application is installed in the terminal 101 and the terminal 102. During the audio and video communication process based on the installed online communication application, the terminal 101 can collect the voice data of the first user and send the collected voice data to the server 102, and the server 102 sends the received voice data to the terminal 102. Correspondingly, the terminal 102 can collect the voice data of the second user and send the collected voice data to the server 102, and the server 102 sends the received voice data to the terminal 101, thereby realizing audio and video communication between the first user on the terminal 101 side and the second user on the terminal 102 side. The terminal 101 and the terminal 102 can be smart phones, tablet computers, laptops, etc., and the embodiments of the present application do not specifically limit the types of the terminals 101 and 102.

[0048] Server 103 can be the backend server of the online communication application used by terminal 101 and terminal 102 for audio and video communication. Server 103 can receive the voice data of the first user sent by terminal 101 and send the received voice data to terminal 102. Correspondingly, server 103 can receive the voice data of the second user sent by terminal 102 and send the received voice data to terminal 101, thereby realizing audio and video communication between the first user on the terminal 101 side and the second user on the terminal 102 side. To ensure the quality of audio and video communication, a voice communication quality assessment model is installed in server 103. During the audio and video communication process, by calling this voice communication quality assessment model, the voice communication quality of each user in the audio and video communication can be evaluated. For example, when server 103 receives the voice data of the first user sent by terminal 101, it will evaluate the voice communication quality of the first user side based on the voice data. Specifically, server 103 may obtain a speech segment to be evaluated from the first user's speech data, process the speech segment to be evaluated to obtain a first embedded feature vector, obtain a pre-stored second embedded feature vector corresponding to a clean speech segment of the first user, concatenate the first embedded feature vector and the second embedded feature vector corresponding to the first user to obtain a third embedded feature vector corresponding to the first user, then invoke a voice communication quality assessment model to process the third embedded feature vector corresponding to the first user to obtain the MOS of the speech segment to be evaluated. Similarly, upon receiving speech data of a second user sent by terminal 102, server 103 may evaluate the speech communication quality of the second user based on the speech data. Specifically, server 103 may obtain a speech segment to be evaluated from the second user's speech data, process the speech segment to be evaluated to obtain a first embedded feature vector, obtain a pre-stored second embedded feature vector corresponding to a clean speech segment of the second user, concatenate the first embedded feature vector and the second embedded feature vector corresponding to the second user to obtain a third embedded feature vector corresponding to the second user, then invoke a voice communication quality assessment model to process the third embedded feature vector corresponding to the second user to obtain the MOS of the speech segment to be evaluated.

[0049] The above-mentioned voice communication quality assessment model can be trained by server 103 or by other servers, and this application does not specifically limit this. Of course, in addition to the voice communication quality assessment model, server 103 can also install other models for voice communication quality assessment. Since these models are all used for voice communication quality assessment, the embodiments of this application combine these models to form a voice communication quality assessment system.

[0050] Referring to FIG. 2 , a voice communication quality assessment system according to an embodiment of the present application is shown. The system includes a feature extraction model, a downsampling model, a voice communication quality assessment model, and the like.

[0051] The feature extraction model extracts the Mel-Spec features of speech segments, generating Mel-Spec feature vectors. Mel-Spec is a frequency domain representation that better aligns with the human auditory perception. Sound is mapped onto the Mel scale using a set of Mel filters. The filters are densely distributed in the low-frequency range and sparsely distributed in the high-frequency range. The Mel-Spectrum is nonlinear, resulting in the same perceived difference between two pairs of frequencies with equal distances on the Mel scale. This means that human perception is linearly related to the Mel scale. In the low-frequency range (1000 Hz), the Mel scale has a nearly linear relationship with normal frequencies, while in the high-frequency range, the relationship is logarithmic.

[0052] The downsampling model is used to downsample the Mel-spectrogram feature vector to obtain an embedded feature vector. The downsampling model can include convolutional layers, pooling layers, and normalization layers.

[0053] The voice communication quality assessment model processes the input embedded feature vector to obtain the Mean of Score (MOS). This model may include a linear mapping layer and a pooling layer. The linear mapping layer performs linear mapping on the input embedded feature vector to obtain a voiceprint feature vector. The pooling layer performs a pooling operation on the voiceprint feature vector to obtain the Mean of Score (MOS). These linear mapping and pooling layers are not comprised solely of a single neural network layer; they incorporate other functional layers, such as a fully connected layer incorporated into the pooling layer.

[0054] This embodiment of the present application provides a method for training a voice communication quality assessment model. Taking a server executing this embodiment of the present application as an example, the server may be the server in FIG1 or another server. Referring to FIG3 , the method flow provided by this embodiment of the present application includes:

[0055] 301. Obtain a sample clean voice segment and a sample to-be-evaluated voice segment from the same user.

[0056] The sample clean voice segments and the sample to-be-evaluated voice segments are not voice segments acquired using speech conversion or text-to-speech conversion methods, but are voice segments collected in actual audio and video communication scenarios. The sample clean voice segments are user voice segments collected in an interference-free audio and video communication scenario. The sample to-be-evaluated voice segments are user voice segments collected in an audio and video communication scenario, which may or may not be subject to interference. The language used by the user in the sample to-be-evaluated voice segments and the sample clean voice segments can be any language, such as Chinese, English, or Japanese, and the sample to-be-evaluated voice segments and the sample clean voice segments can correspond to the same language or different languages. The content and duration of the sample clean voice segments and the sample to-be-evaluated voice segments can be the same or different, and this is not specifically limited in the embodiments of the present application. The sample clean voice segments and the sample to-be-evaluated voice segments are both annotated with a MOS. Generally, the sample clean voice segment has the highest MOS, while the sample to-be-evaluated voice segment has a MOS between the highest and lowest MOS. For example, the MOS score ranges from 1 to 5, the MOS of a clean speech segment may be 5, and the MOS of a speech segment to be evaluated may be 2.

[0057] Furthermore, in an embodiment of the present application, a sample clean speech segment and a sample to-be-evaluated speech segment from the same user are combined into a training sample pair, and multiple training sample pairs are combined into a training data set, and then the voice communication quality evaluation model is trained using the training sample pairs as units.

[0058] Optionally, in order to facilitate subsequent testing of the accuracy of the trained voice communication quality assessment model, the embodiment of the present application collects multiple voice clips of multiple users in an actual audio and video communication scenario, and then combines a sample clean voice clip and a sample voice clip to be evaluated from the same user into a test sample pair, and the multiple test sample pairs constitute a test data set, and then the accuracy of the trained voice communication quality assessment model is verified in units of test sample pairs.

[0059] Considering that there are many types of interference factors in actual audio and video communication scenarios, such as background noise, packet loss, nonlinear processing, low capacity, etc., for different types of interference factors, when constructing the training dataset and the verification dataset, the embodiments of the present application need to ensure that the training dataset and the test dataset include sample speech segments to be evaluated that are collected under different types of interference factors, so as to improve the generalization of the training dataset and the verification dataset, and make the trained voice communication quality assessment model more robust. For example, for the training dataset and the test dataset, the training sample pairs include: 20% of the training sample pairs are training sample pairs with background noise problems in the sample speech segments to be evaluated, 20% of the training sample pairs are training sample pairs with packet loss problems in the sample speech segments to be evaluated, 20% of the training sample pairs are training sample pairs with nonlinear processing problems in the sample speech segments to be evaluated, 20% of the training sample pairs are training sample pairs with low capacity problems in the sample speech segments to be evaluated, and 20% of the training sample pairs are training sample pairs with clean speech segments. By adding clean speech segments as sample speech segments to be evaluated, the bias in subjective ratings can be suppressed.

[0060] 302. Process the sample speech segment to be evaluated to obtain a first sample embedded feature vector.

[0061] The first sample embedded feature vector is used to characterize the voiceprint feature distribution in the frequency domain of the sample speech segment to be evaluated in the audio and video communication scenario at the time of acquisition. Specifically, when processing the sample speech segment to be evaluated to obtain the first sample embedded feature vector, a feature extraction model can be invoked to extract features from the sample speech segment to obtain a first sample Mel-spectrogram feature vector. The downsampling model can then be invoked to downsample the first sample Mel-spectrogram feature vector to obtain the first sample embedded feature vector. The feature extraction model and downsampling model can employ pre-trained models from related technologies.

[0062] Furthermore, the duration of the sample speech segment to be evaluated is usually long, and the computational cost of processing the sample speech segment to be evaluated with a long duration is large and the processing speed is slow. As described in step 301, the duration of the sample speech segment to be evaluated and the sample clean speech segment can be the same or different. If the duration of the sample speech segment to be evaluated is different from the duration of the sample clean speech segment, the length of the embedded feature vector obtained by processing speech segments of different durations must be different. Similarity calculation cannot be performed based on embedded feature vectors of different lengths, while similarity calculation can improve the accuracy of the model. Therefore, in order to reduce the computational cost of the model training process and improve the training speed and accuracy of the model, the embodiment of the present application can pre-process the sample speech segment to be evaluated before processing the sample speech segment. Specifically, the number of sub-segments can be pre-set to a preset number, and then the sample speech segment to be evaluated is randomly segmented using a window of a preset duration. When a preset number of sample speech sub-segments to be evaluated are obtained, segmentation is stopped, and the preset number of sample speech sub-segments to be evaluated can be used as a batch. Then, the feature extraction model is called to extract features from each sample speech sub-segment to be evaluated, and the first sample Mel-spectrogram feature vector corresponding to each sample speech sub-segment to be evaluated is obtained. Then, the downsampling model is called to downsample the first sample Mel-spectrogram feature vector corresponding to each sample speech sub-segment to be evaluated, and the first sample embedded feature vector corresponding to each sample speech sub-segment to be evaluated is obtained. Through the above processing, one sample speech segment to be evaluated corresponds to a preset number of first sample embedded feature vectors.

[0063] Assume that the preset number is B, the preset duration is T, the dimension of the first sample Mel spectrum feature vector is F, and the dimension of the first sample embedded feature vector is H. Then, after the sample speech segment to be evaluated is processed by the preprocessing process and the feature extraction model, [B, T, 1, F] can be obtained. After [B, T, 1, F] is processed by the downsampling model, [B, T, H] can be obtained.

[0064] 303. Process the sample clean speech segment to obtain a second sample embedded feature vector.

[0065] Among them, the second sample embedded feature vector is used to characterize the voiceprint feature distribution of the sample clean voice segment in the frequency domain in the audio and video communication scenario at the time of acquisition. Specifically, when processing the sample clean voice segment to obtain the second sample embedded feature vector, the feature extraction model can be called to perform feature extraction on the sample clean voice segment to obtain the second sample Mel spectrum feature vector, and then the downsampling model can be called to downsample the second sample Mel spectrum feature vector to obtain the second sample embedded feature vector. The feature extraction model and the downsampling model can adopt the models that have been trained in the relevant technology. The dimension of the second sample Mel spectrum feature vector is the same as the dimension of the first sample Mel spectrum feature vector, and the dimension of the second sample embedded feature vector is the same as the dimension of the first sample embedded feature vector.

[0066] Furthermore, the duration of the sample clean speech segment usually obtained is long, and the computational complexity of processing the sample clean speech segment with a long duration is large and the processing speed is slow. As described in step 301, the duration of the sample speech segment to be evaluated and the sample clean speech segment can be the same or different. If the duration of the sample speech segment to be evaluated is different from the duration of the sample clean speech segment, the length of the embedded feature vector obtained by processing speech segments of different durations must be different. Similarity calculation cannot be performed based on embedded feature vectors of different lengths, while similarity calculation can improve the accuracy of the model. Therefore, in order to reduce the computational complexity of the model training process and improve the training speed and accuracy of the model, the embodiment of the present application can pre-process the sample clean speech segment before processing the sample clean speech segment. Specifically, the number of sub-segments can be pre-set to a preset number, and then the sample clean speech segment is randomly segmented using a window of a preset duration. When a preset number of sample clean speech sub-segments are obtained, the segmentation is stopped, and the preset number of sample clean speech sub-segments can be used as a batch. Then the feature extraction model is called to perform feature extraction on each sample clean speech sub-segment to obtain the second sample Mel spectrum feature vector corresponding to each sample clean speech sub-segment. Then the downsampling model is called to downsample the second sample Mel spectrum feature vector corresponding to each sample clean speech sub-segment to obtain the second sample embedded feature vector corresponding to each sample clean speech sub-segment. Through the above processing process, one sample clean speech segment corresponds to a preset number of second sample embedded feature vectors.

[0067] Set the preset number to B, the preset time length to T, the dimension of the second sample Mel spectrum feature vector to F, and the dimension of the second sample embedded feature vector to H. After the sample clean speech segment is processed by the preprocessing process and the feature extraction model, [B, T, 1, F] can be obtained. After [B, T, 1, F] is processed by the downsampling model, [B, T, H] can be obtained.

[0068] 304. Concatenate the first sample embedded feature vector and the second sample embedded feature vector to obtain a third sample embedded feature vector.

[0069] When splicing the first sample embedded feature vector and the second sample embedded feature vector, one of the feature vectors can be spliced ​​behind the other feature vector, such as splicing the first sample embedded feature vector behind the second sample embedded feature vector, or splicing the second sample embedded feature vector behind the first sample embedded feature vector. Of course, if the sample clean speech segment is divided into B sample clean speech sub-segments and the sample to-be-evaluated speech segment is divided into B sample to-be-evaluated speech sub-segments, then when splicing the first sample embedded feature vector and the second sample embedded feature vector, the B second sample embedded feature vectors corresponding to the B sample clean speech sub-segments can be spliced ​​one by one with the B first sample embedded feature vectors corresponding to the B sample to-be-evaluated speech sub-segments to obtain B third sample embedded feature vectors. For example, [B, T, H] corresponding to the sample to-be-evaluated speech segment is spliced ​​with [B, T, H] corresponding to the sample clean speech segment to obtain [B, T, H*2], and the number of dimensions of each third sample embedded feature vector increases.

[0070] The embodiment of the present application enriches the information contained in the third sample embedded feature vector by splicing the first sample embedded feature vector and the second sample embedded feature vector, so that the third sample embedded feature vector integrates the voiceprint features of the user in different audio and video communication scenarios. Training the model based on the third sample embedded feature vector can improve the robustness of the model and enable the model to learn the voiceprint features of clean voice clips.

[0071] 305. Based on the first sample embedded feature vector, the second sample embedded feature vector, the third sample embedded feature vector and the total target loss function, the initial voice communication quality assessment model is trained to obtain a voice communication quality assessment model.

[0072] The initial voice communication quality assessment model is the voice communication quality assessment model to be trained. The initial voice communication quality assessment model has the same structure as the trained voice communication quality assessment model, except for the model parameters at each layer. The training process of the initial voice communication quality assessment model is the process of adjusting the model parameters at each layer of the initial voice communication quality assessment model. The specific training process includes the following steps:

[0073] 3051. Input the first sample embedded feature vector and the second sample embedded feature vector into the initial voice communication quality assessment model respectively. After being processed by the linear mapping layer of the initial voice communication quality assessment model, the first sample voiceprint feature vector and the second sample voiceprint feature vector are output.

[0074] In an embodiment of the present application, the initial voice communication quality assessment model includes a linear mapping layer and a pooling layer following the linear mapping layer. After obtaining a first sample embedded feature vector and a second sample embedded feature vector, the first sample embedded feature vector and the second sample embedded feature vector are respectively input into the initial voice communication quality assessment model. After being processed by the linear mapping layer of the initial voice communication quality assessment model without undergoing a pooling operation in the pooling layer, a first sample voiceprint feature vector and a second sample voiceprint feature vector are obtained. The linear mapping layer employs an attention mechanism to focus more on voiceprint-related features during the linear mapping process, thereby enhancing the voiceprint-related features in the first sample voiceprint feature vector and the second sample voiceprint feature vector, thereby improving the ability to recognize the voiceprint in the first sample voiceprint feature vector.

[0075] 3052. Input the third sample embedded feature vector into the initial voice communication quality assessment model, and output the generated MOS corresponding to the sample speech segment to be assessed.

[0076] In an embodiment of the present application, the third sample embedded feature vector is input into the initial voice communication quality assessment model, and after being processed by the linear mapping layer of the initial voice communication quality assessment model, the third sample voiceprint feature vector is output, and then after being processed by the pooling layer of the initial voice communication quality assessment model, the generated MOS corresponding to the sample voice segment to be evaluated is output.

[0077] 3053. Based on the generated MOS and labeled MOS corresponding to the first sample voiceprint feature vector and the second sample voiceprint feature vector, the sample speech segment to be evaluated, and the total target loss function, the model parameters of the initial voice communication quality assessment model are adjusted to obtain a voice communication quality assessment model.

[0078] Among them, the total target loss function includes the cosine similarity loss function and the quality score loss function. Cosine similarity is used to measure the similarity between two feature vectors. The closer the two feature vectors are, the greater the cosine similarity is. The quality score loss function is used to measure the difference between the MOS generated by the model and the annotated MOS. Specifically, based on the generated MOS and annotated MOS corresponding to the first sample voiceprint feature vector and the second sample voiceprint feature vector, the sample speech segment to be evaluated, and the total target loss function, the model parameters of the initial voice communication quality assessment model are adjusted to obtain the voice communication quality assessment model, including the following steps:

[0079] The first step is to calculate the cosine similarity between the first sample voiceprint feature vector and the second sample voiceprint feature vector to obtain a cosine similarity value.

[0080] Among them, the calculation formula of cosine similarity is:

[0081] in, is the cosine similarity value, is the second sample voiceprint feature vector, is the first sample voiceprint feature vector, and N is the number of dimensions of the second sample voiceprint feature vector (ie, H).

[0082] In the second step, the cosine similarity value is input into the cosine similarity loss function to obtain the cosine similarity loss function value.

[0083] In the third step, the generated MOS and the annotated MOS corresponding to the sample speech segment to be evaluated are input into the quality score loss function to obtain the quality score loss function value.

[0084] The fourth step is to perform weighted calculation on the cosine similarity loss function value and the quality score loss function value to obtain the total loss function value.

[0085] In the fifth step, based on the function value of the total loss function, the model parameters of the initial voice communication quality evaluation model are adjusted to obtain the voice communication quality evaluation model.

[0086] When the function value of the total loss function is obtained, the function value of the total loss function is compared with a preset threshold. If the function value of the total loss function is greater than the preset threshold, the model parameters of the initial voice communication quality assessment model are adjusted, and step 305 is executed based on the voice communication quality assessment model after the model parameters are adjusted until the training cutoff condition is met. The training cutoff condition includes that the function value of the total loss function obtained is less than the preset threshold, or the number of model parameter adjustments reaches a set number of times. The model parameters obtained when the training cutoff condition is met are obtained, and the model parameters are assigned to the initial voice communication quality assessment model to obtain a trained voice communication quality assessment model. The voice communication quality assessment model is used to compare the voiceprint features corresponding to the clean voice segment in the third embedded feature vector with the voiceprint features corresponding to the voice segment to be evaluated, and score the voice segment to be evaluated based on the comparison result and the MOS of the clean voice segment.

[0087] Through the above training process, the voice communication quality assessment model can learn how to compare the voiceprint features of the same user in different audio and video communication scenarios, and how to score the voice clips to be evaluated based on the comparison results. The trained voice communication quality assessment model is more accurate.

[0088] Furthermore, for the cosine similarity loss function and the quality score loss function included in the overall objective loss function, embodiments of the present application can be calculated in parallel in different threads to increase the training speed of the model. Specifically, the cosine similarity loss function value can be calculated in the speaker voiceprint extraction thread, and the quality score loss function value can be calculated in the voice quality score prediction thread. Each thread can call the feature extraction model, the downsampling model, and the initial voice communication quality assessment model to be trained. Since the two loss functions are calculated in parallel, the training time of the model is greatly shortened, and the training efficiency of the voice communication quality assessment model is improved. The training process of the voice communication quality assessment model using different threads will be explained below in conjunction with Figure 4.

[0089] Referring to Figure 4 , in the speaker voiceprint feature extraction thread, the feature extraction model is called to perform feature extraction on the user's sample clean speech segment to obtain a second sample Mel-spectrogram feature vector. The downsampling model is then called to process the second sample Mel-spectrogram feature vector to obtain a second sample embedded feature vector. Once the second sample embedded feature vector is obtained, the speaker voiceprint feature extraction thread can provide the second sample embedded feature vector to the voice quality score prediction thread. The speaker voiceprint feature extraction thread can also call the initial voice communication quality assessment model to be trained to process the second sample embedded feature vector to obtain a second sample voiceprint feature vector. In the voice quality score prediction thread, the feature extraction model can be called to perform feature extraction on the sample speech segment to be evaluated from the same user to obtain a first sample Mel-spectrogram feature vector. The downsampling model is then called to process the first sample Mel-spectrogram feature vector to obtain a first sample embedded feature vector. When the first sample embedded feature vector is obtained, the voice quality score prediction thread can provide it to the speaker voiceprint feature extraction thread, and then in the speaker voiceprint feature extraction thread, the initial voice communication quality assessment model to be trained is called to process the first sample embedded feature vector to obtain the first sample voiceprint feature vector, and then calculate the cosine similarity between the first sample voiceprint feature vector and the second sample voiceprint feature vector to obtain the cosine similarity value, and then obtain the cosine similarity loss function value based on the cosine similarity value. In the speech quality score prediction thread, the second sample embedded feature vector and the first sample embedded feature vector are spliced ​​to obtain a third sample embedded feature vector, and then the initial voice communication quality assessment model to be trained is called to process the third sample embedded feature vector to obtain the generated MOS corresponding to the sample speech segment to be evaluated, and then based on the generated MOS and the labeled MOS corresponding to the sample speech segment to be evaluated, the quality score loss function value is calculated, and then the cosine similarity loss function value and the quality score loss function value are weighted to obtain the function value of the total loss function, and then based on the total loss function value, the model parameters of the initial voice communication quality assessment model are adjusted to obtain a trained voice communication quality assessment model.

[0090] Furthermore, the voice communication quality assessment model in the embodiment of the present application is trained based on sample data of different problems, and can not only provide the MOS of the voice segment to be evaluated, but also identify the problems existing in the current audio and video communication scenario.

[0091] All of the above optional technical solutions can be combined in any way to form optional embodiments of the present application, and will not be described in detail here.

[0092] In the embodiment shown in FIG3 , the feature extraction model and the downsampling model use models that have been trained in the relevant technology. The accuracy of the embedded feature vectors obtained by these two models depends on the accuracy of the model trained by the relevant technology. In order to improve the accuracy of the trained voice communication quality assessment, the embodiment of the present application provides a method for training a voice communication quality assessment model, which trains the feature extraction model, the downsampling model and the voice communication quality assessment model together. Taking the server executing the embodiment of the present application as an example, the server can be the server 103 in FIG1 or other servers. Referring to FIG5 , the method flow provided in the embodiment of the present application includes:

[0093] 501. Obtain a sample clean voice segment and a sample to-be-evaluated voice segment from the same user.

[0094] This step is executed in the same manner as the above step 301. Please refer to the above step 301 for details and will not be repeated here.

[0095] 502. Call the initial feature extraction model to perform feature extraction on the sample speech segment to be evaluated to obtain a first sample Mel-spectrogram feature vector, and call the initial downsampling model to downsample the first sample Mel-spectrogram feature vector to obtain a first sample embedded feature vector.

[0096] 503. Call the initial feature extraction model to perform feature extraction on the sample clean speech segment to obtain a second sample Mel spectrum feature vector, and call the initial downsampling model to downsample the second sample Mel spectrum feature vector to obtain a second sample embedded feature vector.

[0097] 504. Concatenate the first sample embedded feature vector and the second sample embedded feature vector to obtain a third sample embedded feature vector.

[0098] This step is executed in the same manner as the above step 304. Please refer to the above step 304 for details, which will not be repeated here.

[0099] 505. Based on the first sample embedded feature vector, the second sample embedded feature vector, the third sample embedded feature vector and the total target loss function, the initial feature extraction model, the initial downsampling model and the initial voice communication quality assessment model are synchronously trained to obtain the feature extraction model, the downsampling model and the voice communication quality assessment model.

[0100] Specifically, the following methods can be used:

[0101] 5051. Input the first sample embedded feature vector and the second sample embedded feature vector into the initial voice communication quality assessment model respectively. After being processed by the linear mapping layer of the initial voice communication quality assessment model, the first sample voiceprint feature vector and the second sample voiceprint feature vector are output.

[0102] 5052. Input the third sample embedded feature vector into the initial voice communication quality assessment model, and output the generated MOS corresponding to the sample speech segment to be assessed.

[0103] 5053. Based on the generated MOS and labeled MOS corresponding to the first sample voiceprint feature vector and the second sample voiceprint feature vector, the sample speech segment to be evaluated, and the total target loss function, the initial feature extraction model, the initial downsampling model, and the initial voice communication quality assessment model are synchronously trained to obtain the feature extraction model, the downsampling model, and the voice communication quality assessment model.

[0104] The specific steps include:

[0105] The first step is to calculate the cosine similarity between the first sample voiceprint feature vector and the second sample voiceprint feature vector to obtain a cosine similarity value.

[0106] Among them, the calculation formula of cosine similarity is:

[0107] in, is the cosine similarity value, is the second sample voiceprint feature vector, is the first sample voiceprint feature vector, and N is the number of dimensions of the second sample voiceprint feature vector (ie, H).

[0108] In the second step, the cosine similarity value is input into the cosine similarity loss function to obtain the cosine similarity loss function value.

[0109] In the third step, the generated MOS and the annotated MOS corresponding to the sample speech segment to be evaluated are input into the quality score loss function to obtain the quality score loss function value.

[0110] The fourth step is to perform weighted calculation on the cosine similarity loss function value and the quality score loss function value to obtain the total loss function value.

[0111] In the fifth step, based on the function value of the total loss function, the model parameters of the initial feature extraction model, the initial downsampling model and the initial voice communication quality assessment model are adjusted to obtain the feature extraction model, the downsampling model and the voice communication quality assessment model.

[0112] When the function value of the total loss function is obtained, the function value of the total loss function is compared with a preset threshold. If the function value of the total loss function is greater than the preset threshold, the model parameters of the initial feature extraction model, the initial downsampling model, and the initial voice communication quality assessment model are adjusted, and this step is repeated based on the feature extraction model after the model parameters are adjusted, the downsampling model after the model parameters are adjusted, and the voice communication quality assessment model after the model parameters are adjusted until the training cutoff condition is met. The training cutoff condition includes that the function value of the obtained total loss function is less than the preset threshold, or the number of model parameter adjustments reaches a set number of times. The model parameters obtained when the training cutoff condition is met are obtained, and the model parameters are assigned to the initial feature extraction model, the initial downsampling model, and the initial voice communication quality assessment model to obtain a trained feature extraction model, a trained downsampling model, and a trained voice communication quality assessment model.

[0113] Of course, in order to shorten the training time of the feature extraction model, the downsampling model and the voice communication quality assessment model, the cosine similarity loss function and the quality score loss function included in the total target loss function can also be calculated in parallel in different threads. Please refer to the above embodiment for details, which will not be repeated here.

[0114] Furthermore, the voice communication quality assessment model in the embodiment of the present application is trained based on sample data of different problems, and can not only provide the MOS of the voice segment to be evaluated, but also identify the problems existing in the current audio and video communication scenario.

[0115] All of the above optional technical solutions can be combined in any way to form optional embodiments of the present application, and will not be described in detail here.

[0116] This embodiment of the present application provides a method for evaluating voice communication quality. Taking the server 103 in FIG. 1 as an example, the server 103 may be installed with a voice communication quality evaluation model trained in the embodiment shown in FIG. 3 or FIG. 5 . Referring to FIG. 6 , the method flow provided in this embodiment of the present application includes:

[0117] 601. During audio and video communication with a target user, obtain a target user's voice segment to be evaluated.

[0118] The target user is the speaker in an audio or video communication scenario. To ensure communication quality, the server can evaluate the voice communication quality of all parties, including the target user, in real time or at regular intervals. To evaluate the voice communication quality of the target user's audio or video communication, after receiving the target user's voice data, the server can extract a voice segment from the voice data as the voice segment to be evaluated.

[0119] 602. Process the speech segment to be evaluated to obtain a first embedded feature vector.

[0120] The first embedded feature vector is used to represent the voiceprint feature distribution in the frequency domain of the speech segment to be evaluated in the current audio and video communication scenario. When the speech segment to be evaluated is obtained, the server can call the feature extraction model to extract features from the speech segment to obtain a first mel-spectrogram feature vector. The server can then call the downsampling model to process the first mel-spectrogram feature vector to obtain a first embedded feature vector.

[0121] Taking into account the long duration of the voice segment to be evaluated and the high real-time requirements of the audio and video communication scenario, in order to shorten the evaluation time of the voice communication quality of the target user's current audio and video communication, before processing the voice segment to be evaluated, the server can also obtain a target sub-segment of a preset length from the voice segment to be evaluated, and then call the feature extraction model and the downsampling model in sequence to process the target sub-segment to obtain the first embedded feature vector.

[0122] Of course, in order to improve the accuracy of the voice communication quality evaluation results of the target user's current audio and video communication, the server can also use a window of preset length to divide the voice segment to be evaluated, obtain a preset number of target sub-segments, and then call the feature extraction model and the downsampling model in sequence to process the preset number of target sub-segments to obtain a preset number of first embedded feature vectors.

[0123] 603. Obtain a second embedded feature vector.

[0124] The second embedded feature vector is obtained by processing the target user's clean voice segment. The second embedded feature vector is used to characterize the frequency domain voiceprint feature distribution of the clean voice segment in an interference-free audio and video communication scenario. The duration of the voice segment to be evaluated can be the same as or different from the duration of the clean voice segment, and the clean voice segment has the highest voice quality score (MOS).

[0125] In order to be able to evaluate the voice communication quality of users in audio and video communication scenarios, before the user communicates based on an online communication application, a clean voice clip of the user can be obtained, and a feature extraction model can be called to extract features from the clean voice clip of the user to obtain the second mel-spectrogram feature vector corresponding to the user. The downsampling model is then called to process the second mel-spectrogram feature vector corresponding to the user to obtain a second embedded feature vector, and then the second embedded feature vector corresponding to the user is stored. To facilitate the management of the second embedded feature vectors corresponding to different users, when storing the second embedded feature vector corresponding to the user, the server can store the user account registered by the user in the online communication application together. That is, the server can store the correspondence between each user account and the second embedded feature vector.

[0126] Based on the correspondence between each user account and the second embedded feature vector stored in the server, the server may obtain the second embedded feature vector corresponding to the user account of the target user.

[0127] 604. Concatenate the first embedded feature vector and the second embedded feature vector to obtain a third embedded feature vector.

[0128] In a possible implementation, if the number of the first embedded feature vector is one, the first embedded feature vector and the second embedded feature vector may be directly concatenated to obtain a third embedded feature vector.

[0129] In another possible implementation, the number of first embedded feature vectors is a preset number, and the preset number of first embedded feature vectors and second embedded feature vectors may be concatenated to obtain a preset number of third embedded feature vectors.

[0130] 605. Call the voice communication quality evaluation model to process the third embedded feature vector to obtain the MOS of the voice segment to be evaluated.

[0131] In a possible implementation, the number of the third embedded feature vector is one, and the server calls a voice communication quality evaluation model to process the third embedded feature vector to obtain the MOS of the voice segment to be evaluated.

[0132] In another possible implementation, the number of third embedded feature vectors is a preset number, the server calls the voice communication quality evaluation model, processes each third embedded feature vector, obtains a preset number of MOS, and then calculates the average value of the preset number of MOS to obtain an average MOS, and uses the average MOS as the MOS of the voice segment to be evaluated.

[0133] When the MOS of the voice segment to be evaluated is obtained, the server can compare the MOS with the preset score. If the MOS is lower than the preset score, it can be determined that the current voice communication quality of the target user is poor, and then a prompt message is sent to the target user. Upon receiving the prompt message, the target user can suspend the voice conversation until the voice communication quality improves.

[0134] Furthermore, the voice communication quality assessment model invoked by the server is trained using sample data from different problems. It can output not only the MOS of the voice segment being evaluated, but also the problems existing in the current audio and video communication scenario. This allows the server to include factors affecting voice communication quality in the prompt message sent to the target user. The method provided in this embodiment of the application can shorten the time it takes to discover problems, allowing for rapid problem location even if the user does not report them.

[0135] In addition, to improve service quality, during audio and video communication, the server usually uses audio algorithms to improve communication quality. The improved results can be verified by calling the voice communication quality assessment model, thus providing a closed-loop verification tool.

[0136] Please refer to FIG7 , which shows a schematic diagram of the structure of a voice communication quality evaluation device provided in an embodiment of the present application. The device can be implemented by software, hardware, or a combination of both, and becomes all or part of a server. The device includes:

[0137] The acquisition module 701 is configured to acquire a target user's voice segment to be evaluated during the target user's audio and video communication process;

[0138] A processing module 702 is configured to process the speech segment to be evaluated to obtain a first embedded feature vector, where the first embedded feature vector is used to represent the voiceprint feature distribution of the speech segment to be evaluated in the frequency domain in the current audio and video communication scenario;

[0139] Acquisition module 701 is configured to acquire a second embedded feature vector, where the second embedded feature vector is a feature vector obtained by processing a clean voice segment of the target user. The second embedded feature vector is used to represent the voiceprint feature distribution of the clean voice segment in the frequency domain in an interference-free audio and video communication scenario. The clean voice segment has the highest voice quality score (MOS).

[0140] A concatenation module 703 is configured to concatenate the first embedded feature vector and the second embedded feature vector to obtain a third embedded feature vector;

[0141] The processing module 702 is configured to call a voice communication quality assessment model to process the third embedded feature vector to obtain the MOS of the voice segment to be evaluated. The voice communication quality assessment model is used to compare the voiceprint features corresponding to the clean voice segment in the third embedded feature vector with the voiceprint features corresponding to the voice segment to be evaluated, and score the voice segment to be evaluated based on the comparison result and the MOS of the clean voice segment.

[0142] In another embodiment of the present application, the device further comprises:

[0143] The acquisition module 701 is further configured to acquire a sample clean speech segment and a sample to-be-evaluated speech segment from the same user, wherein the sample clean speech segment has the highest MOS;

[0144] The processing module 702 is further configured to process the sample speech segment to be evaluated to obtain a first sample embedded feature vector;

[0145] The processing module 702 is further configured to process the sample clean speech segment to obtain a second sample embedded feature vector;

[0146] The concatenation module 703 is further configured to concatenate the first sample embedded feature vector and the second sample embedded feature vector to obtain a third sample embedded feature vector;

[0147] The training module is configured to train the initial voice communication quality assessment model based on the first sample embedded feature vector, the second sample embedded feature vector, the third sample embedded feature vector and the total target loss function to obtain the voice communication quality assessment model.

[0148] In another embodiment of the present application, the sample speech segment to be evaluated is annotated with MOS, and the training module is configured to input the first sample embedded feature vector and the second sample embedded feature vector into the initial voice communication quality evaluation model respectively, and output the first sample voiceprint feature vector and the second sample voiceprint feature vector after processing by the linear mapping layer of the initial voice communication quality evaluation model; input the third sample embedded feature vector into the initial voice communication quality evaluation model, and output the generated MOS corresponding to the sample speech segment to be evaluated; based on the first sample voiceprint feature vector and the second sample voiceprint feature vector, the generated MOS and annotated MOS corresponding to the sample speech segment to be evaluated, and the total target loss function, the model parameters of the initial voice communication quality evaluation model are adjusted to obtain a voice communication quality evaluation model.

[0149] In another embodiment of the present application, the total target loss function includes a cosine similarity loss function and a quality score loss function, and the training module is configured to calculate the cosine similarity of the first sample voiceprint feature vector and the second sample voiceprint feature vector to obtain a cosine similarity value; input the cosine similarity value into the cosine similarity loss function to obtain a cosine similarity loss function value; input the generated MOS and the labeled MOS corresponding to the sample speech segment to be evaluated into the quality score loss function to obtain a quality score loss function value; perform weighted calculation on the cosine similarity loss function value and the quality score loss function value to obtain a total loss function value; based on the function value of the total loss function, adjust the model parameters of the initial voice communication quality evaluation model to obtain a voice communication quality evaluation model.

[0150] In another embodiment of the present application, the processing module is configured to call a feature extraction model to perform feature extraction on the speech segment to be evaluated to obtain a first Mel-spectrogram feature vector; and call a downsampling model to downsample the first Mel-spectrogram feature vector to obtain a first embedded feature vector.

[0151] In another embodiment of the present application, the device further comprises:

[0152] The acquisition module is further configured to acquire a sample clean speech segment and a sample to-be-evaluated speech segment from the same user, wherein the sample clean speech segment has the highest MOS;

[0153] The processing module is further configured to call the initial feature extraction model to perform feature extraction on the sample speech segment to be evaluated to obtain a first sample Mel spectrum feature vector, and call the initial downsampling model to downsample the first sample Mel spectrum feature vector to obtain a first sample embedded feature vector;

[0154] The processing module is further configured to call the initial feature extraction model to perform feature extraction on the sample clean speech segment to obtain a second sample Mel spectrum feature vector, and call the initial downsampling model to downsample the second sample Mel spectrum feature vector to obtain a second sample embedded feature vector;

[0155] The splicing module is further configured to splice the first sample embedded feature vector and the second sample embedded feature vector to obtain a third sample embedded feature vector;

[0156] The training module is configured to synchronously train the initial feature extraction model, the initial downsampling model and the initial voice communication quality assessment model based on the first sample embedded feature vector, the second sample embedded feature vector, the third sample embedded feature vector and the total target loss function to obtain the feature extraction model, the downsampling model and the voice communication quality assessment model.

[0157] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0158] FIG8 shows a structural block diagram of a server 800 provided by an exemplary embodiment of the present application. Generally, the server 800 includes: a processor 801 and a memory 802 .

[0159] The processor 801 can be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 801 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state; the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 801 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 801 may also include an artificial intelligence processor, which is used to process computing operations related to machine learning.

[0160] The memory 802 may include one or more computer-readable storage media, which may be non-transitory computer-readable storage media, such as CD-ROMs (Compact Disc Read-Only Memory), ROMs, RAMs (Random Access Memory), magnetic tapes, floppy disks, and optical data storage devices. The computer-readable storage media may store at least one computer program, which, when executed, can implement the method for evaluating voice communication quality.

[0161] Of course, the aforementioned server may also include other components, such as input / output interfaces and communication components. The input / output interface provides an interface between the processor and peripheral interface modules, which may be output devices, input devices, etc. The communication components are configured to facilitate wired or wireless communication between the server and other devices.

[0162] Those skilled in the art will appreciate that the structure shown in FIG8 does not limit the server 800 and may include more or fewer components than shown, or combine certain components, or adopt a different component arrangement.

[0163] An embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores at least one computer program, and when the at least one computer program is executed by a processor, the above-mentioned method for evaluating voice communication quality can be implemented.

[0164] An embodiment of the present application provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it can implement the above-mentioned method for evaluating voice communication quality.

[0165] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for evaluating voice communication quality, the method comprising: During the audio and video communication process of the target user, obtaining the target user's voice clip to be evaluated; Processing the speech segment to be evaluated to obtain a first embedded feature vector, where the first embedded feature vector is used to characterize a voiceprint feature distribution of the speech segment to be evaluated in the frequency domain in a current audio and video communication scenario; Obtaining a second embedded feature vector, where the second embedded feature vector is a feature vector obtained by processing a clean voice segment of the target user, and the second embedded feature vector is used to represent the voiceprint feature distribution of the clean voice segment in the frequency domain in an interference-free audio and video communication scenario, wherein the clean voice segment has the highest voice quality score (MOS); concatenating the first embedded feature vector and the second embedded feature vector to obtain a third embedded feature vector; A voice communication quality assessment model is called to process the third embedded feature vector to obtain the MOS of the voice segment to be evaluated. The voice communication quality assessment model is used to compare the voiceprint features corresponding to the clean voice segment in the third embedded feature vector with the voiceprint features corresponding to the voice segment to be evaluated, and score the voice segment to be evaluated based on the comparison result and the MOS of the clean voice segment.

2. The method according to claim 1, wherein Before calling the voice communication quality evaluation model and processing the third embedded feature vector to obtain the MOS of the voice segment to be evaluated, the method further includes: Obtaining a sample clean speech segment and a sample speech segment to be evaluated from the same user, wherein the sample clean speech segment has the highest MOS; Processing the sample speech segment to be evaluated to obtain a first sample embedded feature vector; Processing the sample clean speech segment to obtain a second sample embedded feature vector; concatenating the first sample embedded feature vector and the second sample embedded feature vector to obtain a third sample embedded feature vector; Based on the first sample embedded feature vector, the second sample embedded feature vector, the third sample embedded feature vector and the total target loss function, the initial voice communication quality assessment model is trained to obtain the voice communication quality assessment model.

3. The method according to claim 2, wherein: The sample speech segment to be evaluated is annotated with MOS, and the initial voice communication quality evaluation model is trained based on the first sample embedded feature vector, the second sample embedded feature vector, the third sample embedded feature vector, and the total target loss function to obtain the voice communication quality evaluation model, including: Inputting the first sample embedded feature vector and the second sample embedded feature vector into the initial voice communication quality assessment model respectively, and outputting a first sample voiceprint feature vector and a second sample voiceprint feature vector after being processed by the linear mapping layer of the initial voice communication quality assessment model; Inputting the third sample embedded feature vector into the initial voice communication quality assessment model, and outputting a generated MOS corresponding to the sample speech segment to be assessed; Based on the first sample voiceprint feature vector and the second sample voiceprint feature vector, the generated MOS and the labeled MOS corresponding to the sample speech segment to be evaluated, and the total target loss function, the model parameters of the initial voice communication quality assessment model are adjusted to obtain the voice communication quality assessment model.

4. The method according to claim 3, wherein: The total objective loss function includes a cosine similarity loss function and a quality score loss function. The generated MOS and annotated MOS corresponding to the first sample voiceprint feature vector and the second sample voiceprint feature vector, the sample speech segment to be evaluated, and the total objective loss function are adjusted to obtain the voice communication quality assessment model by adjusting the model parameters of the initial voice communication quality assessment model, including: Calculating the cosine similarity between the first sample voiceprint feature vector and the second sample voiceprint feature vector to obtain a cosine similarity value; Inputting the cosine similarity value into the cosine similarity loss function to obtain a cosine similarity loss function value; Inputting the generated MOS and the annotated MOS corresponding to the sample speech segment to be evaluated into the quality score loss function to obtain a quality score loss function value; Performing weighted calculation on the cosine similarity loss function value and the quality score loss function value to obtain a total loss function value; Based on the function value of the total loss function, the model parameters of the initial voice communication quality assessment model are adjusted to obtain the voice communication quality assessment model.

5. The method according to claim 1, wherein The processing of the speech segment to be evaluated to obtain a first embedded feature vector includes: Calling a feature extraction model to perform feature extraction on the speech segment to be evaluated to obtain a first Mel-spectrogram feature vector; A downsampling model is called to downsample the first Mel-spectrogram feature vector to obtain the first embedded feature vector.

6. The method according to claim 5, wherein: Before calling the voice communication quality evaluation model and processing the third embedded feature vector to obtain the MOS of the voice segment to be evaluated, the method further includes: Obtaining a sample clean speech segment and a sample speech segment to be evaluated from the same user, wherein the sample clean speech segment has the highest MOS; Calling an initial feature extraction model to perform feature extraction on the sample speech segment to be evaluated to obtain a first sample Mel spectrum feature vector, and calling an initial downsampling model to downsample the first sample Mel spectrum feature vector to obtain a first sample embedded feature vector; Calling the initial feature extraction model to perform feature extraction on the sample clean speech segment to obtain a second sample Mel spectrum feature vector, and calling the initial downsampling model to downsample the second sample Mel spectrum feature vector to obtain a second sample embedded feature vector; concatenating the first sample embedded feature vector and the second sample embedded feature vector to obtain a third sample embedded feature vector; Based on the first sample embedded feature vector, the second sample embedded feature vector, the third sample embedded feature vector and the total objective loss function, the initial feature extraction model, the initial downsampling model and the initial voice communication quality assessment model are synchronously trained to obtain the feature extraction model, the downsampling model and the voice communication quality assessment model.

7. A device for evaluating voice communication quality, the device comprising: An acquisition module is configured to acquire a speech segment to be evaluated from a target user during audio and video communication with the target user; a processing module configured to process the speech segment to be evaluated to obtain a first embedded feature vector, wherein the first embedded feature vector is used to represent the voiceprint feature distribution of the speech segment to be evaluated in the frequency domain in the current audio and video communication scenario; The acquisition module is configured to acquire a second embedded feature vector, where the second embedded feature vector is a feature vector obtained by processing a clean voice segment of the target user, the second embedded feature vector being used to characterize a voiceprint feature distribution of the clean voice segment in a frequency domain in an interference-free audio and video communication scenario, and the clean voice segment has the highest voice quality score (MOS); a concatenation module configured to concatenate the first embedded feature vector and the second embedded feature vector to obtain a third embedded feature vector; The processing module is configured to call a voice communication quality assessment model to process the third embedded feature vector to obtain the MOS of the voice segment to be evaluated. The voice communication quality assessment model is used to compare the voiceprint features corresponding to the clean voice segment in the third embedded feature vector with the voiceprint features corresponding to the voice segment to be evaluated, and score the voice segment to be evaluated based on the comparison result and the MOS of the clean voice segment.

8. The device according to claim 7, wherein The device further comprises: The acquisition module is further configured to acquire a sample clean speech segment and a sample speech segment to be evaluated from the same user, wherein the sample clean speech segment has the highest MOS; The processing module is further configured to process the sample speech segment to be evaluated to obtain a first sample embedded feature vector; The processing module is further configured to process the sample clean speech segment to obtain a second sample embedded feature vector; The splicing module is further configured to splice the first sample embedded feature vector and the second sample embedded feature vector to obtain a third sample embedded feature vector; The training module is configured to train the initial voice communication quality assessment model based on the first sample embedded feature vector, the second sample embedded feature vector, the third sample embedded feature vector and the total target loss function to obtain the voice communication quality assessment model.

9. The device according to claim 8, wherein The sample speech segment to be evaluated is annotated with MOS. The training module is configured to input the first sample embedded feature vector and the second sample embedded feature vector into the initial voice communication quality evaluation model respectively, and output the first sample voiceprint feature vector and the second sample voiceprint feature vector after processing by the linear mapping layer of the initial voice communication quality evaluation model; input the third sample embedded feature vector into the initial voice communication quality evaluation model, and output the generated MOS corresponding to the sample speech segment to be evaluated; based on the first sample voiceprint feature vector and the second sample voiceprint feature vector, the generated MOS and the annotated MOS corresponding to the sample speech segment to be evaluated, and the total target loss function, adjust the model parameters of the initial voice communication quality evaluation model to obtain the voice communication quality evaluation model.

10. The device according to claim 9, wherein The total target loss function includes a cosine similarity loss function and a quality score loss function. The training module is configured to calculate the cosine similarity between the first sample voiceprint feature vector and the second sample voiceprint feature vector to obtain a cosine similarity value; and input the cosine similarity value into the cosine similarity loss function to obtain a cosine similarity loss function value. The generated MOS and the labeled MOS corresponding to the sample speech segment to be evaluated are input into the quality score loss function to obtain a quality score loss function value; the cosine similarity loss function value and the quality score loss function value are weightedly calculated to obtain a total loss function value; based on the function value of the total loss function, the model parameters of the initial voice communication quality evaluation model are adjusted to obtain the voice communication quality evaluation model.

11. The device according to claim 7, wherein The processing module is configured to call a feature extraction model to perform feature extraction on the speech segment to be evaluated to obtain a first Mel-spectrogram feature vector; and call a downsampling model to downsample the first Mel-spectrogram feature vector to obtain the first embedded feature vector.

12. The device according to claim 11, wherein The device further comprises: The acquisition module is further configured to acquire a sample clean speech segment and a sample speech segment to be evaluated from the same user, wherein the sample clean speech segment has the highest MOS; The processing module is further configured to call an initial feature extraction model to perform feature extraction on the sample speech segment to be evaluated to obtain a first sample Mel-spectrogram feature vector, and call an initial downsampling model to downsample the first sample Mel-spectrogram feature vector to obtain a first sample embedded feature vector; The processing module is further configured to call the initial feature extraction model to perform feature extraction on the sample clean speech segment to obtain a second sample Mel spectrum feature vector, and call the initial downsampling model to downsample the second sample Mel spectrum feature vector to obtain a second sample embedded feature vector; The splicing module is further configured to splice the first sample embedded feature vector and the second sample embedded feature vector to obtain a third sample embedded feature vector; The training module is configured to synchronously train the initial feature extraction model, the initial downsampling model and the initial voice communication quality assessment model based on the first sample embedded feature vector, the second sample embedded feature vector, the third sample embedded feature vector and the total target loss function to obtain the feature extraction model, the downsampling model and the voice communication quality assessment model.

13. A server comprising a processor and a memory; the memory storing at least one program code; the at least one program code being configured to be called and executed by the processor to implement the method for evaluating voice communication quality according to any one of claims 1 to 6.

14. A computer-readable storage medium, wherein at least one computer program is stored in the computer-readable storage medium, and when the at least one computer program is executed by a processor, the method for evaluating voice communication quality according to any one of claims 1 to 6 can be implemented. 15 . A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the method for evaluating voice communication quality according to claim 1 can be implemented.

Citation Information

Patent Citations

  • Speech quality evaluation method and device

    CN106531190A

  • Non-reference voice quality objective assessment method based on deep learning voice enhancement

    CN107358966A

  • Sound quality detection model training method, sound quality detection method, electronic equipment and medium

    CN114694678A

  • Voice communication quality evaluation method and device, server and storage medium

    CN118038897A

  • Robust intrusive perceptual audio quality assessment based on convolutional neural networks

    WO2022112594A2

Cited By

  • Corpus automatic collection and quality control method and system based on intelligent algorithm

    CN122347962A