Voice quality inspection method, device and equipment and storage medium
By performing interpolation on the model parameters of the single-task speech quality inspection model, and then performing interpolation on the target model parameters, a variety of speech quality inspection results are accurately achieved. This greatly improves the flexibility and accuracy of speech quality inspection, reduces device occupancy and CPU usage, and enhances the user experience.
Patent Information
- Application Number
- CN202411517388.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-28
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2044-10-28
AI Technical Summary
Traditional voice quality inspection methods rely on manual listening and evaluation, which is time-consuming and costly. Furthermore, existing voice quality inspection systems are optimized for single tasks and are difficult to adapt to changing quality inspection needs and business scenarios, resulting in low efficiency and high consumption of computing resources.
By acquiring a single-task speech quality inspection model trained based on a preset neural network model, determining the target model parameters, generating a target speech quality inspection model for processing multiple speech quality inspection tasks, and inputting the speech to be inspected into the model for quality inspection, multiple speech quality inspection results are obtained.
It improves the flexibility and accuracy of voice quality inspection, reduces equipment occupancy, enhances user experience, and enables voice quality inspection results for various voice quality inspection tasks.
Smart Images

Figure CN119418719B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, and particularly relates to a voice quality inspection method and device, equipment and a storage medium. BACKGROUND
[0002] Voice quality inspection is a key link for evaluating and improving the quality of customer service, which involves analyzing customer service dialogue recordings to identify problems for improvement. For example, voice quality inspection includes identifying user emotions, detecting dialogue topics, and detecting dialogue speech rates, etc. Traditional voice quality inspection relies on manual listening evaluation, but this method is time-consuming and costly. Therefore, voice quality inspection technologies based on speech recognition and natural language processing techniques have emerged, including but not limited to emotion detection models, dialogue topic detection models, and dialogue speech rate detection models, and voice quality inspection is performed based on each model. However, these systems are usually optimized for a single task and are difficult to adapt to changing quality inspection requirements and changing business scenarios, and the efficiency of voice quality inspection through each single task is low, and a large amount of computing power and running memory of the detection device is consumed.
[0003] Therefore, how to improve the efficiency and accuracy of voice quality inspection is a problem to be solved at present. SUMMARY
[0004] The main purpose of the present application is to provide a voice quality inspection method, device, equipment and storage medium, which aims to improve the efficiency and accuracy of voice quality inspection.
[0005] In a first aspect, the present application provides a voice quality inspection method, which comprises the following steps:
[0006] Obtaining a single-task voice quality inspection model for at least two different voice quality inspection tasks based on a preset neural network model training;
[0007] According to the model parameters of each single-task voice quality inspection model and the original model parameters of the preset neural network model, determining target model parameters;
[0008] Using the target model parameters as the model parameters of the preset neural network model to obtain a target voice quality inspection model, the target voice quality inspection model is used for processing multiple voice quality inspection tasks;
[0009] Obtaining a target voice to be inspected, and inputting the target voice into the target voice quality inspection model for voice quality inspection to obtain voice quality inspection results of multiple voice quality inspection tasks.
[0010] In a second aspect, the present application further provides a voice quality inspection device, which comprises an acquisition module and a generation module, wherein:
[0011] The acquisition module is configured to acquire single-task voice quality inspection models of at least two different voice quality inspection tasks trained based on a preset neural network model.
[0012] The generation module is configured to determine target model parameters according to model parameters of each single-task voice quality inspection model and original model parameters of the preset neural network model.
[0013] The generation module is further configured to obtain a target voice quality inspection model by taking the target model parameters as model parameters of the preset neural network model, and the target voice quality inspection model is used to process multiple voice quality inspection tasks.
[0014] The acquisition module is further configured to acquire target voice to be inspected.
[0015] The generation module is further configured to input the target voice to the target voice quality inspection model for voice quality inspection to obtain voice quality inspection results of multiple voice quality inspection tasks.
[0016] In a third aspect, the present application further provides a computer device, which comprises a processor, a memory, and a computer program stored in the memory and executable by the processor, wherein the computer program, when executed by the processor, implements the steps of the voice quality inspection method described above.
[0017] In a fourth aspect, the present application further provides a computer readable storage medium, which stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the voice quality inspection method described above.
[0018] The present application provides a voice quality inspection method, device, equipment and storage medium. The present application acquires single-task voice quality inspection models of at least two different voice quality inspection tasks trained based on a preset neural network model, determines target model parameters according to model parameters of each single-task voice quality inspection model and original model parameters of the preset neural network model, obtains a target voice quality inspection model by taking the target model parameters as model parameters of the preset neural network model, acquires target voice to be inspected, and inputs the target voice to the target voice quality inspection model for voice quality inspection to obtain voice quality inspection results of multiple voice quality inspection tasks. In the present application, target model parameters of a model that can be used to process multiple voice quality inspection tasks are generated by model parameters of single-task voice quality inspection models and original model parameters, and the target voice quality inspection model used to process multiple voice quality inspection tasks can be accurately obtained based on the target model parameters. Then, voice quality inspection is performed based on the target voice quality inspection model to obtain voice quality inspection results of multiple voice quality inspection tasks, which greatly improves the flexibility and accuracy of voice quality inspection. BRIEF DESCRIPTION OF DRAWINGS
[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed to be used in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative effort on the basis of these drawings.
[0020] Figure 1 A flowchart of a voice quality inspection method provided by an embodiment of the present application is shown in FIG. 1.
[0021] Figure 2 A flowchart of another voice quality inspection method provided by an embodiment of the present application is shown in FIG. 2.
[0022] Figure 3 A flowchart of a sub-step of the voice quality inspection method in FIG. 1 is shown in FIG. 3. Figure 1
[0023] A flowchart of a sub-step of the voice quality inspection method in FIG. 1 is shown in FIG. 3. Figure 4 A schematic block diagram of a voice quality inspection device provided by an embodiment of the present application is shown in FIG. 4.
[0024] Figure 5 A schematic block diagram of a sub-module of the voice quality inspection device in FIG. 4 is shown in FIG. 5. Figure 4
[0025] Figure 6 A schematic block diagram of a computer device provided by an embodiment of the present application is shown in FIG. 6.
[0026] The implementation, functional features and advantages of the present application will be further described with reference to the embodiments and the drawings. DETAILED DESCRIPTION
[0027] The technical solutions in the embodiments of the present application will be described clearly and completely with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative effort belong to the scope of protection of the present application.
[0028] The flowcharts shown in the drawings are only exemplary, and do not necessarily include all the contents and operations / steps, nor do they have to be executed in the described order. For example, some operations / steps can be decomposed, combined or partially merged, so the actual execution order can be changed according to the actual situation.
[0029] Embodiments of the present application can acquire and process related data based on artificial intelligence technology. Among them, artificial intelligence (AI) is the theory, method, technology and application system of using digital computer or machine controlled by digital computer to simulate, extend and expand human intelligence, perceive environment, acquire knowledge and use knowledge to obtain optimal results.
[0030] The basic technology of artificial intelligence generally includes technologies such as sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction system, mechatronics, etc. The software technology of artificial intelligence mainly includes computer vision technology, robot technology, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning, etc.
[0031] Voice quality inspection is a key link for evaluating and improving the quality of customer service, which involves analyzing customer service dialogue recordings to identify problems for improvement. For example, voice quality inspection includes identifying user emotions, detecting dialogue topics, and detecting dialogue speech rates, etc. Traditional voice quality inspection relies on manual listening evaluation, but this method is time-consuming and costly. Therefore, voice quality inspection technologies using speech recognition and natural language processing techniques have emerged, including but not limited to emotion detection models, dialogue topic detection models, and dialogue speech rate detection models, and voice quality inspection is performed based on each model. However, these systems are usually optimized for a single task and are difficult to adapt to changing quality inspection requirements and changing business scenarios, and the efficiency of voice quality inspection through each single task is low, and a large amount of computing power and running memory of the detection device is consumed.
[0032] To solve the above problems, embodiments of the present application provide a voice quality inspection method, device, equipment and storage medium. The voice quality inspection method includes: obtaining at least two single-task voice quality inspection models for different voice quality inspection tasks trained based on a preset neural network model; determining target model parameters according to model parameters of each single-task voice quality inspection model and original model parameters of the preset neural network model; using the target model parameters as model parameters of the preset neural network model to obtain a target voice quality inspection model, the target voice quality inspection model being used for processing multiple voice quality inspection tasks; obtaining a target voice to be inspected and inputting the target voice into the target voice quality inspection model for voice quality inspection to obtain voice quality inspection results of multiple voice quality inspection tasks.
[0033] The voice quality inspection method can be applied in a computer device, which can be a mobile phone, a tablet computer, a notebook computer, a desktop computer, a personal digital assistant, a wearable device, or other electronic devices.
[0034] Some embodiments of the present application will be described in detail with reference to the drawings. The following examples and features in the examples can be combined with each other without conflict.
[0035] Please refer to Figure 1 , Figure 1 A flowchart of a voice quality inspection method provided by an embodiment of the present application is shown.
[0036] As Figure 1 shown, the voice quality inspection method includes steps S101 to S104.
[0037] Step S101, obtaining single-task voice quality inspection models for at least two different voice quality inspection tasks trained based on a preset neural network model.
[0038] The preset neural network model can be selected according to actual conditions, and embodiments of the present application do not make specific limitations thereon. For example, the preset neural network model includes but is not limited to a convolutional neural network model, a recurrent neural network model, and a recurrent convolutional neural network model. For example, the preset neural network model is ChatGLM3.
[0039] It should be noted that the voice quality inspection task includes but is not limited to emotion recognition, intent detection, topic detection, and speech rate detection, etc. The emotion recognition includes positive emotion, negative emotion, and neutral emotion.
[0040] In some embodiments, as Figure 2 shown, steps S201 to S205 are further included before step S101.
[0041] Step S201, obtaining a sample data set, the sample data set including a plurality of sample data, each sample data including sample voice and identified single-task voice quality inspection results.
[0042] The sample data set includes a plurality of sample data, each sample data including sample voice and identified single-task voice quality inspection results. For example, the voice quality inspection task is to detect emotion, and the identified single-task voice quality inspection results are emotion results corresponding to the sample voice,
[0043] In some embodiments, sample voice is obtained, and the single-task voice quality inspection results of the sample voice are labeled to obtain identified single-task voice quality inspection results. The sample voice and the identified single-task voice quality inspection results are integrated to obtain a sample data. The steps of repeatedly obtaining sample voice, labeling the single-task voice quality inspection results of the sample voice to obtain identified single-task voice quality inspection results, and integrating the sample voice and the identified single-task voice quality inspection results to obtain a sample data are repeated to obtain a sample data set.
[0044] Step S202, selecting one sample data from the sample data set as a target sample data.
[0045] The manner of selecting the sample data can be selected according to actual conditions, and embodiments of the present application do not make specific limitations thereon. For example, one sample data is randomly selected as the target sample data; or the sample data is selected as the target sample data according to the time sequence of generation of the sample data.
[0046] Step S203, inputting the sample voice in the target sample data into a preset neural network model for recognition to obtain a predicted single-task voice quality inspection result.
[0047] The sample voice in the target sample data is inputted into the preset neural network model for recognition to obtain the predicted single-task voice quality inspection result. The sample voice is recognized by the preset neural network model, and the predicted single-task voice quality inspection result can be accurately obtained.
[0048] Step S204, determining whether the preset neural network model converges according to the predicted single-task voice quality inspection result and the identified single-task voice quality inspection result.
[0049] In some embodiments, a loss value of the preset neural network model is determined according to the predicted single-task voice quality inspection result and the identified single-task voice quality inspection result. If the loss value of the preset neural network model is less than or equal to a preset loss value, it is determined that the preset neural network model has converged. If the loss value of the preset neural network model is greater than the preset loss value, it is determined that the preset neural network model has not converged. The preset loss value can be set according to actual conditions, and embodiments of the present application do not make specific limitations thereon. For example, the preset loss value can be 0.002. The loss value of the preset neural network model can accurately know the convergence of the preset neural network model, greatly improving the efficiency and accuracy of training the preset neural network model.
[0050] In some embodiments, the manner of determining the loss value of the preset neural network model according to the predicted single-task voice quality inspection result and the identified single-task voice quality inspection result can be: calculating the cosine similarity of the predicted single-task voice quality inspection result and the identified single-task voice quality inspection result; and determining the value of one minus the cosine similarity as the loss value of the preset neural network model. It should be noted that the loss value of the neural network model can also be calculated in other manners, and the present application does not make specific limitations thereon.
[0051] Step S205, if the preset neural network model does not converge, adjusting the model parameters of the preset neural network model, and cyclically executing the step of selecting one sample data from the sample data set as the target sample data until the preset neural network model converges.
[0052] If the loss value of the preset neural network model is greater than the preset loss value, it is determined that the preset neural network model does not converge, and the model parameters of the preset neural network model are adjusted. One sample data is selected from the sample data set as the target sample data. The sample voice in the target sample data is input into the preset neural network model for recognition to obtain the predicted single-task voice quality inspection result. According to the predicted single-task voice quality inspection result and the identified single-task voice quality inspection result, it is determined whether the preset neural network model converges. If the preset neural network model does not converge, the model parameters of the preset neural network model are adjusted and the previous steps are continued to be cyclically executed until the preset neural network model converges to obtain the single-task voice quality inspection model.
[0053] Step S102, determining the target model parameters according to the model parameters of each single-task voice quality inspection model and the original model parameters of the preset neural network model.
[0054] The model parameters are model parameter coefficients of the single-task voice quality inspection model, and the model parameter coefficients include a plurality of parameter coefficients of each layer of the model. For example, the single-task voice quality inspection model includes an encoder layer, a fully connected layer and a decoder layer, wherein the encoder layer includes 20 parameter coefficients, the fully connected layer includes 50 parameter coefficients, and the decoder layer includes 20 parameter coefficients.
[0055] In some embodiments, the model parameter coefficients of each single-task voice quality inspection model are obtained to obtain a plurality of model parameter coefficients. The model parameters of the preset neural network model are obtained to obtain original model parameters. The target model parameters are determined according to the model parameters of each single-task voice quality inspection model and the original model parameters of the preset neural network model. Through the model parameters of the single-task voice quality inspection model and the original model parameters of the preset neural network model, the model parameters of the model for processing multiple voice quality inspection tasks can be efficiently and accurately determined.
[0056] In some embodiments, as shown in FIG. 10, step S102 includes sub-step S1021 to sub-step S1022. Figure 3
[0057] Sub-step S1021, determining the model parameter difference values of each model parameter and the original model parameter to obtain a plurality of model parameter difference values.
[0058] The model parameters and the original model parameters each include a plurality of parameter coefficients.
[0059] In some embodiments, the parameter coefficients in each model parameter corresponding to the same parameter coefficient positions in the original model parameter are sequentially subtracted to obtain a model parameter difference of each model parameter and the original model parameter, and the model parameter difference includes a sub-parameter difference of each parameter coefficient position. By respectively subtracting the parameter coefficients in each model parameter corresponding to the same parameter coefficient positions in the original model parameter, the plurality of model parameter differences can be accurately obtained.
[0060] For example, the single-task voice quality inspection model includes an emotion recognition model and an intent detection model; the model parameters of the emotion recognition model are obtained to obtain first model parameters; the model parameters of the intent detection model are obtained to obtain second model parameters; the model parameters of the preset neural network model are obtained to obtain original model parameters; wherein the first model parameters include parameter coefficient 1, parameter coefficient 2, parameter coefficient 3, parameter coefficient 4, and parameter coefficient 5; the second model parameters include parameter coefficient 6, parameter coefficient 7, parameter coefficient 8, parameter coefficient 9, and parameter coefficient 10; and the original model parameters include parameter coefficient 11, parameter coefficient 12, parameter coefficient 13, parameter coefficient 14, and parameter coefficient 15.
[0061] The parameter coefficient 1 is subtracted from the parameter coefficient 11 to obtain a sub-parameter difference 1; the parameter coefficient 2 is subtracted from the parameter coefficient 11 to obtain a sub-parameter difference 2; the parameter coefficient 3 is subtracted from the parameter coefficient 11 to obtain a sub-parameter difference 3; the parameter coefficient 4 is subtracted from the parameter coefficient 11 to obtain a sub-parameter difference 4; the parameter coefficient 5 is subtracted from the parameter coefficient 11 to obtain a sub-parameter difference 5; and the sub-parameter difference 1, the sub-parameter difference 2, the sub-parameter difference 3, the sub-parameter difference 4, and the sub-parameter difference 5 are integrated to obtain a first model parameter difference.
[0062] The parameter coefficient 6 is subtracted from the parameter coefficient 11 to obtain a sub-parameter difference 6; the parameter coefficient 7 is subtracted from the parameter coefficient 11 to obtain a sub-parameter difference 7; the parameter coefficient 8 is subtracted from the parameter coefficient 11 to obtain a sub-parameter difference 8; the parameter coefficient 9 is subtracted from the parameter coefficient 11 to obtain a sub-parameter difference 9; the parameter coefficient 10 is subtracted from the parameter coefficient 11 to obtain a sub-parameter difference 10; and the sub-parameter difference 6, the sub-parameter difference 7, the sub-parameter difference 8, the sub-parameter difference 9, and the sub-parameter difference 10 are integrated to obtain a second model parameter difference.
[0063] The sub-step S1022 synthesizes the model parameters of the multi-task voice quality inspection model according to the model parameter differences and the original model parameters to obtain target model parameters.
[0064] In some embodiments, a plurality of sub-parameter differences in each model parameter difference is screened, and the un-screened sub-parameter differences are zeroed to obtain a plurality of candidate model parameter differences; each sub-parameter difference in each candidate model parameter difference is amplified to obtain a plurality of target model parameter differences; and the target model parameters are synthesized according to the plurality of target model parameter differences and the original model parameters to obtain the target model parameters. By synthesizing the target model parameter differences and the original model parameters, the target model parameters can be accurately obtained.
[0065] In some embodiments, the screening manner of the plurality of sub-parameter differences in each model parameter difference can be: performing absolute value processing on the sub-parameter differences at each parameter coefficient position in the model parameter difference to obtain a sub-difference coefficient absolute value corresponding to each sub-parameter difference; sorting the sub-difference coefficient absolute values according to the absolute value size to obtain a sub-difference coefficient sequence; selecting a preset proportion of sub-difference coefficients from the sub-difference coefficient sequence at a preset interval to obtain a plurality of target sub-difference coefficients, the preset proportion being a proportion of the number of target sub-parameter differences to the total number of sub-parameter differences; and obtaining the sub-parameter difference corresponding to each target sub-difference coefficient and determining the obtained sub-parameter difference as the screened target sub-parameter difference. The preset interval and the preset proportion can be set according to actual conditions, and the embodiments of the present application do not make specific limitations thereon, for example, the preset interval is 1 and the preset proportion is 0.2.
[0066] For example, the first model parameter difference includes sub-parameter difference 1, sub-parameter difference 2, sub-parameter difference 3, sub-parameter difference 4, sub-parameter difference 5, sub-parameter difference 6, sub-parameter difference 7, sub-parameter difference 8, sub-parameter difference 9, and sub-parameter difference 10; performing absolute value processing on each sub-parameter difference, sorting the sub-difference coefficient absolute values according to the absolute value size to obtain a sub-difference coefficient sequence of sub-parameter difference 1, sub-parameter difference 5, sub-parameter difference 7, sub-parameter difference 3, sub-parameter difference 2, sub-parameter difference 6, sub-parameter difference 10, sub-parameter difference 4, sub-parameter difference 8, and sub-parameter difference 9; the preset interval is 2 and the preset proportion is 0.3; then the target sub-parameter differences include sub-parameter difference 1, sub-parameter difference 3, and sub-parameter difference 10; or the target sub-parameter differences include sub-parameter difference 5, sub-parameter difference 2, and sub-parameter difference 6; or the target sub-parameter differences include sub-parameter difference 7, sub-parameter difference 6, and sub-parameter difference 8; or the target sub-parameter differences include sub-parameter difference 3, sub-parameter difference 10, and sub-parameter difference 9.
[0067] In some embodiments, the manner of amplifying each sub-parameter difference in each candidate model parameter difference to obtain a plurality of target model parameter differences can be: dividing each sub-parameter difference by the preset ratio to obtain a target model parameter difference corresponding to each sub-parameter difference.
[0068] In some embodiments, the manner of synthesizing model parameters of the multi-task voice quality detection model according to the plurality of target model parameter differences and the original model parameters to obtain target model parameters can be: performing mean processing on sub-parameter differences at the same parameter coefficient position in each target model parameter difference to obtain a total model parameter difference; and performing addition processing on sub-parameter differences at the same parameter coefficient position in the total model parameter difference and the original model parameters to obtain the target model parameters. The target model parameters are accurately obtained by performing addition processing on sub-parameter differences at the same parameter coefficient position in the total model parameter difference and the original model parameters.
[0069] Step S103: taking the target model parameters as model parameters of the preset neural network model to obtain a target voice quality detection model, the target voice quality detection model being used for processing a plurality of voice quality detection tasks.
[0070] In some embodiments, the target model parameters are taken as model parameters of the preset neural network model to obtain a target voice quality detection model, the target voice quality detection model being used for processing a plurality of voice quality detection tasks. By taking the target model parameters as model parameters of the preset neural network model, the target voice quality detection model capable of processing a plurality of voice quality detection tasks can be accurately obtained. By synthesizing the target voice quality detection model capable of processing a plurality of voice quality detection tasks, a plurality of voice quality detection tasks can be completed, the number of models can be reduced, the occupancy rate of the device can be reduced, the use rate of the CPU can be reduced, and the fluency of the device can be improved.
[0071] It should be noted that the single-task voice quality detection model includes an emotion recognition model, an intent detection model, a topic detection model, and a speech rate detection model; and the synthesized target voice quality detection model can perform emotion recognition, intent detection, topic detection, and speech rate detection.
[0072] Step S104: obtaining a target voice to be detected, and inputting the target voice to the target voice quality detection model to perform voice quality detection, thereby obtaining voice quality detection results of a plurality of voice quality detection tasks.
[0073] The target voice to be detected is obtained, and the target voice is input to the target voice quality detection model to perform voice quality detection, thereby obtaining voice quality detection results of a plurality of voice quality detection tasks. The voice quality detection flexibility, efficiency, and accuracy are greatly improved, thereby improving the user experience.
[0074] The speech quality inspection method provided in the above embodiments obtains at least two single-task speech quality inspection models for different speech quality inspection tasks by training a preset neural network model; determines target model parameters based on the model parameters of each single-task speech quality inspection model and the original model parameters of the preset neural network model; uses the target model parameters as the model parameters of the preset neural network model to obtain a target speech quality inspection model, which is used to handle multiple speech quality inspection tasks; obtains the target speech to be inspected and inputs it into the target speech quality inspection model for speech quality inspection to obtain speech quality inspection results for multiple speech quality inspection tasks. In this application, target model parameters that can be used to handle multiple speech quality inspection tasks are generated by using the model parameters of the single-task speech quality inspection model and the original model parameters. Based on the target model parameters, the target speech quality inspection model for handling multiple speech quality inspection tasks can be accurately obtained. Then, speech quality inspection is performed based on the target speech quality inspection model to obtain speech quality inspection results for multiple speech quality inspection tasks, which greatly improves the flexibility and accuracy of speech quality inspection.
[0075] Please see Figure 4 , Figure 4 This is a schematic block diagram of a voice quality inspection device provided in an embodiment of this application.
[0076] like Figure 4 As shown, the voice quality inspection device 300 includes an acquisition module 310 and a generation module 320, wherein:
[0077] The acquisition module 310 is used to acquire a single-task speech quality inspection model that has been trained based on a preset neural network model to obtain at least two different speech quality inspection tasks.
[0078] The generation module 320 is used to determine the target model parameters based on the model parameters of each single-task speech quality inspection model and the original model parameters of the preset neural network model.
[0079] The generation module 320 is further configured to use the target model parameters as model parameters of the preset neural network model to obtain a target speech quality inspection model, which is used to process various speech quality inspection tasks.
[0080] The acquisition module 310 is also used to acquire the target speech to be inspected;
[0081] The generation module 320 is also used to input the target speech into the target speech quality inspection model for speech quality inspection, and obtain speech quality inspection results for various speech quality inspection tasks.
[0082] In some embodiments, such as Figure 5As shown, the generation module 320 includes a first generation submodule 321 and a second generation submodule 322, wherein:
[0083] The first generation submodule 321 is configured to determine model parameter differences between each model parameter and the original model parameter, to obtain a plurality of model parameter differences.
[0084] The second generation submodule 322 is configured to synthesize model parameters of the multi-task voice quality inspection model according to each model parameter difference and the original model parameter, to obtain target model parameters.
[0085] In some embodiments, the first generation submodule 321 is further configured to:
[0086] The parameter coefficients at the same parameter coefficient positions in each model parameter and the original model parameter are sequentially subjected to difference operation, to obtain model parameter differences between each model parameter and the original model parameter, and the model parameter differences include sub-parameter differences at each parameter coefficient position.
[0087] In some embodiments, the second generation submodule 322 is further configured to:
[0088] The plurality of sub-parameter differences in each model parameter difference are screened, and the un-screened sub-parameter differences are subjected to zero processing, to obtain a plurality of candidate model parameter differences.
[0089] Each sub-parameter difference in each candidate model parameter difference is subjected to amplification processing, to obtain a plurality of target model parameter differences.
[0090] The model parameters of the multi-task voice quality inspection model are synthesized according to the plurality of target model parameter differences and the original model parameter, to obtain target model parameters.
[0091] In some embodiments, the second generation submodule 322 is further configured to:
[0092] The sub-parameter differences at each parameter coefficient position in the model parameter difference are subjected to absolute value processing, to obtain sub-difference coefficient absolute values corresponding to each sub-parameter difference.
[0093] The sub-difference coefficient absolute values are sorted according to absolute value sizes, to obtain a sub-difference coefficient sequence.
[0094] A preset proportion of sub-difference coefficients are selected from the sub-difference coefficient sequence at a preset interval, to obtain a plurality of target sub-difference coefficients, and the preset proportion is a proportion of the number of target sub-parameter differences to the total number of sub-parameter differences.
[0095] Obtain a sub-parameter difference corresponding to each target sub-difference coefficient, and determine the obtained sub-parameter difference as a target sub-parameter difference selected.
[0096] In some embodiments, the second generation sub-module 322 is further configured to:
[0097] Perform mean processing on the sub-parameter differences of the same parameter coefficient positions in each target model parameter difference to obtain a total model parameter difference.
[0098] Perform addition processing on the sub-parameter differences of the same parameter coefficient positions in the total model parameter difference and the original model parameter to obtain a target model parameter.
[0099] In some embodiments, the speech quality inspection device 300 is further configured to:
[0100] Obtain a sample data set, the sample data set including a plurality of sample data, the sample data including sample speech and an identified single-task speech quality inspection result;
[0101] Select one sample data from the sample data set as a target sample data;
[0102] Input the sample speech in the target sample data into a preset neural network model for recognition to obtain a predicted single-task speech quality inspection result;
[0103] Determine whether the preset neural network model converges according to the predicted single-task speech quality inspection result and the identified single-task speech quality inspection result;
[0104] If the preset neural network model does not converge, adjust the model parameters of the preset neural network model, and cyclically perform the step of selecting one sample data from the sample data set as a target sample data until the preset neural network model converges.
[0105] It should be noted that, for the convenience and brevity of description, the specific working process of the above speech quality inspection device can refer to the corresponding process in the foregoing speech quality inspection method embodiments, which will not be described herein.
[0106] Please refer to Figure 6 , Figure 6 A structural schematic block diagram of a computer device provided in an embodiment of the present application.
[0107] As Figure 6 shown, the computer device 400 includes a processor 402 and a memory 403 connected through a system bus 401, wherein the memory 403 can include a storage medium and an internal memory.
[0108] The storage medium can store a computer program. The computer program comprises program instructions which, when executed, cause the processor to perform any one of the voice quality inspection methods.
[0109] The processor 402 is configured to provide computing and control capabilities to support the operation of the entire computer device.
[0110] The internal memory provides an environment for the running of the computer program in the storage medium, which, when executed by the processor, causes the processor to perform any one of the voice quality inspection methods.
[0111] Those skilled in the art can understand that, Figure 6 The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0112] It should be understood that the processor 402 can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.
[0113] In one embodiment, the processor 402 is configured to run a computer program stored in the memory to perform the following steps:
[0114] Obtain single-task voice quality inspection models for at least two different voice quality inspection tasks trained based on a preset neural network model;
[0115] According to the model parameters of each single-task voice quality inspection model and the original model parameters of the preset neural network model, determine target model parameters;
[0116] Use the target model parameters as the model parameters of the preset neural network model to obtain a target voice quality inspection model, and the target voice quality inspection model is used to process multiple voice quality inspection tasks;
[0117] Obtain a target voice to be quality inspected, and input the target voice into the target voice quality inspection model to perform voice quality inspection, and obtain voice quality inspection results of multiple voice quality inspection tasks.
[0118] In one embodiment, the processor 402, when implementing the determination of the target model parameter according to the model parameter of each single-task voice quality inspection model and the original model parameter of the preset neural network model, is configured to:
[0119] Determine the model parameter difference between each model parameter and the original model parameter to obtain multiple model parameter differences.
[0120] Synthesize the model parameter of the multi-task voice quality inspection model according to each model parameter difference and the original model parameter to obtain the target model parameter.
[0121] In one embodiment, each model parameter and the original model parameter both include multiple parameter coefficients, and the processor 402, when implementing the determination of the model parameter difference between each model parameter and the original model parameter to obtain multiple model parameter differences, is configured to:
[0122] Perform difference operation on the parameter coefficients at the same parameter coefficient position in each model parameter and the original model parameter in sequence to obtain the model parameter difference between each model parameter and the original model parameter, and the model parameter difference includes a sub-parameter difference at each parameter coefficient position.
[0123] In one embodiment, the processor 402, when implementing the synthesis of the model parameter of the multi-task voice quality inspection model according to each model parameter difference and the original model parameter to obtain the target model parameter, is configured to:
[0124] Filter multiple sub-parameter differences in each model parameter difference, and perform zero processing on the unfiltered sub-parameter differences to obtain multiple candidate model parameter differences.
[0125] Amplify each sub-parameter difference in each candidate model parameter difference to obtain multiple target model parameter differences.
[0126] Synthesize the model parameter of the multi-task voice quality inspection model according to the multiple target model parameter differences and the original model parameter to obtain the target model parameter.
[0127] In one embodiment, the processor 402, when implementing the filtering of multiple sub-parameter differences in each model parameter difference, is configured to:
[0128] perform absolute value processing on a sub-parameter difference value at each parameter coefficient position in the model parameter difference value to obtain a sub-difference coefficient absolute value corresponding to each sub-parameter difference value;
[0129] sort the sub-difference coefficient absolute values according to absolute value sizes to obtain a sub-difference coefficient sequence;
[0130] select a preset proportion of sub-difference coefficients from the sub-difference coefficient sequence at a preset interval to obtain a plurality of target sub-difference coefficients, the preset proportion being a proportion of a number of target sub-parameter difference values to a total number of sub-parameter difference values;
[0131] obtain a sub-parameter difference value corresponding to each target sub-difference coefficient, and determine the obtained sub-parameter difference value as a screened target sub-parameter difference value.
[0132] In one embodiment, the processor 402, when implementing the synthesis of the model parameters of the multi-task voice quality detection model according to the plurality of target model parameter difference values and the original model parameter, is configured to implement:
[0133] perform mean value processing on sub-parameter difference values at the same parameter coefficient position in each target model parameter difference value to obtain a total model parameter difference value;
[0134] perform addition processing on sub-parameter difference values at the same parameter coefficient position in the total model parameter difference value and the original model parameter to obtain a target model parameter.
[0135] In one embodiment, the processor 402 is further configured to implement:
[0136] obtain a sample data set, the sample data set including a plurality of sample data, the sample data including sample voice and an identified single-task voice quality detection result;
[0137] select one sample data from the sample data set as a target sample data;
[0138] input the sample voice in the target sample data to a preset neural network model for recognition to obtain a predicted single-task voice quality detection result;
[0139] determine whether the preset neural network model converges according to the predicted single-task voice quality detection result and the identified single-task voice quality detection result;
[0140] if the preset neural network model does not converge, adjust model parameters of the preset neural network model, and cyclically execute the step of selecting one sample data from the sample data set as a target sample data until the preset neural network model converges.
[0141] It should be noted that the skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the computer device is described above, and the corresponding process in the foregoing voice quality inspection method embodiments can be referred to, which will not be described here.
[0142] The computer readable storage medium in the embodiments of the present application stores a computer program, the computer program includes program instructions, and the method implemented by the program instructions can refer to each embodiment of the voice quality inspection method of the present application.
[0143] The computer readable storage medium can be an internal storage unit of the computer device, for example, a hard disk or a memory of the computer device. The computer readable storage medium can be non-volatile or volatile. The computer readable storage medium can also be an external storage device of the computer device, for example, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc.
[0144] Further, the computer readable storage medium can mainly include a storage program area and a storage data area, wherein the storage program area can store an operating system, at least one application required by a function, etc., and the storage data area can store data created according to the use of the blockchain node, etc.
[0145] The blockchain referred to in the present application is a new application mode of distributed data storage, peer-to-peer transmission, consensus mechanism, encryption algorithm and other computer technologies. The blockchain is essentially a decentralized database, which is a series of data blocks associated using cryptographic methods, each data block contains information of a batch of network transactions, and is used to verify the validity (anti-fake) of the information and generate the next block. The blockchain can include a blockchain underlying platform, a platform product service layer and an application service layer.
[0146] It should be understood that the terms used in the present application specification are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the present application specification, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include the plural forms.
[0147] It should also be understood that, in the specification, the terms "and / or" is used to indicate one or more of the items it conjoins, including the case where one or more of the items is present alone, with the conjoint term describing the items as an "or" list. It should also be understood that, in this document, the term "comprising" is used in the transition sense, encompassing the case where one or more of the items in the list of items following the term are present, along with the case where one or more of those items are not present. It should also be understood that, in this document, the term "comprising" is used in the transition sense, encompassing the case where one or more of the items in the list of items following the term are present, along with the case where one or more of those items are not present. It should also be understood that, in this document, the term "comprising" is used in the transition sense, encompassing the case where one or more of the items in the list of items following the term are present, along with the case where one or more of those items are not present.
[0148] The above-mentioned sequence number of the embodiments of the present application is only for description, and does not represent the advantages and disadvantages of the embodiments. The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of various equivalent modifications or replacements within the technical range disclosed by the present application, and these modifications or replacements should be covered within the protection scope of the present application.
Claims
1. A voice quality inspection method, characterized by, The method comprises: obtaining single-task voice quality inspection models for at least two different voice quality inspection tasks based on a preset neural network model; determining target model parameters according to model parameters of each single-task voice quality inspection model and original model parameters of the preset neural network model; obtaining a target voice quality inspection model by taking the target model parameters as model parameters of the preset neural network model, the target voice quality inspection model being used for processing multiple voice quality inspection tasks; obtaining a target voice to be inspected and inputting the target voice into the target voice quality inspection model for voice quality inspection to obtain voice quality inspection results of multiple voice quality inspection tasks; the determining of the target model parameters according to the model parameters of each single-task voice quality inspection model and the original model parameters of the preset neural network model comprises: each of the model parameters and the original model parameters comprises multiple parameter coefficients, and parameter coefficients at the same parameter coefficient positions in each of the model parameters and the original model parameters are subjected to difference value operation in sequence to obtain model parameter differences of each of the model parameters and the original model parameters, the model parameter differences comprising sub-parameter differences at each parameter coefficient position; each of the model parameter differences is subjected to screening, and sub-parameter differences that are not screened are subjected to zero processing to obtain multiple candidate model parameter differences; each of the sub-parameter differences in each of the candidate model parameter differences is subjected to amplification processing to obtain multiple target model parameter differences; the model parameters of a multi-task voice quality inspection model are synthesized according to the multiple target model parameter differences and the original model parameters to obtain target model parameters.
2. The voice quality monitoring method of claim 1, wherein, the screening of each of the model parameter differences comprises: each of the sub-parameter differences at the same parameter coefficient position in the model parameter difference is subjected to absolute value processing to obtain a sub-difference coefficient absolute value corresponding to each of the sub-parameter differences; sub-difference coefficient sequences are obtained by sorting the sub-difference coefficient absolute values according to absolute value sizes; multiple target sub-difference coefficient absolute values are selected from the sub-difference coefficient sequences at a preset interval and at a preset proportion, the preset proportion being a proportion of a number of target sub-parameter differences to a total number of sub-parameter differences; each target sub-difference coefficient absolute value is obtained, and the obtained target sub-difference coefficient absolute value is determined as a screened target sub-parameter difference.
3. The voice quality monitoring method of claim 1, wherein, the synthesizing of the model parameters of the multi-task voice quality inspection model according to the multiple target model parameter differences and the original model parameters to obtain the target model parameters comprises: sub-parameter differences at the same parameter coefficient positions in each of the target model parameter differences are subjected to mean value processing to obtain a total model parameter difference; the total model parameter difference and sub-parameter differences at the same parameter coefficient positions in the original model parameters are subjected to addition processing to obtain target model parameters.
4. The voice quality monitoring method of any one of claims 1-3, wherein, The method further comprises: obtaining a sample data set, the sample data set comprising multiple sample data, the sample data comprising sample voice and labeled single-task voice quality inspection results; select one sample data from the sample data set as a target sample data; input sample voice in the target sample data into a preset neural network model for recognition to obtain a predicted single-task voice quality inspection result; determine whether the preset neural network model converges according to the predicted single-task voice quality inspection result and the identified single-task voice quality inspection result; if the preset neural network model does not converge, adjust model parameters of the preset neural network model, and cyclically execute the step of selecting one sample data from the sample data set as a target sample data until the preset neural network model converges.
5. A voice quality monitoring device, characterized by, The voice quality inspection device comprises an acquisition module and a generation module, wherein: The acquisition module is configured to acquire single-task voice quality inspection models of at least two different voice quality inspection tasks trained based on a preset neural network model; The generation module is configured to determine target model parameters according to model parameters of each single-task voice quality inspection model and original model parameters of the preset neural network model; The generation module is further configured to use the target model parameters as model parameters of the preset neural network model to obtain a target voice quality inspection model, which is used to process multiple voice quality inspection tasks; The acquisition module is further configured to acquire a target voice to be inspected; The generation module is further configured to input the target voice into the target voice quality inspection model for voice quality inspection to obtain voice quality inspection results of multiple voice quality inspection tasks; The determination of the target model parameters according to the model parameters of each single-task voice quality inspection model and the original model parameters of the preset neural network model comprises: Each of the model parameters and the original model parameters comprises multiple parameter coefficients, the parameter coefficients at the same parameter coefficient positions in each of the model parameters and the original model parameters are subjected to difference value operation in sequence to obtain model parameter differences of each of the model parameters and the original model parameters, and the model parameter differences comprise sub-parameter differences at each parameter coefficient position; Each of the model parameter differences is subjected to screening, and the sub-parameter differences that are not screened are subjected to zero processing to obtain multiple candidate model parameter differences; Each of the sub-parameter differences in each of the candidate model parameter differences is subjected to amplification processing to obtain multiple target model parameter differences; The target model parameters are obtained by synthesizing model parameters of a multi-task voice quality inspection model according to the multiple target model parameter differences and the original model parameters.
6. A computer device, comprising: The computer device comprises a processor, a memory, and a computer program stored on the memory and executable by the processor, wherein the computer program is executed by the processor to implement the steps of the voice quality inspection method according to any one of claims 1 to 4.
7. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, wherein the computer program is executed by the processor to implement the steps of the voice quality inspection method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Voice signal processing model training method, voice signal processing model training device, electronic equipment and storage medium
CN109841220A
Voice quality inspection method and device, quality inspection equipment and readable storage medium
CN111696528A