Voiceprint-based object classification method and device, equipment and storage medium

Voiceprint processing is performed through voice channel detection and pre-trained object classification model, which solves the problem that device differences affect the accuracy of voiceprint matching and achieves higher object classification accuracy.

CN120236566APending Publication Date: 2025-07-01SF TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311873599.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-30
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

In the process of voiceprint extraction, the requirements for recording equipment are relatively strict, and the voiceprints of the same speaking object recorded by different equipment are different, which affects the accuracy of voiceprint matching and the accuracy of object classification.

Method used

By obtaining the target voice data of multiple target objects, performing voice channel detection to obtain voice channel labels, using a pre-trained object classification model for vocal segmentation, voice clustering and object classification, and filtering reference voiceprint data from candidate voiceprint library according to the voice channel label for object classification.

Benefits of technology

It effectively avoids the impact of equipment differences on voiceprint recognition, improves the accuracy of voiceprint matching, and thus improves the accuracy of object classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236566A_ABST
    Figure CN120236566A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a voiceprint-based object classification method and device, equipment and a storage medium, and belongs to the technical field of artificial intelligence. The method comprises the following steps: acquiring target voice data; performing voice channel detection on the target voice to obtain a voice channel label; performing human voice segmentation processing on the target voice according to the human voice segmentation sub-model to obtain human voice segments; wherein the human voice segments are voice segments which do not contain silence; voice clustering processing is carried out on the human voice segments according to the voice clustering sub-models, voice clustering data are obtained, and the voice clustering data comprise object voice segments of the target objects; screening out reference voiceprint data from a preset candidate voiceprint library according to the voice channel label; and carrying out object classification on the object voice segment and the reference voiceprint data according to the object classification sub-model. According to the embodiment of the invention, the accuracy of voiceprint matching can be improved, so that the accuracy of object classification is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular, to a method and device, equipment, and storage medium for object classification based on voiceprint. Background Art

[0002] Speech recognition refers to the process of analyzing speech signals to recognize and understand the content of the words spoken by the speaker. The speech features contained in the speech signal that can characterize the personality information of the speaker. Since the physiological organs used by each person when speaking are different in size and shape, etc., and combined with the differences in factors such as age, personality, and language habits, the voiceprint of each speaker is unique. Voiceprint recognition refers to the process of performing identity recognition and verification by analyzing the voice characteristics of an individual. The object classification method based on voiceprint refers to the process of using voiceprint recognition technology to classify different speaker objects. For example, in a customer service scenario, after the customer service and the user have a conversation using a communication device, a call voice will be generated. By performing object classification on the voiceprint in the call voice, the voice emitted by the customer service can be determined to perform quality inspection on the voice content of the customer service and adjust the customer service voice according to the quality inspection results.

[0003] However, due to the existing technology in the process of voiceprint extraction, the requirements for the recording device are relatively strict, and the voiceprints of the same speaker recorded by different devices will also be different. In this way, the voiceprints of a speaker on different devices are inconsistent, and this inconsistency will affect the accuracy of voiceprint matching, thereby reducing the accuracy of object classification in the voice. Therefore, how to improve the accuracy of object classification in the voice has become an urgent technical problem to be solved. Summary of the Invention

[0004] The main purpose of the embodiments of this application is to propose a method and device, equipment, and storage medium for object classification based on voiceprint, which can improve the accuracy of voiceprint matching, thereby improving the accuracy of object classification.

[0005] To achieve the above object, the first aspect of the embodiments of this application proposes a method for object classification based on voiceprint, and the method includes:

[0006] Obtain target voice data, where the target voice data includes target voices from multiple target objects;

[0007] Perform voice channel detection on the target voice to obtain a voice channel label;

[0008] Obtain a pre-trained object classification model, where the object classification model includes a human voice segmentation sub-model, a voice clustering sub-model, and an object classification sub-model;

[0009] Perform voice segmentation processing on the target voice according to the voice segmentation sub-model to obtain voice segments; wherein, the voice segments are voice segments that do not contain silence;

[0010] Perform voice clustering processing on the voice segments according to the voice clustering sub-model to obtain voice clustering data, and the voice clustering data includes the object voice segments of each target object;

[0011] Screen out reference voiceprint data from a preset candidate voiceprint library according to the voice channel label;

[0012] Perform object classification on the object voice segments and the reference voiceprint data according to the object classification sub-model.

[0013] In some embodiments, performing object classification on the object voice segments and the reference voiceprint data according to the object classification sub-model includes:

[0014] Extract the voiceprint of the object voice segment to obtain object voiceprint data;

[0015] Calculate the voiceprint similarity according to the object voiceprint data and the reference voiceprint data to obtain voiceprint similarity data;

[0016] Determine the object classification label of the object voice segment according to the voiceprint similarity data, and the object classification label indicates the category of the object of the object voice segment.

[0017] In some embodiments, the object classification model further includes a decoder; after performing object classification on the object voice segments and the reference voiceprint data according to the object classification sub-model, the method further includes:

[0018] If the object classification label is a preset target label, obtain the object voice segment with the object classification label to obtain a matching voice segment;

[0019] Perform voice decoding processing on the matching voice segment according to the decoder to obtain a matching object voice text;

[0020] Perform keyword detection on the matching object voice text according to a preset keyword library to obtain detection data.

[0021] In some embodiments, performing voice segmentation processing on the target voice according to the voice segmentation sub-model to obtain voice segments includes;

[0022] Perform voice division on the target voice according to the preset first Mel frequency cepstrum window data to obtain first window voices;

[0023] Extract window features based on the first window voice to obtain first window voice features;

[0024] Perform window voice classification based on the first window voice features to obtain initial window labels;

[0025] Extract window label pairs according to the initial window labels, where the window label pairs include a first window label and a second window label, and the first window label and the second window label are labels of adjacent first window voices;

[0026] If the first window label is a human voice label and the second window label is the human voice label, merge the first window voice of the first window label and the first window voice of the second window label to obtain the human voice speech segment, and the human voice label is used to represent that the first window voice is a human voice label.

[0027] In some embodiments, the performing voice clustering processing on the human voice speech segment according to the voice clustering sub-model to obtain voice clustering data includes:

[0028] Perform voice division on the human voice speech segment according to preset second Mel frequency cepstrum window data to obtain second window voices;

[0029] Extract features from the second window voices to obtain second window voice features;

[0030] Perform feature merging on the second window voice features according to preset segment window data; the segment window length of the segment window data is the sum of the window lengths of a preset number of second Mel frequency cepstrum window data;

[0031] Calculate the segment window mean according to the segment window voice features to obtain the segment window mean;

[0032] Calculate the segment window variance according to the segment window voice features to obtain the segment window variance;

[0033] Perform numerical splicing on the segment window mean and the segment window variance to obtain target segment window data;

[0034] Perform segment clustering according to the target segment window data to obtain the voice clustering data.

[0035] In some embodiments, the performing segment clustering according to the target segment window data to obtain voice clustering data includes:

[0036] Extract window features from the target segment window data to obtain target segment window features;

[0037] Perform spectral clustering based on the target segment window features to obtain segment window clustering labels and segment window segmentation data;

[0038] Perform voice segmentation on the vocal voice segment according to the window segmentation data to obtain the target voice segment;

[0039] Determine the voice clustering data according to the segment window clustering labels and the target voice segment.

[0040] In some embodiments, before filtering the reference voiceprint data from a preset candidate voiceprint library according to the voice channel label, the method further includes: constructing the candidate voiceprint library, specifically including:

[0041] Obtain sample voice data, where the sample voice data includes a sample voice, a sample voice channel label of the sample voice, and a sample object of the sample voice; wherein, the sample voice refers to a voice containing multiple objects;

[0042] Perform vocal segmentation processing on the sample voice according to the vocal segmentation sub-model to obtain sample vocal voice segments;

[0043] Perform voice clustering processing on the sample vocal voice segments according to the voice clustering sub-model to obtain sample object voice segments; wherein, the sample object voice segment refers to a voice segment containing one object;

[0044] Perform voice decoding processing on the sample object voice segment according to the decoder to obtain a sample object voice text;

[0045] Perform text classification on the sample object voice text to obtain sample object classification labels;

[0046] If the sample object classification label is the target label, obtain the sample object voice segment with the sample object classification label to obtain a sample matching voice segment;

[0047] Perform voiceprint extraction on the sample matching voice segment to obtain candidate voiceprint data;

[0048] Construct the candidate voiceprint library according to the candidate voiceprint data, the sample voice channel label, and the sample object.

[0049] To achieve the above object, a second aspect of the embodiments of the present application proposes an object classification device based on voiceprint, and the device includes:

[0050] A data acquisition module, configured to acquire target voice data, where the target voice data includes target voices from multiple target objects;

[0051] A detection module, configured to perform voice channel detection on the target voice to obtain a voice channel label;

[0052] A model acquisition module, configured to acquire a pre-trained object classification model, where the object classification model includes a voice segmentation sub-model, a voice clustering sub-model, and an object classification sub-model;

[0053] A voice segmentation module, configured to perform voice segmentation processing on the target voice according to the voice segmentation sub-model to obtain voice segments of human voices; wherein, the voice segments of human voices are voice segments that do not include silence;

[0054] A voice clustering module, configured to perform voice clustering processing on the voice segments of human voices according to the voice clustering sub-model to obtain voice clustering data, where the voice clustering data includes object voice segments of each target object;

[0055] A screening module, configured to screen out reference voiceprint data from a preset candidate voiceprint library according to the voice channel label;

[0056] A classification module, configured to perform object classification on the object voice segments and the reference voiceprint data according to the object classification sub-model.

[0057] To achieve the above object, a third aspect of the embodiments of the present application proposes a computer device, where the computer device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the method described in any one of the embodiments of the first aspect above is implemented.

[0058] To achieve the above object, a fourth aspect of the embodiments of the present application further proposes a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method described in any one of the embodiments of the first aspect above is implemented.

[0059] The object classification method based on voiceprint proposed in the embodiments of this application determines the voice channel label of the target voice, and filters out the reference voiceprint data from the preset candidate voiceprint library according to the voice channel label, so as to classify the object voice segment in the target voice. Since in the prior art, during the voiceprint extraction process, the requirements for the recording device are relatively strict, and there will also be differences in the voiceprints of the same speaker recorded by different devices. For example, the voices emitted by the microphones of different phones may be distorted to varying degrees. In this way, the voiceprints of a speaker on different devices are inconsistent, and this inconsistency will affect the accuracy of voiceprint matching, thereby reducing the accuracy of object classification in speech. The embodiments of this disclosure obtain the target voice including multiple target objects. By performing voice channel detection on the target voice, the voice channel label of the target voice is obtained. In this way, object classification based on the voice channel can avoid the influence of device differences on voiceprint recognition. Obtain a pre-trained object classification model, which includes a human voice segmentation sub-model, a voice clustering sub-model, and an object classification sub-model. Among them, the human voice segmentation sub-model performs human voice segmentation processing on the target voice to obtain human voice segments. The human voice segment is a voice segment that does not contain silence. The voice clustering sub-model performs voice clustering processing on the human voice segments to obtain voice clustering data. At this time, the voice clustering data includes the object voice segments of each target object. Filter out the reference voiceprint data from the preset candidate voiceprint library according to the voice channel label, and perform object classification on the object voice segments and the reference voiceprint data according to the object classification sub-model. It can be seen that this application performs object classification based on the reference voiceprint data filtered out from the preset candidate voiceprint library according to the voice channel label, which can avoid the influence of device differences on voiceprint recognition, more accurately identify the voiceprint characteristics of the object, improve the accuracy of voiceprint matching, and thus improve the accuracy of object classification. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] Figure 1 is the first flowchart of the object classification method based on voiceprint provided by the embodiments of this application;

[0061] Figure 2 is the structural schematic diagram of the object classification model provided by the embodiments of this application;

[0062] Figure 3 is Figure 1 the flowchart of the specific method of step S140 in

[0063] Figure 4 is the structural schematic diagram of the classification network constructed based on the stacking of TDNN layers and LSTM layers provided by the embodiments of this application;

[0064] Figure 5 is the structural schematic diagram of the TDNN layer provided by the embodiments of this application;

[0065] Figure 6 is Figure 1 a flowchart of the specific method of step S150 in

[0066] Figure 7 a schematic structural diagram of the voice clustering sub - model provided by an embodiment of the present application;

[0067] Figure 8 is Figure 6 a flowchart of the specific method of step S670 in

[0068] Figure 9 a flowchart of constructing a candidate voiceprint library provided by an embodiment of the present application;

[0069] Figure 10 is Figure 1 a flowchart of the specific method of step S170 in

[0070] Figure 11 a second flowchart of the voiceprint - based object classification method provided by an embodiment of the present application;

[0071] Figure 12 a block diagram of the voiceprint - based object classification device provided by an embodiment of the present application;

[0072] Figure 13 a schematic hardware structure diagram of the computer device provided by an embodiment of the present application. Detailed implementation manners

[0073] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0074] It should be noted that although the functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from the module division in the device or the order in the flowchart. Terms such as "first", "second", etc. in the specification, claims and the above - mentioned drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.

[0075] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0076] First, several nouns involved in the present application are analyzed:

[0077] Artificial Intelligence (AI): It is a new technical science that studies and develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence. Artificial intelligence is a branch of computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. The research in this field includes robots, speech recognition, image recognition, natural language processing, and expert systems, etc. Artificial intelligence can simulate the information process of human consciousness and thinking. It also refers to the theory, method, technology, and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.

[0078] Voice channel: It refers to the medium used by people for voice communication. The voice channel usually refers to telephones, Internet voice, walkie-talkies, etc. In voice communication, sound is converted into an electrical signal and transmitted to the receiving end through transmission media such as cables and radio waves. The receiving end then converts the electrical signal back into a sound signal. The quality of the voice channel directly affects the effect and accuracy of voice communication.

[0079] Single-channel voice: It means that the voice channel only contains voice signals from a single sound source. In voice communication, single-channel voice usually refers to the situation where there is only one microphone or only one voice input signal. Single-channel voice and dual-channel voice are two different types of signals used in voice processing and communication.

[0080] Dual-channel voice: It means that the voice channel contains voice signals from two different sound sources. In voice communication, dual-channel voice usually refers to the situation where two microphones or two voice input signals are used.

[0081] Mel Frequency Cepstrum Coefficient (MFCC): It is a technology used for voice signal processing and voice feature extraction. MFCC is a cepstrum coefficient based on the Mel frequency scale and is commonly used in fields such as speech recognition, speaker recognition, and speech emotion recognition. The Mel frequency is proposed based on the auditory characteristics of the human ear and has a non-linear correspondence with the Hertz (Hz) frequency. MFCC is the Hz spectrum feature calculated using this relationship between the Mel frequency and the Hertz frequency.

[0082] Time Delay Neural Network (TDNN): It is a neural network structure composed of a series of convolutional layers. The output of each convolutional layer is processed by a non-linear activation function. Different from the traditional Convolutional Neural Network (CNN), the convolutional operation in TDNN is performed in the time dimension to capture the temporal patterns in the input data. The convolutional kernels in TDNN are shifted in time, thus taking into account the dependencies between different time steps. TDNN layers are often used in speech processing and natural language processing tasks, such as speech recognition, speech synthesis, and language models.

[0083] Long Short-Term Memory layer: It is a variant of the Recurrent Neural Network (RNN), specifically designed for processing sequential data. Compared with the traditional RNN, LSTM introduces a gating mechanism to solve the problem of long-term dependencies. The LSTM layer consists of a series of units, and each unit maintains an internal memory state, and the flow and forgetting of information can be controlled by different gating units. The LSTM layer can effectively handle long-term dependencies, so it performs well in tasks that require considering context information and long-term dependencies, such as language modeling, machine translation, and sentiment analysis.

[0084] Speech recognition refers to the process of analyzing speech signals to identify and understand the content of what the speaker says. Speech features that can characterize the personality information of the speaker are contained in the speech signals. Since the physiological organs used by each person when speaking are different in terms of size and shape, etc., and combined with differences in factors such as age, personality, and language habits, the voiceprint of each speaker is unique. Voiceprint recognition refers to the process of performing identity recognition and verification by analyzing the voice characteristics of an individual. The method of object classification based on voiceprint refers to the process of using voiceprint recognition technology to classify different speaker objects. For example, in a customer service scenario, after the customer service and the user have a conversation using a communication device, a call voice will be generated. By performing object classification on the voiceprint in the call voice, the voice sent by the customer service can be determined to conduct quality inspection on the voice content of the customer service and adjust the customer service voice according to the quality inspection results.

[0085] However, due to the relatively strict requirements for recording devices in the existing technology during the voiceprint extraction process, and there are also differences in the voiceprints of the same speaker recorded by different devices. In this way, the voiceprints of a speaker on different devices are inconsistent, and this inconsistency will affect the accuracy of voiceprint matching, thus reducing the accuracy of object classification in speech. Therefore, how to improve the accuracy of object classification in speech has become an urgent technical problem to be solved.

[0086] Based on this, the embodiments of the present application provide a voiceprint-based object classification method, a voiceprint-based object classification device, a computer device, and a storage medium, which can improve the accuracy of voiceprint matching, thereby improving the accuracy of object classification.

[0087] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.

[0088] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0089] The voiceprint-based object classification method provided by the embodiments of the present application relates to the field of artificial intelligence technology. The voiceprint-based object classification method provided by the embodiments of the present application can be applied to terminals, can also be applied to server sides, or can also be software running on terminals or server sides. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, or a smart watch, etc.; the server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms; the software can be an application that implements the voiceprint-based object classification method, etc., but is not limited to the above forms.

[0090] This application can be used in numerous general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0091] Please refer to Figure 1 , Figure 1 which is an optional flowchart of the voiceprint-based object classification method provided by an embodiment of this application. In some embodiments of this application, the voiceprint-based object classification method specifically includes but is not limited to steps S110 to S170. The following will introduce these seven steps in detail in conjunction with Figure 1 .

[0092] Step S110: Obtain target voice data, where the target voice data includes target voices from multiple target objects;

[0093] Step S120: Perform voice channel detection on the target voice to obtain a voice channel label;

[0094] Step S130: Obtain a pre-trained object classification model, where the object classification model includes a voice segmentation sub-model, a voice clustering sub-model, and an object classification sub-model;

[0095] Step S140: Perform voice segmentation processing on the target voice according to the voice segmentation sub-model to obtain voice segments of human voices; among them, the voice segments of human voices are voice segments that do not contain silence;

[0096] Step S150: Perform voice clustering processing on the voice segments of human voices according to the voice clustering sub-model to obtain voice clustering data, where the voice clustering data includes object voice segments of each target object;

[0097] Step S160: Screen out reference voiceprint data from a preset candidate voiceprint library according to the voice channel label;

[0098] Step S170: Perform object classification on the object voice segments and the reference voiceprint data according to the object classification sub-model.

[0099] In steps S110 to S170 of some embodiments, target voice data including voices from multiple target objects is obtained. Voice channel detection is performed on the target voice to obtain a voice channel label for the target voice. In this way, object classification based on the voice channel can avoid the influence of device differences on voiceprint recognition. An object classification model that has been pre-trained is obtained. The object classification model includes a voice segmentation sub-model, a voice clustering sub-model, and an object classification sub-model. Among them, voice segmentation processing is performed on the target voice according to the voice segmentation sub-model to obtain a voice segment of human voice. The voice segment of human voice is a voice segment that does not include silence. Voice clustering processing is performed on the voice segment of human voice according to the voice clustering sub-model to obtain voice clustering data. At this time, the voice clustering data includes the object voice segments of each target object. Reference voiceprint data is screened from a preset candidate voiceprint library according to the voice channel label, and object classification is performed on the object voice segments and the reference voiceprint data according to the object classification sub-model. It can be seen from this that compared with related technologies, in the process of voiceprint extraction, the requirements for recording devices are relatively strict, and there are also differences in the voiceprints of the same speaker recorded by different devices. In this application, object classification is performed using the reference voiceprint data screened from the preset candidate voiceprint library according to the voice channel label, which can avoid the influence of device differences on voiceprint recognition, more accurately identify the voiceprint characteristics of the object, improve the accuracy of voiceprint matching, and thus improve the accuracy of object classification.

[0100] In step S110 of some embodiments, the target voice data refers to data related to the target voice for which object classification is to be performed. The target voice data includes the target voice. The target voice at this time refers to the voice for which object classification is to be performed, and the target voice includes voices emitted by multiple target objects. For example, in the scenario of express delivery after-sales, the target voice can be the call voice between the express delivery customer service and the user. At this time, the target voice includes the voices of two objects, namely the express delivery customer service and the user.

[0101] It should be noted that for single-channel speech, the sound signal comes from only one sound source. Therefore, when digitized and stored, the speech storage format data has only the speech data of one channel. When storing single-channel speech, the commonly used formats include Waveform Audio File Format (WAV), Moving Picture Experts Group-1 Audio Layer 3 (MP3), Advanced Audio Coding (AAC), etc. For dual-channel speech, the sound signals come from two different sound sources. Therefore, when digitized and stored, the speech storage format data contains data of two channels, corresponding to the left channel and the right channel respectively. When storing dual-channel speech, the commonly used formats also include WAV, MP3, AAC, etc. The difference is that the audio data in these formats contains the sound information of two channels.

[0102] It should be noted that the voiceprint-based object classification method of the present application can be used to assist in voice object classification or voice quality inspection in intelligent customer service, intelligent vehicles, smart homes, etc. For example, in the intelligent customer service scenario of a financial bank, when the intelligent customer service communicates with the target object by voice, the call voice can be continuously recorded and updated. By accurately classifying different objects in the recorded voice, voice detection can be performed based on the voice of the customer service, and the service attitude can be adjusted in a timely manner. It is also possible to detect the voice emotion of the target object (i.e., the user), and according to the emotion change of the target object, prompt the intelligent customer service to call the speech script, that is, call the soothing speech script to comfort the target object, effectively improving the user experience.

[0103] In step S120 of some embodiments, the voice channel detection refers to the process of detecting the voice storage channel information of the target voice to determine whether the target voice is single-channel speech or dual-channel speech at this time. The voice channel label refers to the label used to characterize the channel type to which the target voice belongs. The voice channel label includes a single-channel label and a dual-channel label. The single-channel label indicates that the detected voice is single-channel speech. The dual-channel voice indicates that the detected voice is dual-channel speech.

[0104] It should be noted that the target voice data also includes voice storage format data. The methods for voice channel detection include: performing voice channel detection on the target voice according to the voice storage format data to obtain a voice channel label. Among them, the voice storage format data refers to the format in which the target voice is stored in the corresponding channel of the device when digitizing the voice signal. The voice storage format data includes single-channel storage format and dual-channel storage format. That is to say, according to the voice storage format data, it can be determined whether the target voice is a single-channel voice or a dual-channel voice. Due to the difference in voice storage format, the voiceprints of the same person in single-channel and dual-channel calls will be different. Therefore, object classification based on the voice channel can avoid the influence of device differences on voiceprint recognition.

[0105] It should be noted that in actual voice storage and processing, the selection of appropriate storage format data depends on specific application requirements. For situations that require retaining stereo effects or dual-channel recording, the storage format of dual-channel voice should be selected; while for mono recording or situations that only require mono sound information, the storage format of single-channel voice can be selected.

[0106] It should be noted that the methods for performing voice channel detection on the target voice also include: channel detection (which refers to the method of detecting the channel information in the sound signal of the target voice by analyzing the waveform data of the sound signal of the target voice. Among them, the sound signal of a single channel has only one channel, while the sound signal of a dual channel includes left and right channels. By analyzing the channel information of the sound signal, it can be determined whether the sound is from a single channel or a dual channel), energy detection (which refers to the method of performing channel detection by calculating the energy distribution of the sound signal of the target voice on different channels. Among them, the energy distribution of the sound signal of a dual channel on the left and right channels is usually different, while the energy distribution of the sound signal of a single channel on the two channels will be relatively similar), phase difference detection (which refers to the method of performing channel detection using the phase difference information of the sound signal on different channels. Among them, the phase difference of the sound signal of a dual channel on the left and right channels is usually different, while the phase difference of the sound signal of a single channel on the two channels will be relatively close), spectral feature detection (which refers to the method of performing channel detection by analyzing the features of the sound signal in the frequency domain, such as spectral shape, spectral flatness, etc. Among them, the features of the sound signal of a dual channel in the frequency domain are usually different, while the features of the sound signal of a single channel in the frequency domain will be relatively similar), etc. The specific method of voice channel detection in this application is not limited. In actual applications, multiple methods can also be combined for comprehensive judgment to improve the accuracy and robustness of detection, which will not be elaborated here.

[0107] In step S130 of some embodiments, the object classification model refers to a model used for different objects included in the target voice. Such asFigure 2 As shown in the figure, the object classification model 210 constructed in the present application includes a voice segmentation sub-model, a voice clustering sub-model, a decoder, an object classification sub-model, and a voiceprint extraction sub-model.

[0108] In step S140 of some embodiments, as Figure 2 shown, the voice segmentation sub-model refers to a model used to perform binary classification segmentation on the target voice to divide the silent segment and the human voice segment in the voice. Therefore, according to the voice segmentation sub-model, the target voice is segmented to obtain at least one human voice segment. For example, the target voice is "Hello, I'm customer service agent A. Nice to serve you... Hello, may I ask how to handle service B?", where "..." represents the silent part where no one is speaking in the voice at this time. Therefore, when the target voice is segmented according to the voice segmentation sub-model, the target voice can be divided into multiple voice segments, including voice segment 1 "Hello, I'm customer service agent A. Nice to serve you", voice segment 2 "...", and voice segment 3 "Hello, may I ask how to handle service B". By detecting the type of each voice segment, it is determined whether the voice segment is a silent segment or a human voice segment. Then, the silent segments in the target voice are removed, and only the human voice segments are retained.

[0109] It can be understood that by dividing the silent segments and the human voice segments in the target voice, the object classification model can pay more attention to the human voice segments, thereby improving the recognition accuracy of the voice object. Because the silent segments usually do not contain useful information, excluding them can reduce interference, thereby reducing the length of the voice to be processed and reducing the computational amount, which helps to improve the performance of the object classification model.

[0110] Please refer to Figure 3 , Figure 3 which is the specific flowchart of step S140 provided by the embodiments of the present application. In some embodiments of the present application, step S140 may specifically include but is not limited to steps S310 to S350. The following will introduce these five steps in detail with reference to Figure 3 .

[0111] Step S310: Divide the target voice according to the preset first Mel frequency cepstrum window data to obtain the first window voice;

[0112] Step S320: Extract window features according to the first window voice to obtain the first window voice features;

[0113] Step S330: Classify the window voice according to the first window voice features to obtain the initial window label;

[0114] Step S340: Extract window label pairs according to the initial window labels. The window label pairs include a first window label and a second window label, and the first window label and the second window label are labels of adjacent first window voices.

[0115] Step S350: If the first window label is a human voice label and the second window label is a human voice label, merge the first window voice of the first window label and the first window voice of the second window label to obtain a human voice speech segment. The human voice label is used to indicate that the first window voice is a human voice label.

[0116] In step S310 of some embodiments, the first Mel-frequency cepstrum window data refers to the window parameters used to extract MFCC segments from the audio signal of the target speech. By specifying the window size and step size to extract MFCC segments, it can help to perform more refined feature extraction and analysis on the speech signal. The first window voice refers to an MFCC speech segment of a window extracted from the target speech according to the first Mel-frequency cepstrum window data. Multiple first window voices can be segmented from a target speech.

[0117] It should be noted that the window parameters refer to the parameters of the time window used for frame processing of the speech signal. In a given target speech, a time window of a fixed length will be used to intercept the speech, and analysis will be performed on each window. For example, the window parameter windows = 0.025 seconds (s), which means that the length of each window is 0.025 seconds. This means that in the speech signal, a time window with a length of 0.025 seconds will be used to intercept the speech signal for frame processing.

[0118] It should be noted that the step parameter represents the overlapping part between adjacent windows. For example, the step parameter step = 0.01s means that the step between adjacent windows is 0.01 seconds. In frame processing, there will be a certain overlap between two adjacent windows, and the length of this overlap is the step. In this example, the overlapping part between adjacent windows is 0.01 seconds.

[0119] For example, for a 6 - second target speech, the preset first Mel - frequency cepstrum window data includes window parameters and step parameters. Among them, the window parameter is windows = 0.025 seconds (s), and the step parameter is step = 0.01s. When calculating the number of first - window speeches according to this first Mel - frequency cepstrum window data, it is first necessary to calculate the number of sampling points included in each window. Assuming the sampling rate is 16000Hz, the number of sampling points included in a 0.025 - second window is: 0.025s×16000Hz = 400 sampling points. Similarly, the number of sampling points included in a 0.01 - second step is: 0.01s×16000Hz = 160 sampling points. Therefore, for a 10 - second speech signal, the total number of windows that can be calculated is = (length of speech signal - window length) / step + 1, that is, the number of windows = (10s×16000Hz - 400 sampling points) / 160 sampling points + 1 = 6241 windows. Therefore, through the first Mel - frequency cepstrum window data with windows = 0.025s and step = 0.01s, 6241 first - window speeches can be extracted. The segment length of each first - window speech is 0.025 seconds, and there is an overlap of 0.01 seconds between adjacent segments.

[0120] In step S320 of some embodiments, after determining multiple first - window speeches, window feature extraction is performed on the first - window speeches to obtain first - window speech features (i.e., MFCC features). For example, after performing window feature extraction on the first - window speeches in the target speech, the feature dimension of the window speech features of the target speech is L×M, where L represents the number of first - window speeches (related to the length of the target speech), and M represents the dimension of MFCC itself (which can be 23, 40, 80, etc., and is not specifically limited here). At this time, the dimension of a first - window speech feature is 1×M.

[0121] In step S330 of some embodiments, the initial window label is a binary - classification label, and the initial window label includes a silence label and a human - voice label. The silence label indicates that this first - window speech is a silence segment, and the human - voice label indicates that this first - window speech is a human - voice speech segment.

[0122] It should be noted that the network used for window speech classification of the first - window speech features is a classification network constructed based on the stacking of TDNN layers and LSTM layers. As Figure 4As shown, the classification network 410 constructed based on the stacking of TDNN layers and LSTM layers includes a first TDNN layer, a second TDNN layer, a third TDNN layer, a first LSTM layer, a fourth TDNN layer, a second LSTM layer, a fifth TDNN layer, and an output layer. Therefore, by inputting the window speech features of the target speech of the first window speech features into the classification network constructed based on the stacking of TDNN layers and LSTM layers, the corresponding classification result Y can be output. Y represents a set of initial labels, with a dimension of Lx1, and each dimension corresponds to the initial window label of the first window speech.

[0123] It should be noted that the structure of the TDNN layer adopted in this application can be as Figure 5 shown. This TDNN layer is similar to dilated convolution, so that a large range of outputs from the previous layer can be received with very little computation. In the TDNN, t, t+2, t+5, etc. represent time delays. TDNN is a neural network structure for processing time-series data, which can perform time-delay processing on the input time-series data to capture the time-series information in the data. t represents the input at the current moment, t+2 represents the input 2 time units after the current moment, t+5 represents the input 5 time units after the current moment, and so on. These time-delayed inputs can help the network learn the long-term dependencies in the time-series data, so as to better understand and predict the characteristics of the time-series data.

[0124] In step S340 of some embodiments, the window label pair refers to a label pair composed of the initial window labels corresponding to adjacent first window speeches. For example, the target speech includes 5 first window speeches, namely window speech W1, window speech W2, window speech W3, window speech W4, and window speech W5. After window speech classification, the initial window label T1 corresponding to window speech W1, the initial window label T2 corresponding to window speech W2, the initial window label T3 corresponding to window speech W3, the initial window label T4 corresponding to window speech W4, and the initial window label T5 corresponding to window speech W5 are obtained. At this time, the constructed window label pairs include (initial window label T1, initial window label T2), (initial window label T2, initial window label T3), (initial window label T3, initial window label T4), (initial window label T4, initial window label T5).

[0125] In step S350 of some embodiments, since the duration of the window voice is short, in the subsequent object classification process, complete information cannot be extracted for the judgment of the object. Therefore, it is necessary to exclude the voice segments corresponding to the interfering silence tags to retain the voice segments corresponding to the valid human voice tags. Therefore, if the first window tag is a human voice tag and the second window tag is a human voice tag, the first window voice of the first window tag and the first window voice of the second window tag are merged to obtain a human voice segment.

[0126] It should be noted that if the first window tag is a human voice tag but the second window tag is a silence tag, a segmentation point is determined between these two window tags, and the category of the next tag adjacent to the second window tag is judged. If the next tag adjacent to the second window tag is a human voice tag, the voice segment corresponding to the second window tag is excluded. If the next tag adjacent to the second window tag is a silence tag, then the next tag adjacent to the next tag adjacent to the second window tag is judged until a human voice tag or a voice end marker is detected, and the voice corresponding to the silence tag is excluded. Among them, for the voice segments corresponding to non-adjacent human voice tags, voice merging is not required.

[0127] In some embodiments, in step S150, as Figure 2 shown, the voice clustering sub-model refers to a model used to perform object voice clustering on the target voice to cluster the voice segments of each speaker together. It should be noted that although clustering is performed at this time, it is not clear who the specific object corresponding to each clustering is. In this application, speaker segmentation is performed through the voice clustering sub-model, which can improve the accuracy of voiceprint extraction in practical applications.

[0128] It should be noted that if the voice channel tag indicates that the target voice is a dual-channel voice, it does not need to pass through the voice clustering sub-model. Because in dual-channel voice, the voices of different objects such as the customer service and the customer are in different channels and do not interfere with each other. Therefore, the voices collected from different channels are equivalent to having been clustered, so the voice clustering sub-model 212 only performs clustering segmentation on single-channel voice.

[0129] Please refer to Figure 6 , Figure 6 which is the specific flowchart of step S150 provided by the embodiments of this application. In some embodiments of this application, step S150 may specifically include but is not limited to steps S610 to S670. The following will introduce these seven steps in detail in combination with Figure 6 this.

[0130] Step S610, perform voice division on the human voice segment according to the preset second Mel frequency cepstrum window data to obtain the second window voice;

[0131] Step S620: Extract features from the second-window speech to obtain second-window speech features;

[0132] Step S630: Merge the features of the second-window speech features according to the preset segment-window data to obtain segment-window speech features; the segment-window length of the segment-window data is the sum of the window lengths of a preset number of second Mel-frequency cepstrum window data.

[0133] Step S640: Calculate the segment-window mean according to the segment-window speech features to obtain the segment-window mean;

[0134] Step S650: Calculate the segment-window variance according to the segment-window speech features to obtain the segment-window variance;

[0135] Step S660: Concatenate the numerical values of the segment-window mean and the segment-window variance to obtain the target segment-window data;

[0136] Step S670: Perform segment clustering according to the target segment-window data to obtain speech clustering data.

[0137] In step S610 of some embodiments, the second Mel-frequency cepstrum window data refers to the window parameters used in the speech clustering sub-model to extract MFCC segments from the audio signal of the target speech. The second Mel-frequency cepstrum window data and the above-mentioned first Mel-frequency cepstrum window data have the same parameters and functions, except for the differences in the sub-models applied. The second-window speech refers to an MFCC speech segment of one window extracted from each human voice speech segment according to the second Mel-frequency cepstrum window data. For example, the window parameter in the second Mel-frequency cepstrum window data is windows = 1.5 seconds, and the step parameter is step = 0.75 seconds.

[0138] It should be noted that, as Figure 7 shown, the speech clustering sub-model 710 includes a first TDNN layer, a second TDNN layer, a third TDNN layer, a fourth TDNN layer, a fifth TDNN layer, a time segment encoding (stats) layer, a sixth TDNN layer, and a clustering output layer. The TDNN layers in this application adopt the same structure and will not be elaborated here.

[0139] In step S620 of some embodiments, the present application may first use the first TDNN layer, the second TDNN layer, the third TDNN layer, the fourth TDNN layer, and the fifth TDNN layer to process in sequence to extract features from the second-window speech to obtain second-window speech features. That is to say, these five TDNN layers analyze the MFCC features at the frame level.

[0140] In step S630 of some embodiments, since the MFCC features obtained based on the second Mel-frequency cepstrum window data are features extracted from the speech in a relatively small window, the features presented by the sound in such a short window may be unstable. To accurately divide the speech segments of different objects, the second-window speech features are merged according to the preset segment window data to obtain the segment-window speech features. The preset segment window data is used to represent the data after merging multiple windows. The segment window length of the segment window data is the sum of the window lengths of a preset number of second Mel-frequency cepstrum window data. For example, if the preset number corresponding to the window segment "chunk" in the preset segment window data is 100, it means that the features within 100 windows are merged. If the length of each window is 0.025 seconds, then the segment window length is 0.025×100 = 2.5 seconds. The preset segment window data can be flexibly set according to actual needs, and the set preset segment window data can distinguish the speech segments of different objects. Herein, "chunk" generally refers to dividing the speech signal into shorter time periods, and each time period is usually called a "chunk" or a "frame".

[0141] In steps S640 and S650 of some embodiments, in the time segment encoding layer, feature extraction at the segment level is performed on the second-window speech features within a chunk, that is, by calculating the mean value and variance corresponding to each second-window speech feature in the previous-layer segment-window speech features, the segment-window mean value and the segment-window variance are obtained. At this time, both the segment-window mean value and the segment-window variance are respectively a preset number. Then, the segment-window mean value and the segment-window variance are numerically concatenated to obtain the target segment window data, which is input into the sixth TDNN layer.

[0142] In steps S660 and S670 of some embodiments, the sixth TDNN layer is equivalent to a feature extraction layer, and its output serves as the feature representation of each speaker, and then clustering is performed on each feature representation. For example, if the speaking objects in the target speech include two, then the class value k in the clustering output layer is 2. Segment clustering is performed according to the target segment window data, and the obtained speech clustering data includes two types of sounds, that is, corresponding to two objects.

[0143] Therefore, through the speech clustering sub-model of the present application, the speech segments belonging to different objects in the human voice speech segments can be clustered together to achieve the division of speech segments of different speaking objects and improve the accuracy of subsequent object classification.

[0144] Please refer to Figure 8 , Figure 8 which is the specific flowchart of step S670 provided by the embodiments of the present application. In some embodiments of the present application, step S670 may specifically include but is not limited to steps S810 to S840. The following combinesFigure 8 These four steps will be introduced in detail.

[0145] Step S810: Extract window features from the target segment window data to obtain the target segment window features.

[0146] Step S820: Perform spectral clustering based on the target segment window features to obtain segment window clustering labels and segment window segmentation data.

[0147] Step S830: Segment the vocal speech segments according to the window segmentation data to obtain object speech segments.

[0148] Step S840: Determine the speech clustering data according to the segment window clustering labels and the object speech segments.

[0149] In steps S810 and S820 of some embodiments, window features are extracted from the target segment window data according to the sixth TDNN layer to obtain the target segment window features. The target segment window features refer to the features extracted after frame processing of the vocal speech segments. Spectral clustering is performed based on the target segment window features to obtain segment window clustering labels and segment window segmentation data. The segment window clustering labels are used to represent the objects corresponding to the segment window clustering labels. For example, object A. However, at this time, the segment window clustering labels can only be used to represent the clustering of different objects, but it is not clear which object the object specifically refers to. The segment window segmentation data refers to the time segmentation points for dividing different objects in the vocal speech segments. Therefore, according to this time segmentation point, the speech segments of different objects can be further divided from the vocal speech segments.

[0150] It should be noted that the spectral clustering of the target segment window features can be performed using the similarity matrix of the data, which is suitable for processing non-linear and non-convex data structures, or other spectral clustering methods can also be used, which are not specifically limited herein.

[0151] In steps S830 to S840 of some embodiments, the vocal speech segments are segmented according to the window segmentation data to obtain object speech segments. For example, a vocal speech segment is "Okay, may I help you with anything else? Not for now". According to the process of spectral clustering in the above embodiments, "Okay, may I help you with anything else?" belongs to object 1, and "Not for now" belongs to object 2, and the segment window segmentation data is after "help you". The vocal speech segment is segmented according to the window segmentation data to obtain the object speech segment of object 1 as "Okay, may I help you with anything else?", and the object speech segment of object 2 as "Not for now". In this way, the object speech segments of object 1 can be clustered together, and the object speech segments of object 2 can be clustered together to determine the speech clustering data of each object.

[0152] In step S160 of some embodiments, to avoid the influence of device differences on voiceprint recognition, the present application constructs a voiceprint library according to the speech of different channel types. Therefore, it is necessary to match the corresponding target voiceprint library from the preset candidate voiceprint libraries according to the voice channel label, and multiple reference voiceprint data are stored in the target voiceprint library, and each voiceprint data corresponds to a reference object. For example, in the scenario of customer service voice quality inspection, the candidate voiceprint library stores the voiceprint data of multiple customer service representatives. By matching the voiceprint of the customer in the target voice with the reference voiceprint data in the target voiceprint library, the corresponding customer service representative is determined.

[0153] It can be understood that there are two candidate voiceprint libraries, one is the voiceprint library corresponding to the single-channel label, and the other is the voiceprint library corresponding to the dual-channel label. The voiceprint library corresponding to the label that is the same as the voice channel label is used as the target voiceprint library.

[0154] Please refer to Figure 9 , Figure 9 which is another flowchart of the object classification method based on voiceprint provided by the embodiments of the present application. In some embodiments of the present application, before step S160, the object classification method based on voiceprint provided by the embodiments of the present application further includes: constructing a candidate voiceprint library. Specifically, constructing the candidate voiceprint library may include but is not limited to steps S910 to S980. The following will introduce these eight steps in detail in combination with Figure 9 this.

[0155] Step S910, obtaining sample speech data, where the sample speech data includes sample speech, the sample speech channel label of the sample speech, and the sample object of the sample speech; among them, the sample speech refers to the speech containing multiple objects;

[0156] Step S920, performing voice segmentation processing on the sample speech according to the voice segmentation sub-model to obtain sample voice segments of human voices;

[0157] Step S930, performing voice clustering processing on the sample voice segments of human voices according to the voice clustering sub-model to obtain sample object voice segments; among them, the sample object voice segment refers to the voice segment containing one object;

[0158] Step S940, performing voice decoding processing on the sample object voice segments according to the decoder to obtain sample object voice texts;

[0159] Step S950, performing text classification on the sample object voice texts to obtain sample object classification labels;

[0160] Step S960, if the sample object classification label is the target label, obtaining the sample object voice segments with the sample object classification label to obtain sample matching voice segments;

[0161] Step S970: Extract the voiceprint from the sample-matched voice segment to obtain candidate voiceprint data;

[0162] Step S980: Construct a candidate voiceprint library based on the candidate voiceprint data, the sample voice channel label, and the sample object.

[0163] In step S910 of some embodiments, the sample voice is the same as the target voice in the above embodiments, and the sample voice channel label is the same as the voice channel label in the above embodiments. However, here it is for training purposes and will not be elaborated further. The sample object refers to the object to be matched included in the target voice, such as a customer service object, a delivery object, etc.

[0164] In step S920 and step S930 of some embodiments, step S920 is similar to step S140 of the above embodiments, and step S930 is similar to step S150 of the above embodiments. However, here it is for training purposes and will not be elaborated further.

[0165] In step S940 of some embodiments, during the training process of constructing the voiceprint library, since it is not yet clear which voices are spoken by the objects to be matched, text classification needs to be performed on the text corresponding to the voices. The input of the decoder is the sample object voice segment containing only one speaker, and the output is the corresponding sample object voice text.

[0166] In step S950 of some embodiments, the sample object classification label is used to represent the identity label of the object corresponding to the text. For example, in a customer service scenario, the sample object classification labels include a customer service label and a customer label. Since there are always standard phrases for customer service, such as an opening phrase (Hello, this is **, glad to serve you), a closing phrase (Wish you a happy life, goodbye), and an invitation phrase (If you are satisfied with this service, **). Therefore, when performing text classification on the sample object voice text, by performing entity annotation on the sample voice text and searching for key phrases throughout the text to identify the speaker's identity. For example, if the sample object voice text T1 is "Hello, this is **, glad to serve you", it is determined that this text belongs to the opening text of the customer service based on entity annotation and keyword search. Similarly, the sample object voice text T2 "My friend has a question to ask" belongs to the customer's inquiry statement. The sample object voice text T3 "Okay, what else can I do for you" belongs to the customer service's initiative to offer help. The sample object voice text T4 "I have no questions" belongs to the customer's answer. The sample object voice text T5 "Wish you a happy life, goodbye" belongs to the closing phrase of the customer service. Because after classification, it is determined that the text under the customer service label includes the sample object voice text T1, the sample object voice text T3, and the sample object voice text T5, and the text under the customer label includes the sample object voice text T2 and the sample object voice text T4.

[0167] In step S960 of some embodiments, the target label refers to the label of the object to be matched. For example, when constructing a candidate voiceprint library based on the voiceprint of the customer service, the target label is the customer service label. If the sample object classification label is the target label, obtain the sample object voice segment with the sample object classification label to obtain the sample matching voice segment. That is, use the sample object voice segment corresponding to the target label as the sample matching voice segment.

[0168] In steps S970 to S980 of some embodiments, the candidate voiceprint data refers to the voiceprint data extracted from the sample matching voice segment. According to the correspondence between the candidate voiceprint data, the sample voice channel label, and the sample object, construct a candidate voiceprint library. One candidate voiceprint library corresponds to one sample voice channel label, and one sample object includes a set of candidate voiceprint data under one sample voice channel label. By classifying the object based on the reference voiceprint data filtered from the preset candidate voiceprint library according to the voice channel label, the influence of device differences on voiceprint recognition can be avoided, the voiceprint characteristics of the object can be more accurately recognized, the accuracy of voiceprint matching can be improved, and thus the accuracy of object classification can be improved.

[0169] In step S170 of some embodiments, after determining the reference voiceprint data, object classification is performed on the object voice segment and the reference voiceprint data according to the object classification sub-model to determine the category of the object of the object voice segment.

[0170] Please refer to Figure 10 , Figure 10 which is the specific flowchart of step S170 provided by the embodiments of the present application. In some embodiments of the present application, step S170 may specifically include but is not limited to steps S1010 to S1030. The following will introduce these three steps in detail in combination with Figure 10 this.

[0171] Step S1010: Extract the voiceprint of the object voice segment to obtain the object voiceprint data;

[0172] Step S1020: Calculate the voiceprint similarity according to the object voiceprint data and the reference voiceprint data to obtain the voiceprint similarity data;

[0173] Step S1030: Determine the object classification label of the object voice segment according to the voiceprint similarity data, and the object classification label indicates the category of the object of the object voice segment.

[0174] In steps S1010 to S1030 of some embodiments, the object voiceprint data refers to the voiceprint data extracted from the object voice segment according to the voiceprint extraction sub-model. The present application does not limit the specific method of voiceprint extraction. The voiceprint similarity data is a numerical value used to represent the similarity degree between the object voiceprint data and the reference voiceprint data. The voiceprint similarity calculation can adopt cosine similarity calculation, Euclidean distance, etc. If the voiceprint similarity data is greater than the preset classification threshold, it can be determined that the object voiceprint data and the reference voiceprint data are the same, that is, the object voice segment is the voice issued by the object corresponding to the reference voiceprint data. The preset classification threshold is used to represent the threshold for determining object classification, and can be 95%, 96%, etc. The object classification label indicates the category of the object of the object voice segment. For example, it can be determined whether the voice segment is the voice issued by the customer service. In the customer service scenario, the object classification label includes a customer service label and a customer label.

[0175] Please refer to Figure 11 , Figure 11 which is another flowchart of the object classification method based on voiceprint provided by the embodiments of the present application. In some embodiments of the present application, after step S170, the object classification method based on voiceprint provided by the embodiments of the present application may specifically further include but is not limited to steps S1110 to S1130. The following will introduce these three steps in detail in combination with Figure 11 this.

[0176] Step S1110, if the object classification label is a preset target label, obtain the object voice segment with the object classification label to obtain a matching voice segment;

[0177] Step S1120, perform voice decoding processing on the matching voice segment according to the decoder to obtain a matching object voice text;

[0178] Step S1130, perform keyword detection on the matching object voice text according to the preset keyword library to obtain detection data.

[0179] In steps S1110 to S1130 of some embodiments, if the object classification label is a preset target label, obtain the object voice segment with the object classification label to obtain a matching voice segment. For example, in the process of customer voice quality inspection, the target label is the customer service label. If it is determined that the object classification label is the preset target label, the object voice segment with the object classification label is used as the matching voice segment. Perform voice decoding processing on the matching voice segment according to the decoder to obtain a matching object voice text. Perform keyword detection on the matching object voice text according to the preset keyword library to obtain detection data. The detection data is used to represent the data indicating that the matching object voice text matches the keywords in the keyword library. Therefore, subsequent reminder or reward operations can be performed according to the detection data.

[0180] It should be noted that in the process of customer voice quality inspection, the keyword library can be used to represent the sensitive word library that the customer service is not allowed to say, or the word library that the customer service can be rewarded for saying. For example, if the customer service says a relatively sensitive statement, the corresponding customer service can be reminded in time to make adjustments. Therefore, in the process of voice quality inspection, this solution usually performs behavior detection on both the customer service and the customer. In the voice quality inspection system, when detecting a call, the voiceprint of the customer service extracted in advance can be used to verify the voiceprint of the person speaking, and the role of the speaker can be judged in real time. If the customer service says a service prohibited word, the system can remind the customer service to pay attention to the service attitude at any time.

[0181] The embodiment of the present invention establishes the corresponding relationship between the voice and the actual role through the voice segmentation sub-model, the voice clustering sub-model, the decoder combining speech-to-text, and the text role classification, and performs object classification according to the reference voiceprint data selected from the preset candidate voiceprint library according to the voice channel label, which can avoid the influence of device differences on voiceprint recognition, more accurately identify the voiceprint characteristics of the object, improve the accuracy of voiceprint matching, and thus improve the accuracy of object classification.

[0182] Please refer to Figure 12 , Figure 12It is a block diagram of a voiceprint-based object classification device provided by some embodiments of the present application. In some embodiments, the voiceprint-based object classification device may specifically include a data acquisition module 1210, a detection module 1220, a model acquisition module 1230, a human voice segmentation module 1240, a voice clustering module 1250, a screening module 1260, and a classification module 1270.

[0183] The data acquisition module 1210 is configured to acquire target voice data, where the target voice data includes target voices from multiple target objects;

[0184] The detection module 1220 is configured to perform a voice channel detection on the target voice to obtain a voice channel label of the target voice;

[0185] The model acquisition module 1230 is configured to acquire a pre-trained object classification model, where the object classification model includes a human voice segmentation sub-model, a voice clustering sub-model, and an object classification sub-model;

[0186] The human voice segmentation module 1240 is configured to perform a human voice segmentation process on the target voice according to the human voice segmentation sub-model to obtain human voice segments; where the human voice segments are voice segments that do not include silence;

[0187] The voice clustering module 1250 is configured to perform a voice clustering process on the human voice segments according to the voice clustering sub-model to obtain voice clustering data, where the voice clustering data includes object voice segments of each target object;

[0188] The screening module 1260 is configured to screen out reference voiceprint data from a preset candidate voiceprint library according to the voice channel label;

[0189] The classification module 1270 is configured to perform object classification on the object voice segments and the reference voiceprint data according to the object classification sub-model.

[0190] It should be noted that the voiceprint-based object classification device of the embodiments of the present application is used to implement the above-mentioned voiceprint-based object classification method. The voiceprint-based object classification device of the embodiments of the present application corresponds to the foregoing voiceprint-based object classification method. For the specific processing process, please refer to the foregoing voiceprint-based object classification method, and details are not described herein one by one.

[0191] The embodiments of the present application also provide a computer device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the above-mentioned voiceprint-based object classification method of the embodiments of the present application.

[0192] The computer device can be any intelligent terminal including a mobile phone, a tablet computer, a personal digital assistant (PDA), an in-vehicle computer, etc.

[0193] The following Figure 13 introduces the computer device of the embodiment of the present application in detail.

[0194] As Figure 13 , Figure 13 schematically shows the hardware structure of the computer device of another embodiment. The computer device includes:

[0195] A processor 1310, which can be implemented by using a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present application;

[0196] A memory 1320, which can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 1320 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1320 and are called by the processor 1310 to execute the voiceprint-based object classification method of the embodiments of the present application;

[0197] An input / output interface 1330, which is used to implement information input and output;

[0198] A communication interface 1340, which is used to implement communication interaction between this device and other devices, and can implement communication through a wired method (such as USB, network cable, etc.) or through a wireless method (such as a mobile network, WIFI, Bluetooth, etc.);

[0199] A bus 1350, which transmits information between the various components of the device (such as the processor 1310, the memory 1320, the input / output interface 1330, and the communication interface 1340);

[0200] Among them, the processor 1310, the memory 1320, the input / output interface 1330, and the communication interface 1340 are communicatively connected to each other inside the device through the bus 1350.

[0201] The embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the method for object classification based on voiceprint in the above embodiment.

[0202] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include high-speed random access memory, and can also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory optionally includes a memory remotely disposed relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0203] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art can know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.

[0204] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown in the figures, or combine some steps, or different steps.

[0205] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0206] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices can be implemented as software, firmware, hardware, and appropriate combinations thereof.

[0207] In the description of this application and the above-mentioned drawings, terms such as "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order different from those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0208] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist simultaneously. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or a similar expression means any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0209] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical, or other forms.

[0210] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0211] In addition, in each embodiment of the present application, each functional unit can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0212] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes multiple instructions for causing an electronic device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in the various embodiments of the present application. The aforementioned storage medium includes: various media that can store programs, such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.

[0213] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the rights of the embodiments of the present application. Any modification, equivalent replacement, and improvement made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall be within the scope of the rights of the embodiments of the present application.

Claims

1. A method for object classification based on voiceprint, characterized in that The method includes: Obtaining target voice data, where the target voice data includes target voices from multiple target objects; Performing voice channel detection on the target voice to obtain a voice channel label; Obtaining a pre-trained object classification model, where the object classification model includes a voice segmentation sub-model, a voice clustering sub-model, and an object classification sub-model; Performing voice segmentation processing on the target voice according to the voice segmentation sub-model to obtain voice segments of human voices; wherein, the voice segments of human voices are voice segments that do not include silence; Performing voice clustering processing on the voice segments of human voices according to the voice clustering sub-model to obtain voice clustering data, where the voice clustering data includes object voice segments of each target object; Selecting reference voiceprint data from a preset candidate voiceprint library according to the voice channel label; Performing object classification on the object voice segments and the reference voiceprint data according to the object classification sub-model.

2. The method according to claim 1, characterized in that, The performing object classification on the object voice segments and the reference voiceprint data according to the object classification sub-model includes: Extracting a voiceprint from the object voice segments to obtain object voiceprint data; Calculating a voiceprint similarity according to the object voiceprint data and the reference voiceprint data to obtain voiceprint similarity data; Determining an object classification label of the object voice segments according to the voiceprint similarity data, where the object classification label indicates the category of the object of the object voice segments.

3. The method according to claim 2, wherein The object classification model further includes a decoder; after performing object classification on the object voice segments and the reference voiceprint data according to the object classification sub-model, the method further includes: If the object classification label is a preset target label, obtaining the object voice segments with the object classification label to obtain matching voice segments; Performing voice decoding processing on the matching voice segments according to the decoder to obtain a matching object voice text model; Performing keyword detection on the matching object voice text according to a preset keyword library to obtain detection data.

4. The method according to claim 1, wherein The performing voice segmentation processing on the target voice according to the voice segmentation sub-model to obtain voice segments of human voices includes; Dividing the target voice according to preset first Mel-frequency cepstrum window data to obtain first window voices; Extracting window features from the first window voices to obtain first window voice features; Performing window voice classification according to the first window voice features to obtain initial window labels; Extracting window label pairs according to the initial window labels, where the window label pairs include a first window label and a second window label, and the first window label and the second window label are labels of adjacent first window voices; If the first window label is a human voice label and the second window label is the human voice label, merging the first window voice with the first window label and the first window voice with the second window label to obtain the voice segments of human voices, where the human voice label is used to indicate that the first window voice is a human voice label.

5. The method according to claim 1, wherein Performing voice clustering processing on the human voice speech segment according to the voice clustering sub-model to obtain voice clustering data, including: Performing voice division on the human voice speech segment according to preset second Mel-frequency cepstrum window data to obtain second window voices; Performing feature extraction according to the second window voices to obtain second window voice features; Performing feature merging on the second window voice features according to preset segment window data to obtain segment window voice features; the segment window length of the segment window data is the sum of the window lengths of a preset number of the second Mel-frequency cepstrum window data; Calculating the segment window mean according to the segment window voice features to obtain the segment window mean; Calculating the segment window variance according to the segment window voice features to obtain the segment window variance; Performing numerical splicing on the segment window mean and the segment window variance to obtain target segment window data; Performing segment clustering according to the target segment window data to obtain the voice clustering data.

6. The method according to claim 5, characterized in that Performing segment clustering according to the target segment window data to obtain voice clustering data, including: Performing window feature extraction on the target segment window data to obtain target segment window features; Performing spectral clustering according to the target segment window features to obtain segment window clustering labels and segment window segmentation data; Performing voice segmentation on the human voice speech segment according to the window segmentation data to obtain the object voice segment; Determining the voice clustering data according to the segment window clustering labels and the object voice segment.

7. The method according to claim 3, characterized in that, Before screening reference voiceprint data from a preset candidate voiceprint library according to the voice channel label, the method further includes: constructing the candidate voiceprint library, specifically including: Obtaining sample voice data, where the sample voice data includes a sample voice, the sample voice channel label of the sample voice, and the sample object of the sample voice; wherein, the sample voice refers to a voice containing multiple objects; Performing human voice segmentation processing on the sample voice according to the human voice segmentation sub to obtain sample human voice speech segments; Performing voice clustering processing on the sample human voice speech segments according to the voice clustering sub-model to obtain sample object voice segments; wherein, the sample object voice segment refers to a voice segment containing one object; Performing voice decoding processing on the sample object voice segments according to the decoder to obtain sample object voice texts; Performing text classification on the sample object voice texts to obtain sample object classification labels; If the sample object classification label is the target label, obtaining the sample object voice segments with the sample object classification label to obtain sample matching voice segments; Performing voiceprint extraction on the sample matching voice segments to obtain candidate voiceprint data; Constructing the candidate voiceprint library according to the candidate voiceprint data, the sample voice channel label, and the sample object.

8. An object classification device based on voiceprint, characterized in that, The device includes: A data acquisition module, configured to acquire target voice data, where the target voice data includes target voices from multiple target objects; A detection module, configured to perform a voice channel detection on the target voice to obtain a voice channel label of the target voice; A model acquisition module, configured to acquire a pre-trained object classification model, where the object classification model includes a voice segmentation sub-model, a voice clustering sub-model, and an object classification sub-model; A voice segmentation module, configured to perform a voice segmentation process on the target voice according to the voice segmentation sub-model to obtain a voice segment of human voice; wherein, the voice segment of human voice is a voice segment without silence; A voice clustering module, configured to perform a voice clustering process on the voice segment of human voice according to the voice clustering sub-model to obtain voice clustering data, where the voice clustering data includes an object voice segment of each target object; A screening module, configured to screen out reference voiceprint data from a preset candidate voiceprint library according to the voice channel label; A classification module, configured to perform an object classification on the object voice segment and the reference voiceprint data according to the object classification sub-model.

9. A computer device, characterized in that, The computer device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Call bill retrieval method and device, equipment, medium and product

    CN121509916A