Voice Processing Method, Device, Electronic Device and Storage Medium
By using delay neural networks and enhancement matrix in multi-speaker audio to improve feature similarity, combined with path integration clustering algorithms, the problem of poor sound spectrum separation in the existing technology is solved, and more efficient audio clip separation is achieved.
Patent Information
- Application Number
- CN202111268983.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-29
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2041-10-29
AI Technical Summary
In the audio separation of multiple speakers, the clustering algorithm results are difficult to guide the feature extraction process, resulting in poor sound spectrum separation effect.
By obtaining the initial acoustic features of the input audio, using a time-delay neural network to extract features, and improving feature similarity through enhanced matrix and probability linear discriminant analysis model, and combining path integration clustering algorithm for sound spectrum feature clustering.
The audio clip separation effect of different speakers is improved, and the clustering accuracy and separation effect of sound spectrum characteristics are enhanced.
Smart Images

Figure CN114023349B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and more particularly to the field of speech recognition technology. Specifically, it relates to a speech processing method, apparatus, electronic device, and storage medium. Background Art
[0002] Artificial intelligence is a discipline that studies how to make a computer simulate certain thinking processes and intelligent behaviors of humans (such as learning, reasoning, thinking, planning, etc.). It has both hardware-level technologies and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing; artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech recognition technology, natural language processing technology, and machine learning / deep learning, big data processing technology, and knowledge graph technology.
[0003] In scenarios such as intelligent customer service, conference discussions, interview conversations, public security interrogations, and variety shows, there are often multiple speakers on a single sound channel. Therefore, it is generally necessary to separate the human voices from the recorded audio through a system and then perform targeted analysis. In this process, the voice segments of different speakers can be separated by clustering the spectral features extracted from the audio according to the given number of speakers or a certain clustering threshold.
[0004] The methods described in this section are not necessarily methods that have been previously envisioned or adopted. Unless otherwise specified, no method described in this section should be considered prior art solely because it is included in this section. Similarly, unless otherwise specified, the problems mentioned in this section should not be considered to have been recognized in any prior art. Summary of the Invention
[0005] The present disclosure provides a speech processing method, apparatus, electronic device, and storage medium.
[0006] According to an aspect of the present disclosure, there is provided a speech processing method, including: obtaining a plurality of initial spectral features of an input audio; enhancing the plurality of initial spectral features to obtain corresponding enhanced spectral features; for each enhanced spectral feature among the plurality of enhanced spectral features, determining the similarity between the enhanced spectral feature and other enhanced spectral features; and clustering the plurality of enhanced spectral features based on the similarity between the plurality of enhanced spectral features, where the enhanced spectral features belonging to the same category belong to the same speaker.
[0007] According to another aspect of the present disclosure, a method for training a speech processing model is provided, including: obtaining a plurality of sample initial spectrogram features of a sample audio, where the sample audio includes at least two sample audio segments corresponding to at least two speakers; enhancing the plurality of sample initial spectrogram features to obtain corresponding plurality of sample enhanced spectrogram features; for each sample enhanced spectrogram feature among the plurality of sample enhanced spectrogram features, determining the similarity between the sample enhanced spectrogram feature and other sample enhanced spectrogram features; and clustering the plurality of sample enhanced spectrogram features based on the similarity between the plurality of sample enhanced spectrogram features, where the enhanced spectrogram features belonging to the same category belong to the same speaker; dividing the sample audio based on the clustering result to obtain at least two predicted audio segments corresponding to the at least two speakers; and adjusting parameters of the speech processing model based on the difference between the sample audio segments and the predicted audio segments.
[0008] According to another aspect of the present disclosure, a speech processing device is provided, including: an initial feature obtaining unit configured to obtain a plurality of initial spectrogram features of an input audio; an enhancing unit configured to enhance the plurality of initial spectrogram features to obtain corresponding plurality of enhanced spectrogram features; a similarity determining unit configured to, for each enhanced spectrogram feature among the plurality of enhanced spectrogram features, determine the similarity between the enhanced spectrogram feature and other enhanced spectrogram features; and a clustering unit configured to cluster the plurality of enhanced spectrogram features based on the similarity between the plurality of enhanced spectrogram features, where the enhanced spectrogram features belonging to the same category belong to the same speaker.
[0009] According to another aspect of the present disclosure, a device for training a speech processing model is provided, including: an initial feature obtaining unit configured to obtain a plurality of sample initial spectrogram features of a sample audio, where the sample audio includes at least two sample audio segments corresponding to at least two speakers; an enhancing unit configured to enhance the plurality of sample initial spectrogram features to obtain corresponding plurality of sample enhanced spectrogram features; a similarity determining unit configured to, for each sample enhanced spectrogram feature among the plurality of sample enhanced spectrogram features, determine the similarity between the sample enhanced spectrogram feature and other sample enhanced spectrogram features; and a clustering unit configured to cluster the plurality of sample enhanced spectrogram features based on the similarity between the plurality of sample enhanced spectrogram features, where the enhanced spectrogram features belonging to the same category belong to the same speaker; a prediction unit configured to divide the sample audio based on the clustering result to obtain at least two predicted audio segments corresponding to the at least two speakers; and a training unit configured to adjust parameters of the speech processing model based on the difference between the sample audio segments and the predicted audio segments.
[0010] According to another aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is caused to execute the method as described above.
[0011] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method as described above.
[0012] According to another aspect of the present disclosure, there is provided a computer program product, including a computer program, wherein the computer program, when executed by a processor, implements the method as described above.
[0013] According to one or more embodiments of the present disclosure, audio segments of different speakers can be accurately obtained from audio in which there are multiple speakers.
[0014] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The drawings exemplarily show embodiments and form a part of the specification, and are used together with the written description of the specification to explain the exemplary implementation manners of the embodiments. The shown embodiments are only for illustrative purposes and do not limit the scope of the claims. In all the drawings, the same reference numerals refer to similar but not necessarily identical elements.
[0016] Figure 1 A schematic diagram of an exemplary system in which the various methods and apparatuses described herein can be implemented according to embodiments of the present disclosure is shown.
[0017] Figure 2 An exemplary flowchart of a speech processing method according to an embodiment of the present disclosure is shown.
[0018] Figure 3A An exemplary structure of a directed graph according to an embodiment of the present disclosure is shown.
[0019] Figures 3B - 3H An example of an increasing path when categories A and B are merged is shown.
[0020] Figure 4 An exemplary architecture diagram of a speech processing model according to an embodiment of the present disclosure is shown.
[0021] Figure 5An exemplary flowchart of a method for training a speech processing model according to an embodiment of the present disclosure is shown.
[0022] Figure 6 A block diagram of a speech processing device according to an embodiment of the present disclosure is shown.
[0023] Figure 7 A block diagram of a device for training a speech processing model according to an embodiment of the present disclosure is shown.
[0024] Figure 8 A block diagram of an electronic device that can be applied to an embodiment of the present disclosure is shown. Detailed implementation manners
[0025] The exemplary embodiments of the present disclosure are described below in conjunction with the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0026] In the present disclosure, unless otherwise specified, the terms "first", "second", etc. are used to describe various elements and are not intended to limit the positional relationship, temporal relationship, or importance relationship of these elements. Such terms are only used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of the element, and in certain cases, based on the context description, they may also refer to different instances.
[0027] In the description of various examples in the present disclosure, the terms used are only for the purpose of describing specific examples and are not intended to be limiting. Unless the context clearly indicates otherwise, if the number of elements is not specifically limited, the element may be one or more. In addition, the term "and / or" used in the present disclosure covers any one of the listed items and all possible combinations.
[0028] In the related art, in order to separate audio segments of different speakers present in an audio, a method of extracting the spectral features of the speakers from the audio and clustering the spectral features can be adopted to process the audio. A commonly used algorithm for extracting the spectral features of speakers can be a time delay neural network (TDNN). A commonly used clustering algorithm can be hierarchical clustering (AHC), etc.
[0029] However, the solution for separating speakers adopted in the related art is not end-to-end. In the related art, the result of the clustering algorithm cannot be used to guide the feature extraction process for obtaining the spectral features to be clustered, so it is difficult to improve the effect of spectral separation.
[0030] In view of the above problems, according to one aspect of the present disclosure, a voice processing method is provided. Embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.
[0031] Figure 1 FIG. shows a schematic diagram of an exemplary system 100 in which the various methods and apparatuses described herein may be implemented according to embodiments of the present disclosure. Referring to Figure 1 , the system 100 includes one or more client devices 101, 102, 103, 104, 105, and 106, a server 120, and one or more communication networks 110 that couple the one or more client devices to the server 120. The client devices 101, 102, 103, 104, 105, and 106 may be configured to execute one or more application programs.
[0032] In embodiments of the present disclosure, the server 120 may run one or more services or software applications that enable the execution of the voice processing method of the present disclosure.
[0033] In certain embodiments, the server 120 may also provide other services or software applications that may include non-virtual environments and virtual environments. In certain embodiments, these services may be provided as web-based services or cloud services, for example, provided to users of the client devices 101, 102, 103, 104, 105, and / or 106 under a software as a service (SaaS) model.
[0034] In Figure 1 the configuration shown, the server 120 may include one or more components that implement the functions performed by the server 120. These components may include software components, hardware components, or combinations thereof that may be executed by one or more processors. Users operating the client devices 101, 102, 103, 104, 105, and / or 106 may in turn utilize one or more client applications to interact with the server 120 to utilize the services provided by these components. It should be understood that various different system configurations are possible, which may be different from the system 100. Therefore, Figure 1 is an example of a system for implementing the various methods described herein and is not intended to be limiting.
[0035] Users may use the client devices 101, 102, 103, 104, 105, and / or 106 to generate voice data for the voice processing method of the present disclosure. The client device may provide an interface that enables the user of the client device to interact with the client device. The client device may also output information to the user via the interface. Although Figure 1Only six client devices are depicted, but those skilled in the art will be able to understand that the present disclosure can support any number of client devices.
[0036] Client devices 101, 102, 103, 104, 105, and / or 106 can include various types of computing devices, such as portable handheld devices, general-purpose computers (such as personal computers and laptop computers), workstation computers, wearable devices, gaming systems, thin clients, various messaging devices, sensors, or other sensing devices, etc. These computing devices can run various types and versions of software applications and operating systems, such as MICROSOFT Windows, APPLE iOS, UNIX-like operating systems, Linux, or Linux-like operating systems (such as GOOGLE Chrome OS); or include various mobile operating systems, such as MICROSOFT Windows Mobile OS, iOS, Windows Phone, Android. Portable handheld devices can include cellular phones, smartphones, tablets, personal digital assistants (PDAs), etc. Wearable devices can include head-mounted displays and other devices. Gaming systems can include various handheld gaming devices, Internet-enabled gaming devices, etc. Client devices are capable of executing various different applications, such as various Internet-related applications, communication applications (such as email applications), short message service (SMS) applications, and can use various communication protocols.
[0037] Network 110 can be any type of network known to those skilled in the art, which can support data communication using any one of a variety of available protocols (including but not limited to TCP / IP, SNA, IPX, etc.). By way of example only, one or more networks 110 can be a local area network (LAN), an Ethernet-based network, token ring, wide area network (WAN), the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a public switched telephone network (PSTN), an infrared network, a wireless network (such as Bluetooth, WIFI), and / or any combination of these and / or other networks.
[0038] Server 120 may include one or more general-purpose computers, dedicated server computers (such as PC (Personal Computer) servers, UNIX servers), blade servers, mainframes, server clusters, or any other suitable arrangement and / or combination. Server 120 may include one or more virtual machines running a virtual operating system, or other computing architectures involving virtualization (such as one or more flexible pools of logical storage devices that can be virtualized to maintain virtual storage devices of the server). In various embodiments, Server 120 may run one or more services or software applications that provide the functions described below.
[0039] The computing units in Server 120 may run one or more operating systems including any of the above operating systems as well as any commercially available server operating systems. Server 120 may also run any one of a variety of additional server applications and / or middleware applications, including HTTP servers, FTP servers, CGI servers, JAVA servers, database servers, etc.
[0040] In some embodiments, Server 120 may include one or more applications to analyze and combine data feeds and / or event updates received from users of client devices 101, 102, 103, 104, 105, and 106. Server 120 may also include one or more applications to display data feeds and / or real-time events via one or more display devices of client devices 101, 102, 103, 104, 105, and 106.
[0041] In some embodiments, Server 120 may be a server of a distributed system, or a server combined with a blockchain. Server 120 may also be a cloud server, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology. A cloud server is a host product in the cloud computing service system, which solves the defects of high management difficulty and weak business scalability existing in traditional physical hosts and virtual private server (VPS, Virtual Private Server) services.
[0042] System 100 may also include one or more databases 130. In some embodiments, these databases may be used to store data and other information. For example, one or more of databases 130 may be used to store information such as audio files and video files. Databases 130 may reside at various locations. For example, the databases used by server 120 may be local to server 120, or may be remote from server 120 and may communicate with server 120 via a network-based or dedicated connection. Databases 130 may be of different types. In some embodiments, the databases used by server 120 may be, for example, relational databases. One or more of these databases may store, update, and retrieve data to and from the databases in response to commands.
[0043] In some embodiments, one or more of databases 130 may also be used by an application to store application data. The databases used by the application may be different types of databases, such as key-value repositories, object repositories, or conventional repositories supported by a file system.
[0044] Figure 1 System 100 can be configured and operated in various ways to enable the application of the various methods and apparatuses described in this disclosure.
[0045] Figure 2 An exemplary flowchart of a voice processing method according to an embodiment of this disclosure is shown. It can be executed by the client or server shown in Figure 1 the method 200 shown in Figure 2 .
[0046] In voice processing method 200, as shown in Figure 2 , in step S202, a plurality of initial spectrogram features of the input audio can be obtained. In step S204, the plurality of initial spectrogram features can be enhanced to obtain corresponding enhanced spectrogram features. In step S206, for each enhanced spectrogram feature among the plurality of enhanced spectrogram features, the similarity between the enhanced spectrogram feature and other enhanced spectrogram features is determined. In step S208, the plurality of enhanced spectrogram features can be clustered based on the similarity between the plurality of enhanced spectrogram features, where the enhanced spectrogram features belonging to the same category belong to the same speaker.
[0047] Using the voice processing method provided by the embodiments of this disclosure, the initial spectrogram features obtained by the feature extraction model can be enhanced, thereby improving the clustering effect for spectrogram features, and thus improving the discrimination effect for different speakers in the input audio.
[0048] The methods provided by the embodiments of this disclosure will be described in detail below.
[0049] In step S202, multiple initial acoustic spectral features of the input audio can be obtained.
[0050] In some embodiments, each audio frame in the input audio can be processed based on a time-delay neural network (TDNN) to obtain multiple initial acoustic spectral features. Without departing from the principles of the present disclosure, any other acoustic spectral feature extraction model or a variant of the acoustic spectral feature extraction model can also be used to process the input audio to obtain the above-mentioned multiple initial acoustic spectral features.
[0051] In some implementation manners, an acoustic spectral feature extraction model composed of a multi-layer TDNN, a fully connected layer, and an activation layer can be used to process the features of each audio frame in the input audio. The probability of each predicted speaker is output by the model. Among them, the features of each audio frame can be Mel-frequency cepstral coefficients (MFCC) features, perceptual linear prediction (PLP) features, or Fbank features, etc. The dimension of the features of the audio frame can be 20 dimensions or any other appropriate dimension. Among them, the duration of the audio frame can be 25 ms, and the frame shift can be 10 ms. When processing the features of each audio frame, a certain amount of context can be carried. For example, when it is desired to determine the initial acoustic spectral feature of audio frame A, the audio frame features of the two frames before and after the audio frame can be input into the model together, and the result output by the last TDNN layer is used as the initial acoustic spectral feature of this audio frame. The dimension of the initial acoustic spectral feature can be 128 dimensions or any other appropriate dimension. In some examples, the initial acoustic spectral feature can be determined by means of multi-frame accumulation. For example, the result of the initial acoustic spectral feature can be calculated every 50 frames.
[0052] In some embodiments, the parameters of the above-mentioned acoustic spectral feature extraction model can be obtained through pre-training. For example, open-source data such as Aishell, Librispeech, and SRE04-08 can be used to train the model and determine the parameters in the acoustic spectral feature extraction model.
[0053] In step S204, the multiple initial acoustic spectral features can be enhanced to obtain corresponding multiple enhanced acoustic spectral features.
[0054] In some embodiments, for each of the multiple initial acoustic spectral features, at least one enhancement matrix can be used to process the initial acoustic spectral feature to obtain a corresponding enhanced acoustic spectral feature. In some implementations, the first enhancement matrix and the second enhancement matrix can be used to process the initial acoustic spectral feature to obtain a corresponding enhanced acoustic spectral feature. In some examples, before further processing the feature processed by the first enhancement matrix using the second enhancement matrix, unit length normalization processing can also be performed on the feature processed by the first enhancement matrix. Among them, the first enhancement matrix can be determined based on the whitening model, and the second enhancement matrix can be determined based on the principal component analysis model. In some examples, the first enhancement matrix and the second enhancement matrix can also include bias coefficients. It can be understood that without departing from the principles of the present disclosure, more or fewer enhancement matrices can also be used to perform enhancement processing on the initial acoustic spectral features.
[0055] Among them, the whitening model can be used to remove the redundancy of the initial acoustic spectral features, reduce the correlation between each dimension in the initial acoustic spectral features, and make the parameters of each dimension have the same variance. The principal component analysis model can be used to reduce the feature dimension without losing feature information, thereby reducing the complexity of model calculation. In some examples, the dimension of the initial acoustic spectral feature can be 128 dimensions, while the dimension of the enhanced acoustic spectral feature can be 10 dimensions.
[0056] In step S206, for each of the multiple enhanced acoustic spectral features, determine the similarity between the enhanced acoustic spectral feature and other enhanced acoustic spectral features.
[0057] In some embodiments, the probability linear discriminant analysis (PLDA) model can be used to process the enhanced acoustic spectral feature and another enhanced acoustic spectral feature to obtain the similarity between the enhanced acoustic spectral feature and another enhanced acoustic spectral feature.
[0058] In some examples, the similarity matrix s between the first enhanced acoustic spectral feature and the second enhanced acoustic spectral feature can be determined based on the following formula:
[0059]
[0060] Among them, μ1 represents the first enhanced acoustic spectral feature, μ2 represents the second enhanced acoustic spectral feature, and Q and P represent two analysis matrices used in PLDA. Among them, the parameters of the matrices P and Q can be determined based on the PLDA model pre-trained using open-source data. Among them, PLDA can be used to measure the similarity between the acoustic spectral features of two speakers.
[0061] In step S208, multiple enhanced acoustic spectral features can be clustered based on the similarity between them, where the enhanced acoustic spectral features belonging to the same category belong to the same speaker.
[0062] In some embodiments, step S208 may include determining the path weights between each enhanced acoustic spectral feature based on the similarity between the enhanced acoustic spectral features determined in step S206. According to the path weights, path integral clustering (PIC) can be performed on the multiple enhanced acoustic spectral features obtained in step S204 to cluster the multiple enhanced acoustic spectral features into a predetermined number of categories.
[0063] In some implementations, the PIC model is based on a directed graph structure. Among them, each node in the directed graph represents an enhanced acoustic spectral feature, and each edge between the nodes represents the similarity between two enhanced acoustic spectral features. The directed graph of the PIC model can be represented by the weights between the samples that indicate the enhanced acoustic spectral feature μ i and μ i that are the closest. In some examples, the weights of the edges between the nodes in the PIC model can be determined based on the following formula:
[0064]
[0065] where s(i, j) represents the similarity s calculated based on formula (1) between the feature μ i and the feature μ j , and represents the set of the K samples that are the closest to the feature μ i .
[0066] Figure 3A FIG. shows an exemplary structure of a directed graph according to an embodiment of the present disclosure. As Figure 3A shown, the directed graph 300 may include nodes 1, 2, 3, and 4, and the edges between the nodes represent the similarity between the two end nodes.
[0067] In Figure 3A the example shown, all possible paths in the directed graph 300 may include P 12 , P 13 , P 13 P 32 , P 32 , P 43 , P 43 P 32 . The weight of the edge P ij between node i and node j can be calculated based on formula (2). Further, it can be based on P 12 +P 13 +P 13 P32 +P 32 +P 43 +P 43 P 32 The sum of the weights of all paths in the directed graph 300 is calculated by summing the results of 12 , P 13 , P 13 P 32 , P 32 , P 43 , P 43 P 32 The weight value of .
[0068] For a directed graph, the total weight of the graph can be obtained by adding up the weights of all paths in the graph. The larger the sum, the more stable the graph.
[0069] When clustering N enhanced sound spectrum features, each enhanced sound spectrum feature can be clustered together based on the similarity between the features. For example, you can select an enhanced sound spectrum feature μ that is not clustered. i , and select the enhanced spectral features that are related to μ from the remaining enhanced spectral features that have not been clustered. i The enhanced spectral feature μ with the highest similarity j , and μ i and μ j Then, except for the clustered features μ i and μ j Repeat the above steps for other enhanced sound spectrum features except N until all enhanced sound spectrum features are processed. In this way, the N enhanced sound spectrum features can be clustered into a number of classes that is half the total number of samples (when N is an odd number, the total number of classes can be half the total number of samples plus one). The two features in each class can be represented as a directed graph, and the weight of the edge connecting the two features can be calculated using formula (2).
[0070] The obtained graphs of each class can then be merged. Whether graph B and graph A can be merged depends on whether the affinity between graph B and graph A is greater than the affinity between all other graphs and graph A. In some embodiments, the affinity between two graphs can be represented based on the path score that increases after the two graphs are merged.
[0071] Figures 3B - 3H An example of an addition path when category A and category B are merged is shown.
[0072] like Figure 3BAs shown, if Category A and Category B are merged, the increasing path when merging Category A and Category B can be determined based on the edge from Node 3 of Category A pointing to Node 2 of Category B and the edge from Node 1 of Category B pointing to Node 1 of Category A. Among them, establishing the edge from Node 3 of Category A pointing to Node 2 of Category B indicates that for each node in Category A, the similarity between Node 2 of Category B and Node 3 of Category A is the highest. Similarly, establishing the edge from Node 1 of Category B pointing to Node 1 of Category A indicates that for each node in Category B, the similarity between Node 1 of Category A and Node 1 of Category B is the highest.
[0073] Furthermore, Figures 3C - 3E shows the increasing path for Category A after merging Category A and Category B. In Figures 3C - 3E three increasing paths after merging are shown in a dashed line. Figures 3F - 3H shows the increasing path for Category B after merging Category A and Category B. In Figures 3F - 3H three increasing paths after merging are shown in a dashed line. By calculating the path scores corresponding to the increasing paths for Category A and Category B respectively, the increasing path score after merging Category A and Category B can be obtained, that is, the sum of the path scores corresponding to the increasing paths for Category A and Category B.
[0074] Using Figures 3B - 3H the method shown in, the increasing path score when Category A is merged with other categories different from Category B can be further calculated, and finally the category with the highest increasing path score is selected for merging with Category A.
[0075] By continuously repeating the above operation of merging categories, N enhanced acoustic spectrum features can be clustered into a predetermined number of categories. In some embodiments, taking the number of speakers as 2 as an example, the above operation of merging categories can be repeated until the N enhanced acoustic spectrum features are clustered into 2 categories. When the number of speakers is other values (such as any natural number greater than 2), the above operation of merging categories can also be repeated until the N enhanced acoustic spectrum features are clustered into the same number of categories as the speakers.
[0076] By using the path score calculated based on formula (2) to represent the affinity between different categories, feature clustering can be effectively achieved based on the similarity between features.
[0077] In some embodiments, Method 200 may further include partitioning the input audio based on the clustering result to obtain audio segments corresponding to the speakers. Based on the clustering result, it can be determined that the acoustic spectrum features belonging to the same category correspond to the same speaker. Therefore, the audio segments in the input audio belonging to the same speaker can be determined based on the audio frames corresponding to the acoustic spectrum features.
[0078] Figure 4 An exemplary architecture diagram of a speech processing model according to an embodiment of the present disclosure is shown. The speech processing method 200 described in combination can be implemented using the speech processing model 400 shown in Figure 4 Figure 2 .
[0079] As Figure 4 shown, the speech processing model 400 may include a speech feature extraction unit 410, a feature enhancement unit 420, a similarity determination unit 430, and a clustering unit 440.
[0080] The speech feature extraction unit 410 can be used to extract multiple initial spectrogram features x from the input audio. In some embodiments, the speech feature extraction unit may be a spectrogram feature extraction model composed of a multi-layer TDNN network, a fully connected layer, and an activation layer. In some implementation manners, the parameters of the speech feature extraction unit are pre-trained using open-source data. The parameters of the speech feature extraction unit are not adjusted during the training process of the speech processing model 400.
[0081] The feature enhancement unit 420 can be used to perform feature enhancement on the initial spectrogram features output by the speech feature extraction unit 410 to obtain enhanced spectrogram features μ. In some embodiments, the initial spectrogram features can be processed using a first enhancement matrix and a second enhancement matrix to obtain corresponding enhanced spectrogram features. In some examples, before further processing the features processed by the first enhancement matrix using the second enhancement matrix, unit length normalization processing can also be performed on the features processed by the first enhancement matrix. Among them, the first enhancement matrix can be determined based on a whitening model, and the second enhancement matrix can be determined based on a principal component analysis model. During the training process of the speech processing model 400, the parameters of the feature enhancement unit can be adjusted through the clustering results to further improve the clustering effect.
[0082] The similarity determination unit 430 can be used to determine the correlation between multiple enhanced spectrogram features output by the feature enhancement unit 420. The similarity between the first enhanced spectrogram feature and the second enhanced spectrogram feature can be calculated based on the above formula (1). Among them, during the training process of the speech processing model 400, the parameters of the analysis matrices P and Q used in formula (1) can be adjusted through the clustering results to further improve the clustering effect.
[0083] The clustering unit 440 can use the similarity output by the similarity unit 430 to cluster multiple enhanced spectrogram features output by the feature enhancement unit 420. The clustering unit 440 can be based on the combination of Figure 2 The described step S208 clusters the enhanced acoustic spectral features to obtain a predetermined number of categories, where each category corresponds to the same speaker. Based on the audio frames corresponding to the acoustic spectral features in each category, audio segments corresponding to the same speaker can be determined from the input audio. By comparing the differences between the audio segments of the same speaker obtained based on the output of the clustering unit 440 and the audio segments corresponding to different speakers in reality, the parameters of the feature enhancement unit 420 and the similarity determination unit 430 can be adjusted, so as to guide the parameters of the feature enhancement unit and the parameters of the similarity determination unit based on the clustering result, and improve the clustering effect.
[0084] Figure 5 FIG. shows an exemplary flowchart of a method for training a speech processing model according to an embodiment of the present disclosure. The training method 500 shown in can be executed by the client or server shown in. The method shown in can be used to train the parameters of the speech processing model shown in. Figure 1 executed by the client or server shown in Figure 5 The training method 500 shown in Figure 5 The method shown in can be used to Figure 4 train the parameters of the speech processing model shown in.
[0085] In step S502, a plurality of sample initial acoustic spectral features of the sample audio can be obtained, where the sample audio includes at least two sample audio segments corresponding to at least two speakers. In step S504, the plurality of sample initial acoustic spectral features can be enhanced to obtain corresponding plurality of sample enhanced acoustic spectral features. In step S506, for each sample enhanced acoustic spectral feature among the plurality of sample enhanced acoustic spectral features, the similarity between the sample enhanced acoustic spectral feature and other sample enhanced acoustic spectral features can be determined. In step S508, the plurality of sample enhanced acoustic spectral features can be clustered based on the similarity between the plurality of sample enhanced acoustic spectral features, where the enhanced acoustic spectral features belonging to the same category belong to the same speaker.
[0086] The sample audio can be processed by using the speech processing method 200 described in combination with Figure 2 to obtain the clustering result of the sample audio, which will not be elaborated here.
[0087] In step S510, the sample audio can be divided based on the clustering result to obtain at least two predicted audio segments corresponding to at least two speakers.
[0088] In step S512, the parameters of the speech processing model can be adjusted based on the differences between the sample audio segments and the predicted audio segments.
[0089] As described above, in step S502, each audio frame in the sample audio can be processed based on a pre-trained time-delay neural network to obtain multiple sample initial spectrogram features. The open-source data Aishell, Librispeech, and SRE04-08 can be used to train the model and determine the parameters in the time-delay neural network. Step S512 may not include adjusting the parameters of the time-delay neural network based on the difference between the sample audio segment and the predicted audio segment. When pre-training the time-delay neural network, cross-entropy can be used as the loss function.
[0090] As described above, in step S506, the PLDA model can be used to process the sample enhanced spectrogram features and another sample enhanced spectrogram feature to obtain the similarity between the sample enhanced spectrogram feature and the another sample enhanced spectrogram feature. The PLDA model can be formed by a first analysis matrix P and a second analysis matrix Q. Among them, the initial parameters of the first analysis matrix P and the second analysis matrix Q can be obtained through pre-training. In some embodiments, the features extracted by a pre-trained spectrogram feature extraction model can be used to pre-train the first analysis matrix P and the second analysis matrix Q in the PLDA model to obtain the initial parameters of the first analysis matrix P and the second analysis matrix Q. Similarly, the features extracted by a pre-trained spectrogram feature extraction model can be used for pre-training to obtain the parameters of the whitening model and the principal component analysis model as the initial parameters of the first enhancement matrix and the second enhancement matrix.
[0091] In some embodiments, the parameters of the first analysis matrix and the second analysis matrix are adjusted based on the difference between the sample audio segment and the predicted audio segment. Thus, the similarity analysis between features can be guided based on the clustering result, thereby improving the accuracy of calculating the similarity between features, and thus further improving the clustering effect of the speech processing model.
[0092] As described above, in step 504, for each sample initial spectrogram feature among the multiple sample initial spectrogram features, at least one enhancement matrix can be used to process the sample initial spectrogram feature to obtain the corresponding sample enhanced spectrogram feature. For example, the first enhancement matrix and the second enhancement matrix are used to process the initial spectrogram feature to obtain the corresponding sample enhanced spectrogram feature. In some examples, the initial parameters of the first enhancement matrix can be based on the parameters of a pre-trained whitening model, and the initial parameters of the second enhancement matrix can be based on the parameters of a pre-trained principal component analysis model. By setting the initial parameters of the first enhancement matrix and the second enhancement matrix to the parameters of the pre-trained whitening model and the pre-trained principal component analysis model respectively, the first enhancement matrix and the second enhancement matrix can respectively achieve the enhancement effects of the whitening model and the principal component analysis model.
[0093] In some embodiments, the parameters of the first enhancement matrix and the second enhancement matrix may be adjusted based on the difference between the sample audio segment and the predicted audio segment, so that the enhancement effects of the first enhancement matrix and the second enhancement matrix are more suitable for the current application scenario.
[0094] In some embodiments, step S512 may include: calculating the difference between the sample audio segment and the predicted audio segment by using the loss function of binary cross entropy, and adjusting the parameters of the speech processing model based on the difference output by the loss function. For example, the parameters of the first analysis matrix, the second analysis matrix, the first enhancement matrix, and the second enhancement matrix may be adjusted backward by using the difference output by the loss function.
[0095] By using the method provided in the present disclosure, the parameters in the speech processing model can be adjusted based on the clustering result, so that the clustering result can guide the parameter training in the speech training model, thereby further improving the clustering effect, that is, improving the recognition accuracy of the audio segments of multiple speakers existing in the same audio.
[0096] According to another aspect of the present disclosure, a speech processing apparatus is also provided. Figure 6 The block diagram of a speech processing apparatus 600 according to an embodiment of the present disclosure is shown.
[0097] As Figure 6 shown, the speech processing apparatus 600 may include an initial feature acquisition unit 610, an enhancement unit 620, a similarity determination unit 630, and a clustering unit 640.
[0098] The initial feature acquisition unit 610 may be configured to acquire multiple initial spectrogram features of the input audio. The enhancement unit 620 may be configured to enhance the multiple initial spectrogram features to obtain corresponding multiple enhanced spectrogram features. The similarity determination unit 630 may be configured to determine the similarity between each enhanced spectrogram feature among the multiple enhanced spectrogram features and other enhanced spectrogram features. The clustering unit 640 may be configured to cluster the multiple enhanced spectrogram features based on the similarity between the multiple enhanced spectrogram features, where the enhanced spectrogram features belonging to the same category belong to the same speaker.
[0099] The operations performed by the above units 610 to 640 correspond to steps S202 to S208 described with reference to Figure 2 and thus the details of each aspect will not be described again.
[0100] According to another aspect of the present disclosure, a device for training a speech processing model is also provided. Figure 7 The block diagram of a device 700 for training a speech processing model according to an embodiment of the present disclosure is shown.
[0101] As shown Figure 7 in FIG. 700, the apparatus may include an initial feature acquisition unit 710, an enhancement unit 720, a similarity determination unit 730, a clustering unit 740, a prediction unit 750, and a training unit 760.
[0102] The initial feature acquisition unit 710 may be configured to acquire a plurality of sample initial spectrogram features of the sample audio, where the sample audio includes at least two sample audio segments corresponding to at least two speakers. The enhancement unit 720 may be configured to enhance the plurality of sample initial spectrogram features to obtain corresponding plurality of sample enhanced spectrogram features. The similarity determination unit 730 may be configured to determine, for each sample enhanced spectrogram feature among the plurality of sample enhanced spectrogram features, the similarity between the sample enhanced spectrogram feature and other sample enhanced spectrogram features. The clustering unit 740 may be configured to cluster the plurality of sample enhanced spectrogram features based on the similarity between the plurality of sample enhanced spectrogram features, where the enhanced spectrogram features belonging to the same category belong to the same speaker. The prediction unit 750 may be configured to divide the sample audio based on the clustering result to obtain at least two predicted audio segments corresponding to the at least two speakers. The training unit 760 may be configured to adjust the parameters of the speech processing model based on the difference between the sample audio segments and the predicted audio segments.
[0103] The operations performed by the above units 710 to 760 correspond to the steps S502 to S512 described with reference to Figure 5 and thus details of each aspect thereof will not be elaborated further.
[0104] According to another aspect of the present disclosure, there is also provided an electronic device, including at least one processor; and a memory communicatively connected to the at least one processor, where the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is caused to execute the method as described above.
[0105] According to another aspect of the present disclosure, there is also provided a non-transitory computer-readable storage medium storing computer instructions, where the computer instructions are used to cause a computer to execute the method as described above.
[0106] According to another aspect of the present disclosure, there is also provided a computer program product, including a computer program, where the computer program, when executed by a processor, implements the method as described above.
[0107] Referring to Figure 8, a block diagram of an electronic device 800 that can be a server or a client of the present disclosure will now be described. It is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0108] As Figure 8 shown, the electronic device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the electronic device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0109] Multiple components in the electronic device 800 are connected to the I / O interface 805, including: an input unit 806, an output unit 807, a storage unit 808, and a communication unit 809. The input unit 806 can be any type of device that can input information into the electronic device 800. The input unit 806 can receive input digital or character information, and generate key signal inputs related to the user settings and / or function controls of the electronic device, and can include, but is not limited to, a mouse, a keyboard, a touch screen, a trackpad, a trackball, a joystick, a microphone, and / or a remote control. The output unit 807 can be any type of device that can present information, and can include, but is not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 808 can include, but is not limited to, magnetic disks, optical disks. The communication unit 809 allows the electronic device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks, and can include, but is not limited to, a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as Bluetooth TM devices, 802.11 devices, WiFi devices, WiMax devices, cellular communication devices, and / or the like.
[0110] The computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 executes the various methods and processes described above, such as the speech processing method. For example, in some embodiments, the speech processing method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the speech processing method described above can be executed. Alternatively, in other embodiments, the computing unit 801 can be configured to execute the speech processing method by any other suitable means (e.g., by means of firmware).
[0111] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0112] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0113] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0114] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).
[0115] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0116] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is generated by computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, a server of a distributed system, or a server incorporating a blockchain.
[0117] It should be understood that the various forms of processes shown above can be used, with steps reordered, added or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitation is imposed herein.
[0118] In the technical solutions of this disclosure, the acquisition, storage and application of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs. The intention of this disclosure is that personal information data should be managed and processed in a way that minimizes the risk of inadvertent or unauthorized use and access. The risk is minimized by restricting data collection and deleting data when it is no longer needed. It should be noted that all information related to personnel in this application is collected with the knowledge and consent of the personnel.
[0119] Although the embodiments or examples of this disclosure have been described with reference to the accompanying drawings, it should be understood that the above methods, systems and devices are merely exemplary embodiments or examples, and the scope of the present invention is not limited by these embodiments or examples, but is only limited by the authorized claims and their equivalent scope. Various elements in the embodiments or examples can be omitted or replaced by their equivalent elements. In addition, the steps can be executed in an order different from that described in this disclosure. Further, the various elements in the embodiments or examples can be combined in various ways. Importantly, with the evolution of technology, many of the elements described herein can be replaced by equivalent elements that emerge after this disclosure.
Claims
1. A voice processing method, comprising: Obtaining a plurality of initial spectrogram features of an input audio; Enhancing the plurality of initial spectrogram features to obtain corresponding enhanced spectrogram features; For each enhanced spectrogram feature among the plurality of enhanced spectrogram features, determining the similarity between the enhanced spectrogram feature and other enhanced spectrogram features; And Clustering the plurality of enhanced spectrogram features based on the similarity between the plurality of enhanced spectrogram features, wherein the enhanced spectrogram features belonging to the same category belong to the same speaker, Wherein, enhancing the plurality of initial spectrogram features to obtain corresponding enhanced spectrogram features includes: For each initial spectrogram feature among the plurality of initial spectrogram features, processing the initial spectrogram feature by using a first enhancement matrix and a second enhancement matrix to obtain a corresponding enhanced spectrogram feature, wherein the first enhancement matrix is determined based on a whitening model, and the second enhancement matrix is determined based on a principal component analysis model.
2. The speech processing method according to claim 1, wherein, For each enhanced spectrogram feature among the plurality of enhanced spectrogram features, determining the similarity between the enhanced spectrogram feature and other enhanced spectrogram features includes: Processing the enhanced spectrogram feature and another enhanced spectrogram feature based on a probabilistic linear discriminant analysis model to obtain the similarity between the enhanced spectrogram feature and the other enhanced spectrogram feature.
3. The voice processing method according to claim 1, wherein, Clustering the plurality of enhanced spectrogram features based on the similarity between the plurality of enhanced spectrogram features includes: Determining path weights between the respective enhanced spectrogram features based on the similarity; Performing path integration clustering on the plurality of enhanced spectrogram features according to the path weights to cluster the plurality of enhanced spectrogram features into a predetermined number of categories.
4. The voice processing method according to claim 1, wherein, Obtaining a plurality of initial spectrogram features of an input audio includes: Processing each audio frame in the input audio based on a time delay neural network to obtain the plurality of initial spectrogram features.
5. The voice processing method according to claim 1, further comprising: Dividing the input audio based on the clustering result to obtain audio segments corresponding to speakers.
6. A method for training a voice processing model, comprising: Obtaining a plurality of sample initial spectrogram features of a sample audio, wherein the sample audio includes at least two sample audio segments corresponding to at least two speakers; Enhancing the plurality of sample initial spectrogram features to obtain corresponding sample enhanced spectrogram features; For each sample enhanced spectrogram feature among the plurality of sample enhanced spectrogram features, determining the similarity between the sample enhanced spectrogram feature and other sample enhanced spectrogram features; Clustering the plurality of sample enhanced spectrogram features based on the similarity between the plurality of sample enhanced spectrogram features, wherein the enhanced spectrogram features belonging to the same category belong to the same speaker; Dividing the sample audio based on the clustering result to obtain at least two predicted audio segments corresponding to the at least two speakers; And Adjusting the parameters of the voice processing model based on the difference between the sample audio segments and the predicted audio segments, Wherein, enhancing the plurality of sample initial spectrogram features to obtain corresponding sample enhanced spectrogram features includes: For each of the multiple initial acoustic spectral features of the samples, the initial acoustic spectral feature is processed using a first enhancement matrix and a second enhancement matrix to obtain a corresponding enhanced acoustic spectral feature of the sample, where the initial parameters of the first enhancement matrix are based on a pre-trained whitening model, and the initial parameters of the second enhancement matrix are based on a pre-trained principal component analysis model.
7. The method according to claim 6, wherein, For each of the multiple enhanced acoustic spectral features of the samples, determining the similarity between the enhanced acoustic spectral feature of the sample and the enhanced acoustic spectral features of other samples includes: Processing the enhanced acoustic spectral feature of the sample and another enhanced acoustic spectral feature based on a probabilistic linear discriminant analysis model to obtain the similarity between the enhanced acoustic spectral feature of the sample and the other enhanced acoustic spectral feature.
8. The method according to claim 7, wherein, The probabilistic linear discriminant analysis model includes a first analysis matrix and a second analysis matrix, and the initial parameters of the first analysis matrix and the second analysis matrix are obtained through pre-training.
9. The method according to claim 8, wherein, Adjusting the parameters of the speech processing model based on the difference between the sample audio segment and the predicted audio segment includes: Adjusting the parameters of the first analysis matrix and the second analysis matrix based on the difference between the sample audio segment and the predicted audio segment.
10. The method according to claim 6, wherein Obtaining multiple initial acoustic spectral features of a sample audio includes: Processing each audio frame in the sample audio based on a pre-trained time-delay neural network to obtain the multiple initial acoustic spectral features.
11. The method according to claim 6, wherein, Adjusting the parameters of the speech processing model based on the difference between the sample audio segment and the predicted audio segment includes: Calculating the difference between the sample audio segment and the predicted audio segment using a loss function of binary cross-entropy, Adjusting the parameters of the speech processing model based on the difference output by the loss function.
12. A speech processing device, comprising: An initial feature acquisition unit configured to acquire multiple initial acoustic spectral features of an input audio; An enhancement unit configured to enhance the multiple initial acoustic spectral features to obtain corresponding multiple enhanced acoustic spectral features; A similarity determination unit configured to, for each of the multiple enhanced acoustic spectral features, determine the similarity between the enhanced acoustic spectral feature and the enhanced acoustic spectral features of other samples; And A clustering unit configured to cluster the multiple enhanced acoustic spectral features based on the similarity between the multiple enhanced acoustic spectral features, where the enhanced acoustic spectral features belonging to the same category belong to the same speaker, where enhancing the multiple initial acoustic spectral features to obtain corresponding multiple enhanced acoustic spectral features includes: For each of the multiple initial acoustic spectral features, processing the initial acoustic spectral feature using a first enhancement matrix and a second enhancement matrix to obtain a corresponding enhanced acoustic spectral feature, where the first enhancement matrix is determined based on a whitening model, and the second enhancement matrix is determined based on a principal component analysis model.
13. A device for training a speech processing model, comprising: An initial feature acquisition unit configured to acquire multiple sample initial acoustic spectral features of a sample audio, where the sample audio includes at least two sample audio segments corresponding to at least two speakers; An enhancement unit configured to enhance the initial acoustic spectral features of the plurality of samples to obtain corresponding enhanced acoustic spectral features of the plurality of samples; A similarity determination unit configured to determine, for each enhanced acoustic spectral feature of the plurality of enhanced acoustic spectral features, the similarity between the enhanced acoustic spectral feature and other enhanced acoustic spectral features; A clustering unit configured to cluster the plurality of enhanced acoustic spectral features based on the similarities between the plurality of enhanced acoustic spectral features, wherein the enhanced acoustic spectral features belonging to the same category belong to the same speaker; A prediction unit configured to partition the sample audio based on the clustering result to obtain at least two predicted audio segments corresponding to the at least two speakers; And A training unit configured to adjust the parameters of the speech processing model based on the difference between the sample audio segment and the predicted audio segment, wherein enhancing the initial acoustic spectral features of the plurality of samples to obtain corresponding enhanced acoustic spectral features of the plurality of samples includes: For each initial acoustic spectral feature of the plurality of initial acoustic spectral features, processing the initial acoustic spectral feature with a first enhancement matrix and a second enhancement matrix to obtain a corresponding enhanced acoustic spectral feature, wherein the initial parameters of the first enhancement matrix are based on a pre-trained whitening model, and the initial parameters of the second enhancement matrix are based on a pre-trained principal component analysis model.
14. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-11.
15. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-11.
16. A computer program product, comprising a computer program, wherein, The computer program, when executed by a processor, implements the method according to any one of claims 1-11.
Citation Information
Patent Citations
Speaker confirmation method and system for coping with complex acoustic environment and storage medium
CN111986679A