Audio signal processing, conference recording and presentation methods, devices, systems, and media

By segmenting audio signals in multi-person speaking scenarios and combining duration and voiceprint features for hierarchical clustering, the problem of high speaker identification error rate in existing technologies is solved, achieving more accurate speaker identification and labeling.

CN114792522BActive Publication Date: 2026-01-02ALIBABA GROUP HOLDING LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202110105959.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-01-26
Publication Date
2026-01-02
Estimated Expiration
2041-01-26

AI Technical Summary

Technical Problem

In multi-person speaking scenarios, existing technologies for speaker identification based on voiceprint features have a high false positive rate, especially with low accuracy under noise interference and emotional changes.

Method used

The audio signal is segmented into multiple segments by identifying speaker change points. The audio segments are then hierarchically clustered based on their duration and voiceprint features to identify and add user tags.

Benefits of technology

It improves the accuracy and efficiency of speaker identification, reduces errors caused by audio segments with unstable voiceprint features, and ensures the accuracy of user tagging results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114792522B_ABST
    Figure CN114792522B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide an audio signal processing, conference recording and presentation method, device, system and medium. In the embodiments of the present application, for the audio signal of a multi-person speaking scene, the audio signal is first cut into multiple audio segments based on speaker change points, and then the multiple audio segments are clustered in multiple levels according to the time length and voiceprint features of the multiple audio segments to identify the audio segments corresponding to the same speaker and add user labels. Among them, instead of simply using voiceprint features for clustering, the time length and voiceprint features of the audio segments are combined for hierarchical clustering. Hierarchical clustering can first cluster the audio segments with more stable voiceprint features. Compared with clustering all audio segments at the same time, hierarchical clustering can reduce the error caused by audio segments with unstable voiceprint features, can more accurately identify the audio segments corresponding to the same speaker, improve the identification efficiency, and the user label result is more accurate.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of audio signal processing, and in particular to an audio signal processing method, a conference recording and presentation method, a device, a system and a medium. BACKGROUND

[0002] In a multi-speaker scene such as a conference or a court trial, in order to meet the demand for recording the conference content, some products with voice collection functions are usually used, such as a sound pickup device, a voice recorder, etc., to collect voice signals in the multi-speaker scene in real time. Based on the voice signals collected by these products, the speaking content in the multi-speaker scene can be directly queried based on the voice signals, or the voice signals can be transcribed into text for query.

[0003] In order to facilitate understanding of the speaker information corresponding to the speaking content when querying, after the voice signals are collected, the speaker needs to be identified, that is, it is identified which speaking content is spoken by which speaker. In the prior art, a neural network model is used to extract the voiceprint features in the voice signals, and the speaking content corresponding to the same speaker is distinguished according to the voiceprint features.

[0004] However, in actual application, there may be strong noise interference in the multi-speaker scene, and the voiceprint features of the speaker may also change due to emotional influence, which will cause misjudgment of the recognition result based on the voiceprint features and low recognition accuracy. SUMMARY

[0005] Aspects of the present application provide an audio signal processing method, a conference recording and presentation method, a device, a system and a medium, which can more accurately identify audio segments corresponding to the same speaker and improve the efficiency of identification.

[0006] The audio signal processing method provided by the embodiments of the present application comprises: identifying a speaker change point in an audio signal collected in a multi-speaker scene; dividing the audio signal into a plurality of audio segments according to the speaker change point, and extracting voiceprint features of the plurality of audio segments; performing hierarchical clustering on the plurality of audio segments according to the time length and the voiceprint features of the plurality of audio segments, to obtain audio segments corresponding to the same speaker; and adding the same user label to the audio segments corresponding to the same speaker, to obtain an audio signal with a user label.

[0007] The embodiment of the present application further provides an audio signal processing method, comprising: performing sound source positioning on an audio signal collected in a multi-person speaking scenario to obtain a change point of a sound source position; dividing the audio signal into a plurality of audio segments according to the change point of the sound source position, and extracting voiceprint features of the plurality of audio segments; clustering the plurality of audio segments according to the voiceprint features of the plurality of audio segments and the sound source position to obtain audio segments corresponding to the same speaker; and adding the same user mark to the audio segments corresponding to the same speaker to obtain an audio signal with the user mark added.

[0008] The embodiment of the present application further provides a conference recording method, comprising: collecting an audio signal in a multi-person conference scenario, and identifying a speaker change point in the audio signal; dividing the audio signal into a plurality of audio segments according to the speaker change point, and extracting voiceprint features of the plurality of audio segments; hierarchically clustering the plurality of audio segments according to the time length and the voiceprint features of the plurality of audio segments to obtain audio segments corresponding to the same speaker; adding the same user mark to the audio segments corresponding to the same speaker, and generating conference recording information according to the audio signal with the user mark added, wherein the conference recording information comprises a conference identifier.

[0009] The embodiment of the present application further provides a conference recording presentation method, comprising: receiving a conference review request, wherein the conference review request comprises a conference identifier to be presented; obtaining conference recording information to be presented according to the conference identifier; and presenting the conference recording information, wherein the conference recording information is generated according to an audio signal with a user mark added in a multi-person conference scenario; wherein the same user mark is added to audio segments corresponding to the same speaker among a plurality of audio segments divided according to a speaker change point in the audio signal, and the audio segments corresponding to the same speaker are obtained by hierarchically clustering the plurality of audio segments according to the time length and the voiceprint features of the plurality of audio segments. The embodiment of the present application further provides an audio processing system, comprising: a sound pickup device and a server device; the sound pickup device is deployed in a multi-person speaking scenario, and is configured to collect an audio signal in a multi-person speaking scenario, identify a speaker change point in the audio signal, divide the audio signal into a plurality of audio segments according to the speaker change point, and extract voiceprint features corresponding to the plurality of audio segments; and the server device is configured to hierarchically cluster the plurality of audio segments according to the time length and the voiceprint features of the plurality of audio segments to obtain audio segments corresponding to the same speaker, add the same user mark to the audio segments corresponding to the same speaker, and obtain an audio signal with the user mark added.

[0010] The embodiment of the present application also provides an audio processing system, comprising: a pickup device and a server device; the pickup device is arranged in a multi-person speaking scene and is used for collecting an audio signal in the multi-person speaking scene, identifying a speaker change point in the audio signal, and cutting the audio signal into a plurality of audio segments according to the speaker change point; and the server device is used for extracting a voiceprint feature corresponding to the plurality of audio segments, performing hierarchical clustering on the plurality of audio segments according to the time length and the voiceprint feature of the plurality of audio segments, obtaining audio segments corresponding to the same speaker, and adding the same user mark to the audio segments corresponding to the same speaker to obtain an audio signal with the user mark.

[0011] The embodiment of the present application also provides a pickup device, comprising: a processor and a memory; the memory is used for storing a computer program; the processor is coupled with the memory and is used for executing the computer program to identify a speaker change point in an audio signal collected in a multi-person speaking scene, cut the audio signal into a plurality of audio segments according to the speaker change point, and extract a voiceprint feature of the plurality of audio segments; perform hierarchical clustering on the plurality of audio segments according to the time length and the voiceprint feature of the plurality of audio segments to obtain audio segments corresponding to the same speaker; and add the same user mark to the audio segments corresponding to the same speaker to obtain an audio signal with the user mark.

[0012] The embodiment of the present application also provides a pickup device, comprising: a processor and a memory; the memory is used for storing a computer program; the processor is coupled with the memory and is used for executing the computer program to perform sound source positioning on an audio signal collected in a multi-person speaking scene to obtain a change point of a sound source position, cut the audio signal into a plurality of audio segments according to the change point of the sound source position, and extract a voiceprint feature of the plurality of audio segments; perform clustering on the plurality of audio segments according to the voiceprint feature and the sound source position of the plurality of audio segments to obtain audio segments corresponding to the same speaker; and add the same user mark to the audio segments corresponding to the same speaker to obtain an audio signal with the user mark.

[0013] The embodiment of the present application also provides a server device, comprising: a processor and a memory; the memory is used for storing a computer program; the processor is coupled with the memory and is used for executing the computer program to receive a plurality of audio segments and corresponding voiceprint features sent by a pickup device, perform hierarchical clustering on the plurality of audio segments according to the time length and the voiceprint feature of the plurality of audio segments to obtain audio segments corresponding to the same speaker, and add the same user mark to the audio segments corresponding to the same speaker to obtain an audio signal with the user mark.

[0014] The embodiment of the present application further provides a server device, comprising: a processor and a memory; the memory is used for storing a computer program; the processor is coupled with the memory and is used for executing the computer program, so as to: receive a plurality of audio segments sent by a sound pickup device; extract a voiceprint feature corresponding to the plurality of audio segments; perform hierarchical clustering on the plurality of audio segments according to the time length and the voiceprint feature of the plurality of audio segments, so as to obtain audio segments corresponding to the same speaker; and add the same user mark to the audio segments corresponding to the same speaker, so as to obtain audio signals with the user mark added.

[0015] The embodiment of the present application further provides a computer readable storage medium storing a computer program, when the computer program is executed by a processor, the processor is caused to implement the steps in the methods provided by the embodiment of the present application.

[0016] The embodiment of the present application further provides a computer program product, comprising computer programs / instructions, when the computer programs / instructions are executed by a processor, the processor is caused to implement the steps in the methods provided by the embodiment of the present application.

[0017] In the embodiment of the present application, for the audio signals in the multi-person speaking scene, the audio signals are first cut into a plurality of audio segments based on the speaker change points, then hierarchical clustering is performed on the plurality of audio segments according to the time length and the voiceprint feature of the plurality of audio segments, the audio segments corresponding to the same speaker are identified and the user mark is added. Among them, the hierarchical clustering is performed by combining the time length and the voiceprint feature of the audio segments, instead of simply using the voiceprint feature for clustering. Compared with clustering all audio segments at the same time, the hierarchical clustering can reduce the error caused by the audio segments with unstable voiceprint features, can more accurately identify the audio segments corresponding to the same speaker, improve the identification efficiency, and the user mark result is more accurate. BRIEF DESCRIPTION OF DRAWINGS

[0018] The accompanying drawings, which are included to provide a further understanding of the present application, constitute a part of the present application, the illustrative embodiments of the present application and the description thereof serve to explain the present application, and do not constitute improper limitations on the present application. In the drawings:

[0019] Figure 1a A flowchart of an audio signal processing method provided by an exemplary embodiment of the present application is shown in the figure;

[0020] Figure 1b A flowchart of another audio signal processing method provided by an exemplary embodiment of the present application is shown in the figure;

[0021] Figure 1c A flowchart of another audio signal processing method provided by an exemplary embodiment of the present application is shown in the figure;

[0022] Figure 2a This is a schematic diagram illustrating the clustering of audio segments in each layer;

[0023] Figure 2b This is a schematic diagram illustrating the clustering of audio segments in each layer;

[0024] Figure 2c This is a schematic diagram illustrating the clustering of audio segments in the first layer;

[0025] Figure 3a This is a diagram illustrating the usage of a microphone in a multi-person conference scenario.

[0026] Figure 3b This is a schematic diagram illustrating the usage of a microphone in a business cooperation and negotiation scenario.

[0027] Figure 3c This is a schematic diagram illustrating the usage of a sound pickup device in a teaching setting.

[0028] Figure 3d A flowchart illustrating a meeting recording method provided for an exemplary embodiment of this application;

[0029] Figure 3e A flowchart illustrating a meeting record presentation method provided for an exemplary embodiment of this application;

[0030] Figure 4a A schematic diagram of the structure of an audio processing system provided for an exemplary embodiment of this application;

[0031] Figure 4b A schematic diagram of another audio processing system provided as an exemplary embodiment of this application;

[0032] Figure 5 A schematic diagram of the structure of a sound pickup device provided for an exemplary embodiment of this application;

[0033] Figure 6 This is a schematic diagram of the structure of a server device provided for an exemplary embodiment of this application. Detailed Implementation

[0034] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0035] In actual applications, in a multi-speaker speaking scenario, there can be strong noise interference, and the voiceprint features of the speakers can also change due to emotions, which can cause misjudgment of the recognition result based on the voiceprint features and low recognition accuracy. To solve this problem, in some embodiments of the present application, for an audio signal in a multi-speaker speaking scenario, the audio signal is first divided into multiple audio segments based on speaker change points, and then the multiple audio segments are clustered in multiple levels according to the lengths and voiceprint features of the multiple audio segments to identify audio segments corresponding to the same speaker and add user labels. Among them, instead of simply using voiceprint features for clustering, the lengths and voiceprint features of the audio segments are combined for hierarchical clustering. Hierarchical clustering can first cluster audio segments with more stable voiceprint features, which can reduce errors caused by audio segments with unstable voiceprint features compared to clustering all audio segments at the same time, can more accurately identify audio segments corresponding to the same speaker, improve the efficiency of recognition, and the user label result is more accurate.

[0036] The technical solutions provided by the embodiments of the present application are described in detail below with reference to the accompanying drawings.

[0037] Figure 1a A flowchart of an audio signal processing method provided by an exemplary embodiment of the present application is shown in FIG. Figure 1a As shown in the figure, the method comprises:

[0038] 101a, identifying speaker change points in an audio signal collected in a multi-speaker speaking scenario;

[0039] 102a, dividing the audio signal into multiple audio segments according to the speaker change points, and extracting voiceprint features of the multiple audio segments;

[0040] 103a, clustering the multiple audio segments in multiple levels according to the lengths and voiceprint features of the multiple audio segments to obtain audio segments corresponding to the same speaker;

[0041] 104a, adding the same user label to the audio segments corresponding to the same speaker to obtain an audio signal with user labels added.

[0042] In this embodiment, the speaker change point refers to the position point in the audio signal that distinguishes different speakers, that is, the position of the speaker change event, which can be one or more, for example, two, three or five. In this embodiment, the identification method of the speaker change point is not limited, which is illustrated below.

[0043] For example, a speaker change point in an audio signal can be identified by a voice activity detection (VAD) technique. An endpoint in VAD refers to a critical point of change between a silent signal and a valid speech signal. For an audio signal collected in a multi-person speech scene, a VAD technique can be used to find the start point and end point of each speech segment, distinguish the speech period and non-speech period, and further remove silence, noise, etc. In this embodiment, the pause duration between the start point and the end point can be used to determine the speaker change point. For example, if the pause time interval between the start point and the end point is greater than a set threshold, the positions of the speech endpoints (i.e., the start point and the end point) in this case can be regarded as the speaker change point.

[0044] For another example, the voiceprint feature of the audio signal collected in the multi-person speech scene can also be extracted, and the position point where the voiceprint changes in the audio signal can be determined as the speaker change point according to the change of the voiceprint feature in the audio signal. Alternatively, the VAD technique and the voiceprint feature can be combined, and the start point and end point corresponding to each speech period detected by the VAD are further combined with the voiceprint features at the adjacent start point and end point. If the voiceprint features at the adjacent start point and end point change, the positions of the speech endpoints (i.e., the start point and the end point) can be determined as the speaker change point.

[0045] For another example, when collecting an audio signal, the sound source location can be positioned based on a microphone array to obtain a change point of the sound source location, and the speaker change point in the audio signal can be determined according to the change point of the sound source location. For example, in a speech scene where the position of each speaker is fixed, the change point of the sound source location can be determined as the speaker change point.

[0046] Of course, in some speech scenes, the speaker can move, i.e., the position of the speaker is not fixed. For this case, the sound source positioning and the VAD technique can be combined, the change point of the sound source location is located by the sound source positioning technique, and the start point and end point corresponding to each speech period in the audio signal are determined by the VAD technique. The change point of the sound source location is corrected according to the start point and end point determined by the VAD, so as to obtain an accurate speaker change point. Specifically, the change point of the sound source location can be aligned with the VAD detection result on a time axis, and it is determined whether there is a detected speech endpoint, such as a start point or an end point, within a certain time before and after the change point of the sound source location. If so, the position of the speech endpoint can be determined as the speaker change point. By the above method, the speaker change point can be more accurately determined, and the speech recognition result can be more accurately truncated, so as to avoid the phenomenon of losing the beginning or the end of a word.

[0047] In this embodiment, the audio signal can be divided into multiple speech segments according to the speaker change points. For example, for an audio signal, the start position is recorded as A1, and the end position is recorded as A2. In the case that one speaker change point B1 is identified in the audio signal, the audio signal can be divided into audio segment A1-> B1 and audio segment B1-> A2 according to the speaker change point B1.

[0048] In this embodiment, after the multiple audio segments are divided according to the speaker change points, the voiceprint features of the multiple audio segments can be extracted. The voiceprint features can be represented by feature vectors. The voiceprint features are the characteristics of the audio segments, and the voiceprint features of the audio segments corresponding to different speakers are generally different. In this embodiment, the implementation of extracting the voiceprint features of the multiple speech segments is not limited. For example, a neural network model for extracting voiceprint features can be pre-trained, and the pre-trained neural network model can be used to extract the voiceprint features of the multiple speech segments. The neural network model can be, but is not limited to, a model based on Mel-scale Frequency Cepstral Coefficients (MFCC) or Gaussian Mixture Model-Universal Background Model (GMM-UBM).

[0049] In the embodiment, each of the plurality of audio segments segmented by the speaker change point corresponds to a speaker, and different audio segments can correspond to the same speaker or different speakers. In an application requiring adding user labels to the audio segments, it is necessary to identify audio segments corresponding to the same speaker, so as to add the same user label to the audio segments corresponding to the same speaker. In order to more accurately identify the voice segments corresponding to the same speaker, in the embodiment, the plurality of audio segments can be clustered based on the voiceprint features of the plurality of audio segments, so as to cluster the audio segments with the same voiceprint features together. In the embodiment, the audio segments with the same or similar voiceprint features are regarded as the audio segments corresponding to the same user. In addition, due to the speaking habits, manners and special needs of the speakers and other factors, there can be particularly short speeches in the multi-person speaking scene, such as "um", "ah", "yes", "good" and the like, so that there can be some shorter audio segments segmented. The longer the duration of the audio segment is, the more stable the corresponding voiceprint feature is, and vice versa. For example, the voiceprint features of "ah" spoken by user A and "ah" spoken by user B are not very obvious. In view of this, in the embodiment, the duration of the audio segment is further considered, and the plurality of audio segments are hierarchically clustered in combination with the duration of the plurality of audio segments. Hierarchical clustering refers to the process of clustering the layered audio segments layer by layer after layering the plurality of audio segments, so as to give full play to the advantages of longer audio segments and reduce the interference that can be caused by shorter audio segments. Therefore, in the embodiment, after obtaining the plurality of audio segments and the voiceprint features thereof, the plurality of audio segments are hierarchically clustered according to the duration and voiceprint features of the plurality of audio segments, so as to obtain the audio segments corresponding to the same speaker.

[0050] Since the longer the time length of an audio segment is, the more stable the corresponding voiceprint feature is, based on this, in some optional embodiments of the present application, the multiple audio segments can be layered according to the time lengths of the multiple audio segments, to obtain multiple layers of audio segments corresponding to different time length ranges; and the multiple layers of audio segments are hierarchically clustered according to the voiceprint features of the multiple audio segments, in the order from long to short time length range, to obtain at least one clustering result, each of which includes audio segments corresponding to the same speaker. In the hierarchical clustering, not only the voiceprint features are used to cluster the multiple audio segments, but also the time lengths of the audio segments are combined, the audio segments with longer time length range are clustered according to the voiceprint features first, and then the audio segments with shorter time length are clustered according to the voiceprint features, in the process of clustering the audio segments with shorter time length, it is needed to judge whether the audio segments with shorter time length belong to the result clustered by the audio segments with longer time length, and if not, a new clustering result can be established, and so on, to complete the clustering of all layers of audio segments. In this way, the hierarchical clustering in the order from long to short time length can take the clustering result of the audio segments with longer time length as the main one, reduce the recognition error caused by the unstable voiceprint features of the audio segments with shorter time length, and improve the accuracy of recognizing the audio segments corresponding to the same speaker.

[0051] In the present embodiment, after obtaining the audio segments corresponding to the same speaker, the same user mark can be added to the audio segments corresponding to the same speaker, to obtain the audio signal with added user marks. In the present embodiment, the implementation of adding the user mark is not limited. For example, a speech segment with a user mark can be inserted before each audio segment, for example, the speech segment "user C1 please speak" can be inserted before the audio segment corresponding to user C1. For another example, the same user mark point can be added to the audio segments corresponding to the same speaker on the audio track, for example, the red mark point is added to the audio segment corresponding to speaker C2, the green mark point is added to the audio segment corresponding to speaker C3, and the yellow mark point is added to the audio segment corresponding to speaker E, and so on.

[0052] In the embodiments of the present application, for the audio signal in a multi-speaker speaking scenario, the audio signal is first cut into multiple audio segments based on the speaker change points, then the multiple audio segments are hierarchically clustered according to the time lengths and voiceprint features of the multiple audio segments, to identify the audio segments corresponding to the same speaker and add the user marks. In this process, not only the voiceprint features are used for clustering, but also the time lengths and voiceprint features of the audio segments are combined for hierarchical clustering. The hierarchical clustering can cluster the audio segments with more stable voiceprint features first, compared to clustering all the audio segments at the same time, the hierarchical clustering can reduce the error caused by the audio segments with unstable voiceprint features, can more accurately identify the audio segments corresponding to the same speaker, improve the efficiency of identification, and the user mark result is more accurate.

[0053] In this embodiment, the implementation of layering the plurality of audio segments according to the time lengths of the plurality of audio segments to obtain the multi-layer audio segments corresponding to different time length ranges is not limited. In an optional embodiment, a quantity threshold of each layer can be set, the plurality of audio segments can be sorted according to the time lengths, and the sorted plurality of audio segments can be layered according to the quantity threshold of each layer set in advance to obtain the multi-layer audio segments. In another optional embodiment, a time length threshold of each layer can be set in advance, and the plurality of audio segments can be layered according to the time lengths of the plurality of audio segments and the time length threshold of each layer set in advance to obtain the multi-layer audio segments corresponding to different time length ranges; wherein the smaller the layer number is, the greater the corresponding time length threshold is, and the time length of the audio segment in each layer is greater than or equal to the time length threshold of the layer. For example, the audio segments with a time length greater than 20s can be divided into a first layer, the audio segments with a time length of 10s-20s can be divided into a second layer, the audio segments with a time length of 5s-10s can be divided into a third layer, and the audio segments with a time length less than 5s can be divided into a fourth layer.

[0054] In this embodiment, after obtaining the multi-layer audio segments, the implementation of layering the multi-layer audio segments to obtain at least one clustering result is not limited. Details are described below.

[0055] In an optional embodiment, the audio segments in each layer can be clustered according to the voiceprint features corresponding to the audio segments in each layer to obtain a clustering result of each layer; and the clustering results of adjacent two layers can be clustered in turn according to the voiceprint features of the clustering results of each layer in the order from small to large layer number to obtain at least one clustering result. The clustering result of each layer can be one or multiple, for example, 2, 3 or 5, etc., which is not limited. For example, Figure 2aAs shown, according to the time length of the plurality of audio segments, the plurality of audio segments are divided into three layers, according to the voiceprint features of the audio segments of each layer, the audio segments of each layer are clustered to obtain the clustering results of each layer, the first layer has two clustering results D1 and D2, the second layer has three clustering results D3, D4 and D5, and the third layer has two clustering results D6 and D7; then, according to the voiceprint features of the second layer clustering results, the second layer clustering results are clustered to the first layer clustering results D1 or D2, wherein, according to the voiceprint features of the clustering results D3, D4 and D5, it can be judged whether the clustering results D3, D4 and D5 can be clustered into the clustering results D1 and D2, assuming that the clustering results D3 and D4 can be clustered into the clustering results D1 to obtain the clustering results E1, and the clustering result D5 can be clustered into the clustering result D2 to obtain the clustering result E2, so that after the second layer clustering results are clustered to the first layer clustering results, two clustering results E1 and E2 can be obtained; finally, according to the voiceprint features of the third layer clustering results, the third layer clustering results are clustered to the existing clustering results E1 and E2, wherein, according to the voiceprint features of the clustering results D6 and D7, it can be judged whether the clustering results D6 and D7 can be clustered into the clustering results E1 or E2, assuming that the clustering result D6 can be clustered into the clustering result E1 to obtain the clustering result E3, and the clustering result D7 can be clustered into the clustering result E2 to obtain the clustering result E4, finally two clustering results E3 and E4 are obtained, that is, two audio segments corresponding to the speakers are obtained.

[0056] In another optional embodiment, the audio segments of the first layer are first clustered, and then based on the clustering results of the first layer, the audio segments of each layer are clustered to the existing clustering results in order from small to large. Specifically, first, for the audio segments in the first layer, the audio segments in the first layer are clustered according to the voiceprint features corresponding to the audio segments in the first layer to obtain at least one clustering result; then, for the audio segments in the non-first layer, in order from small to large according to the number of layers, the audio segments in the non-first layer are clustered to the existing clustering results according to the voiceprint features corresponding to the audio segments in the non-first layer; and if there are remaining audio segments in the non-first layer that have not been clustered into the existing clustering results, the remaining audio segments are clustered according to the voiceprint features corresponding to the remaining audio segments to generate new clustering results, until each audio segment on all layers is clustered into a clustering result. The entire hierarchical clustering process is illustrated below by taking the audio segments divided into three layers as an example.

[0057] As Figure 2bAs shown, first, according to the voiceprint features of the audio segments of the first layer, the audio segments of the first layer are clustered to obtain two clustering results F1 and F2, the clustering result F1 contains the audio segment g1 and the audio segment g2, and the clustering result F2 contains the audio segment g3; then, the audio segments of the second layer are clustered to the existing clustering results F1 and F2 of the first layer, wherein the second layer contains three audio segments, which are the audio segment g4, the audio segment g5 and the audio segment g6, then according to the voiceprint features of the audio segment g4, the audio segment g5 and the audio segment g6, it is judged whether the audio segment g4, the audio segment g5 and the audio segment g6 can be clustered into the clustering result F1 or F2, assuming that the audio segment g5 and the audio segment g6 are clustered into the clustering result F2, and the audio segment g4 cannot be clustered into the clustering results F1 and F2 of the first layer, then the audio segment g4 is separately taken as a clustering result F3, in this way, after the audio segments of the second layer are clustered to the clustering results of the first layer, three clustering results F1, F2 and F3 are obtained; finally, the clustering results of the third layer are clustered to the existing clustering results F1, F2 and F3, wherein the third layer contains two audio segments, which are the audio segment g7 and the audio segment g8; according to the voiceprint features of the audio segment g7 and the audio segment g8, it is judged whether the audio segment g7 and the audio segment g8 can be clustered into the existing clustering results F1, F2 or F3, assuming that the audio segment g7 is clustered into the clustering result F1, and the audio segment g8 is clustered into the clustering result F2; finally, three clustering results F1, F2 and F3 can be obtained, that is, the audio segments corresponding to the three speakers.

[0058] In the embodiment, the implementation of clustering the plurality of audio segments is not limited, for example, but not limited to: K-means clustering, mean shift clustering, density-based clustering (DBSCAN), expectation maximization (EM) clustering based on Gaussian mixture model (GMM), agglomerative hierarchical clustering or graph community detection clustering, etc.

[0059] In the embodiment, the implementation of clustering the audio segments in the first layer according to the voiceprint features corresponding to the audio segments in the first layer to obtain at least one clustering result is not limited. The implementation of clustering the audio segments in the first layer according to the voiceprint features corresponding to the audio segments in the first layer to obtain at least one clustering result includes: in the case that the first layer contains at least two audio segments, calculating the overall similarity between the at least two audio segments in the first layer according to the voiceprint features corresponding to the at least two audio segments in the first layer, and optionally, the voiceprint feature similarity between the at least two audio segments can be taken as the overall similarity between the at least two audio segments; dividing the at least two audio segments in the first layer into at least one clustering result according to the overall similarity between the at least two audio segments in the first layer; and calculating the clustering center of the at least one clustering result according to the voiceprint features corresponding to the audio segments contained in the at least one clustering result, and the clustering center includes a center voiceprint feature. Specifically, for any audio segment in the first layer, the overall similarity between the audio segment and other audio segments in the first layer can be calculated according to the voiceprint feature corresponding to the audio segment and the voiceprint features corresponding to the other audio segments in the first layer; if there is a target audio segment in the other audio segments in the first layer that has an overall similarity to the audio segment satisfying a set similarity condition, the audio segment and the target audio segment are clustered to obtain a target clustering result, and the clustering center of the target clustering result is updated according to the voiceprint features corresponding to the audio segment and the target audio segment. Further, it can be calculated whether the target clustering result and the remaining audio segments in the first layer can be clustered, and for the remaining audio segments that cannot be clustered into the target clustering result, the remaining audio segments can be clustered according to the voiceprint features corresponding to the remaining audio segments to generate a new clustering result, until all the audio segments on the first layer are clustered into a clustering result.

[0060] For example, in the case that three audio segments are included in the first layer, the three audio segments are audio segment h1, audio segment h2 and audio segment h3, the voiceprint feature similarity of audio segment h1 and audio segment h2 can be calculated first, the voiceprint feature similarity is taken as the overall similarity between the two audio segments h1 and h2, if the overall similarity meets the set condition, it is considered that audio segment h1 and audio segment h2 come from the same speaker, audio segment h1 and audio segment h2 can be clustered in a clustering result H1, and the clustering center of the clustering result H1, that is, the center voiceprint feature, is calculated. For example, the voiceprint feature of audio segment h1 can be directly taken as the center voiceprint feature, the voiceprint feature of audio segment h2 can also be taken as the center voiceprint feature, or the voiceprint feature of audio segment h1 and the voiceprint feature of audio segment h2 can be averaged to obtain the center voiceprint feature, which is not limited; after the clustering result H1 is obtained, the similarity of the center voiceprint feature of the clustering result H1 and the voiceprint feature of audio segment h3 can be calculated, the voiceprint feature similarity is taken as the overall similarity between the clustering result H1 and audio segment h3, if the overall similarity meets the set condition, it is considered that the clustering result H1 and audio segment h3 come from the same speaker, then the clustering result H1 and audio segment h3 can be clustered into a clustering result H2, and the center voiceprint feature of the clustering result H2 is calculated according to the voiceprint features of the clustering result H1 and audio segment h3; if the similarity threshold does not meet the set condition, it is considered that the clustering result H1 and audio segment h3 are not from the same speaker, then audio segment h3 can be taken as a clustering result H3 alone.

[0061] In the embodiment, the implementation of clustering the audio segments in the non-first layer to the existing clustering results in order according to the number of layers from small to large, according to the voiceprint features corresponding to the audio segments in the non-first layer is not limited, for example, for each audio segment in any non-first layer, the overall similarity between the audio segment and the existing clustering result is calculated according to the voiceprint feature corresponding to the audio segment and the clustering center of the existing clustering result; if there is a target clustering result in the existing clustering result which meets the set similarity condition with the overall similarity of the audio segment, the audio segment is added to the target clustering result, and the clustering center of the target clustering result is updated according to the voiceprint feature corresponding to the audio segment.

[0062] In the embodiment, the implementation of updating the cluster center of the target clustering result according to the voiceprint feature of the audio segment is not limited, in an optional embodiment, the voiceprint features of the audio segments contained in the target clustering result are directly averaged to obtain a new center voiceprint feature as the updated cluster center of the target clustering result. In another optional embodiment, the layers to which the audio segments contained in the target clustering result belong are determined, different layers are set with different weights, and the smaller the layer is, the greater the corresponding weight is; the voiceprint features corresponding to the audio segments are weighted and summed according to the weights corresponding to the layers to which the audio segments belong, to obtain a new center voiceprint feature as the updated cluster center of the target clustering result. For example, the target clustering result contains audio segment j1 and audio segment j2 of the first layer, and audio segment j3 of the second layer, when calculating the cluster center, the audio segments of the first layer are set with a weight k1, and the audio segments of the second layer are set with a weight k2, k1>k2 and k1+k2=1, then the center voiceprint feature of the target clustering result is: (the voiceprint feature of j1)*k1+(the voiceprint feature of j2)*k1+(the voiceprint feature of j3)*k2.

[0063] In the embodiment of the application, in a specific multi-person speaking scene, for example, a multi-person conference, a specific speaker can usually speak at his own seat or the like, and the position of the speaker usually does not change during the conference. Therefore, whether a speaker change event exists can be determined by identifying a sudden change in the sound source direction. Based on this, the embodiment of the application further provides an audio signal processing method, as shown in Figure 1b The method comprises the following steps.

[0064] 101b, performing sound source positioning on the audio signal collected in the multi-person speaking scene to obtain a change point of the sound source position;

[0065] 102b, dividing the audio signal into a plurality of audio segments according to the change point of the sound source position, and extracting voiceprint features of the plurality of audio segments;

[0066] 103b, performing hierarchical clustering on the plurality of audio segments according to the time length, the voiceprint features and the sound source position of the plurality of audio segments, to obtain audio segments corresponding to the same speaker;

[0067] 104b, adding the same user mark to the audio segments corresponding to the same speaker to obtain an audio signal with a user mark.

[0068] In the embodiment, the audio signal is subjected to sound source positioning to obtain a change point of sound source position; the audio signal is segmented according to the change point of sound source position to obtain a plurality of audio segments, each audio segment corresponding to a unique sound source position, for the case where the position of the speaker does not change, it can be considered that each sound source position corresponds to a speaker, that is, each audio segment corresponds to a speaker, for the case where the position of the speaker changes, the audio segment of the speaker can be divided into two audio segments, one audio segment corresponds to before the position change, and one audio segment corresponds to after the position change, at this time, each audio segment also corresponds to a speaker.

[0069] In the embodiment, after the audio segments are segmented into a plurality of audio segments according to the change point of sound source position, the voiceprint features of the plurality of audio segments can be extracted, for the implementation of extracting voiceprint features, refer to the foregoing embodiments, which will not be repeated here.

[0070] In the embodiment, each of the plurality of audio segments segmented by the change point of sound source position corresponds to a speaker, different audio segments can correspond to the same speaker or different speakers. In the application of adding user labels to audio segments, it is necessary to identify audio segments corresponding to the same speaker, so as to add the same user label to the audio segments corresponding to the same speaker. In order to more accurately identify the voice segments corresponding to the same speaker, in the embodiment, the plurality of audio segments can be clustered based on the voiceprint features of the plurality of audio segments, and the audio segments with the same voiceprint features are clustered together as much as possible. In the embodiment, the audio segments with the same or similar voiceprint features are regarded as the audio segments corresponding to the same user. Further, the sound source position corresponding to the audio segment can also be combined, if the voiceprint features of two audio segments are the same or similar and come from the same sound source position, the probability that the two audio segments correspond to the same user will be higher. In addition, considering that the longer the length of the audio segment is, the more stable the corresponding voiceprint feature is, and vice versa, the shorter the length of the audio segment is, the lower the stability of the corresponding voiceprint feature is, and the less obvious the distinguishability is. Therefore, in the embodiment, the length of the audio segment is further considered, and the plurality of audio segments are hierarchically clustered in combination with the lengths of the plurality of audio segments, the hierarchical clustering refers to the process of clustering the layered audio segments layer by layer after the plurality of audio segments are layered, so as to fully exert the advantage of the longer audio segment and reduce the interference possibly caused by the shorter audio segment. Therefore, in the embodiment, after obtaining the plurality of audio segments and the voiceprint features and sound source positions thereof, the plurality of audio segments are hierarchically clustered according to the lengths, voiceprint features and sound source positions of the plurality of audio segments, to obtain the audio segments corresponding to the same speaker.

[0071] In an optional embodiment of the present application, an implementation of hierarchical clustering of a plurality of audio segments according to time length, voiceprint features and sound source positions of the plurality of audio segments comprises: hierarchically clustering the plurality of audio segments according to time length of the plurality of audio segments to obtain a plurality of layers of audio segments corresponding to different time length ranges; and hierarchically clustering the plurality of layers of audio segments according to voiceprint features and sound source positions of the plurality of audio segments in order of time length ranges from long to short to obtain at least one clustering result, each of which includes audio segments corresponding to the same speaker.

[0072] The implementation of hierarchically clustering the plurality of audio segments according to time length of the plurality of audio segments to obtain a plurality of layers of audio segments corresponding to different time length ranges can refer to the foregoing embodiments and will not be described here. In the present embodiment, the implementation of hierarchically clustering the plurality of layers of audio segments according to voiceprint features and sound source positions of the plurality of audio segments in order of time length ranges from long to short to obtain at least one clustering result is not limited. The following is an example.

[0073] In an optional embodiment, the audio segments in each layer can be clustered according to voiceprint features and sound source positions of the audio segments in each layer to obtain a clustering result of each layer; and the clustering results of adjacent two layers can be clustered in order of layer number from small to large according to voiceprint features and sound source positions of the clustering results of each layer to obtain at least one clustering result.

[0074] In another optional embodiment, the audio segments in the first layer can be clustered first, and then the audio segments in each layer can be clustered into the existing clustering results in order of layer number from small to large based on the clustering result of the first layer. Specifically, first, the audio segments in the first layer can be clustered according to voiceprint features and sound source positions of the audio segments in the first layer to obtain at least one clustering result; then, the audio segments in the non-first layer can be clustered into the existing clustering results in order of layer number from small to large according to voiceprint features and sound source positions of the audio segments in the non-first layer; and if there are remaining audio segments in the non-first layer that have not been clustered into the existing clustering results, the remaining audio segments can be clustered according to voiceprint features and sound source positions of the remaining audio segments to generate a new clustering result, until each audio segment on all layers is clustered into at least one clustering result.

[0075] In an optional embodiment of the present application, the implementation of clustering the audio segments in the first layer according to the voiceprint features and the sound source positions corresponding to the audio segments in the first layer includes: if the first layer only includes one audio segment, the audio segment itself forms a clustering result; if the first layer includes at least two audio segments, in the case that the first layer includes at least two audio segments, the overall similarity between the at least two audio segments in the first layer is calculated according to the voiceprint features and the sound source positions corresponding to the at least two audio segments in the first layer. For example, the voiceprint feature similarity between the at least two audio segments can be calculated first, and then the sound source position similarity between the at least two audio segments is calculated, the voiceprint feature similarity and the sound source position similarity are weighted to obtain the overall similarity between the at least two audio segments in the first layer. Further, the at least two audio segments in the first layer can be divided into at least one clustering result according to the overall similarity between the at least two audio segments in the first layer. Further, the clustering center of the at least one clustering result needs to be calculated according to the voiceprint features and the sound source positions corresponding to the audio segments included in the at least one clustering result, the clustering center includes the center voiceprint feature and the center sound source position, and provides a basis for clustering the audio segments other than the first layer to the at least one clustering result. For example, for each clustering result, the average of the voiceprint features corresponding to the audio segments included in the clustering result can be taken as the center voiceprint feature of the clustering result, and the average of the sound source positions corresponding to the audio segments included in the clustering result can be taken as the center sound source position of the clustering result. For another example, the voiceprint feature of any audio segment included in the clustering result can be directly taken as the center voiceprint feature of the clustering result, and the sound source position of any audio segment included in the clustering result can be directly taken as the center sound source position of the clustering result.

[0076] As Figure 2cAs shown, the audio segments of the first layer include: audio segment ml-audio segment m6. The overall similarity between any two audio segments can be calculated, and two audio segments with an overall similarity higher than a set similarity threshold (e.g., 90%) can be clustered. For example, the overall similarity threshold between audio segment ml and audio segment m3 is 91%, the overall similarity threshold between audio segment m2 and audio segment m4 is 93%, and the overall similarity threshold between audio segment m3 and audio segment m6 is 95%. Thus, audio segment ml and audio segment m3 can be clustered to obtain clustering result Ml, audio segment m2 and audio segment m4 can be clustered to obtain clustering result M2, and audio segment m3 and audio segment m6 can be clustered to obtain clustering result M3. The clustering centers of clustering result Ml, clustering result M2, and clustering result M3 are calculated respectively. The overall similarity between two clustering results is calculated according to the clustering centers of the two clustering results. If the overall similarity exceeds a set threshold (e.g., 90%), the two clustering results will be further clustered. For example, the overall similarity between clustering result Ml and clustering result M2 is 90%, the overall similarity between clustering result Ml and clustering result M3 is 85%, and the overall similarity between clustering result M2 and clustering result M3 is 80%. Thus, clustering result Ml and clustering result M2 are further clustered to obtain clustering result M4, and clustering result M3 is taken as a separate clustering result. Finally, the audio segments of the first layer obtain two clustering results M3 and M4.

[0077] Further optionally, for any audio segment in a non-first layer, a process of clustering the audio segment into an existing clustering result includes: calculating the overall similarity between the audio segment and the existing clustering result according to the voiceprint feature and the sound source position corresponding to the audio segment and the clustering center of the existing clustering result. If there is a target clustering result in the existing clustering result with an overall similarity to the audio segment satisfying a set similarity condition, it can be considered that the audio segment and the audio segments in the target clustering result come from the same speaker. The audio segment is added to the target clustering result, and the clustering center of the target clustering result is updated according to the voiceprint feature and the sound source position corresponding to the audio segment.

[0078] For the target clustering result, when a new audio segment is added to the target clustering result, the clustering center of the target clustering result can be updated in the following manner, but not limited thereto. For example, the voiceprint features of all audio segments included in the target clustering result can be averaged, and the average value can be taken as the center voiceprint feature of the clustering center of the target clustering result. The sound source positions of the audio segments in the target clustering result can be averaged, and the average value can be taken as the center sound source position of the clustering center of the target clustering result. For another example, the layers to which each audio segment included in the target clustering result belongs can be determined, different weights can be set for different layers, and the smaller the layer number is, the greater the corresponding weight is. The corresponding voiceprint features of each audio segment included in the target clustering result can be weighted and summed according to the weights corresponding to the layers to which the audio segments belong, to obtain a new center voiceprint feature. The corresponding sound source positions of each audio segment included in the target clustering result can be weighted and summed according to the weights corresponding to the layers to which the audio segments belong, to obtain a new center sound source position. The new center voiceprint feature and the new center sound source position form the updated clustering center of the target clustering result.

[0079] In the embodiments of the present application, for the audio signal in a multi-person speaking scene, the audio signal is first cut into multiple audio segments based on the sound source position, and then the multiple audio segments are clustered in multiple levels according to the time length, voiceprint feature and sound source position of the multiple audio segments, to identify the audio segments corresponding to the same speaker and add user labels. Among them, instead of simply using voiceprint features for clustering, the sound source position, voiceprint feature and hierarchical aggregation are combined, wherein the sound source position can accurately segment the audio signal, and the hierarchical aggregation can reduce the influence of short voice on the recognition result. On this basis, the voiceprint feature is used to identify the audio segments corresponding to the same speaker, which can greatly improve the recognition efficiency and make the user label result more accurate.

[0080] The embodiments also provide an audio signal processing method, as shown in Figure 1c The method comprises:

[0081] 101c, performing sound source positioning on the audio signal collected in the multi-person speaking scene to obtain a change point of the sound source position;

[0082] 102c, cutting the audio signal into multiple audio segments according to the change point of the sound source position, and extracting voiceprint features of the multiple audio segments;

[0083] 103c, clustering the multiple audio segments according to the voiceprint features and the sound source positions of the multiple audio segments to obtain audio segments corresponding to the same speaker;

[0084] 104c, adding the same user label to the audio segments corresponding to the same speaker to obtain an audio signal with user labels added.

[0085] The dividing the audio signal into a plurality of audio segments according to the change point of the sound source position comprises: taking the change point of the sound source position as a speaker change point, thereby dividing the audio signal into a plurality of audio segments; or, in combination with a VAD technology, detecting a start point and an end point of the audio signal by using the VAD technology; correcting the change point of the sound source position according to the start point and the end point to obtain a speaker change point, and then dividing the audio signal into a plurality of audio segments according to the speaker change point.

[0086] In an optional embodiment, the clustering the plurality of audio segments according to the voiceprint features and the sound source positions of the plurality of audio segments to obtain audio segments corresponding to a same speaker comprises: layering the plurality of audio segments according to time lengths of the plurality of audio segments to obtain a plurality of audio segments corresponding to different time length ranges; and performing hierarchical clustering on the plurality of audio segments according to the voiceprint features and the sound source positions of the plurality of audio segments in an order from long to short time length range to obtain at least one clustering result, each of which includes audio segments corresponding to a same speaker.

[0087] In the embodiments of the present application, for an audio signal in a multi-speaker speaking scenario, the audio signal is first divided into a plurality of audio segments based on speaker change points, and then hierarchical clustering is performed on the plurality of audio segments according to time lengths and voiceprint features of the plurality of audio segments to identify audio segments corresponding to a same speaker and add user labels. In this embodiment, the hierarchical clustering is performed on the plurality of audio segments according to the time lengths and the voiceprint features of the plurality of audio segments, instead of only using the voiceprint features for clustering. The hierarchical clustering can first cluster audio segments with more stable voiceprint features, which can reduce errors caused by audio segments with unstable voiceprint features compared to clustering all audio segments at the same time, and can more accurately identify audio segments corresponding to a same speaker, improve the efficiency of identification, and make the user label result more accurate.

[0088] The audio signal processing method provided by the embodiments of the present application can be applied to various multi-speaker speaking scenarios, such as multi-speaker conference scenarios, business negotiation scenarios, or teaching scenarios. In these application scenarios, the sound pickup device of the embodiments of the present application is deployed in these scenarios to collect audio signals in the multi-speaker speaking scenarios and implement other functions described in the above method embodiments and the following system embodiments of the present application. In order to have a better collection effect and facilitate sound source positioning of the audio signals, the placement position of the sound pickup device can be reasonably determined according to the specific deployment of the multi-speaker speaking scenario. For example, as shown in FIG. 1, in a multi-speaker conference scenario, the sound pickup device is deployed in the center of the conference table, and a plurality of speakers are distributed at different positions of the sound pickup device to facilitate picking up the speech of each speaker. Figure 3a As shown in FIG. 2, in a business negotiation scenario, the sound pickup device is deployed on a table, and a plurality of speakers are distributed at different positions of the sound pickup device to facilitate picking up the speech of each speaker. Figure 3bAs shown, in the business cooperation negotiation scene, the first business party and the second business party are opposite to each other, and the conference organizer is located between the first business party and the second business party, responsible for organizing the negotiation between the two parties. The pickup device is deployed at the central position of the conference organizer, the first business party and the second business party, and the first business party, the second business party and the conference organizer are in different positions of the pickup device, which facilitates the pickup of the pickup device. As shown Figure 3c As shown, in the teaching scene, the pickup device is deployed on the teaching table, and the teacher and the student are located in different positions of the pickup device, which facilitates the pickup of the voice of the teacher and the student.

[0089] The application of the above-mentioned audio signal processing method in the multi-person conference scene is taken as an example. The conference recording can be performed for the multi-person conference scene, and the conference recording can be further presented or reproduced. As shown Figure 3d As shown, the conference recording method provided by the example embodiment of the present application comprises the following steps:

[0090] 301d, collecting an audio signal in a multi-person conference scene, and identifying a speaker change point in the audio signal;

[0091] 302d, according to the speaker change point, the audio signal is divided into a plurality of audio segments, and the voiceprint features of the plurality of audio segments are extracted;

[0092] 303d, according to the time length and voiceprint features of the plurality of audio segments, the plurality of audio segments are hierarchically clustered to obtain audio segments corresponding to the same speaker;

[0093] 304d, the audio segments corresponding to the same speaker are added with the same user mark to obtain an audio signal with user mark added;

[0094] 305d, generating conference recording information according to the audio signal with user mark added, the conference recording information comprising a conference identifier.

[0095] The detailed description of steps 301d-304d can be referred to the foregoing embodiments, which will not be repeated here. In this embodiment, the focus is on step 305d. Specifically, after obtaining the audio signal with user mark added, the conference record information can be generated according to the audio signal with user mark added, and the corresponding conference identifier is added to the conference record information, which has uniqueness and can uniquely identify a multi-person conference. In an optional embodiment, the audio signal with user mark added can be directly used as the conference record information. In another optional embodiment, the audio signal with user mark added can be converted into text information with speaker information, which can include the content in the format similar to but not limited to the following: A speaker: xxxx; B speaker: yyy; and the like, and then the text information with speaker information is used as the conference record information. Regardless of the form of conference record information, the conference scene can be reproduced based on the conference record information, facilitating the query or review of the conference content.

[0096] Figure 3e A flowchart of a conference record presentation method provided by an exemplary embodiment of the present application is shown in FIG. 3, which includes the following steps. Figure 3e

[0097] 301e, receiving a conference review request, the conference review request including a conference identifier to be presented;

[0098] 302e, obtaining conference record information to be presented according to the conference identifier;

[0099] 303e, presenting the conference record information, which is generated according to the audio signal with user mark added in the multi-person conference scene; wherein the audio segments corresponding to the same speaker are added with the same user mark, and the audio segments corresponding to the same speaker are obtained by hierarchical clustering of the audio segments according to the time length and voiceprint features of the audio segments.

[0100] In this embodiment, for the multi-person conference scene, the conference record can be made, and the conference record process is as follows: collecting the audio signal in the multi-person conference scene, identifying the speaker change point in the audio signal; dividing the audio signal into a plurality of audio segments according to the speaker change point in the audio signal, and extracting the voiceprint features of the plurality of audio segments; then, hierarchical clustering of the plurality of audio segments is performed according to the time length and voiceprint features of the audio segments, to obtain the audio segments corresponding to the same speaker; the same user mark is added to the audio segments corresponding to the same speaker, to obtain the audio signal with user mark added; the conference record information is generated according to the audio signal with user mark added, and the corresponding conference identifier is added to the conference record information. The related process of the conference record can be referred to the foregoing embodiments, which will not be repeated here.​

[0101] After obtaining the conference record information, the relevant conference content can be consulted through the conference record information, and then the conference consultation service is provided to the outside. Based on this, a conference consultation request can be received, and the conference identifier to be presented is carried in the request. Based on the conference identifier and the conference identifier in each conference record information, the conference record information to be presented can be obtained therefrom, and the conference record information is presented. Optionally, if the conference record information is an audio signal added with a user mark, the audio signal added with the user mark can be played through a player, or the audio signal added with the user mark can be converted into text information and then displayed. If the conference record information is text information with speaker information converted from the audio signal added with the user mark, the text information with the speaker information can be displayed through a display, or the text information with the speaker information can be played through a player. In this way, the query or consultation demand of the conference content can be met.

[0102] In addition, it should be noted that, since the conference record information embodies the speaker information or the corresponding user mark, when the conference record is consulted, the conference content corresponding to a certain speaker can be consulted or played back alone, instead of the information of multiple speakers being mixed together, thereby improving the identification degree of the conference speaker and the conference content. For example, in addition to the conference identifier to be presented, the speaker information or the user mark can also be included in the conference consultation request, and the speaker information and the user mark have a corresponding relationship. In this way, according to the conference identifier, the conference record information to be presented can be obtained; according to the speaker information or the user mark, the part of the conference content corresponding to the speaker information or the user mark in the conference record information can be obtained, and the part of the conference content corresponding to the speaker information or the user mark is presented.

[0103] It should be noted that the method provided in the embodiments of the present application can be completely completed by the sound pickup device, or a part of the function can be implemented on the server device, and no limitation is made to this. The sound pickup device can be implemented as a sound recorder, a sound bar, a sound recorder or a sound pickup device, or can be implemented as a terminal device or an audio and video conference device with a sound recording function. Based on this, the present embodiment provides an audio processing system, which describes the process of implementing the audio signal processing method based on the sound pickup device and the server device. As shown in the Figure 4a The audio processing system 400 includes a sound pickup device 401 and a server device 402. The audio processing system 400 can be applied to a multi-person speaking scene, for example Figure 3a a multi-person conference scene as shown in Figure 3b a business cooperation negotiation scene as shown in Figure 3c a teaching scene as shown in and the like. In these scenes, the sound pickup device 401 can cooperate with the server device 402 to implement the above-mentioned method embodiments of the present application, and theFigure 3a to Figure 3c The server device 402 is not shown in the multi-person speaking scenario.

[0104] The pickup device 401 of the embodiment has function modules such as a power-on button, an adjustment button, a microphone array, and a speaker, and further optionally includes a display screen. The pickup device 401 can implement functions such as automatic recording, MP3 playing, FM frequency modulation, digital camera, telephone recording, timing recording, external transcription, a repeating machine, or editing. As shown in the figure, the pickup device 401 can collect an audio signal in a multi-person speaking scenario, identify a speaker change point in the audio signal, cut the audio signal into a plurality of audio segments according to the speaker change point, extract a voiceprint feature corresponding to the plurality of audio segments, and send the plurality of audio segments and the voiceprint feature corresponding thereto to the server device 402. Figure 4a

[0105] In the embodiment, the server device 402 can receive the plurality of audio segments and the voiceprint feature corresponding thereto sent by the pickup device 401, perform hierarchical clustering on the plurality of audio segments according to the time length and the voiceprint feature of the plurality of audio segments, to obtain audio segments corresponding to a same speaker, and add a same user mark to the audio segments corresponding to the same speaker, to obtain an audio signal with the user mark added.

[0106] In the embodiment, the pickup device 401 can pick up an audio signal in a multi-person speaking scenario by using a microphone array. Based on the intensity of a same sound signal picked up by microphones at different positions in the microphone array, the sound source position of the sound signal can be calculated. Based on this, in an optional embodiment of the present application, the pickup device 401 can perform sound source positioning on the audio signal to obtain a change point of the sound source position when identifying a speaker change point in the audio signal. According to the change point of the sound source position, the speaker change point in the audio signal is determined, and further, a plurality of audio segments can be cut according to the speaker change point, each audio corresponding to a unique sound source position. Accordingly, when the server device 402 performs hierarchical clustering on the plurality of audio segments according to the time length and the voiceprint feature of the plurality of audio segments, to obtain audio segments corresponding to a same speaker, the server device 402 can perform hierarchical clustering on the plurality of audio segments according to the time length, the voiceprint feature, and the sound source position of the plurality of audio segments, to obtain the audio segments corresponding to the same speaker.

[0107] In the embodiment, as shown in the figure, an audio processing system is also provided, Figure 4b Figure 4b The difference between the embodiment shown in the figure and the embodiment shown in the figure is that, in the embodiment shown in the figure, the process of extracting the voiceprint feature corresponding to the plurality of audio segments is implemented on the pickup device 401, while in the embodiment shown in the figure, the process of extracting the voiceprint feature corresponding to the plurality of audio segments is implemented on the server device 402. Figure 4a Figure 4a Figure 4b ​​​​In this embodiment, the process of extracting the voiceprint features corresponding to the plurality of audio segments is implemented on the server device 402, and other contents Figure 4a are the same or similar to those shown in FIG. 2, and details can be referred to the foregoing embodiments, which will not be described herein. Figure 4b

[0108] In this embodiment, after the server device 402 adds the same user mark to the audio segments corresponding to the same speaker, the server device 402 can store the audio signals with the added user mark for subsequent query and use. In an optional embodiment, as shown in FIG. 3, the audio processing system further includes a transcription device 403. The server device 402 can send the audio signals with the added user mark to the transcription device 403. The transcription device 403 receives the audio signals with the added user mark, converts the audio signals with the added user mark into text information with the user mark, and returns the text information with the user mark to the server device 402 or stores the text information with the user mark in a database 406. Further, as shown in FIG. 3, the audio processing system further includes a query terminal 404. The query terminal 404 can send a first query request to the server device 402. The first query request includes a user mark to be queried. The server device 402 receives the first query request, obtains text information corresponding to the user mark to be queried from the text information with the user mark, and returns the text information to the query terminal 404. Figure 4a Figure 4a In another optional embodiment, as shown in FIG. 4, after the server device 402 generates the audio signals with the user mark, the server device 402 can output the audio signals with the user mark to an upper application on the server device 402. For example, the upper application can be a remote conference application or a social application, etc. The upper application can obtain user information in a multi-person speaking scene, for example, identification information of the user, such as name, nickname, or voiceprint feature, etc. The upper application can associate the user information with the audio signals with the user mark. The association manner of the user information and the audio signals with the user mark is not limited. For example, the upper application stores a corresponding relationship between the user mark and the user information. Based on the corresponding relationship, the user information corresponding to the user mark can be found, and the user information is associated with the audio signals with the user mark.

[0109] Further, as shown in FIG. 4, the query terminal 404 can send a second query request to the server device 402. The second query request includes an audio segment to be queried. The server device 402 receives the second query request, extracts a user mark corresponding to the audio segment to be queried from the audio signals with the added user mark, and returns the user mark and / or user information corresponding to the user mark to the query terminal 404. Figure 4a

[0110] Further, as shown in FIG. 4, the query terminal 404 can send a second query request to the server device 402. The second query request includes an audio segment to be queried. The server device 402 receives the second query request, extracts a user mark corresponding to the audio segment to be queried from the audio signals with the added user mark, and returns the user mark and / or user information corresponding to the user mark to the query terminal 404. Figure 4a

[0111] ​​​​In yet another alternative embodiment, such as Figure 4b As shown, the audio processing system also includes a playback device 405. The playback device 405 can send an audio signal acquisition request to the server device 402. Based on the request, the server device 402 can output the audio signal with added user tags to the playback device 405. The playback device 405 receives and plays the audio signal with added user tags.

[0112] It should be noted that the execution subject of each step of the method provided in the above embodiments can be the same device, or the method can be executed by different devices. For example, the execution subject of steps 101a to 103a can be device A; or the execution subject of steps 101a and 102a can be device A, and the execution subject of step 103a can be device B; and so on.

[0113] Furthermore, in some processes described in the above embodiments and accompanying drawings, multiple operations appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order they appear herein, or they may be executed in parallel. The operation numbers, such as 101a, 102a, etc., are merely used to distinguish different operations and do not represent any execution order. Additionally, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel. It should be noted that the descriptions such as "first" and "second" in this document are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to different types.

[0114] Figure 5 This is a schematic diagram of the structure of a sound pickup device provided for an exemplary embodiment of this application. Figure 5 As shown, the pickup device includes a processor 55 and a memory 54.

[0115] Memory 54 is used to store computer programs and can be configured to store various other data to support operation on the pickup device. Examples of this data include instructions for any application or method used to operate on the pickup device, contact data, phonebook data, messages, pictures, videos, etc.

[0116] The memory 54 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk.

[0117] The processor 55, coupled with the memory 54, is configured to execute a computer program in the memory 54 to identify a speaker change point in an audio signal collected in a multi-person speaking scenario, segment the audio signal into a plurality of audio segments according to the speaker change point, extract a voiceprint feature of each of the plurality of audio segments, perform hierarchical clustering on the plurality of audio segments according to the time length and the voiceprint feature of each of the plurality of audio segments to obtain audio segments corresponding to a same speaker, and add a same user label to the audio segments corresponding to the same speaker to obtain an audio signal with the user label added.

[0118] The above process can be completed on the sound pickup device, or part of the functions can be performed on a server device, for example, the extraction of the voiceprint feature of each of the plurality of audio segments, the hierarchical clustering of the plurality of audio segments according to the time length and the voiceprint feature of each of the plurality of audio segments to obtain audio segments corresponding to a same speaker, and the addition of a same user label to the audio segments corresponding to the same speaker to obtain an audio signal with the user label added.

[0119] In an optional embodiment, when the processor 55 performs hierarchical clustering on the plurality of audio segments according to the time length and the voiceprint feature of each of the plurality of audio segments to obtain audio segments corresponding to a same speaker, the processor 55 is specifically configured to: segment the plurality of audio segments according to the time length to obtain a plurality of layers of audio segments corresponding to different time length ranges; and perform hierarchical clustering on the plurality of layers of audio segments according to the voiceprint feature of each of the plurality of audio segments in a time length range order from long to short to obtain at least one clustering result, each of which includes audio segments corresponding to a same speaker.

[0120] In an optional embodiment, when the processor 55 segments the plurality of audio segments according to the time length to obtain a plurality of layers of audio segments corresponding to different time length ranges, the processor 55 is specifically configured to: segment the plurality of audio segments according to the time length and a pre-set time length threshold of each layer to obtain a plurality of layers of audio segments corresponding to different time length ranges; and wherein the smaller the number of layers is, the larger the corresponding time length threshold is, and the time length of each audio segment in each layer is greater than or equal to the time length threshold of the layer.

[0121] In an optional embodiment, when the processor 55 performs hierarchical clustering on the multi-layer audio segments according to the voiceprint features corresponding to the plurality of audio segments and in the order of time length range from long to short to obtain at least one clustering result, the processor 55 is specifically configured to: for the audio segments in the first layer, perform clustering on the audio segments in the first layer according to the voiceprint features corresponding to the audio segments in the first layer to obtain at least one clustering result; for the audio segments in the non-first layer, in the order of layer number from small to large, sequentially perform clustering on the audio segments in the non-first layer to the existing clustering results according to the voiceprint features corresponding to the audio segments in the non-first layer; and if there are remaining audio segments in the non-first layer that are not clustered into the existing clustering results, perform clustering on the remaining audio segments according to the voiceprint features corresponding to the remaining audio segments to generate a new clustering result, until each audio segment on all layers is clustered into a clustering result.

[0122] In an optional embodiment, when the processor 55 identifies the speaker change point in the audio signal collected in the multi-person speaking scenario, the processor 55 is specifically configured to: perform sound source positioning on the audio signal to obtain a change point of the sound source position; determine the speaker change point in the audio signal according to the change point of the sound source position; and each audio segment divided by the speaker change point corresponds to a unique sound source position.

[0123] In an optional embodiment, when the processor 55 performs hierarchical clustering on the multi-layer audio segments according to the voiceprint features corresponding to the plurality of audio segments and in the order of time length range from long to short to obtain at least one clustering result, the processor 55 is specifically configured to: perform hierarchical clustering on the multi-layer audio segments according to the voiceprint features corresponding to the plurality of audio segments and the sound source positions and in the order of time length range from long to short to obtain at least one clustering result, each clustering result including audio segments corresponding to the same speaker.

[0124] In an optional embodiment, when the processor 55 performs hierarchical clustering on the multi-layer audio segments according to the voiceprint features corresponding to the plurality of audio segments and the sound source positions and in the order of time length range from long to short to obtain at least one clustering result, the processor 55 is specifically configured to: for the audio segments in the first layer, perform clustering on the audio segments in the first layer according to the voiceprint features corresponding to the audio segments in the first layer and the sound source positions to obtain at least one clustering result; for the audio segments in the non-first layer, in the order of layer number from small to large, sequentially perform clustering on the audio segments in the non-first layer to the existing clustering results according to the voiceprint features corresponding to the audio segments in the non-first layer and the sound source positions; and if there are remaining audio segments in the non-first layer that are not clustered into the existing clustering results, perform clustering on the remaining audio segments according to the voiceprint features corresponding to the remaining audio segments and the sound source positions to generate a new clustering result, until each audio segment on all layers is clustered into a clustering result.

[0125] In an optional embodiment, for the audio segments in the first layer, when the processor 55 clusters the audio segments in the first layer according to the voiceprint features and the sound source positions corresponding to the audio segments in the first layer to obtain at least one clustering result, specifically for: in the case that the first layer contains at least two audio segments, calculating the overall similarity between the at least two audio segments in the first layer according to the voiceprint features and the sound source positions corresponding to the at least two audio segments in the first layer; dividing the at least two audio segments in the first layer into at least one clustering result according to the overall similarity between the at least two audio segments in the first layer; and respectively calculating the clustering centers of the at least one clustering result according to the voiceprint features and the sound source positions corresponding to the audio segments contained in the at least one clustering result, the clustering center including a center voiceprint feature and a center sound source position.

[0126] In an optional embodiment, for the audio segments in the non-first layer, when the processor 55 clusters the audio segments in the non-first layer to the existing clustering results in order according to the voiceprint features and the sound source positions corresponding to the audio segments in the non-first layer in the order of the layer number from small to large, specifically for: for each audio segment in any non-first layer, calculating the overall similarity between the audio segment and the existing clustering results according to the voiceprint feature and the sound source position corresponding to the audio segment and the clustering centers of the existing clustering results; if there is a target clustering result in the existing clustering results that meets the set similarity condition with the overall similarity of the audio segment, adding the audio segment to the target clustering result, and updating the clustering center of the target clustering result according to the voiceprint feature and the sound source position corresponding to the audio segment.

[0127] In an optional embodiment, when the processor 55 updates the clustering center of the target clustering result according to the voiceprint feature and the sound source position corresponding to the audio segment, specifically for: determining the layer number to which each audio segment contained in the target clustering result belongs, wherein different layer numbers correspond to different weights, and the smaller the layer number, the greater the corresponding weight; and respectively performing weighted summation on the voiceprint features and the sound source positions corresponding to each audio segment according to the weights corresponding to the layer number to which each audio segment belongs, to obtain a new center voiceprint feature and a new center sound source position as the updated clustering center of the target clustering result.

[0128] In an optional embodiment, when the processor 55 performs hierarchical clustering on the multi-layer audio segments according to the voiceprint features and the sound source positions corresponding to the plurality of audio segments in the order of the time length range from long to short to obtain at least one clustering result, specifically for: clustering the audio segments in each layer according to the voiceprint features corresponding to the audio segments in each layer to obtain the clustering result of each layer; and in the order of the layer number from small to large, clustering the clustering results of adjacent two layers according to the voiceprint features of the clustering results of each layer to obtain at least one clustering result.

[0129] In an optional embodiment, the processor 55 is further configured to output the audio signal with the added user mark to a transcription device, so that the transcription device converts the audio signal with the added user mark into text information with the user mark; or output the audio signal with the added user mark to a playback device, so that the playback device plays the audio signal with the added user mark; or output the audio signal with the added user mark to an upper-layer application, so that the upper-layer application obtains user information corresponding to the user mark and associates the user information with the audio segment with the user mark.

[0130] In an optional embodiment, the processor 55 is further configured to receive a first query request including a user mark to be queried, obtain text information corresponding to the user mark to be queried from the text information with the user mark, and return the text information to a query end that initiates the first query request; or receive a second query request including an audio segment to be queried, extract a user mark corresponding to the audio segment to be queried from the audio signal with the added user mark, and return the user mark and / or user information corresponding to the user mark to a query end that initiates the second query request.

[0131] Detailed descriptions of the operations can refer to the descriptions in the foregoing method embodiments, which will not be repeated here.

[0132] Further, as shown in Figure 5 , the sound pickup device further includes a communication component 56, a display 57, a power supply component 58, an audio component 59, and other components. Figure 5 Some components are only schematically shown in the sound pickup device, and it does not mean that the sound pickup device only includes the components shown in Figure 5 .

[0133] Correspondingly, the embodiments of the present application further provide a computer readable storage medium storing a computer program, when the computer program is executed by a processor, the processor can implement the steps in the method embodiments shown in Figure 1a and Figure 1b which can be executed by the sound pickup device.

[0134] Correspondingly, the embodiments of the present application further provide a computer program product including a computer program / instruction, when the computer program / instruction is executed by a processor, the processor can implement the steps in the method embodiments shown in Figure 1a and Figure 1b which can be executed by the sound pickup device.

[0135] The embodiments of the present application further provide a sound pickup device, and the implementation structure of the sound pickup device is the same as or similar to the implementation structure of the sound pickup device shown in Figure 5 , and the implementation structure of the sound pickup device can refer to the structure implementation of the sound pickup device shown in Figure 5 . The sound pickup device provided in the embodiments can refer to the sound pickup device shown in Figure 5The difference between the pickup devices in the illustrated embodiments mainly lies in different functions implemented by the processors executing the computer programs stored in the memories. For the pickup device provided in the embodiments, the processor thereof executes the computer program stored in the memory, which can be used for: performing sound source positioning on the audio signals collected in a multi-person speaking scenario to obtain a change point of the sound source position; dividing the audio signals into a plurality of audio segments according to the change point of the sound source position, and extracting voiceprint features of the plurality of audio segments; clustering the plurality of audio segments according to the voiceprint features of the plurality of audio segments and the sound source position to obtain audio segments corresponding to the same speaker; and adding the same user mark to the audio segments corresponding to the same speaker to obtain audio signals with the user mark added. For detailed descriptions of the operations, refer to the descriptions in the foregoing method embodiments, which will not be repeated here.

[0136] The foregoing process can be all completed on the pickup device, or part of the functions can be placed on the server device to be executed, for example, the process of extracting the voiceprint features of the plurality of audio segments, clustering the plurality of audio segments according to the voiceprint features of the plurality of audio segments and the sound source position to obtain audio segments corresponding to the same speaker, and adding the same user mark to the audio segments corresponding to the same speaker to obtain audio signals with the user mark added, which can be completed by the server device.

[0137] Correspondingly, the embodiments of the present application also provide a computer readable storage medium storing a computer program, which, when executed by a processor, causes the processor to be able to implement Figure 1c The steps in the method embodiments illustrated above can be executed by the pickup device.

[0138] Correspondingly, the embodiments of the present application also provide a computer program product, which includes computer programs / instructions, which, when executed by a processor, causes the processor to be able to implement Figure 1c The steps in the method embodiments illustrated above can be executed by the pickup device.

[0139] Figure 6 A structural schematic diagram of a server device provided for the exemplary embodiments of the present application is shown in FIG. 6. As shown in the figure, the server device includes a processor 65 and a memory 64. Figure 6

[0140] The memory 64 is used for storing computer programs, and can be configured to store other various data to support operations on the server device. Examples of the data include instructions of any application program or method used for operating on the server device, contact data, phonebook data, messages, pictures, videos, etc.

[0141] ​The memory 64 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk.

[0142] Processor 65, coupled to memory 64, executes a computer program in memory 64 for: receiving multiple audio segments and their corresponding voiceprint features sent by the audio device; performing hierarchical clustering of the multiple audio segments according to their duration and voiceprint features to obtain audio segments corresponding to the same speaker; and adding the same user tag to the audio segments corresponding to the same speaker to obtain an audio signal with the added user tag. Detailed descriptions of each operation can be found in the foregoing method embodiments and will not be repeated here.

[0143] Furthermore, such as Figure 6 As shown, the server-side device also includes other components such as a communication component 66 and a power supply component 68. Figure 6 The diagram only shows a portion of the components and does not imply that the server-side device only includes... Figure 6 The components shown.

[0144] This application also provides a server-side device, the implementation structure of which is similar to... Figure 6 The implementation structure of the server-side devices shown is the same or similar, and can be referred to. Figure 6 The structural implementation of the server-side device is shown. The server-side device provided in this embodiment is similar to... Figure 6 The main difference between the server devices in the illustrated embodiments lies in the different functions implemented by the computer programs stored in the memory executed by the processor. For the server device provided in this embodiment, the computer programs stored in the memory executed by its processor can be used to: receive multiple audio segments sent by the audio receiving device; extract the voiceprint features corresponding to the multiple audio segments; perform hierarchical clustering of the multiple audio segments based on their duration and voiceprint features to obtain audio segments corresponding to the same speaker; and add the same user tag to the audio segments corresponding to the same speaker to obtain an audio signal with the added user tag. Detailed descriptions of each operation can be found in the descriptions in the foregoing method embodiments, and will not be repeated here.

[0145] Accordingly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, enables the processor to implement the steps that can be executed by a server device in the audio signal processing method embodiments.

[0146] Correspondingly, the embodiment of the present application also provides a computer program product, comprising computer programs / instructions, which, when executed by a processor, enable the processor to implement each step in the method embodiment of the audio signal processing method shown in the above.

[0147] In addition to the above device, the embodiment of the present application also provides a conference recording device, comprising a memory and a processor; the memory is used to store a computer program; the processor is coupled with the processor and is used to execute the computer program stored in the memory, so as to: collect an audio signal in a multi-person conference scene, identify a speaker change point in the audio signal; according to the speaker change point, the audio signal is divided into a plurality of audio segments, and a voiceprint feature of the plurality of audio segments is extracted; according to the time length and the voiceprint feature of the plurality of audio segments, the plurality of audio segments are hierarchically clustered to obtain audio segments corresponding to the same speaker; the same user mark is added to the audio segments corresponding to the same speaker to obtain an audio signal with a user mark added; conference recording information is generated according to the audio signal with the user mark added, and the conference recording information comprises a conference identifier.

[0148] The embodiment of the present application also provides a conference recording presentation device, comprising a memory and a processor; the memory is used to store a computer program; the processor is coupled with the processor and is used to execute the computer program stored in the memory, so as to: receive a conference review request, the conference review request containing a conference identifier to be presented; according to the conference identifier, conference recording information to be presented is obtained; the conference recording information is presented, and the conference recording information is generated according to an audio signal with a user mark added in a multi-person conference scene; wherein the same user mark is added to the audio segments corresponding to the same speaker in the plurality of audio segments divided according to the speaker change point in the audio signal, and the audio segments corresponding to the same speaker are obtained by hierarchically clustering the plurality of audio segments according to the time length and the voiceprint feature of the plurality of audio segments.

[0149] Correspondingly, the embodiment of the present application also provides a computer readable storage medium storing a computer program, which, when executed by a processor, enables the processor to implement each step in the method embodiment shown in the above. Figure 3d or Figure 3e each step in the method embodiment shown in the above.

[0150] Correspondingly, the embodiment of the present application also provides a computer program product, comprising computer programs / instructions, which, when executed by a processor, enable the processor to implement each step in the method embodiment shown in the above. Figure 3d or Figure 3e each step in the method embodiment shown in the above.

[0151] the aboveFigure 5 and Figure 6 The communication component in the aforementioned

[0152] The display in the aforementioned Figure 5 The display in the aforementioned

[0153] The power supply component in the aforementioned Figure 5 and Figure 6 The power supply component in the aforementioned

[0154] The power supply component in the aforementioned Figure 5 The audio component in the aforementioned

[0155] Those skilled in the art will understand that embodiments of the present application can be provided as a method, a system or a computer program product. Accordingly, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment or an embodiment combining software and hardware aspects. Furthermore, the present application can take the form of a computer program product on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage devices, etc.) embodying computer readable program code thereon.

[0156] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks or in the flowchart one or more blocks.

[0157] These computer program instructions can also be stored in a computer readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer readable memory produce an article of manufacture including instructions which implement the function specified in the flowchart block or blocks or in the flowchart one or more blocks.

[0158] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks or in the flowchart one or more blocks.

[0159] In one typical configuration, the computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0160] The memory can include non-persistent memory and / or persistent memory, such as flash memory, or other non-volatile memory. The memory is an example of computer readable media. The computer readable media can also include transmission media or signals.

[0161] Computer-readable media includes permanent and non-permanent, movable and non-movable media that can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible to a computing device. According to the definition herein, computer-readable media does not include transitory media such as modulated data signals and carriers.

[0162] It should also be noted that the terms "comprising", "containing", or any other variant thereof are intended to cover a non-exclusive inclusion, such that a process, method, article or apparatus that comprises a list of elements does not only include those elements, but can also include other elements not expressly listed or inherent to such process, method, article or apparatus. Without more limitations, the element defined by the statement "comprising a" does not exclude the presence of additional identical elements in the process, method, article or apparatus that includes the element.

[0163] The above only describes the embodiments of the present application and is not intended to limit the present application. Those skilled in the art can make various changes and modifications to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the scope of the claims of the present application.

Claims

1. An audio signal processing method, characterized in that, include: Identify speaker change points in audio signals collected in multi-person speaking scenarios; The audio signal is divided into multiple audio segments based on the speaker change points, and the voiceprint features of the multiple audio segments are extracted. Based on the duration and voiceprint features of the multiple audio segments, hierarchical clustering is performed on the multiple audio segments to obtain audio segments corresponding to the same speaker; the hierarchical clustering is performed on the multiple audio segments in order of duration from longest to shortest. Add the same user tag to audio segments corresponding to the same speaker to obtain audio signals with added user tags.

2. The method according to claim 1, characterized in that, Based on the duration and voiceprint features of the multiple audio segments, hierarchical clustering is performed on the multiple audio segments to obtain audio segments corresponding to the same speaker, including: Based on the duration of the multiple audio segments, the multiple audio segments are layered to obtain multi-layered audio segments corresponding to different duration ranges; Based on the voiceprint features corresponding to the multiple audio segments, the multi-layer audio segments are hierarchically clustered in order of duration from longest to shortest to obtain at least one clustering result. Each clustering result includes audio segments corresponding to the same speaker.

3. The method according to claim 2, characterized in that, Based on the duration of the multiple audio segments, the multiple audio segments are layered to obtain multi-layered audio segments corresponding to different duration ranges, including: Based on the duration of the multiple audio segments and the preset duration thresholds for each layer, the multiple audio segments are layered to obtain multi-layer audio segments corresponding to different duration ranges. Among them, the smaller the number of layers, the larger the corresponding duration threshold, and the duration of the audio segment in each layer is greater than or equal to the duration threshold of that layer.

4. The method according to claim 3, characterized in that, Based on the voiceprint features corresponding to the multiple audio segments, the multi-layered audio segments are hierarchically clustered in descending order of duration to obtain at least one clustering result, including: For the audio segments in the first layer, the audio segments in the first layer are clustered according to the corresponding voiceprint features to obtain at least one clustering result; For audio segments outside the first layer, cluster them according to their corresponding voiceprint features, in ascending order of layer number, and then cluster them into the existing clustering results. If there are remaining audio segments in the first layer that have not been clustered into the existing clustering results, then the remaining audio segments are clustered according to the voiceprint features corresponding to the remaining audio segments to generate new clustering results, until every audio segment in all layers is clustered into a single clustering result.

5. The method according to claim 3, characterized in that, Identify speaker change points in audio signals captured in multi-person speaking scenarios, including: The audio signal is used to locate the sound source in order to obtain the change point of the sound source location; Based on the change point of the sound source location, the speaker change point in the audio signal is determined; wherein each audio segment segmented by the speaker change point corresponds to a unique sound source location.

6. The method according to claim 5, characterized in that, Based on the voiceprint features corresponding to the multiple audio segments, the multi-layered audio segments are hierarchically clustered in descending order of duration to obtain at least one clustering result, including: Based on the voiceprint features and sound source locations corresponding to the multiple audio segments, the multi-layer audio segments are hierarchically clustered in descending order of duration to obtain at least one clustering result. Each clustering result includes audio segments corresponding to the same speaker.

7. The method according to claim 6, characterized in that, Based on the voiceprint features and sound source locations corresponding to the multiple audio segments, hierarchical clustering is performed on the multi-layered audio segments in descending order of duration to obtain at least one clustering result, including: For the audio segments in the first layer, the audio segments in the first layer are clustered according to the voiceprint features and sound source locations corresponding to the audio segments in the first layer, and at least one clustering result is obtained; For audio segments outside the first layer, cluster them according to their layer number (from smallest to largest) and their corresponding speaker features and sound source locations; and If there are remaining audio segments in the first layer that have not been clustered into the existing clustering results, then the remaining audio segments are clustered according to the voiceprint features and sound source locations corresponding to the remaining audio segments to generate new clustering results, until every audio segment in all layers is clustered into a single clustering result.

8. The method according to claim 7, characterized in that, For the audio segments in the first layer, based on the corresponding voiceprint features and sound source locations, the audio segments in the first layer are clustered to obtain at least one clustering result, including: If the first layer contains at least two audio segments, calculate the overall similarity between the at least two audio segments in the first layer based on the voiceprint features and sound source locations corresponding to the at least two audio segments in the first layer; Based on the overall similarity between at least two audio segments in the first layer, the at least two audio segments in the first layer are divided into at least one clustering result; and Based on the voiceprint features and sound source locations corresponding to the audio segments contained in the at least one clustering result, the cluster centers of the at least one clustering result are calculated respectively, and the cluster centers include the central voiceprint features and the central sound source location.

9. The method according to claim 8, characterized in that, For audio segments outside the first layer, in ascending order of layer number, the audio segments outside the first layer are clustered into the existing clustering results, based on their corresponding speaker features and sound source locations, including: For any audio segment that is not in the first layer, calculate the overall similarity between the audio segment and the existing clustering results based on the corresponding voiceprint features, sound source location, and cluster center of the existing clustering results. If there is a target clustering result in the existing clustering results that has an overall similarity that meets the set similarity conditions with the audio segment, the audio segment is added to the target clustering result, and the cluster center of the target clustering result is updated according to the voiceprint features and sound source location corresponding to the audio segment.

10. The method according to claim 9, characterized in that, The cluster centers of the target clustering result are updated based on the voiceprint features and sound source location corresponding to the audio segment, including: Determine the layer number to which each audio segment in the target clustering result belongs, where different layer numbers correspond to different weights, and the smaller the layer number, the greater the corresponding weight; Based on the weights corresponding to the layers to which each audio segment belongs, the voiceprint features and sound source locations corresponding to each audio segment are weighted and summed to obtain new central voiceprint features and new central sound source locations, which are then used as the updated cluster centers of the target clustering results.

11. The method according to claim 6, characterized in that, Based on the voiceprint features and sound source locations corresponding to the multiple audio segments, hierarchical clustering is performed on the multi-layered audio segments in descending order of duration to obtain at least one clustering result, including: Based on the voiceprint features corresponding to the audio segments in each layer, the audio segments in each layer are clustered to obtain the clustering results for each layer; Following the order of increasing layer number, clustering is performed on adjacent layers based on the voiceprint characteristics of each layer's clustering results to obtain at least one clustering result.

12. The method according to any one of claims 1-11, characterized in that, Also includes: The audio signal with added user tags is output to the transcription device, so that the transcription device can convert the audio signal with added user tags into text information with user tags; or The audio signal with added user tags is output to the playback device so that the playback device can play the audio signal with user tags; or The audio signal with added user tags is output to the upper-layer application so that the upper-layer application can obtain the user information corresponding to the user tags and associate it with the audio segments with user tags.

13. The method according to claim 12, characterized in that, Also includes: Receive a first query request, the first query request including a user tag to be queried, obtain the text information corresponding to the user tag to be queried from the text information with the user tag, and return it to the query terminal that initiated the first query request; or Receive a second query request, the second query request including an audio segment to be queried, extract the user tag corresponding to the audio segment to be queried from the audio signal with added user tags, and return the user tag and / or the user information corresponding to the user tag to the query terminal that initiated the second query request.

14. An audio signal processing method, characterized in that, include: Sound source localization is performed on audio signals collected in multi-person speaking scenarios to obtain the change points of sound source location; The audio signal is divided into multiple audio segments based on the change points of the sound source location, and the voiceprint features of the multiple audio segments are extracted. Based on the voiceprint features and sound source locations of the multiple audio segments, the multiple audio segments are clustered to obtain audio segments corresponding to the same speaker; the clustering is a hierarchical clustering of the multiple audio segments in descending order of duration; Add the same user tag to audio segments corresponding to the same speaker to obtain audio signals with added user tags.

15. The method according to claim 14, characterized in that, Based on the voiceprint features and sound source locations of the multiple audio segments, the multiple audio segments are clustered to obtain audio segments corresponding to the same speaker, including: Based on the duration of the multiple audio segments, the multiple audio segments are layered to obtain multi-layered audio segments corresponding to different duration ranges; Based on the voiceprint features and sound source locations corresponding to the multiple audio segments, the multi-layer audio segments are hierarchically clustered in descending order of duration to obtain at least one clustering result. Each clustering result includes audio segments corresponding to the same speaker.

16. A method for taking meeting minutes, characterized in that, include: Collect audio signals from multi-person conference scenarios and identify speaker change points in the audio signals; The audio signal is divided into multiple audio segments based on the speaker change points, and the voiceprint features of the multiple audio segments are extracted. Based on the duration and voiceprint features of the multiple audio segments, hierarchical clustering is performed on the multiple audio segments to obtain audio segments corresponding to the same speaker; the hierarchical clustering is performed on the multiple audio segments in order of duration from longest to shortest. Add the same user tag to audio segments corresponding to the same speaker to obtain audio signals with user tags added; Meeting record information is generated based on the audio signal with added user tags, and the meeting record information includes a meeting identifier.

17. A method for presenting meeting minutes, characterized in that, include: Receive a meeting access request, the meeting access request containing the meeting identifier to be presented; Based on the meeting identifier, obtain the meeting record information to be presented; The meeting record information is presented, which is generated based on audio signals with user tags added in a multi-person meeting scenario; Among the multiple audio segments segmented based on the speaker change points in the audio signal, the audio segments corresponding to the same speaker have the same user tag added. The audio segments corresponding to the same speaker are obtained by hierarchical clustering of the multiple audio segments based on their duration and voiceprint features. The hierarchical clustering is performed on the multiple audio segments in order of their duration from longest to shortest.

18. An audio processing system, characterized in that, include: Sound pickup equipment and server equipment; The sound pickup device is deployed in a multi-person speaking scenario to collect audio signals in the multi-person speaking scenario, identify speaker change points in the audio signals, divide the audio signals into multiple audio segments based on the speaker change points, and extract the voiceprint features corresponding to the multiple audio segments. The server device is used to perform hierarchical clustering of the multiple audio segments based on their duration and voiceprint features to obtain audio segments corresponding to the same speaker; the hierarchical clustering is performed on the multiple audio segments in descending order of duration; and the same user tag is added to the audio segments corresponding to the same speaker to obtain audio signals with added user tags.

19. The system according to claim 18, characterized in that, The sound pickup device is specifically used for: locating the sound source of the audio signal to obtain the change point of the sound source position; determining the speaker change point in the audio signal based on the change point of the sound source position; wherein each audio segment segmented by the speaker change point corresponds to a unique sound source position; The server-side device is specifically used to: perform hierarchical clustering of the multiple audio segments based on their duration, voiceprint features, and sound source location, so as to obtain audio segments corresponding to the same speaker.

20. An audio processing system, characterized in that, include: Sound pickup equipment and server equipment; The sound pickup device is deployed in a multi-person speaking scenario to collect audio signals in the multi-person speaking scenario, identify speaker change points in the audio signals, and divide the audio signals into multiple audio segments based on the speaker change points; The server-side device is used to extract the voiceprint features corresponding to the multiple audio segments, and perform hierarchical clustering on the multiple audio segments according to their duration and voiceprint features to obtain audio segments corresponding to the same speaker; the hierarchical clustering is performed on the multiple audio segments in descending order of duration; and the same user tag is added to the audio segments corresponding to the same speaker to obtain audio signals with added user tags.

21. A sound pickup device, characterized in that, include: Processor and memory; The memory is used to store computer programs; The processor is coupled to the memory and is used to execute the computer program for: identifying speaker change points in audio signals collected in a multi-person speaking scenario; segmenting the audio signal into multiple audio segments based on the speaker change points, and extracting the voiceprint features of the multiple audio segments; Based on the duration and voiceprint features of the multiple audio segments, hierarchical clustering is performed on the multiple audio segments to obtain audio segments corresponding to the same speaker; the hierarchical clustering is performed on the multiple audio segments in descending order of duration; the same user tag is added to the audio segments corresponding to the same speaker to obtain audio signals with added user tags.

22. A sound pickup device, characterized in that, include: Processor and memory; The memory is used to store computer programs; The processor is coupled to the memory and is used to execute the computer program for: locating the sound source of an audio signal acquired in a multi-person speaking scenario to obtain the change point of the sound source location; dividing the audio signal into multiple audio segments based on the change point of the sound source location, and extracting the voiceprint features of the multiple audio segments; Based on the voiceprint features and sound source locations of the multiple audio segments, the multiple audio segments are clustered to obtain audio segments corresponding to the same speaker; the clustering is performed hierarchically according to the length of the multiple audio segments from longest to shortest; the same user tag is added to the audio segments corresponding to the same speaker to obtain audio signals with added user tags.

23. A server-side device, characterized in that, include: Processor and memory; The memory is used to store computer programs; The processor, coupled to the memory, is used to execute the computer program for: receiving multiple audio segments and their corresponding voiceprint features sent by the audio device; performing hierarchical clustering on the multiple audio segments according to their duration and voiceprint features to obtain audio segments corresponding to the same speaker; the hierarchical clustering is performed on the multiple audio segments in descending order of duration; and adding the same user tag to the audio segments corresponding to the same speaker to obtain an audio signal with added user tags.

24. A server-side device, characterized in that, include: Processor and memory; The memory is used to store computer programs; The processor is coupled to the memory and is used to execute the computer program for: receiving multiple audio segments sent by the audio device; extracting the voiceprint features corresponding to the multiple audio segments; performing hierarchical clustering on the multiple audio segments according to their duration and voiceprint features to obtain audio segments corresponding to the same speaker; the hierarchical clustering is performed on the multiple audio segments in descending order of duration; and adding the same user tag to the audio segments corresponding to the same speaker to obtain an audio signal with added user tags.

25. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it causes the processor to perform the steps of the method according to any one of claims 1-17.

26. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, the processor is caused to perform the steps of the method according to any one of claims 1-17.

Citation Information

Patent Citations

  • Voice analysis method and device

    CN111613249A

  • Conference summary transcription method and device and storage medium

    CN112037791A

  • Conference summary automatic generation method for video conference

    CN112165599A