Speaker recognition method based on bipartite graph matching and electronic equipment

By adopting a speaker recognition method based on binary graph matching in multi-speaker scenarios, the problems of insufficient real-time and high complexity in the prior art are solved, and efficient and accurate speaker recognition is achieved, which is suitable for dynamically changing long-term voice scenarios.

CN119993166AActive Publication Date: 2025-05-13BEIJING TONGXIANG QIANFANG TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510162083.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2025-05-13
Estimated Expiration
2045-02-13

AI Technical Summary

Technical Problem

The prior art has problems such as insufficient real-time performance, high complexity of long audio processing, difficulty in distinguishing multiple speakers, and poor adaptability of dynamically changing scenarios in multi-speaker scenarios.

Method used

Using a speaker recognition method based on binary graph matching, the audio clips of the target duration are split by real-time audio streams, the voiceprint embedding vector is extracted, the embedding matrix is ​​constructed, and the dichotomy is used to accelerate feature comparison, and the computing resource allocation is optimized in combination with the Top-K strategy.

Benefits of technology

It significantly reduces the computational complexity and improves recognition accuracy. It is suitable for complex long-term voice scenarios, especially in scenes where more speakers are frequently alternating.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993166A_ABST
    Figure CN119993166A_ABST
Patent Text Reader

Abstract

The invention discloses a bipartite graph matching-based speaker recognition method and electronic equipment, and belongs to the technical field of audio recognition. The method comprises the following steps: acquiring an audio stream, and segmenting the audio stream into continuous target duration audio clips in real time; determining a voiceprint embedding vector corresponding to each audio clip of the target duration; determining an embedding matrix based on the voiceprint embedding vector; determining a target voiceprint feature corresponding to each target duration audio clip according to a bipartite graph method; calculating the similarity between the voiceprint embedding vectors; and grouping the voiceprint embedding vectors based on the similarity between the voiceprint embedding vectors, and clustering the voiceprint embedding vectors corresponding to the same speaker to obtain a clustering result and generate a target recognition report. According to the invention, the calculation complexity can be reduced and the identification precision can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the technical field of audio recognition, and in particular relates to a speaker recognition method and electronic device based on bipartite graph matching. Background Art

[0002] Speaker recognition technology is a technology that recognizes and distinguishes speakers by analyzing unique physiological and behavioral characteristics in voice signals. This field has made significant progress in recent years and has demonstrated important value in applications such as security verification, smart assistants, and voice transcription. However, existing technologies mostly extract and compare voiceprint features of the entire audio segment to identify the speaker. This method usually performs well in single-speaker or short audio clips, but in long-term, multi-speaker scenarios, such as complex application scenarios such as conference records, phone calls, and video voice recognition, it faces problems and challenges such as lack of real-time performance, high complexity in long audio processing, difficulty in distinguishing multiple speakers, and poor adaptability to dynamically changing scenarios.

[0003] In order to solve the above problems, a speaker recognition method and electronic device based on bipartite graph matching are proposed in this application. Summary of the invention

[0004] In order to address the deficiencies of the prior art, the present application provides a speaker recognition method and electronic device based on bipartite graph matching to solve the problems and challenges of the prior art recognition methods for multiple people speaking, such as insufficient real-time performance, high complexity in long audio processing, difficulty in distinguishing multiple speakers, and poor adaptability to dynamically changing scenes.

[0005] The technical effects to be achieved by this application are achieved through the following solutions:

[0006] In a first aspect, the present application provides a speaker recognition method based on bipartite graph matching, the method comprising:

[0007] Acquire an audio stream, and divide the audio stream into continuous audio segments of target duration in real time, wherein the audio stream is audio of multiple people speaking;

[0008] Based on each of the target duration audio segments, determining a voiceprint embedding vector corresponding to each of the target duration audio segments;

[0009] Determining an embedding matrix based on the voiceprint embedding vector, wherein the embedding matrix is ​​a structured representation of the voiceprint embedding vector;

[0010] Determine the target voiceprint feature corresponding to each audio segment of the target duration according to the bipartite graph method;

[0011] Based on the embedding matrix and the target voiceprint feature, calculating the similarity between each of the voiceprint embedding vectors;

[0012] grouping the voiceprint embedding vectors based on the similarity between the voiceprint embedding vectors, clustering the voiceprint embedding vectors corresponding to the same speaker, and obtaining a clustering result;

[0013] Based on the clustering results and each of the target duration audio segments, each person's speaking time period and duration are generated, and a target recognition report is generated.

[0014] In some embodiments, determining the target voiceprint feature corresponding to each audio segment of the target duration according to the bipartite graph method includes:

[0015] Each of the target duration audio clips and the known voiceprint features are respectively used as nodes on both sides of a bipartite graph to construct a corresponding bipartite graph structure; wherein the known voiceprint features are dynamically changing in real time, and new voiceprint features acquired in real time are added to the known voiceprint features during the recognition process;

[0016] The first voiceprint feature that best matches each of the target duration audio segments is searched in the bipartite graph structure, and the first voiceprint feature is determined as the target voiceprint feature corresponding to each of the target duration audio segments; that is, the first voiceprint feature is determined from the known voiceprint features.

[0017] In some embodiments, the similarity between the voiceprint embedding vectors is calculated based on the cosine similarity method.

[0018] In some embodiments, the voiceprint embedding vectors corresponding to the same speaker are clustered using the K-means method or the hierarchical clustering method.

[0019] In some embodiments, determining the voiceprint embedding vector corresponding to each audio segment of the target duration based on each audio segment of the target duration includes:

[0020] Using a deep learning model, a voiceprint embedding vector corresponding to each audio segment of the target duration is determined from each audio segment of the target duration.

[0021] In some embodiments, the deep learning model includes: a convolutional neural network or a long short-term memory network.

[0022] In some embodiments, the target recognition report is used to present dynamic features of the interaction between speakers.

[0023] In some embodiments, the target duration of the audio segment is 25 seconds, 30 seconds, or 35 seconds.

[0024] In a second aspect, the present application provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements any one of the aforementioned methods for speaker recognition based on bipartite graph matching when executing the computer program.

[0025] In a third aspect, the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement any of the aforementioned methods for speaker recognition based on bipartite graph matching.

[0026] The speaker recognition method and electronic device based on bipartite graph matching provided by the present application adopt a real-time update and dynamic adaptation method in a multi-speaker scenario, extract high-frequency feature fragments through voiceprint sampling, use bipartite matching to accelerate the feature comparison process, and then combine the Top-K strategy to optimize computing resource allocation. This method can significantly reduce the computational complexity while improving recognition accuracy, and is particularly suitable for complex and long-term speech scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] In order to more clearly illustrate the embodiments of the present application or the existing technical solutions, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0028] Figure 1 This is a flow chart of a speaker recognition method based on bipartite graph matching in one embodiment of the present application;

[0029] Figure 2 It is a schematic block diagram of an electronic device in an embodiment of the present application. DETAILED DESCRIPTION

[0030] In order to make the purpose, technical solution and advantages of the present application clearer, the technical solution of the present application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present application.

[0031] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in one or more embodiments of the present application should be understood by people with ordinary skills in the field to which the present application belongs. The "first", "second" and similar words used in one or more embodiments of the present application do not indicate any order, quantity or importance, but are only used to distinguish different components. "Include" or "comprise" and similar words mean that the elements or objects appearing before the word cover the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may also change accordingly.

[0032] Related audio recognition methods face the following core problems and challenges:

[0033] 1. Lack of real-time performance

[0034] Traditional speaker recognition methods usually analyze and output results only after collecting and processing complete audio data. This post-processing method causes significant time delays in the entire recognition process, making it difficult to meet real-time requirements. In dynamic scenarios, real-time identification and updating of speaker identities are particularly critical. For example, in video conferencing, real-time identification of speakers is the basis for accurate subtitles and instant responses, and traditional methods cannot meet this requirement.

[0035] 2. Long audio processing is complex

[0036] In the processing of long audio data, a complete analysis of all speech segments consumes a lot of computing resources. Especially in multi-speaker scenarios, the complexity of feature extraction and matching processes increases exponentially. The length and complexity of audio data significantly increase the amount of processing computation, which places higher demands on hardware resources and algorithm efficiency. For resource-constrained devices (such as mobile terminals), this high computing burden is even more unbearable.

[0037] 3. Difficulty distinguishing between multiple speakers

[0038] In scenarios where multiple speakers take turns speaking, background noise, speech overlap, and similarity of speech features can significantly affect the accuracy of speaker recognition. For example, in a conference room or on a phone call, background noise may mask the speaker's speech features, and when the voiceprints of different speakers are highly similar, traditional methods are difficult to accurately distinguish. In particular, when multiple speakers frequently take turns speaking, traditional methods are prone to misjudgment or omission.

[0039] 4. Poor adaptability to dynamically changing scenes

[0040] In long-duration speech scenarios, speakers may change dynamically over time, such as when new speakers join or when existing speakers’ speech features change (such as speech speed or intonation). Traditional global feature extraction methods are usually difficult to quickly adapt to new speech data because they rely on pre-extracted fixed patterns. This limitation can easily lead to delayed or distorted recognition results in practical applications.

[0041] 5. Limitations of Deep Learning Technology

[0042] In recent years, deep learning technology has achieved excellent performance in processing short audio clips. For example, Convolutional Neural Networks (CNN) and Recurrent Neural Networks (RNN) have demonstrated high accuracy in voiceprint feature extraction and classification. However, these models still face many challenges in processing long audio and real-time recognition. Deep learning models usually require a large amount of training data and high computing power, and in multi-speaker scenarios, how to dynamically update the model and efficiently process overlapping speech remains an unresolved problem.

[0043] In response to the above challenges, this application development attempts to overcome them by performing real-time speaker analysis every 30 seconds, combining a binary matching algorithm, a Top-K strategy, and voiceprint sampling technology. This includes using binary matching to quickly narrow down the range of possible voiceprint candidates, and using a Top-K strategy to prioritize the most distinctive voiceprint features for analysis. Voiceprint sampling technology captures local audio features in a short period of time, enabling real-time updates and dynamic adaptation. For example, in a multi-speaker scenario, high-frequency feature fragments are extracted through voiceprint sampling, binary matching is used to accelerate the feature comparison process, and the Top-K strategy is combined to optimize the allocation of computing resources. This can significantly reduce computational complexity while improving recognition accuracy, and is especially suitable for complex, long-duration speech scenarios.

[0044] The speaker recognition method based on bipartite graph matching in this application has the following advantages:

[0045] 1. Use binary matching algorithm to improve processing efficiency

[0046] This application combines a binary matching algorithm that can quickly narrow down the range of voiceprint candidates and greatly improve the efficiency of voiceprint feature matching. When faced with complex, long-term, multi-speaker audio data, the binary matching algorithm significantly reduces the computational complexity by reducing the number of candidates that need to be compared. This technology is particularly important in dynamic, real-time application scenarios, and can identify and update the speaker's identity in a short period of time, solving the real-time shortcomings of traditional methods.

[0047] 2. Top-K strategy optimizes feature selection

[0048] In view of the diversity of voiceprint features in multi-speaker scenarios, this application introduces the Top-K strategy, which gives priority to the most discriminative voiceprint features for analysis at each voiceprint comparison. In this way, this application not only improves the accuracy of recognition, but also optimizes the allocation of computing resources and avoids excessive calculation of low-discriminative features. Especially in complex scenarios, the Top-K strategy can effectively reduce the computing burden and improve the real-time response capability of the speaker recognition method based on bipartite graph matching in this application.

[0049] 3. Voiceprint sampling technology realizes dynamic adaptation

[0050] This application introduces voiceprint sampling technology, which enables the speaker recognition method based on bipartite graph matching of this application to dynamically update and adapt to new voice data by capturing local audio features in a short period of time. This is particularly important for scenarios where multiple speakers speak alternately, and can distinguish different speakers in real time, avoiding recognition errors caused by voice overlap or changes in speech speed in traditional methods. Voiceprint sampling technology improves the adaptability of the speaker recognition method based on bipartite graph matching of this application in long-term and changing environments, and solves the limitation that traditional global feature extraction methods cannot quickly adapt to dynamic changes.

[0051] 4. Optimize the allocation of computing resources

[0052] When faced with long audio data, traditional methods often consume a lot of computing resources. This application effectively reduces the amount of computation for each feature extraction and comparison by combining voiceprint sampling technology with bipartite matching and Top-K strategy, thereby reducing the hardware resource requirements of the speaker recognition method based on bipartite graph matching in this application. This optimization is particularly prominent on devices with limited computing resources such as mobile terminals, allowing complex speaker recognition tasks to run efficiently in low-resource environments.

[0053] 5. Real-time and multi-speaker adaptability

[0054] In application scenarios such as conferences or phone calls where multiple speakers frequently alternate, the technical solution of this application can update the speaker identity in real time and accurately distinguish overlapping speech. Combining bipartite matching and Top-K strategy, the speaker recognition method based on bipartite graph matching in this application can quickly adapt to different speaker changes, effectively solving the problems of poor adaptability and incorrect recognition of existing technologies in dynamic and complex scenarios.

[0055] Various non-limiting implementations of the present application are described in detail below in conjunction with the accompanying drawings.

[0056] First, refer to Figure 1, the speaker recognition method based on bipartite graph matching of the present application is described in detail.

[0057] The present application provides a speaker recognition method based on bipartite graph matching, the method comprising:

[0058] S1: Acquire an audio stream, and divide the audio stream into continuous audio segments of target duration in real time, wherein the audio stream is audio of multiple people speaking;

[0059] S2: Based on each audio segment of the target duration, determine a voiceprint embedding vector corresponding to each audio segment of the target duration;

[0060] S3: determining an embedding matrix based on the voiceprint embedding vector, wherein the embedding matrix is ​​a structured representation of the voiceprint embedding vector;

[0061] S4: determining a target voiceprint feature corresponding to each audio segment of the target duration according to a bipartite graph method;

[0062] S5: Calculating the similarity between each of the voiceprint embedding vectors based on the embedding matrix and the target voiceprint feature;

[0063] S6: grouping the voiceprint embedding vectors based on the similarity between the voiceprint embedding vectors, clustering the voiceprint embedding vectors corresponding to the same speaker, and obtaining a clustering result;

[0064] S7: Generate each person's speaking time period and duration based on the clustering results and each of the target duration audio segments, and generate a target recognition report.

[0065] The technical solution of this application not only makes up for the shortcomings of the existing technology in terms of real-time performance, computational complexity and multi-speaker adaptability, but also improves the accuracy and efficiency of speaker recognition. The method of this application will provide more accurate and efficient technical support for applications such as smart conferences, real-time translation, and speech transcription, and promote the widespread application of speaker recognition technology in complex application scenarios.

[0066] Exemplarily, a continuous audio segment of target duration may be obtained through a voiceprint embedding and extraction module, specifically including:

[0067] The audio stream is divided into continuous audio segments of target duration, such as 30 seconds, using high-precision timeline segmentation technology to ensure that each segment contains enough voice information for subsequent processing. In some embodiments, the target duration audio segment is 25 seconds, 30 seconds, or 35 seconds.

[0068] In some embodiments, determining the target voiceprint feature corresponding to each audio segment of the target duration according to the bipartite graph method includes:

[0069] Each of the target duration audio clips and the known voiceprint features are respectively used as nodes on both sides of a bipartite graph to construct a corresponding bipartite graph structure; wherein the known voiceprint features are dynamically changing in real time, and new voiceprint features acquired in real time are added to the known voiceprint features during the recognition process;

[0070] The first voiceprint feature that best matches each of the target duration audio segments is searched in the bipartite graph structure, and the first voiceprint feature is determined as the target voiceprint feature corresponding to each of the target duration audio segments.

[0071] In some embodiments, the similarity between the voiceprint embedding vectors is calculated based on the cosine similarity method.

[0072] In some embodiments, the voiceprint embedding vectors corresponding to the same speaker are clustered using the K-means method or the hierarchical clustering method to achieve accurate recognition in a multi-speaker environment.

[0073] In some embodiments, determining the voiceprint embedding vector corresponding to each audio segment of the target duration based on each audio segment of the target duration includes:

[0074] Using a deep learning model, a voiceprint embedding vector corresponding to each audio segment of the target duration is determined from each audio segment of the target duration. The voiceprint embedding vector can effectively represent the unique acoustic characteristics of the speaker. The extracted voiceprint embedding vector can be normalized to unify the numerical range, eliminate the scale differences between different audio segments, and ensure the stability and accuracy of subsequent algorithms.

[0075] In some embodiments, the deep learning model includes: a convolutional neural network or a long short-term memory network.

[0076] After determining the voiceprint embedding vector corresponding to each of the target duration audio segments, the corresponding voiceprint embedding vector and its potential label (the label is used to identify the identity or category of the speaker to which each audio segment belongs) can be generated based on the timestamp information of each target duration audio segment, ensuring that the time correspondence between the label and each target duration audio segment is accurate.

[0077] For example, an optimization method such as the Hungarian algorithm can be used to find the best matching pair in the bipartite graph structure, so as to achieve accurate matching of each audio segment with the most likely voiceprint feature, thereby improving the accuracy and efficiency of recognition.

[0078] In some embodiments, the target recognition report is used to present dynamic features of the interaction between speakers.

[0079] For example, the clustering results of the target-length audio clips are subjected to time series analysis to generate the speaking time period and duration of each speaker, fully presenting the dynamic characteristics of multi-speaker interaction. The time series analysis results are integrated to output a detailed multi-speaker recognition report (i.e., target recognition report). This application supports subsequent application scenarios, such as meeting records, customer service analysis, etc., and improves the practicality and intelligence level of the speaker recognition method based on bipartite graph matching in this application.

[0080] This application brings many beneficial effects in voiceprint recognition and audio analysis in a multi-speaker environment. Compared with the existing technology, it has the following significant advantages:

[0081] 1. Improve the accuracy and stability of voiceprint recognition

[0082] By using advanced deep learning models (such as CNN and LSTM) to extract high-dimensional voiceprint embedding vectors and combining them with feature normalization, this solution significantly improves the accuracy and consistency of voiceprint recognition. The normalized embedding vectors effectively eliminate the numerical differences between different audio clips, ensuring the stable operation of subsequent matching and clustering algorithms.

[0083] 2. Efficient recognition capability in multi-speaker environments

[0084] By using a bipartite graph matching analysis module and an optimized matching algorithm (such as the Hungarian algorithm), this solution can accurately distinguish and identify the voiceprint features of different speakers in a complex multi-speaker environment. By constructing an embedding matrix to describe the similarity relationship between audio clips, the speaker recognition method based on bipartite graph matching in this application can efficiently process large amounts of audio data and adapt to scenarios where multiple speakers speak at the same time.

[0085] 3. Enhanced Similarity Measurement and Cluster Analysis

[0086] By using the cosine similarity metric and advanced clustering algorithms (such as K-means and hierarchical clustering), this solution can accurately quantify the similarity between voiceprint embedding vectors and effectively classify audio clips of the same speaker. This not only improves the accuracy of clustering, but also enhances the performance of the speaker recognition method based on bipartite graph matching in this application when processing high-dimensional voiceprint data.

[0087] 4. Comprehensive time series analysis and comprehensive identification

[0088] By performing time series analysis on the matching results of all audio clips, this solution can generate the speaking time period and duration of each speaker, and provide detailed dynamic information of multi-speaker interaction. This comprehensive recognition capability makes the speaker recognition method based on bipartite graph matching in this application more practical and intelligent in practical applications, such as meeting records, customer service analysis and other scenarios.

[0089] 5. Modular design, flexible and scalable

[0090] This application adopts a modular technical architecture, and each module (voiceprint embedding extraction, dynamic label and embedding matrix construction, bipartite graph matching analysis, similarity measurement and clustering, multi-speaker comprehensive recognition) is independent of each other and can be flexibly combined to facilitate system upgrades and functional expansion. Users can flexibly adjust and optimize each module according to specific needs to improve the adaptability and scalability of the speaker recognition method based on bipartite graph matching in this application.

[0091] 6. Optimized processing efficiency and resource utilization

[0092] Through precise audio segmentation and efficient feature extraction methods, this application greatly improves processing efficiency while ensuring recognition accuracy. The optimized design of feature standardization and embedding matrix construction reduces the consumption of computing resources and improves the performance of the speaker recognition method based on bipartite graph matching in large-scale audio data processing.

[0093] This application significantly improves the accuracy, efficiency and adaptability of voiceprint recognition by combining advanced voiceprint recognition technology and optimization algorithms in a multi-speaker environment. These beneficial effects not only enhance the practical value of voiceprint recognition in practical applications, but also demonstrate the technical innovation and broad application prospects of this application in the fields of audio analysis and multi-speaker recognition.

[0094] It should be noted that the method of one or more embodiments of the present application can be performed by a single device, such as a computer or server. The method of this embodiment can also be applied to a distributed scenario and completed by multiple devices cooperating with each other. In the case of such a distributed scenario, one of the multiple devices can only perform one or more steps in the method of one or more embodiments of the present application, and the multiple devices will interact with each other to complete the described method.

[0095] It should be noted that the above describes specific embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0096] Based on the same inventive concept, corresponding to any of the above-mentioned embodiments and methods, the present application also discloses an electronic device;

[0097] Specifically, Figure 2 The hardware structure diagram of an electronic device for a speaker recognition method based on bipartite graph matching provided in this embodiment is shown, and the device may include: a processor 410, a memory 420, an input / output interface 430, a communication interface 440, and a bus 450. The processor 410, the memory 420, the input / output interface 430, and the communication interface 440 are connected to each other in communication within the device through the bus 450.

[0098] The processor 410 can be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application.

[0099] The memory 420 may be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 420 may store an operating system and other application programs. When the technical solution provided in the embodiment of the present application is implemented by software or firmware, the relevant program code is stored in the memory 420 and is called and executed by the processor 410.

[0100] The input / output interface 430 is used to connect the input / output module to realize information input and output. The input / output module can be configured in the device as a component (not shown in the figure), or it can be externally connected to the device to provide corresponding functions. The input device may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device may include a display, a speaker, a vibrator, an indicator light, etc.

[0101] The communication interface 440 is used to connect a communication module (not shown) to realize communication interaction between the device and other devices. The communication module can realize communication through a wired mode (for example, USB, network cable, etc.) or a wireless mode (for example, mobile network, WIFI, Bluetooth, etc.).

[0102] The bus 450 includes a path that transmits information between the various components of the device (eg, the processor 410 , the memory 420 , the input / output interface 430 , and the communication interface 440 ).

[0103] It should be noted that, although the above device only shows the processor 410, the memory 420, the input / output interface 430, the communication interface 440 and the bus 450, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, it can be understood by those skilled in the art that the above device may also only include the components necessary for implementing the embodiment of the present application, and does not necessarily include all the components shown in the figure.

[0104] The electronic device in the above embodiment is used to implement the corresponding speaker recognition method based on bipartite graph matching in any of the above embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be described in detail here.

[0105] Based on the same inventive concept, corresponding to any of the above-mentioned embodiments and methods, one or more embodiments of the present application further provide a computer-readable storage medium, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute the speaker recognition method based on bipartite graph matching as described in any of the above embodiments.

[0106] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, tape disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device.

[0107] The computer instructions stored in the storage medium of the above embodiment are used to enable the computer to execute the speaker recognition method based on bipartite graph matching as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0108] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of the present application (including the claims) is limited to these examples. In line with the concept of the present application, the technical features in the above embodiments or different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of different aspects of one or more embodiments of the present application as described above, which are not provided in detail for the sake of simplicity.

[0109] In addition, to simplify the description and discussion, and in order not to make one or more embodiments of the present application difficult to understand, the known power / ground connections to the integrated circuit (IC) chip and other components may or may not be shown in the provided drawings. In addition, the device can be shown in the form of a block diagram to avoid making one or more embodiments of the present application difficult to understand, and this also takes into account the fact that the details of the implementation of these block diagram devices are highly dependent on the platform on which one or more embodiments of the present application will be implemented (that is, these details should be fully within the scope of understanding of those skilled in the art). In the case of elaborating specific details (e.g., circuits) to describe exemplary embodiments of the present application, it is obvious to those skilled in the art that one or more embodiments of the present application can be implemented without these specific details or when these specific details are changed. Therefore, these descriptions should be considered illustrative rather than restrictive.

[0110] Although the present application has been described in conjunction with specific embodiments of the present application, many replacements, modifications and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may use the embodiments discussed.

[0111] One or more embodiments of the present application are intended to cover all such substitutions, modifications and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of one or more embodiments of the present application should be included in the scope of protection of the present application.

Claims

1. A speaker recognition method based on bipartite graph matching, characterized in that: include: Acquire an audio stream, and divide the audio stream into continuous audio segments of target duration in real time, wherein the audio stream is audio of multiple people speaking; Based on each of the target duration audio segments, determining a voiceprint embedding vector corresponding to each of the target duration audio segments; Determining an embedding matrix based on the voiceprint embedding vector, wherein the embedding matrix is ​​a structured representation of the voiceprint embedding vector; Determine the target voiceprint feature corresponding to each audio segment of the target duration according to the bipartite graph method; Based on the embedding matrix and the target voiceprint feature, calculating the similarity between each of the voiceprint embedding vectors; grouping the voiceprint embedding vectors based on the similarity between the voiceprint embedding vectors, clustering the voiceprint embedding vectors corresponding to the same speaker, and obtaining a clustering result; Based on the clustering results and each of the target duration audio segments, each person's speaking time period and duration are generated, and a target recognition report is generated.

2. The speaker recognition method based on bipartite graph matching according to claim 1, characterized in that: The step of determining the target voiceprint feature corresponding to each audio segment of the target duration according to the bipartite graph method includes: Each of the target duration audio clips and the known voiceprint features are respectively used as nodes on both sides of a bipartite graph to construct a corresponding bipartite graph structure; wherein the known voiceprint features are dynamically changing in real time, and new voiceprint features acquired in real time are added to the known voiceprint features during the recognition process; The first voiceprint feature that best matches each of the target duration audio segments is searched in the bipartite graph structure, and the first voiceprint feature is determined as the target voiceprint feature corresponding to each of the target duration audio segments.

3. The speaker recognition method based on bipartite graph matching according to claim 1 or 2, characterized in that: The similarity between each of the voiceprint embedding vectors is calculated based on the cosine similarity method.

4. The speaker recognition method based on bipartite graph matching according to claim 3, characterized in that: The voiceprint embedding vectors corresponding to the same speaker are clustered using the K-means method or the hierarchical clustering method.

5. The speaker recognition method based on bipartite graph matching according to claim 4, characterized in that: The step of determining, based on each of the target duration audio segments, a voiceprint embedding vector corresponding to each of the target duration audio segments comprises: Using a deep learning model, a voiceprint embedding vector corresponding to each audio segment of the target duration is determined from each audio segment of the target duration.

6. The speaker recognition method based on bipartite graph matching according to claim 5, characterized in that: The deep learning model includes: a convolutional neural network or a long short-term memory network.

7. The speaker recognition method based on bipartite graph matching according to claim 1, characterized in that: The target recognition report is used to present the dynamic characteristics of the interaction between the speakers.

8. The speaker recognition method based on bipartite graph matching according to claim 7, characterized in that: The target duration of the audio clip is 25 seconds, 30 seconds, or 35 seconds.

9. An electronic device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the speaker recognition method based on bipartite graph matching as described in any one of claims 1 to 8 when executing the computer program.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the speaker recognition method based on bipartite graph matching as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Speaker recognition method and device based on clustering, equipment and storage medium

    CN113851136A

  • Speaker audio separation method, terminal equipment and storage medium

    CN115691474A

  • Speaker verification methods and apparatus

    US20170061968A1

  • Trial-based calibration for audio-based identification, recognition, and detection system

    US20190013013A1

  • Audio data processing method, device and system

    WO2022062471A1

Cited By

  • Zero-configuration adaptive speaker recognition method and system

    CN120708626A

  • A zero-configuration adaptive speaker recognition method and system

    CN120708626B