Voiceprint-based memory representation method and system

By using a voiceprint representation memory-based method to dynamically adjust the order of voiceprint feature extraction, the problems of resource waste and low recognition efficiency in existing technologies are solved, and efficient recognition in complex environments is achieved.

CN121075338BActive Publication Date: 2026-02-03GUANGZHOU SIZHENG ELECTRONIC TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511625359.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-07
Publication Date
2026-02-03
Estimated Expiration
2045-11-07

AI Technical Summary

Technical Problem

Existing voiceprint recognition technology is inefficient in real-time and low-power scenarios, and the fixed feature extraction sequence cannot adapt to changes in user status and scenario, resulting in wasted resources and low recognition efficiency.

Method used

The method adopts a voiceprint representation memory approach, which uses a customized, periodically updated general voiceprint representation extraction sequence and a dynamic voiceprint representation extraction sequence to dynamically adjust the feature extraction order according to the user's voiceprint characteristics, prioritizes the extraction of high-discrimination features, and combines an early termination mechanism to improve recognition efficiency.

Benefits of technology

It improves the screening efficiency and recognition accuracy of voiceprint recognition, reduces resource consumption, adapts to changes in user status and scenario, and enhances the robustness and efficiency of recognition in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121075338B_ABST
    Figure CN121075338B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of sound pickup recognition, in particular to a sound pickup recognition method and system based on voiceprint representation memory, which combines general voiceprint representation extraction sequence and dynamic sound feature extraction sequence through setting, avoids the resource waste of traditional technology "one-time extraction of all features", overcomes the limitation of "fixed sequence cannot cope with real-time differences", and further reduces the computing cost by cooperating with the early termination mechanism (stopping when recognizing the unique user), and finally realizes the comprehensive improvement of recognition robustness, efficiency and accuracy in complex interference environment.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of sound recognition, more particularly, it relates to a sound recognition method and system based on voiceprint representation memory. BACKGROUND

[0002] The existing voiceprint recognition technology generally adopts the mode of "extracting all voiceprint representations in full amount", that is, all voiceprint representations including MFCC, differential MFCC, spectral entropy, short-time energy, pitch frequency, etc. are extracted without discrimination from the input sound signal, and then subsequent comparison and analysis are carried out. However, different voiceprint representations have significant differences in the efficiency of "distinguishing users", and part of the representations have weak distinguishing ability in the current scene or user group, but still need to consume resources such as computing power and storage for extraction and calculation, which not only wastes resources in the feature extraction link, but also directly reduces the overall efficiency of sound recognition due to redundant calculation, especially difficult to meet the needs of real-time and low-power scenarios such as embedded devices and mobile terminals.

[0003] Even if part of the technology adopts "fixed feature extraction sequence", it can only mechanically extract voiceprint representations according to the preset priority, and cannot adapt to the dynamic changes of the scene and the differences of the user group. On the one hand, the distinguishing efficiency of voiceprint representation will fluctuate with the change of user state (such as health, emotion) or even the size of user group, and the fixed sequence is difficult to respond to these dynamic factors; on the other hand, in the presence of multiple candidate users (grabbing users), the fixed sequence lacks specificity to "feature differences among the current candidate users", and still extracts in full amount or fixed priority, which leads to the failure to prioritize the use of high-value features that are more critical for distinguishing, intensifies resource consumption, and makes it difficult to quickly narrow down the candidate range, further restricting the improvement of recognition efficiency.

[0004] Based on the above, the present application provides a sound recognition method and system based on voiceprint representation memory. SUMMARY

[0005] In view of the deficiencies of the prior art, the purpose of the present application is to provide a sound recognition method and system based on voiceprint representation memory.

[0006] To achieve the above-mentioned purpose, the present application provides the following technical solutions:

[0007] The sound recognition method based on voiceprint representation memory comprises the following steps:

[0008] Step 1: Customizing a general voiceprint representation extraction sequence that can be updated regularly;

[0009] Step two: whenever a sound signal is collected, the collected sound signal is sequentially extracted according to the general voiceprint characteristic extraction sequence, and each extracted voiceprint characteristic is labeled as an extracted characteristic. The feature vector of the extracted characteristic in the sound signal is obtained.

[0010] Step three: the feature vectors of the extracted characteristic of each user in the user voiceprint characteristic memory library are obtained, and the characteristic vector similarity of the extracted characteristic of each user is further obtained. According to the comparison result of the characteristic vector similarity and the appearance vector similarity threshold, it is determined whether the user is a grabbing user.

[0011] Step four: when there is a grabbing user, it is determined whether there is a unique characteristic user among all grabbing users. If there is a unique characteristic user, it is determined that the sound signal belongs to the unique characteristic user. If there is no unique characteristic user, a dynamic voiceprint characteristic extraction sequence is customized, and the voiceprint characteristics are sequentially extracted according to the dynamic voiceprint characteristic extraction sequence.

[0012] Step five: when there is no grabbing user, the next voiceprint characteristic of the sound signal is extracted according to the voiceprint characteristic extraction sequence.

[0013] Further, the customization process of the general voiceprint characteristic extraction sequence is as follows: a user voiceprint characteristic memory library is constructed, the characteristic mutual difference highlight index of each type of voiceprint characteristic in all sound signals before the current time of the system is obtained, all types of voiceprint characteristics are sequentially sorted according to the value of the characteristic mutual difference highlight index from large to small, and the sorting sequence is labeled as the general voiceprint characteristic extraction sequence.

[0014] Further, the acquisition process of the characteristic mutual difference highlight index of a type of voiceprint characteristic is as follows: a type of voiceprint characteristic is selected, the feature vectors of the type of voiceprint characteristic of all users in the user voiceprint characteristic memory library are obtained, the feature vectors of the type of voiceprint characteristic of each two users are combined into a voiceprint characteristic comparison combination, and then multiple voiceprint characteristic comparison combinations are generated. The characteristic vector similarity of each voiceprint characteristic comparison combination is obtained, the characteristic vector similarities of all voiceprint characteristic comparison combinations are summed and averaged to calculate the average value of the characteristic vector similarity , the appearance vector similarity threshold is set, when the characteristic vector similarity of a voiceprint characteristic comparison combination is lower than the appearance vector similarity threshold, the voiceprint characteristic comparison combination is marked as a voiceprint characteristic mutual difference combination, the total number of voiceprint characteristic mutual difference combinations is compared with the total number of voiceprint characteristic comparison combinations to calculate the ratio of voiceprint characteristic mutual difference , the characteristic mutual difference highlight index of the type of voiceprint characteristic is calculated by . 、 are characteristic weight coefficients.

[0015] Further, the user voiceprint characteristic memory bank contains each voiceprint characteristic of all users, and the construction process is as follows: the voiceprint characteristic type stored in the explicit user voiceprint characteristic memory bank is determined, the valid speech of each user is collected through a microphone, each type of voiceprint characteristic of the valid speech of each user is extracted, the feature vector of each type of voiceprint characteristic of each user is obtained, a user-voiceprint characteristic storage structure is adopted, each user corresponds to an independent storage unit, and thus the user voiceprint characteristic memory bank is constructed.

[0016] Further, when there is a captured user, it is determined whether there is a unique characteristic user among all captured users, and the specific process is as follows: the extraction characteristic similarity of each captured user is obtained, the voiceprint characteristic collocation index of each captured user is further obtained, the first two captured users with high voiceprint characteristic collocation index values are marked as first users, the absolute difference of the voiceprint characteristic collocation indexes of the two first users is calculated, the voiceprint characteristic collocation gap value is calculated, the voiceprint characteristic collocation gap threshold is set, when the voiceprint characteristic collocation gap value is higher than the voiceprint characteristic collocation gap threshold, the first user with the highest voiceprint characteristic collocation index value is marked as the unique characteristic user, and when the voiceprint characteristic collocation gap value is not higher than the voiceprint characteristic collocation gap threshold, it is determined that there is no unique characteristic user among all captured users.

[0017] Further, the extraction characteristic similarity of the captured user is obtained in the following process: a captured user is selected, the characteristic vector similarity of each extraction characteristic of the captured user is obtained, the sum average value of the characteristic vector similarities of all extraction characteristics is calculated, and the extraction characteristic similarity of the captured user is calculated.

[0018] Further, the voiceprint characteristic collocation index of the captured user is obtained in the following process: a captured user is selected, the extraction characteristic similarity of the captured user is obtained, the remaining captured users are marked as comparison users, the captured user is compared with each comparison user, the absolute difference of the extraction characteristic similarities of the compared captured user and comparison user is calculated, a plurality of extraction characteristic similarity difference values are calculated, the sum average value of all extraction characteristic similarity difference values is calculated, the average extraction characteristic similarity difference value is calculated, the extraction characteristic similarity of the captured user and the average extraction characteristic similarity difference value are summed, and the voiceprint characteristic collocation index of the captured user is calculated.

[0019] Further, the customization process of the dynamic voiceprint characteristic extraction sequence is as follows: the characteristic dynamic optimization capture index of the remaining type of voiceprint characteristic is obtained, all voiceprint characteristics are sorted in descending order according to the value of the characteristic dynamic optimization capture index, and the sorted sequence is marked as the dynamic voiceprint characteristic extraction sequence.

[0020] Further, the acquisition process of the dynamic optimization capture index of a type of voiceprint representation is as follows: selecting a type of voiceprint representation, acquiring the feature vectors of each capture user for the type of voiceprint representation from the user voiceprint representation memory bank, combining the feature vectors of each two capture users for the type of voiceprint representation into a target representation comparison combination, acquiring the representation vector similarity of each target representation comparison combination, further determining the individual comparison difference index of each capture user, setting an individual comparison difference threshold, when the individual comparison difference index of the capture user is lower than the individual comparison difference threshold, marking the capture user as a representation independent user, when the individual comparison difference index of the capture user is not lower than the individual comparison difference threshold, marking the capture user as a non-representation independent user, performing ratio calculation on the total number of non-representation independent users and the total number of representation independent users, calculating the non-representation ratio PPdc, performing sum average calculation on the individual comparison difference indexes of all non-representation independent users, calculating the non-representation independent difference index FCGA, acquiring n general voiceprint representation extraction sequences updated before the current time of the system, acquiring the sorting sequence number of the type of voiceprint representation in each general voiceprint representation extraction sequence, performing sum average calculation on all sorting sequence numbers of the type of voiceprint representation, calculating the average sorting sequence number Wac, performing pairwise comparison on all sorting sequence numbers of the type of voiceprint representation, performing absolute difference calculation on the two compared sorting sequence numbers, calculating the sorting sequence number fluctuation value, performing sum average calculation on all sorting sequence number fluctuation values, calculating the average sorting sequence number fluctuation value Ydd, and calculating the dynamic optimization capture index of the type of voiceprint representation by wherein h1 is a first auxiliary coefficient and h2 is a second auxiliary coefficient.

[0021] The determination process of the individual comparison difference index of the capture user is as follows: selecting a capture user, acquiring all target representation comparison combinations containing the capture user and marking all target representation comparison combinations as individual comparison combinations, performing sum average calculation on the representation vector similarities of all individual comparison combinations, and calculating the individual comparison difference index of the capture user.

[0022] Further, the sound pickup recognition system based on the voiceprint representation memory includes a general voiceprint representation sequence module, a capture user determination module, and a sound pickup recognition analysis module.

[0023] Compared with the prior art, the present application has the following beneficial effects:

[0024] ​The system and method of the present application ensure the general sequence to preferentially extract high-discrimination features by characterizing different voiceprint features to quantify the discrimination ability of users, improve the screening efficiency of initial identification from the source, and regularly update the general sequence to continuously optimize the long-term changes of user voiceprints (such as age and state differences) and the expansion of the user scale of the memory bank, avoiding the recognition accuracy decay caused by the insufficient scene adaptability of the traditional fixed sequence.

[0025] When there are multiple users to be captured but no unique user is found, a customized dynamic voiceprint feature extraction sequence is used to accurately focus on the features that have the strongest discrimination ability for the current candidate user and preferentially extract high-value information. This targeted extraction strategy not only reduces the resource consumption of invalid feature calculation, but also quickly narrows down the candidate range, effectively improving the recognition accuracy and efficiency in a multi-user mixed scenario.

[0026] Through the setting of the general voiceprint feature extraction sequence and the dynamic voiceprint feature extraction sequence, in the voiceprint feature memory identification process, the combination of the two avoids the resource waste of the traditional technology of "one-time extraction of all features", overcomes the limitations of "fixed sequence cannot cope with real-time differences", and further reduces the calculation cost by cooperating with the early termination mechanism (stopping when the unique user is identified), ultimately achieving comprehensive improvement of recognition robustness, efficiency, and accuracy in complex interference environments. BRIEF DESCRIPTION OF DRAWINGS

[0027] Fig. 1 The principle flowchart of the voiceprint feature memory-based sound pickup identification method;

[0028] Fig. 2 The principle diagram of constructing the user voiceprint feature memory bank;

[0029] Fig. 3 The principle flowchart of determining whether there is a unique user. DETAILED DESCRIPTION

[0030] Embodiment one: reference Figs. 1 to 3 The sound pickup identification system based on voiceprint feature memory includes a general voiceprint feature sequence module, a user capture determination module, and a sound pickup identification analysis module.

[0031] Step one: customize the general voiceprint feature extraction sequence and update the general voiceprint feature extraction sequence based on a fixed period (the update duration of the fixed period is set and adjusted according to the mobility of the user).

[0032] Step 2: Whenever a sound signal is collected, the sound signal is sequentially extracted according to the general voiceprint representation extraction sequence. (Before extraction, the sound signal needs to be preprocessed. Preprocessing methods include noise suppression (using algorithms such as adaptive filtering and wavelet threshold denoising to eliminate environmental noise and device background noise, especially for low signal-to-noise ratio scenarios, highlighting the spectral characteristics of human voice signals) and endpoint detection (distinguishing between "effective speech segments" (containing voiceprint information) and "non-speech segments" (such as silence and sudden noise) by using the short-time energy and zero-crossing rate of the sound signal, retaining only the speech segments for subsequent processing, reducing redundant calculations). For each voiceprint representation extracted, the voiceprint representation is marked as the extracted representation, and the feature vector of the extracted representation in the sound signal is obtained.

[0033] Step 3: Obtain the feature vector of the extracted representation for each user in the user voiceprint representation memory bank, and calculate the similarity between the feature vector of the extracted representation and the feature vector of the extracted representation for each user. Further obtain the representation vector similarity of the extracted representation for each user, and mark users whose representation vector similarity is higher than the representation vector similarity threshold as crawling users (if it is not higher, no marking is required).

[0034] Step 4: When there is a captured user, determine whether there is a unique representative user among all captured users. If there is a unique representative user, determine that the sound signal belongs to the unique representative user (and stop the current sound pickup and recognition process). If there is no unique representative user, customize the dynamic voiceprint representation extraction sequence and extract the voiceprint representation sequentially according to the dynamic voiceprint representation extraction sequence.

[0035] Step 5: If no user is captured, continue to extract features from the next voiceprint representation of the sound signal according to the voiceprint representation extraction sequence.

[0036] The customization process of the general voiceprint representation extraction sequence is as follows: construct a user voiceprint representation memory, obtain the representation dissimilarity index of each type of voiceprint representation in all sound signals before the current system time, sort all types of voiceprint representations in descending order of the value of the representation dissimilarity index, and label the sorted sequence as the general voiceprint representation extraction sequence. The voiceprint representation includes MFCC, differential MFCC, spectral entropy, short-time energy, fundamental frequency, etc.

[0037] The process of customizing the dynamic voiceprint representation extraction sequence is as follows: Obtain the dynamic optimization capture index of the other types of voiceprint representations (except for the current and previous extracted representations, for example: if the general voiceprint representation extraction sequence is MFCC→Differential MFCC→Spectral Entropy→Short-time Energy→Fundamental Frequency, if the extracted representations already included MFCC, Differential MFCC and Spectral Entropy, then only the dynamic optimization capture index of the short-time energy and fundamental frequency voiceprint representations need to be obtained). Sort all voiceprint representations in descending order of the value of the dynamic optimization capture index, and label the sorted sequence as the dynamic voiceprint representation extraction sequence.

[0038] The process of obtaining the dynamic optimization crawling index for a type of voiceprint representation is as follows: Select a type of voiceprint representation, retrieve the feature vectors of each crawling user for that type of voiceprint representation from the user voiceprint representation memory, combine the feature vectors of every two crawling users for that type of voiceprint representation into a target representation comparison combination, and obtain the representation vector similarity of each target representation comparison combination (the acquisition process is the same as the acquisition process of the representation vector similarity of the voiceprint representation comparison combination, and will not be repeated here). Further determine the individual comparison difference index for each crawling user, and set an individual comparison difference threshold (the individual comparison difference threshold is designed based on the statistical distribution of historical data). When the individual comparison difference index of a crawling user is lower than the individual comparison difference threshold, the crawling user is marked as a representation-independent user; when the individual comparison difference index of a crawling user is not lower than the individual comparison difference threshold, the crawling user is marked as a non-representation-independent user. The total number of non-representation-independent users is then compared with the number of representation-independent users. The total number of users is used to calculate the non-representational ratio PPdc. The individual contrast difference indices of all non-representational independent users are summed and averaged to calculate the non-representational independent difference index FCGA. The n general voiceprint representation extraction sequences updated before the current system time are obtained. The ranking number of the voiceprint representation of that type in each general voiceprint representation extraction sequence is obtained (in the general voiceprint representation extraction sequence MFCC→Differential MFCC→Spectral Entropy→Short-Time Energy→Fundamental Frequency, the ranking number of MFCC is 1, the ranking number of Differential MFCC is 2, and so on). The average ranking number Wac is calculated by summing and averaging all ranking numbers of that type of voiceprint representation. All ranking numbers of that type of voiceprint representation are compared pairwise, and the absolute difference between the two compared ranking numbers is calculated to calculate the ranking number fluctuation value. The average ranking number fluctuation value Ydd is calculated by summing and averaging all ranking number fluctuation values. Calculate the dynamic optimization index for capturing this type of voiceprint representation. Where h1 is the first auxiliary coefficient and h2 is the second auxiliary coefficient, h1+h2=1. Since the long-term position of the historical sorting number is more important, the value of h1 can be 0.7 and the value of h2 can be 0.3.

[0039] The process for determining the individual contrast difference index of a crawled user is as follows: Select a crawled user, obtain all target representation comparison combinations that contain that crawled user (e.g., target representation comparison combination 1 consists of the voiceprint representations of user a and user b; if user a is selected, then target representation comparison combination 1 is obtained), and label them all as individual contrast combinations. Calculate the average sum of the representation vector similarities of all individual contrast combinations to obtain the individual contrast difference index of the crawled user.

[0040] When there are crawled users, determine whether there is a unique representative user among all crawled users. The specific process is as follows: Obtain the extracted representation similarity of each crawled user, and further obtain the voiceprint representation matching index of each crawled user. Mark the two crawled users with the highest voiceprint representation matching index values ​​as first-time users. Calculate the absolute difference between the voiceprint representation matching indices of the two first-time users to obtain the voiceprint representation matching gap value. Set a voiceprint representation matching gap threshold (the voiceprint representation matching gap threshold is designed based on the statistical distribution of historical data). When the voiceprint representation matching gap value is higher than the voiceprint representation matching gap threshold, mark the first-time user with the highest voiceprint representation matching index value as a unique representative user. When the voiceprint representation matching gap value is not higher than the voiceprint representation matching gap threshold, determine that there is no unique representative user among all crawled users.

[0041] The process of obtaining the extracted representation similarity of crawled users is as follows: Select a crawled user, obtain the representation vector similarity of each extracted representation of the crawled user (including the current and previous ones), sum and average the representation vector similarities of all extracted representations, and calculate the extracted representation similarity of the crawled user.

[0042] The process of obtaining the voiceprint representation matching index of a crawled user is as follows: Select a crawled user, obtain the extracted representation similarity of the crawled user, mark the remaining crawled users as comparison users, compare the crawled user with each comparison user, calculate the absolute difference between the extracted representation similarity of the crawled user and the comparison user, calculate multiple extracted representation similarity differences, sum and average all extracted representation similarity differences to calculate the average extracted representation similarity difference, sum the extracted representation similarity of the crawled user with the average extracted representation similarity difference to calculate the voiceprint representation matching index of the crawled user.

[0043] The process of obtaining the representation dissimilarity index of a type of voiceprint representation is as follows: Select a type of voiceprint representation, obtain the feature vectors of that type of voiceprint representation for all users in the user's voiceprint representation memory (e.g., if MFCC is selected, obtain the MFCC feature vectors of all users in the user's voiceprint representation memory), combine the feature vectors of that type of voiceprint representation for every two users into a voiceprint representation comparison combination, and then generate multiple voiceprint representation comparison combinations. Obtain the representation vector similarity of each voiceprint representation comparison combination (taking MFCC as an example, combine the MFCC feature vectors of user a and user b into a voiceprint representation comparison combination, label user a's MFCC feature vector as A, and user b's MFCC feature vector as B, and obtain the similarity using the formula...). The similarity of the representation vectors of all voiceprint representation comparison combinations is calculated by summing and averaging the similarity of the representation vectors of all voiceprint representation comparison combinations, and the mean of the face similarity is calculated. Set a similarity threshold for the representation vector (the similarity threshold is designed based on the statistical distribution of historical data; for example, if collecting MFCC similarity from 1000 different users, take the "75th percentile" or "80th percentile" of the distribution as the threshold—ensuring that the similarity of more than 90% of real different users is below this threshold, thus avoiding "mistakenly classifying different users as similar" from a data perspective). When the representation vector similarity of a voiceprint representation comparison combination is lower than the representation vector similarity threshold, the voiceprint representation comparison combination is marked as a voiceprint representation dissimilar combination (if it is not lower, no marking is required). Calculate the ratio of the total number of voiceprint representation dissimilar combinations to the total number of voiceprint representation comparison combinations to determine the proportion of dissimilar representations. ,pass Calculate the distinctiveness index of the voiceprint representation for this type; , All are characterizing weight coefficients; because The proportion of users whose voiceprint representation can effectively distinguish should be given higher weight. It is a supplementary indicator (reflecting the overall trend), and its weight can be slightly lower, therefore The value can be 0.6. The value can be 0.4.

[0044] The user voiceprint representation memory contains the voiceprint features of every user. The construction process is as follows: The user voiceprint representation memory must store the following voiceprint representation types (including but not limited to MFCC (time-frequency domain features extracted based on the nonlinear perception of frequency by the human ear (Mel scale), reflecting the spectral envelope of sound), differential MFCC (first / second-order difference of MFCC, reflecting the rate of change of MFCC between adjacent frames, capturing the dynamic transition of sound in the time dimension), spectral entropy (frequency domain features, by quantifying the "disorder" of the spectral distribution, distinguishing between "regular human voice spectrum" and "chaotic noise spectrum," enhancing anti-interference ability), short-time energy (time domain features, by calculating the energy value of each frame of speech signal, reflecting the change in sound intensity), and fundamental frequency (the continuous change curve of fundamental frequency (F0) over time, directly reflecting the rise and fall of pitch)). Effective speech from each user is collected via microphone (e.g., needle-like...). For user a, valid speech is collected in a quiet indoor environment. Silence and truncated samples are removed using endpoint detection (VAD) to ensure that the percentage of valid speech is ≥90%. For each user's valid speech, various types of voiceprint representations are extracted, resulting in feature vectors for each type of voiceprint representation. (Frame-level features of each voiceprint representation are extracted: MFCC (e.g., 13-dimensional / frame), Differential MFCC (13-dimensional / frame), Spectral Entropy (1-dimensional / frame), Short-Time Energy (1-dimensional / frame), and Fundamental Frequency (1-dimensional / frame). The frame sequences of each type of voiceprint representation are aggregated into segment-level features: variable-length frame sequences are converted to fixed dimensions (e.g., MFCC is calculated by taking the 13-dimensional mean, resulting in 13-dimensional segment features, which are the features of MFCC)) A user-voiceprint representation storage structure is adopted, with each user corresponding to an independent storage unit, thus constructing a user voiceprint representation memory bank.

[0045] Example 2: A voiceprint representation memory-based voice recognition system, including a general voiceprint representation sequence module, a user capture and determination module, and a voice recognition and analysis module.

[0046] The Universal Voiceprint Representation Sequence Module is used to customize the universal voiceprint representation extraction sequence and update the universal voiceprint representation extraction sequence based on a fixed period.

[0047] The user capture and determination module extracts voiceprint representations sequentially according to a general voiceprint representation extraction sequence whenever a sound signal is collected. For each extracted voiceprint representation, it is marked as an extracted representation. The feature vector of the extracted representation in the sound signal is obtained, and the feature vectors of the extracted representation for each user in the user voiceprint representation memory are obtained. The similarity between the extracted feature vector and the feature vector of the extracted representation for each user is calculated, and the representation vector similarity for the extracted representation for each user is further obtained. Users whose representation vector similarity is higher than the representation vector similarity threshold are marked as captured users.

[0048] The voice recognition and analysis module, when a user is captured, determines whether there is a unique user among all captured users. If a unique user exists, it determines that the sound signal belongs to that unique user (and stops the current voice recognition process). If no unique user exists, it customizes a dynamic voiceprint representation extraction sequence and extracts voiceprint representations sequentially according to the dynamic voiceprint representation extraction sequence. When no user is captured, it continues to extract features from the next voiceprint representation of the sound signal according to the voiceprint representation extraction sequence.

[0049] The above formulas are all dimensionless calculations, and the preset parameters in the formulas should be set by those skilled in the art according to the actual situation.

[0050] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.

[0051] It should be understood that in the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0052] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0053] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0054] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0055] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0056] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A voice recognition method based on voiceprint representation memory, characterized in that, Includes the following steps: Step 1: Customize a general voiceprint representation extraction sequence that can be updated regularly; Step 2: Whenever a sound signal is collected, the sound signal is sequentially extracted according to the general voiceprint representation extraction sequence. Each extracted voiceprint representation is marked as the extracted representation, and the feature vector of the extracted representation in the sound signal is obtained. Step 3: Obtain the feature vector of the extracted representation for each user in the user voiceprint representation memory bank, and further obtain the representation vector similarity of the extracted representation for each user. Based on the comparison results of the representation vector similarity and the representation vector similarity threshold, determine whether the user is the crawling user. Step 4: When there are captured users, determine whether there is a unique user among all captured users. If there is a unique user, determine that the sound signal belongs to the unique user. If there is no unique user, customize the dynamic voiceprint representation extraction sequence and extract the voiceprint representation sequentially according to the dynamic voiceprint representation extraction sequence. The customization process of dynamic voiceprint representation extraction sequence is as follows: obtain the dynamic optimization capture index of the other types of voiceprint representations in the general voiceprint representation extraction sequence, excluding the extracted representations that have already been extracted; sort all voiceprint representations in descending order of the value of the dynamic optimization capture index; and label the sorted sequence as dynamic voiceprint representation extraction sequence. The process of obtaining the dynamic optimization crawling index for a type of voiceprint representation is as follows: Select a type of voiceprint representation; retrieve the feature vectors of each crawling user for that type of voiceprint representation from the user voiceprint representation memory; combine the feature vectors of every two crawling users for that type of voiceprint representation into a target representation comparison combination; obtain the representation vector similarity of each target representation comparison combination; further determine the individual contrast difference index for each crawling user; set an individual contrast difference threshold; when the individual contrast difference index of a crawling user is lower than the individual contrast difference threshold, the crawling user is marked as a representation-independent user; when the individual contrast difference index of a crawling user is not lower than the individual contrast difference threshold, the crawling user is marked as a non-representation-independent user; and then compare the total number of non-representation-independent users with... The non-representation ratio (PPdc) is calculated by comparing the total number of unique users. The non-representation independent difference index (FCGA) is calculated by summing and averaging the individual comparison differences of all non-representation unique users. The n general voiceprint representation extraction sequences updated before the current system time are obtained. The ranking number of the voiceprint representation of that type in each general voiceprint representation extraction sequence is obtained. The average ranking number (Wac) is calculated by summing and averaging all ranking numbers of that type of voiceprint representation. All ranking numbers of that type of voiceprint representation are compared pairwise, and the absolute difference between the two compared ranking numbers is calculated to obtain the ranking number fluctuation value. The average ranking number fluctuation value (Ydd) is calculated by summing and averaging all ranking number fluctuation values. Calculate the dynamic optimization index for capturing this type of voiceprint representation. Where h1 is the first auxiliary coefficient and h2 is the second auxiliary coefficient; The process of determining the individual contrast difference index of crawled users is as follows: Select a crawled user, obtain all target representation comparison combinations containing that crawled user and label them as individual contrast combinations, sum and average the representation vector similarity of all individual contrast combinations, and calculate the individual contrast difference index of the crawled user. Step 5: If no user is captured, continue to extract features from the next voiceprint representation of the sound signal according to the voiceprint representation extraction sequence.

2. The voice recognition method based on voiceprint representation memory according to claim 1, characterized in that, The customization process of the general voiceprint representation extraction sequence is as follows: construct a user voiceprint representation memory, obtain the representation dissimilarity index of each type of voiceprint representation in all sound signals before the current system time, sort all types of voiceprint representations in descending order of the representation dissimilarity index value, and label the sorted sequence as the general voiceprint representation extraction sequence.

3. The voice recognition method based on voiceprint representation memory according to claim 2, characterized in that, The process of obtaining the representation dissimilarity index of a type of voiceprint representation is as follows: Select a type of voiceprint representation, obtain the feature vectors of the voiceprint representation of that type for all users in the user voiceprint representation memory bank, combine the feature vectors of each pair of users' voiceprint representations of that type into a voiceprint representation comparison combination, and then generate multiple voiceprint representation comparison combinations. Obtain the representation vector similarity of each voiceprint representation comparison combination, and sum and average the representation vector similarities of all voiceprint representation comparison combinations to calculate the mean of representation similarity. A similarity threshold for representation vectors is set. When the similarity of the representation vectors of a voiceprint representation comparison combination is lower than the threshold, the voiceprint representation comparison combination is marked as a voiceprint representation dissimilar combination. The ratio of the total number of voiceprint representation dissimilar combinations to the total number of voiceprint representation comparison combinations is calculated to determine the proportion of representation dissimilarities. ,pass Calculate the distinctiveness index of the voiceprint representation for this type; , All of these are character weight coefficients.

4. The voice recognition method based on voiceprint representation memory according to claim 2, characterized in that, The user voiceprint representation memory contains the voiceprint features of all users. The construction process is as follows: Determine the types of voiceprint representations that the user voiceprint representation memory needs to store, collect the effective speech of each user through a microphone, extract the voiceprint representations of each type from the effective speech of each user, and then obtain the feature vectors of each type of voiceprint representation for each user. Adopt the user-voiceprint representation storage structure, with each user corresponding to an independent storage unit, and then construct the user voiceprint representation memory.

5. The voice recognition method based on voiceprint representation memory according to claim 1, characterized in that, When there are crawled users, determine whether there is a unique representative user among all crawled users. The specific process is as follows: obtain the extracted representation similarity of each crawled user, further obtain the voiceprint representation matching index of each crawled user, mark the two crawled users with the highest voiceprint representation matching index as first users, calculate the absolute difference of the voiceprint representation matching indices of the two first users, calculate the voiceprint representation matching gap value, set the voiceprint representation matching gap threshold, when the voiceprint representation matching gap value is higher than the voiceprint representation matching gap threshold, mark the first user with the highest voiceprint representation matching index as the unique representative user, when the voiceprint representation matching gap value is not higher than the voiceprint representation matching gap threshold, determine that there is no unique representative user among all crawled users; The process of obtaining the voiceprint representation matching index of a crawled user is as follows: Select a crawled user, obtain the extracted representation similarity of the crawled user, mark the remaining crawled users as comparison users, compare the crawled user with each comparison user, calculate the absolute difference between the extracted representation similarity of the crawled user and the comparison user, calculate multiple extracted representation similarity differences, sum and average all extracted representation similarity differences to calculate the average extracted representation similarity difference, sum the extracted representation similarity of the crawled user with the average extracted representation similarity difference to calculate the voiceprint representation matching index of the crawled user.

6. The voice recognition method based on voiceprint representation memory according to claim 5, characterized in that, The process of obtaining the extracted representation similarity of crawled users is as follows: Select a crawled user, obtain the representation vector similarity of each extracted representation of the crawled user, sum and average the representation vector similarities of all extracted representations, and calculate the extracted representation similarity of the crawled user.

7. A voiceprint-based recognition system, applied to the voiceprint-based recognition method according to any one of claims 1-6, characterized in that, It includes a general voiceprint representation sequence module, a user capture and determination module, and a voice pickup and recognition analysis module; The Universal Voiceprint Representation Sequence Module is used to customize the universal voiceprint representation extraction sequence and update the universal voiceprint representation extraction sequence based on a fixed period. The user capture and determination module extracts voiceprint representations sequentially according to the general voiceprint representation extraction sequence whenever a sound signal is collected. For each extracted voiceprint representation, it marks it as an extracted representation, obtains the feature vector of the extracted representation in the sound signal, obtains the feature vector of the extracted representation of each user in the user voiceprint representation memory, and calculates the similarity between the feature vector of the extracted representation and the feature vector of the extracted representation of each user. It further obtains the representation vector similarity of the extracted representation of each user, and marks users whose representation vector similarity is higher than the representation vector similarity threshold as captured users. The voice recognition and analysis module, when a user is captured, determines whether there is a unique user among all captured users. If a unique user exists, the sound signal is determined to belong to that unique user. If no unique user exists, a customized dynamic voiceprint extraction sequence is used to extract voiceprint features sequentially according to the dynamic voiceprint extraction sequence. When no user is captured, the next voiceprint feature of the sound signal is extracted according to the voiceprint extraction sequence.

Citation Information

Patent Citations

  • Streaming speech recognition method and system, equipment and storage medium

    CN118038874A

  • Newborn cry recognition system based on machine learning

    CN119920261A