Pickup recognition method and system based on voiceprint representation memory

By employing a voiceprint-based recognition method that utilizes customized and dynamic feature extraction sequences, the problems of resource waste and insufficient adaptability in existing technologies are solved, achieving efficient and low-cost voiceprint recognition.

CN121075338AActive Publication Date: 2025-12-05GUANGZHOU SIZHENG ELECTRONIC TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511625359.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-07
Publication Date
2025-12-05
Estimated Expiration
2045-11-07

AI Technical Summary

Technical Problem

Existing voiceprint recognition technologies have shortcomings in terms of resource consumption and recognition efficiency. In particular, they are difficult to meet the requirements of real-time performance and low power consumption in embedded devices and mobile terminals. Furthermore, fixed feature extraction sequences cannot adapt to dynamic changes in scenarios and differences in user groups, resulting in low recognition efficiency.

Method used

A voice recognition method based on voiceprint representation memory is adopted. By customizing a general voiceprint representation extraction sequence and a dynamic sound feature extraction sequence that can be updated regularly, the user identity is determined according to the similarity of user feature vectors in the voiceprint representation memory and the similarity threshold, and high-value information is extracted first in multi-user scenarios.

Benefits of technology

It improves the efficiency of initial identification and screening, reduces resource consumption, quickly narrows down the candidate range, improves identification accuracy and efficiency, reduces computational costs, and achieves robust identification under complex interference environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121075338A_ABST
    Figure CN121075338A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of pickup recognition, in particular to a pickup recognition method and system based on voiceprint characterization memory.The system and method are characterized in that a general voiceprint characterization extraction sequence and a dynamic sound feature extraction sequence are set, and in the pickup recognition process of voiceprint characterization memorization, the voiceprint characterization memorization is carried out; the combination of the two not only avoids the resource waste of'one-time extraction of full features' in the traditional technology, but also overcomes the limitation that'a fixed sequence cannot cope with real-time differences', and further reduces the calculation cost by cooperating with an early termination mechanism (stopping when a unique user is identified). And finally, comprehensive improvement of recognition robustness, efficiency and precision in a complex interference environment is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of voice recognition technology, and more specifically, to a voice recognition method and system based on voiceprint representation memory. Background Technology

[0002] Current voiceprint recognition technologies generally adopt a "full extraction of all voiceprint representations" approach. This means that for the input sound signal, all voiceprint representations, including MFCC, differential MFCC, spectral entropy, short-time energy, and fundamental frequency, are extracted without discrimination before subsequent comparative analysis. However, different voiceprint representations have significantly different effectiveness in "distinguishing users." Some representations have weak distinguishing ability in the current scenario or user group, yet they still require computational and storage resources for extraction and calculation. This not only wastes resources in the feature extraction stage but also directly reduces the overall efficiency of voice recognition due to redundant calculations, making it particularly difficult to meet the needs of real-time and low-power scenarios (such as embedded devices and mobile terminals).

[0003] Even when some technologies employ "fixed feature extraction sequences," they can only mechanically extract voiceprint representations according to preset priorities, failing to adapt to dynamic changes in scenarios and differences among user groups. On the one hand, the distinguishing effectiveness of voiceprint representations fluctuates with changes in user status (such as health and mood) and even the size of the user group, making it difficult for fixed sequences to respond to these dynamic factors. On the other hand, in scenarios with multiple candidate users (captured users), fixed sequences lack specificity regarding "feature differences among current candidate users," still extracting all or fixed priorities. This results in the failure to prioritize the use of high-value features that are more critical for differentiation, exacerbating resource consumption and making it difficult to quickly narrow down the candidate range, further restricting the improvement of recognition efficiency.

[0004] Based on the above, this application proposes a voice recognition method and system based on voiceprint representation memory. Summary of the Invention

[0005] In view of the shortcomings of the existing technology, the purpose of this invention is to provide a voice recognition method and system based on voiceprint representation memory.

[0006] To achieve the above objectives, the present invention provides the following technical solution: The voiceprint-based recognition method includes the following steps: Step 1: Customize a general voiceprint representation extraction sequence that can be updated regularly; Step 2: Whenever a sound signal is collected, the sound signal is sequentially extracted according to the general voiceprint representation extraction sequence. Each extracted voiceprint representation is marked as the extracted representation, and the feature vector of the extracted representation in the sound signal is obtained. Step 3: Obtain the feature vector of the extracted representation for each user in the user voiceprint representation memory bank, and further obtain the representation vector similarity of the extracted representation for each user. Based on the comparison results of the representation vector similarity and the representation vector similarity threshold, determine whether the user is the crawling user. Step 4: When there are captured users, determine whether there is a unique user among all captured users. If there is a unique user, determine that the sound signal belongs to the unique user. If there is no unique user, customize a dynamic sound feature extraction sequence and extract the voiceprint representation sequentially according to the dynamic sound feature extraction sequence. Step 5: If no user is captured, continue to extract the next sound feature of the sound signal according to the sound feature extraction sequence.

[0007] Furthermore, the customization process of the general voiceprint representation extraction sequence is as follows: construct a user voiceprint representation memory bank, obtain the representation dissimilarity index of each type of voiceprint representation in all sound signals before the current system time, sort all types of voiceprint representations in descending order of the representation dissimilarity index value, and label the sorted sequence as the general voiceprint representation extraction sequence.

[0008] Furthermore, the process of obtaining the representation dissimilarity salience index of a type of voiceprint representation is as follows: Select a type of voiceprint representation, obtain the feature vectors of the same type of voiceprint representation for all users in the user voiceprint representation memory bank, combine the feature vectors of each pair of users' voiceprint representations of the same type into a voiceprint representation comparison combination, and then generate multiple voiceprint representation comparison combinations. Obtain the representation vector similarity of each voiceprint representation comparison combination, and sum and average the representation vector similarities of all voiceprint representation comparison combinations to calculate the mean of representation similarity. A similarity threshold for representation vectors is set. When the similarity of the representation vectors of a voiceprint representation comparison combination is lower than the threshold, the voiceprint representation comparison combination is marked as a voiceprint representation dissimilar combination. The ratio of the total number of voiceprint representation dissimilar combinations to the total number of voiceprint representation comparison combinations is calculated to determine the proportion of representation dissimilarities. ,pass Calculate the distinctiveness index of the voiceprint representation for this type; , All of these are character weight coefficients.

[0009] Furthermore, the user voiceprint representation memory contains the voiceprint features of all users. The construction process is as follows: the voiceprint representation types to be stored in the user voiceprint representation memory are defined; valid speech of each user is collected through a microphone; voiceprint representations of each type are extracted from the valid speech of each user; feature vectors of each type of voiceprint representation of each user are obtained; a user-voiceprint representation storage structure is adopted, with each user corresponding to an independent storage unit, thereby constructing the user voiceprint representation memory.

[0010] Furthermore, when there are crawled users, it is determined whether there is a unique representative user among all crawled users. The specific process is as follows: obtain the extracted representation similarity of each crawled user, further obtain the voiceprint representation matching index of each crawled user, mark the top two crawled users with the highest voiceprint representation matching index as first-time users, calculate the absolute difference between the voiceprint representation matching indices of the two first-time users, calculate the voiceprint representation matching gap value, set the voiceprint representation matching gap threshold, when the voiceprint representation matching gap value is higher than the voiceprint representation matching gap threshold, mark the first-time user with the highest voiceprint representation matching index as a unique representative user, when the voiceprint representation matching gap value is not higher than the voiceprint representation matching gap threshold, it is determined that there is no unique representative user among all crawled users.

[0011] Furthermore, the process of obtaining the extracted representation similarity of crawled users is as follows: Select a crawled user, obtain the representation vector similarity of each extracted representation of the crawled user, sum and average the representation vector similarities of all extracted representations to calculate the extracted representation similarity of the crawled user.

[0012] Furthermore, the process of obtaining the voiceprint representation matching index of the crawled user is as follows: Select a crawled user, obtain the extracted representation similarity of the crawled user, mark the remaining crawled users as comparison users, compare the crawled user with each comparison user, calculate the absolute difference between the extracted representation similarity of the crawled user and the comparison user, calculate multiple extracted representation similarity differences, sum and average all extracted representation similarity differences to calculate the average extracted representation similarity difference, sum the extracted representation similarity of the crawled user with the average extracted representation similarity difference to calculate the voiceprint representation matching index of the crawled user.

[0013] Furthermore, the customization process of the dynamic voice feature extraction sequence is as follows: obtain the dynamic optimization capture index of the other types of voiceprint representations, sort all voiceprint representations in descending order of the value of the dynamic optimization capture index, and label the sorted sequence as the dynamic voice feature extraction sequence.

[0014] Furthermore, the process of obtaining the dynamic optimization crawling index for a type of voiceprint representation is as follows: Select a type of voiceprint representation; retrieve the feature vectors of each crawling user for that type of voiceprint representation from the user voiceprint representation memory; combine the feature vectors of every two crawling users for that type of voiceprint representation into a target representation comparison combination; obtain the representation vector similarity of each target representation comparison combination; further determine the individual comparison difference index for each crawling user; set an individual comparison difference threshold; when the individual comparison difference index of a crawling user is lower than the individual comparison difference threshold, the crawling user is marked as a representation-independent user; when the individual comparison difference index of a crawling user is not lower than the individual comparison difference threshold, the crawling user is marked as a non-representation-independent user; and calculate the total number of non-representation-independent users. The ratio of the quantity to the total number of independent users is calculated to obtain the non-representation ratio PPdc. The summation and average of the individual contrast differences of all independent users are then calculated to obtain the non-representation independent difference index FCGA. The n general voiceprint representation extraction sequences updated before the current system time are obtained. The ranking number of the voiceprint representation of that type in each general voiceprint representation extraction sequence is obtained. The summation and average of all ranking numbers of that type of voiceprint representation is calculated to obtain the average ranking number Wac. All ranking numbers of that type of voiceprint representation are compared pairwise, and the absolute difference between the two compared ranking numbers is calculated to obtain the ranking number fluctuation value. The summation and average of all ranking number fluctuation values ​​is then calculated to obtain the average ranking number fluctuation value Ydd. Calculate the dynamic optimization index for capturing this type of voiceprint representation. Where h1 is the first auxiliary coefficient and h2 is the second auxiliary coefficient; The process for determining the individual contrast difference index of a crawled user is as follows: Select a crawled user, obtain all target representation comparison combinations that contain that crawled user and label them all as individual contrast combinations, sum and average the representation vector similarities of all individual contrast combinations, and calculate the individual contrast difference index of that crawled user.

[0015] Furthermore, the voiceprint representation memory-based voice recognition system includes a general voiceprint representation sequence module, a user capture and determination module, and a voice recognition and analysis module.

[0016] Compared with the prior art, the present invention has the following beneficial effects: The system and method of the present invention quantify the ability of different voiceprint representations to distinguish users by representing the dissimilarity prominence index, ensuring that the general sequence extracts high-discrimination features first, thereby improving the screening efficiency of the initial recognition from the source. The general sequence is updated regularly and can be continuously optimized with the long-term changes of user voiceprints (such as age and state differences) and the expansion of the memory bank user scale, avoiding the recognition accuracy decay caused by insufficient scene adaptability of traditional fixed sequences. When multiple users exist but no unique representative user is found, a customized dynamic voice feature extraction sequence can accurately focus on the features that have the strongest distinguishing ability for the current candidate user, prioritizing the extraction of high-value information. This targeted extraction strategy not only reduces the resource consumption of invalid feature calculations but also quickly narrows down the candidate range, effectively improving the recognition accuracy and efficiency in multi-user overlapping scenarios. By setting up a general voiceprint representation extraction sequence and a dynamic sound feature extraction sequence, the combination of the two avoids the resource waste of the traditional technology of "extracting all features at once" and overcomes the limitation that "fixed sequences cannot cope with real-time differences". At the same time, with the early termination mechanism (stopping when a unique user is identified), the computational cost is further reduced, and the comprehensive improvement of recognition robustness, efficiency and accuracy in complex interference environments is finally achieved. Attached Figure Description

[0017] Figure 1 This is a flowchart illustrating the principle of the voice recognition method based on voiceprint representation memory. Figure 2 A schematic diagram illustrating the construction principle of the user's voiceprint representation memory database; Figure 3 A flowchart illustrating the principle of determining whether a unique user identifier exists. Detailed Implementation

[0018] Example 1: Refer to Figures 1 to 3 The voiceprint representation memory-based voice recognition system includes: a general voiceprint representation sequence module, a user capture and determination module, and a voice recognition and analysis module.

[0019] Step 1: Customize the general voiceprint representation extraction sequence and update the general voiceprint representation extraction sequence based on a fixed period (the update duration of the fixed period is set and adjusted according to the user's mobility).

[0020] Step 2: Whenever a sound signal is collected, the sound signal is sequentially extracted according to the general voiceprint representation extraction sequence. (Before extraction, the sound signal needs to be preprocessed. Preprocessing methods include noise suppression (using algorithms such as adaptive filtering and wavelet threshold denoising to eliminate environmental noise and device background noise, especially for low signal-to-noise ratio scenarios, highlighting the spectral characteristics of human voice signals) and endpoint detection (distinguishing between "effective speech segments" (containing voiceprint information) and "non-speech segments" (such as silence and sudden noise) by using the short-time energy and zero-crossing rate of the sound signal, retaining only the speech segments for subsequent processing, reducing redundant calculations). For each voiceprint representation extracted, the voiceprint representation is marked as the extracted representation, and the feature vector of the extracted representation in the sound signal is obtained.

[0021] Step 3: Obtain the feature vector of the extracted representation for each user in the user voiceprint representation memory bank, and calculate the similarity between the feature vector of the extracted representation and the feature vector of the extracted representation for each user. Further obtain the representation vector similarity of the extracted representation for each user, and mark users whose representation vector similarity is higher than the representation vector similarity threshold as crawling users (if it is not higher, no marking is required).

[0022] Step 4: When there is a captured user, determine whether there is a unique user among all captured users. If there is a unique user, determine that the sound signal belongs to the unique user (and stop the current sound pickup and recognition process). If there is no unique user, customize a dynamic sound feature extraction sequence and extract the voiceprint representation sequentially according to the dynamic sound feature extraction sequence.

[0023] Step 5: If no user is captured, continue to extract the next sound feature of the sound signal according to the sound feature extraction sequence.

[0024] The customization process of the general voiceprint representation extraction sequence is as follows: construct a user voiceprint representation memory, obtain the representation dissimilarity index of each type of voiceprint representation in all sound signals before the current system time, sort all types of voiceprint representations in descending order of the value of the representation dissimilarity index, and label the sorted sequence as the general voiceprint representation extraction sequence. The voiceprint representation includes MFCC, differential MFCC, spectral entropy, short-time energy, fundamental frequency, etc.

[0025] The process of customizing dynamic sound feature extraction sequences is as follows: Obtain the dynamic optimization capture index of other types of voiceprint representations (in addition to the current and previous extracted representations, for example: if the general voiceprint representation extraction sequence is MFCC→differential MFCC→spectral entropy→short-time energy→fundamental frequency; if the extracted representations already included MFCC, differential MFCC and spectral entropy, then only the dynamic optimization capture index of the short-time energy and fundamental frequency voiceprint representations need to be obtained). Sort all voiceprint representations in descending order of the value of the dynamic optimization capture index, and label the sorted sequence as a dynamic sound feature extraction sequence.

[0026] The process of obtaining the dynamic optimization crawling index for a type of voiceprint representation is as follows: Select a type of voiceprint representation, retrieve the feature vectors of each crawling user for that type of voiceprint representation from the user voiceprint representation memory, combine the feature vectors of every two crawling users for that type of voiceprint representation into a target representation comparison combination, and obtain the representation vector similarity of each target representation comparison combination (the acquisition process is the same as the acquisition process of the representation vector similarity of the voiceprint representation comparison combination, and will not be repeated here). Further determine the individual comparison difference index for each crawling user, and set an individual comparison difference threshold (the individual comparison difference threshold is designed based on the statistical distribution of historical data). When the individual comparison difference index of a crawling user is lower than the individual comparison difference threshold, the crawling user is marked as a representation-independent user; when the individual comparison difference index of a crawling user is not lower than the individual comparison difference threshold, the crawling user is marked as a non-representation-independent user. The total number of non-representation-independent users is then compared with the number of representation-independent users. The total number of users is used to calculate the non-representational ratio PPdc. The individual contrast difference indices of all non-representational independent users are summed and averaged to calculate the non-representational independent difference index FCGA. The n general voiceprint representation extraction sequences updated before the current system time are obtained. The ranking number of the voiceprint representation of that type in each general voiceprint representation extraction sequence is obtained (in the general voiceprint representation extraction sequence MFCC→Differential MFCC→Spectral Entropy→Short-Time Energy→Fundamental Frequency, the ranking number of MFCC is 1, the ranking number of Differential MFCC is 2, and so on). The average ranking number Wac is calculated by summing and averaging all ranking numbers of that type of voiceprint representation. All ranking numbers of that type of voiceprint representation are compared pairwise, and the absolute difference between the two compared ranking numbers is calculated to calculate the ranking number fluctuation value. The average ranking number fluctuation value Ydd is calculated by summing and averaging all ranking number fluctuation values. Calculate the dynamic optimization index for capturing this type of voiceprint representation. Where h1 is the first auxiliary coefficient and h2 is the second auxiliary coefficient, h1+h2=1. Since the long-term position of the historical sorting number is more important, the value of h1 can be 0.7 and the value of h2 can be 0.3.

[0027] The process for determining the individual contrast difference index of a crawled user is as follows: Select a crawled user, obtain all target representation comparison combinations that contain that crawled user (e.g., target representation comparison combination 1 consists of the voiceprint representations of user a and user b; if user a is selected, then target representation comparison combination 1 is obtained), and label them all as individual contrast combinations. Calculate the average sum of the representation vector similarities of all individual contrast combinations to obtain the individual contrast difference index of the crawled user.

[0028] When there are crawled users, determine whether there is a unique representative user among all crawled users. The specific process is as follows: Obtain the extracted representation similarity of each crawled user, and further obtain the voiceprint representation matching index of each crawled user. Mark the two crawled users with the highest voiceprint representation matching index values ​​as first-time users. Calculate the absolute difference between the voiceprint representation matching indices of the two first-time users to obtain the voiceprint representation matching gap value. Set a voiceprint representation matching gap threshold (the voiceprint representation matching gap threshold is designed based on the statistical distribution of historical data). When the voiceprint representation matching gap value is higher than the voiceprint representation matching gap threshold, mark the first-time user with the highest voiceprint representation matching index value as a unique representative user. When the voiceprint representation matching gap value is not higher than the voiceprint representation matching gap threshold, determine that there is no unique representative user among all crawled users.

[0029] The process of obtaining the extracted representation similarity of crawled users is as follows: Select a crawled user, obtain the representation vector similarity of each extracted representation of the crawled user (including the current and previous ones), sum and average the representation vector similarities of all extracted representations, and calculate the extracted representation similarity of the crawled user.

[0030] The process of obtaining the voiceprint representation matching index of a crawled user is as follows: Select a crawled user, obtain the extracted representation similarity of the crawled user, mark the remaining crawled users as comparison users, compare the crawled user with each comparison user, calculate the absolute difference between the extracted representation similarity of the crawled user and the comparison user, calculate multiple extracted representation similarity differences, sum and average all extracted representation similarity differences to calculate the average extracted representation similarity difference, sum the extracted representation similarity of the crawled user with the average extracted representation similarity difference to calculate the voiceprint representation matching index of the crawled user.

[0031] The process of obtaining the representation dissimilarity index of a type of voiceprint representation is as follows: Select a type of voiceprint representation, obtain the feature vectors of that type of voiceprint representation for all users in the user's voiceprint representation memory (e.g., if MFCC is selected, obtain the MFCC feature vectors of all users in the user's voiceprint representation memory), combine the feature vectors of that type of voiceprint representation for every two users into a voiceprint representation comparison combination, and then generate multiple voiceprint representation comparison combinations. Obtain the representation vector similarity of each voiceprint representation comparison combination (taking MFCC as an example, combine the MFCC feature vectors of user a and user b into a voiceprint representation comparison combination, label user a's MFCC feature vector as A, and user b's MFCC feature vector as B, and obtain the similarity using the formula...). The similarity of the representation vectors of all voiceprint representation comparison combinations is calculated by summing and averaging the similarity of the representation vectors of all voiceprint representation comparison combinations, and then calculating the mean similarity of the representation vectors. Set a similarity threshold for the representation vector (the similarity threshold is designed based on the statistical distribution of historical data; for example, if collecting MFCC similarity from 1000 different users, take the "75th percentile" or "80th percentile" of the distribution as the threshold—ensuring that the similarity of more than 90% of real different users is below this threshold, thus avoiding "mistakenly classifying different users as similar" from a data perspective). When the representation vector similarity of a voiceprint representation comparison combination is lower than the representation vector similarity threshold, the voiceprint representation comparison combination is marked as a voiceprint representation dissimilar combination (if it is not lower, no marking is required). Calculate the ratio of the total number of voiceprint representation dissimilar combinations to the total number of voiceprint representation comparison combinations to determine the proportion of dissimilar representations. ,pass Calculate the distinctiveness index of the voiceprint representation for this type; , All are characterizing weight coefficients, because; The proportion of users whose voiceprint representation can effectively distinguish should be given higher weight. It is a supplementary indicator (reflecting the overall trend), and its weight can be slightly lower. The value can be 0.6. The value can be 0.4.

[0032] The user voiceprint representation memory contains the voiceprint features of every user. The construction process is as follows: The user voiceprint representation memory must store the following voiceprint representation types (including but not limited to MFCC (time-frequency domain features extracted based on the nonlinear perception of frequency by the human ear (Mel scale), reflecting the spectral envelope of sound), differential MFCC (first / second-order difference of MFCC, reflecting the rate of change of MFCC between adjacent frames, capturing the dynamic transition of sound in the time dimension), spectral entropy (frequency domain features, by quantifying the "disorder" of the spectral distribution, distinguishing between "regular human voice spectrum" and "chaotic noise spectrum," enhancing anti-interference ability), short-time energy (time domain features, by calculating the energy value of each frame of speech signal, reflecting the change in sound intensity), and fundamental frequency (the continuous change curve of fundamental frequency (F0) over time, directly reflecting the rise and fall of pitch)). Effective speech from each user is collected via microphone (e.g., needle-like...). For user a, valid speech is collected in a quiet indoor environment. Silence and truncated samples are removed using endpoint detection (VAD) to ensure that the percentage of valid speech is ≥90%. For each user's valid speech, various types of voiceprint representations are extracted, resulting in feature vectors for each type of voiceprint representation. (Frame-level features of each voiceprint representation are extracted: MFCC (e.g., 13-dimensional / frame), Differential MFCC (13-dimensional / frame), Spectral Entropy (1-dimensional / frame), Short-Time Energy (1-dimensional / frame), and Fundamental Frequency (1-dimensional / frame). The frame sequences of each type of voiceprint representation are aggregated into segment-level features: variable-length frame sequences are converted to fixed dimensions (e.g., MFCC is calculated by taking the 13-dimensional mean, resulting in 13-dimensional segment features, which are the features of MFCC)) A user-voiceprint representation storage structure is adopted, with each user corresponding to an independent storage unit, thus constructing a user voiceprint representation memory bank.

[0033] Example 2: A voiceprint representation memory-based voice recognition system, including a general voiceprint representation sequence module, a user capture and determination module, and a voice recognition and analysis module.

[0034] The Universal Voiceprint Representation Sequence Module is used to customize the universal voiceprint representation extraction sequence and update the universal voiceprint representation extraction sequence based on a fixed period.

[0035] The user capture and determination module extracts voiceprint representations sequentially according to a general voiceprint representation extraction sequence whenever a sound signal is collected. For each extracted voiceprint representation, it is marked as an extracted representation. The feature vector of the extracted representation in the sound signal is obtained, and the feature vectors of the extracted representation for each user in the user voiceprint representation memory are obtained. The similarity between the extracted feature vector and the feature vector of the extracted representation for each user is calculated, and the representation vector similarity for the extracted representation for each user is further obtained. Users whose representation vector similarity is higher than the representation vector similarity threshold are marked as captured users.

[0036] The voice recognition and analysis module, when a user is captured, determines whether there is a unique user among all captured users. If a unique user exists, it determines that the sound signal belongs to that unique user (and stops the current voice recognition process). If no unique user exists, it customizes a dynamic sound feature extraction sequence and extracts voiceprint representations sequentially according to the dynamic sound feature extraction sequence. When no user is captured, it continues to extract features for the next sound feature of the sound signal according to the sound feature extraction sequence.

[0037] The above formulas are all dimensionless calculations, and the preset parameters in the formulas should be set by those skilled in the art according to the actual situation.

[0038] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.

[0039] It should be understood that in the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0040] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0041] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0042] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0043] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0044] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A sound pickup recognition method based on a voiceprint characteristic memory, characterized in that, The method comprises the following steps: Step 1: customizing a general voiceprint representation extraction sequence that can be updated regularly; Step 2: whenever a voice signal is collected, sequentially extracting voiceprint representations from the collected voice signal according to the general voiceprint representation extraction sequence, marking each extracted voiceprint representation as an extraction representation, and obtaining a feature vector of the extraction representation in the voice signal; Step 3: obtaining feature vectors of the extraction representation of each user in the user voiceprint representation memory, further obtaining a representation vector similarity of the extraction representation of each user, and determining whether the user is a grabbing user according to a comparison result of the representation vector similarity and a representation vector similarity threshold; Step 4: when there is a grabbing user, determining whether there is a unique representation user among all grabbing users, if there is a unique representation user, determining that the voice signal belongs to the unique representation user, if there is no unique representation user, customizing a dynamic voice feature extraction sequence, and sequentially extracting voiceprint representations according to the dynamic voice feature extraction sequence; Step 5: when there is no grabbing user, continuing to extract the next voice feature of the voice signal according to the voice feature extraction sequence.

2. The voiceprint-based memory characterization pickup identification method according to claim 1, characterized in that, The customization process of the general voiceprint representation extraction sequence is as follows: constructing a user voiceprint representation memory, obtaining a representation mutual difference highlight index of each type of voiceprint representation in all voice signals before the current time of the system, sequentially sorting all types of voiceprint representations according to the values of the representation mutual difference highlight index from large to small, and marking the sorting sequence as the general voiceprint representation extraction sequence. 3.The voiceprint-based memory characterization pickup identification method according to claim 2, characterized in that, The acquisition process of the representation mutual difference highlight index of one type of voiceprint representation is as follows: selecting one type of voiceprint representation, acquiring the feature vectors of the type of voiceprint representation of all users in the user voiceprint representation memory, combining the feature vectors of the type of voiceprint representation of each two users into a voiceprint representation comparison combination, and then generating a plurality of voiceprint representation comparison combinations; acquiring the representation vector similarity of each voiceprint representation comparison combination; performing sum average calculation on the representation vector similarities of all voiceprint representation comparison combinations to calculate the average value of the representation vector similarity , setting a representation vector similarity threshold value, marking the voiceprint representation comparison combination as a voiceprint representation mutual difference combination when the representation vector similarity of the voiceprint representation comparison combination is lower than the representation vector similarity threshold value, performing ratio calculation on the total number of voiceprint representation mutual difference combinations and the total number of voiceprint representation comparison combinations to calculate the proportion of representation mutual difference , calculating the representation mutual difference highlight index of the type of voiceprint representation by ; , are representation weight coefficients. 4.The voiceprint-based memory characterization pickup identification method of claim 2, wherein, The user voiceprint representation memory contains each voiceprint feature of all users, and the construction process is as follows: determining the types of voiceprint representations that need to be stored in the user voiceprint representation memory, collecting valid speech of each user through a microphone, extracting each type of voiceprint representation from the valid speech of each user, further obtaining a feature vector of each type of voiceprint representation of each user, using a user-voiceprint representation storage structure, each user corresponding to an independent storage unit, and then constructing the user voiceprint representation memory. 5.The voiceprint-based memory characterization pickup identification method according to claim 1, wherein, When there is a grabbing user, it is determined whether there is a unique representation user among all grabbing users, and the specific process is as follows: obtaining an extraction representation similarity of each grabbing user, further obtaining a voiceprint representation collocation index of each grabbing user, marking the first two grabbing users with the highest voiceprint representation collocation index as first users, calculating the absolute difference of the voiceprint representation collocation index of the two first users to obtain a voiceprint representation collocation gap value, setting a voiceprint representation collocation gap threshold, when the voiceprint representation collocation gap value is higher than the voiceprint representation collocation gap threshold, marking the first user with the highest voiceprint representation collocation index as the unique representation user, and when the voiceprint representation collocation gap value is not higher than the voiceprint representation collocation gap threshold, determining that there is no unique representation user among all grabbing users.

6. The voiceprint-based memory characterization pickup identification method according to claim 5, characterized in that, The extraction representation similarity of the grabbing user is obtained as follows: selecting a grabbing user, obtaining a representation vector similarity of each extraction representation of the grabbing user, performing sum average value calculation on the representation vector similarities of all extraction representations, and calculating the extraction representation similarity of the grabbing user. 7.The voiceprint-based memory characterization pickup identification method according to claim 5, characterized in that, The process of obtaining the voiceprint representation matching index of the user is as follows: selecting a user, obtaining the extraction representation similarity of the user, marking the remaining users as comparison users, comparing the user with each comparison user, calculating the absolute difference of the extraction representation similarity of the comparison user and the comparison user, calculating the average extraction representation similarity difference, summing the extraction representation similarity of the user and the average extraction representation similarity difference to obtain the voiceprint representation matching index of the user. 8.The voiceprint-based memory characterization pickup identification method of claim 1, wherein, The customization process of the dynamic sound feature extraction sequence is as follows: obtaining the representation dynamic optimization capture index of the remaining type voiceprint representation, sorting all voiceprint representations in descending order according to the value of the representation dynamic optimization capture index, and labeling the sorted sequence as the dynamic sound feature extraction sequence. 9.The voiceprint-based memory characterization pickup identification method of claim 1, wherein, The acquisition process of the representation dynamic optimization crawling index of one type of voiceprint representation is as follows: selecting one type of voiceprint representation, obtaining the feature vectors of each crawling user for the type of voiceprint representation from the user voiceprint representation memory bank, combining the feature vectors of each two crawling users for the type of voiceprint representation into a target representation comparison combination, obtaining the representation vector similarity of each target representation comparison combination, further determining the individual comparison difference index of each crawling user, setting an individual comparison difference threshold, when the individual comparison difference index of the crawling user is lower than the individual comparison difference threshold, marking the crawling user as a representation independent user, when the individual comparison difference index of the crawling user is not lower than the individual comparison difference threshold, marking the crawling user as a non-representation independent user, calculating the ratio of the total number of non-representation independent users to the total number of representation independent users, calculating the non-representation ratio PPdc, calculating the sum average of the individual comparison difference indexes of all non-representation independent users, calculating the non-representation independent difference index FCGA, obtaining n general voiceprint representation extraction sequences updated before the current time of the system, obtaining the sorting sequence number of the type of voiceprint representation in each general voiceprint representation extraction sequence, calculating the average sorting sequence number Wac by summing and averaging all sorting sequence numbers of the type of voiceprint representation, comparing all sorting sequence numbers of the type of voiceprint representation in pairs, calculating the absolute difference value of the two compared sorting sequence numbers to calculate the sorting sequence number fluctuation value, calculating the average sorting sequence number fluctuation value Ydd by summing and averaging all sorting sequence number fluctuation values, and calculating the representation dynamic optimization crawling index of the type of voiceprint representation by h1*Ydd+h2*FCGA wherein h1 is the first auxiliary coefficient and h2 is the second auxiliary coefficient. The process of determining the individual comparison difference index of the user is as follows: selecting a user, obtaining all target representation comparison combinations containing the user and marking them as individual comparison combinations, summing the representation vector similarity of all individual comparison combinations to calculate the individual comparison difference index of the user.

10. The sound recognition system based on voiceprint characteristic memory, applied to the sound recognition method based on voiceprint characteristic memory according to any one of claims 1-9, characterized in that, It comprises a general voiceprint representation sequence module, a user capture judgment module, and a sound recognition analysis module.

Citation Information

Patent Citations

  • Streaming speech recognition method and system, equipment and storage medium

    CN118038874A

  • Newborn cry recognition system based on machine learning

    CN119920261A

  • Voice verification circuit for validating the identity of an unknown person

    US5054083A