Identity recognition method, device, electronic device and storage medium
By calculating the acoustic and prosthetic feature scores of the audio signal and combining interactive features, the problem of insufficient feature fusion in the prior art is solved, and the accuracy and reliability of speaker recognition are improved.
Patent Information
- Application Number
- CN202510221598.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-27
AI Technical Summary
In the prior art, the acoustic and prologue characteristics lack an effective fusion mechanism, resulting in a decrease in the accuracy of speaker recognition.
By extracting the acoustic characteristics and prosody characteristics of the audio signal, each score (Signal-to-noise ratio score, stability score, consistency score, reliability score), and based on these scores, the feature quality score is determined, and the target feature is determined based on the interactive characteristics of the acoustic characteristics and prosody characteristics are determined.
Effectively fusion of acoustic and pronunciation characteristics improves the accuracy and reliability of speaker recognition.
Smart Images

Figure CN119694321B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of identity recognition technology, and in particular to an identity recognition method, device, electronic equipment and storage medium. Background Art
[0002] As an important branch of biometric identification, speaker recognition technology has broad application prospects in many fields. In the field of identity recognition, this technology can be used in scenarios such as financial transactions and access control systems to provide a safe and convenient way to authenticate identities. In the field of security monitoring, speaker recognition can assist in identity confirmation. In the field of intelligent interaction, this technology provides a personalized voice interaction experience for smart homes and in-vehicle systems. In addition, in the fields of judicial evidence collection and medical diagnosis, speaker recognition also shows unique application value.
[0003] However, existing methods have a gap in feature expression, lack an effective fusion mechanism for acoustic features and prosodic features, and cannot fully utilize multimodal information, thus reducing the accuracy of speaker recognition. Summary of the invention
[0004] The present invention provides an identity recognition method, device, electronic device and storage medium to solve the defects in the prior art that there is a split in feature expression, acoustic features and prosodic features lack an effective fusion mechanism, multimodal information cannot be fully utilized, and the accuracy of speaker recognition is reduced.
[0005] The present invention provides an identity recognition method, comprising the following steps:
[0006] respectively obtaining target features of at least two audio signals of the person to be tested;
[0007] Based on all the target features, determining the identity recognition result of the person to be tested;
[0008] The target feature of the audio signal is determined based on the following steps:
[0009] Extracting acoustic features and prosodic features of the audio signal, determining a first score of the acoustic features, and determining a second score of the prosodic features; the first score includes a signal-to-noise ratio score and a stability score, the signal-to-noise ratio score is used to reflect the signal-to-noise ratio of the audio signal, and the stability score is used to reflect the stability of the audio signal; the second score includes a consistency score and a reliability score, the consistency score is used to reflect the consistency of the audio signal with other audio signals, and the reliability score is used to reflect the reliability of the audio signal;
[0010] A feature quality score is determined based on the first score and the second score, and the target feature is determined based on the feature quality score and an interaction feature of the acoustic feature and the prosodic feature.
[0011] According to an identity recognition method provided by the present invention, the step of acquiring the acoustic features and the prosodic features comprises:
[0012] Determining the acoustic feature based on the Mel-frequency cepstral coefficient feature, fundamental frequency feature and energy feature of the audio signal;
[0013] The prosody feature is determined based on the intonation feature, the rhythm feature and the duration feature of the audio signal.
[0014] According to an identity recognition method provided by the present invention, the Mel frequency cepstral coefficient feature, the fundamental frequency feature and the energy feature are all extracted from the audio signal based on an adaptive window size;
[0015] The intonation feature, the rhythm feature and the duration feature are all extracted from the audio signal based on an adaptive window size;
[0016] The step of determining the adaptive window size comprises:
[0017] Determining an adaptive coefficient based on local energy, zero crossing rate, and sound change rate;
[0018] The adaptive window size is determined based on the adaptive coefficient and the basic window size.
[0019] According to an identity recognition method provided by the present invention, determining the identity recognition result of the person to be tested based on all the target features includes:
[0020] Inputting all the target features into a graph structure to obtain graph structure features output by the graph structure;
[0021] Extracting bottom-level features, middle-level features and high-level features of the graph structure features respectively, and determining multi-layer fusion features based on the bottom-level features, the middle-level features and the high-level features;
[0022] Based on the correlation measurement and redundancy measurement of the nodes in the graph structure, feature screening is performed on the multi-layer fusion features to obtain screening features;
[0023] Inputting the screening features into a structure preserving encoder to obtain all encoding features output by the structure preserving encoder;
[0024] The identity recognition result is determined based on the similarities between all the coding features.
[0025] According to an identity recognition method provided by the present invention, determining the identity recognition result based on the similarity between all the coding features includes:
[0026] Determining a similarity score based on the local structural similarity, global distribution similarity and semantic similarity between all the coding features, and a weight coefficient; the weight coefficient is dynamically adjusted based on an adaptive mechanism;
[0027] Based on the similarity score, the identity recognition result is determined.
[0028] According to an identity recognition method provided by the present invention, determining the identity recognition result based on the similarity score includes:
[0029] Based on the similarity score, determining a confidence score;
[0030] The identity recognition result is determined based on the similarity score and the confidence score.
[0031] According to an identity recognition method provided by the present invention, the training step of the structure-preserving encoder includes:
[0032] Obtaining an initial structure preserving encoder, a sample screening feature, and a prior distribution of the sample screening feature;
[0033] Inputting the sample screening feature into the initial structure preserving encoder to obtain the prediction coding feature input by the initial structure preserving encoder;
[0034] Determining a global distribution loss based on an actual distribution corresponding to the predicted coding feature and the prior distribution;
[0035] Based on the global distribution loss, parameters of the initial structure preserving encoder are iterated to obtain the structure preserving encoder.
[0036] According to an identity recognition method provided by the present invention, the method of performing parameter iteration on the initial structure preserving encoder based on the global distribution loss to obtain the structure preserving encoder includes:
[0037] Obtaining a first sample screening feature and a second sample screening feature;
[0038] Inputting the first sample screening feature into the initial structure preserving encoder to obtain a first encoding feature output by the initial structure preserving encoder;
[0039] Inputting the second sample screening feature into the initial structure preserving encoder to obtain a second encoding feature output by the initial structure preserving encoder;
[0040] determining a local structural loss based on a difference between the first encoding feature and the second encoding feature;
[0041] Acquire a sample audio signal and a speaker label of the sample audio signal;
[0042] Using the initial structure-maintaining encoder as an initial generator, and generating a target coding feature corresponding to the sample audio signal based on the initial generator;
[0043] Determining, based on the initial discriminator, the identity authentication category corresponding to the target coding feature;
[0044] Based on the difference between the identity verification category and the speaker label, a discriminative loss is determined, and based on the discriminative loss, the local structure loss and the global distribution loss, a target loss is determined, and parameters of the initial structure-preserving encoder are iterated based on the target loss to obtain the structure-preserving encoder.
[0045] According to an identity recognition method provided by the present invention, the edge weight matrix of the graph structure is determined based on the acoustic similarity, rhythmic similarity, temporal similarity and contextual similarity between two nodes in the graph structure;
[0046] The acoustic similarity is determined based on acoustic features between two nodes in the graph structure;
[0047] The prosodic similarity is determined based on prosodic features between any two nodes in the graph structure;
[0048] The temporal similarity is determined based on the timestamps corresponding to each two nodes in the graph structure;
[0049] The context similarity is determined based on the context relevance between any two nodes in the graph structure.
[0050] The present invention also provides an identity recognition device, comprising:
[0051] An acquisition unit, used to respectively acquire target features of at least two audio signals of the person to be tested;
[0052] A determination unit, used to determine the identity recognition result of the person to be tested based on all the target features;
[0053] The target feature of the audio signal is determined based on the following steps:
[0054] Extracting acoustic features and prosodic features of the audio signal, determining a first score of the acoustic features, and determining a second score of the prosodic features; the first score includes a signal-to-noise ratio score and a stability score, the signal-to-noise ratio score is used to reflect the signal-to-noise ratio of the audio signal, and the stability score is used to reflect the stability of the audio signal; the second score includes a consistency score and a reliability score, the consistency score is used to reflect the consistency of the audio signal with other audio signals, and the reliability score is used to reflect the reliability of the audio signal;
[0055] A feature quality score is determined based on the first score and the second score, and the target feature is determined based on the feature quality score and an interaction feature of the acoustic feature and the prosodic feature.
[0056] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, any of the above-mentioned identity recognition methods is implemented.
[0057] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, and the computer program, when executed by a processor, implements any of the above-mentioned identity recognition methods.
[0058] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the identity recognition method described above is implemented.
[0059] The identity recognition method, device, electronic device and storage medium provided by the present invention respectively obtain the target features of at least two audio signals of the person to be tested, and then determine the identity recognition result of the person to be tested based on all the target features; wherein the target features of the audio signal are determined based on the following steps: extracting the acoustic features and prosodic features of the audio signal, determining the first score of the acoustic features, and determining the second score of the prosodic features; the first score includes a signal-to-noise ratio score and a stability score, and the second score includes a consistency score and a reliability score, and finally, based on the first score and the second score, determining the feature quality score, and based on the feature quality score and the interactive features of the acoustic features and the prosodic features, determining the target features. This process makes full use of multimodal information such as the signal-to-noise ratio score, the stability score, the consistency score and the reliability score, effectively fuses the acoustic features and the prosodic features, and determines the identity recognition result based on all the target features obtained by the fusion, thereby improving the accuracy and reliability of speaker recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0061] Figure 1 This is one of the flow charts of the identity recognition method provided by the present invention.
[0062] Figure 2 It is a flow chart of the feature extraction module provided by the present invention.
[0063] Figure 3 It is a schematic diagram of the map construction process provided by the present invention.
[0064] Figure 4 It is a schematic diagram of the process of multi-level feature extraction and fusion provided by the present invention.
[0065] Figure 5 It is a flow chart of the graph entropy analysis and key feature screening module provided by the present invention.
[0066] Figure 6 It is a flow chart of the encoding module provided by the present invention.
[0067] Figure 7 It is a flow chart of the similarity calculation and determination module provided by the present invention.
[0068] Figure 8 This is the second flow chart of the identity recognition method provided by the present invention.
[0069] Fig. 9 It is a structural schematic diagram of the identity recognition device provided by the present invention.
[0070] Fig.10 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0071] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0072] The terms "first", "second", etc. in the present invention are used to distinguish similar objects, rather than to describe a specific order or precedence. It should be understood that the terms used in this way can be interchangeable where appropriate, so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of the same type.
[0073] In the related technology, in the speaker verification system, feature extraction is the front-end module of the entire system, and its performance directly determines the upper limit of subsequent processing. Traditional feature extraction methods face fundamental challenges: speaker features show different characteristics at different time scales, from millisecond-level vocal tract features to sentence-level prosodic features, all of which contain important identity information. How to effectively extract and organize these multi-scale features becomes a key issue.
[0074] Based on the above problems, the present invention provides an identity recognition method. Figure 1 It is one of the flow charts of the identity recognition method provided by the present invention, such as Figure 1 As shown, the method includes step 110 and step 120.
[0075] Step 110: Obtain target features of at least two audio signals of the person to be tested respectively.
[0076] Specifically, target features of at least two audio signals of the person to be tested can be obtained respectively, wherein the at least two audio signals of the person to be tested can be obtained through a sound pickup device, where the sound pickup device can be a smart phone, a tablet computer, or a smart appliance, such as a stereo, a television, and an air conditioner. After the sound pickup device obtains voice data through a microphone array, it can also amplify and reduce noise on the voice data. The embodiment of the present invention does not specifically limit this.
[0077] Step 120, determining the identity recognition result of the person to be tested based on all the target features;
[0078] The target feature of the audio signal is determined based on the following steps:
[0079] Extracting acoustic features and prosodic features of the audio signal, determining a first score of the acoustic features, and determining a second score of the prosodic features; the first score includes a signal-to-noise ratio score and a stability score, the signal-to-noise ratio score is used to reflect the signal-to-noise ratio of the audio signal, and the stability score is used to reflect the stability of the audio signal; the second score includes a consistency score and a reliability score, the consistency score is used to reflect the consistency of the audio signal with other audio signals, and the reliability score is used to reflect the reliability of the audio signal;
[0080] A feature quality score is determined based on the first score and the second score, and the target feature is determined based on the feature quality score and an interaction feature of the acoustic feature and the prosodic feature.
[0081] Specifically, after all audio signals are acquired, the identity recognition result of the person to be tested can be determined based on all target features. Specifically, the acoustic features and rhythmic features of the audio signal are extracted to determine the first score of the acoustic feature. The first score includes a signal-to-noise ratio score and a stability score. The signal-to-noise ratio score is used to reflect the signal-to-noise ratio of the audio signal, and the stability score is used to reflect the stability of the audio signal. The formula of the first score is as follows:
[0082]
[0083] in, Indicates the first score, Represents the acoustic characteristics, represents the signal-to-noise ratio score, Represents the stability score.
[0084] It should be noted that the signal-to-noise ratio evaluation function The quality of the feature is measured by the ratio of signal power to noise power, and the stability evaluation function The standard deviation or coefficient of variation measures how much a feature varies across samples, with lower values indicating greater stability.
[0085] Among them, the acoustic features can be determined based on Mel-frequency cepstral coefficient features, fundamental frequency features and energy features.
[0086] Then, a second score of the prosody feature may be determined, the second score including a consistency score and a reliability score, the consistency score is used to reflect the consistency of the audio signal with other audio signals, and the reliability score is used to reflect the reliability of the audio signal. The formula for the second score is as follows:
[0087]
[0088]
[0089]
[0090] in, represents the second score, Represents rhythmic features, represents the consistency score, represents the reliability score, For the i reference templates, N is the number of templates, is the variance of the current feature in the short time window, is the attenuation coefficient.
[0091] Among them, the prosodic feature can be determined based on the intonation feature, rhythm feature and duration feature of the audio signal.
[0092] Further, after determining the first score and the second score, a feature quality score may be determined based on the first score and the second score, as shown in the following formula:
[0093]
[0094] in, represents the feature quality score, and is the weight coefficient, Indicates the first score, Score for the second.
[0095] Further, after determining the feature quality score, the interactive features of the acoustic features and the prosodic features, as well as the feature quality score , determine the target features of any audio signal , the formula is as follows:
[0096]
[0097] in, is the feature fusion function, defined as:
[0098]
[0099] in, , Acoustic characteristics and rhythmic features interactive features.
[0100] Step 120: Determine the identity recognition result of the person to be tested based on all the target features.
[0101] Specifically, after all target features are obtained, the identity recognition result of the person to be tested can be determined based on all target features. For example, two target features can be compared, and the identity recognition result of the person to be tested can be determined based on the similarity between the two target features.
[0102] The method provided by the embodiment of the present invention obtains the target features of at least two audio signals of the person to be tested respectively, and then determines the identity recognition result of the person to be tested based on all the target features; wherein the target features of the audio signal are determined based on the following steps: extracting the acoustic features and prosodic features of the audio signal, determining the first score of the acoustic features, and determining the second score of the prosodic features; the first score includes a signal-to-noise ratio score and a stability score, and the second score includes a consistency score and a reliability score, and finally, based on the first score and the second score, determining the feature quality score, and based on the feature quality score and the interactive features of the acoustic features and the prosodic features, determining the target features. This process makes full use of multimodal information such as the signal-to-noise ratio score, the stability score, the consistency score, and the reliability score, effectively fuses the acoustic features and the prosodic features, and determines the identity recognition result based on all the target features obtained by the fusion, thereby improving the accuracy and reliability of speaker recognition.
[0103] Based on the above embodiment, the step of acquiring the acoustic features and the prosodic features includes:
[0104] Step 210, determining the acoustic feature based on the Mel-frequency cepstral coefficient feature, fundamental frequency feature and energy feature of the audio signal;
[0105] Step 220: Determine the prosodic feature based on the intonation feature, rhythm feature, and duration feature of the audio signal.
[0106] Specifically, the acoustic features are determined based on the Mel frequency cepstral coefficient features, fundamental frequency features and energy features of the audio signal. The formula is as follows:
[0107]
[0108]
[0109] in, Represents the acoustic characteristics, represents the Mel frequency cepstral coefficient feature, Represents the fundamental frequency feature (Fundamental Frequency), Represents energy characteristics, is any audio signal, representing the language waveform data in the time domain, is a window function, is the Mel filter bank, is the discrete cosine transform, is the fundamental frequency extraction function, It is an energy calculation function, and FFT() represents Fast Fourier Transform (FFT), which converts the time domain signal into a frequency domain energy spectrum.
[0110] It should be noted that the Mel-Frequency Cepstral Coefficient feature is obtained by performing MFCC (Mel-Frequency Cepstral Coefficients) on any audio signal. The fundamental frequency feature refers to the lowest frequency component in the audio signal, which determines the pitch of the sound and is one of the most characteristic frequency components in the sound.
[0111] Then, based on the intonation features, rhythm features, and duration features of the audio signal, the prosodic features are determined, and the formula is as follows:
[0112]
[0113]
[0114] in, Represents rhythmic features, represents the intonation feature at time t, represents the rhythm characteristics at time t, represents the duration feature at time t, is the intonation extraction function, is the rhythm extraction function, is the duration feature extraction function, Represents the fundamental frequency of any audio signal at time t, which is used to describe the pitch of speech. represents the energy at time t, describing the intensity of speech, represents the zero-crossing rate, represents the speech segment at time t, i.e., the audio signal, Represents the context weight.
[0115] In the related technology, in the speaker verification system, feature extraction is the front-end module of the entire system, and its performance directly determines the upper limit of subsequent processing. Traditional feature extraction methods face three fundamental challenges: first, the speech signal itself has a highly time-varying characteristic, and fixed feature extraction strategies are difficult to accurately capture this dynamic change; second, speaker features show different characteristics at different time scales, from millisecond-level vocal tract features to sentence-level prosodic features, all of which contain important identity information. How to effectively extract and organize these multi-scale features has become a key issue; third, in actual application environments, background noise, channel distortion and other factors will seriously affect the quality of features. How to extract robust feature representation is a problem that needs to be solved urgently.
[0116] Based on the above embodiment, the Mel frequency cepstral coefficient feature, the fundamental frequency feature and the energy feature are all extracted from the audio signal based on an adaptive window size;
[0117] The intonation feature, the rhythm feature and the duration feature are all extracted from the audio signal based on an adaptive window size;
[0118] The step of determining the adaptive window size comprises:
[0119] Determining an adaptive coefficient based on local energy, zero crossing rate, and sound change rate;
[0120] The adaptive window size is determined based on the adaptive coefficient and the basic window size.
[0121] Specifically, Figure 2 It is a flow chart of the feature extraction module provided by the present invention, such as Figure 2 As shown, for any audio signal (input speech waveform), after passing through the adaptive window module, that is, based on the adaptive window size, the Mel-frequency cepstral coefficient feature (MFCC), fundamental frequency feature (F0) and energy feature (Energy) of any audio signal are extracted respectively.
[0122] The step of determining the adaptive window size includes:
[0123] Determining an adaptive coefficient based on local energy, zero crossing rate, and sound change rate;
[0124] Based on the adaptive coefficient and the basic window size, the adaptive window size is determined as follows:
[0125] make Indicates the adaptive window size, that is, time Window size at:
[0126]
[0127] in, is the base window size, is the adaptive coefficient, which is calculated by the following formula:
[0128]
[0129] in, is the local energy, is the zero crossing rate, is the sound change rate, is the adaptive mapping function.
[0130] Accordingly, the intonation feature, rhythm feature and duration feature of any audio signal can be extracted respectively based on the adaptive window size.
[0131] Finally, feature quality assessment is performed based on acoustic features and prosodic features, that is, SNR assessment and stability assessment are performed based on acoustic features, and consistency assessment and reliability assessment are performed based on prosodic features.
[0132] The method provided by the embodiment of the present invention adopts a variable analysis window in the time dimension, and the size of the window is automatically adjusted according to the local stability of the speech signal. In areas where the signal energy changes dramatically, the system automatically reduces the window to capture rapidly changing features; in relatively stable areas, a larger window is used to obtain a more stable feature representation. This adaptive mechanism enables the system to accurately capture key changes in the speech signal while maintaining computational efficiency, thereby dynamically perceiving the local characteristics of the speech signal, adaptively adjusting the feature extraction strategy, and improving the stability of feature extraction.
[0133] The method provided by the embodiment of the present invention designs a window size adjustment strategy based on signal energy and zero-crossing rate, and designs a feature compensation mechanism to handle discontinuities caused by window changes.
[0134] In the related technologies, in speaker verification systems, how to effectively organize and express the complex correlation between features is a key issue. Traditional feature organization methods often use simple vector concatenation or statistical aggregation, which has three significant limitations: first, it cannot effectively capture the dynamic correlation patterns between features, resulting in the loss of important structural information; second, the organizational structure of features is too static and rigid, making it difficult to adapt to the time-varying characteristics of speaker features; finally, the lack of hierarchical expression of feature relationships makes it difficult for the system to understand the deep dependencies between features. In other words, the description of feature correlation is not comprehensive, often only considering the feature similarity of a single dimension, ignoring the complex interactive relationships between features.
[0135] To address these fundamental issues, the embodiments of the present invention innovatively introduce a dynamic graph structure to express and organize features. The core idea of this graph-based expression is to represent the features of each key time point as nodes in the graph, and to characterize the multi-dimensional association relationship between features through carefully designed edge weights. The uniqueness of this design is that it can not only express the direct association between features, but also naturally reflect the transitive relationship between features through the connection structure of the graph.
[0136] At the node construction level, we break through the traditional uniform sampling method and propose an adaptive node generation strategy based on information importance. The system evaluates the discriminative contribution of the features at each time point in real time and prioritizes those time points with high information content as nodes in the graph. This selective node generation mechanism significantly improves the information efficiency of the graph structure while also reducing the computational complexity of subsequent processing.
[0137] In terms of edge weight design, the embodiment of the present invention constructs a multi-dimensional similarity measurement method. Unlike traditional methods that only consider the Euclidean distance or cosine similarity between feature vectors, the method of the embodiment of the present invention simultaneously considers multiple dimensions such as acoustic similarity, prosodic correlation, temporal dependency, and contextual consistency. These similarities in different dimensions are fused through a dynamic weight mechanism, and the weight distribution will be adaptively adjusted according to the characteristics of the current speech segment and the requirements of the verification task.
[0138] Based on the above embodiment, the step 120 of determining the identity recognition result of the person to be tested based on all the target features includes:
[0139] Step 121, inputting all the target features into a graph structure to obtain graph structure features output by the graph structure;
[0140] Step 122, respectively extracting bottom-level features, middle-level features, and high-level features of the graph structure features, and determining multi-layer fusion features based on the bottom-level features, the middle-level features, and the high-level features;
[0141] Step 123, based on the correlation measurement and redundancy measurement of the nodes in the graph structure, feature screening is performed on the multi-layer fusion features to obtain screening features;
[0142] Step 124, inputting the screening features into a structure preserving encoder to obtain all encoding features output by the structure preserving encoder;
[0143] Step 125, determining the identity recognition result based on the similarities between all the coding features.
[0144] Specifically, all target features are input into the graph structure to obtain graph structure features of the graph structure output.
[0145] Figure 3 It is a schematic diagram of the map construction process provided by the present invention, such as Figure 3 As shown in the figure, firstly, key time points are extracted from the speech feature data, then the associations between points are established, and finally the most important features and associations are retained. Specifically, the construction process of the graph structure is as follows:
[0146] make Represents the generated node set:
[0147]
[0148] in, For time The eigenvector at is the node importance score, calculated as follows:
[0149]
[0150] in, is the characteristic significance, is the local difference, is the context importance, is the weight coefficient.
[0151] Here, the edge weight matrix of the graph structure is determined based on the acoustic similarity, prosodic similarity, temporal similarity, and contextual similarity between two nodes in the graph structure;
[0152] Among them, the acoustic similarity is determined based on the acoustic features between two nodes in the graph structure;
[0153] The prosodic similarity is determined based on the prosodic features between two nodes in the graph structure;
[0154] The temporal similarity is determined based on the timestamps corresponding to each pair of nodes in the graph structure;
[0155] The context similarity is determined based on the context relevance between two nodes in the graph structure. The formula is as follows:
[0156] Edge Weight Matrix The calculation process:
[0157]
[0158] in, represents the u-th similarity calculation function, Represents two nodes in the graph, is the adaptive weight;
[0159] The similarity calculation includes:
[0160] Acoustic Similarity:
[0161] in, Representation Node The acoustic eigenvector of Representation Node The acoustic eigenvector of is the scale parameter of the acoustic feature;
[0162] Rhythmic Similarity:
[0163] in, Representation Node The rhythmic feature vector of Representation Node The prosodic feature vector of
[0164] Timing similarity:
[0165] in, Representation Node The timestamp, that is, the start and end time of the speech segment, Representation Node timestamp, is the time series attenuation coefficient;
[0166] Contextual Similarity:
[0167] in, Represents two nodes in a graph.
[0168] It should be noted that the design ideas of the above-mentioned edge weight calculation module include three points, namely, building a multi-dimensional similarity calculation framework, designing an adaptive weight fusion mechanism, and introducing an edge weight dynamic adjustment strategy.
[0169] The embodiment of the present invention also involves a graph structure optimization module. The design idea of the graph structure optimization module is: designing a dynamic graph pruning strategy, building an edge weight redistribution mechanism, and implementing a structural redundancy elimination algorithm, which is as follows:
[0170] Optimized graph structure It can be expressed as:
[0171]
[0172] Among them, node optimization:
[0173] Edge set optimization:
[0174] here: Reservation threshold for nodes, is the dynamic edge weight threshold, calculated as follows:
[0175]
[0176] in, is the basic threshold, is the adjustment coefficient, is the graph complexity evaluation function.
[0177] Overall formal expression:
[0178] Output graph structure of graph construction module It can be expressed as:
[0179]
[0180] in, For a node collection: , For edge sets: , is the edge weight matrix: .
[0181] In summary, at the feature representation level, the present invention breaks the separation of acoustic features and prosodic features in traditional methods and designs a unified feature extraction framework. This framework simultaneously processes speech features at different time scales through a multi-scale analysis network. At the bottom layer, the system extracts basic acoustic parameters such as fundamental frequency, energy, and spectral envelope; at the middle layer, acoustic patterns at the phoneme and syllable levels are captured through feature combination; at the high level, attention is paid to the long-term characteristics of the speaker, such as intonation curves and speaking styles. This hierarchical design ensures that features at different scales can be fully expressed.
[0182] In the related technologies, how to effectively integrate feature information at different levels is the key to improving system performance in speaker verification. Traditional feature fusion methods have four main deficiencies: first, the scale differences and expression differences between features at different levels make direct fusion prone to information imbalance; second, the feature fusion process lacks selectivity and easily introduces redundant or irrelevant information; third, the fusion strategy is usually static and cannot be dynamically adjusted according to the quality of the input features; finally, the fused features often lose the hierarchical structure information of the original features.
[0183] Based on in-depth thinking about these issues, this module proposes an innovative hierarchical feature extraction and fusion framework. The core concept of this framework is to achieve organic fusion of features while maintaining the independence of features at each layer by building a multi-level feature extraction network. This design breaks the flat thinking of traditional feature fusion and builds a three-dimensional feature expression system.
[0184] In the underlying feature extraction stage, the system focuses on basic acoustic features, such as spectrum features, energy features, etc. The innovation of this layer is the introduction of an adaptive feature enhancement mechanism that can dynamically adjust the feature extraction parameters according to the local characteristics of the signal. In particular, when processing speech segments with severe noise or channel distortion, the system automatically enhances the feature extraction strength of key frequency bands to ensure the reliability of the underlying features.
[0185] The feature extraction in the middle layer focuses on capturing acoustic events and speech patterns. In this layer, we design an innovative context-aware network that can effectively integrate temporal context information and identify discriminative acoustic patterns. The network achieves selective utilization of context information through the attention mechanism and avoids interference from irrelevant information.
[0186] High-level feature extraction focuses on the overall feature expression of the speaker. The core of this layer is an adaptive feature aggregation network that can extract stable speaker features from long-term sequences. In particular, we introduce a feature enhancement mechanism based on contrastive learning, which improves the discriminability and robustness of features by constructing feature representations from different perspectives.
[0187] At the feature fusion level, we broke the traditional practice of simple superposition or splicing and designed a dynamic feature fusion network. This network achieves adaptive fusion of features by learning the importance weights of features at different levels. More importantly, the fusion process maintains the hierarchical properties of the features, which enables the system to flexibly access feature information at different levels as needed in subsequent processing.
[0188] In order to ensure the quality of fused features, we also introduced an innovative feature quality assessment mechanism. This mechanism can evaluate the reliability of features at each layer in real time and dynamically adjust the fusion strategy based on the evaluation results. This adaptive fusion method significantly improves the robustness of the system in complex environments.
[0189] Specifically, the bottom-level features, middle-level features and high-level features of the graph structure features can be extracted respectively, and multi-layer fusion features can be determined based on the bottom-level features, middle-level features and high-level features. This can be achieved through a multi-level feature extraction and fusion module as follows:
[0190] Underlying features The extraction process:
[0191]
[0192] Among them, the extraction of each sub-feature is:
[0193] in, is the input feature map, is the convolution kernel set, is the attention matrix, is the convolution operation, It is the underlying feature extraction function.
[0194] Mid-level features The extraction process:
[0195]
[0196] in:
[0197] here: is the underlying feature set, is the context window, is the middle-level attention matrix, is the feature aggregation function, It is the middle-level feature extraction function.
[0198] High-level features The extraction process:
[0199]
[0200] in:
[0201] here: is the middle-level feature set, is the global context, is the high-level attention matrix, is the feature transformation function, It is a high-level feature extraction function.
[0202] Fusion Features The generation process:
[0203]
[0204] in, Represents the underlying features, Represents the middle-level features, Represents high-level features, Represents the weight matrix of each layer feature.
[0205] The fusion function is defined as:
[0206] here: is the weight matrix of each layer feature, is the residual connection, is the activation function.
[0207] Overall formal expression:
[0208] The output features of the multi-level feature extraction and fusion module, i.e., multi-layer fusion features It can be expressed as:
[0209]
[0210] in, is the overall feature extraction and fusion function, is the input fusion feature, is the context information, For global information.
[0211] Figure 4 It is a schematic diagram of the process of multi-level feature extraction and fusion provided by the present invention, such as Figure 4As shown, the bottom-level feature extraction includes the extraction of vocal tract features, fundamental frequency features, energy features and spectrum features; the middle-level feature extraction includes the extraction of phoneme-level features and syllable-level features; the high-level feature extraction includes the extraction of intonation curves and speaking styles; finally, the bottom-level features, middle-level features and high-level features are fused.
[0212] In speaker verification systems, how to select the most representative and discriminative feature set from a large number of features is a key challenge. Traditional feature selection methods are usually based on simple statistical measures or fixed selection rules. These methods have the following core problems: first, they often ignore the complex correlation between features, resulting in the selected feature set being individually significant but redundant as a whole; second, the feature importance assessment is too static and cannot adapt to the dynamic changes of speakers and environments; third, feature selection without theoretical guidance may lose key discriminative information, affecting the overall performance of the system.
[0213] Based on these observations, this module (graph entropy analysis and key feature screening module) innovatively introduces graph theory and information entropy theory into the feature selection process. This design is based on the following in-depth thinking:
[0214] Systematic expression of feature associations: There are complex dependencies between speaker features, which can be naturally expressed through graph structures. By mapping the feature space to a graph structure, we can use powerful tools from graph theory to analyze the interaction patterns and information flow between features. The graph structure can not only express the direct associations between features, but also reflect the indirect influences between features, providing a theoretical basis for comprehensive feature importance assessment.
[0215] Dynamic characteristics of entropy: Information entropy, as a core tool for measuring uncertainty, has good theoretical properties. By introducing a dynamic entropy calculation mechanism, the system can capture the time-varying characteristics of feature importance. Especially when processing speaker features, the change in entropy value can reflect the change in the distinguishing ability of features in different speech segments, providing a reliable metric for adaptive feature selection.
[0216] Multi-level information flow analysis: By building a hierarchical entropy propagation network, the system can analyze the information contribution of features at different abstract levels. This design allows the feature selection process to not only consider local statistical characteristics, but also grasp the global structural information, thereby achieving more intelligent feature screening.
[0217] Furthermore, after obtaining the multi-layer fusion features, the multi-layer fusion features can be screened based on the correlation measurement and redundancy measurement of the nodes in the graph structure to obtain the screening features, which can be implemented through the above-mentioned graph entropy analysis and key feature screening module, as follows:
[0218] Figure 5is a flow chart of the graph entropy analysis and key feature screening module provided by the present invention, such as Figure 5 As shown in the figure, the design idea of the node entropy evaluator is to build a multi-dimensional entropy calculation framework, design a local-global entropy fusion mechanism, and implement a dynamic entropy update strategy, as follows:
[0219] node The entropy value calculate:
[0220]
[0221] in, , and They represent the weight coefficients of local entropy value, global entropy value and temporal entropy value respectively.
[0222] The local entropy value is:
[0223] Global entropy value:
[0224] Time series entropy value:
[0225] here: is the local transition probability, is the class conditional probability, is the time series state probability.
[0226] The design idea of the graph structure entropy analyzer is to build a hierarchical entropy propagation network, design entropy aggregation and dispersion mechanisms, and implement adaptive propagation control, as follows:
[0227] Graph Structural Entropy Calculation:
[0228]
[0229] in, is a global balancing factor used to control the contribution of edge entropy to the graph structure entropy. represents vertex entropy, represents edge entropy.
[0230] The vertex entropy is:
[0231] Edge entropy:
[0232] Entropy Propagation Update:
[0233] here: is the probability of node importance, is the edge weight probability, is the normalized edge weight, It is used to adjust the proportion of entropy that a node retains when updating its entropy value. It is the entropy value of neighbor nodes in local propagation, which is used to dynamically adjust the importance of nodes.
[0234] The design ideas of the feature screening optimizer are: designing an entropy-based feature sorting strategy, building a multi-constrained feature selection mechanism, and implementing adaptive threshold adjustment, as follows:
[0235] Feature selection process:
[0236]
[0237] in, : candidate feature subset, It means to maximize the objective function while satisfying the constraints.
[0238] The score function is:
[0239] in, Represents the weight coefficient, which is used to adjust the balance between the correlation measure and the redundancy measure.
[0240] Correlation measures:
[0241] Redundancy measures:
[0242] in, is the selected feature subset, For the complete feature set, is the mutual information measure.
[0243] Overall formal expression:
[0244] The output of the graph entropy analysis and feature screening module can be expressed as:
[0245]
[0246] Among them, G is the input graph structure, H is the entropy matrix, is the adaptive threshold.
[0247] In speaker verification systems, the encoding module undertakes the key task of mapping high-dimensional feature space to low-dimensional discriminant space. Traditional encoding methods have three fundamental problems: first, it is difficult to maintain the discriminative information of features during dimensionality reduction, and effective information loss often occurs; second, the encoding results are too sensitive to small perturbations of the input features, resulting in insufficient system robustness; third, the lack of a structured constraint mechanism makes it difficult for the encoding space to form a good geometric structure.
[0248] This module proposes an adaptive encoding framework based on structure preservation. The core idea of this framework is to maintain the key structure of the feature space while reducing the dimension by introducing multiple constraint mechanisms. Specifically, we ensure the effectiveness of encoding from three levels: at the local structure level, by designing the neighborhood preservation loss, we ensure that the neighbor relationship of similar samples in the encoding space is not destroyed; at the global structure level, by introducing the manifold alignment mechanism, we maintain the overall shape of the sample distribution; at the semantic level, by constructing the speaker perception loss, we ensure that the encoding result has strong discriminability.
[0249] This multi-level structure-preserving strategy not only improves the reliability of encoding, but also provides a better feature basis for subsequent similarity calculations. In particular, by considering both local structure and global distribution during the encoding process, the system can better handle nonlinear changes in speaker characteristics and improve its adaptability in complex scenarios. At the same time, the introduced adaptive mechanism enables the encoding process to dynamically adjust the mapping strategy according to the characteristics of the input features, further enhancing the robustness of the system.
[0250] That is, the screening features can be input into the structure preserving encoder to obtain the encoding features output by the structure preserving encoder.
[0251] The core of the structure-preserving encoder is a nonlinear mapping function φ(x), which maps the input feature x to a low-dimensional space, formally expressed as:
[0252]
[0253] in, represents a parameterized mapping network, is the network parameter.
[0254] Here, the training steps of the structure-preserving encoder include:
[0255] An initial structure preserving encoder, sample screening features, and a priori distribution of the sample screening features are obtained, and then the sample screening features are input into the initial structure preserving encoder to obtain a predicted coding feature of the initial structure preserving encoder input.
[0256] After obtaining the predicted coding features, the global distribution loss can be determined based on the actual distribution corresponding to the predicted coding features and the prior distribution. The formula for the global distribution loss is as follows:
[0257]
[0258] in, represents the global distribution loss, represents the maximum mean difference, and are the actual distribution and target distribution in the encoding space, namely represents the actual distribution corresponding to the predicted encoding feature, represents the prior distribution (target distribution).
[0259] After obtaining the global distribution loss, the parameters of the initial structure preserving encoder can be iterated based on the global distribution loss, and the initial structure preserving encoder after completing the parameter iteration is used as the structure preserving encoder.
[0260] Based on the above embodiment, the step of performing parameter iteration on the initial structure preserving encoder based on the global distribution loss to obtain the structure preserving encoder includes:
[0261] Step 310, obtaining a first sample screening feature and a second sample screening feature;
[0262] Step 320, inputting the first sample screening feature into the initial structure preserving encoder to obtain a first encoding feature output by the initial structure preserving encoder;
[0263] Step 330, inputting the second sample screening feature into the initial structure preserving encoder to obtain a second encoding feature output by the initial structure preserving encoder;
[0264] Step 340, determining a local structural loss based on a difference between the first coding feature and the second coding feature;
[0265] Step 350, obtaining a sample audio signal and a speaker label of the sample audio signal;
[0266] Step 360, using the initial structure-maintaining encoder as an initial generator, and generating a target coding feature corresponding to the sample audio signal based on the initial generator;
[0267] Step 370, determining the identity authentication category corresponding to the target coding feature based on the initial discriminator;
[0268] Step 380, determining a discriminative loss based on the difference between the identity verification category and the speaker label, determining a target loss based on the discriminative loss, the local structure loss and the global distribution loss, and iterating parameters of the initial structure preserving encoder based on the target loss to obtain the structure preserving encoder.
[0269] Specifically, Figure 6 It is a flow chart of the encoding module provided by the present invention, such as Figure 6As shown, first, the first sample screening feature and the second sample screening feature are obtained, and then the first sample screening feature is input into the initial structure preserving encoder to obtain the first coding feature output by the initial structure preserving encoder, and the second sample screening feature is input into the initial structure preserving encoder to obtain the second coding feature output by the initial structure preserving encoder.
[0270] After obtaining the first coding feature and the second coding feature, the local structural loss can be determined based on the difference between the first coding feature and the second coding feature. The formula of the local structural loss is as follows:
[0271]
[0272] in, represents local structural loss, represents the first encoding feature, represents the second encoding feature, Represents sample pairs The similarity weight of .
[0273] It can be understood that the greater the difference between the first coding feature and the second coding feature, the greater the local structure loss; and the smaller the difference between the first coding feature and the second coding feature, the smaller the local structure loss.
[0274] Furthermore, a sample audio signal and a speaker label of the sample audio signal are obtained, and the initial structure-preserving encoder is used as an initial generator, and a target coding feature corresponding to the sample audio signal is generated based on the initial generator.
[0275] Then, based on the initial discriminator, the identity authentication category corresponding to the target encoding feature is determined.
[0276] Finally, based on the difference between the identity verification category and the speaker label, the discriminative loss is determined. The formula of the discriminative loss is as follows:
[0277]
[0278] in, represents the discriminative loss, Discriminator network, Indicates the speaker label.
[0279] Furthermore, a target loss can be determined based on the discriminative loss, the local structure loss, and the global distribution loss, and parameters of the initial structure preserving encoder can be iterated based on the target loss, and the initial structure preserving encoder after the parameter iteration is used as the structure preserving encoder.
[0280] Here, based on the discriminative loss, local structure loss, and global distribution loss, the formula for determining the target loss is as follows:
[0281]
[0282] in, represents the target loss, represents local structural loss, represents the global distribution loss, represents the discriminative loss, is the regularization term, is the adaptive weight coefficient.
[0283] In addition, the embodiment of the present invention can also perform adaptive coding optimization. In order to achieve the adaptability of the coding process, a dynamic weight adjustment mechanism is introduced:
[0284] Weight update strategy:
[0285] in, is the smoothing factor, is an adaptive function, is the loss value at the current moment.
[0286] Furthermore, in order to improve the stability of the coding, the embodiment of the present invention designs a coding space normalization strategy:
[0287] Normalization function:
[0288] in, is the scaling factor, through the temperature parameter Dynamic Adjustment:
[0289] The training of the encoding module adopts an end-to-end optimization method, which includes the following key steps:
[0290] Parameter optimization strategy: Parameter initialization uses a pre-trained encoder network to accelerate the convergence process. The optimization process uses the Adam optimizer, and the learning rate uses a cosine annealing strategy:
[0291]
[0292] in, Indicates The current learning rate at the training step, represents the initial learning rate (the preset baseline learning rate), Indicates the number of iterations of the current training. Represents the total number of iterations (a complete annealing cycle), and cos() represents the cosine function, which is used to periodically adjust the learning rate.
[0293] The weight coefficient of the loss function is dynamically adjusted based on the performance on the validation set:
[0294]
[0295] in, Indicates The dynamic weight coefficients of the loss terms (such as the weights of different tasks in multi-task learning), Indicates The performance indicator of the loss item on the validation set, SoftMax() represents the normalization function.
[0296] The convergence criterion uses the KL divergence change of the encoding space distribution:
[0297]
[0298] in, The model represents the hidden layer distribution, Represents the convergence threshold.
[0299] In the related technology, in the speaker verification system, the similarity calculation and judgment module is the last link of the decision chain, and its performance directly determines the verification accuracy of the system. There are three key problems in the traditional similarity calculation and judgment method: first, the simple distance measurement cannot fully capture the nonlinear distribution characteristics of the speaker's characteristics, resulting in inaccurate judgment results; second, the fixed judgment threshold is difficult to adapt to the verification needs in different scenarios, reducing the practicality of the system; third, the lack of a reliable confidence assessment mechanism makes it difficult to quantify the credibility of the verification results.
[0300] Based on in-depth thinking about these issues, this module proposes an adaptive multi-level similarity calculation and judgment framework. The core concept of this framework is to achieve more reliable identity recognition by building a multi-dimensional similarity measurement system and combining it with an adaptive decision-making mechanism. At the similarity calculation level, we are no longer limited to a single Euclidean distance or cosine similarity, but have designed a comprehensive measurement system that includes local structural similarity, global distribution similarity, and semantic similarity. This multi-dimensional similarity calculation method can more comprehensively characterize the similarity of speaker characteristics and provide a more accurate basis for judgment.
[0301] At the decision strategy level, we introduced an adaptive decision mechanism based on confidence. This mechanism achieves adaptive adjustment of the decision threshold by dynamically evaluating multiple factors such as the verification environment, feature quality, and similarity distribution. In particular, we designed an innovative confidence evaluation framework that can provide a reliable confidence score for each verification result, which is of great significance for risk control in practical applications.
[0302] Figure 7is a flow chart of the similarity calculation and determination module provided by the present invention, such as Figure 7 As shown, in the embodiment of the present invention, the similarity score can be determined based on the local structural similarity, global distribution similarity and semantic similarity between all coding features, and the weight coefficient, wherein the weight coefficient is dynamically adjusted based on the adaptive mechanism. In other words, multi-dimensional similarity calculation is performed based on the input feature pair.
[0303] The formula for the similarity score is as follows:
[0304]
[0305] Among them, the local structure similarity Defined as:
[0306] in, represents the first encoding feature, represents the second encoding feature, Represents standard deviation.
[0307] Global distribution similarity Use distribution alignment to calculate:
[0308] in, represents the output distribution corresponding to the first encoded feature, Represents the output distribution corresponding to the second encoded feature.
[0309] Semantic Similarity Through deep feature extraction network calculate:
[0310] in, represents the semantic feature corresponding to the first encoding feature, Indicates the semantic feature corresponding to the second encoding feature.
[0311] The weight coefficient ω is dynamically adjusted through an adaptive mechanism:
[0312] in, is the environment perception function, Represents context information.
[0313] It should be noted that , and Applicable to both Adaptive mechanism.
[0314] After obtaining the similarity score, the identity authentication result may be determined based on the similarity score. For example, the confidence score may be determined based on the similarity score, and then the identity authentication result may be determined based on the similarity score and the confidence score.
[0315] Specifically, confidence evaluation adopts a multi-factor fusion approach, that is, fusion with feature quality score, formally expressed as:
[0316]
[0317] in, Score the quality of the feature, is the similarity score, For environmental condition assessment, is a nonlinear mapping function.
[0318] The confidence update is calculated recursively:
[0319] in, is the smoothing factor, which is adjusted dynamically based on the verification performance.
[0320] Adaptive decision strategy:
[0321] The decision strategy is based on a comprehensive decision of similarity score and confidence:
[0322]
[0323] in, represents the decision function, is the adaptive threshold function:
[0324] in, represents the basic judgment threshold, Represents the adjustment coefficient.
[0325] Decision risk assessment is achieved by minimizing expected risk:
[0326] in, is the loss function, is the true label, represents the decision function.
[0327] In summary, this method effectively captures the time-varying characteristics and complex correlations of speaker features by introducing a dynamic graph structure. At the same time, a multi-level feature extraction and fusion mechanism is designed to achieve the organic integration of acoustic features and prosodic features. In addition, the adaptive mechanism of the present invention significantly improves the robustness of the system in complex environments. These innovations not only improve the accuracy of speaker verification, but also enhance the practicality and reliability of the system, providing better technical solutions for applications in related fields.
[0328] Furthermore, in order to enhance the robustness of features, the present invention introduces a feature enhancement mechanism based on contrastive learning. This mechanism constructs different forms of data perturbations to train the system to learn stable feature representations that are independent of noise and channel changes. In particular, an innovative contrast loss function is designed, which not only considers the similarity between samples, but also introduces the constraint of temporal consistency to ensure that the extracted features remain consistent in the time dimension.
[0329] At the implementation level, the system adopts a modular design approach, decomposing the feature extraction process into three key steps: window division, feature calculation, and feature enhancement. Each step is equipped with a corresponding quality assessment mechanism. By real-time monitoring of the reliability of features, the system can adjust the extraction strategy in a timely manner to ensure the quality of the output features. At the same time, we have also designed an adaptive allocation mechanism for computing resources to dynamically adjust the allocation of computing resources according to the complexity of the input signal, optimizing computing efficiency while ensuring feature quality.
[0330] This multi-level, adaptive feature extraction framework not only improves the expressiveness of features, but also significantly enhances the adaptability of the system in complex environments. Through system verification, the module has shown excellent stability in various noise environments and channel conditions, providing high-quality feature input for subsequent processing modules. More importantly, this design idea breaks the limitations of traditional feature extraction methods and provides a new technical paradigm for the field of speech signal processing.
[0331] Based on any of the above embodiments, the system adopts a hierarchical module design, including six core functional modules that cooperate with each other: feature extraction module, graph construction module, multi-level feature extraction and fusion module, graph entropy analysis and key feature screening module, encoding module and similarity calculation and judgment module. These modules form a complete speaker verification processing link through a carefully designed interaction mechanism.
[0332] Figure 8 This is the second flow chart of the identity recognition method provided by the present invention, such as Figure 8As shown in the figure, the input layer first receives the original audio signal. These signals are then passed to the feature extraction layer, where the feature extraction module extracts acoustic and rhythmic features. In the graph construction layer, the graph construction module converts the extracted features into dynamic feature graphs. The multi-level feature extraction and fusion module of the feature fusion layer is responsible for completing the hierarchical extraction and fusion of features. In the feature optimization layer, the graph entropy analysis and key feature screening module evaluates and screens the importance of features. Finally, the encoding module and the similarity calculation and judgment module of the decision layer complete the final verification decision.
[0333] Information transmission between modules follows strict protocol design:
[0334] The feature extraction module first provides the original feature vector to the graph construction module to provide basic data support for the subsequent graph structure construction.
[0335] The graph construction module passes the generated dynamic feature graph to the multi-level feature extraction and fusion module to support hierarchical feature extraction.
[0336] The fusion results of the multi-level feature extraction and fusion module will be passed to the graph entropy analysis and key feature screening module for feature importance evaluation.
[0337] The selected key features are converted into low-dimensional embedding vectors by the encoding module. These embedding vectors are finally used by the similarity calculation and judgment module to generate verification results.
[0338] Module collaborative decision-making mechanism:
[0339] In order to ensure the scientific nature of the collaborative decision-making of each module, this system has established a multi-module collaborative decision-making mechanism based on confidence. This mechanism is implemented through a mathematical model:
[0340]
[0341] in, Represents the final verification result. Represents the weight coefficient of each module, Reflects the verification results of each module Confidence score of .
[0342] This mechanism can effectively balance the decision recommendations of each module and ultimately achieve the optimal verification result.
[0343] In summary, this method captures the time-varying characteristics and complex correlations of speaker features by constructing a dynamic graph structure, combining a multi-level feature extraction and fusion mechanism, integrating acoustic features and rhythmic features, and solving the problems of insufficient feature dynamics and insufficient use of multimodal information in traditional methods. The system includes six core modules: feature extraction, graph construction, multi-level feature fusion, graph entropy analysis, encoding and judgment. The feature extraction module uses adaptive windows and contrastive learning to enhance robustness, the graph construction module optimizes feature expression through node importance evaluation and multi-dimensional similarity measurement, and the graph entropy analysis and key feature screening module screens key discriminant features based on information entropy theory. In addition, the encoding module realizes a robust mapping from high dimensions to low dimensions through structural preservation constraints, and the similarity calculation and judgment module integrates multi-dimensional similarity and confidence evaluation to achieve adaptive decision-making. The present invention significantly improves the accuracy, robustness and practicality of speaker verification in complex environments, and is suitable for multiple fields such as financial security, intelligent interaction, and security monitoring.
[0344] The identity recognition device provided by the present invention is described below. The identity recognition device described below and the identity recognition method described above can be referenced to each other.
[0345] Based on any of the above embodiments, the present invention provides an identity recognition device, Fig. 9 is a schematic diagram of the structure of the identity recognition device provided by the present invention, such as Fig. 9 As shown, the device comprises:
[0346] An acquisition unit 910 is used to respectively acquire target features of at least two audio signals of the person to be tested;
[0347] A determination unit 920, configured to determine an identity recognition result of the person to be tested based on all the target features;
[0348] The target feature of the audio signal is determined based on the following steps:
[0349] Extracting acoustic features and prosodic features of the audio signal, determining a first score of the acoustic features, and determining a second score of the prosodic features; the first score includes a signal-to-noise ratio score and a stability score, the signal-to-noise ratio score is used to reflect the signal-to-noise ratio of the audio signal, and the stability score is used to reflect the stability of the audio signal; the second score includes a consistency score and a reliability score, the consistency score is used to reflect the consistency of the audio signal with other audio signals, and the reliability score is used to reflect the reliability of the audio signal;
[0350] A feature quality score is determined based on the first score and the second score, and the target feature is determined based on the feature quality score and an interaction feature of the acoustic feature and the prosodic feature.
[0351] The device provided by the embodiment of the present invention obtains the target features of at least two audio signals of the person to be tested respectively, and then determines the identity recognition result of the person to be tested based on all the target features; wherein the target features of the audio signal are determined based on the following steps: extracting the acoustic features and prosodic features of the audio signal, determining the first score of the acoustic features, and determining the second score of the prosodic features; the first score includes a signal-to-noise ratio score and a stability score, and the second score includes a consistency score and a reliability score, and finally, based on the first score and the second score, determining the feature quality score, and based on the feature quality score and the interactive features of the acoustic features and the prosodic features, determining the target features. This process makes full use of multimodal information such as the signal-to-noise ratio score, the stability score, the consistency score, and the reliability score, effectively fuses the acoustic features and the prosodic features, and determines the identity recognition result based on all the target features obtained by the fusion, thereby improving the accuracy and reliability of speaker recognition.
[0352] Based on any of the above embodiments, it further includes a feature acquisition unit, wherein the feature acquisition unit is specifically used to:
[0353] Determining the acoustic feature based on the Mel-frequency cepstral coefficient feature, fundamental frequency feature and energy feature of the audio signal;
[0354] The prosody feature is determined based on the intonation feature, the rhythm feature and the duration feature of the audio signal.
[0355] Based on any of the above embodiments, the Mel-frequency cepstral coefficient feature, the fundamental frequency feature and the energy feature are all extracted from the audio signal based on an adaptive window size;
[0356] The intonation feature, the rhythm feature and the duration feature are all extracted from the audio signal based on an adaptive window size;
[0357] The adaptive window size determination unit is specifically used for:
[0358] Determining an adaptive coefficient based on local energy, zero crossing rate, and sound change rate;
[0359] The adaptive window size is determined based on the adaptive coefficient and the basic window size.
[0360] Based on any of the above embodiments, the determining unit 920 specifically includes:
[0361] A first input unit, used to input all the target features into a graph structure to obtain graph structure features output by the graph structure;
[0362] An extraction unit, used to extract the bottom-level features, middle-level features and high-level features of the graph structure features respectively, and determine multi-layer fusion features based on the bottom-level features, the middle-level features and the high-level features;
[0363] A feature screening unit, used to screen the multi-layer fusion features based on the correlation measurement and redundancy measurement of the nodes in the graph structure to obtain screening features;
[0364] A second input unit, used for inputting the screening features into a structure preserving encoder to obtain all encoding features output by the structure preserving encoder;
[0365] A similarity calculation unit is used to determine the identity recognition result based on the similarities between all the coding features.
[0366] Based on any of the above embodiments, the similarity calculation unit specifically includes:
[0367] A similarity score determination unit is used to determine a similarity score based on the local structural similarity, global distribution similarity and semantic similarity between all the coding features, and a weight coefficient; the weight coefficient is dynamically adjusted based on an adaptive mechanism;
[0368] The identity recognition result determining unit is used to determine the identity recognition result based on the similarity score.
[0369] Based on any of the above embodiments, the unit for determining the identity recognition result is specifically used to:
[0370] Based on the similarity score, determining a confidence score;
[0371] The identity recognition result is determined based on the similarity score and the confidence score.
[0372] Based on any of the above embodiments, a training unit is further included, and the training unit specifically includes:
[0373] Acquiring a sample unit, for acquiring an initial structure preserving encoder, a sample screening feature, and a prior distribution of the sample screening feature;
[0374] A sample input unit, used for inputting the sample screening feature into the initial structure preserving encoder to obtain a prediction coding feature input by the initial structure preserving encoder;
[0375] A global distribution loss determination unit, configured to determine a global distribution loss based on an actual distribution corresponding to the predicted coding feature and the prior distribution;
[0376] A parameter iteration unit is used to perform parameter iteration on the initial structure preserving encoder based on the global distribution loss to obtain the structure preserving encoder.
[0377] Based on any of the above embodiments, the parameter iteration unit is specifically used to:
[0378] Obtaining a first sample screening feature and a second sample screening feature;
[0379] Inputting the first sample screening feature into the initial structure preserving encoder to obtain a first encoding feature output by the initial structure preserving encoder;
[0380] Inputting the second sample screening feature into the initial structure preserving encoder to obtain a second encoding feature output by the initial structure preserving encoder;
[0381] determining a local structural loss based on a difference between the first encoding feature and the second encoding feature;
[0382] Acquire a sample audio signal and a speaker label of the sample audio signal;
[0383] Using the initial structure-maintaining encoder as an initial generator, and generating a target coding feature corresponding to the sample audio signal based on the initial generator;
[0384] Determining, based on the initial discriminator, the identity authentication category corresponding to the target coding feature;
[0385] Based on the difference between the identity verification category and the speaker label, a discriminative loss is determined, and based on the discriminative loss, the local structure loss and the global distribution loss, a target loss is determined, and parameters of the initial structure-preserving encoder are iterated based on the target loss to obtain the structure-preserving encoder.
[0386] Based on any of the above embodiments, the edge weight matrix of the graph structure is determined based on acoustic similarity, prosodic similarity, temporal similarity and contextual similarity between two nodes in the graph structure;
[0387] The acoustic similarity is determined based on acoustic features between two nodes in the graph structure;
[0388] The prosodic similarity is determined based on prosodic features between any two nodes in the graph structure;
[0389] The temporal similarity is determined based on the timestamps corresponding to each two nodes in the graph structure;
[0390] The context similarity is determined based on the context relevance between any two nodes in the graph structure.
[0391] Fig.10 is a schematic diagram of the structure of the electronic device provided by the present invention, such as Fig.10 As shown, the electronic device may include: a processor (processor) 1010 , a communication interface (Communications Interface) 1020 , a memory (memory) 1030 and a communication bus 1040 , wherein the processor 1010 , the communication interface 1020 , and the memory 1030 communicate with each other via the communication bus 1040 . The processor 1010 can call the logic instructions in the memory 1030 to execute the identity recognition method, which includes: respectively obtaining the target features of at least two audio signals of the person to be tested; determining the identity recognition result of the person to be tested based on all the target features; wherein the target features of the audio signal are determined based on the following steps: extracting the acoustic features and prosodic features of the audio signal, determining a first score of the acoustic features, and determining a second score of the prosodic features; the first score includes a signal-to-noise ratio score and a stability score, the signal-to-noise ratio score is used to reflect the signal-to-noise ratio of the audio signal, and the stability score is used to reflect the stability of the audio signal; the second score includes a consistency score and a reliability score, the consistency score is used to reflect the consistency of the audio signal with other audio signals, and the reliability score is used to reflect the reliability of the audio signal; based on the first score and the second score, determine the feature quality score, and based on the feature quality score and the interactive features of the acoustic features and the prosodic features, determine the target feature.
[0392] In addition, the logic instructions in the above-mentioned memory 1030 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.
[0393] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the identity recognition method provided by the above methods, the method including: respectively obtaining target features of at least two audio signals of the person to be tested; based on all the target features, determining the identity recognition result of the person to be tested; wherein the target feature of the audio signal is determined based on the following steps: extracting the acoustic features and prosodic features of the audio signal, determining a first score of the acoustic feature, and determining a second score of the prosodic feature; the first score includes a signal-to-noise ratio score and a stability score, the signal-to-noise ratio score is used to reflect the signal-to-noise ratio of the audio signal, and the stability score is used to reflect the stability of the audio signal; the second score includes a consistency score and a reliability score, the consistency score is used to reflect the consistency of the audio signal with other audio signals, and the reliability score is used to reflect the reliability of the audio signal; based on the first score and the second score, a feature quality score is determined, and based on the feature quality score and the interactive features of the acoustic feature and the prosodic feature, the target feature is determined.
[0394] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which is implemented when the processor executes the identity recognition method provided by the above-mentioned methods, the method comprising: respectively obtaining target features of at least two audio signals of the person to be tested; determining the identity recognition result of the person to be tested based on all the target features; wherein the target features of the audio signal are determined based on the following steps: extracting the acoustic features and prosodic features of the audio signal, determining a first score of the acoustic features, and determining a second score of the prosodic features; the first score includes a signal-to-noise ratio score and a stability score, the signal-to-noise ratio score is used to reflect the signal-to-noise ratio of the audio signal, and the stability score is used to reflect the stability of the audio signal; the second score includes a consistency score and a reliability score, the consistency score is used to reflect the consistency of the audio signal with other audio signals, and the reliability score is used to reflect the reliability of the audio signal; based on the first score and the second score, a feature quality score is determined, and based on the feature quality score and the interactive features of the acoustic features and the prosodic features, the target features are determined. The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0395] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0396] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An identity recognition method, characterized in that: include: respectively obtaining target features of at least two audio signals of the person to be tested; Wherein, the target features of the audio signal include: Extracting acoustic features and prosodic features of the audio signal, determining a first score of the acoustic features, and determining a second score of the prosodic features; the first score includes a signal-to-noise ratio score and a stability score; the second score includes a consistency score and a reliability score; the stability score is determined by a stability evaluation function, and the stability evaluation function measures the degree of variation of features in different samples by a standard deviation or a coefficient of variation; The formula for the reliability score is as follows: ; in, represents the reliability score, Represents rhythmic features, is the variance of the current feature in the short time window, is the attenuation coefficient; The formula for the consistency score is as follows: ; in, represents the consistency score, Represents rhythmic features, For the i reference templates, N is the number of templates; Determine a feature quality score based on the first score and the second score, and determine the target feature based on the feature quality score and an interaction feature between the acoustic feature and the prosodic feature; Based on all the target features, determining the identity recognition result of the person to be tested includes: Inputting all the target features into a graph structure to obtain graph structure features output by the graph structure; Extracting bottom-level features, middle-level features and high-level features of the graph structure features respectively, and determining multi-layer fusion features based on the bottom-level features, the middle-level features and the high-level features; Based on the correlation measurement and redundancy measurement of the nodes in the graph structure, feature screening is performed on the multi-layer fusion features to obtain screening features; Inputting the screening features into a structure preserving encoder to obtain all encoding features output by the structure preserving encoder; Determining a similarity score based on the local structural similarity, global distribution similarity and semantic similarity between all the coding features, and a weight coefficient; the weight coefficient is dynamically adjusted based on an adaptive mechanism; Determining the identity recognition result based on the similarity score; The graph structure The formula is as follows: ; in, For a node collection: , For edge sets: , is the edge weight matrix: ; The core of the structure-preserving encoder is a nonlinear mapping function , nonlinear mapping function Mapping the input feature x to a low-dimensional space can be formally expressed as: ; in, represents a parameterized mapping network, is the network parameter.
2. The identity recognition method according to claim 1, characterized in that: The step of acquiring the acoustic features and the prosodic features comprises: Determining the acoustic feature based on the Mel-frequency cepstral coefficient feature, fundamental frequency feature and energy feature of the audio signal; The prosody feature is determined based on the intonation feature, the rhythm feature and the duration feature of the audio signal.
3. The identity recognition method according to claim 2, characterized in that: The Mel frequency cepstral coefficient feature, the fundamental frequency feature and the energy feature are all extracted from the audio signal based on an adaptive window size; The intonation feature, the rhythm feature and the duration feature are all extracted from the audio signal based on an adaptive window size; The step of determining the adaptive window size comprises: Determining an adaptive coefficient based on local energy, zero crossing rate, and sound change rate; The adaptive window size is determined based on the adaptive coefficient and the basic window size.
4. The identity recognition method according to claim 1, characterized in that: The determining the identity recognition result based on the similarity score includes: Based on the similarity score, determining a confidence score; The identity recognition result is determined based on the similarity score and the confidence score.
5. The identity recognition method according to claim 1, characterized in that: The structure maintains the encoder training steps, including: Obtaining an initial structure preserving encoder, a sample screening feature, and a prior distribution of the sample screening feature; Inputting the sample screening feature into the initial structure preserving encoder to obtain the prediction coding feature input by the initial structure preserving encoder; Determining a global distribution loss based on an actual distribution corresponding to the predicted coding feature and the prior distribution; Based on the global distribution loss, parameters of the initial structure preserving encoder are iterated to obtain the structure preserving encoder.
6. The identity recognition method according to claim 5, characterized in that: The step of performing parameter iteration on the initial structure preserving encoder based on the global distribution loss to obtain the structure preserving encoder comprises: Obtaining a first sample screening feature and a second sample screening feature; Inputting the first sample screening feature into the initial structure preserving encoder to obtain a first encoding feature output by the initial structure preserving encoder; Inputting the second sample screening feature into the initial structure preserving encoder to obtain a second encoding feature output by the initial structure preserving encoder; determining a local structural loss based on a difference between the first encoding feature and the second encoding feature; Acquire a sample audio signal and a speaker label of the sample audio signal; Using the initial structure-maintaining encoder as an initial generator, and generating a target coding feature corresponding to the sample audio signal based on the initial generator; Determining, based on the initial discriminator, the identity authentication category corresponding to the target coding feature; Based on the difference between the identity verification category and the speaker label, a discriminative loss is determined, and based on the discriminative loss, the local structure loss and the global distribution loss, a target loss is determined, and parameters of the initial structure-preserving encoder are iterated based on the target loss to obtain the structure-preserving encoder.
7. The identity recognition method according to claim 1, characterized in that: The edge weight matrix of the graph structure is determined based on acoustic similarity, prosodic similarity, temporal similarity and contextual similarity between two nodes in the graph structure; The acoustic similarity is determined based on acoustic features between two nodes in the graph structure; The prosodic similarity is determined based on prosodic features between any two nodes in the graph structure; The temporal similarity is determined based on the timestamps corresponding to each two nodes in the graph structure; The context similarity is determined based on the context relevance between any two nodes in the graph structure.
8. An identity recognition device, characterized in that: include: An acquisition unit, used to respectively acquire target features of at least two audio signals of the person to be tested; Wherein, the target features of the audio signal include: Extracting acoustic features and prosodic features of the audio signal, determining a first score of the acoustic features, and determining a second score of the prosodic features; the first score includes a signal-to-noise ratio score and a stability score; the second score includes a consistency score and a reliability score; the stability score is determined by a stability evaluation function, and the stability evaluation function measures the degree of variation of features in different samples by a standard deviation or a coefficient of variation; The formula for the reliability score is as follows: ; in, represents the reliability score, Represents rhythmic features, is the variance of the current feature in the short time window, is the attenuation coefficient; The formula for the consistency score is as follows: ; in, represents the consistency score, Represents rhythmic features, For the i reference templates, N is the number of templates; Determine a feature quality score based on the first score and the second score, and determine the target feature based on the feature quality score and an interaction feature between the acoustic feature and the prosodic feature; Inputting all the target features into a graph structure to obtain graph structure features output by the graph structure; Extracting bottom-level features, middle-level features and high-level features of the graph structure features respectively, and determining multi-layer fusion features based on the bottom-level features, the middle-level features and the high-level features; Based on the correlation measurement and redundancy measurement of the nodes in the graph structure, feature screening is performed on the multi-layer fusion features to obtain screening features; Inputting the screening features into a structure preserving encoder to obtain all encoding features output by the structure preserving encoder; Determining a similarity score based on the local structural similarity, global distribution similarity and semantic similarity between all the coding features, and a weight coefficient; the weight coefficient is dynamically adjusted based on an adaptive mechanism; Determining the identity recognition result based on the similarity score; The graph structure The formula is as follows: ; in, For a node collection: , For edge sets: , is the edge weight matrix: ; The core of the structure-preserving encoder is a nonlinear mapping function , nonlinear mapping function Mapping the input feature x to a low-dimensional space can be formally expressed as: ; in, represents a parameterized mapping network, is the network parameter.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the identity recognition method according to any one of claims 1 to 7 is implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the identity recognition method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Voiceprint recognition method, device and keyless safety lock system and implementing method
CN104575492A
Speech scoring method and device, electronic device, and storage medium
CN109256152A