Feature acquisition method and device

By preprocessing and feature extraction of ear thread data and extracting ear thread features using filters, the problem of low reliability of ear thread features in the prior art is solved, and higher identity recognition accuracy is achieved.

CN119993194AActive Publication Date: 2025-05-13NANJING ZGMICRO CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510247201.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-05-13
Estimated Expiration
2045-03-04

AI Technical Summary

Technical Problem

In the prior art, the reliability of extracting ear thread features is low and it is difficult to meet the application scenarios of high safety requirements.

Method used

By preprocessing the user's target ear pattern data, including noise reduction processing, pre-emphasis processing, frame processing and window processing, it is converted into spectrum data, and the band energy value is extracted based on the pre-set N target filters, and feature conversion is performed to obtain M ear pattern feature data.

Benefits of technology

The reliability of ear thread feature extraction is improved, and the user identity is distinguished by ear thread features that reflect the personalized ear thread characteristics of each user, forming an identity identifier for uniquely indicating each user.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993194A_ABST
    Figure CN119993194A_ABST
Patent Text Reader

Abstract

The invention provides a feature acquisition method and device, and particularly relates to the technical field of voice signal processing, and the method comprises the steps: converting the preprocessed target earprint data of a user into frequency spectrum data, and extracting the energy of the frequency spectrum data based on N preset target filters, and obtaining N frequency band energy values, the center frequencies of the N target filters are symmetrically arranged by taking a set reference frequency as a center, the reference frequency is contained in a frequency band corresponding to a frequency response peak value of earprint data of a target user group, and the target user group comprises the user; the frequency band energy value is obtained based on convolution operation of the corresponding target filter and the frequency spectrum data in the signal part of the corresponding frequency band; the N frequency band energy values are subjected to feature conversion, M earprint feature data are obtained, and the identity label of the user comprises the M earprint feature data. According to the method, more reliable earprint features are extracted from the target earprint data of the user through the N target filters.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of speech signal processing, and in particular to a feature acquisition method and device. Background Art

[0002] Among the existing biometric technologies, voiceprint recognition has been widely used in scenarios such as identity authentication and voice control. However, voiceprint recognition is greatly affected by environmental noise and the speaker's state (such as emotions, health, etc.), and it is difficult to fully meet application requirements in scenarios with high security requirements. In contrast, earprint recognition has higher stability and individual specificity by analyzing the characteristics of ear canal echo signals, and has gradually become a potential complementary or alternative technology.

[0003] But for now, the reliability of earprint features extracted based on related technologies is lower than expected. Summary of the invention

[0004] The purpose of the present application is to provide a feature acquisition method and device for solving the technical problem of low reliability of earprint features extracted by related technologies.

[0005] In a first aspect, an embodiment of the present application provides a feature acquisition method, the method comprising:

[0006] Preprocessing the target earprint data of the user to obtain preprocessed data, wherein the preprocessing includes at least one of noise reduction processing, pre-emphasis processing, frame processing and windowing processing;

[0007] The preprocessed data is converted into spectrum data, and the energy of the spectrum data is respectively extracted based on N preset target filters to obtain N frequency band energy values, wherein the center frequencies of the N target filters are symmetrically arranged with a set reference frequency as the center, the reference frequency is included in the frequency band corresponding to the frequency response peak of the earprint data of the target user group, the target user group includes the user, and the frequency band energy value is obtained based on the convolution operation of the corresponding target filter and the signal part of the spectrum data in the corresponding frequency band;

[0008] Feature conversion is performed on the N frequency band energy values ​​respectively to obtain M earprint feature data, and the user's identity includes the M earprint feature data, wherein N and M are integers greater than 1.

[0009] In a second aspect, an embodiment of the present application further provides a feature acquisition device, the device comprising:

[0010] A preprocessing module, used to preprocess the target earprint data of the user to obtain preprocessed data, wherein the preprocessing includes at least one of noise reduction processing, pre-emphasis processing, frame processing and windowing processing;

[0011] an energy extraction module, for converting the preprocessed data into spectrum data, and extracting the energy of the spectrum data based on N preset target filters, respectively, to obtain N frequency band energy values, wherein the center frequencies of the N target filters are symmetrically arranged with a set reference frequency as the center, the reference frequency is included in a frequency band corresponding to a frequency response peak of the earprint data of a target user group, the target user group includes the user, and the frequency band energy value is obtained based on a convolution operation of a corresponding target filter and a signal portion of the spectrum data in a corresponding frequency band;

[0012] The feature conversion module is used to perform feature conversion on the N frequency band energy values ​​respectively to obtain M earprint feature data, and the user's identity includes the M earprint feature data, wherein N and M are integers greater than 1.

[0013] In a third aspect, the present disclosure provides an electronic device, comprising a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the method described in the first aspect.

[0014] In a fourth aspect, the present disclosure provides a computer program product, comprising computer instructions, which, when executed by a processor, implement the steps of the method described in the first aspect.

[0015] In the present application, a reference frequency is determined in the frequency band corresponding to the frequency response peak of the earprint data of the target user group, and the center frequencies of N target filters are symmetrically set around the reference frequency. More reliable earprint features are extracted from the target earprint data of the user through the N target filters, and the user identity is distinguished by the earprint features that reflect the personalized earprint characteristics of each user, thereby forming an identity identifier that uniquely indicates each user. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 is a flowchart of a feature acquisition method provided by an embodiment of the present disclosure;

[0017] Figure 2 It is a flowchart of an MFCC feature extraction method based on earprint characteristics provided by an embodiment of the present disclosure;

[0018] Figure 3 is a schematic diagram of a change in the frequency response amplitude of the left and right ears of a user provided by an embodiment of the present disclosure;

[0019] Figure 4 is a distribution diagram of a Mel filter bank provided by an embodiment of the present disclosure;

[0020] Figure 5is a structural schematic diagram of a feature acquisition device provided by an embodiment of the present disclosure;

[0021] Figure 6 It is a schematic diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0022] The following will be combined with the drawings in the embodiments of the present disclosure to clearly and completely describe the technical solutions in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, rather than all of the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present disclosure.

[0023] The present application embodiment provides a feature acquisition method, such as Figure 1 As shown, the method includes:

[0024] Step 101: pre-process the target earprint data of the user to obtain pre-processed data.

[0025] The preprocessing includes at least one of noise reduction processing, pre-emphasis processing, frame processing and windowing processing.

[0026] In the present application, the target earprint data of the user is specifically the echo signal of the user's ear canal, wherein the echo signal of the user's ear canal can be collected by any target device (such as headphones) with audio playback and recording functions.

[0027] For example, when the target device is set to headphones, the process of collecting the echo signal of the user's ear canal is as follows: the user wears the headphones, and the headphones play a 2-second Gaussian white noise. During the playback process, the headphones simultaneously collect the echo signal of the Gaussian white noise after it is reflected by the user's left ear canal and / or right ear canal.

[0028] Noise reduction processing is used to remove noise from the echo signal to improve the clarity and comprehensibility of the signal portion of the echo signal associated with the user's personalized characteristics. In this application, noise is not only used to indicate signal interference caused by the external environment, but also used to indicate the signal portion of the echo signal that cannot be used to express the user's personalized characteristics.

[0029] Pre-emphasis processing is used to enhance the energy of the high-frequency part of the echo signal to compensate for the loss of the high-frequency part of the echo signal in the subsequent processing process, to adapt to the characteristics of different people's earprint differences concentrated in the high-frequency part of the echo signal, and to improve the clarity and comprehensibility of the signal part of the echo signal associated with the user's personalized characteristics (indicating the earprint differences of different users).

[0030] The frame processing is used to divide the continuous echo signal into short time frames so that the echo signal remains stable within each short time frame, so as to facilitate the subsequent operation of extracting earprint feature data from the echo signal.

[0031] Windowing is used to smooth the edge of each short time frame after framing to reduce the discontinuity of the frame edge and suppress the risk of spectrum leakage, making the subsequent frequency domain conversion and analysis of the echo signal more accurate.

[0032] Step 102: convert the preprocessed data into spectrum data, and extract the energy of the spectrum data based on N preset target filters to obtain N frequency band energy values.

[0033] Among them, the center frequencies of the N target filters are symmetrically arranged with a set reference frequency as the center, the reference frequency is included in the frequency segment corresponding to the frequency response peak of the earprint data of the target user group, the target user group includes the user, and the frequency band energy value is obtained based on the convolution operation of the corresponding target filter and the signal part of the spectrum data in the corresponding frequency band.

[0034] In the present application, the target user group can be a set of several users who meet the set user selection rules, wherein the set user selection rules can be adaptively adjusted according to actual needs, for example: users engaged in specific professions are determined as users in the target user group, or users located in a specific space are determined as users in the target user group, etc.

[0035] The earprint data of the target user group includes the echo signal of the left ear canal of each user in the target user group and / or the echo signal of the right ear canal of each user in the target user group. The echo signals referred to here can also be collected by the aforementioned target device.

[0036] It should be understood that the frequency response peak of the echo signal of the left ear canal of each user in the target user group and / or the frequency response peak of the echo signal of the right ear canal of each user in the target user group constitute a peak data set; in the peak data set, the proportion of data in the frequency band corresponding to the frequency response peak of the earprint data of the target user group is greater than a set ratio threshold (such as 0.9, 0.95, 0.99, etc.)

[0037] In one example, the frequency response peak of the echo signal of the left ear canal of each user in the target user group can be taken as the center, and the set frequency band width can be used as the distance from the center to the edge to obtain the frequency band corresponding to the frequency response peak of the echo signal of the left ear canal of each user in the target user group. In a similar manner, the frequency band corresponding to the frequency response peak of the echo signal of the right ear canal of each user in the target user group can also be obtained; the frequency band corresponding to the frequency response peak of the echo signal of the left ear canal of each user in the target user group and the frequency band corresponding to the frequency response peak of the echo signal of the right ear canal of each user in the target user group are deintersected or unioned to obtain the frequency band corresponding to the frequency response peak of the earprint data of the target user group.

[0038] In another example, the frequency response peak of the echo signal of the left ear canal and / or the average of the frequency response peak of the echo signal of the right ear canal of each user in the target user group can be used as the target center, and the set frequency band width can be used as the distance from the target center to the edge to obtain the frequency band corresponding to the frequency response peak of the earprint data of the target user group.

[0039] Exemplarily, the reference frequency can be randomly selected within the frequency segment corresponding to the frequency response peak of the earprint data of the target user group, or it can be the median point frequency within the frequency segment corresponding to the frequency response peak of the earprint data of the target user group, or it can be the frequency of the corresponding frequency response peak point within the frequency segment corresponding to the frequency response peak of the earprint data of the target user group.

[0040] Step 103: Perform feature conversion on the N frequency band energy values ​​respectively to obtain M earprint feature data.

[0041] The user's identity includes the M earprint feature data, where N and M are integers greater than 1.

[0042] In the present application, a reference frequency is determined in the frequency band corresponding to the frequency response peak of the earprint data of the target user group, and the center frequencies of N target filters are symmetrically set around the reference frequency. More reliable earprint features are extracted from the target earprint data of the user through the N target filters, and the user identity is distinguished by the earprint features that reflect the personalized earprint characteristics of each user, thereby forming an identity identifier that uniquely indicates each user.

[0043] The M earprint feature data may directly constitute the identity of the user, or may be combined with other identity features of the user (such as voiceprint features, facial features, etc.) to form the identity of the user.

[0044] Exemplarily, based on the user's identity identifier, the user's identity registration in the identity information database can be completed, and the user's identity authentication can also be completed. This application does not limit the specific application of the user's identity identifier.

[0045] In one embodiment, the N frequency band energy values ​​are respectively subjected to feature conversion to obtain M earprint feature data, including:

[0046] The N frequency band energy values ​​are screened based on the noise frequency interval to obtain N' target frequency band energy values, wherein the N' target frequency band energy values ​​include: among the N frequency band energy values, all frequency band energy values ​​that are outside the noise frequency interval, within the noise frequency interval, the similarity between the first frequency response curve and the second frequency response curve is less than or equal to a set difference threshold, the first frequency response curve is used to represent the frequency response trend of the earprint data of the left ear of the target user group with frequency, and the second frequency response curve is used to represent the frequency response trend of the earprint data of the right ear of the target user group with frequency;

[0047] The N' target frequency band energy values ​​are respectively subjected to feature conversion to obtain the M earprint feature data, where N' is a positive integer less than or equal to N, wherein each earprint feature data is obtained based on the N' target frequency band energy values ​​and corresponding feature dimension parameters, and different earprint feature data have different corresponding feature dimension parameters.

[0048] In this embodiment, the noise frequency interval is also the frequency interval where the aforementioned noise is located. Based on the above settings, the interference caused by the noise is suppressed, and the reliability of the M earprint feature data finally obtained is further improved.

[0049] It should be pointed out that by determining the noise frequency interval based on the similarity comparison between curves, the frequency response characteristics reflected by the earprint data of each user in the target user group can be fully and effectively utilized to avoid interference from extreme point values ​​and ensure the accuracy of the determined noise frequency interval.

[0050] Exemplarily, the frequency response of the left ear of each user in the target user group at different frequencies and the frequency response of the right ear of each user in the target user group at different frequencies can be processed by curve fitting to obtain the first frequency response curve and the second frequency response curve respectively.

[0051] In one embodiment, the performing feature conversion on the N' target frequency band energy values ​​to obtain the M earprint feature data includes:

[0052] Performing logarithmic transformation on the N' target frequency band energy values ​​respectively to obtain N' logarithmic data, where the logarithmic data is the logarithmic value of the sum of the corresponding target frequency band energy value and the set parameter;

[0053] According to the set M feature dimension parameters, discrete cosine transform (DCT) is performed on the N' logarithmic data respectively to obtain the M earprint feature data.

[0054] In this embodiment, the processing based on logarithmic transformation is used to achieve compression of a wide range of energy values, thereby suppressing the excessive dominance of relatively stronger frequency components in the target frequency band energy values ​​over the feature expression; while the processing based on discrete cosine transformation can compress the overall information of the logarithmic data while highlighting the earprint features contained in the logarithmic data, thereby increasing the accuracy and reliability of the M earprint feature data obtained.

[0055] The setting parameter is used to avoid the problem of zero logarithmic value. In the application, a very small floating point number can be used as the setting parameter to avoid the problem of zero logarithmic value while reducing the numerical impact of the addition of the setting parameter on the logarithmic value finally calculated.

[0056] By setting the feature dimension parameters, dimensional information of a specific dimension can be extracted from the N' target frequency band energy values ​​to make full use of the N' target frequency band energy values ​​to form comprehensive and accurate earprint feature data; it should be noted that as the value of M increases, the M earprint feature data will have a more accurate representation of the user's earprint features, but the calculation complexity will also increase accordingly. In the application, the specific value of M can be adaptively selected according to actual needs, such as 12 or 13.

[0057] In one embodiment, among the N target filters, the bandwidth of each target filter is positively correlated with a distance parameter corresponding to the target filter, and the distance parameter is used to indicate a frequency difference between a center frequency of the corresponding target filter and the reference frequency.

[0058] In this embodiment, based on the above settings, N target filters are used to simulate the changes in the audio perception ability of the human ear as much as possible (the audio perception ability of the human ear is strongest near the reference frequency, and as the distance from the reference frequency increases, the audio perception ability gradually decreases), so that the energy values ​​of the N frequency bands are more accurate and effective.

[0059] In one embodiment, the target filter is a Mel filter, the N target filters include a first target filter and a second target filter, the center frequency of the first target filter is higher than the reference frequency, and the center frequency of the second target filter is lower than the reference frequency;

[0060] Among them, the central Mel frequencies of multiple first target filters are distributed at equidistant intervals.

[0061] In this embodiment, the plurality of first target filters are evenly set by using the Mel frequency as a reference to avoid the problem of uneven characteristic distribution of the echo signal at the conventional frequency (specifically: the earprint characteristics carried by the echo signal at the conventional frequency are densely distributed in the low frequency band and sparsely distributed in the high frequency band), thereby ensuring the accuracy of the frequency band energy values ​​extracted by the plurality of first target filters; and referring to the settings of the plurality of first target filters, the plurality of second target filters are mirrored with the reference frequency as the mirror center, which can ensure the accuracy of the frequency band energy values ​​extracted by the plurality of second target filters while making the settings of the plurality of second target filters more convenient.

[0062] For example, if the reference frequency is set to 4000, the first target filter and the second target filter satisfy the following relationship:

[0063] a n -4000=4000-b n

[0064] Wherein, a is the center frequency of the first target filter, b is the center frequency of the second target filter, and n is the index of the corresponding filter.

[0065] In one embodiment, the preprocessing of the earprint data of the user to obtain preprocessed data includes:

[0066] According to a preset short-time energy threshold and a zero-crossing rate threshold, a noise reduction process is performed on a plurality of earprint audio signals included in the earprint data to obtain at least one target audio signal, wherein the short-time energy of the target audio signal is greater than or equal to the short-time energy threshold, and the zero-crossing rate of the target audio signal is less than or equal to the zero-crossing rate threshold;

[0067] Among the at least one target audio signal, performing pre-emphasis processing on each of the target audio signals to obtain at least one pre-emphasized audio signal, wherein the pre-emphasis processing includes first-order difference processing;

[0068] Performing frame processing on the at least one pre-emphasized audio signal to obtain frame signal data;

[0069] Performing windowing processing on the frame signal data based on a set window function to obtain the preprocessed data, wherein the window function includes a Hamming window function;

[0070] The frequency spectrum data is obtained by fast Fourier transform based on the preprocessed data.

[0071] In this embodiment, based on the setting of the short-time energy threshold and the zero-crossing rate threshold, the unvoiced segments in the earprint data are identified and filtered, and the voiced segments that can more accurately reflect the harmonic characteristics of the ear canal and facilitate feature extraction are retained.

[0072] For ease of understanding, the following examples are provided:

[0073] like Figure 2 As shown, the present application also provides a Mel-frequency Cepstral Coefficients (MFCC) feature extraction method based on earprint characteristics, the method comprising the following steps:

[0074] Step 201: Collect user earprint data.

[0075] Specifically, the user's ear canal echo signal is collected as earprint data, and the collection work can be performed using the aforementioned target device.

[0076] Step 202: VAD extracts voiced sound signals.

[0077] Specifically, the ear canal echo signal collected in step 201 is framed, the short-time energy and zero-crossing rate of each frame signal are calculated, and by setting thresholds of the short-time energy and the zero-crossing rate (which can be understood as short-time energy thresholds and zero-crossing rate thresholds), the unvoiced segment signals are identified and filtered, and only the voiced segment signals with higher short-time energy and lower zero-crossing rate are retained. Since the retained voiced part can more accurately reflect the harmonic characteristics of the ear canal, it is convenient for feature extraction and the extracted earprint features can be subsequently used for identity recognition.

[0078] Step 203: pre-emphasis.

[0079] Specifically, a pre-emphasis filter is applied to the voiced segment signal retained in step 202, and a first-order difference operation can be performed on the voiced end signal, and the formula is:

[0080] y[n]=x[n]-α·x[n-1]

[0081] Among them, x[n] represents the current sampling point of the input signal, x[n-1] represents the previous sampling point, and α is the pre-emphasis coefficient, which is set to 0.97 and is used to control the degree of enhancement of the high-frequency part. The key signal in the earprint signal usually has more energy distribution in the high-frequency part, and the earprint differences between different people are also mainly concentrated in the high-frequency part. Therefore, through pre-emphasis processing, the high-frequency part of the voiced segment signal can be made more prominent.

[0082] Step 204: Frame division.

[0083] Specifically, the non-stationary signal processed in step 203 is divided into short time frames of 16ms. At this time, each frame signal after framing is considered to be stationary, the frame shift can be set to 10ms, and each frame signal has 37.5% overlap to ensure continuity between frames and reduce information loss between frames.

[0084] Step 205: Add window.

[0085] Specifically, a window function is applied to each frame of signal samples after framing to reduce the distortion effect of the signal at the frame boundary. Here, a Hamming window can be used, and the definition formula is as follows:

[0086]

[0087] Where ω[n] is the weight of the window function, N is the number of sample points, and n is the sample number in the frame. By multiplying the window function point by point with each frame of signal, the signal boundary can be smoothed and the spectrum leakage effect can be reduced.

[0088] Step 206: FFT transformation.

[0089] Specifically, a fast Fourier transform (FFT) is performed on the windowed signal of each frame to convert the time domain signal into a frequency domain signal.

[0090] The formula is as follows:

[0091]

[0092] Where X[k] is the complex value of the kth frequency component of the frequency domain signal, x[n] is the nth sampling point of the time domain signal, and N is the number of sampling points per frame of the signal, where N = 512. Through FFT calculation, the spectrum amplitude and phase information of each frame of the signal can be obtained.

[0093] Step 207: Set a 4-8kHz filter.

[0094] Specifically, the earprint spectrum obtained in step 206 is used to set a Mel filter with a frequency of 4000 Hz as a starting point and a frequency of 8000 Hz as a cutoff point.

[0095] The user's earprint data statistics can be referenced Figure 3 , Figure 3In the range of 2000Hz to 8000Hz, the frequency response amplitude of earprint data is relatively large. Therefore, in the frequency range of 4000Hz to 8000Hz, 7 Mel filters (which can be understood as multiple first target filters) are set according to the Mel scale, denoted as F = [F0, F1, F2, F3, ..., F7, F8], where F0 corresponds to 4000Hz frequency, F8 corresponds to 8000Hz frequency, and the middle 7 filters are evenly distributed on the Mel frequency to better capture the key frequency characteristics in the ear canal echo. Mel frequency (f Mel ) and the Hertz frequency (f) is given by the following formula:

[0096]

[0097] Step 208: Set a 0-4kHz filter.

[0098] The Mel filter set in step 207 is mirror-flipped along the 4000 Hz frequency axis to set a Mel filter within a frequency range of 0 to 4000 Hz.

[0099] Specifically, in the frequency range of 0 Hz to 4000 Hz, according to the filter F=[F0, F1, F2, F3, ..., F7, F8] in step 7, the filter f=[f8, f7, f6, ..., f1, f0] is set, where the relationship between F and f is:

[0100] A m -4000=4000-B m

[0101] Among them, A is the frequency corresponding to filter F, B is the frequency corresponding to filter f, and m is the set filter index. These filters are densely distributed around the 4000Hz frequency. The bandwidth of each filter increases as the distance between the center point of the filter and the 4000Hz frequency increases. The distribution of Mel filters in the frequency range of 0-8000Hz can be referred to Figure 4 .

[0102] Step 209: Convolution calculates frequency band energy.

[0103] A filter whose center point (which can be understood as the center frequency) is located in a critical frequency range (which can be understood as a frequency interval excluding a noise frequency interval) among the Mel filters set in the steps 207 and 208 is selected as a filter for feature extraction, and the spectrum is decomposed into an energy distribution adapted to the Mel scale.

[0104] Specifically, a Mel filter with a center frequency in the frequency range of 2000 to 8000 Hz is selected for subsequent Mel cepstral coefficient extraction, and the filter F is finally obtained. total = [f5, f4, f3, ..., f1, F0, F1, F2, ..., F7], after applying the Mel filter to perform weighted processing on the spectrum, the output of each filter is obtained. The output of each filter represents the energy of the signal in the Mel frequency band.

[0105] By performing a convolution operation on the spectrum signal obtained by FFT and the Mel filter, the spectrum signal can be converted from a linear frequency scale to a Mel scale, and the energy value of each Mel frequency band (which can be understood as the aforementioned frequency band energy value) can be obtained.

[0106] Step 210: Logarithmic transformation.

[0107] Specifically, the energy value output by each Mel filter in step 210 is logarithmically transformed. Assuming e n is the energy output of the nth Mel filter, and the energy after logarithmic transformation is expressed as:

[0108] E n =log(e n +ε)

[0109] Among them, E is the energy value after logarithmic transformation, and ε is a very small floating point number (which can be understood as the aforementioned setting parameter) to avoid the logarithmic zero value problem. By performing logarithmic transformation on the energy value of each frequency band, a wide range of energy values ​​can be effectively compressed so that stronger frequency components do not overly dominate the feature representation.

[0110] Step 211: DCT transformation.

[0111] The logarithmic Mel frequency band energy obtained in step 210 is processed using discrete cosine transform (DCT), the data is converted from the frequency domain back to the cepstral domain, the first 13 Mel cepstral coefficients are extracted, and the 13-dimensional earprint MFCC features are obtained.

[0112] Specifically, DCT maps the logarithmic Mel band energy from the frequency domain back to the cepstrum domain, which is usually more suitable for speech processing and feature extraction. It can effectively compress information and highlight the main features of the signal. Given a logarithmic Mel band energy sequence of length 13, E = [E1, E2, E3, ..., E 13 ], the transformation formula of DCT is as follows:

[0113]

[0114] Among them, C k is the kth coefficient in the cepstrum domain, En is the nth logarithmic Mel band energy, N is the number of Mel bands, M is the MFCC dimension, where N=13, M=13, and the k value in the above formula can be understood as the parameter representation of the aforementioned feature dimension parameter in the current example. The cepstrum domain can compress the noise component in the spectral information and highlight the core pattern of the speech signal, making the feature more stable and discriminative in the subsequent recognition process. Through the above calculation, the 13-dimensional Mel cepstrum coefficient can be obtained, which is the extracted earprint MFCC feature, that is, the earprint feature referred to in step 212.

[0115] Based on the above method, the X-Vector model is selected as the identity recognition model, and the error rate and accuracy of the identity recognition model training phase are tested. The performance of the extracted MFCC features for identity recognition is shown in Table 1. Based on Table 1, it can be seen that compared with the MFCC feature extraction method that directly refers to the voiceprint, the present application can effectively improve the accuracy of identity recognition.

[0116] Table 1

[0117] Features used Accuracy (%) Equal error rate (%) Common ear print features 96.3866 1.8067 This application MFCC features 96.6114 1.6943

[0118] See also Figure 5 , Figure 5 is a feature acquisition device provided by an embodiment of the present disclosure, such as Figure 5 As shown, the feature acquisition device 500 includes:

[0119] A preprocessing module 501 is used to preprocess the target earprint data of the user to obtain preprocessed data, wherein the preprocessing includes at least one of noise reduction processing, pre-emphasis processing, frame processing and windowing processing;

[0120] An energy extraction module 502, used for converting the pre-processed data into spectrum data, and extracting the energy of the spectrum data based on N preset target filters, respectively, to obtain N frequency band energy values, wherein the center frequencies of the N target filters are symmetrically arranged with a set reference frequency as the center, the reference frequency is included in the frequency band corresponding to the frequency response peak of the earprint data of the target user group, the target user group includes the user, and the frequency band energy value is obtained based on a convolution operation of the corresponding target filter and the signal part of the spectrum data in the corresponding frequency band;

[0121] The feature conversion module 503 is used to perform feature conversion on the N frequency band energy values ​​respectively to obtain M earprint feature data. The user's identity includes the M earprint feature data, wherein N and M are integers greater than 1.

[0122] In one embodiment, the feature conversion module 503 includes:

[0123] a screening unit, configured to screen the N frequency band energy values ​​based on a noise frequency interval to obtain N' target frequency band energy values, wherein the N' target frequency band energy values ​​include: among the N frequency band energy values, all frequency band energy values ​​that are outside the noise frequency interval, within the noise frequency interval, a similarity between a first frequency response curve and a second frequency response curve that is less than or equal to a set difference threshold, the first frequency response curve being used to represent a frequency response trend of the earprint data of the left ear of the target user group as a function of frequency, and the second frequency response curve being used to represent a frequency response trend of the earprint data of the right ear of the target user group as a function of frequency;

[0124] The feature conversion unit is used to perform feature conversion on the N' target frequency band energy values ​​respectively to obtain the M earprint feature data, where N' is a positive integer less than or equal to N, wherein each earprint feature data is obtained based on the N' target frequency band energy values ​​and corresponding feature dimension parameters, and different earprint feature data have different corresponding feature dimension parameters.

[0125] In one embodiment, the feature conversion unit is specifically used to:

[0126] Performing logarithmic transformation on the N' target frequency band energy values ​​respectively to obtain N' logarithmic data, where the logarithmic data is the logarithmic value of the sum of the corresponding target frequency band energy value and the set parameter;

[0127] According to the set M feature dimension parameters, discrete cosine transform is performed on the N' logarithmic data respectively to obtain the M earprint feature data.

[0128] In one embodiment, among the N target filters, the bandwidth of each target filter is positively correlated with a distance parameter corresponding to the target filter, and the distance parameter is used to indicate the frequency difference between the center frequency of the corresponding target filter and the reference frequency.

[0129] and / or,

[0130] The target filter is a Mel filter, the N target filters include a first target filter and a second target filter, the center frequency of the first target filter is higher than the reference frequency, the center frequency of the second target filter is lower than the reference frequency, wherein the center Mel frequencies of the first target filters are distributed at equidistant intervals;

[0131] and / or,

[0132] The preprocessing module 501 is specifically used for:

[0133] According to a preset short-time energy threshold and a zero-crossing rate threshold, a noise reduction process is performed on a plurality of earprint audio signals included in the earprint data to obtain at least one target audio signal, wherein the short-time energy of the target audio signal is greater than or equal to the short-time energy threshold, and the zero-crossing rate of the target audio signal is less than or equal to the zero-crossing rate threshold;

[0134] Among the at least one target audio signal, performing pre-emphasis processing on each of the target audio signals to obtain at least one pre-emphasized audio signal, wherein the pre-emphasis processing includes first-order difference processing;

[0135] Performing frame processing on the at least one pre-emphasized audio signal to obtain frame signal data;

[0136] Performing windowing processing on the frame signal data based on a set window function to obtain the preprocessed data, wherein the window function includes a Hamming window function;

[0137] The frequency spectrum data is obtained by fast Fourier transform based on the preprocessed data.

[0138] The feature acquisition device 500 provided in the embodiment of the present disclosure can implement each process in the above-mentioned feature acquisition method embodiment, and will not be described again here to avoid repetition.

[0139] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device and a readable storage medium.

[0140] Figure 6 A schematic block diagram of an example electronic device 600 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.

[0141] like Figure 6As shown, the device 600 includes a computing unit 601, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 to a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the device 600 can also be stored. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0142] A number of components in the device 600 are connected to the I / O interface 605, including: an input unit 606, such as a keyboard, a mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a disk, an optical disk, etc.; and a communication unit 609, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 609 allows the device 600 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0143] The computing unit 601 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSP), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 601 performs the various methods and processes described above, such as feature acquisition methods. For example, in some embodiments, the feature acquisition method may be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as a storage unit 608. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the feature acquisition method described above may be performed. Alternatively, in other embodiments, the computing unit 601 may be configured to execute the feature acquisition method in any other appropriate manner (eg, by means of firmware).

[0144] Various embodiments of the systems and techniques described above herein may be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0145] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0146] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0147] As used herein, the term "machine-readable medium" refers to any computer program product, apparatus, and / or device (e.g., disk, optical disk, memory, programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.

[0148] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0149] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0150] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.

[0151] The present application also provides a computer program product, including computer instructions, which, when executed by a processor, implement the above Figure 1 The various processes of the method embodiment shown can achieve the same technical effect, and will not be described again here to avoid repetition.

[0152] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.

[0153] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. A feature acquisition method, characterized in that: The method comprises: Preprocessing the target earprint data of the user to obtain preprocessed data, wherein the preprocessing includes at least one of noise reduction processing, pre-emphasis processing, frame processing and windowing processing; The preprocessed data is converted into spectrum data, and the energy of the spectrum data is respectively extracted based on N preset target filters to obtain N frequency band energy values, wherein the center frequencies of the N target filters are symmetrically arranged with a set reference frequency as the center, the reference frequency is included in the frequency band corresponding to the frequency response peak of the earprint data of the target user group, the target user group includes the user, and the frequency band energy value is obtained based on the convolution operation of the corresponding target filter and the signal part of the spectrum data in the corresponding frequency band; Feature conversion is performed on the N frequency band energy values ​​respectively to obtain M earprint feature data, and the user's identity includes the M earprint feature data, wherein N and M are integers greater than 1.

2. The method according to claim 1, characterized in that The performing feature conversion on the N frequency band energy values ​​to obtain M earprint feature data respectively includes: The N frequency band energy values ​​are screened based on the noise frequency interval to obtain N' target frequency band energy values, wherein the N' target frequency band energy values ​​include: among the N frequency band energy values, all frequency band energy values ​​that are outside the noise frequency interval, within the noise frequency interval, the similarity between the first frequency response curve and the second frequency response curve is less than or equal to a set difference threshold, the first frequency response curve is used to represent the frequency response trend of the earprint data of the left ear of the target user group with frequency, and the second frequency response curve is used to represent the frequency response trend of the earprint data of the right ear of the target user group with frequency; The N' target frequency band energy values ​​are respectively subjected to feature conversion to obtain the M earprint feature data, where N' is a positive integer less than or equal to N, wherein each earprint feature data is obtained based on the N' target frequency band energy values ​​and corresponding feature dimension parameters, and different earprint feature data have different corresponding feature dimension parameters.

3. The method according to claim 2, characterized in that The performing feature conversion on the N' target frequency band energy values ​​respectively to obtain the M earprint feature data includes: Performing logarithmic transformation on the N' target frequency band energy values ​​respectively to obtain N' logarithmic data, where the logarithmic data is the logarithmic value of the sum of the corresponding target frequency band energy value and the set parameter; According to the set M feature dimension parameters, discrete cosine transform is performed on the N' logarithmic data respectively to obtain the M earprint feature data.

4. The method according to claim 1, characterized in that Among the N target filters, the bandwidth of each target filter is positively correlated with a distance parameter corresponding to the target filter, and the distance parameter is used to indicate a frequency difference between a center frequency of the corresponding target filter and the reference frequency.

5. The method according to claim 1, characterized in that The preprocessing of the earprint data of the user to obtain preprocessed data includes: According to a preset short-time energy threshold and a zero-crossing rate threshold, a noise reduction process is performed on a plurality of earprint audio signals included in the earprint data to obtain at least one target audio signal, wherein the short-time energy of the target audio signal is greater than or equal to the short-time energy threshold, and the zero-crossing rate of the target audio signal is less than or equal to the zero-crossing rate threshold; Among the at least one target audio signal, performing pre-emphasis processing on each of the target audio signals to obtain at least one pre-emphasized audio signal, wherein the pre-emphasis processing includes first-order difference processing; Performing frame processing on the at least one pre-emphasized audio signal to obtain frame signal data; Performing windowing processing on the frame signal data based on a set window function to obtain the preprocessed data, wherein the window function includes a Hamming window function; The frequency spectrum data is obtained by fast Fourier transform based on the preprocessed data.

6. The method according to claim 1, characterized in that The target filter is a Mel filter, the N target filters include a first target filter and a second target filter, the center frequency of the first target filter is higher than the reference frequency, and the center frequency of the second target filter is lower than the reference frequency; Among them, the central Mel frequencies of multiple first target filters are distributed at equidistant intervals.

7. A feature acquisition device, characterized in that: The device comprises: A preprocessing module, used to preprocess the target earprint data of the user to obtain preprocessed data, wherein the preprocessing includes at least one of noise reduction processing, pre-emphasis processing, frame processing and windowing processing; an energy extraction module, for converting the preprocessed data into spectrum data, and extracting the energy of the spectrum data based on N preset target filters, respectively, to obtain N frequency band energy values, wherein the center frequencies of the N target filters are symmetrically arranged with a set reference frequency as the center, the reference frequency is included in a frequency band corresponding to a frequency response peak of the earprint data of a target user group, the target user group includes the user, and the frequency band energy value is obtained based on a convolution operation of a corresponding target filter and a signal portion of the spectrum data in a corresponding frequency band; The feature conversion module is used to perform feature conversion on the N frequency band energy values ​​respectively to obtain M earprint feature data, and the user's identity includes the M earprint feature data, wherein N and M are integers greater than 1.

8. The device according to claim 7, characterized in that The feature conversion module comprises: a screening unit, configured to screen the N frequency band energy values ​​based on a noise frequency interval to obtain N' target frequency band energy values, wherein the N' target frequency band energy values ​​include: among the N frequency band energy values, all frequency band energy values ​​that are outside the noise frequency interval, within the noise frequency interval, a similarity between a first frequency response curve and a second frequency response curve that is less than or equal to a set difference threshold, the first frequency response curve being used to represent a frequency response trend of the earprint data of the left ear of the target user group as a function of frequency, and the second frequency response curve being used to represent a frequency response trend of the earprint data of the right ear of the target user group as a function of frequency; The feature conversion unit is used to perform feature conversion on the N' target frequency band energy values ​​respectively to obtain the M earprint feature data, where N' is a positive integer less than or equal to N, wherein each earprint feature data is obtained based on the N' target frequency band energy values ​​and corresponding feature dimension parameters, and different earprint feature data have different corresponding feature dimension parameters.

9. The device according to claim 8, characterized in that The feature conversion unit is specifically used for: Performing logarithmic transformation on the N' target frequency band energy values ​​respectively to obtain N' logarithmic data, where the logarithmic data is the logarithmic value of the sum of the corresponding target frequency band energy value and the set parameter; According to the set M feature dimension parameters, discrete cosine transform is performed on the N' logarithmic data respectively to obtain the M earprint feature data.

10. The device according to claim 7, characterized in that Among the N target filters, the bandwidth of each target filter is positively correlated with the distance parameter corresponding to the target filter, and the distance parameter is used to indicate the frequency difference between the center frequency of the corresponding target filter and the reference frequency. and / or, The target filter is a Mel filter, the N target filters include a first target filter and a second target filter, the center frequency of the first target filter is higher than the reference frequency, the center frequency of the second target filter is lower than the reference frequency, wherein the center Mel frequencies of the first target filters are distributed at equidistant intervals; and / or, The preprocessing module is specifically used for: According to a preset short-time energy threshold and a zero-crossing rate threshold, a noise reduction process is performed on a plurality of earprint audio signals included in the earprint data to obtain at least one target audio signal, wherein the short-time energy of the target audio signal is greater than or equal to the short-time energy threshold, and the zero-crossing rate of the target audio signal is less than or equal to the zero-crossing rate threshold; Among the at least one target audio signal, performing pre-emphasis processing on each of the target audio signals to obtain at least one pre-emphasized audio signal, wherein the pre-emphasis processing includes first-order difference processing; Performing frame processing on the at least one pre-emphasized audio signal to obtain frame signal data; Performing windowing processing on the frame signal data based on a set window function to obtain the preprocessed data, wherein the window function includes a Hamming window function; The frequency spectrum data is obtained by fast Fourier transform based on the preprocessed data.

Citation Information

Patent Citations

  • Transformer voiceprint feature controllable precision extraction and recognition method and system based on MFCC

    CN115223576A

  • Open wearable acoustic device and active noise reduction method

    US20250037695A1