A feature acquisition method and apparatus
Patent Information
- Application Number
- CN202510247201.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2026-09-01
- Estimated Expiration
- 2045-03-04
AI Technical Summary
[0004]本申请的目的在于提供一种特征获取方法及装置,用于解决相关技术提取的耳纹特征的可靠性低的技术问题
[0015]在本申请中,在目标用户群的耳纹数据的频率响应峰值所对应频率段中确定基准频率,并以基准频率为中心将N个目标滤波器的中心频率对称设置,从通过N个目标滤波器从用户的目标耳纹数据中提取更加可靠的耳纹特征,并通过体现每个用户的个性化耳纹特点的耳纹特征来进行用户身份的区分,进而形成用于唯一指示每个用户的身份标识。
Smart Images

Figure CN119993194B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of speech signal processing, specifically to a feature acquisition method and apparatus. Background Technology
[0002] Among existing biometric technologies, voiceprint recognition has been widely used in scenarios such as identity authentication and voice control. However, voiceprint recognition is greatly affected by environmental noise and the speaker's state (such as emotion and health), making it difficult to fully meet application requirements in scenarios with high security needs. In contrast, earprint recognition, by analyzing the characteristics of the ear canal echo signal, has higher stability and individual specificity, and is therefore gradually becoming a potential supplementary or alternative technology.
[0003] However, at present, the reliability of ear print features extracted based on related technologies is lower than expected. Summary of the Invention
[0004] The purpose of this application is to provide a feature acquisition method and apparatus to solve the technical problem of low reliability of ear print features extracted by related technologies.
[0005] In a first aspect, embodiments of this application provide a feature acquisition method, the method comprising:
[0006] The user's target earprint data is preprocessed to obtain preprocessed data, wherein the preprocessing includes at least one of noise reduction processing, pre-emphasis processing, frame segmentation processing, and windowing processing;
[0007] The preprocessed data is converted into spectrum data, and the energy of the spectrum data is extracted based on N pre-set target filters to obtain N frequency band energy values. The center frequencies of the N target filters are symmetrically set with a set reference frequency as the center. The reference frequency is contained in the frequency band corresponding to the frequency response peak of the earprint data of the target user group. The target user group includes the users. The frequency band energy values are obtained by convolving the corresponding target filter and the spectrum data in the signal part of the corresponding frequency band.
[0008] The N frequency band energy values are respectively subjected to feature transformation to obtain M earprint feature data. The user's identity identifier includes the M earprint feature data, where N and M are integers greater than 1.
[0009] Secondly, embodiments of this application also provide a feature acquisition device, the device comprising:
[0010] The preprocessing module is used to preprocess the user's target earprint data to obtain preprocessed data, wherein the preprocessing includes at least one of noise reduction processing, pre-emphasis processing, frame segmentation processing, and windowing processing;
[0011] An energy extraction module is used to convert the preprocessed data into spectrum data, and extract the energy of the spectrum data based on N pre-set target filters to obtain N frequency band energy values. The center frequencies of the N target filters are symmetrically set with a set reference frequency as the center. The reference frequency is contained in the frequency band corresponding to the frequency response peak of the earprint data of the target user group. The target user group includes the users. The frequency band energy values are obtained by convolving the corresponding target filter and the spectrum data in the signal part of the corresponding frequency band.
[0012] The feature conversion module is used to perform feature conversion on the N frequency band energy values respectively to obtain M earprint feature data. The user's identity identifier includes the M earprint feature data, where N and M are integers greater than 1.
[0013] Thirdly, this disclosure provides an electronic device including a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, performs the steps of the method described in the first aspect.
[0014] Fourthly, this disclosure provides a computer program product including computer instructions that, when executed by a processor, implement the steps of the method described in the first aspect.
[0015] In this application, a reference frequency is determined in the frequency range corresponding to the frequency response peak of the earprint data of the target user group, and the center frequencies of N target filters are symmetrically set with the reference frequency as the center. More reliable earprint features are extracted from the target earprint data of the user through the N target filters, and the user identity is distinguished by earprint features that reflect the personalized earprint characteristics of each user, thereby forming an identity identifier that uniquely indicates each user. Attached Figure Description
[0016] Figure 1 This is a schematic flowchart of a feature acquisition method provided in an embodiment of this disclosure;
[0017] Figure 2 This is a schematic flowchart of an MFCC feature extraction method based on ear print characteristics provided in an embodiment of this disclosure;
[0018] Figure 3 This is a schematic diagram illustrating the variation of frequency response amplitude in a user's left and right ears, provided in an embodiment of this disclosure.
[0019] Figure 4 This is a schematic diagram of the distribution of a Mel filter bank provided in an embodiment of this disclosure;
[0020] Figure 5This is a schematic diagram of the structure of a feature acquisition device provided in an embodiment of this disclosure;
[0021] Figure 6 This is a schematic diagram of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0022] The technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this disclosure. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.
[0023] This application provides a feature acquisition method, such as... Figure 1 As shown, the method includes:
[0024] Step 101: Preprocess the user's target earprint data to obtain preprocessed data.
[0025] The preprocessing includes at least one of noise reduction processing, pre-emphasis processing, frame splitting processing, and windowing processing.
[0026] In this application, the user's target earprint data is specifically the echo signal of the user's ear canal. The echo signal of the user's ear canal can be collected by any target device (such as headphones) with audio playback and recording functions.
[0027] For example, when the target device is set as headphones, the process of collecting the echo signal in the user's ear canal is as follows: the user wears headphones, and the headphones play a 2-second Gaussian white noise. During the playback, the headphones simultaneously collect the echo signal after the Gaussian white noise is reflected through the user's left ear canal and / or right ear canal.
[0028] Noise reduction processing is used to remove noise from echo signals in order to improve the clarity and understandability of the signal portion of the echo signal associated with the user's personalized characteristics. In this application, noise is used not only to indicate signal interference caused by the external environment, but also to indicate the signal portion of the echo signal that cannot be used to represent the user's personalized characteristics.
[0029] Pre-emphasis processing is used to enhance the energy of the high-frequency components in the echo signal to compensate for the loss of the high-frequency components in the echo signal during subsequent processing. This is to adapt to the characteristic that the differences in earprints of different people are concentrated in the high-frequency components of the echo signal, and to improve the clarity and intelligibility of the signal portion of the echo signal that is associated with the user's personalized characteristics (indicating the differences in earprints of different users).
[0030] Framing is used to divide continuous echo signals into short time frames so that the echo signals remain stable within each short time frame, facilitating subsequent operations to extract earprint feature data from the echo signals.
[0031] Windowing is used to smooth the edges of each short time frame after framing to reduce the discontinuity of frame edges, suppress the risk of spectral leakage, and make the subsequent frequency domain conversion and analysis of the echo signal more accurate.
[0032] Step 102: Convert the preprocessed data into spectrum data, and extract the energy of the spectrum data based on N pre-set target filters to obtain N frequency band energy values.
[0033] The center frequencies of the N target filters are symmetrically set around a set reference frequency. The reference frequency is contained within the frequency band corresponding to the frequency response peak of the earprint data of the target user group. The target user group includes the users. The frequency band energy value is obtained by convolving the corresponding target filter and the spectrum data in the signal part of the corresponding frequency band.
[0034] In this application, the target user group can be a set of users that meet the set user selection rules. The set user selection rules can be adaptively adjusted according to actual needs. For example, users engaged in a specific profession can be identified as users in the target user group, or users located in a specific space can be identified as users in the target user group.
[0035] The earprint data of the target user group includes the echo signal of the left ear canal of each user in the target user group and / or the echo signal of the right ear canal of each user in the target user group. The echo signal referred to here can also be acquired by the aforementioned target device.
[0036] It should be understood that the peak frequency response of the echo signal from the left ear canal of each user in the target user group and / or the peak frequency response of the echo signal from the right ear canal of each user in the target user group constitute a peak data set; in the peak data set, the proportion of data located within the frequency band corresponding to the peak frequency response of the earprint data of the target user group is greater than a set proportion threshold (such as 0.9, 0.95, 0.99, etc.).
[0037] In one example, the frequency response peak of the echo signal from the left ear canal of each user in the target user group can be used as the center, and the set frequency band width can be used as the distance from the center to the edge to obtain the frequency band corresponding to the frequency response peak of the echo signal from the left ear canal of each user in the target user group. A similar method can be used to obtain the frequency band corresponding to the frequency response peak of the echo signal from the right ear canal of each user in the target user group. The frequency bands corresponding to the frequency response peak of the echo signal from the left ear canal of each user in the target user group and the frequency bands corresponding to the frequency response peak of the echo signal from the right ear canal of each user in the target user group are then intersected or combined to obtain the frequency band corresponding to the frequency response peak of the earprint data of the target user group.
[0038] In another example, the average of the frequency response peak values of the echo signals from the left ear canal and / or the right ear canal of each user in the target user group can be used as the target center, and the set frequency band width can be used as the distance from the target center to the edge, so as to obtain the frequency band corresponding to the frequency response peak values of the earprint data of the target user group.
[0039] For example, the reference frequency can be randomly selected within the frequency range corresponding to the peak frequency response of the earprint data of the target user group, or it can be the median frequency within the frequency range corresponding to the peak frequency response of the earprint data of the target user group, or it can be the frequency of the corresponding peak frequency response point within the frequency range corresponding to the peak frequency response of the earprint data of the target user group.
[0040] Step 103: Perform feature transformation on the N frequency band energy values respectively to obtain M earprint feature data.
[0041] The user's identity identifier includes the M earprint feature data, where N and M are integers greater than 1.
[0042] In this application, a reference frequency is determined in the frequency range corresponding to the frequency response peak of the earprint data of the target user group, and the center frequencies of N target filters are symmetrically set with the reference frequency as the center. More reliable earprint features are extracted from the target earprint data of the user through the N target filters, and the user identity is distinguished by earprint features that reflect the personalized earprint characteristics of each user, thereby forming an identity identifier that uniquely indicates each user.
[0043] M earprint feature data can directly constitute the user's identity identifier, or they can be combined with other user identity features (such as voiceprint features, facial features, etc.) to form the user's identity identifier.
[0044] For example, the user's identity can be registered in the identity information database and the user's identity can be verified based on the user's identity identifier. This application does not limit the specific application of the user's identity identifier.
[0045] In one embodiment, the step of performing feature transformation on the N frequency band energy values to obtain M earprint feature data includes:
[0046] Based on the noise frequency range, the N frequency band energy values are filtered to obtain N' target frequency band energy values. The N' target frequency band energy values include all frequency band energy values located outside the noise frequency range. Within the noise frequency range, the similarity between the first frequency response curve and the second frequency response curve is less than or equal to a set difference threshold. The first frequency response curve is used to represent the frequency response trend of the earprint data of the left ear of the target user group with frequency. The second frequency response curve is used to represent the frequency response trend of the earprint data of the right ear of the target user group with frequency.
[0047] The N' target frequency band energy values are respectively subjected to feature transformation to obtain the M earprint feature data, where N' is a positive integer less than or equal to N. Each earprint feature data is obtained based on the N' target frequency band energy values and the corresponding feature dimension parameters. Different earprint feature data correspond to different feature dimension parameters.
[0048] In this embodiment, the noise frequency range is the same as the frequency range where the aforementioned noise is located. Based on the above settings, the interference caused by noise is suppressed, and the reliability of the final M earprint feature data is further improved.
[0049] It should be noted that determining the noise frequency range based on the similarity comparison between curves can fully and effectively utilize the frequency response characteristics reflected by the earprint data of each user in the target user group, thereby avoiding interference from extreme point values and ensuring the accuracy of the determined noise frequency range.
[0050] For example, the frequency response of the left ear of each user in the target user group at different frequencies and the frequency response of the right ear of each user in the target user group at different frequencies can be processed by curve fitting to obtain the first frequency response curve and the second frequency response curve, respectively.
[0051] In one embodiment, the step of performing feature transformation on the N' target frequency band energy values to obtain the M earprint feature data includes:
[0052] Logarithmic transformation is performed on the N' target frequency band energy values to obtain N' logarithmic data, where the logarithmic data is the logarithmic value of the sum of the corresponding target frequency band energy value and the set parameter;
[0053] Based on the set M feature dimension parameters, the N' logarithmic data are subjected to Discrete Cosine Transform (DCT) respectively to obtain the M earprint feature data.
[0054] In this embodiment, the processing based on logarithmic transformation is used to compress a wide range of energy values and suppress the excessive dominance of relatively stronger frequency components in the target frequency band energy values on feature representation; while the processing based on discrete cosine transform can compress the overall information of logarithmic data while highlighting the earprint features contained in the logarithmic data, thereby increasing the accuracy and reliability of the obtained M earprint feature data.
[0055] The setting parameter is used to avoid the problem of the logarithm being zero. In the application, a very small floating-point number can be used as the setting parameter to reduce the impact of adding the setting parameter on the final calculated logarithm value while avoiding the problem of the logarithm being zero.
[0056] Setting the feature dimension parameter allows you to extract specific dimension information from the N' target frequency band energy values, so as to make full use of the N' target frequency band energy values and form comprehensive and accurate earprint feature data. It should be noted that as the value of M increases, the M earprint feature data will more accurately represent the features of the user's earprint, but the computational complexity will also increase accordingly. In the application, you can make an adaptive choice on the specific value of M according to actual needs, such as 12 or 13.
[0057] In one embodiment, among the N target filters, the bandwidth of each target filter is positively correlated with the distance parameter corresponding to the target filter, the distance parameter being used to indicate the frequency difference between the center frequency of the corresponding target filter and the reference frequency.
[0058] In this embodiment, based on the above settings, the N target filters are used to simulate as much as possible the changes in the human ear's audio perception ability (referring to the situation where the human ear's audio perception ability is strongest near the reference frequency, and gradually decreases as the distance from the reference frequency increases), so that the N frequency band energy values are more accurate and effective.
[0059] In one embodiment, the target filter is a Mel filter, and the N target filters include a first target filter and a second target filter, wherein the center frequency of the first target filter is higher than the reference frequency, and the center frequency of the second target filter is lower than the reference frequency;
[0060] Among them, the center Mel frequencies of multiple first target filters are distributed at equal intervals.
[0061] In this embodiment, multiple first target filters are uniformly set using the Mel frequency as a reference to avoid the problem of uneven feature distribution in the echo signal at conventional frequencies (specifically, the earprint features carried by the echo signal are densely distributed in the low-frequency band and sparsely distributed in the high-frequency band at conventional frequencies), thus ensuring the accuracy of the frequency band energy values extracted by the multiple first target filters. Furthermore, by mirroring the multiple second target filters with the reference frequency as the mirror center, the accuracy of the frequency band energy values extracted by the multiple second target filters can be ensured, while making the setting of the multiple second target filters more convenient.
[0062] For example, if the reference frequency is set to 4000, then the first target filter and the second target filter satisfy the following relationship:
[0063] a n -4000 = 4000 - b n
[0064] Where a is the center frequency of the first target filter, b is the center frequency of the second target filter, and n is the index of the corresponding filter.
[0065] In one embodiment, the preprocessing of the user's earprint data to obtain preprocessed data includes:
[0066] Based on a preset short-time energy threshold and a zero-crossing rate threshold, the earprint data includes multiple earprint audio signals that are denoised to obtain at least one target audio signal. The short-time energy of the target audio signal is greater than or equal to the short-time energy threshold, and the zero-crossing rate of the target audio signal is less than or equal to the zero-crossing rate threshold.
[0067] In the at least one target audio signal, pre-emphasis processing is performed on each of the target audio signals to obtain at least one pre-emphasis audio signal, wherein the pre-emphasis processing includes first-order differential processing;
[0068] The at least one pre-emphasized audio signal is subjected to frame segmentation processing to obtain frame signal data;
[0069] The frame signal data is windowed based on a set window function to obtain the preprocessed data. The window function includes a Hamming window function.
[0070] The spectral data is obtained by fast Fourier transform based on the preprocessed data.
[0071] In this embodiment, based on the setting of short-time energy threshold and zero-crossing rate threshold, the unvoiced segment in earprint data is identified and filtered, while the voiced segment, which can more accurately reflect the harmonic characteristics of the ear canal and is convenient for feature extraction, is retained.
[0072] For ease of understanding, the following example is provided:
[0073] like Figure 2 As shown, this application also provides a method for extracting Mel-frequency cepstral coefficients (MFCC) features based on earprint characteristics, which includes the following steps:
[0074] Step 201: Collect user earprint data.
[0075] Specifically, the method involves collecting the user's ear canal echo signal as earprint data, and the aforementioned target device can be used to perform the collection work.
[0076] Step 202: VAD extracts voiced signals.
[0077] Specifically, the ear canal echo signal acquired in step 201 is processed by framing, and the short-time energy and zero-crossing rate of each frame signal are calculated. By setting thresholds for short-time energy and zero-crossing rate (which can be understood as short-time energy threshold and zero-crossing rate threshold), unvoiced signal segments are identified and filtered, and only voiced signal segments with high short-time energy and low zero-crossing rate are retained. Since the retained voiced part can more accurately reflect the harmonic characteristics of the ear canal, it is convenient for feature extraction and subsequent use of the extracted earprint features for identity recognition.
[0078] Step 203, Pre-weighting.
[0079] Specifically, applying a pre-emphasis filter to the voiced segment signal retained in step 202 allows for a first-order difference operation on the voiced segment signal, the formula of which is:
[0080] y[n] = x[n] - α·x[n-1]
[0081] Where x[n] represents the current sampling point of the input signal, x[n-1] represents the previous sampling point, and α is the pre-emphasis coefficient with a value of 0.97, used to control the enhancement level of the high-frequency part. Key signals in earprint signals typically have a higher energy distribution in the high-frequency part, and the differences in earprints among different people are also mainly concentrated in the high-frequency part. Therefore, pre-emphasis processing can make the high-frequency part of the voiced signal more prominent.
[0082] Step 204: Frame division.
[0083] Specifically, the non-stationary signal processed in step 203 is divided into short frames of 16ms. At this time, each frame signal after framing is considered stationary. The frame shift can be set to 10ms. Each frame signal has 37.5% overlap to ensure the continuity between frames and reduce the loss of information between frames.
[0084] Step 205: Add a window.
[0085] Specifically, a window function is applied to each frame of signal samples after framing to reduce signal distortion at frame boundaries. A Hamming window can be used here, and its formula is defined as follows:
[0086]
[0087] Where ω[n] is the weight of the window function, N is the number of sample points, and n is the sample index within the frame. By multiplying the window function point by point with the signal of each frame, the boundaries of the signal can be smoothed and the spectral leakage effect can be reduced.
[0088] Step 206: FFT transformation.
[0089] Specifically, a Fast Fourier Transform (FFT) is performed on the windowed signal of each frame to convert the time-domain signal into a frequency-domain signal.
[0090] The formula is as follows:
[0091]
[0092] Where X[k] is the complex value of the k-th frequency component of the frequency domain signal, x[n] is the n-th sampling point of the time domain signal, and N is the number of sampling points per frame of the signal, where N = 512. The spectral amplitude and phase information of each frame of the signal can be obtained through FFT calculation.
[0093] Step 207: Set the 4-8kHz filter.
[0094] Specifically, the earprint spectrum obtained in step 206 is set with a Mel filter, starting at 4000Hz and ending at 8000Hz.
[0095] User earprint data statistics can be referenced. Figure 3 , Figure 3Within the 2000Hz–8000Hz range, the frequency response amplitude of earprint data is relatively large. Therefore, within the 4000Hz–8000Hz frequency range, seven Mel filters (which can be understood as multiple first-target filters) are set according to the Mel scale, denoted as F = [F0, F1, F2, F3, ..., F7, F8], where F0 corresponds to the 4000Hz frequency, F8 corresponds to the 8000Hz frequency, and the seven filters in the middle are evenly distributed at the Mel frequencies to better capture key frequency features in the ear canal echo. Mel frequency (f Mel The relationship between (f) and the Hertzian frequency (f) is given by the following formula:
[0096]
[0097] Step 208: Set the 0-4kHz filter.
[0098] The Mel filter set in step 207 is mirrored along the 4000Hz frequency axis, and a Mel filter in the frequency range of 0 to 4000Hz is set.
[0099] Specifically, within the frequency range of 0Hz to 4000Hz, based on the filter F = [F0, F1, F2, F3, ..., F7, F8] from step 7, the filter f = [f8, f7, f6, ..., f1, f0] is set, where the relationship between F and f is as follows:
[0100] A m -4000=4000-B m
[0101] Where A is the frequency corresponding to filter F, B is the frequency corresponding to filter f, and m is the set filter index. These filters are densely distributed around the 4000Hz frequency. The bandwidth of each filter increases with the distance between the filter center point and the 4000Hz frequency. The distribution of Mel filters in the 0-8000Hz frequency range can be referenced. Figure 4 .
[0102] Step 209: Convolution to calculate frequency band energy.
[0103] The filter whose center point (which can be understood as the center frequency) of the Mel filter set in steps 207 and 208 is located in the key frequency range (which can be understood as the frequency range other than the noise frequency range) is selected as the filter for feature extraction, and the spectrum is decomposed into an energy distribution adapted to the Mel scale.
[0104] Specifically, a Mel filter with a center frequency in the range of 2000–8000 Hz is selected for subsequent Mel cepstral coefficient extraction, ultimately yielding filter F. total =[f5,f4,f3,...,f1,F0,F1,F2,...,F7], and the output of each filter is obtained by weighting the spectrum using Mel filters. The output of each filter represents the energy of the signal in that Mel frequency band.
[0105] By convolving the spectrum signal obtained by FFT with the Mel filter, the spectrum signal can be transformed from the linear frequency scale to the Mel scale, and the energy value of each Mel frequency band can be obtained (which can be understood as the aforementioned frequency band energy value).
[0106] Step 210: Logarithmic transformation.
[0107] Specifically, a logarithmic transformation is performed on the energy value output by each Mel filter in step 210, assuming e n Let the energy output of the nth Mel filter be the energy after logarithmic transformation, then the energy is expressed as:
[0108] E n =log(e n +ε)
[0109] Here, E is the energy value after logarithmic transformation, and ε is a very small floating-point number (which can be understood as the aforementioned setting parameter) used to avoid the problem of zero logarithmic values. By performing a logarithmic transformation on the energy value of each frequency band, a large range of energy values can be effectively compressed, so that stronger frequency components do not excessively dominate the feature representation.
[0110] Step 211, DCT transformation.
[0111] The logarithmic Mel frequency band energy obtained in step 210 is processed using discrete cosine transform (DCT) to convert the data from the frequency domain back to the cepstral domain, and the first 13 Mel cepstral coefficients are extracted to obtain 13-dimensional earprint MFCC features.
[0112] Specifically, DCT maps the logarithmic Mel-band energy from the frequency domain back to the cepstral domain. The cepstral domain is generally more suitable for speech processing and feature extraction, effectively compressing information and highlighting the main features of the signal. Given a logarithmic Mel-band energy sequence of length 13, E = [E1, E2, E3, ..., E...], ... 13 The transformation formula for DCT is as follows:
[0113]
[0114] Among them, C k It is the k-th coefficient in the cepstral domain, En This represents the energy of the nth logarithmic Mel frequency band, where N is the number of Mel frequency bands and M is the MFCC dimension. Here, N = 13 and M = 13. The value of k in the above formula can be understood as the parameter representation of the aforementioned feature dimension parameters in the current example. The cepstral domain can compress the noise components in the spectral information, highlighting the core pattern of the speech signal, making the features more stable and discriminative in the subsequent recognition process. Through the above calculation, the 13-dimensional Mel cepstral coefficients can be obtained, which are the extracted earprint MFCC features, i.e., the earprint features referred to in step 212.
[0115] Based on the above method, the X-Vector model was selected as the identity recognition model. The error rate and accuracy of the identity recognition model during the training phase were tested. The performance of the extracted MFCC features for identity recognition is shown in Table 1. As can be seen from Table 1, compared with the MFCC feature extraction method that directly refers to voiceprints, this application can effectively improve the accuracy of identity recognition.
[0116] Table 1
[0117] Common ear print features 96.3866 1.8067 This application features MFCC characteristics. 96.6114 1.6943
[0118] See Figure 5 , Figure 5 This is a feature acquisition device provided in an embodiment of the present disclosure, such as... Figure 5 As shown, the feature acquisition device 500 includes:
[0119] The preprocessing module 501 is used to preprocess the user's target earprint data to obtain preprocessed data, wherein the preprocessing includes at least one of noise reduction processing, pre-emphasis processing, frame segmentation processing and windowing processing;
[0120] The energy extraction module 502 is used to convert the preprocessed data into spectrum data, and extract the energy of the spectrum data based on N pre-set target filters to obtain N frequency band energy values. The center frequencies of the N target filters are symmetrically set with a set reference frequency as the center. The reference frequency is contained in the frequency band corresponding to the frequency response peak of the earprint data of the target user group. The target user group includes the user. The frequency band energy values are obtained by convolving the corresponding target filter and the spectrum data in the signal part of the corresponding frequency band.
[0121] The feature conversion module 503 is used to perform feature conversion on the N frequency band energy values respectively to obtain M earprint feature data. The user's identity identifier includes the M earprint feature data, where N and M are integers greater than 1.
[0122] In one embodiment, the feature conversion module 503 includes:
[0123] A filtering unit is used to filter the N frequency band energy values based on the noise frequency range to obtain N' target frequency band energy values, wherein the N' target frequency band energy values include: all frequency band energy values located outside the noise frequency range among the N frequency band energy values, and within the noise frequency range, the similarity between the first frequency response curve and the second frequency response curve is less than or equal to a set difference threshold, wherein the first frequency response curve is used to represent the frequency response trend of the earprint data of the left ear of the target user group with frequency, and the second frequency response curve is used to represent the frequency response trend of the earprint data of the right ear of the target user group with frequency;
[0124] The feature conversion unit is used to perform feature conversion on the N' target frequency band energy values respectively to obtain the M earprint feature data, where N' is a positive integer less than or equal to N. Each earprint feature data is obtained based on the N' target frequency band energy values and the corresponding feature dimension parameters. Different earprint feature data have different feature dimension parameters.
[0125] In one embodiment, the feature conversion unit is specifically used for:
[0126] Logarithmic transformation is performed on the N' target frequency band energy values to obtain N' logarithmic data, where the logarithmic data is the logarithmic value of the sum of the corresponding target frequency band energy value and the set parameter;
[0127] Based on the set M feature dimension parameters, the N' logarithmic data are subjected to discrete cosine transform to obtain the M earprint feature data.
[0128] In one embodiment, among the N target filters, the bandwidth of each target filter is positively correlated with the distance parameter corresponding to the target filter, the distance parameter indicating the frequency difference between the center frequency of the corresponding target filter and the reference frequency.
[0129] And / or,
[0130] The target filter is a Mel filter, and the N target filters include a first target filter and a second target filter. The center frequency of the first target filter is higher than the reference frequency, and the center frequency of the second target filter is lower than the reference frequency. The center Mel frequencies of the multiple first target filters are distributed at equal intervals.
[0131] And / or,
[0132] The preprocessing module 501 is specifically used for:
[0133] Based on a preset short-time energy threshold and a zero-crossing rate threshold, the earprint data includes multiple earprint audio signals that are denoised to obtain at least one target audio signal. The short-time energy of the target audio signal is greater than or equal to the short-time energy threshold, and the zero-crossing rate of the target audio signal is less than or equal to the zero-crossing rate threshold.
[0134] In the at least one target audio signal, pre-emphasis processing is performed on each of the target audio signals to obtain at least one pre-emphasis audio signal, wherein the pre-emphasis processing includes first-order differential processing;
[0135] The at least one pre-emphasized audio signal is subjected to frame segmentation processing to obtain frame signal data;
[0136] The frame signal data is windowed based on a set window function to obtain the preprocessed data. The window function includes a Hamming window function.
[0137] The spectral data is obtained by fast Fourier transform based on the preprocessed data.
[0138] The feature acquisition device 500 provided in this embodiment can implement the various processes in the above feature acquisition method embodiments, and will not be described again here to avoid repetition.
[0139] According to embodiments of this disclosure, this disclosure also provides an electronic device and a readable storage medium.
[0140] Figure 6 A schematic block diagram of an example electronic device 600 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0141] like Figure 6As shown, device 600 includes a computing unit 601, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 602 or a computer program loaded from storage unit 608 into random access memory (RAM) 603. The RAM 603 may also store various programs and data required for the operation of device 600. The computing unit 601, ROM 602, and RAM 603 are interconnected via bus 604. An input / output (I / O) interface 605 is also connected to bus 604.
[0142] Multiple components in device 600 are connected to I / O interface 605, including: input unit 606, such as keyboard, mouse, etc.; output unit 607, such as various types of monitors, speakers, etc.; storage unit 608, such as disk, optical disk, etc.; and communication unit 609, such as network card, modem, wireless transceiver, etc. Communication unit 609 allows device 600 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0143] The computing unit 601 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. The computing unit 601 performs the various methods and processes described above, such as feature acquisition methods. For example, in some embodiments, the feature acquisition method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 608. In some embodiments, part or all of the computer program may be loaded and / or installed on device 600 via ROM 602 and / or communication unit 609. When the computer program is loaded into RAM 603 and executed by the computing unit 601, one or more steps of the feature acquisition method described above may be performed. Alternatively, in other embodiments, the computing unit 601 may be configured to perform a feature acquisition method by any other suitable means (e.g., by means of firmware).
[0144] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0145] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0146] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0147] As used herein, the term "machine-readable medium" refers to any computer program product, device, and / or apparatus (e.g., disk, optical disk, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including machine-readable media that receive machine instructions as machine-readable signals. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0148] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0149] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0150] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0151] This application also provides a computer program product, including computer instructions, which, when executed by a processor, implement the above-described... Figure 1 The various processes of the method embodiments shown can achieve the same technical effect, and will not be described again here to avoid repetition.
[0152] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0153] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A feature acquisition method, characterized in that, The method includes: The user's target earprint data is preprocessed to obtain preprocessed data, wherein the preprocessing includes at least one of noise reduction processing, pre-emphasis processing, frame segmentation processing, and windowing processing; The preprocessed data is converted into spectral data, and the energy of the spectral data is extracted based on N pre-set target filters to obtain N frequency band energy values. Among the N target filters, the bandwidth of each target filter is positively correlated with the distance parameter corresponding to the target filter. The distance parameter is used to indicate the frequency difference between the center frequency and the reference frequency of the corresponding target filter. The center frequencies of the N target filters are symmetrically set around the set reference frequency, which is contained in the frequency band corresponding to the frequency response peak of the earprint data of the target user group. The target user group includes the users. The frequency band energy values are obtained by convolving the corresponding target filter and the spectral data in the signal portion of the corresponding frequency band. The N frequency band energy values are respectively subjected to feature transformation to obtain M earprint feature data. The user's identity identifier includes the M earprint feature data, where N and M are integers greater than 1. The step of performing feature transformation on the N frequency band energy values to obtain M earprint feature data includes: Based on the noise frequency range, the N frequency band energy values are filtered to obtain N' target frequency band energy values. The N' target frequency band energy values include all frequency band energy values located outside the noise frequency range. Within the noise frequency range, the similarity between the first frequency response curve and the second frequency response curve is less than or equal to a set difference threshold. The first frequency response curve is used to represent the frequency response trend of the earprint data of the left ear of the target user group with frequency. The second frequency response curve is used to represent the frequency response trend of the earprint data of the right ear of the target user group with frequency. The N' target frequency band energy values are respectively subjected to feature transformation to obtain the M earprint feature data, where N' is a positive integer less than or equal to N. Each earprint feature data is obtained based on the N' target frequency band energy values and the corresponding feature dimension parameters. Different earprint feature data correspond to different feature dimension parameters.
2. The method according to claim 1, characterized in that, The step of performing feature transformation on the N' target frequency band energy values to obtain the M earprint feature data includes: Logarithmic transformation is performed on the N' target frequency band energy values to obtain N' logarithmic data, where the logarithmic data is the logarithmic value of the sum of the corresponding target frequency band energy value and the set parameter; Based on the set M feature dimension parameters, the N' logarithmic data are subjected to discrete cosine transform to obtain the M earprint feature data.
3. The method according to claim 1, characterized in that, The preprocessing of the user's earprint data to obtain preprocessed data includes: Based on a preset short-time energy threshold and a zero-crossing rate threshold, the earprint data includes multiple earprint audio signals that are denoised to obtain at least one target audio signal. The short-time energy of the target audio signal is greater than or equal to the short-time energy threshold, and the zero-crossing rate of the target audio signal is less than or equal to the zero-crossing rate threshold. In the at least one target audio signal, pre-emphasis processing is performed on each of the target audio signals to obtain at least one pre-emphasis audio signal, wherein the pre-emphasis processing includes first-order differential processing; The at least one pre-emphasized audio signal is subjected to frame segmentation processing to obtain frame signal data; The frame signal data is windowed based on a set window function to obtain the preprocessed data. The window function includes a Hamming window function. The spectral data is obtained by fast Fourier transform based on the preprocessed data.
4. The method according to claim 1, characterized in that, The target filter is a Mel filter, and the N target filters include a first target filter and a second target filter. The center frequency of the first target filter is higher than the reference frequency, and the center frequency of the second target filter is lower than the reference frequency. Among them, the center Mel frequencies of multiple first target filters are distributed at equal intervals.
5. A feature acquisition device, characterized in that, The device includes: The preprocessing module is used to preprocess the user's target earprint data to obtain preprocessed data, wherein the preprocessing includes at least one of noise reduction processing, pre-emphasis processing, frame segmentation processing, and windowing processing; An energy extraction module is used to convert the preprocessed data into spectral data and extract the energy of the spectral data based on N pre-set target filters to obtain N frequency band energy values. In the N target filters, the bandwidth of each target filter is positively correlated with the distance parameter corresponding to the target filter. The distance parameter indicates the frequency difference between the center frequency and the reference frequency of the corresponding target filter. The center frequencies of the N target filters are symmetrically set around the set reference frequency, which is contained within the frequency band corresponding to the frequency response peak of the earprint data of the target user group. The target user group includes the users. The frequency band energy values are obtained by convolving the corresponding target filter and the spectral data in the corresponding frequency band. The feature conversion module is used to perform feature conversion on the N frequency band energy values respectively to obtain M earprint feature data. The user's identity identifier includes the M earprint feature data, where N and M are integers greater than 1. The feature conversion module includes: A filtering unit is used to filter the N frequency band energy values based on the noise frequency range to obtain N' target frequency band energy values, wherein the N' target frequency band energy values include: all frequency band energy values located outside the noise frequency range among the N frequency band energy values, and within the noise frequency range, the similarity between the first frequency response curve and the second frequency response curve is less than or equal to a set difference threshold, wherein the first frequency response curve is used to represent the frequency response trend of the earprint data of the left ear of the target user group with frequency, and the second frequency response curve is used to represent the frequency response trend of the earprint data of the right ear of the target user group with frequency; The feature conversion unit is used to perform feature conversion on the N' target frequency band energy values respectively to obtain the M earprint feature data, where N' is a positive integer less than or equal to N. Each earprint feature data is obtained based on the N' target frequency band energy values and the corresponding feature dimension parameters. Different earprint feature data have different feature dimension parameters.
6. The apparatus according to claim 5, characterized in that, The feature transformation unit is specifically used for: Logarithmic transformation is performed on the N' target frequency band energy values to obtain N' logarithmic data, where the logarithmic data is the logarithmic value of the sum of the corresponding target frequency band energy value and the set parameter; Based on the set M feature dimension parameters, the N' logarithmic data are subjected to discrete cosine transform to obtain the M earprint feature data.
7. The apparatus according to claim 5, characterized in that, The target filter is a Mel filter, and the N target filters include a first target filter and a second target filter. The center frequency of the first target filter is higher than the reference frequency, and the center frequency of the second target filter is lower than the reference frequency. The center Mel frequencies of the multiple first target filters are distributed at equal intervals. And / or, The preprocessing module is specifically used for: Based on a preset short-time energy threshold and a zero-crossing rate threshold, the earprint data includes multiple earprint audio signals that are denoised to obtain at least one target audio signal. The short-time energy of the target audio signal is greater than or equal to the short-time energy threshold, and the zero-crossing rate of the target audio signal is less than or equal to the zero-crossing rate threshold. In the at least one target audio signal, pre-emphasis processing is performed on each of the target audio signals to obtain at least one pre-emphasis audio signal, wherein the pre-emphasis processing includes first-order differential processing; The at least one pre-emphasized audio signal is subjected to frame segmentation processing to obtain frame signal data; The frame signal data is windowed based on a set window function to obtain the preprocessed data. The window function includes a Hamming window function. The spectral data is obtained by fast Fourier transform based on the preprocessed data.