Control Method for Headphone Noise Reduction

By obtaining the ambient audio signal, position information and attitude information of the headphones, and using the improved Transformer model and proxy attention mechanism to dynamically adjust the noise reduction parameters, the problem of low scene recognition efficiency and accuracy of the adaptive noise reduction method of wireless headphones is solved, reducing power consumption, improving user experience and adaptability.

CN119110198BActive Publication Date: 2025-07-11JIANGXI RUISHENG ELECTRONIC CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411159685.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-22
Publication Date
2025-07-11
Estimated Expiration
2044-08-22

AI Technical Summary

Technical Problem

The existing wireless headphone adaptive noise reduction method has low efficiency and accuracy in scene recognition, high power consumption, and cannot effectively adapt to the variable noise environment and user diversified needs.

Method used

By obtaining the ambient audio signal, position information and attitude information of the headset, using the improved Transformer model and agent attention mechanism, combined with the power status of the headset, the noise reduction parameters are dynamically adjusted to optimize the noise reduction effect.

Benefits of technology

It improves the recognition efficiency and accuracy of headphones in complex noise environments, reduces power consumption, improves user experience, and enhances adaptability to diverse usage scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119110198B_ABST
    Figure CN119110198B_ABST
Patent Text Reader

Abstract

The present invention discloses a control method for headphone noise reduction, which includes obtaining an environmental audio signal of the scene where the headphone is currently located; detecting whether the current position information of the headphone can be obtained; if not, determining the target scene type according to the environmental audio signal; if so, obtaining the current position information and remaining power parameter of the headphone; detecting whether the remaining power parameter is greater than a preset power threshold; if not, determining the target scene type according to the position information and the environmental audio signal; if so, obtaining the current attitude information of the headphone; determining the target scene type based on the environmental audio signal, position information, attitude information and an improved Transformer model with a proxy attention mechanism configured in advance; obtaining the noise reduction parameter corresponding to the target scene type as the target noise reduction parameter. This application divides different noise reduction parameter acquisition modes based on the position information and power state, optimizes the acquisition method of the target noise reduction parameter, and improves the efficiency and accuracy of adaptive noise reduction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of headphone noise reduction, and particularly to a control method for headphone noise reduction. Background Art

[0002] With the rapid development of electronic technology, wireless headphones have become an essential item for people's daily travel. When facing various relatively noisy environments, people usually choose to wear wireless headphones to reduce the impact of environmental noise.

[0003] Different noise environments have different requirements for the degree of noise reduction. To better cope with the changing noise environments, some existing wireless headphones adopt an adaptive noise reduction method to reduce the impact of environmental noise. However, the existing adaptive noise reduction methods cannot relatively accurately and efficiently identify the noise scenarios, resulting in a low matching degree of the noise reduction degree requirements, and the existing adaptive noise reduction has a high power consumption, thus affecting the user experience.

[0004] In view of this, it is necessary to provide a control method for headphone noise reduction to solve the above problems. Summary of the Invention

[0005] In view of the deficiencies in the prior art, the present invention provides a control method for headphone noise reduction, aiming to solve the technical problems of low efficiency and accuracy in scene recognition of headphone adaptive noise reduction, unreasonable power consumption configuration, and low adaptability to diverse user needs and complex usage scenarios.

[0006] To achieve the above object, a first aspect of the present invention provides a control method for headphone noise reduction, which includes:

[0007] Obtain the environmental audio signal of the current scene where the headphone is located;

[0008] Detect whether the current position information of the headphone can be obtained;

[0009] If not, determine the target scene type according to the environmental audio signal;

[0010] If so, obtain the current position information and remaining battery parameter of the headphone;

[0011] Detect whether the remaining battery parameter is greater than a preset battery threshold;

[0012] If not, determine the target scene type according to the position information and the environmental audio signal;

[0013] If so, obtain the current attitude information of the headphone;

[0014] Determine the target scene type based on the environmental audio signal, location information, attitude information, and a preset improved Transformer model configured with an agent attention mechanism;

[0015] Obtain the noise reduction parameter corresponding to the target scene type as the target noise reduction parameter.

[0016] In a preferred embodiment, the step of determining the target scene type according to the environmental audio signal includes:

[0017] Obtain the first spectrum and the second spectrum according to the environmental audio signal;

[0018] Based on the first spectrum and a preset first extraction mode, obtain the amplitude feature matrix;

[0019] Based on the second spectrum and a preset second extraction mode, obtain a number of energy eigenvalues;

[0020] Obtain the similarity coefficient between the amplitude feature matrix and the comparison matrix of the preset scene types;

[0021] Based on the similarity coefficient, screen out the specific scene types from the scene types;

[0022] Determine the target scene type according to the energy eigenvalues, similarity coefficient, and specific scene type.

[0023] In a preferred embodiment, the step of obtaining the amplitude feature matrix based on the first spectrum and a preset first extraction mode includes:

[0024] Obtain the preset interception time width, the first interception frequency width, and the initial matrix;

[0025] Based on the first spectrum, interception time width, and first interception frequency width, obtain a number of first target spectra;

[0026] Sequentially associate the first target spectra with the elements of the initial matrix one by one;

[0027] Calculate the amplitude feature parameters of the first target spectra respectively;

[0028] According to the magnitudes of the amplitude feature parameters, adjust the positions of the corresponding elements on the same preset target dimension in the initial matrix to obtain the amplitude feature matrix.

[0029] In a preferred embodiment, the step of determining the target scene type according to the energy eigenvalues, similarity coefficient, and specific scene type includes:

[0030] Obtain the preset energy interval of the specific scene type;

[0031] Sequentially determine whether the energy eigenvalues fall into the corresponding energy intervals;

[0032] Obtain the number of falls of the energy eigenvalue falling into the energy interval of each specific scenario type;

[0033] Based on the number of falls and the similarity coefficient, determine the target scenario type from the specific scenario types.

[0034] In a preferred embodiment, the steps of determining the target scenario type according to the location information and the environmental audio signal include:

[0035] Obtain the key location information from the location information;

[0036] Obtain the environmental audio spectrum from the environmental audio signal;

[0037] Based on the key location information, determine the initial scenario type and the associated scenario type;

[0038] According to the environmental audio spectrum and the preset extraction mode, obtain a number of energy eigenvalues;

[0039] According to the energy eigenvalues, determine the target scenario type from the initial scenario type and the associated scenario type.

[0040] In a preferred embodiment, the steps of determining the initial scenario type and the associated scenario type based on the key location information include:

[0041] Obtain the adaptation value of the key location information and the preset scenario type;

[0042] Based on the adaptation value, determine the initial scenario type from the preset scenario types;

[0043] Obtain the association value between the initial scenario type and the remaining scenario types;

[0044] Based on the association value and the preset first threshold, obtain the associated scenario type from the remaining scenario types.

[0045] In a preferred embodiment, the steps of determining the target scenario type from the initial scenario type and the associated scenario type according to the energy eigenvalues include:

[0046] Obtain the energy interval of the initial scenario type and the energy interval of the associated scenario type;

[0047] Obtain the number of falls of the energy eigenvalue falling into the energy interval of the initial scenario type and the energy interval of the associated scenario type;

[0048] Based on the number of falls, obtain the target scenario type.

[0049] In a preferred embodiment, the steps of determining the target scene type based on the environmental audio signal, location information, attitude information, and an improved Transformer model configured with a proxy attention mechanism include:

[0050] Extract an audio feature sequence from the environmental audio signal;

[0051] Extract an auxiliary feature sequence from the location information and attitude information;

[0052] Input the audio feature sequence and the auxiliary feature sequence into the improved Transformer model configured with a proxy attention mechanism to obtain the matching values of each scene type;

[0053] Obtain the target scene type according to the magnitudes of the matching values.

[0054] In a preferred embodiment, the steps of inputting the audio feature sequence and the auxiliary feature sequence into the improved Transformer model configured with a proxy attention mechanism to obtain the matching values of each scene type include:

[0055] Perform mapping and encoding on the audio feature sequence and the auxiliary feature sequence to respectively obtain an audio enhancement sequence and an auxiliary enhancement sequence;

[0056] Based on the audio enhancement sequence, the auxiliary enhancement sequence, and a preset proxy attention mechanism, obtain an enhanced feature matrix;

[0057] Based on a preset fully connected layer and the enhanced feature matrix, obtain the matching values of each scene type.

[0058] In a preferred embodiment, the steps of obtaining an enhanced feature matrix based on the audio enhancement sequence, the auxiliary enhancement sequence, and a preset proxy attention mechanism include:

[0059] According to the audio enhancement sequence and a preset transformation weight matrix, respectively obtain a first matrix, a second matrix, and a third matrix;

[0060] According to the auxiliary enhancement sequence and a preset proxy transformation matrix, obtain a proxy matrix;

[0061] Perform interactive calculations on the proxy matrix, the second matrix, and the third matrix to obtain a fusion matrix;

[0062] Perform interactive calculations on the first matrix, the proxy matrix, and the fusion matrix to obtain an attention weight matrix;

[0063] Perform weighted summation on the attention weight matrix and the third matrix to obtain the enhanced feature matrix.

[0064] The second aspect of the present invention provides a headset, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the control method for headset noise reduction described in any one of the above are implemented.

[0065] The third aspect of the present invention provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps of the control method for headset noise reduction described in any one of the above are implemented.

[0066] The beneficial effects of the present invention are as follows: It divides modes for obtaining noise reduction parameters according to different parameter information through the current position information and remaining power of the headset. When the position information is not obtained, it directly uses the environmental audio signal for scene recognition to quickly determine the scene type; after obtaining the position information, if the power does not meet the requirements, it combines the environmental audio signal and the position information to achieve a quick preliminary screening and secondary repeated screening of the scene type; if the power meets the requirements, it uses an improved Transformer model and multi-source feature information to identify and classify complex audio scene types while maintaining a relatively low computational complexity; rationally configures the noise reduction power consumption of the headset, optimizes the way of obtaining noise reduction parameters, improves the efficiency and accuracy of adaptive noise reduction, improves the adaptability of the headset to diverse user usage requirements and usage scenarios, and improves the user experience. Description of the Drawings

[0067] Figure 1 It is the first flowchart of the control method for headset noise reduction disclosed in the embodiments of the present invention;

[0068] Figure 2 It is the second flowchart of the control method for headset noise reduction disclosed in the embodiments of the present invention;

[0069] Figure 3 It is the third flowchart of the control method for headset noise reduction disclosed in the embodiments of the present invention;

[0070] Figure 4 It is the fourth flowchart of the control method for headset noise reduction disclosed in the embodiments of the present invention;

[0071] Figure 5 It is the module structure diagram of the headset disclosed in the embodiments of the present invention. Detailed Embodiments

[0072] In the present invention, the terms "arranged", "provided with", and "connected" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral structure; it may be a mechanical connection or an electrical connection; it may be directly connected, or indirectly connected through an intermediate medium, or there may be internal communication between two devices, components, or parts. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0073] The terms "first" and "second" are only used for descriptive purposes and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In the description of the present application, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically defined.

[0074] Furthermore, in addition to being able to represent an orientation or positional relationship, some of the above terms may also be used to represent other meanings. For example, the term "upper" may also be used to represent a certain attachment relationship or connection relationship in some cases. For those of ordinary skill in the art, the specific meanings of these terms in the present invention can be understood according to specific circumstances.

[0075] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0076] The following is the content of the first aspect of the present invention:

[0077] Please refer to Figure 1 , in this embodiment, the steps of the control method for headphone noise reduction include:

[0078] S1. Obtain the environmental audio signal of the current scene where the headphone is located.

[0079] Among them, the headphone mentioned in this embodiment is a headphone with an active noise reduction function. The headphone can be a wireless headphone or a wired headphone, and it includes a processor and an audio output unit. The processor can generate an audio signal with an opposite phase to the environmental noise through the obtained noise reduction parameters, thereby canceling the external environmental noise.

[0080] The environmental audio signal can be obtained through the acquisition microphone array provided on the headphone, or through the acquisition microphone array on the terminal device connected to the headphone. Specifically, the way for the acquisition microphone array to obtain the environmental audio signal can be real-time acquisition or periodic acquisition.

[0081] S2. Detect whether the current position information of the earphone can be obtained;

[0082] S3. If the current position information of the earphone cannot be obtained, determine the target scene type according to the environmental audio signal.

[0083] Among them, the position information of the earphone can be the position information centered on the earphone or the position information centered on the terminal device connected to the earphone. Correspondingly, the acquisition method of the position information of the earphone can be direct acquisition or indirect acquisition. Specifically, a positioning module can be set inside the earphone, and the current position information of the earphone can be directly obtained through this positioning module; it can also be indirectly obtained through a terminal device (such as a mobile phone, a smart watch, etc.) with a positioning module set inside, and approximate the position information of the terminal device as the position information of the earphone. During the daily use of the earphone, in most cases, the earphone is connected to a terminal device for use. Preferably, the position information of the terminal device can be selected to indirectly obtain the position information of the earphone. Specifically, by detecting the output situation of the positioning module inside the earphone or the output situation of the terminal connected to the earphone, the acquisition situation of the current position information of the earphone can be determined.

[0084] By detecting whether the current position information of the earphone can be obtained, it can be determined whether the scene type where the earphone is currently located can be recognized in combination with the position information, so as to trigger a mode of obtaining target noise reduction parameters based on different types of parameter information. After detecting that the current position information of the earphone is not obtained, the environmental audio signal can be converted into an environmental audio spectrum, and then the corresponding audio feature information can be extracted from the environmental audio spectrum to determine the type of the scene where the earphone is currently located from the preset scene types. Specifically, the environmental audio signal can be converted into a single type of environmental audio spectrum or different types of environmental audio spectra. Correspondingly, the audio features extracted from the environmental audio spectrum can also be single type of audio feature information or multi-type of audio feature information. Therefore, the target scene type can be determined based on single type of audio feature information or multi-type of audio feature information. Preferably, the method of determining the target scene type based on multi-type of audio feature information is adopted. It is easy to understand that through multi-type of audio feature information, the target scene type can be determined relatively more accurately.

[0085] S4. If the current position information of the earphone can be obtained, obtain the current position information of the earphone and the remaining battery parameter;

[0086] S5. Detect whether the remaining battery parameter is greater than the preset battery threshold;

[0087] S6. If the remaining battery parameter is not greater than the battery threshold, determine the target scene type according to the position information and the environmental audio signal.

[0088] Among them, the power threshold is a power parameter preset for distinguishing different noise reduction parameter acquisition modes, which can be set according to the actual design requirements.

[0089] After obtaining the current position information and remaining power parameter of the earphone, compare the remaining power parameter with the power threshold, and then enable the corresponding target noise reduction parameter acquisition mode according to the comparison result of the remaining power parameter and the power threshold.

[0090] When it is determined that the remaining power parameter is not greater than the power threshold, it means that the current power of the earphone is relatively low, and the battery life of the earphone should be ensured first. The target noise reduction parameter acquisition mode that gives priority to ensuring the use of the earphone can be enabled. After triggering the corresponding target noise reduction parameter acquisition mode, the required audio feature information and position feature information are respectively extracted from the position information and the environmental audio signal, and then the obtained audio feature information is used to determine the target scene type from the preset scene types.

[0091] Specifically, after obtaining the audio feature information and position feature information, a progressive screening method of comparing the feature information successively can be adopted to determine the target scene type; or a screening method of comparing the feature information in parallel and then adding weights can be adopted to determine the target scene type, which can be selected according to the actual design requirements.

[0092] It is easy to understand that by combining the position information and the environmental audio signal, the acquisition method of the target noise reduction parameter can be optimized, the calculation amount of scene recognition can be reduced, the efficiency and accuracy of adaptive noise reduction can be improved, the power consumption of the earphone can be reduced to a certain extent, the battery life of the earphone can be guaranteed, and the user experience can be improved.

[0093] S7. If the remaining power parameter is greater than the power threshold, obtain the current attitude information of the earphone;

[0094] S8. Based on the environmental audio signal, position information, attitude information and a preset improved Transformer model configured with a proxy attention mechanism, determine the target scene type;

[0095] S9. Obtain the noise reduction parameters corresponding to the target scene type as the target noise reduction parameters.

[0096] Among them, the improved Transformer model uses the encoding module of the Transformer model and introduces a proxy attention mechanism to replace the original self-attention mechanism in the Transformer model, which can realize more accurate and efficient recognition and classification of complex sound scenes by using multi-source feature information. The proxy attention mechanism consists of two softmax attention operations, and a set of additional proxy vectors A with fewer parameters is introduced into the original ternary attention paradigm (Q, K, V) of the attention mechanism to form a new quaternary proxy attention paradigm (Q, A, K, V).

[0097] When it is determined that the remaining power parameter is greater than the power threshold, it means that the current power of the earphone is relatively large, which can support the earphone to obtain the target noise reduction parameter acquisition mode with higher recognition and analysis accuracy. After triggering the corresponding target noise reduction parameter acquisition mode, the current attitude information of the earphone is obtained. Specifically, the current attitude information of the earphone can be obtained through an attitude sensor set inside the earphone, and the attitude sensor includes but is not limited to an acceleration sensor, a gyroscope, a magnetic sensor, etc.

[0098] After obtaining the current attitude information of the earphone, the current environmental audio signal, position information and attitude information of the earphone are preprocessed to obtain a feature sequence that can be input into the improved Transformer model from the environmental audio signal, position information and attitude information. Specifically, the feature sequence can be a sequence containing a single type of feature information or a fusion sequence that fuses multiple types of feature information, which can be selected according to the actual design requirements. Correspondingly, the acquisition method of the feature sequence can be the corresponding feature sequences respectively extracted from the denoised environmental audio signal, position information and attitude information, or a fusion sequence that fuses the required features extracted from the denoised environmental audio signal, position information and attitude information according to the design requirements.

[0099] After obtaining the required feature sequence, the feature sequence is input into the improved Transformer model. The improved Transformer model first further processes the feature sequence to transform it into an enhanced sequence with high dimensions and sorting information. The obtained enhanced sequence then passes through the proxy attention mechanism of the improved Transformer model to obtain an enhanced feature matrix containing weight information. Then, through the transformation and calculation of the enhanced feature matrix by the fully connected layer, the matching values of each preset scene type can be obtained, and then the target scene type can be determined according to the matching values.

[0100] It can be understood that in the case where the current remaining power of the earphone is relatively high, by improving the Transformer model and multi-source feature information, the function of identifying and classifying complex audio scenarios where the earphone is located is realized. By using the proxy attention mechanism of the improved Transformer model, while maintaining a relatively low computational complexity, the integration and recognition of multi-source feature information are realized, the accuracy of dynamic adjustment of recognition attention weights is improved, and thus the efficiency and accuracy of the earphone in recognizing complex audio scenarios are improved. Furthermore, the matching degree of the earphone's adaptive noise reduction is improved, the effect of the earphone's adaptive noise reduction is improved, and the user experience is improved.

[0101] After determining the corresponding target noise reduction parameter acquisition mode according to the current position information and remaining power parameter of the earphone, start the corresponding target noise reduction parameter acquisition mode. Based on the required parameter information, determine the target scene type, and then obtain the pre-set noise reduction parameters corresponding to the target scene type, and the target noise reduction parameters adapted to the current scene where the earphone is located can be obtained.

[0102] It can be understood that in this application, the mode of obtaining noise reduction parameters according to different parameter information is divided through the current position information and remaining power of the earphone. When the position information is not obtained, the environmental audio signal is directly used for scene recognition to quickly determine the scene type. After the position information is obtained, if the power does not meet the requirements, the environmental audio signal and the position information are combined to achieve a quick preliminary screening and secondary repeated screening of the scene type. If the power meets the requirements, the improved Transformer model and multi-source feature information are used to realize the recognition and classification of complex audio scene types while maintaining a relatively low computational complexity. The noise reduction power consumption of the earphone is reasonably configured, the acquisition method of noise reduction parameters is optimized, the efficiency and accuracy of adaptive noise reduction are improved, the adaptability of the earphone to diverse user usage requirements and usage scenarios is improved, and the user experience is improved.

[0103] Further, please refer to Figure 2 , in one embodiment, the step S3 of determining the target scene type according to the environmental audio signal includes:

[0104] S31. Obtain a first frequency spectrum and a second frequency spectrum according to the environmental audio signal;

[0105] S32. Based on the first frequency spectrum and a preset first extraction mode, obtain an amplitude feature matrix;

[0106] S33. Based on the second frequency spectrum and a preset second extraction mode, obtain a number of energy eigenvalues;

[0107] S34. Obtain the similarity coefficient between the amplitude feature matrix and a comparison matrix of a preset scene type;

[0108] S35. Screen out a specific scene type from the scene types based on the similarity coefficient;

[0109] S36. Determine the target scene type according to the energy eigenvalue, the similarity coefficient, and the specific scene type.

[0110] Among them, both the first frequency spectrum and the second frequency spectrum are audio frequency spectrums obtained by converting the environmental audio signal, and the frequency spectrum types of the first frequency spectrum and the second frequency spectrum are different. The first frequency spectrum and the second frequency spectrum can be selected from frequency spectrum types such as the spectrum obtained by short-time Fourier transform, the Mel spectrum, etc. Preferably, the first frequency spectrum is the spectrum obtained by short-time Fourier transform, and the second frequency spectrum is the Mel spectrum.

[0111] The first extraction mode is an extraction logic method for extracting features set according to the characteristics of the first frequency spectrum. The amplitude feature matrix is a matrix obtained from the first frequency spectrum based on the first extraction mode, which can, to a certain extent, represent the characteristic situation of the amplitude intensity change of the environmental audio signal in each frequency band at the overall level.

[0112] The second extraction mode is an extraction logic method for extracting features set according to the characteristics of the second frequency spectrum. The energy eigenvalue is the energy parameter information corresponding to the frequency band in the second frequency spectrum, which can be used to represent the energy situation of the frequency band.

[0113] The comparison matrix is a matrix extracted from the pre-trained scene types and can be used to match and compare with the amplitude feature matrix. The similarity coefficient is a reflection value of the similarity degree between the amplitude feature matrix and the comparison matrix. The specific scene type is the scene type whose matching degree between the comparison matrix and the amplitude feature matrix meets the requirements.

[0114] Specifically, after converting the environmental audio signal into the first frequency spectrum, the first extraction mode is started, and the first frequency spectrum can be segmented based on two dimensions of time and frequency to obtain the required first target frequency spectrum. Specifically, it can be to intercept the first frequency spectrum from the time dimension first and then from the frequency dimension to obtain the first target frequency spectrum; it can also be to intercept the first frequency spectrum from the frequency dimension first and then from the time dimension to obtain the first target frequency spectrum.

[0115] After obtaining the required first target frequency spectrum, calculate the amplitude parameters corresponding to each first target frequency spectrum respectively. Specifically, the amplitude can be the maximum value of the amplitude within the corresponding first target frequency spectrum, or the average value of the amplitude within the corresponding first target frequency spectrum, or the variance of the amplitude within the corresponding first target frequency spectrum. Preferably, the amplitude is the average value of the amplitude within the first target frequency spectrum. It can be understood that selecting the average value can better represent the overall situation of the amplitude within the first target frequency spectrum and can effectively reduce the interference of mutation signals.

[0116] After obtaining the first target spectrum and the corresponding amplitude parameters, the amplitude feature matrix can be obtained by numbering the first target spectrum and arranging the numbers based on the position of the first target spectrum in the first spectrum to obtain an initial matrix. Then, according to the magnitude of the amplitude and the preset target dimension, the positions of the elements in the same target dimension in the initial matrix are adjusted to obtain the amplitude feature matrix. For example, if the first spectrum has the time dimension as rows (or the abscissa) and the frequency dimension as columns (or the ordinate), based on the arrangement of the time dimension and frequency dimension of the first spectrum, the numbers of the first target spectrum are arranged to obtain the initial matrix. Then, the column dimension of the initial matrix is the target dimension, and the elements in the same column are sorted according to the magnitude of the amplitude corresponding to each number to obtain the amplitude feature matrix.

[0117] It is easy to understand that the arrangement of the elements in the amplitude feature matrix to a certain extent reflects the amplitude intensity change between different frequency bands under the same time dimension. Therefore, the characteristic situation of the amplitude intensity change of the environmental audio signal in the current scene where the headset is located can be determined on the overall level. Furthermore, according to the similarity between the amplitude feature matrix and the comparison matrix, the matching degree between the current scene where the headset is located and the scene type can be determined.

[0118] Another way to obtain the amplitude feature matrix is to preset an initial matrix. The number of elements in this initial matrix is the same as the number of the first target spectra, and the arrangement of the elements in the initial matrix is the same as the position arrangement of the first target spectra in the first spectrum. Then, the first target spectra are sequentially associated with the elements in the preset initial matrix one by one, and then according to the magnitude of the amplitude corresponding to the first target spectra, the arrangement positions of the elements in the same target dimension in the initial matrix are adjusted to obtain the amplitude feature matrix.

[0119] After obtaining the second spectrum, the frequency band range relatively sensitive to the human ear can be intercepted, and then according to the preset frequency width, the intercepted spectrum is intercepted again to obtain a number of second target spectra. Or, it can be intercepted first with the preset frequency width, and then the spectrum located in the frequency band range relatively sensitive to the human ear is obtained to get a number of second target spectra.

[0120] After obtaining a number of second target spectra, the energy characteristic values corresponding to the second target spectra are calculated respectively. The energy characteristic value can be the maximum energy value in the second target spectrum or the average energy value of the second target spectrum. Preferably, the energy characteristic value is the average energy value, which can better reflect the energy distribution of the specific frequency band of the second target spectrum.

[0121] After obtaining the amplitude feature matrix, the preset comparison matrix for each scenario type can be obtained, and then based on the calculation formula of the similarity coefficient, the similarity coefficient between the amplitude feature matrix and the comparison matrix can be calculated item by item. Specifically, the dimension similarity coefficient value between the amplitude feature matrix and the comparison matrix on the same preset target dimension can be calculated first, and then based on the dimension similarity coefficient, the similarity coefficient between the amplitude feature matrix and the comparison matrix can be obtained. The calculation formula of the similarity coefficient can be the cosine similarity calculation formula, or the Euclidean distance calculation formula, etc., and can be selected according to the actual design requirements. After obtaining the dimension similarity coefficient, the average value of each dimension similarity coefficient can be calculated as the similarity coefficient between the amplitude feature matrix and the comparison matrix, or each dimension similarity coefficient can be weighted calculated to obtain the similarity coefficient between the amplitude feature matrix and the comparison matrix, and can be selected according to the actual design requirements.

[0122] After obtaining the similarity coefficient between the amplitude feature matrix and each comparison matrix, the matching degree between the current scenario of the earphone and each preset scenario type can be determined, and then the required specific scenario type can be screened out from the preset scenario types according to whether the similarity coefficient meets the preset requirements. It is easy to understand that the specific scenario type can be screened from the preset scenario types according to the magnitude of the similarity coefficient, or the required number of similarity coefficients can be obtained, and then the preset number of specific scenario types can be screened out from the preset scenario types.

[0123] After obtaining the energy eigenvalue, similarity coefficient and specific scenario type, the energy interval corresponding to the frequency band of each preset specific scenario type can be obtained, and then it can be determined whether the energy eigenvalue is within the corresponding energy interval. After obtaining the situation where the energy eigenvalue falls within the energy interval, the target scenario type can be screened progressively from the specific scenario type according to the falling situation of the energy eigenvalue and the magnitude of the similarity coefficient; or the target scenario type can be screened from the specific scenario type by combining the falling situation of the energy eigenvalue, the similarity coefficient and the preset weight coefficient; or it can be a combination of the above two methods to screen the target scenario type from the specific scenario type, and can be selected according to the actual design requirements.

[0124] It can be understood that by using the recognition and screening method of bispectrum and double extraction, a progressive recognition and screening of the scenario type is carried out from different feature angles. The similarity coefficient between the amplitude feature matrix and the comparison amplitude feature matrix is used to preliminarily screen the scenario type from the overall level, and then combined with the energy eigenvalue and the similarity coefficient, the specific scenario type is repeatedly screened from the local level of the sensitive frequency band to determine the target scenario type, which optimizes the acquisition method of the earphone noise reduction parameters, improves the efficiency and accuracy of scenario recognition, reduces the calculation amount of scenario recognition, reduces the power consumption of the earphone, and improves the user experience.

[0125] Further, in one embodiment, step S32 of obtaining the amplitude feature matrix based on the first spectrum and a preset first extraction mode includes:

[0126] S321. Obtain a preset truncation time width, a first truncation frequency width, and an initial matrix;

[0127] S322. Based on the first spectrum, the truncation time width, and the first truncation frequency width, obtain a number of first target spectra;

[0128] S323. Sequentially associate each first target spectrum with the elements of the initial matrix one by one;

[0129] S324. Calculate the amplitude feature parameters of the first target spectra respectively;

[0130] S325. According to the magnitudes of the amplitude feature parameters, adjust the positions of the corresponding elements in the same preset target dimension of the initial matrix to obtain the amplitude feature matrix.

[0131] Wherein, the truncation time width is the frame truncation unit for segmenting the first spectrum in the time dimension. The first truncation frequency width is the frequency truncation unit for segmenting the first spectrum in the frequency dimension. The parameter magnitudes of the truncation time width and the first truncation frequency width can be set according to the actual design requirements. The first target spectrum is the spectrum obtained after the first spectrum is truncated in both the time and frequency dimensions. The initial matrix is a preset matrix that can be used to associate a number of first target spectra. The target dimension is the row dimension or the column dimension, and the type of this target dimension depends on the dimension where the time dimension is located in the first spectrum.

[0132] After obtaining the first spectrum, the corresponding first extraction mode is triggered to obtain a preset truncation time width, a first truncation frequency width, and an initial matrix. According to the truncation time width and the first truncation frequency width, the first spectrum is segmented in the time dimension and the frequency dimension to obtain a number of first target spectra.

[0133] After obtaining a number of first target spectra, according to the arrangement pattern of the first target spectra in the first spectrum, the first target spectra are sequentially associated with the elements of the initial matrix one by one. That is, after intercepting in the time dimension and the frequency dimension, the arrangement of the first target spectra in the first spectrum can be regarded as a specific matrix, and the elements at the same positions in this specific matrix and the initial matrix are associated one by one. For example, in the first spectrum, the first target spectra in the first column are A1, A2, A3 in sequence, the first target spectra in the second column are B1, B2, B3 in sequence, and the first target spectra in the third column are C1, C2, C3 in sequence. In the initial matrix, the elements in the first column are X1, X2, X3 in sequence, the elements in the second column are Y1, Y2, Y3 in sequence, and the elements in the third column are Z1, Z2, Z3 in sequence. Then, A1, A2, A3 can be associated with X1, X2, X3 one by one, B1, B2, B3 can be associated with Y1, Y2, Y3 one by one, and C1, C2, C3 can be associated with Z1, Z2, Z3 one by one.

[0134] Calculate the amplitude characteristic parameters of the first target spectra respectively. The amplitude characteristic parameter can be the average amplitude of the first target spectrum, or the maximum amplitude of the first target spectrum, or the amplitude variance of the first target spectrum, and can be selected according to the actual design requirements. After calculating the amplitude characteristic parameters of each first target spectrum, adjust the positions of the corresponding elements in the initial matrix on the same preset target dimension according to the magnitudes of the amplitude characteristic parameters. Specifically, if the time dimension is the row dimension in the first spectrum, the target dimension is the column dimension; if the time dimension is the column dimension in the first spectrum, the target dimension is the row dimension. After adjusting the positions of the elements in the initial matrix, the amplitude characteristic matrix can be obtained.

[0135] It can be understood that by adjusting the positions of the elements in the initial matrix according to the magnitudes of the amplitude characteristic parameters of the first target spectra to obtain the amplitude characteristic matrix, the characteristic situation of the amplitude intensity change in different frequency bands of the environmental audio signal under the same time dimension can be relatively clearly reflected, and then the characteristic situation of the amplitude intensity change at the overall level can be reflected, improving the efficiency of feature information extraction and the efficiency of subsequent scene recognition and screening.

[0136] Further, in one embodiment, the steps S33 of obtaining a number of energy eigenvalues based on the second spectrum and a preset second extraction mode include:

[0137] S331. Obtain a preset sensitive frequency range and a second interception frequency width;

[0138] S332. Based on the second spectrum, the sensitive frequency range and the second interception frequency width, obtain a number of second target spectra;

[0139] S333. Calculate the energy characteristic parameters of the second target spectrum respectively to obtain the energy characteristic value.

[0140] Among them, the sensitive frequency range is a specific frequency range for feature extraction preset according to the human ear's auditory characteristics. Preferably, the sensitive frequency range is from 1000 Hz to 4000 Hz. The second truncation bandwidth is the truncation unit for further truncating the spectrum within the sensitive frequency range from the frequency dimension. The second target spectrum is the spectrum obtained after the second spectrum is truncated twice. The energy characteristic parameter is the energy characteristic information extracted from the second target spectrum.

[0141] After obtaining the second spectrum, the corresponding second extraction mode is triggered to obtain the preset sensitive frequency range and the second truncation bandwidth, and then the second spectrum is truncated according to the sensitive frequency range and the second truncation bandwidth. Specifically, the second spectrum can be processed according to the sensitive frequency range first to obtain the sensitive spectrum, and then the sensitive spectrum is processed using the second truncation bandwidth to obtain a number of second target spectra. It can also be that the second spectrum is processed according to the second truncation bandwidth first to obtain a number of second sub-spectra, and then the second sub-spectra within the sensitive frequency range are obtained from the number of second sub-spectra as the second target spectrum. It is easy to understand that since the human ear has different sensitivities to different frequencies, based on the sensitivity difference of frequencies, obtaining the spectrum of the frequency range relatively sensitive to the human ear from the second spectrum as the second target spectrum can obtain a feature extraction range that is more in line with the human ear's auditory characteristics and has a more obvious impact on the noise reduction effect, reducing the amount of feature extraction required, improving the efficiency of subsequent feature extraction, and at the same time reducing the calculation of subsequent scene recognition and improving the efficiency and accuracy of scene recognition.

[0142] After obtaining a number of second target spectra, the energy characteristic parameters within the second target spectrum are extracted, and then the energy characteristic value is calculated. The energy characteristic value can be the average energy within the second target spectrum, or the maximum energy within the second target spectrum, or the energy variance value within the second target spectrum, which can be selected according to the actual design requirements. Preferably, the energy characteristic value is the average energy within the second target spectrum. It is easy to understand that when the energy characteristic value is the average energy within the second target spectrum, it can better reflect the overall energy situation of the second target spectrum and effectively reduce the impact of energy mutations.

[0143] It can be understood that by setting the sensitive frequency range and the second truncation bandwidth, a second target spectrum that is more in line with the human ear's auditory characteristics can be obtained, and then an energy characteristic value that is more in line with the human ear's auditory characteristics and has a more obvious impact on the noise reduction effect can be obtained, reducing the calculation amount of feature extraction, improving the efficiency and accuracy of feature extraction, and further improving the efficiency and accuracy of subsequent scene recognition.

[0144] Further, in a preferred embodiment, step S34 of obtaining the similarity coefficient between the amplitude feature matrix and the comparison matrix of the preset scene type includes:

[0145] S341. Obtain the comparison matrix of the preset scene type;

[0146] S342. According to the cosine similarity formula, obtain a number of initial similarity coefficients between the amplitude feature matrix and the comparison matrix in the preset target dimension;

[0147] S343. Calculate the average value of the initial similarity coefficients to obtain the similarity coefficient between the amplitude feature matrix and the comparison matrix.

[0148] Wherein, the target dimension is the row dimension or the column dimension, and the type of the target dimension depends on the dimension where the time dimension is located in the first spectrum. Specifically, if the time dimension is the row dimension in the first spectrum, the target dimension is the column dimension; if the time dimension is the column dimension in the first spectrum, the target dimension is the row dimension.

[0149] After obtaining the amplitude feature matrix, the comparison matrix of the preset scene type can be obtained, and the similarity coefficients between the amplitude feature matrix and each comparison matrix can be calculated respectively. Taking the time dimension as the row dimension in the first spectrum and the target dimension as the column dimension as an example, specifically, according to the cosine similarity formula, the initial similarity coefficients between the amplitude feature matrix and the comparison matrix in the column dimension are calculated respectively, that is, the similarity degree between the corresponding column vectors of the amplitude feature matrix and the comparison matrix is calculated. After obtaining the initial similarity coefficients of each column vector, the average value of the initial similarity coefficients is calculated respectively, and the similarity coefficient between the amplitude feature matrix and the comparison matrix can be obtained. According to the same calculation method, the similarity coefficients between the amplitude feature matrix and each comparison matrix can be obtained in turn.

[0150] It can be understood that by calculating the initial similarity coefficients between the amplitude feature matrix and the comparison matrix, and using the average value of the initial similarity coefficients as the similarity coefficient between the amplitude feature matrix and the comparison matrix, the similarity degree between the amplitude feature matrix and the comparison matrix can be relatively better reflected, and the efficiency and accuracy of the subsequent preliminary screening of scene recognition can be improved.

[0151] Further, in a preferred embodiment, step S35 of screening out a specific scene type from the scene types based on the similarity coefficient includes:

[0152] S351. Obtain the maximum value in the similarity coefficients as the maximum similarity coefficient;

[0153] S352. Obtain the corresponding comparison value according to the maximum similarity coefficient;

[0154] S353. Determine one by one whether the similarity coefficient is less than the comparison value;

[0155] S354. If not, mark the similarity coefficient as the target similarity coefficient.

[0156] S353. According to the target similarity coefficient, screen out a specific scenario type from the scenario types.

[0157] Among them, the comparison value is a parameter used to compare with the similarity coefficient to screen out the target similarity coefficient that meets the requirements.

[0158] Specifically, after obtaining the similarity coefficients, the similarity coefficients can be compared with each other to determine the maximum value among the similarity coefficients, and then the maximum similarity coefficient is obtained. After obtaining the maximum similarity coefficient, the maximum similarity coefficient is compared with a number of preset comparison intervals. Each comparison interval is preset with a corresponding comparison value, and the larger the endpoint value of the comparison interval, the relatively larger the corresponding comparison value. After determining the comparison interval in which the maximum similarity coefficient falls, the comparison value corresponding to this comparison interval can be obtained, and then the similarity coefficient is compared with the obtained comparison value, and then the target similarity coefficient that meets the requirements can be determined.

[0159] It is easy to understand that when the value of the maximum similarity coefficient is relatively large, the value of the obtained comparison value is also relatively large, which can reduce the number of subsequent obtained target similarity coefficients to a certain extent, and then reduce the number of specific scenario types, improving the efficiency of subsequent scenario recognition. When the value of the maximum similarity coefficient is relatively small, the value of the obtained comparison value is also relatively small, and then the number of obtained target similarity coefficients can be increased to a certain extent, and then the number of specific scenario types is increased to improve the accuracy of subsequent scenario recognition.

[0160] It can be understood that by obtaining the corresponding comparison value according to the maximum similarity coefficient, the number of obtained target similarity coefficients can be dynamically adjusted, and then the number of subsequent obtained specific scenario types can be controlled according to the value of the maximum similarity coefficient, so as to improve the efficiency and accuracy of subsequent scenario type recognition to a certain extent.

[0161] Furthermore, in one embodiment, the steps S36 of determining the target scenario type according to the energy eigenvalue, the similarity coefficient and the specific scenario type include:

[0162] S361. Obtain the preset energy interval of the specific scenario type.

[0163] S362. Determine in turn whether the energy eigenvalue falls into the corresponding energy interval.

[0164] S363. Obtain the number of times the energy eigenvalue falls into the energy intervals of each specific scenario type.

[0165] S364. Determine the target scene type from specific scene types based on the falling-in number and the similarity coefficient.

[0166] Among them, the energy interval is the energy fluctuation interval of the corresponding frequency band in the scene type, which represents the acceptable range of the energy fluctuation of the scene type in the corresponding frequency band. The falling-in number is the number of energy eigenvalues falling into the corresponding energy interval.

[0167] Each scene type is preset with an energy interval for comparison. After obtaining a number of energy eigenvalues and a specific scene type, the energy interval corresponding to the specific scene type can be obtained, and then the obtained energy eigenvalues are compared item by item with the energy intervals corresponding to each specific scene type, so as to obtain the falling-in number of the energy eigenvalues falling into the corresponding energy interval.

[0168] Since the energy eigenvalues are extracted based on specific frequency bands to which the human ear is relatively sensitive, there may be a situation where the falling-in numbers of similar scene types are the same. Therefore, the falling-in number and the similarity coefficient can be combined to determine the target scene type from specific scene types. Specifically, the method for determining the target scene type can be to calculate the weights of the falling-in number and the similarity coefficient, and then determine the target scene type from specific scene types, or to use the falling-in number and the similarity coefficient for conditional screening, and then determine the target scene type from specific scene types, or to combine the above two methods to determine the target scene type from specific scene types.

[0169] It can be understood that by obtaining the falling-in number of the energy eigenvalues falling into the energy interval of a specific scene type, the matching degree between the current scene where the earphone is located and the specific scene type in the sensitive frequency band can be obtained relatively quickly. Then, by combining the falling-in number and the similarity coefficient, the specific scene type can be re-screened from the local level and the overall level, so as to obtain the target scene type, optimizing the method for obtaining the target scene type and improving the efficiency and accuracy of the re-screening of the scene type.

[0170] In a preferred embodiment, step S364 of determining the target scene type from specific scene types based on the falling-in number and the similarity coefficient includes:

[0171] S3641. Obtain the maximum falling-in number in the falling-in numbers and the maximum similarity coefficient in the similarity coefficients;

[0172] S3642. Determine whether the maximum falling-in number and the maximum similarity coefficient correspond to the same specific scene type;

[0173] S3643. If so, mark the specific scene type as the target scene type;

[0174] S3644. If not, obtain the reference value corresponding to the specific scenario type based on the number of falls, the similarity coefficient, and the preset weight coefficient;

[0175] S3645. Mark the specific scenario type corresponding to the maximum reference value in the reference values as the target scenario type.

[0176] Among them, the reference value is the adaptation degree value of the specific scenario type and the current scenario type of the earphone calculated based on the number of falls and the similarity coefficient under the weight coefficient, and can be used to screen out the target scenario type from the specific scenario types.

[0177] After obtaining the number of falls and the similarity coefficient, compare the numbers of falls with each other to obtain the maximum number of falls, and compare the similarity coefficients with each other to obtain the maximum similarity coefficient. After obtaining the maximum number of falls and the maximum similarity coefficient, the specific scenario type corresponding to the maximum number of falls and the specific scenario type corresponding to the maximum similarity coefficient can be obtained, and then it is determined whether the maximum number of falls and the maximum similarity coefficient correspond to the same specific scenario type.

[0178] If the maximum number of falls and the maximum similarity coefficient both correspond to the same specific scenario type, mark this specific scenario type as the target scenario type. It is easy to understand that when the maximum number of falls and the maximum similarity coefficient both correspond to the same specific scenario type, it means that the matching degree of this specific scenario type is the highest at both the local level and the overall level. Therefore, this specific scenario type can be marked as the target scenario type.

[0179] If the maximum number of falls and the maximum similarity coefficient do not correspond to the same specific scenario type, it is necessary to determine the reference value of the specific scenario type based on the weight coefficient, comprehensively considering the number of falls and the similarity coefficient. Specifically, the way to obtain the weight coefficient can be to directly obtain the preset weight coefficient, or to obtain the corresponding weight coefficient by combining the maximum number of falls and the maximum similarity coefficient. Preferably, the method of combining the maximum number of falls and the maximum similarity coefficient to obtain the corresponding weight coefficient is adopted to ensure the accuracy of the obtained reference value.

[0180] After determining the corresponding weight coefficient, the reference value corresponding to the specific scenario type can be calculated according to the number of falls, the similarity coefficient, and the weight coefficient. After obtaining the reference value of the specific scenario type, screen out the maximum reference value from the reference values, and mark the specific scenario type corresponding to the maximum reference value as the target scenario type.

[0181] It can be understood that by first screening based on the correspondence between the maximum number of falling-in and the maximum similarity coefficient, and then screening according to the specific situation of the reference value, it is possible to relatively well re-screen the specific scenario type from the local level and the overall level, thereby obtaining the target scenario type, optimizing the acquisition method of the target scenario type, and improving the efficiency and accuracy of the scenario type re-screening.

[0182] Further, please refer to Figure 3 , in one embodiment, the step S6 of determining the target scenario type according to the position information and the environmental audio signal includes:

[0183] S61. Obtain key position information from the position information;

[0184] S62. Obtain the environmental audio spectrum from the environmental audio signal;

[0185] S63. Based on the key position information, determine the initial scenario type and the associated scenario type;

[0186] S64. According to the environmental audio spectrum and the preset extraction mode, obtain a number of energy feature values;

[0187] S65. According to the energy feature values, determine the target scenario type from the initial scenario type and the associated scenario type.

[0188] Among them, the key position information is the marker parameter information used to reflect the position feature situation of the current scenario where the earphone is located. The initial scenario type is the scenario type obtained through the key position information from the preset scenario types for subsequent identification and determination. The associated scenario type is the scenario type that has a certain degree of similarity with the sound features of the determined initial scenario type.

[0189] There is a geographical information database for initially determining the scenario type according to the key position information pre-set in the earphone system or the APP supporting the earphone. By comparing the extracted key position information, the corresponding scenario type can be confirmed. Similarly to the acquisition method of the position information, the extraction method of the key position information can be extracted by the earphone or by the terminal device connected to the earphone.

[0190] After extracting the key position information from the position information, the initial scenario type and the associated scenario type can be determined respectively through the key position information, that is, taking the adaptation degree of the key position information as the division basis, and then obtaining the initial scenario type and the associated scenario type respectively; or the initial scenario type can be determined first through the key position information, and then the associated scenario type can be determined according to the association degree between the initial scenario type and the other remaining scenario types.

[0191] After obtaining the environmental audio spectrum, feature extraction can be performed on the environmental audio spectrum based on a preset extraction mode. This extraction mode can be to extract the overall frequency band of the environmental audio spectrum, or to extract a specific frequency band of the environmental audio signal.

[0192] Preferably, the extraction mode is to extract a specific frequency band of the environmental audio signal. It can be understood that since the human ear has different sensitivities to different frequencies, and both the initial scene type and the associated scene type are scene types obtained through preliminary screening, therefore, by adopting the method of extracting a specific frequency band of the environmental audio spectrum, the characteristics of the human ear can be combined, and the feature information of the frequency band to which the human ear is relatively sensitive can be extracted more emphatically. While improving the accuracy of feature extraction, the amount of feature information to be extracted is reduced, thereby improving the efficiency and accuracy of screening using feature information in subsequent scene recognition.

[0193] Based on the preset extraction mode, several required intercepted spectra can be intercepted from the environmental audio spectrum, and then the corresponding energy feature values can be extracted from the obtained several intercepted spectra. The energy feature value can be the average value of the energy values of each frequency on the intercepted spectrum, or the maximum value of the energy values of each frequency on the intercepted spectrum, or the variance value of the energy values of each frequency on the intercepted spectrum.

[0194] After obtaining the energy feature value, according to the energy feature value, the way to determine the target scene type from the initial scene type and the associated scene type can be to obtain the energy intervals of the initial scene type and the associated scene type, then compare the energy feature value with the corresponding energy intervals, and then screen the target scene type according to the number of times the energy feature value falls within the energy intervals; it can also be based on a preset weight formula of the energy feature value, calculate the energy integration value, and then compare it with a preset comparison interval to determine the target scene type. After determining the target scene type, the noise reduction parameters corresponding to the target scene type can be obtained, thereby realizing the adaptive noise reduction function of the earphone.

[0195] It can be understood that by using the key position information to obtain the initial scene type and the associated scene type, the scene type can be preliminarily screened relatively quickly, reducing the number of scene types that need to be recognized. Then, combined with several energy feature values, the initial scene type and the associated scene type can be further re-screened more emphatically to obtain the target scene type, thereby optimizing the way to obtain the target noise reduction parameters, reducing the calculation amount of scene recognition, improving the efficiency and accuracy of adaptive noise reduction, reducing the power consumption of the earphone to a certain extent, and improving the user experience.

[0196] Further, in an embodiment, the step S63 of determining the initial scene type and the associated scene type based on the key position information includes:

[0197] S631. Obtain the adaptation value of the key position information and the preset scene type;

[0198] S632. Determine the initial scene type from the preset scene types based on the adaptation value;

[0199] S633. Obtain the association value between the initial scene type and the remaining scene types;

[0200] S634. Obtain the associated scene type from the remaining scene types based on the association value and the preset first threshold.

[0201] Wherein, the adaptation value is the adaptation degree value of the key position information and each scene type, which can be used to quantify the adaptation degree between each scene type and the key position information. The association value reflects the similarity degree of the characteristics of each scene type in the spectrum. This association value is a parameter generated by pre-training and analyzing each scene type, and each scene type generates a corresponding association value for the remaining scene types. The first threshold is a numerical value used for screening the association value, which can screen the association value to obtain the associated scene type that meets the requirements.

[0202] Specifically, when obtaining the adaptation value of the key position information and the scene type, weight relationships can be assigned to specific parameters in the key position information to determine the adaptation degree value of the key position information and each scene type, and then determine the initial scene type that meets the requirements. It is easy to understand that the number of initial scene types that can be determined through the key position information can be one or several, and can be set according to the actual design requirements to control the number of obtained initial scene types.

[0203] It can be understood that by using the method of screening the initial scene type with the adaptation value, the initial scene type can be screened in a quantitative manner, which can better control the number of obtained initial scene types. At the same time, using the adaptation value method can also better control the number of obtained subsequent associated scene types.

[0204] After obtaining the adaptation value, the method of determining the initial scene type from the scene types according to the adaptation value can be to set a comparison value for screening, and then screen out the initial scene type that meets the requirements; it can also be to screen out the scene type with the highest adaptation degree in the adaptation value as the initial scene type, which can be selected according to the actual design requirements. Preferably, the scene type corresponding to the maximum value in the adaptation value is marked as the initial scene type.

[0205] Further, in a preferred embodiment, the step S634 of obtaining the associated scene type from the remaining scene types based on the association value and the preset first threshold includes:

[0206] S6341. Obtain the maximum value among the adaptation values as the maximum adaptation value;

[0207] S6342. Obtain the preset adaptation value interval in which the maximum adaptation value falls and use it as the target adaptation value interval;

[0208] S6343. Obtain the preset first threshold corresponding to the target adaptation value interval;

[0209] S6344. Determine whether the associated value is less than the first threshold;

[0210] S6345. If not, mark the scenario type corresponding to the associated value as the associated scenario type.

[0211] Among them, the adaptation value interval is a pre-set comparison interval for determining the first threshold corresponding to the maximum adaptation value. Each adaptation value interval corresponds to a first threshold. The larger the endpoint of the adaptation value interval, the larger the corresponding first threshold set. The number and interval size of the adaptation value intervals can be set according to the actual design requirements.

[0212] Specifically, after obtaining the maximum adaptation value, it is possible to first determine the adaptation value interval in which the maximum adaptation value falls, and then obtain the corresponding first threshold according to the fallen adaptation value interval. When the maximum adaptation value is relatively high, a relatively large first threshold can be obtained. When the maximum adaptation value is relatively low, a relatively small first threshold can be obtained.

[0213] It is easy to understand that obtaining the corresponding first threshold according to the situation of the maximum adaptation value can realize dynamically adjusting the difficulty of obtaining the associated scenario type according to the adaptation degree between the key position information and the scenario type, and then realize controlling the number of obtained associated scenario types according to the situation of the adaptation value. That is, when the maximum adaptation value is relatively high, the size of the first threshold can be appropriately increased, which can effectively reduce the acquisition amount of the associated scenario model and improve the subsequent scenario recognition efficiency; when the maximum adaptation value is relatively low, the size of the first threshold can be appropriately reduced, and relatively more associated scenario models can be obtained, thereby ensuring the accuracy of subsequent scenario recognition.

[0214] Further, in one embodiment, the step S64 of obtaining a number of energy feature values according to the environmental audio spectrum and the preset extraction mode includes:

[0215] S641. Obtain the preset sensitive interception frequency band and the preset interception frequency width;

[0216] S642. Obtain a number of target spectra according to the sensitive interception frequency band, the interception frequency width and the environmental audio spectrum;

[0217] S643. Calculate the energy characteristic parameters of the target spectrum respectively to obtain the corresponding energy characteristic values.

[0218] Among them, the sensitive intercept band is a specific frequency range for feature extraction preset according to the human ear auditory characteristics. Preferably, the sensitive intercept band is from 1000 Hz to 4000 Hz. The intercept bandwidth is an intercept unit that further intercepts the spectrum within the sensitive intercept band from the frequency dimension. The target spectrum is the spectrum obtained after the environmental audio spectrum is intercepted twice and is the spectrum calculation unit of the energy characteristic value.

[0219] After obtaining the environmental audio spectrum, the corresponding sensitive intercept band and intercept bandwidth are acquired, and then the environmental audio spectrum is intercepted according to the obtained sensitive intercept band and intercept bandwidth. Specifically, the environmental audio spectrum can be processed according to the sensitive intercept band first to obtain the sensitive spectrum, and then the sensitive spectrum is processed using the intercept bandwidth to further obtain a number of target spectra. Or the environmental audio spectrum can be processed according to the intercept bandwidth first to obtain a number of sub-spectra, and then the sub-spectra within the sensitive intercept band are obtained from the number of sub-spectra as the target spectrum.

[0220] After obtaining a number of target spectra, the energy characteristic parameters within the target spectra are extracted, and then the energy characteristic values are calculated. The energy characteristic value can be the average energy within the target spectrum, or the maximum value of the energy within the target spectrum, or the energy variance value within the target spectrum, and can be selected according to the actual design requirements. Preferably, the energy characteristic value is the average energy of the target spectrum. It is easy to understand that the energy characteristic value being the average energy within the target spectrum can better reflect the overall energy situation of the target spectrum and can effectively reduce the influence of energy mutations.

[0221] It can be understood that by setting the sensitive intercept band and intercept bandwidth, a target spectrum more in line with the human ear auditory characteristics can be obtained, and then an energy characteristic value more in line with the human ear auditory characteristics and having a more obvious influence on the noise reduction effect can be extracted, reducing the computational amount of feature extraction, improving the efficiency and accuracy of feature extraction, and further improving the efficiency and accuracy of subsequent scene recognition.

[0222] Further, in one embodiment, the step S65 of determining the target scene type from the initial scene type and the associated scene type according to the energy characteristic value includes:

[0223] S651. Obtain the energy interval of the initial scene type and the energy interval of the associated scene type;

[0224] S652. Obtain the number of times the energy characteristic value falls within the energy interval of the initial scene type and the energy interval of the associated scene type;

[0225] S653. Obtain the target scene type based on the falling-in number.

[0226] Among them, the energy interval is the energy fluctuation interval of the corresponding frequency band in the scene type, which represents the acceptable range of the energy fluctuation of the scene type in the corresponding frequency band. The falling-in number is the number of energy eigenvalues falling into the corresponding energy interval.

[0227] After obtaining the energy eigenvalues, obtain the energy intervals of the initial scene type and the associated scene type, and then compare the energy eigenvalues item by item with the energy intervals corresponding to the initial scene type and the associated scene type, so as to obtain the falling-in number of the energy eigenvalues falling into the corresponding energy interval. After obtaining the falling-in numbers corresponding to the initial scene type and the associated scene type, the target scene type can be screened according to the specific situation of the falling-in numbers. Specifically, the scene type corresponding to the largest falling-in number can be used as the target scene type by comparing the falling-in numbers; or the corresponding situation can be divided according to the number of the largest falling-in numbers, and then the target scene type can be obtained; or the key position information and the adaptability of the scene type can be combined for comprehensive screening, and then the target scene type can be obtained.

[0228] In a preferred embodiment, step S653 of obtaining the target scene type based on the falling-in number includes:

[0229] S6531. Obtain the value of the number of the largest falling-in numbers in the falling-in numbers;

[0230] S6532. If the value is 1, mark the scene type corresponding to the largest falling-in number as the target scene type;

[0231] S6533. If the value is greater than 1, obtain the adaptation value between the scene type corresponding to the largest falling-in number and the key position information;

[0232] S6534. Mark the scene type corresponding to the maximum value in the adaptation values as the target scene type.

[0233] After obtaining the falling-in numbers, compare the falling-in numbers with each other to obtain the largest falling-in number, and then the value of the number of the largest falling-in numbers can be obtained. When the value of the number of the largest falling-in numbers is only one, it means that among the initial scene type and the associated scene type, there is a scene type with a relatively prominent degree of compliance of the energy eigenvalues in the human ear highly sensitive frequency band. Therefore, the scene type corresponding to the largest falling-in number can be directly used as the target scene type. When the value of the number of the largest falling-in numbers is greater than one, the target scene type can be comprehensively determined by combining the adaptation values of the key position information. By comparing the sizes of the adaptation values between the scene types corresponding to the largest falling-in numbers, the maximum adaptation value is determined, and the scene type corresponding to the largest falling-in number and the maximum adaptation value is used as the target scene type.

[0234] It can be understood that by comparing the energy eigenvalue with the energy interval and based on the number of times the energy eigenvalue falls within the energy interval, the target scene type can be determined, thereby further re-screening the initial scene type and the associated scene type, optimizing the acquisition method of the target scene type, and improving the efficiency and accuracy of the scene type re-screening.

[0235] Further, please refer to Figure 4 , in one embodiment, based on the environmental audio signal, location information, attitude information, and an improved Transformer model configured with a proxy attention mechanism, the steps S8 for determining the target scene type include:

[0236] S81. Extract an audio feature sequence from the environmental audio signal;

[0237] S82. Extract an auxiliary feature sequence from the location information and the attitude information;

[0238] S83. Input the audio feature sequence and the auxiliary feature sequence into the improved Transformer model configured with a proxy attention mechanism to obtain the matching values of each scene type;

[0239] S84. Obtain the target scene type according to the magnitudes of the matching values.

[0240] Among them, the audio feature sequence is data that can be input into the improved Transformer model and represents the sound characteristics of the current scene where the earphone is located, which is obtained by extraction and transformation from the environmental audio signal. The auxiliary feature sequence is data that can be input into the improved Transformer model and represents the current state characteristics of the earphone, which is obtained by extracting feature information from the location information and the attitude information respectively and then transforming and fusing them.

[0241] After the acquisition microphone array provided in the earphone collects the environmental audio signal of the current scene where the earphone is located, it is necessary to preprocess the environmental audio signal to remove redundant or irrelevant interference information. After obtaining the preprocessed environmental audio signal, according to the actual design requirements, the obtained audio signal can be converted into a spectrogram of the required type, and then the obtained spectrogram is segmented into several sub-spectrograms according to the preset requirements, and audio features are extracted from the several sub-spectrograms. After obtaining the audio features, the extracted audio features are combined into an audio feature sequence.

[0242] After obtaining the position information of the current scene where the earphone is located and the current attitude information of the earphone, it is necessary to preprocess the position information and the attitude information respectively to remove redundant or irrelevant interference information. After completing the preprocessing of the position information and the attitude information, the required position features and attitude features are extracted from the position information and the attitude information respectively, and the obtained position features and attitude features are combined respectively to obtain a position feature sequence and an attitude feature sequence. Then, the two feature sequences are fused, and thus the required auxiliary feature sequence can be obtained.

[0243] After inputting the audio feature sequence and the auxiliary feature sequence into the improved Transformer model, the audio feature sequence and the auxiliary feature sequence are further processed to be transformed into an audio enhancement sequence and an auxiliary enhancement sequence with high dimensions and sorting information. The obtained audio enhancement sequence and auxiliary enhancement sequence then pass through the proxy attention mechanism of the improved Transformer model, and thus an enhanced feature matrix containing weight information is obtained. Then, through the transformation and calculation of the enhanced feature matrix by the fully connected layer, the matching values of each scene type can be obtained.

[0244] Specifically, in the process of calculating the proxy attention weight, the audio enhancement sequence will transform into Q, K, and V matrices with different weights and functions, and the auxiliary enhancement sequence will transform into a proxy matrix A, that is, generate the corresponding four-element proxy attention paradigm (Q, A, K, V). Then, the first softmax attention operation is performed using A, K, and V to obtain a fusion matrix V' that fuses the feature information in both the audio enhancement sequence and the auxiliary enhancement sequence. Then, the second softmax attention operation is performed using Q, A, and V' to obtain the required attention weight matrix. Then, the weighted sum of the attention weight matrix and V is used to obtain the enhanced feature matrix containing weight information.

[0245] Then, based on the obtained enhanced feature matrix and the fully connected layer, the matching values of each scene type can be determined. Specifically, after inputting the enhanced feature matrix into the fully connected layer, the fully connected layer will perform feature mapping on the enhanced feature matrix, that is, the fully connected layer will perform a linear transformation on each feature vector in the enhanced feature matrix and output a vector with the same number of elements as the number of scene types. Each element of this vector corresponds to the prediction score of a scene type. After completing the feature mapping and obtaining the prediction scores, the softmax activation function is used to transform the obtained prediction scores into a probability distribution between 0 and 1, and thus the matching values of each scene type are obtained. Then, the maximum value in the matching values is selected, and the scene type corresponding to this maximum matching value is the target scene type. Thus, the noise reduction parameters of the target scene type can be obtained to realize the adaptive noise reduction function of the earphone.

[0246] It can be understood that by improving the Transformer model and multi-source feature information, the function of recognizing and classifying the complex audio scenes where the earphones are located is realized. By using the proxy attention mechanism of the improved Transformer model, while maintaining a relatively low computational complexity, the integration and recognition of multi-source feature information are realized, the accuracy of the dynamic adjustment of the recognition attention weight is improved, thereby improving the efficiency and accuracy of the earphones in recognizing complex audio scenes, further improving the matching degree of the earphones' adaptive noise reduction, improving the effect of the earphones' adaptive noise reduction, and improving the user experience.

[0247] Further, in one embodiment, the step S81 of extracting the audio feature sequence from the environmental audio signal includes:

[0248] S811. Preprocess the environmental audio signal to obtain a denoised audio signal;

[0249] S812. Convert the denoised audio signal into an environmental audio spectrum;

[0250] S813. Divide the environmental audio spectrum into several target spectra;

[0251] S814. Sequentially extract audio features from the target spectra to obtain an audio feature sequence.

[0252] After the environmental audio signal of the current scene where the earphones are located is collected, since the environmental audio signal may be mixed with interference signals such as background noise, it is necessary to perform preprocessing such as filtering on the environmental audio signal to obtain a denoised audio signal.

[0253] After obtaining the denoised audio signal, in order to further extract audio features, it is necessary to convert the denoised audio signal into an environmental audio spectrum capable of extracting the required audio features. It is easy to understand that the type of the environmental audio spectrum can be selected according to the design requirements, and the environmental audio spectrum includes but is not limited to the environmental audio spectrum obtained by short-time Fourier transform, Mel spectrum, etc. Preferably, since the human ear has different sensitivities to different frequencies of sound, in order to extract audio features more in line with the auditory characteristics of the human ear, the denoised audio signal can be converted into a Mel spectrum.

[0254] After obtaining the environmental audio spectrum, it is necessary to divide the environmental audio spectrum into several target spectra to facilitate the subsequent extraction of audio features. Specifically, the division method of the environmental audio spectrum can be linear division or non-linear division. Preferably, the non-linear division method is adopted to divide the environmental audio spectrum, intercepting with a relatively denser width in the frequency band where the human ear is relatively more sensitive and with a relatively sparser width in the frequency band where the human ear is relatively less sensitive, so that the obtained target spectra are more in line with the auditory characteristics of the human ear in the overall distribution, thereby improving the accuracy of subsequent feature extraction.

[0255] After obtaining a number of target spectra, audio features are sequentially extracted from the target spectra, and the obtained audio features are sorted and combined to obtain an audio feature sequence. Specifically, a feature encoder such as a convolutional neural network or a recurrent neural network can be used to extract high-order features, or the spectral features, statistical features, etc. of the sub-spectra can be directly extracted to obtain audio features. Preferably, a feature encoder such as a convolutional neural network or a recurrent neural network can be used to extract high-order features as audio features.

[0256] Further, in one embodiment, the step S82 of extracting the auxiliary feature sequence from the position information and the attitude information includes:

[0257] S821. Preprocess the position information and the attitude information respectively to obtain denoised position information and denoised attitude information;

[0258] S822. Extract position features from the denoised position information to obtain a position feature sequence;

[0259] S823. Extract attitude features from the denoised attitude information to obtain an attitude feature sequence;

[0260] S824. Perform feature fusion on the position feature sequence and the attitude feature sequence to obtain an auxiliary feature sequence.

[0261] Among them, the position feature is the key information in the position information that can be used to represent the position characteristics of the current scene of the earphone to a certain extent. The attitude feature is the key information in the attitude feature that can be used to represent the current state of the earphone to a certain extent.

[0262] Specifically, after obtaining the current position information and attitude information of the earphone, since the position information and the attitude information may both contain interference information or repetitive and meaningless information, it is necessary to denoise the position information and the attitude information to obtain denoised position information and denoised attitude information.

[0263] After obtaining the denoised position information and denoised attitude information, feature extraction can be performed on the denoised position information and denoised attitude information according to preset requirements. Specifically, the position features include but are not limited to geographical location information, scene type information, point of interest information, etc.; the attitude feature information includes but is not limited to the wearing attitude information of the earphone and the acceleration information of the earphone, etc. It can be understood that the position features extracted from the denoised position information can be only single-category position features or several categories of position features; similarly, the attitude features can also be only single-category attitude features or several categories of attitude features, which can be selected according to actual design requirements.

[0264] After obtaining the position feature and the attitude feature, the position features can be combined to obtain a position feature sequence, and the attitude features can be combined to obtain an attitude feature sequence. After fusing the obtained position feature sequence and the attitude feature sequence, an auxiliary feature sequence can be obtained. Specifically, the methods of fusing the position feature sequence and the attitude feature sequence include, but are not limited to, direct splicing, weighted splicing, etc. between sequences, and can be selected according to the design requirements. Preferably, the weighted splicing method is adopted, which can better divide the weights between the position information and the attitude information according to the actual situation, and then improve the adaptability of the subsequent adjustment of the attention weight.

[0265] It can be understood that by obtaining the auxiliary feature sequence obtained by fusing the position feature sequence and the attitude feature sequence, the characteristic relationship between the current position of the earphone and the earphone attitude can be better reflected from the overall level, enriching the recognition information amount of the subsequent sound scene classification, assisting in adjusting the proxy attention weight, improving the accuracy of the proxy attention weight, and thus improving the accuracy of the sound scene classification.

[0266] Furthermore, in a preferred embodiment, the step S822 of extracting position features from the denoised position information to obtain a position feature sequence includes:

[0267] S8221. Extract the current geographical location information of the earphone from the denoised position information;

[0268] S8222. Determine the initial information of the scene type according to the geographical location information and the preset geographical information database;

[0269] S8223. Encode and transform the initial information of the scene type to obtain a position feature sequence.

[0270] Among them, the geographical location information includes the longitude and latitude information and altitude information of the current position of the earphone. The initial information of the scene type is the information of the possible scene types where the earphone is currently located initially determined by the geographical location information of the earphone.

[0271] After extracting the current longitude and latitude information and altitude information of the earphone from the denoised position information, through the geographical information database, the current scene type where the earphone is located can be initially determined by using the geographical location information of the earphone. It can be understood that after interacting the geographical location information with the geographical information database, the geographical information database can map the initial scene types that the current geographical location information of the earphone may conform to. Due to the complexity of the actual scene, the initial scene type information initially determined by the geographical location information may be a single scene type or multiple scene types. To ensure the comprehensiveness of the position feature sequence, all the initially determined eligible scene types should be obtained as the initial scene types.

[0272] Specifically, the interaction between the geographical location information and the geographical information database can be realized through a terminal connected to the earphone. The way to realize the interaction can be cloud online interaction, that is, the initial information of the scene type is determined by real-time data exchange between the terminal device and the cloud. The way to realize the interaction can also be offline interaction, that is, the geographical information database can be pre-downloaded into the APP of the earphone of the terminal. Preferably, the interaction method is offline interaction, which can better adapt to the actual use situation and improve the efficiency of obtaining the initial information of the scene type to a certain extent.

[0273] In a preferred embodiment, the user can, according to their own usage needs, pre-select and download the geographical information database of the required area in the earphone APP of the terminal. When the terminal receives a signal that needs to determine the initial information of the scene type, it can directly determine the initial information of the scene type from the already downloaded geographical information database, and then relatively quickly determine the initial information of the scene type corresponding to the current geographical location information. At the same time, after the earphone is connected to the terminal, it will automatically trigger the detection of whether the geographical information database needs to be updated. If so, it will be updated automatically to ensure the accuracy of the subsequent determined initial scene type.

[0274] After determining the initial information of the scene type, it is also necessary to perform encoding conversion on the obtained initial scene type to obtain a position feature sequence, so as to ensure that the position feature sequence can be input into the improved Transformer model and recognized and analyzed by the improved Transformer model.

[0275] It can be understood that the position feature sequence determined by the initial information of the scene type can be used to assist in adjusting the recognition weight of the proxy attention mechanism of the improved Transformer model according to the initial information of the scene type, improve the accuracy of subsequent calculation of the proxy attention weight, and thus improve the accuracy and efficiency of complex sound scene classification.

[0276] Furthermore, in a preferred embodiment, the step S823 of extracting the pose feature from the denoised pose information to obtain the pose feature sequence includes:

[0277] S8231. Extract the current acceleration feature of the earphone from the denoised pose information;

[0278] S8232. Perform encoding conversion on the acceleration feature to obtain the pose feature sequence.

[0279] Among them, the acceleration feature can be used to reflect the current motion state of the earphone, that is, through the acceleration feature, it can be determined whether the earphone is currently in a dynamic or static state as a whole, and then the current motion state of the user can be determined.

[0280] Specifically, acceleration data containing the acceleration components of the earphone in the three directions of the X-axis, Y-axis, and Z-axis are obtained from the denoised pose information, and then these acceleration data are subjected to feature extraction to obtain acceleration features. It can be understood that the acceleration features include, but are not limited to, peak acceleration, average acceleration, acceleration variance, etc., and can be selected and extracted according to the actual design requirements.

[0281] After obtaining the required acceleration features, the obtained acceleration features are encoded and transformed to obtain a pose feature sequence, so as to ensure that the pose feature sequence can be input into the improved Transformer model and recognized and analyzed by the improved Transformer model.

[0282] It can be understood that the pose feature sequence determined by the acceleration features can determine the current motion state of the user, and realize the auxiliary adjustment of the recognition weight of the proxy attention mechanism of the improved Transformer model according to the current motion state of the user, improve the accuracy of subsequent calculation of the proxy attention weight, and further improve the accuracy and efficiency of classifying complex sound scenes.

[0283] Further, in one embodiment, the step S83 of inputting the audio feature sequence and the auxiliary feature sequence into the improved Transformer model configured with the proxy attention mechanism to obtain the matching values of each scene type includes:

[0284] S831. Map and encode the audio feature sequence and the auxiliary feature sequence to obtain an audio enhancement sequence and an auxiliary enhancement sequence respectively;

[0285] S832. Based on the audio enhancement sequence, the auxiliary enhancement sequence and the preset proxy attention mechanism, obtain an enhanced feature matrix;

[0286] S833. Based on the preset fully connected layer and the enhanced feature matrix, obtain the matching values of each scene type.

[0287] Among them, the audio enhancement sequence and the auxiliary enhancement sequence are high-dimensional data with sorting information obtained after the transformation of the audio feature sequence and the auxiliary feature sequence respectively. The enhanced feature matrix is data with weight information obtained after calculation by the proxy attention mechanism. The matching value is the matching degree between the preset scene type and the current sound scene where the earphone is located. The larger the matching value, the higher the matching degree; otherwise, the matching degree is lower.

[0288] After inputting the audio feature sequence and the auxiliary feature sequence into the improved Transformer model, the audio feature sequence and the auxiliary feature sequence need to go through the embedding layer for linear mapping respectively, mapping each feature vector in the audio feature sequence and the auxiliary feature sequence into a higher-dimensional embedding space, and then obtaining the audio embedding vector sequence and the auxiliary embedding vector sequence respectively. After obtaining the audio embedding vector sequence and the auxiliary embedding vector sequence, since the improved Transformer model does not have built-in position information, it is necessary to add position encoding to indicate the sorting of the feature vectors in the feature sequence. After completing the position encoding respectively, the audio enhanced sequence and the auxiliary enhanced sequence can be obtained.

[0289] After inputting the audio enhanced sequence and the auxiliary enhanced sequence into the proxy attention mechanism in the improved Transformer model, the proxy attention mechanism then performs two softmax attention operations using the audio enhanced sequence and the auxiliary enhanced sequence respectively. From the perspective of multi-source feature information fusion at the overall level, an initial attention weight matrix is obtained, and then the initial attention matrix is normalized to obtain the attention weight matrix. Further, weighted summation is performed using the attention weight matrix to obtain an enhanced feature matrix containing the weight information of each feature vector.

[0290] Further, in a preferred embodiment, based on the audio enhanced sequence, the auxiliary enhanced sequence and the preset proxy attention mechanism, step S832 of obtaining the enhanced feature matrix includes:

[0291] S8321. According to the audio enhanced sequence and the preset conversion weight matrix, obtain the first matrix, the second matrix and the third matrix respectively;

[0292] S8322. According to the auxiliary enhanced sequence and the preset proxy conversion matrix, obtain the proxy matrix;

[0293] S8323. Perform interactive calculation on the proxy matrix, the second matrix and the third matrix to obtain the fusion matrix;

[0294] S8324. Perform interactive calculation on the first matrix, the proxy matrix and the fusion matrix to obtain the attention weight matrix;

[0295] S8325. Perform weighted summation on the attention weight matrix and the third matrix to obtain the enhanced feature matrix.

[0296] Among them, the conversion weight matrix is obtained through training in advance for converting the audio enhanced sequence into data available for the proxy attention mechanism. The conversion weight matrix includes the first conversion matrix W A corresponding to the first matrix, the second matrix and the third matrix respectively, the second conversion matrix W K and the third conversion matrix WV The proxy transformation matrix is data obtained in advance through training for converting an auxiliary enhancement sequence into a proxy matrix applied to the proxy attention mechanism.

[0297] Specifically, after the audio enhancement sequence is input into the proxy attention mechanism, the first transformation matrix W A , the second transformation matrix W K and the third transformation matrix W V are respectively subjected to transformation calculations with the audio enhancement sequence to obtain the first matrix, the second matrix, and the third matrix. After the auxiliary enhancement sequence is input into the proxy attention mechanism, the auxiliary enhancement sequence is calculated with the proxy transformation matrix to obtain the proxy matrix. The number of parameters of the proxy matrix is much less than that of the first matrix.

[0298] After the conversion of the audio enhancement sequence and the auxiliary enhancement sequence is completed, the first softmax attention operation is performed. The proxy matrix is used to replace the first matrix as the query matrix, the second matrix as the key matrix, and the third matrix as the value matrix to realize the interaction between the proxy matrix, the second matrix, and the third matrix, so as to aggregate the information in the audio enhancement sequence, and then fuse the aggregated information with the information in the auxiliary enhancement sequence. Specifically, each proxy vector in the proxy matrix is multiplied by each key vector in the second matrix to obtain the first attention score that can reflect the strength of the correlation between different positions in the sequence, and then the first attention score is normalized to obtain the first attention weight. Then, the first attention weight is weighted and summed with the value vectors in the third matrix to obtain the fusion matrix.

[0299] It can be understood that by using the proxy matrix with fewer parameters to replace the first matrix as the query matrix, the direct interaction calculation between the first matrix and the second matrix can be avoided, so as to reduce the computational complexity. At the same time, a multi-source feature information fusion mechanism is established to integrate audio feature information, position feature information, and pose feature information, preparing for the recognition and classification of subsequent complex sound scenes.

[0300] After obtaining the information output by the first attention operation, the proxy matrix broadcasts the obtained information back to the first matrix. Using the first matrix as the query matrix, the proxy matrix as the key matrix, and the fusion matrix as the value matrix, a second softmax attention operation is performed to achieve the interaction between the first matrix, the proxy matrix, and the fusion matrix. Similarly to the first softmax attention operation, first calculate the second attention score using the first matrix and the proxy matrix, then normalize the second attention score to obtain the second attention weight, and then perform weighted summation using the second attention weight and the fusion matrix to obtain the attention weight matrix. After obtaining the attention weight matrix, perform weighted summation using the attention weight matrix and the third matrix to obtain the enhanced feature matrix containing the weight relationship of each feature vector.

[0301] It can be understood that the proxy attention mechanism can effectively process and integrate multi-source data information in a complex acoustic environment, achieving a good balance between computational complexity and classification accuracy, and significantly improving the accuracy and efficiency in identifying complex audio scenes.

[0302] In summary, the present application divides the mode of obtaining noise reduction parameters according to different parameter information through the current position information and remaining battery power of the earphone. When the position information is not obtained, the environmental audio signal is directly used for scene recognition to quickly determine the scene type; after obtaining the position information, if the battery power does not meet the requirements, the environmental audio signal and the position information are combined to achieve a quick preliminary screening and repeated screening of the scene type; if the battery power meets the requirements, an improved Transformer model and multi-source feature information are used to identify and classify complex audio scene types while maintaining low computational complexity; the noise reduction power consumption of the earphone is reasonably configured, the acquisition method of noise reduction parameters is optimized, the efficiency and accuracy of adaptive noise reduction are improved, the adaptability of the earphone to diverse user usage requirements and usage scenarios is improved, and the user experience is enhanced.

[0303] The following is the content of the second aspect of the present invention:

[0304] The present invention provides an earphone, as Figure 5 shown. The earphone includes a memory 10, a processor 20, and program instructions 30 for controlling earphone noise reduction stored on the memory 10 and executable on the processor 20. When the program instructions 30 for controlling earphone noise reduction are executed by the processor 20, the foregoing method for controlling earphone noise reduction is implemented.

[0305] In some embodiments, the processor may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips. The processor is generally used to control the overall operation of the earphone. In this embodiment, the processor is used to run the program code stored in the readable storage medium or process data.

[0306] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a readable storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which may be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods of the various embodiments of the present invention.

[0307] The following is the content of the third aspect of the present invention:

[0308] The present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-described earphone noise reduction control method are implemented.

[0309] The above are only the specific embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and retouches can be made, and these improvements and retouches should also be regarded as the protection scope of the present application.

Claims

1. A control method for headphone noise reduction, characterized in that, Including: Obtain the environmental audio signal of the current scene where the earphone is located; Detect whether the current position information of the earphone can be obtained; If not, determine the target scene type according to the environmental audio signal; If so, obtain the current position information of the earphone and the remaining power parameter; Detect whether the remaining power parameter is greater than a preset power threshold; If not, determine the target scene type according to the position information and the environmental audio signal; If so, obtain the current attitude information of the earphone; Based on the environmental audio signal, position information, attitude information and a preset improved Transformer model configured with a proxy attention mechanism, determine the target scene type; Obtain the noise reduction parameter corresponding to the target scene type as the target noise reduction parameter; The step of determining the target scene type based on the environmental audio signal, position information, attitude information and a preset improved Transformer model configured with a proxy attention mechanism includes: Extract an audio feature sequence from the environmental audio signal; Extract an auxiliary feature sequence from the position information and the attitude information; Input the audio feature sequence and the auxiliary feature sequence into an improved Transformer model configured with a proxy attention mechanism to obtain the matching values of each scene type; Obtain the target scene type according to the magnitudes of the matching values; The step of extracting the auxiliary feature sequence from the position information and the attitude information includes: Preprocess the position information and the attitude information respectively to obtain denoised position information and denoised attitude information; Extract position features from the denoised position information to obtain a position feature sequence; Extract attitude features from the denoised attitude information to obtain an attitude feature sequence; Fuse the position feature sequence and the attitude feature sequence to obtain an auxiliary feature sequence.

2. The control method for noise reduction of an earphone according to claim 1, characterized in that, The step of determining the target scene type according to the environmental audio signal includes: Obtain a first frequency spectrum and a second frequency spectrum according to the environmental audio signal; Based on the first frequency spectrum and a preset first extraction mode, obtain an amplitude feature matrix; Based on the second frequency spectrum and a preset second extraction mode, obtain a number of energy feature values; Obtain the similarity coefficient between the amplitude feature matrix and a comparison matrix of a preset scene type; Based on the similarity coefficient, screen out a specific scene type from the scene types; Determine the target scene type according to the energy feature values, the similarity coefficient and the specific scene type.

3. The control method for noise reduction of an earphone according to claim 2, wherein The step of obtaining the amplitude feature matrix based on the first frequency spectrum and a preset first extraction mode includes: Obtain a preset interception time width, a first interception frequency width and an initial matrix; Based on the first frequency spectrum, the interception time width and the first interception frequency width, obtain a number of first target frequency spectra; Sequentially associate the first target frequency spectra with the elements of the initial matrix one by one; Calculate the amplitude feature parameters of the first target frequency spectra respectively; According to the magnitudes of the amplitude feature parameters, adjust the positions of the corresponding elements on the same preset target dimension in the initial matrix to obtain an amplitude feature matrix.

4. The control method for headphone noise reduction according to claim 2, wherein, The step of determining the target scene type according to the energy feature values, the similarity coefficient and the specific scene type includes: Obtain the preset energy range for the specific scenario type; Determine in sequence whether the energy eigenvalue falls into the corresponding energy range; Obtain the number of falls of the energy eigenvalue into the energy ranges of each specific scenario type; Based on the number of falls and the similarity coefficient, determine the target scenario type from the specific scenario types.

5. A control method for noise reduction of an earphone according to claim 1, characterized in that, The step of determining the target scenario type according to the position information and the environmental audio signal includes: Obtain the key position information from the position information; Obtain the environmental audio spectrum from the environmental audio signal; Based on the key position information, determine the initial scenario type and the associated scenario type; According to the environmental audio spectrum and the preset extraction mode, obtain a number of energy eigenvalues; According to the energy eigenvalues, determine the target scenario type from the initial scenario type and the associated scenario type.

6. The control method for noise reduction of an earphone according to claim 5, characterized in that, The step of determining the initial scenario type and the associated scenario type based on the key position information includes: Obtain the adaptation value between the key position information and the preset scenario type; Based on the adaptation value, determine the initial scenario type from the preset scenario types; Obtain the association value between the initial scenario type and the remaining scenario types; Based on the association value and the preset first threshold, obtain the associated scenario type from the remaining scenario types.

7. A control method for headphone noise reduction according to claim 6, characterized in that, The step of determining the target scenario type from the initial scenario type and the associated scenario type according to the energy eigenvalues includes: Obtain the energy range of the initial scenario type and the energy range of the associated scenario type; Obtain the number of falls of the energy eigenvalue into the energy range of the initial scenario type and the energy range of the associated scenario type; Based on the number of falls, obtain the target scenario type.

8. A control method for headphone noise reduction according to claim 1, characterized in that, The step of inputting the audio feature sequence and the auxiliary feature sequence into the improved Transformer model configured with the proxy attention mechanism to obtain the matching values of each scenario type includes: Perform mapping and encoding on the audio feature sequence and the auxiliary feature sequence to obtain an audio enhancement sequence and an auxiliary enhancement sequence respectively; Based on the audio enhancement sequence, the auxiliary enhancement sequence and the preset proxy attention mechanism, obtain an enhanced feature matrix; Based on the preset fully connected layer and the enhanced feature matrix, obtain the matching values of each scenario type.

9. The control method for noise reduction of an earphone according to claim 8, wherein, The step of obtaining the enhanced feature matrix based on the audio enhancement sequence, the auxiliary enhancement sequence and the preset proxy attention mechanism includes: According to the audio enhancement sequence and the preset conversion weight matrix, obtain the first matrix, the second matrix and the third matrix respectively; According to the auxiliary enhancement sequence and the preset proxy conversion matrix, obtain the proxy matrix; Perform interactive calculation on the proxy matrix, the second matrix and the third matrix to obtain a fusion matrix; Perform interactive calculation on the first matrix, the proxy matrix and the fusion matrix to obtain an attention weight matrix; Perform weighted summation on the attention weight matrix and the third matrix to obtain the enhanced feature matrix.

Citation Information

Patent Citations

  • Earphone noise reduction method and device, storage medium and electronic equipment

    CN112423175A

  • Active noise reduction method and device, chip, earphone and storage medium

    CN115474121A