Adaptive noise reduction method for earphone, earphone and storage medium
By obtaining position, time and ambient audio information in the headset, determining the scene type and calculating matching values to obtain noise reduction parameters, the problems of low scene recognition efficiency, poor accuracy and high power consumption in the adaptive noise reduction technology of the existing headset are solved, achieving more efficient and more accurate noise reduction effects and lower power consumption.
Patent Information
- Application Number
- CN202411987405.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-13
AI Technical Summary
The existing headphone adaptive noise reduction technology has shortcomings in scene recognition efficiency and accuracy, and it consumes a high power, which affects the user experience.
By obtaining the position information, time information and environmental audio signals of the headphones, determining the initial scene type and associated scene type based on the key position information, filtering and determining the scene model with time information, and calculating the matching value using the scene recognition model and recognition weight data, and then obtaining noise reduction parameters.
It improves the efficiency and accuracy of adaptive noise reduction of headphones, reduces power consumption, and improves user experience.
Smart Images

Figure CN119996887A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of headphone noise reduction, and in particular to a headphone adaptive noise reduction method, a headphone and a storage medium. Background Art
[0002] With the rapid development of electronic technology, wireless headphones have become an indispensable item for people's daily travel. When facing various relatively noisy environments, people usually choose to wear wireless headphones to reduce the impact of environmental noise.
[0003] Different noise environments have different requirements for noise reduction. To better cope with the changing noise environment, some existing wireless headphones use adaptive noise reduction to reduce the impact of environmental noise. However, the existing adaptive noise reduction method cannot identify the noise scene relatively accurately and efficiently, which leads to a low matching degree of noise reduction requirements, and the existing adaptive noise reduction has high power consumption, which in turn affects the user experience.
[0004] In view of this, it is necessary to provide a method for adaptive noise reduction of headphones, headphones and storage media to solve the above problems. Summary of the invention
[0005] In view of the deficiencies in the prior art, the present invention provides a method for adaptive noise reduction of headphones, headphones and a storage medium, which aim to solve the technical problems of low scene recognition efficiency and accuracy and excessive power consumption of adaptive noise reduction of headphones.
[0006] To achieve the above-mentioned purpose, a first aspect of the present invention provides a method for adaptive noise reduction of headphones, comprising:
[0007] Get the current headset's location information, time information, and ambient audio signal;
[0008] Determining whether the location information contains key location information;
[0009] If yes, then determining the initial scene type and the associated scene type based on the key location information;
[0010] Based on the initial scene type, the associated scene type and the time information, a judgment scene model is obtained;
[0011] Obtaining a matching value based on the ambient audio signal, the determination scene model, the preset recognition weight data and the scene recognition model;
[0012] Based on the matching value, the noise reduction parameter is obtained.
[0013] In a preferred embodiment, the step of determining the initial scene type and the associated scene type based on the key location information further includes:
[0014] Determine an initial scene type from preset scene types according to the key position information;
[0015] Get the associated values of the initial scene type and other scene types;
[0016] Based on the comparison between the association value and a preset first threshold, the associated scene type is screened out.
[0017] In a preferred embodiment, the step of determining the initial scene type from preset scene types according to the key position information includes:
[0018] Obtaining the adaptation value of key location information and preset scene type;
[0019] From the adaptation values, obtain the maximum adaptation value;
[0020] According to the maximum adaptation value, the corresponding scene type is obtained as the initial scene type.
[0021] In a preferred embodiment, based on the comparison between the association value and the preset first threshold, the step of screening out the associated scene type includes:
[0022] Obtain the maximum adaptation value between key location information and preset scene type;
[0023] Determine a preset adaptation value interval within which the maximum adaptation value falls, and obtain a first threshold corresponding to the adaptation value interval;
[0024] Determining whether the correlation value is less than a first threshold;
[0025] If not, the scene type corresponding to the associated value is marked as the associated scene type.
[0026] In a preferred embodiment, the step of obtaining the determination scene model based on the initial scene type, the associated scene type and the time information includes:
[0027] Based on the initial scene type and the associated scene type, a specific scene model is obtained;
[0028] According to the time information, the judgment scene model is obtained by filtering from the specific scene models.
[0029] In a preferred embodiment, the step of obtaining a matching value based on the ambient audio signal, the determination scene model, the preset recognition weight data and the scene recognition model includes:
[0030] Perform feature extraction processing on the ambient audio signal to obtain the ambient audio spectrum;
[0031] Based on the preset scene recognition module and recognition weight data, the ambient audio spectrum is compared with the judgment scene model to obtain the corresponding matching value.
[0032] In a preferred embodiment, the step of obtaining the noise reduction parameter based on the matching value includes:
[0033] Get the maximum value among the matching values as the target matching value;
[0034] According to the target matching value, a corresponding judgment scene model is obtained as a target scene model;
[0035] Based on the target scene model, the corresponding noise reduction parameters are obtained.
[0036] In a preferred embodiment, before the step of obtaining the position information, time information and ambient audio signal of the earphone, the method further includes:
[0037] Detecting the degree of change of the ambient audio signal;
[0038] If the degree of change of the ambient audio signal exceeds a preset degree value, the acquisition of the position information, time information and ambient audio signal of the earphone is started.
[0039] A second aspect of the present invention provides a headset, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements any of the above steps of the method for adaptive noise reduction of the headset when executing the computer program.
[0040] A third aspect of the present invention provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps of any of the above-mentioned methods for adaptive noise reduction of headphones are implemented.
[0041] The beneficial effects of the method for adaptive noise reduction of headphones, headphones and storage medium provided by the present invention are: by utilizing key position information and setting the degree of association, the initial scene type and the associated scene type are determined, and the number of scene types that need to be recognized is preliminarily reduced. Combined with time information, the judgment scene type is further determined, and the number of scene models that need to be recognized is further reduced. Combined with the scene recognition model and the recognition weight data, the matching value of each judgment scene model and the current scene of the headphones can be obtained, and the noise reduction parameters required for adaptive noise reduction can be obtained based on the matching value, thereby reducing the calculation amount of scene recognition, effectively improving the efficiency and accuracy of adaptive noise reduction of the headphones, reducing the power consumption of the headphones to a certain extent, and improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 A schematic diagram of a first process flow of a method for adaptive noise reduction of headphones disclosed in an embodiment of the present invention;
[0043] Figure 2A second flow chart of the method for adaptive noise reduction of headphones disclosed in an embodiment of the present invention;
[0044] Figure 3 A third flow chart of the method for adaptive noise reduction of headphones disclosed in an embodiment of the present invention;
[0045] Figure 4 A fourth flow chart of the method for adaptive noise reduction of headphones disclosed in an embodiment of the present invention;
[0046] Figure 5 A fifth flow chart of the method for adaptive noise reduction of headphones disclosed in an embodiment of the present invention;
[0047] Figure 6 A sixth flow chart of the method for adaptive noise reduction of headphones disclosed in an embodiment of the present invention;
[0048] Figure 7 This is a schematic diagram of the module structure of the earphone disclosed in an embodiment of the present invention. DETAILED DESCRIPTION
[0049] In the present invention, the terms "disposed", "provided with" and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral structure; it can be a mechanical connection, or an electrical connection; it can be a direct connection, or an indirect connection through an intermediate medium, or it can be an internal connection between two devices, elements or components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0050] The terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of the features. In the description of this application, "plurality" means at least two, such as two, three, etc., unless otherwise clearly and specifically defined.
[0051] In addition, some of the above terms may be used to express other meanings in addition to indicating orientation or positional relationship. For example, the term "on" may also be used to express a certain dependency or connection relationship in some cases. For those skilled in the art, the specific meanings of these terms in the present invention can be understood according to specific circumstances.
[0052] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0053] With the rapid development of electronic technology, wireless headphones have become an indispensable item for people's daily travel. When faced with various relatively noisy environments, people usually choose to wear wireless headphones to reduce the impact of environmental noise. Different noise scenes have different requirements for the degree of noise reduction of headphones. Some existing wireless headphones use adaptive noise reduction to reduce environmental noise, but their scene recognition efficiency and accuracy are relatively low, the output noise reduction parameters cannot be well adapted to the current scene type, and the power consumption of the headphones is relatively high, which affects the user experience. The present invention provides a method for adaptive noise reduction of headphones, headphones and a storage medium, which solve the technical problems of low scene recognition efficiency and accuracy of adaptive noise reduction of headphones, and excessive power consumption.
[0054] The following is the content of the first aspect of the present invention:
[0055] Please refer to Figure 1 In this embodiment, the method for adaptive noise reduction of headphones includes the following steps:
[0056] S1. Obtaining the current earphone position information, time information and ambient audio signal.
[0057] Among them, the earphones mentioned in this embodiment are earphones with active noise reduction function. The earphones can be wireless earphones or wired earphones. They include a processor and an audio output unit. The processor can generate an audio signal with a phase opposite to the ambient noise through the acquired noise reduction parameters, thereby offsetting the external ambient noise.
[0058] The location information of the headset can be location information with the headset as the positioning center, or it can be location information with the positioning center of the terminal device connected to the headset. Accordingly, the location information of the headset can be obtained directly or indirectly. Specifically, a positioning module can be set inside the headset, and the location information of the current headset can be directly obtained through the positioning module; it can also be indirectly obtained through a terminal device (such as a mobile phone, smart watch, etc.) with a positioning module set inside, and the location information of the terminal device is approximately equivalent to the location information of the headset. In the daily use of the headset, the headset is mostly connected to a terminal device for use. Preferably, the location information of the terminal device can be used to indirectly obtain the location information of the headset.
[0059] Similarly, the time information can be obtained through a timing module provided inside the headset, or through a timing module on a terminal device connected to the headset, which will not be described in detail here.
[0060] The ambient audio signal can be obtained by a collection microphone set on the headset. Specifically, the collection microphone on the headset can obtain the sound signal of the environment where the headset is located in real time, or can collect the sound signal of the environment where the headset is located within a preset period of time.
[0061] S2, determining whether the location information contains key location information;
[0062] S3. If yes, determine the initial scene type and the associated scene type based on the key location information;
[0063] S4. Obtain a determination scene model based on the initial scene type, the associated scene type and the time information.
[0064] Among them, the key position information is the feature information in the position information, which can reflect the current scene of the headset to a certain extent. The initial scene type is the scene type confirmed by the key position information, the associated scene type is the scene type whose degree of association with the initial scene type meets the requirements, and there is a certain degree of similarity between the environmental sound features of the associated scene type and the environmental sound features of the initial scene type. The judgment scene model is an audio feature model that is pre-divided and trained for scene matching comparison.
[0065] In the daily use of headphones by users, the noise intensity and noise characteristics of different usage scenarios are different, and the noise intensity and noise characteristics of the same usage scenario may also be different in different time periods. For example, the scene where users use headphones can be an office area or a shopping mall area, and the sound characteristics of the office area in different time periods are different, such as the sound characteristics of the office area during office hours and the office area during lunch break time.
[0066] In order to improve the accuracy of scene recognition and the adaptability of noise reduction requirements, different scene types can be divided first, and then each scene type can be further divided based on time information to obtain a scene model based on the time information of each scene type. Specifically, several scene types can be divided, such as office scene A, home scene B, etc., and then different scene models can be classified for each scene type based on each time period, and corresponding noise reduction parameters can be configured for each scene model. For example, office scene A can be divided into scene models of different time periods such as a1 and a2, and different noise reduction parameters such as X1 and X2 can be configured respectively. It can be understood that different scene types can be divided into different numbers of scene models according to different time intervals, and the same scene type can also be divided into scene models of different time interval lengths. The division can be selected according to the actual situation, which improves the adaptability of the scene model and thus improves the accuracy of scene recognition.
[0067] Specifically, corresponding to the method of obtaining the location information, the key location information can be determined by the headset or by the terminal device connected to the headset. Preferably, after the terminal device connected to the headset obtains the location information, the terminal device can perform feature extraction and determination to determine whether the location information contains the key location information. It is easy to understand that when the key location information can be extracted from the location information, it means that the scene where the headset is currently located has obvious characteristics, so the corresponding scene type can be preliminarily and quickly determined through the key location information, that is, the initial scene type is obtained.
[0068] After obtaining the key position information and determining the initial scene type, it is possible to further determine the associated scene type that has a certain degree of similarity in sound characteristics with the initial scene type. The associated scene type can be obtained through key position information or through the initial scene type. Specifically, after obtaining the key position information, the corresponding associated position information can be obtained based on the correlation of the key position information, and then the corresponding associated scene type can be determined; or the corresponding scene type that meets the correlation requirements can be obtained based on the correlation of the initial scene type, and then the associated scene type can be obtained. It can be understood that by obtaining the associated scene type, on the basis of the preliminary determination of the scene type based on the position information, the scene type with a certain degree of similarity can be further expanded, thereby improving the accuracy of scene recognition of adaptive noise reduction.
[0069] After determining the initial scene type and the associated scene type, the scene models related to the initial scene type and the associated scene type can be obtained, and then divided and screened in combination with the time information to obtain the scene models corresponding to the initial scene type and the associated scene type in the current time period, and then obtain the judgment scene model for subsequent matching and comparison. Specifically, first obtain the specific scene model based on the initial scene type and the associated scene type; then filter the judgment scene model from the specific scene model according to the time information.
[0070] It can be understood that by combining time information to screen and obtain the judgment scene model for matching and comparison, the number of scene models used for matching and comparison can be effectively reduced, the efficiency and accuracy of scene recognition of adaptive noise reduction can be further improved, the workload of subsequent matching and comparison can be reduced, and the power consumption of the headphones can be reduced to a certain extent.
[0071] S5. Obtain a matching value based on the environmental audio signal, the determination scene model, the preset recognition weight data and the scene recognition model.
[0072] Among them, the recognition weight data is the emphasis of the human ear when perceiving the ambient audio, and the recognition weight data can be obtained by pre-training and detection. The matching value is used to reflect the degree of matching between the ambient audio of the current scene in which the headset is located and the judgment scene model. When the matching value is higher, the matching degree between the ambient audio of the current scene and the judgment scene model is higher; when the matching value is lower, the matching degree between the ambient audio of the current scene and the judgment model is lower. The scene recognition model can be a convolutional neural network model, a recurrent neural network model, or a combination of the two, or a Transformer scene recognition model, which is not limited in the embodiments of the present application.
[0073] After obtaining the judgment scene model, the ambient audio signal needs to be preprocessed to extract a feature sequence that can be applied to the scene recognition model. After the ambient audio signal is processed, the processed ambient audio signal is matched and compared with the judgment scene model according to the scene recognition model and recognition weight data to obtain the matching value of each scene model.
[0074] Preferably, the scene recognition model uses the Transformer scene recognition model, please refer to Figure 5 , based on the environmental audio signal, the determination scene model, the preset recognition weight data and the scene recognition model, the step S5 of obtaining the matching value includes:
[0075] S51, performing feature extraction processing on the ambient audio signal to obtain an ambient audio spectrum;
[0076] S52: Based on the preset scene recognition module and recognition weight data, compare the ambient audio spectrum with the determination scene model to obtain a corresponding matching value.
[0077] After obtaining the ambient audio signal in the time domain, the ambient audio signal is first framed and windowed to obtain a more continuous ambient audio signal, and then the ambient audio signal is short-time Fourier transformed to obtain a linear spectrum of the ambient audio in the frequency domain. The linear spectrum of the ambient audio in the frequency domain is then input into the Mel filter group to obtain a Mel nonlinear spectrum, and then the Mel nonlinear spectrum is logarithmically processed to obtain an ambient audio spectrum that can be used for scene recognition.
[0078] In the process of frame segmentation and windowing, the frame length of each frame can be 15ms to 30ms, preferably 25ms. The window function can be a rectangular window, a Hamming window, etc., which is not limited here. Preferably, the window function is a Hamming window.
[0079] Specifically, a Hamming window of 25 ms is used, and frames are extracted every 10 ms, that is, adjacent frames overlap by 15 ms. It can be understood that the frame processing method can smoothly process sudden change signals, avoid leakage, and improve the continuity of environmental audio signals.
[0080] The linear spectrum of the ambient audio obtained after the short-time Fourier transform of the ambient audio signal is not enough to reflect the characteristics of human auditory perception, so it is necessary to pass through the Mel filter group to obtain the Mel nonlinear spectrum. It can be understood that through the filtering effect of the Mel filter group, the frequency components that do not match the human auditory perception can be filtered out, and the frequency components that match the human auditory perception can be retained, which can improve the accuracy of subsequent scene recognition. The required ambient audio spectrum can be obtained by logarithmically processing the Mel nonlinear spectrum. It can be understood that logarithmically processing the Mel nonlinear spectrum can reduce the calculation amount of subsequent scene recognition to a certain extent, thereby improving the efficiency of scene recognition.
[0081] Furthermore, the recognition weight data is used as the composition data of the attention baseline focus of the compiled semantics in the Transformer scene recognition model. After obtaining the ambient audio spectrum and inputting it into the Transformer scene recognition model, the Transformer scene recognition model extracts features from the ambient audio spectrum based on the recognition weight data, and then obtains the ambient audio high-order feature sequence. The Transformer scene recognition model calls the required judgment scene model, matches and compares the ambient audio high-order feature sequence with the judgment scene model, and then obtains the matching value corresponding to each judgment scene model based on the calculation of the linear classifier.
[0082] It is understandable that by processing the ambient audio signals, calling the judgment scene model and selecting the Transformer scene recognition model, the number of scene models that need to be recognized and judged in the scene recognition process can be effectively reduced, thereby reducing unnecessary calculations and improving the efficiency and accuracy of scene recognition, thereby improving the accuracy of the adaptive noise reduction of the headphones, reducing the power consumption of the headphones to a certain extent, and improving the user experience.
[0083] S6. Obtain noise reduction parameters based on the matching value.
[0084] Among them, the noise reduction parameters include noise reduction filter parameters and acoustic EQ parameters. After obtaining the noise reduction parameters, the processor of the headset can generate an audio signal with a phase opposite to the ambient noise based on the noise reduction parameters, thereby offsetting the external ambient noise.
[0085] Specifically, after obtaining the matching value, the preset comparison threshold of the matching value can be compared according to the matching value, and different noise reduction mode situations can be divided based on the comparison situation. The scene model corresponding to the required matching value can also be directly called based on the size of the matching values to obtain the corresponding noise reduction parameters.
[0086] Please refer to Figure 6In a preferred embodiment, step S6 of obtaining noise reduction parameters based on the matching value includes:
[0087] S61, obtaining the maximum value among the matching values as the target matching value;
[0088] S62, obtaining a corresponding determination scene model according to the target matching value as a target scene model;
[0089] S63: Based on the target scene model, obtain corresponding noise reduction parameters.
[0090] It is easy to understand that the target matching value reflects the best matching degree between the judgment scene model and the ambient audio of the scene where the headset is located. The judgment scene model is obtained based on key position information and time information. By determining the judgment scene model, the approximate range of the scene where the headset is currently located has been preliminarily determined. Then, by screening the matching value and obtaining the target matching value, the target scene model with the highest matching degree with the current scene of the headset can be obtained, and then the corresponding noise reduction parameters can be obtained, which optimizes the method of obtaining the noise reduction parameters.
[0091] It can be understood that the present application determines the initial scene type and the associated scene type by utilizing key position information and setting the degree of association, thereby preliminarily reducing the number of scene types that need to be recognized. Combined with the time information, the judgment scene type is further determined, and the number of scene models that need to be recognized is further reduced. Combined with the scene recognition model and the recognition weight data, the matching value of each judgment scene model and the current scene of the headset can be obtained, and the noise reduction parameters required for adaptive noise reduction can be obtained based on the matching value, thereby reducing the calculation amount of scene recognition, effectively improving the efficiency and accuracy of the adaptive noise reduction of the headset, reducing the power consumption of the headset to a certain extent, and improving the user experience.
[0092] For further information, please refer to Figure 2 In one embodiment, step S3 of determining the initial scene type and the associated scene type based on the key location information includes:
[0093] S31, determining an initial scene type from preset scene types according to the key position information;
[0094] S32, obtaining the association value between the initial scene type and other scene types;
[0095] S33: Filter out associated scene types based on a comparison between the associated value and a preset first threshold.
[0096] The correlation value reflects the similarity between the sound features of each scene type, and the correlation value can be obtained in advance by comparing the sound features of the scene models in each scene type and undergoing a large amount of scene model training. The first threshold is a preset comparison value, which is mainly used to screen the correlation value to obtain the associated scene type that meets the requirements.
[0097] Specifically, the method of determining the initial scene type using the key position information can be to identify the key position information and directly determine the corresponding initial scene type, or to determine the degree of adaptation of the key position information to each scene type, and then determine the initial scene type whose degree of adaptation meets the requirements, such as obtaining the scene type with the highest degree of adaptation as the initial scene type, or obtaining the scene type with a degree of adaptation exceeding a preset degree as the initial scene type. It is easy to understand that the number of initial scene types that can be determined by the key position information can be one or more.
[0098] After the initial scene type is determined, the association value between the initial scene type and other scene types can be obtained. In the process of screening out the associated scene types based on the comparison between the association value and the first threshold, the situation of the first threshold will affect the determination of the associated scene type. The first threshold can be set by setting a fixed value or by establishing a database of the first threshold, and then calling a suitable first threshold based on the specific situation.
[0099] It can be understood that through key position information, the pre-set scene types are preliminarily screened to obtain the initial scene type, and then combined with the association value of the initial scene type, the associated scene type with similarity is obtained, which ensures the efficiency and accuracy of subsequent scene recognition, and can effectively reduce the number of subsequent scene recognitions, reduce unnecessary calculations, and reduce the power consumption of the headphones.
[0100] For further information, please refer to Figure 3 In a preferred embodiment, the step S31 of determining the initial scene type from the preset scene types according to the key position information includes:
[0101] S311, obtaining an adaptation value between key location information and a preset scene type;
[0102] S312. Obtain a maximum adaptation value from the adaptation values;
[0103] S313. Obtain a corresponding scene type according to the maximum adaptation value as an initial scene type.
[0104] Among them, the adaptation value reflects the degree of adaptation between the key position information and the scene type. The larger the adaptation value, the higher the degree of adaptation between the scene type and the key position information; the smaller the adaptation value, the lower the degree of adaptation between the scene type and the key position information.
[0105] Specifically, each scene type is pre-set with a corresponding trigger information element. After obtaining the key position information, the key position information can be matched with the trigger information element of each scene type, and then the adaptation value corresponding to each scene type can be output.
[0106] It is easy to understand that by obtaining the maximum adaptation value from the obtained adaptation values and obtaining the initial scene type in a preferential manner, the efficiency of scene recognition can be improved to a certain extent by reducing the number of scene types subsequently used for scene recognition while ensuring the accuracy of subsequent scene recognition.
[0107] For further information, please refer to Figure 4 In a preferred embodiment, based on the comparison between the correlation value and the preset first threshold, the step S33 of screening out the associated scene type includes:
[0108] S331, obtaining a maximum adaptation value between key location information and a preset scene type;
[0109] S332: Determine the adaptation value interval into which the maximum adaptation value falls, and obtain a first threshold corresponding to the adaptation value interval;
[0110] S333, determining whether the correlation value is less than a first threshold;
[0111] S334: If not, mark the scene type corresponding to the associated value as an associated scene type.
[0112] Each adaptation value interval corresponds to a first threshold value, and the number of adaptation value intervals and first threshold values can be preset according to design requirements.
[0113] Specifically, after obtaining the maximum adaptation value, the adaptation value interval in which the maximum adaptation value falls can be determined first, and then the corresponding first threshold value can be obtained according to the adaptation value interval in which it falls. When the maximum adaptation value is relatively high, a first threshold value with a relatively large value is obtained, and when the maximum adaptation value is relatively low, a first threshold value with a relatively small value is obtained.
[0114] It is easy to understand that by obtaining the corresponding first threshold based on the maximum adaptation value, it is possible to dynamically adjust the difficulty of obtaining associated scene types according to the degree of adaptation of the key position information and the scene type, thereby controlling the number of associated scene types obtained according to the adaptation value. That is, when the maximum adaptation value is relatively high, the size of the first threshold can be appropriately increased, which can effectively reduce the amount of associated scene models obtained, thereby improving the efficiency of subsequent scene recognition; when the maximum adaptation value is relatively low, the size of the first threshold can be appropriately reduced, and relatively more associated scene models can be obtained, thereby ensuring the accuracy of subsequent scene recognition.
[0115] Furthermore, in one embodiment, before the step of obtaining the location information, time information and ambient audio signal of the earphone, the method further includes:
[0116] Detecting the degree of change of the ambient audio signal;
[0117] If the degree of change of the ambient audio signal exceeds a preset degree value, the acquisition of the position information, time information and ambient audio signal of the earphone is started.
[0118] The degree of change of the ambient audio signal may be any one or a combination of any of the indicators such as the frequency characteristics and energy characteristics of the ambient audio.
[0119] It can be understood that before triggering the scene recognition of adaptive noise reduction, the adaptive noise reduction function will be triggered by detecting the degree of change of the ambient audio signal. Only when the degree of change of the ambient audio signal exceeds the preset value will the scene recognition be started. This can effectively avoid the frequent triggering of scene recognition of adaptive noise reduction and the frequent false triggering and switching of noise reduction parameters, which can effectively reduce the power consumption of headphones and improve the user experience.
[0120] To summarize, the present application determines the initial scene type and the associated scene type by utilizing key position information and setting the degree of association, thereby preliminarily reducing the types of scenes that need to be recognized. Combined with time information, the judgment scene type is further determined, and the number of scene models that need to be recognized is further reduced. Combined with the scene recognition model and the recognition weight data, the matching value of each judgment scene model and the current scene of the headset can be obtained, and the noise reduction parameters required for adaptive noise reduction can be obtained based on the matching value, thereby reducing the calculation amount of scene recognition, effectively improving the efficiency and accuracy of the adaptive noise reduction of the headset, reducing the power consumption of the headset to a certain extent, and improving the user experience.
[0121] The following is the content of the second aspect of the present invention:
[0122] The present invention provides an earphone, such as Figure 7As shown, the headset includes a memory 10, a processor 20, and a method program instruction 30 for adaptive noise reduction of the headset stored in the memory 10 and executable on the processor 20. When the method program instruction 30 for adaptive noise reduction of the headset is executed by the processor 20, the aforementioned method for adaptive noise reduction of the headset is implemented.
[0123] In some embodiments, the processor may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor is generally used to control the overall operation of the headset. In this embodiment, the processor is used to run program codes stored in a readable storage medium or process data.
[0124] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a readable storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods of various embodiments of the present invention.
[0125] The following is the content of the third aspect of the present invention:
[0126] The present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned headphone adaptive noise reduction method are implemented.
[0127] The above is only a specific implementation method of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A method for adaptive noise reduction of headphones, characterized in that: include: Obtain the current location information, time information and ambient audio signal of the headset; Determining whether the location information contains key location information; If yes, determining the initial scene type and the associated scene type based on the key location information; Based on the initial scene type, the associated scene type and the time information, a determination scene model is obtained; Obtaining a matching value based on the environmental audio signal, the determination scene model, the preset recognition weight data and the scene recognition model; Based on the matching value, a noise reduction parameter is obtained.
2. The method for adaptive noise reduction of headphones according to claim 1, characterized in that: The step of determining the initial scene type and the associated scene type based on the key position information further includes: Determining the initial scene type from preset scene types according to the key position information; Obtaining association values between the initial scene type and other scene types; Based on the comparison between the association value and a preset first threshold, the associated scene type is screened out.
3. The method for adaptive noise reduction of headphones according to claim 2, characterized in that: The step of determining the initial scene type from preset scene types according to the key position information includes: Obtaining an adaptation value between the key location information and a preset scene type; From the adaptation values, obtaining a maximum adaptation value; According to the maximum adaptation value, a corresponding scene type is obtained as an initial scene type.
4. The method for adaptive noise reduction of headphones according to claim 2, characterized in that: The step of screening out the associated scene type based on the comparison between the associated value and a preset first threshold comprises: Obtaining a maximum adaptation value between the key position information and a preset scene type; Determine a preset adaptation value interval within which the maximum adaptation value falls, and obtain a first threshold corresponding to the adaptation value interval; Determining whether the correlation value is less than the first threshold; If not, the scene type corresponding to the associated value is marked as an associated scene type.
5. The method for adaptive noise reduction of headphones according to claim 1, characterized in that: The step of obtaining a determination scene model based on the initial scene type, the associated scene type and the time information comprises: Based on the initial scene type and the associated scene type, obtaining a specific scene model; The determination scene model is obtained by screening from the specific scene models according to the time information.
6. The method for adaptive noise reduction of headphones according to claim 1, characterized in that: The step of obtaining a matching value based on the ambient audio signal, the determination scene model, the preset recognition weight data and the scene recognition model comprises: Performing feature extraction processing on the ambient audio signal to obtain an ambient audio spectrum; Based on a preset scene recognition module and recognition weight data, the ambient audio spectrum is compared with the determination scene model to obtain a corresponding matching value.
7. The method for adaptive noise reduction of headphones according to claim 1, characterized in that: The step of obtaining the noise reduction parameter based on the matching value comprises: Obtaining the maximum value among the matching values as the target matching value; According to the target matching value, a corresponding determination scene model is obtained as a target scene model; Based on the target scene model, corresponding noise reduction parameters are obtained.
8. The method for adaptive noise reduction of headphones according to claim 1, characterized in that: Before the step of obtaining the position information, time information and ambient audio signal of the earphone, the step further includes: Detecting the degree of change of the ambient audio signal; If the degree of change of the ambient audio signal exceeds a preset degree value, the acquisition of the position information, time information and ambient audio signal of the earphone is started.
9. A headset comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the headphone adaptive noise reduction method according to any one of claims 1 to 8 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method for adaptive noise reduction of headphones according to any one of claims 1 to 8 are implemented.