Adaptive active noise reduction method, earphone and storage medium
By using position information and time information in wireless headphones to optimize the acquisition method of noise reduction parameters, the problem of insufficient efficiency and accuracy of scene recognition and high power consumption of adaptive noise reduction in the prior art is solved, and more efficient and more accurate noise reduction effect and lower power consumption are achieved.
Patent Information
- Application Number
- CN202411987883.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-13
AI Technical Summary
The adaptive noise reduction method of existing wireless headphones is insufficient in scene recognition efficiency and accuracy, and it consumes a high power, which affects the user experience.
By obtaining the position information and time information of the headphones, determine whether it is located in a preset specific area and obtain the corresponding noise reduction parameters; if it is not in a specific area, the judgment scene model is obtained based on the key position information and time information, and the noise reduction parameters are obtained based on the environmental audio signal and the preset scene recognition model and recognition weight data.
It effectively reduces the frequency of adaptive noise reduction triggering scene recognition, improves the matching of noise reduction parameters, reduces the power consumption of the headphones, and improves the user experience.
Smart Images

Figure CN119996888A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of headphone noise reduction, and in particular to a method for adaptive active noise reduction, a headphone and a storage medium. Background Art
[0002] With the rapid development of electronic technology, wireless headphones have become an indispensable item for people's daily travel. When facing various relatively noisy environments, people usually choose to wear wireless headphones to reduce the impact of environmental noise.
[0003] Different noise environments have different requirements for noise reduction. To better cope with the changing noise environment, some existing wireless headphones use adaptive noise reduction to reduce the impact of environmental noise. However, the existing adaptive noise reduction method cannot identify the noise scene relatively accurately and efficiently, which leads to a low matching degree of noise reduction requirements, and the existing adaptive noise reduction has high power consumption, which in turn affects the user experience.
[0004] In view of this, it is necessary to provide an adaptive active noise reduction method, earphones and storage medium to solve the above problems. Summary of the invention
[0005] In view of the deficiencies in the prior art, the present invention provides a method, earphones and storage medium for adaptive active noise reduction, which aim to solve the technical problems of low scene recognition efficiency and accuracy and excessive power consumption of earphone adaptive noise reduction.
[0006] To achieve the above-mentioned purpose, a first aspect of the present invention provides a method for adaptive active noise reduction, which comprises:
[0007] Get the current location and time information of the headset;
[0008] Based on the location information, determining whether the headset is located in a preset specific area;
[0009] If yes, then combine the time information to obtain the preset specific noise reduction parameters;
[0010] If not, determining whether the location information contains key location information;
[0011] If yes, key position information and ambient audio signals are obtained;
[0012] Based on key location information and time information, obtain the judgment scene model;
[0013] The corresponding noise reduction parameters are obtained according to the ambient audio signal, the judgment scene model, the preset scene recognition model and the preset recognition weight data.
[0014] In a preferred embodiment, the step of determining whether the earphone is located in a preset specific area based on the location information includes:
[0015] Obtain the positioning point, specific point and range radius in the location information;
[0016] Calculate the distance between the positioning point and the specific point to obtain the relative distance;
[0017] Determine whether the relative distance is greater than the range radius;
[0018] If not, the headset is located within a preset specific area.
[0019] In a preferred embodiment, the step of obtaining the preset specific noise reduction parameters in combination with the time information includes:
[0020] Get the preset time interval corresponding to a specific area;
[0021] According to the time information, the time interval in which the time information is located is obtained by filtering and taking it as the specific time interval;
[0022] According to a specific time interval, the corresponding specific noise reduction parameters are obtained.
[0023] In a preferred embodiment, the step of determining whether the earphone is in a preset specific area based on the location information includes:
[0024] Detect whether a specific noise reduction mode is activated;
[0025] If yes, then based on the location information, determine whether the headset is located in a preset specific area;
[0026] If not, determine whether the location information contains key location information.
[0027] In a preferred embodiment, based on the key position information and time information, the step of obtaining the determination scene model includes:
[0028] Determine an initial scene type from preset scene types according to the key position information;
[0029] Get the associated values of the initial scene type and other scene types;
[0030] Based on the comparison between the correlation value and a preset first threshold, the correlation scene type is screened out;
[0031] A determination scene model is obtained according to the initial scene type, the associated scene type and the time information.
[0032] In a preferred embodiment, the step of determining the initial scene type from preset scene types according to the key position information includes:
[0033] Obtaining the adaptation value of key location information and preset scene type;
[0034] From the adaptation values, obtain the maximum adaptation value;
[0035] According to the maximum adaptation value, the corresponding scene type is obtained as the initial scene type.
[0036] In a preferred embodiment, based on the comparison between the association value and the preset first threshold, the step of screening out the associated scene type includes:
[0037] Obtain the maximum adaptation value between the key location information and the preset scene type;
[0038] Screening to obtain a preset adaptation value interval in which the maximum adaptation value falls, and obtaining a first threshold value corresponding to the adaptation value interval;
[0039] Determining whether the correlation value is less than a first threshold;
[0040] If not, the scene type corresponding to the associated value is marked as the associated scene type.
[0041] In a preferred embodiment, the step of obtaining corresponding noise reduction parameters according to the ambient audio signal, the determination scene model, the preset scene recognition model and the preset recognition weight data includes:
[0042] Perform feature extraction processing on the ambient audio signal to obtain the ambient audio spectrum;
[0043] Based on the preset Transformer scene recognition module and recognition weight data, the ambient audio spectrum is compared with the judgment scene model to obtain the corresponding matching value;
[0044] A maximum matching value is obtained from the matching values, and a noise reduction parameter corresponding to the judgment scene model corresponding to the maximum matching value is obtained.
[0045] A second aspect of the present invention provides a headset, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and when the processor executes the computer program, the steps of any one of the above-mentioned methods of adaptive active noise reduction are implemented.
[0046] A third aspect of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of any of the above-mentioned methods for adaptive active noise reduction are implemented.
[0047] The method, earphone and storage medium for adaptive active noise reduction provided by the present invention have the beneficial effects of: by setting a specific area and optimizing the method for obtaining the judgment scene model, based on a preset trigger order, the method for adaptively obtaining noise reduction parameters is optimized, and by first determining whether the earphone is in a specific area, the frequency of triggering scene recognition for adaptive noise reduction can be effectively reduced, and at the same time, the user can obtain specific noise reduction parameters that are more in line with their actual usage needs; after the earphone is not in the specific area, based on the key position information, the initial position information and the associated position information are screened out, and then combined with the time information, the judgment scene model is obtained, and then combined with the scene recognition model and the recognition weight data, the matching value of each judgment scene model and the current scene of the earphone can be obtained, and then the noise reduction parameters required for adaptive noise reduction can be obtained based on the matching value, which can effectively reduce the number of scene models used for matching and comparison, and at the same time improve the efficiency and accuracy of scene recognition of adaptive noise reduction, reduce the amount of calculation for subsequent matching and comparison, reduce the power consumption of the earphone to a certain extent, and improve the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 A first flow chart of the method for adaptive active noise reduction disclosed in an embodiment of the present invention;
[0049] Figure 2 A second flow chart of the method for adaptive active noise reduction disclosed in an embodiment of the present invention;
[0050] Figure 3 A third flow chart of the method for adaptive active noise reduction disclosed in an embodiment of the present invention;
[0051] Figure 4 A fourth flow chart of the method for adaptive active noise reduction disclosed in an embodiment of the present invention;
[0052] Figure 5 A fifth flow chart of the method for adaptive active noise reduction disclosed in an embodiment of the present invention;
[0053] Figure 6 A sixth flow chart of the method for adaptive active noise reduction disclosed in an embodiment of the present invention;
[0054] Figure 7 This is a schematic diagram of the module structure of the earphone disclosed in an embodiment of the present invention. DETAILED DESCRIPTION
[0055] In the present invention, the terms "disposed", "provided with" and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral structure; it can be a mechanical connection, or an electrical connection; it can be a direct connection, or an indirect connection through an intermediate medium, or it can be an internal connection between two devices, elements or components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0056] The terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of the features. In the description of this application, "plurality" means at least two, such as two, three, etc., unless otherwise clearly and specifically defined.
[0057] In addition, some of the above terms may be used to express other meanings in addition to indicating orientation or positional relationship. For example, the term "on" may also be used to express a certain dependency or connection relationship in some cases. For those skilled in the art, the specific meanings of these terms in the present invention can be understood according to specific circumstances.
[0058] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0059] With the rapid development of electronic technology, wireless headphones have become an indispensable item for people's daily travel. When faced with various relatively noisy environments, people usually choose to wear wireless headphones to reduce the impact of environmental noise. Different noise scenes have different requirements for the degree of noise reduction of headphones. Some existing wireless headphones use adaptive noise reduction to reduce environmental noise, but their scene recognition efficiency and accuracy are relatively low, the output noise reduction parameters cannot be well adapted to the current scene type, and the power consumption of the headphones is relatively high, which affects the user experience. The present invention provides a method, headphones and storage medium for adaptive active noise reduction, which solve the technical problems of low scene recognition efficiency and accuracy of adaptive noise reduction of headphones and excessive power consumption.
[0060] The following is the content of the first aspect of the present invention:
[0061] The earphones mentioned in this embodiment are earphones with active noise reduction function. The earphones can be wireless earphones or wired earphones. They include a processor and an audio output unit. The processor can generate an audio signal with a phase opposite to the ambient noise through the acquired noise reduction parameters, thereby offsetting the external ambient noise.
[0062] Please refer to Figure 1 In this embodiment, the steps of the method for adaptive active noise reduction include:
[0063] S1. Obtain the current location information and time information of the headset.
[0064] Among them, the location information of the earphone can be the location information with the earphone as the positioning center, or it can be the location information with the positioning center of the terminal device connected to the earphone. Accordingly, the location information of the earphone can be obtained directly or indirectly. Specifically, a positioning module can be set inside the earphone, and the location information of the current earphone can be directly obtained through the positioning module; it can also be indirectly obtained through a terminal device (such as a mobile phone, smart watch, etc.) with a positioning module inside, and the location information of the terminal device is approximately equivalent to the location information of the earphone. In the daily use of the earphone, the earphone is mostly connected to a terminal device for use. Preferably, the location information of the terminal device can be used to indirectly obtain the location information of the earphone.
[0065] Similarly, the time information can be obtained through a timing module provided inside the headset, or through a timing module on a terminal device connected to the headset, which will not be described in detail here.
[0066] S2. determining whether the headset is located in a preset specific area based on the location information;
[0067] S3. If yes, obtain preset specific noise reduction parameters in combination with the time information.
[0068] The specific area is the usage area range defined by the user according to his / her own usage. The specific area can be a high-frequency usage scenario in which the user uses the headset, or a fixed usage scenario randomly defined by the user. The number of specific areas defined is not limited here. The specific noise reduction parameters are the noise reduction parameters pre-configured by the user for the specific area defined by the user according to his / her own usage habits.
[0069] When the user sets a specific area, the user can define the specific area through the control software on the terminal, or can assist in defining the specific area by operating the headset. The definition method of the specific area can be selected according to the design requirements or triggered according to the user's choice. Preferably, the control software can be used to define the specific area.
[0070] Specifically, when the user uses the control software on the terminal to demarcate a specific area, the user can obtain a range map through the control software and directly mark the specific area on the range map. The control software can record the relevant parameter information of the relevant specific area through measurement.
[0071] When the user uses headphones to assist in demarcating a specific area, the user needs to be within the specific area and trigger the recording mode of the headphones. When the user is near the center of the specific area and touches the headphones, the terminal records the first position point. The user moves to the farthest edge of the specific area and touches the headphones again, the terminal records the second position point. The specific area can be obtained based on the first position point and the second position point.
[0072] When the user sets the specific noise reduction parameters corresponding to a specific area, the user can divide the specific noise reduction parameters corresponding to different time periods for the specific area. It is easy to understand that in different time periods, the user's requirements for noise reduction in a specific area may be different. Therefore, when the user sets the specific noise reduction parameters, the user can divide the specific noise reduction parameters corresponding to different time periods according to his own usage needs.
[0073] When determining whether the headset is within a preset specific area, the determination is usually made based on the relationship between the distance from the headset to the specific area and the determination distance of the specific area. Since the location information can be directly obtained through the headset or indirectly obtained through the terminal device, the determination distance can be obtained directly through the range radius of the specific area or by correcting the range radius of the specific area. The determination distance acquisition method can be set according to the design requirements.
[0074] It can be understood that by pre-setting specific areas and specific noise reduction parameters, the noise reduction parameters can be automatically adjusted after the headphones enter a specific area, which optimizes the method of obtaining the noise reduction parameters, reduces the frequency of scene recognition through audio to obtain noise reduction parameters, and reduces the amount of calculation required to obtain the noise reduction parameters, thereby reducing the power consumption of the headphones and improving the user experience.
[0075] S4. If not, determine whether the location information contains key location information;
[0076] S5. If yes, obtain key position information and ambient audio signals;
[0077] S6. Obtaining a judgment scene model based on key location information and time information;
[0078] S7. Obtain corresponding noise reduction parameters according to the ambient audio signal, the determination scene model, the preset scene recognition model and the preset recognition weight data.
[0079] Among them, the key position information is the feature information in the position information, which can reflect the situation of the current scene of the headset to a certain extent. The judgment scene model is an audio feature model that is pre-divided and trained for scene matching and comparison. The ambient audio signal is the sound information of the current scene of the headset, which can be obtained through the collection microphone set on the headset. Specifically, the collection microphone on the headset can obtain the sound signal of the environment in which the headset is located in real time, or it can collect the sound signal of the environment in which the headset is located within a preset period. The recognition weight data is the emphasis of the human ear when perceiving the ambient audio, and the recognition weight data can be obtained by pre-training and detection. The scene recognition model can be a convolutional neural network model, a recurrent neural network model, or a combination of the two, or a Transformer scene recognition model, which is not limited to this in the embodiment of the present application.
[0080] In the daily use of headphones by users, the noise intensity and noise characteristics of different usage scenarios are different, and the noise intensity and noise characteristics of the same usage scenario may also be different in different time periods. For example, the scene where users use headphones can be an office area or a shopping mall area, and the sound characteristics of the office area in different time periods are different, such as the sound characteristics of the office area during office hours and the office area during lunch break time.
[0081] In order to improve the accuracy of scene recognition and the adaptability of noise reduction requirements, different scene types can be divided according to key location information, and then each scene type can be further divided in combination with time information, so that each scene type includes at least one scene model corresponding to a time period. Specifically, several scene types can be divided, such as office scene A, home scene B, etc., and then different scene models are classified for each scene type in combination with each time period, and corresponding noise reduction parameters are configured for each scene model. For example, office scene A can be divided into scene models of different time periods such as a1 and a2, and different noise reduction parameters such as X1 and X2 are configured respectively. It can be understood that different scene types can be divided into different numbers of scene models according to different time intervals, and the same scene type can also be divided into scene models of different time interval lengths. The division can be selected according to the actual situation, which improves the adaptability of the scene model and thus improves the accuracy of scene recognition.
[0082] Specifically, corresponding to the method of obtaining the position information, the key position information can be determined by the headset or by the terminal device connected to the headset. Preferably, after the terminal device connected to the headset obtains the position information, the terminal device can perform feature extraction and determination to determine whether the position information contains the key position information. It is easy to understand that when the key position information can be extracted from the position information, it means that the scene in which the headset is currently located has obvious characteristics, so the corresponding scene type can be preliminarily and quickly determined through the key position information, and then combined with the time information, the required scene model can be obtained from the determined scene type, and then the required determination scene model can be obtained.
[0083] The method of obtaining and determining the scene model through key position information and time information can be to directly determine the scene type through the key position information, and then obtain the corresponding scene type from the determined scene type in combination with the time information, or to determine the initial scene type through the key position information, and then obtain the scene type that meets the association requirements with the initial scene type, and then obtain the corresponding scene model from the aforementioned scene type in combination with the time information.
[0084] After obtaining the judgment scene model, the ambient audio signal needs to be preprocessed to extract audio data that can be applied to the scene recognition model. After obtaining the aforementioned audio data, the processed ambient audio signal is matched and compared with the judgment scene model according to the scene recognition model and recognition weight data to obtain the matching value of each scene model.
[0085] It can be understood that by combining key position information with time information to screen out the judgment scene models used for matching and comparison, the number of scene models used for matching and comparison can be effectively reduced, the efficiency and accuracy of scene recognition for adaptive noise reduction can be further improved, the amount of calculation for subsequent matching and comparison can be reduced, and the power consumption of the headphones can be reduced to a certain extent.
[0086] This embodiment optimizes the method of adaptively acquiring noise reduction parameters based on a preset trigger order by setting a specific area and optimizing the method of acquiring the judgment scene model. It first determines whether the headset is in a specific area, which can effectively reduce the frequency of triggering scene recognition for adaptive noise reduction, and at the same time enable users to obtain specific noise reduction parameters that are more in line with their actual usage needs. After the headset is not in the specific area, the judgment scene model is screened based on key position information and time information, and then the matching value of each judgment scene model and the current scene of the headset is obtained by combining the scene recognition model and the recognition weight data. Then, the noise reduction parameters required for adaptive noise reduction can be obtained based on the matching value, which can effectively reduce the number of scene models used for matching and comparison, and at the same time improve the efficiency and accuracy of scene recognition of adaptive noise reduction, reduce the amount of calculation for subsequent matching and comparison, reduce the power consumption of the headset to a certain extent, and improve the user experience.
[0087] For further information, please refer to Figure 2 In one embodiment, step S2 of determining whether the headset is located in a preset specific area based on the location information includes:
[0088] S21, obtaining the positioning point, the specific point of the specific area and the range radius in the location information;
[0089] S22, calculating the distance between the positioning point and the specific point to obtain a relative distance;
[0090] S23, determining whether the relative distance is greater than the range radius;
[0091] S24: If not, the headset is located in a preset specific area.
[0092] Among them, the positioning point is the positioning center of the current position of the headset, and the specific point is a point within the central range of the specific area, preferably the central point. The range radius is the distance extending outward from the specific point to the farthest edge of the specific area. The range of the specific area can be determined by the specific point and the range radius. The range size of different specific areas may be different, so when determining whether the headset is in a specific area, it is necessary to combine the range radius to determine.
[0093] Specifically, in the actual use process of the user, the user may set one specific area or multiple specific areas based on his own use needs, that is, there is at least one specific point and range radius. Therefore, when determining whether the headset is located in the preset specific area, it is necessary to obtain all the set specific points and range radii, and then calculate the relative distance between the positioning point and the obtained specific point, and then compare the relative distance with the corresponding range radius to determine whether the headset is in the preset specific area.
[0094] The relative distance can be calculated using the distance formula between two points. It is easy to understand that after the relative distance is calculated, by directly comparing the size relationship between the relative distance and the range radius, it is possible to quickly determine whether the headset is located in a specific area. On the basis of having high judgment accuracy, the amount of calculation is effectively reduced, a faster response can be achieved, and the desired result can be determined.
[0095] For further information, please refer to Figure 3 In this embodiment, after determining that the earphone is in a preset specific area, the step S3 of obtaining the preset specific noise reduction parameter in combination with the time information includes:
[0096] S31, obtaining a preset time interval corresponding to a specific area;
[0097] S32, according to the time information, filtering and obtaining the time interval in which the time information is located as the specific time interval;
[0098] S33. Obtain corresponding specific noise reduction parameters according to the specific time interval.
[0099] Among them, the time interval is a time period designated by the user according to his / her own usage requirements. Each time interval is pre-set with a corresponding specific noise reduction parameter. The size of the time interval can be adjusted according to the user's usage requirements. If the user does not divide the corresponding time interval for a specific area, it is regarded as a time interval, and any time information is located in this time interval. For example, the user can set two time intervals a and b for a specific area A. The interval size of a and the interval size of b can be the same or different. Time interval a corresponds to noise reduction parameter x, and time interval b corresponds to noise reduction parameter y. The specific time interval is the time interval in which the current time is located.
[0100] Specifically, after determining that the headphones are located in a specific area, the preset time intervals corresponding to the specific area can be obtained, and then the time information can be compared with the endpoint moments of the time intervals one by one to determine the time interval in which the time information falls, and the time interval can be marked as a specific time interval. The noise reduction parameters corresponding to the specific time interval can then be obtained through the specific time interval.
[0101] It can be understood that by dividing the time interval, the matching of specific noise reduction parameters and time information can be achieved. By determining the time interval in which the time information is located, the efficiency and accuracy of obtaining specific noise reduction parameters corresponding to a specific area at the current time can be improved.
[0102] Further, in a preferred embodiment, the step of determining whether the earphone is in a preset specific area based on the location information includes:
[0103] Detect whether a specific noise reduction mode is activated;
[0104] If yes, then based on the location information, determine whether the headset is located in a preset specific area;
[0105] If not, determine whether the location information contains key location information.
[0106] Among them, the specific noise reduction mode is a noise reduction mode that determines whether the headset is in a preset specific area based on location information to obtain specific noise reduction parameters.
[0107] It is understandable that the specific area and specific noise reduction parameters are pre-set by the user according to their actual needs. After completing the settings, the user can choose whether to start the corresponding specific noise reduction mode according to their own situation. When the specific noise reduction mode is started, in the process of obtaining the noise reduction parameters, it will give priority to judging whether the headset is in the preset specific area based on the location information; when the specific noise reduction mode is not started, in the process of obtaining the noise reduction parameters, it will directly judge whether the location information contains key location information to perform adaptive noise reduction scene recognition. By controlling the start and stop of the specific noise reduction mode, the user can adjust the way of obtaining noise reduction parameters for adaptive noise reduction according to their own usage needs, thereby increasing the operability of user use and improving the user experience.
[0108] For further information, please refer to Figure 4 In one embodiment, based on the key position information and time information, step S6 of obtaining the determination scene model includes:
[0109] S61, determining an initial scene type from preset scene types according to the key position information;
[0110] S62, obtaining association values between the initial scene type and other scene types;
[0111] S63, based on the comparison between the correlation value and a preset first threshold, screening out the correlation scene type;
[0112] S64: Obtain a determination scene model according to the initial scene type, the associated scene type and the time information.
[0113] Among them, the initial scene type is the scene type confirmed by the key position information, the associated scene type is the scene type whose association with the initial scene type meets the requirements, and there is a certain degree of similarity between the environmental sound characteristics of the associated scene type and the environmental sound characteristics of the initial scene type. The association value reflects the similarity between the sound characteristics of each scene type, and the association value can be obtained in advance by comparing the sound characteristics of the scene model in each scene type and undergoing a large amount of scene model training. The first threshold is a preset comparison value, which is mainly used to screen the association value to obtain the associated scene type that meets the requirements.
[0114] The initial scene type is determined by the adaptation value between the key position information and each preset scene type. Specifically, the maximum adaptation value can be obtained by comparing the sizes of each adaptation value, and the scene model corresponding to the maximum adaptation value can be used as the initial scene type; or each adaptation value can be compared with a preset comparison value to screen out the adaptation value that meets the requirements, and the scene model corresponding to the adaptation value that meets the requirements can be used as the initial scene type. It is easy to understand that based on the above statement, the number of initial scene types obtained may be different depending on the method of obtaining the initial scene type using the key position information, that is, the number of initial scene types may be one or several, and the method of determining the required initial scene type can be selected according to the actual design requirements.
[0115] After the initial scene type is determined, the association value between the initial scene type and other scene types can be obtained. In the process of screening out the associated scene types based on the comparison of the association value and the first threshold, the situation of the first threshold will affect the determination of the associated scene type. The first threshold can be set by setting a fixed value or by establishing a database of the first threshold, and then calling a suitable first threshold based on the specific situation. It can be understood that through the key position information, the pre-set scene types are preliminarily screened to obtain the initial scene type, and then combined with the comparison of the association values of the initial scene type and other scene types, the associated scene types with similarities are obtained, which ensures the efficiency and accuracy of subsequent scene recognition, and can effectively reduce the number of scene types for subsequent scene recognition, reduce unnecessary calculations, and reduce the power consumption of the headphones.
[0116] After determining the initial scene type and the associated scene type, the scene models in the initial scene type and the associated scene type can be obtained, and then divided and screened in combination with the time information to obtain the scene models corresponding to the initial scene type and the associated scene type in the current time period, and then obtain the judgment scene model for subsequent matching and comparison. Specifically, first obtain the specific scene model based on the initial scene type and the associated scene type; then filter the judgment scene model from the specific scene model according to the time information.
[0117] It can be understood that by combining time information to screen and obtain the judgment scene model for matching and comparison, the number of scene models used for matching and comparison can be effectively reduced, the efficiency and accuracy of scene recognition of adaptive noise reduction can be further improved, the workload of subsequent matching and comparison can be reduced, and the power consumption of the headphones can be reduced to a certain extent.
[0118] Further, in a preferred embodiment, the step of determining the initial scene type from preset scene types according to the key position information includes:
[0119] Obtaining the adaptation value of key location information and preset scene type;
[0120] From the adaptation values, obtain the maximum adaptation value;
[0121] According to the maximum adaptation value, the corresponding scene type is obtained as the initial scene type.
[0122] Among them, the adaptation value reflects the degree of adaptation between the key position information and the scene type. The larger the adaptation value, the higher the degree of adaptation between the scene type and the key position information; the smaller the adaptation value, the lower the degree of adaptation between the scene type and the key position information.
[0123] Specifically, each scene type is pre-set with a corresponding trigger information element. After obtaining the key position information, the key position information can be matched with the trigger information element of each scene type, and then the adaptation value corresponding to each scene type can be output.
[0124] It is easy to understand that by obtaining the maximum adaptation value from the obtained adaptation values and obtaining the initial scene type in a preferential manner, the number of scene types subsequently used for scene recognition is reduced while ensuring the accuracy of subsequent scene recognition, which can improve the efficiency of scene recognition to a certain extent.
[0125] For further information, please refer to Figure 5 In one embodiment, based on the comparison between the association value and the preset first threshold, the step S63 of screening out the associated scene type includes:
[0126] S631, obtaining a maximum adaptation value between key location information and a preset scene type;
[0127] S632: Filter and obtain a preset adaptation value interval in which the maximum adaptation value falls, and obtain a first threshold corresponding to the adaptation value interval;
[0128] S633, determining whether the correlation value is less than a first threshold;
[0129] S634: If not, mark the scene type corresponding to the associated value as an associated scene type.
[0130] Each adaptation value interval corresponds to a first threshold value, and the number of adaptation value intervals and first threshold values can be predetermined according to design requirements.
[0131] Specifically, after obtaining the maximum adaptation value, the adaptation value interval in which the maximum adaptation value falls can be determined first, and then the corresponding first threshold value can be obtained according to the adaptation value interval in which it falls. When the maximum adaptation value is relatively high, a first threshold value with a relatively large value is obtained, and when the maximum adaptation value is relatively low, a first threshold value with a relatively small value is obtained.
[0132] It is easy to understand that by obtaining the corresponding first threshold based on the maximum adaptation value, it is possible to dynamically adjust the difficulty of obtaining associated scene types according to the degree of adaptation of the key position information and the scene type, thereby controlling the number of associated scene types obtained according to the adaptation value. That is, when the maximum adaptation value is relatively high, the size of the first threshold can be appropriately increased, which can effectively reduce the amount of associated scene models obtained, thereby improving the efficiency of subsequent scene recognition; when the maximum adaptation value is relatively low, the size of the first threshold can be appropriately reduced, and relatively more associated scene models can be obtained, thereby ensuring the accuracy of subsequent scene recognition.
[0133] For further information, please refer to Figure 6 In one embodiment, the scene recognition model uses a Transformer scene recognition model, and the step S7 of obtaining the corresponding noise reduction parameters according to the ambient audio signal, the determination scene model, the preset scene recognition model and the preset recognition weight data includes:
[0134] S71, performing feature extraction processing on the ambient audio signal to obtain an ambient audio spectrum;
[0135] S72, based on a preset Transformer scene recognition module, comparing the ambient audio spectrum with the determination scene model to obtain a corresponding matching value;
[0136] S73. Obtain a maximum matching value from the matching values, and obtain a noise reduction parameter corresponding to the determination scene model corresponding to the maximum matching value.
[0137] After obtaining the ambient audio signal in the time domain, the ambient audio signal is first framed and windowed to obtain a more continuous ambient audio signal, and then the ambient audio signal is short-time Fourier transformed to obtain a linear spectrum of the ambient audio in the frequency domain. The linear spectrum of the ambient audio in the frequency domain is then input into the Mel filter group to obtain a Mel nonlinear spectrum, and then the Mel nonlinear spectrum is logarithmically processed to obtain an ambient audio spectrum that can be used for scene recognition.
[0138] In the process of frame segmentation and windowing, the frame length of each frame can be 15ms to 30ms, preferably 25ms. The window function can be a rectangular window, a Hamming window, etc., which is not limited here. Preferably, the window function is a Hamming window.
[0139] Specifically, a Hamming window of 25 ms is used, and frames are extracted every 10 ms, that is, adjacent frames overlap by 15 ms. It can be understood that the frame processing method can smoothly process sudden change signals, avoid leakage, and improve the continuity of environmental audio signals.
[0140] The linear spectrum of the ambient audio obtained after the short-time Fourier transform of the ambient audio signal is not enough to reflect the characteristics of human auditory perception, so it is necessary to pass through the Mel filter group to obtain the Mel nonlinear spectrum. It can be understood that through the filtering effect of the Mel filter group, the frequency components that do not match the human auditory perception can be filtered out, and the frequency components that match the human auditory perception can be retained, which can improve the accuracy of subsequent scene recognition. The required ambient audio spectrum can be obtained by logarithmically processing the Mel nonlinear spectrum. It can be understood that logarithmically processing the Mel nonlinear spectrum can reduce the calculation amount of subsequent scene recognition to a certain extent, thereby improving the efficiency of scene recognition.
[0141] Furthermore, the recognition weight data is used as the composition data of the attention baseline focus of the compiled semantics in the Transformer scene recognition model. After obtaining the ambient audio spectrum and inputting it into the Transformer scene recognition model, the Transformer scene recognition model extracts features from the ambient audio spectrum based on the recognition weight data, and then obtains the ambient audio high-order feature sequence. The Transformer scene recognition model calls the required judgment scene model, matches and compares the ambient audio high-order feature sequence with the judgment scene model, and then obtains the matching value corresponding to each judgment scene model based on the calculation of the linear classifier.
[0142] After obtaining the matching values corresponding to each judgment scene model, the maximum matching value is obtained from the matching values to obtain the maximum matching value, and then the corresponding judgment scene model is obtained according to the maximum matching value as the target scene model; based on the target scene model, the corresponding noise reduction parameters are obtained.
[0143] It is understandable that by processing the ambient audio signals, calling the judgment scene model and selecting the Transformer scene recognition model, the number of scene models that need to be recognized and judged in the scene recognition process can be effectively reduced, thereby reducing unnecessary calculations and improving the efficiency and accuracy of scene recognition, thereby improving the accuracy of the adaptive noise reduction of the headphones, reducing the power consumption of the headphones to a certain extent, and improving the user experience.
[0144] To summarize, the present application optimizes the method of adaptively acquiring noise reduction parameters based on a preset trigger order by setting specific areas and optimizing the method of acquiring judgment scene models. First, by determining whether the headset is in a specific area, the frequency of triggering scene recognition for adaptive noise reduction can be effectively reduced, and at the same time, users can obtain specific noise reduction parameters that are more in line with their actual usage needs; after the headset is not in a specific area, the initial position information and the associated position information are screened out based on the key position information, and then combined with the time information to obtain the judgment scene model, and then combined with the scene recognition model and the recognition weight data, the matching value of each judgment scene model and the current scene of the headset can be obtained, and then the noise reduction parameters required for adaptive noise reduction can be obtained based on the matching value, which can effectively reduce the number of scene models used for matching and comparison, while improving the efficiency and accuracy of scene recognition for adaptive noise reduction, reducing the amount of calculation for subsequent matching and comparison, reducing the power consumption of the headset to a certain extent, and improving the user experience.
[0145] The following is the content of the second aspect of the present invention:
[0146] The present invention provides an earphone, such as Figure 7 As shown, the headset includes a memory 10, a processor 20, and a method program instruction 30 for adaptive active noise reduction stored in the memory 10 and executable on the processor 20. When the method program instruction 30 for adaptive active noise reduction is executed by the processor 20, the aforementioned method for adaptive active noise reduction is implemented.
[0147] In some embodiments, the processor may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor is generally used to control the overall operation of the headset. In this embodiment, the processor is used to run program codes stored in a readable storage medium or process data.
[0148] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a readable storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods of various embodiments of the present invention.
[0149] The following is the content of the third aspect of the present invention:
[0150] The present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned adaptive active noise reduction method are implemented.
[0151] The above is only a specific implementation method of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A method for adaptive active noise reduction, characterized in that: include: Get the current location and time information of the headset; Based on the location information, determining whether the headset is located in a preset specific area; If yes, then obtaining preset specific noise reduction parameters in combination with the time information; If not, determining whether the location information contains key location information; If yes, then obtaining the key position information and the ambient audio signal; Based on the key position information and the time information, obtaining a determination scene model; Corresponding noise reduction parameters are obtained according to the environmental audio signal, the determination scene model, the preset scene recognition model and the preset recognition weight data.
2. The method of adaptive active noise reduction according to claim 1, characterized in that: The step of determining whether the headset is located in a preset specific area based on the location information includes: Obtaining the positioning point in the location information, the specific point and the range radius of the specific area; Calculating the distance between the positioning point and the specific point to obtain a relative distance; Determining whether the relative distance is greater than the range radius; If not, the earphone is located in a preset specific area.
3. The method of adaptive active noise reduction according to claim 1, characterized in that: The step of obtaining a preset specific noise reduction parameter in combination with the time information includes: Obtaining a preset time interval corresponding to the specific area; According to the time information, a time interval in which the time information is located is obtained by screening as a specific time interval; According to the specific time interval, a corresponding specific noise reduction parameter is obtained.
4. The method for adaptive active noise reduction according to claim 1, characterized in that: The step of determining whether the headset is within a preset specific area based on the location information includes: Detect whether a specific noise reduction mode is activated; If yes, determining whether the headset is located in a preset specific area based on the location information; If not, it is determined whether the location information contains key location information.
5. The method of adaptive active noise reduction according to claim 1, characterized in that: The step of acquiring a determination scene model based on the key position information and the time information comprises: Determining an initial scene type from preset scene types according to the key position information; Obtaining association values between the initial scene type and other scene types; Based on the comparison between the association value and a preset first threshold, filtering out the associated scene type; The determination scene model is obtained according to the initial scene type, the associated scene type and the time information.
6. The method for adaptive active noise reduction according to claim 5, characterized in that: The step of determining the initial scene type from preset scene types according to the key position information includes: Obtaining an adaptation value between the key location information and a preset scene type; Obtaining a maximum adaptation value from the adaptation values; According to the maximum adaptation value, a corresponding scene type is obtained as the initial scene type.
7. The method of adaptive active noise reduction according to claim 5, characterized in that: The step of screening out the associated scene type based on the comparison between the associated value and a preset first threshold comprises: Obtaining a maximum adaptation value between the key position information and a preset scene type; Screening to obtain a preset adaptation value interval in which the maximum adaptation value falls, and obtaining a first threshold value corresponding to the adaptation value interval; Determining whether the correlation value is less than the first threshold; If not, the scene type corresponding to the associated value is marked as the associated scene type.
8. The method of adaptive active noise reduction according to claim 1, characterized in that: The step of acquiring corresponding noise reduction parameters according to the ambient audio signal, the determination scene model, the preset scene recognition model and the preset recognition weight data comprises: Performing feature extraction processing on the ambient audio signal to obtain an ambient audio spectrum; Based on a preset Transformer scene recognition module and recognition weight data, the ambient audio spectrum is compared with the determination scene model to obtain a corresponding matching value; A maximum matching value is obtained from the matching values, and a noise reduction parameter corresponding to the determination scene model corresponding to the maximum matching value is obtained.
9. A headset comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method for adaptive active noise reduction according to any one of claims 1 to 8 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method for adaptive active noise reduction according to any one of claims 1 to 8 are implemented.