Intelligent noise reduction mode switching method and system

By working in tandem with noise level assessment and environmental scene recognition models, the noise reduction mode is adjusted in real time, solving the problem of poor noise reduction effect in static settings and achieving high-quality recording and personalized user experience in complex environments.

CN120375871BActive Publication Date: 2025-11-14SHENZHEN LONGXINWEI SEMICON TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510873606.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-11-14
Estimated Expiration
2045-06-27

AI Technical Summary

Technical Problem

The existing static settings for noise reduction parameters cannot meet the personalized and diverse needs of users, and cannot be adjusted in real time according to scene changes, resulting in poor noise reduction effect in complex and ever-changing environments, affecting recording quality and user experience.

Method used

By working together with a noise level assessment model and an environmental scene recognition model, environmental data is acquired in real time, noise levels and scene types are analyzed, and noise reduction modes are automatically switched to adapt to different environments and user needs.

Benefits of technology

It enables precise adjustment of noise reduction modes in complex environments, ensuring that recording quality is not affected by environmental changes, improving user experience and device adaptability, and meeting personalized needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120375871B_ABST
    Figure CN120375871B_ABST
Patent Text Reader

Abstract

This invention relates to the field of noise reduction technology, and in particular to an intelligent noise reduction mode switching method and system. The method involves: acquiring real-time environmental data of the user device's current environment after recording begins; analyzing the real-time environmental data using a noise level assessment model to obtain the real-time noise level of the user device's current environment; correlating the real-time noise level with the amplitude of noise signal changes in the real-time environmental data; inputting the real-time environmental data into an environmental scene recognition model to obtain the real-time scene type corresponding to the user device's current environment; and automatically switching the user device to the corresponding target noise reduction mode based on the real-time noise level and the real-time scene type to ensure the recording quality of the user device. By analyzing environmental data from different perspectives through the noise level assessment model and the environmental scene recognition model, environmental adaptability is improved, noise reduction effect is enhanced, and recording quality is ensured to be unaffected by environmental changes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of noise reduction technology, and in particular to an intelligent noise reduction mode switching method and system. Background Technology

[0002] In many current audio processing devices and systems, noise reduction has become a key factor in improving user experience. However, the fact that noise reduction parameters are mostly statically set has led to a series of problems that urgently need to be solved, severely restricting the optimization of noise reduction effects and the improvement of user experience.

[0003] Different scenarios present vastly different acoustic environments. In a quiet library, ambient noise primarily originates from the soft sounds of turning pages, whispers, and the faint hum of equipment. In this case, noise reduction focuses on precisely suppressing these minute noises while preserving as much detail as possible of the original sound, ensuring that reading and note-taking sounds are clearly identifiable. However, on bustling streets, the noise is intense, with cars roaring, horns blaring, crowds clamoring, and various vehicles clashing. This noise is not only high-intensity but also has a wide frequency range and changes rapidly. In factory workshops, the high-intensity, regular low-frequency noise generated by continuously running machines dominates. Faced with such diverse scenarios, statically set noise reduction parameters are insufficient to meet the specific needs of different environments. Fixed noise reduction parameters cannot effectively adapt to the unique frequency characteristics, intensity variations, and dynamic range of noise in different scenarios. This leads to over-noise reduction in some scenarios, causing severe sound distortion and loss of important information; while in other scenarios, insufficient noise reduction results in noise still significantly interfering with recording or voice communication quality.

[0004] Real-world environments are often in a state of dynamic change. For example, when a user is riding the subway, the ambient noise increases dramatically as they move from a relatively quiet platform into a noisy carriage; or while walking outdoors, the noise intensity and frequency characteristics change drastically as they move from an open plaza to a construction site. In these rapidly changing scenarios, static noise reduction parameters cannot respond in time, leading to significant deviations in noise reduction performance at the moment of scene transition. At the moment of increased noise, because the noise reduction parameters fail to adjust synchronously, the noise enters the audio acquisition equipment without obstruction, severely impacting recording quality or speech clarity; conversely, when the noise suddenly decreases, the previously high-intensity noise reduction settings may overprocess the audio signal, making the sound hollow and unnatural.

[0005] In related technologies, static noise reduction parameter settings cannot meet the personalized and diverse needs of users. Different users have significantly different perceptions and preferences for sound. Some users are extremely sensitive to noise and want to achieve the ultimate noise reduction effect in any scenario; while others value preserving the ambiance of ambient sound to better blend into their surroundings. Furthermore, users' noise reduction needs vary depending on the usage scenario, such as studying, working, or entertaining. However, existing static noise reduction parameter settings cannot be flexibly adjusted according to individual user differences and different usage scenarios, preventing users from obtaining the optimal audio experience based on their own needs and significantly reducing user satisfaction and loyalty to the device or system.

[0006] Clearly, the static settings of existing noise reduction parameters severely limit the adaptability, real-time performance, and ability to meet users' personalized needs in complex and ever-changing environments. There is an urgent need for a solution that can modify noise reduction parameters in real time according to scene changes in order to improve the quality of audio processing and user experience. Summary of the Invention

[0007] This invention addresses the technical problems existing in the prior art by providing an intelligent noise reduction mode switching method and system. It is used to analyze environmental data from different perspectives through the collaborative work of a noise level assessment model and an environmental scene recognition model, effectively improving the adaptability of user equipment to the environment, helping user equipment maintain the optimal noise reduction effect, and ensuring that the recording quality is not affected by environmental changes.

[0008] In a first aspect, embodiments of this application provide an intelligent noise reduction mode switching method, including:

[0009] After recording is started, real-time environmental data of the user device’s current environment is acquired; the real-time environmental data includes at least: ambient sound data, spatial location data, device usage status, ambient temperature and humidity data, air pressure data, and electromagnetic interference intensity data.

[0010] The real-time environmental data is analyzed using a noise level assessment model to obtain the real-time noise level of the user equipment's current environment; the real-time noise level is correlated with the amplitude of noise signal changes in the real-time environmental data.

[0011] The real-time environmental data is input into the environmental scene recognition model to obtain the real-time scene type corresponding to the current environment of the user device; the real-time scene type is determined based on the real-time spatial changes reflected by any one or more of the spatial location data, device usage status, environmental temperature and humidity data, air pressure data, and electromagnetic interference intensity data, as well as the noise signal change trend in the environmental sound data.

[0012] Based on the real-time noise level and the real-time scene type, the user device is automatically switched to the corresponding target noise reduction mode to ensure the recording quality of the user device.

[0013] Secondly, embodiments of this application provide an intelligent noise reduction mode switching system, which includes the following units:

[0014] The acquisition unit is used to acquire real-time environmental data of the user device's current environment after recording is started; the real-time environmental data includes at least: ambient sound data, spatial location data, device usage status, ambient temperature and humidity data, air pressure data, and electromagnetic interference intensity data.

[0015] The analysis unit is used to perform noise analysis on the real-time environmental data through a noise level assessment model to obtain the real-time noise level in the environment where the user equipment is currently located; the real-time noise level is related to the amplitude of noise signal changes in the real-time environmental data.

[0016] The identification unit is used to input the real-time environmental data into the environmental scene identification model to obtain the real-time scene type corresponding to the current environment of the user device; the real-time scene type is determined based on the real-time spatial changes reflected by any one or more of the spatial location data, device usage status, environmental temperature and humidity data, air pressure data, and electromagnetic interference intensity data, as well as the noise signal change trend in the environmental sound data.

[0017] The switching unit is used to automatically switch the user device to the corresponding target noise reduction mode based on the real-time noise level and the real-time scene type to ensure the recording quality of the user device.

[0018] Thirdly, embodiments of this application provide an electronic device, the electronic device comprising:

[0019] At least one processor, memory, and input / output unit;

[0020] The memory is used to store computer programs, and the processor is used to call the computer programs stored in the memory to execute the intelligent noise reduction mode switching method of the first aspect.

[0021] Fourthly, a computer-readable storage medium is provided, comprising instructions that, when executed on a computer, cause the computer to perform the intelligent noise reduction mode switching method of the first aspect.

[0022] The beneficial effects of this invention are: it provides an intelligent noise reduction mode switching method and system. In this technical solution, after recording begins, real-time environmental data of the user device's current environment is acquired; the real-time environmental data includes at least: ambient sound data, spatial location data, device usage status, ambient temperature and humidity data, air pressure data, and electromagnetic interference intensity data. Then, the real-time environmental data is analyzed using a noise level assessment model to obtain the real-time noise level of the user device's current environment; the real-time noise level is correlated with the amplitude of noise signal changes in the real-time environmental data. Next, the real-time environmental data is input into an environmental scene recognition model to obtain the real-time scene type corresponding to the user device's current environment; the real-time scene type is determined based on the real-time spatial changes reflected by any one or more of the spatial location data, device usage status, ambient temperature and humidity data, air pressure data, and electromagnetic interference intensity data, as well as the noise signal change trend in the ambient sound data. Finally, based on the real-time noise level and the real-time scene type, the user device is automatically switched to the corresponding target noise reduction mode to ensure the recording quality of the user device.

[0023] In this embodiment, the noise level assessment model and the environmental scene recognition model work together to analyze environmental data from different perspectives, effectively improving the adaptability of user equipment to the environment, helping user equipment maintain the best noise reduction effect, and ensuring that the recording quality is not affected by environmental changes. Attached Figure Description

[0024] Figure 1 This is a flowchart illustrating an intelligent noise reduction mode switching method according to an embodiment of this application;

[0025] Figure 2 This is a schematic diagram of the structure of an intelligent noise reduction mode switching system according to an embodiment of this application;

[0026] Figure 3 This is a schematic diagram of the structure of an electronic device according to an embodiment of this application;

[0027] Figure 4 This is a schematic diagram of the structure of a media device according to an embodiment of this application. Detailed Implementation

[0028] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0029] In the description of this application, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the stated features. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.

[0030] In the description of this application, the term "for example" is used to mean "used as an example, illustration, or description." Any embodiment described as "for example" in this application is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use the invention. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that the invention can be made without using these specific details. In other instances, well-known structures and processes will not be described in detail to avoid obscuring the description of the invention with unnecessary detail. Therefore, the invention is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed in this application.

[0031] This application provides an intelligent noise reduction mode switching method and system. In this technical solution, after recording begins, real-time environmental data of the user device's current environment is acquired. This real-time environmental data includes at least: ambient sound data, spatial location data, device usage status, ambient temperature and humidity data, air pressure data, and electromagnetic interference intensity data. Then, the real-time environmental data is analyzed using a noise level assessment model to obtain the real-time noise level of the user device's current environment. The real-time noise level is correlated with the amplitude of noise signal changes in the real-time environmental data. Next, the real-time environmental data is input into an environmental scene recognition model to obtain the real-time scene type corresponding to the user device's current environment. The real-time scene type is determined based on the real-time spatial changes reflected by any one or more of the spatial location data, device usage status, ambient temperature and humidity data, air pressure data, and electromagnetic interference intensity data, as well as the noise signal change trend in the ambient sound data. Finally, based on the real-time noise level and the real-time scene type, the user device is automatically switched to the corresponding target noise reduction mode to ensure the recording quality of the user device.

[0032] In this embodiment, the noise level assessment model and the environmental scene recognition model work together to analyze environmental data from different perspectives, effectively improving the adaptability of user equipment to the environment, helping user equipment maintain the best noise reduction effect, and ensuring that the recording quality is not affected by environmental changes.

[0033] The intelligent noise reduction mode switching scheme provided in this application embodiment can also be executed by an electronic device, which can be a server, server cluster, or cloud server. The electronic device can also be a terminal device such as a mobile phone, computer, tablet computer, wearable device, or dedicated device (such as a dedicated terminal device with an intelligent noise reduction mode switching system). These electronic devices can also carry the chip described in the above embodiments. Alternatively, these electronic devices can also install a service program for executing the intelligent noise reduction mode switching scheme.

[0034] Figure 1 This is a schematic diagram of an intelligent noise reduction mode switching method provided in an embodiment of this application, as shown below. Figure 1 As shown, the method includes the following steps:

[0035] 101. After recording is started, obtain real-time environmental data of the user's current environment;

[0036] 102. The real-time environmental data is analyzed using a noise level assessment model to obtain the real-time noise level in the environment where the user equipment is currently located; the real-time noise level is correlated with the amplitude of noise signal changes in the real-time environmental data.

[0037] 103. Input the real-time environmental data into the environmental scene recognition model to obtain the real-time scene type corresponding to the current environment of the user device;

[0038] 104. Based on the real-time noise level and the real-time scene type, automatically switch the user device to the corresponding target noise reduction mode to ensure the recording quality of the user device.

[0039] In this embodiment of the application, the real-time scene type is determined based on the real-time spatial changes reflected by any one or more of the spatial location data, device usage status, ambient temperature and humidity data, air pressure data, and electromagnetic interference intensity data, as well as the noise signal change trend in the ambient sound data.

[0040] For example, a user activates the recording function while traveling by high-speed rail. At this time, the spatial location data shows that the device is moving at high speed, and the changes in latitude and longitude show a trend consistent with the high-speed rail's route.

[0041] The device's usage status indicates that it is placed on a small table and is in recording mode. The ambient temperature and humidity data are relatively stable, with the temperature maintained at 26 degrees Celsius, the humidity at approximately 40%, and the air pressure data also stable at around 101 kPa.

[0042] Electromagnetic interference intensity data indicates the presence of some electronic equipment interference, which is consistent with the operation of numerous electronic devices within high-speed train carriages. Environmental sound data shows a continuous low-frequency rumbling sound, interspersed with the rhythmic sound of the train rubbing against the tracks, and occasional fluctuations in the sounds of train attendants' announcements and passengers' conversations. Specific analysis methods are detailed below and will not be elaborated upon here.

[0043] Based on this real-time environmental data, the system determined the real-time scene type to be inside a high-speed train carriage. The rapid movement and specific trajectory of the spatial location data, combined with the device's placement, initially pointed to a scene inside a means of transportation. The stability of environmental temperature and humidity, air pressure, and specific electromagnetic interference intensity further narrowed down the possibilities. The changing trends of noise signals in the environmental sound data, such as continuous low-frequency rumbling and regular friction sounds, became key evidence for confirming the scene as a high-speed train carriage, thus achieving accurate scene recognition and laying the foundation for subsequently matching an appropriate noise reduction mode.

[0044] In the above steps, on the one hand, by analyzing real-time environmental data, the real-time noise level and scene type can be accurately obtained. Based on this, the system can automatically select and switch to the most suitable target noise reduction mode for different noise conditions and scenes. For example, in a noisy construction site, the system will automatically turn on the strong noise reduction mode, effectively filtering out a large amount of high-intensity environmental noise, ensuring that the recorded sound is clearly distinguishable; while in quiet scenes such as libraries, turning off the noise reduction function can avoid audio distortion introduced by the noise reduction algorithm, completely preserving every detail of the original sound, thereby comprehensively improving the recording quality and meeting users' needs for high-quality recording in various complex environments. Moreover, the real-time noise level and environmental scene type are dynamically obtained based on real-time environmental data. This means that as the user's environment changes, such as moving from a quiet indoor environment to a noisy outdoor street, the system can promptly capture changes in noise level and scene and quickly adjust the noise reduction mode. Compared to traditional fixed noise reduction modes or noise reduction methods that require manual switching, this method can always maintain the optimal noise reduction effect, ensuring that the recording quality is not affected by environmental changes.

[0045] Furthermore, in this embodiment, the user does not need to manually adjust the noise reduction mode during recording; the entire analysis, judgment, and switching process can be completed automatically. This greatly simplifies the operation process and lowers the user's barrier to entry, providing significant convenience, especially for users unfamiliar with audio technology or noise reduction functions. Users only need to focus on the recording content itself, without being distracted by noise reduction-related operations, significantly improving the user's experience and focus during recording. When determining the target noise reduction mode and generating noise reduction mode parameters, in addition to considering real-time noise levels and scene types, personalized customization can also be achieved by incorporating user attributes. For example, users of different ages and professions have different sensitivities and needs for sound; the system can tailor the most suitable noise reduction mode for each user based on their historical usage data and preference settings. This personalized service not only improves user satisfaction with the product but also enhances user loyalty.

[0046] On the other hand, the embodiments of this application rely on real-time environmental data for analysis and decision-making, fully demonstrating the intelligent characteristics of data-driven approaches. Through the collection and analysis of large amounts of environmental data, the noise level assessment model and the environmental scene recognition model can continuously optimize and improve their accuracy and adaptability. With the increase in data volume and the continuous iteration of the model, the system can make more accurate judgments and responses to various complex environments, realizing the transformation from a simple mode to intelligent environmental perception and adaptive processing.

[0047] Furthermore, this embodiment of the application effectively integrates the intelligent advantages of multiple models through the collaborative work of a noise level assessment model and an environmental scene recognition model. Environmental data is analyzed from different perspectives, and the results are then synthesized, providing a comprehensive and accurate basis for determining the target noise reduction mode. This model fusion approach not only improves the system's intelligence level but also provides a useful reference for the design and development of other intelligent systems.

[0048] It is worth noting that the embodiments of this application can be widely applied in multiple fields. In the education field, it can be used for classroom recording, lecture recording, and other scenarios to ensure that students can clearly record the teacher's lecture content; in the business field, it can meet the needs of meeting recording, interview recording, and other scenarios, providing reliable support for business communication and information recording; in the medical field, it can be used for doctor-patient communication recording, medical record recording, and other scenarios, helping to improve the quality and efficiency of medical services. Through its application in different fields, its technical and social value is further expanded.

[0049] This application's embodiments can be integrated with various user devices, including mobile phones, voice recorders, tablets, and professional recording equipment. Simply integrating this intelligent noise reduction mode switching system enables intelligent upgrades to the noise reduction function. Furthermore, with continuous technological advancements and the emergence of new environmental data types, the system can easily expand and upgrade its functionality by updating the noise level assessment model and environmental scene recognition model to adapt to more complex environments and diverse user needs.

[0050] As an optional embodiment, a noise reduction mode adjustment command issued by the user device can also be received; the noise reduction mode adjustment command is constructed based on parameter adjustment information input by the user in the user device. Furthermore, based on the noise reduction mode adjustment command and the historical scene type of the user device, a user individual preference library corresponding to the user device is established; wherein, the user individual preference library stores parameter adjustment strategies corresponding to the user device under various historical scene types. Next, after automatically switching the user device to the corresponding target noise reduction mode using the noise reduction mode parameters in step 104, a corresponding noise reduction mode parameter adjustment strategy can be selected from the user individual preference library, and based on the selected parameter adjustment strategy, the noise reduction mode parameters in the target noise reduction mode are adjusted to ensure that the recording quality matches the user's individual preferences.

[0051] In the above embodiments, by collecting user-initiated noise reduction mode adjustment commands, the system gains a deeper understanding of users' personalized needs for noise reduction modes in different scenarios. These commands are constructed based on parameter adjustment information input by users on their devices, reflecting users' preferences for noise reduction intensity, sound detail preservation, and other aspects. Simultaneously, this information is compiled and summarized into a user individual preference database, taking into account the historical scene types of the user's device. In this database, each historical scene type corresponds to a specific set of parameter adjustment strategies, which are a concentrated expression of user preferences. When the system automatically switches the user's device to the target noise reduction mode based on the real-time noise level and scene type, it further retrieves the parameter adjustment strategy matching the current noise reduction mode from the user individual preference database. Then, based on this strategy, the parameters of the target noise reduction mode are optimized and adjusted.

[0052] Through the above embodiments, it is evident that traditional noise reduction modes are often generic and cannot meet the diverse needs of different users. However, the embodiments described above can personalize the noise reduction mode based on the user's past operations and preferences, ensuring that the recording quality matches the user's unique auditory requirements. Whether the user seeks ultimate quietness or desires to retain more environmental sound details, these preferences can be met. Secondly, this enhances the product's adaptability and competitiveness. Among the numerous recording devices on the market, products that offer personalized services are more likely to stand out. By meeting users' personalized needs, user satisfaction and loyalty are increased. Furthermore, it facilitates continuous product optimization. By analyzing data from individual user preference databases, product developers can identify general user preference trends, thereby enabling targeted improvements to noise reduction algorithms and functional settings in subsequent product upgrades, allowing the product to continuously evolve and better adapt to market demands.

[0053] After recording is started, in step 101, real-time environmental data of the user device's current environment is obtained.

[0054] In this embodiment, the real-time environmental data includes at least: ambient sound data, spatial location data, device usage status, ambient temperature and humidity data, air pressure data, and electromagnetic interference intensity data. For example, in an outdoor park scene in the early morning, the recording device collects real-time environmental data after being turned on. Ambient sound data includes birdsong, the rustling of leaves in the breeze, and the conversations of people exercising in the distance; spatial location data shows that the device is near a lawn in the park, with latitude and longitude coordinates [specific coordinate values]; device usage status indicates that it is in handheld recording mode, with screen brightness at 50% and volume at default values; ambient temperature and humidity data show a temperature of 20 degrees Celsius and relative humidity of 60%; air pressure data shows a relatively stable 101.325 kPa; electromagnetic interference intensity data shows a slight interference from nearby streetlight circuits, but the overall data is within the normal range. This data is collected in real time, providing a comprehensive and rich information foundation for subsequent noise level assessment and environmental scene recognition, facilitating accurate switching of intelligent noise reduction modes, ensuring that recording quality is not excessively affected by environmental factors, making the recorded sound clear and realistic, and providing users with a good recording experience.

[0055] As an optional embodiment, the noise level assessment model includes at least an extraction layer, an analysis layer, and an output layer. Based on this structure, in step 102, the real-time environmental data is analyzed using the noise level assessment model to obtain the real-time noise level of the user equipment's current environment, including:

[0056] An extraction layer extracts environmental audio features from the real-time environmental data; an analysis layer performs real-time signal amplitude analysis on the environmental audio features to obtain the corresponding environmental noise variation amplitude features; and an output layer classifies the analyzed environmental noise variation amplitude features according to a preset strategy to obtain the real-time noise level.

[0057] For example, suppose a user is in a bustling market. After turning on the recording device, real-time environmental data is fed into a noise level assessment model. The extraction layer first filters out environmental audio features from the real-time environmental data, which contains various information. Audio information such as the sounds of vendors hawking their wares, the hustle and bustle of the crowd, and vehicle horns are all accurately extracted.

[0058] Next, the analysis layer begins its work, performing real-time signal amplitude analysis on these extracted environmental audio features. Through complex algorithms, it calculates in detail the amplitude changes of each sound signal, such as the amplitude difference between a sudden high-decibel hawking sound and the relatively calm conversation of a crowd, thereby obtaining the corresponding environmental noise amplitude characteristics and clearly presenting the fluctuations of noise at different times.

[0059] Finally, the output layer classifies the environmental noise variation characteristics obtained from the analysis layer according to a preset strategy. If the preset strategy divides the noise level into three levels: low, medium, and high, after comparison and judgment, the output layer will determine the current real-time noise level of the environment as "high" because the market environmental noise varies greatly and frequently.

[0060] From a technical perspective, this noise level assessment model can efficiently and accurately analyze environmental noise. The extraction layer ensures the acquisition of key audio information without being interfered with by other non-audio environmental data. The analysis layer deeply analyzes audio features, providing solid data support for accurately judging noise levels. The classification mechanism of the output layer transforms complex analysis results into intuitive and easy-to-understand noise level grades, facilitating the rapid switching of appropriate noise reduction modes based on different levels. This greatly improves the response speed and noise reduction effect of the entire intelligent noise reduction system, ensuring a high-quality recording experience for users in various complex environments.

[0061] As an optional embodiment, in the above steps, the environmental audio features are subjected to real-time signal amplitude analysis processing through the analysis layer to obtain the corresponding environmental noise change amplitude features, including:

[0062] The environmental audio features are subjected to noise separation processing to obtain multiple environmental noise signals; the separated multiple environmental noise signals are subjected to channel separation, and the peak points of each independent channel noise signal are identified to determine the peak points in each independent channel noise signal; the real-time signal amplitude of each independent channel noise signal is determined based on the peak points to obtain the corresponding environmental noise change amplitude features.

[0063] Specifically, ambient audio is typically a complex signal composed of various noises mixed together. Through noise separation processing, signal processing algorithms and models are used to decompose the mixed ambient audio features into multiple relatively independent ambient noise signals based on the characteristics of different noises, such as frequency and timbre, for subsequent more detailed analysis. Each separated ambient noise signal undergoes channel separation, further decomposing it into noise signals for each independent channel. Then, a specific algorithm is used to identify peak values ​​in each independent channel noise signal, finding the maximum value points, i.e., peak points. These peak points reflect the strongest amplitude of the noise signal in that channel at a specific moment. Based on the found peak points, the real-time signal amplitude of each independent channel noise signal at different times can be determined. The changes in these amplitude values ​​constitute the ambient noise variation amplitude characteristics, thus comprehensively and meticulously depicting the dynamic changes of ambient noise.

[0064] This allows for in-depth analysis of various noise components in ambient audio, accurately capturing the amplitude variations of each type of noise. This provides a detailed and reliable data foundation for accurately assessing ambient noise levels, making the system's understanding of ambient noise more precise. It can handle complex and varied ambient noise, effectively processing both mixtures of different noise types and variations in noise distribution across different channels, enhancing the system's adaptability and stability in various scenarios. Accurate amplitude characteristics of ambient noise variations help the intelligent noise reduction system more precisely match and adjust noise reduction modes, adopting more appropriate noise reduction strategies for noise of different amplitudes, thereby significantly improving noise reduction performance and recording quality.

[0065] Furthermore, in the above steps, channel separation is performed on the separated multi-channel environmental noise signals, and peak identification is performed on each of the separated independent channel noise signals to determine the peak points in each independent channel noise signal. This can be achieved as follows:

[0066] Multiple ambient noise signals are segmented using a sliding time window to obtain multiple ambient noise signal segments under different sliding time windows. These segments are then separated according to channel type to obtain individual channel noise signals for each ambient noise signal under each sliding time window. Channel types include at least left stereo, right stereo, and channels at different dry / wet sound levels. Quartile calculations are performed on each individual channel noise signal to obtain its quartile position and range. Anomaly detection is performed based on the quartile positions and ranges of each individual channel noise signal to obtain abnormal signal values. These abnormal signal values ​​are processed to obtain optimized individual channel noise signals. Hilbert transform and amplitude envelope calculations are then performed on the optimized individual channel noise signals to obtain their amplitude envelopes. Peak identification is performed based on the amplitude envelopes of each individual channel noise signal to obtain its peak points.

[0067] In the steps described above, a sliding time-series window is a technique for dividing a continuous time-series signal into multiple overlapping or non-overlapping segments. For multiple ambient noise signals, a sliding time-series window can divide them into multiple shorter signal segments. The size of the sliding window and the sliding step size can be adjusted according to specific needs. For example, if the window size is set to w samples and the sliding step size is s samples, then for a signal of length n, multiple signal segments of length w will be generated, with adjacent segments separated by s samples. This segmentation method helps to transform long-series signals into multiple short-series signals, facilitating subsequent processing and analysis, and enabling better capture of the local features of the signal at different time periods.

[0068] This is because environmental noise signals are usually continuous, but the characteristics of noise signals may differ in different time segments. By using a sliding time window, continuous signals can be localized in the time dimension, making it easier to analyze and process local signals.

[0069] This effectively reduces the complexity of signal processing by breaking down long signal sequences into multiple short segments, allowing subsequent processing to be performed within local time windows. This facilitates parallel or distributed processing and improves computational efficiency. It also helps capture local features of the signal within different time windows, avoiding the loss of local information that might be lost when processing the entire long signal sequence uniformly, thus providing a finer-grained data foundation for subsequent operations such as channel separation and anomaly detection.

[0070] Next, the multiple ambient noise signal segments under various sliding timing windows are separated according to channel type. Specifically, based on different channel types (such as left stereo channel, right stereo channel, and channels under different dry / wet tone levels), the multiple ambient noise signal segments under the segmented sliding timing windows are separated. This is based on the physical or acoustic characteristics of the channels; noise signals from different channels have different characteristics, and separation allows for targeted analysis of the noise signals from different channels. For stereo signals, the left and right channels may receive sound from different directions, resulting in different signal characteristics; for channels under different dry / wet tone levels, the frequency and amplitude characteristics of their signals will also differ. This separation facilitates subsequent individual processing of signals from different channels.

[0071] This enables independent analysis of different channels, allowing the system to more accurately identify noise characteristics in each channel, avoid mutual interference between signals from different channels, and improve the accuracy of noise analysis. Different post-processing strategies can be selected for different channels based on their characteristics. For example, different noise reduction algorithms or parameters may be used for channels at different dry and wet sound levels, providing more detailed information for the final noise reduction mode switching.

[0072] Then, quartiles are calculated for each independent channel noise signal to obtain the quartile positions and ranges for each channel. Quartiles are statistical measures that divide data into four equal parts. The first quartile, Q1, indicates that 25% of the data are less than this value, the third quartile, Q3, indicates that 75% of the data are less than this value, and IQR = Q3 - Q1. Calculating the quartiles for each independent channel noise signal helps to understand the data distribution. By sorting the noise signals for each channel, finding Q1 and Q3, and calculating the IQR, we can understand the central tendency and dispersion of the data. This method helps to detect the distribution range of the data and the potential range of outliers, because data exceeding the range of Q1 - 1.5 * IQR to Q3 + 1.5 * IQR may be outliers.

[0073] Therefore, the statistical characteristics of the vocal tract noise signal can be quickly understood through the quartile range, including the degree of signal dispersion and the central location of data, providing a basis for subsequent anomaly detection. This helps to identify possible outliers, prepare for subsequent processing of anomalous signals, and avoid interference from anomalous signals on the overall signal analysis.

[0074] Furthermore, anomaly detection is performed based on the quartile positions and quartile ranges corresponding to the noise signals of each independent channel, yielding the abnormal signal values ​​contained within each independent channel's noise signal. Specifically, based on the calculated Q1, Q3, and IQR, signal values ​​below Q1 - 1.5 * IQR or above Q3 + 1.5 * IQR are considered abnormal signal values. This is a statistical anomaly detection method. In noise signals, anomalies may be caused by sudden noise interference or equipment malfunctions, affecting the normal analysis of the noise signal. For example, in a normal ambient noise signal, a sudden appearance of extremely large or small signal amplitudes may be due to transient interference from equipment, which can be detected using this method.

[0075] This effectively filters out abnormal signals that may interfere with normal signal analysis, improving the accuracy of subsequent noise signal analysis and making subsequent processing and analysis more stable and reliable. It provides a basis for noise signal purification; by removing or correcting these abnormal signals, signal quality can be improved.

[0076] Furthermore, abnormal signal values ​​in each independent channel noise signal are processed to obtain optimized noise signals for each independent channel. Specifically, different processing methods can be used for detected abnormal signal values. For example, they can be replaced with Q1 or Q3, or interpolated using adjacent signal values, or they can be directly deleted. The goal is to eliminate the influence of abnormal signals on subsequent analysis, making the signal smoother and more normal. For example, if an abnormal signal value is much higher than the normal range, replacing it with Q3 can make the amplitude distribution of the signal more reasonable, avoiding the impact of abnormal values ​​on subsequent analysis and processing.

[0077] This approach yields smoother and more normal signals, avoiding interference from outliers in subsequent signal processing steps and improving the accuracy and stability of signal processing. It also helps improve the reliability of subsequent analysis and processing results based on these signals, providing a higher-quality data foundation for subsequent Hilbert transform and amplitude envelope calculations.

[0078] Furthermore, Hilbert transform and amplitude envelope calculations are performed on the optimized noise signals of each independent channel to obtain the amplitude envelope of each independent channel noise signal. It can be understood that the Hilbert transform is a method for converting a real signal into an analytic signal. For a real signal x(t), its Hilbert transform H(x(t)) can be calculated using an integral formula. The analytic signal z(t) = x(t) + j * H(x(t)), and the amplitude envelope A(t) can be calculated using A(t) = sqrt(x(t)^2 + H(x(t))^2). The Hilbert transform can extract the instantaneous amplitude information of a signal. For noise signals, its amplitude envelope reflects the amplitude variation trend of the noise signal. For environmental noise signals, its amplitude envelope can more clearly represent the amplitude variation of the noise signal at different time points and better reflect the actual energy variation of the noise. Especially for non-stationary noise signals, it can better capture the amplitude modulation characteristics of the signal.

[0079] This allows for the extraction of the amplitude envelope of the noise signal, which more accurately reflects the amplitude changes of the noise signal and provides more intuitive and accurate information for amplitude analysis. It also provides more effective data for subsequent peak identification because the amplitude envelope contains amplitude information of the noise signal and reflects the energy and intensity characteristics of the noise better than the original signal, thus helping to better understand the characteristics of the noise signal.

[0080] Finally, peak identification is performed based on the amplitude envelope of each independent channel noise signal to obtain the peak points in each independent channel noise signal. For the amplitude envelope signal of each channel, the peak point is the local maximum value in the amplitude envelope signal. By traversing the amplitude envelope signal, those points that are larger than their immediate neighbors are identified as peak points. These peak points reflect the maximum amplitude of the noise signal in that channel within a local time range and are important characteristic points of the noise signal intensity in that channel. For example, for an amplitude envelope [1, 2, 3, 2, 1], 3 is the peak point, representing the maximum amplitude of the noise signal within that time window. In this way, the peak points of the noise signal in each channel can be accurately found, providing key information for subsequent noise analysis, such as the maximum intensity of the noise and its timing. This helps in analyzing the characteristics of noise in different channels, provides a basis for peak point-based noise level assessment and noise reduction mode switching, improves the performance of the entire intelligent noise reduction system, and makes the selection of noise reduction modes and parameter adjustments more targeted and accurate.

[0081] Through the above series of steps, multiple environmental noise signals can be analyzed and processed in detail, from channel separation and anomaly detection to peak identification, providing comprehensive and accurate information for intelligent noise reduction systems, which helps to achieve more precise noise reduction mode switching and better noise reduction effect.

[0082] As an optional embodiment, it is assumed that the environmental scene recognition model includes at least: a preprocessing layer, a feature extraction layer, a construction and fusion layer, a prediction layer, and an output layer. Based on this, in step 103, the real-time environmental data is input into the environmental scene recognition model to obtain the real-time scene type corresponding to the current environment of the user device, including:

[0083] The real-time environmental data is preprocessed through a preprocessing layer; multi-dimensional real-time environmental features are extracted from the real-time environmental data through a feature extraction layer; the multi-dimensional real-time environmental features include at least: spatial location features, equipment usage status features, ambient temperature and humidity features, air pressure features, electromagnetic interference intensity features, and ambient sound features; a fusion layer is constructed to project the multi-dimensional real-time environmental features into a multi-dimensional analysis space as corresponding feature space points, and point cloud type matching is performed on the feature point cloud formed by the feature space points to determine the candidate scene type; a prediction layer is used to predict the scene adaptability of the candidate scene type to obtain the adaptability probability corresponding to the candidate scene type; and an output layer is used to determine the real-time scene type based on the adaptability probability corresponding to the candidate scene type.

[0084] In the above model, the preprocessing layer preprocesses the real-time environmental data, including but not limited to data cleaning, removing missing and outlier values. For example, values ​​in environmental temperature and humidity data that are significantly outside the reasonable range are corrected or deleted. Data normalization unifies data from different dimensions to the same units and value range, such as normalizing air pressure data and electromagnetic interference intensity data to the [0, 1] interval for subsequent model processing. Data encoding may also be performed, converting categorical data such as equipment usage status into numerical form for easier model recognition.

[0085] This improves data quality, making the data more suitable for model processing, reduces interference from outlier data, enhances model stability and accuracy, and avoids model learning bias caused by data inconsistency.

[0086] The feature extraction layer extracts various features from real-time environmental data. For example, it extracts spatial location features from spatial location data, such as latitude and longitude, altitude, and whether it is indoors; it extracts equipment usage status features from equipment usage status, such as whether the equipment is stationary or moving, and whether the screen is on; it extracts environmental temperature and humidity features from environmental temperature and humidity data, such as the rate of temperature change and the average humidity; it extracts air pressure features from air pressure data, such as the range of air pressure fluctuations; it extracts electromagnetic interference intensity features from electromagnetic interference intensity data, such as the frequency and intensity trend of interference; and it extracts environmental sound features from environmental sound data, such as the frequency distribution and energy spectrum of sound.

[0087] In this way, the raw environmental data is transformed into a feature representation that is valuable for scene recognition, highlighting the key information related to the scene in the data, providing a rich information foundation for subsequent scene recognition, and enabling the model to better distinguish different scenes based on these features.

[0088] A fusion layer is constructed to project multi-dimensional real-time environmental features into a multi-dimensional analysis space, forming feature space points, with each feature corresponding to one dimension in the space. These points constitute a feature point cloud, which is matched with predefined point cloud templates representing different scene types to determine candidate scene types. For example, the current feature point cloud is compared with predefined point clouds representing "office scene" and "outdoor street scene" to calculate similarity and identify scene types with high similarity as candidates.

[0089] By fusing multi-dimensional features and using feature point cloud matching, possible scene types are initially screened out. By comprehensively considering information from multiple dimensions, the accuracy and comprehensiveness of scene recognition are improved, and the possibility of misjudgment is reduced.

[0090] The prediction layer performs scene adaptability prediction on candidate scene types, utilizing the knowledge learned by the model to evaluate the adaptability probability of each candidate scene type in the current environment. This may involve using machine learning algorithms, such as neural networks and decision trees, to train the model based on historical data, enabling it to determine the probability of each candidate scene type based on input features. Thus, a quantified adaptability probability is assigned to each candidate scene type, further clarifying the degree of matching between each candidate scene and the current environment. This provides a more convincing basis for ultimately determining the real-time scene type and improves the reliability of scene recognition.

[0091] The output layer determines the real-time scene type based on the adaptation probability corresponding to the candidate scene types. Typically, the candidate scene type with the highest adaptation probability is selected as the final real-time scene type output. If multiple candidate scene types have similar probabilities, other strategies can be combined, such as further analyzing feature details or referring to historical scene data to make a decision. This provides a clear real-time scene type result, completing the entire environmental scene recognition process. It provides an accurate basis for subsequent switching of intelligent noise reduction modes based on scene type, enabling the intelligent noise reduction system to accurately adapt to different environments.

[0092] For example, the pre-configured noise reduction mode selection strategy is based on extensive experimental data and real-world user feedback, summarizing the most suitable noise reduction mode for different combinations of noise levels and scene types. For instance, in a quiet library scene, even with occasional slight page-turning sounds, the overall noise level is low, making a light noise reduction mode suitable to avoid excessive noise reduction affecting sound details. Conversely, in a noisy factory workshop, where noise levels are high and continuous, a strong noise reduction mode is needed to ensure clear recording. By matching the real-time acquired noise level and scene type with the conditions in the strategy, the target noise reduction mode can be determined.

[0093] This method can quickly and accurately match the appropriate noise reduction mode for different environments, greatly improving the targeting of noise reduction. It avoids applying noise reduction in inappropriate modes, such as using strong noise reduction in a quiet environment which can cause sound distortion, or using light noise reduction in a noisy environment which cannot effectively eliminate noise, thus ensuring that the recording quality is always at a high level.

[0094] Furthermore, user attributes include factors such as age, hearing preferences, and usage habits. Users of different ages have different sensitivities to sound; for example, older people may prefer to retain more sound details. Users with different hearing preferences have different requirements for noise reduction intensity and frequency range. Usage habits also affect parameter settings; for example, users who frequently record meetings may want to emphasize the human voice frequency range. Parameters are generated by comprehensively considering these factors, combined with real-time noise levels and scene type. For example, for young users who prefer clear human voices in high-noise environments such as airport waiting halls (scene type), the parameters will focus on enhancing the filtering of high-frequency noise while preserving the clarity of the human voice frequency range.

[0095] This enables personalized customization of noise reduction mode parameters. Compared to generic parameter settings, parameters generated after considering user attributes better meet the unique needs of each user, improving the user experience. This allows users to obtain recording results that meet their expectations in different environments, enhancing the product's applicability and user satisfaction.

[0096] In the steps described above, the device's internal audio processing module receives the generated noise reduction mode parameters. These parameters control the operation of the audio processing algorithm, such as adjusting the filter cutoff frequency, gain, and noise reduction intensity. For example, if the parameters specify that noise in a specific frequency band should be attenuated by 20dB in a strong noise reduction mode, the audio processing module will process the input audio signal according to these parameters, switching from the current noise reduction state to the target noise reduction mode. This ensures that the user's device can accurately and efficiently switch to a noise reduction mode that suits the current environment and the user's needs. It enables the device to quickly adapt to environmental changes, continuously providing users with a high-quality recording environment, ensuring that recording quality is not affected by the device switching process, and maintaining the continuity and stability of the recording.

[0097] As an optional embodiment, based on the real-time noise level and the real-time scene type, automatically switching the user device to the corresponding target noise reduction mode to ensure the recording quality of the user device includes:

[0098] Based on the real-time noise level and the real-time scene type, a target noise reduction mode is determined using a pre-configured noise reduction mode selection strategy; based on user attributes, the real-time noise level, and the real-time scene type, noise reduction mode parameters matching the target noise reduction mode are generated; and the user device is switched to the corresponding target noise reduction mode using the noise reduction mode parameters.

[0099] In the steps described above, the target noise reduction mode is determined based on the real-time noise level and scene type. The pre-configured noise reduction mode selection strategy is based on extensive experimental data and real-world user feedback, summarizing the most suitable noise reduction modes for different combinations of noise levels and scene types. For example, in a quiet library scene, even with occasional slight page-turning sounds, the overall noise level is low, making a light noise reduction mode suitable to avoid excessive noise reduction affecting sound details. In contrast, in a noisy factory workshop, where noise levels are high and persistent, a strong noise reduction mode is needed to ensure clear recording. By matching the real-time noise level and scene type with the conditions in the strategy, the target noise reduction mode can be determined. This method can quickly and accurately match appropriate noise reduction modes for different environments, greatly improving the targeting of noise reduction. It avoids using inappropriate modes for noise reduction, such as using strong noise reduction in a quiet environment leading to sound distortion, or using light noise reduction in a noisy environment failing to effectively eliminate noise, thus ensuring that the recording quality remains at a high level.

[0100] Then, noise reduction mode parameters are generated based on user attributes, real-time noise levels, and real-time scene types. User attributes include factors such as age, hearing preferences, and usage habits. Users of different ages have different sensitivities to sound; for example, older people may prefer to retain more sound details. Users with different hearing preferences have different requirements for noise reduction intensity and frequency range. Usage habits also affect parameter settings; for example, users who frequently record meetings may want to emphasize the human voice frequency band. Parameters are generated by comprehensively considering these factors in conjunction with real-time noise levels and scene types. For example, for young users who prefer clear human voices in high-noise environments such as airport waiting halls (scene type), the parameters will focus on enhancing the filtering of high-frequency noise while preserving the clarity of the human voice frequency band.

[0101] This allows for personalized customization of noise reduction mode parameters. Compared to generic parameter settings, parameters generated after considering user attributes better meet the unique needs of each user, improving the user experience. This allows users to obtain recording results that meet their expectations in different environments, enhancing the product's applicability and user satisfaction.

[0102] Finally, the user device is switched to the target noise reduction mode using noise reduction mode parameters. The device's internal audio processing module receives the generated noise reduction mode parameters. These parameters control the operation of the audio processing algorithm, such as adjusting the filter cutoff frequency, gain, and noise reduction intensity. For example, if the parameters specify that noise in a certain frequency band should be attenuated by 20dB in strong noise reduction mode, the audio processing module will process the input audio signal according to these parameters, thus switching from the current noise reduction state to the target noise reduction mode.

[0103] This ensures that user devices can accurately and efficiently switch to noise reduction modes that suit the current environment and user needs. It enables devices to quickly adapt to environmental changes, continuously providing users with a high-quality recording environment, ensuring that recording quality is not affected by device switching, and maintaining the continuity and stability of recording.

[0104] The following is an example of an intelligent noise reduction mode switching method in a voice recorder scenario:

[0105] Scene 1: Office Scene

[0106] Once the recorder is turned on, its built-in sensors begin to operate. The microphone collects ambient sound data, detecting occasional conversations with colleagues, keyboard clicks, and a faint hum from the air conditioner; the overall sound is relatively stable and the volume is moderate. The position sensor indicates the device is in a fixed location in the office, stationary on a desk. The ambient temperature and humidity sensor reports a temperature of 25°C and humidity of 50%. The barometric pressure sensor shows the air pressure is stable near standard atmospheric pressure. The electromagnetic interference intensity data indicates the presence of weak and stable electromagnetic interference from office equipment.

[0107] The noise level assessment model analyzed the ambient sound data and concluded that the noise level was low. Meanwhile, the environmental scene recognition model, based on the collected data, determined that the current scene was an office.

[0108] Based on a pre-set strategy, a light noise reduction mode is suitable for low-noise office environments. The neural network model combines noise levels and scene type to output corresponding light noise reduction mode parameters. For example, it sets a slower update rate for noise estimation in spectral subtraction to avoid overestimating noise, while fine-tuning the Wiener filter parameters to suppress low-frequency background noise to a certain extent without excessively affecting speech frequencies. The recorder automatically switches to light noise reduction mode based on these parameters, reducing background noise while preserving as much detail as possible in conversations and keyboard typing, ensuring clear and natural recording quality that meets the requirements for voice recording in office settings.

[0109] Scene 2: Street Scene

[0110] The user walked onto the street with the recorder, and the microphone captured a large amount of noise, including car sounds, horns, and crowd noise, with significant variations in sound intensity and frequency. The location sensor indicated that the device was in motion, and its latitude and longitude were constantly changing. Ambient temperature and humidity data showed a temperature of 30°C, humidity of 40%, and air pressure close to standard atmospheric pressure, but the intensity of electromagnetic interference increased due to the presence of more electronic devices in the vicinity.

[0111] The noise level assessment model analysis indicates that the current noise level is high, and the environmental scene recognition model identifies it as a street scene.

[0112] For high-noise street scenes, a strong noise reduction mode is selected. The neural network model generates parameters for the strong noise reduction mode based on the noise level and scene type. For spectral subtraction, the noise estimation update speed is accelerated to adapt to the rapidly changing noise environment; for Wiener filtering, the attenuation coefficient for high-frequency noise is increased. The recorder quickly switches to strong noise reduction mode based on these parameters, effectively filtering out most street noise, highlighting useful audio such as human voices, and ensuring that the recorded content remains clearly identifiable in noisy environments.

[0113] Scenario 3: Adaptive optimization based on user habits

[0114] Over a period of use, the voice recorder recorded the user's operating habits regarding noise cancellation mode in different scenarios. For example, in a gym setting, the user frequently manually adjusted the noise cancellation mode to appropriately reduce the volume of music and gym equipment while preserving the sounds of conversations around them. By analyzing these actions, the voice recorder learned the user's preference for preserving a certain amount of ambient sound in a gym setting.

[0115] When the system detects a gym environment with high noise levels again, the neural network model, in addition to considering the scene and noise level, also incorporates the user's habit to generate specific noise reduction mode parameters. Based on a strong noise reduction mode, the parameters are fine-tuned to reduce the noise reduction intensity in certain ambient audio segments. This effectively reduces noise interference while still meeting the user's need to retain some ambient sound, further enhancing the user experience.

[0116] In this embodiment, the noise level assessment model and the environmental scene recognition model work together to analyze environmental data from different perspectives, effectively improving the adaptability of user equipment to the environment, helping user equipment maintain the best noise reduction effect, and ensuring that the recording quality is not affected by environmental changes.

[0117] Figure 2 This is a schematic diagram of the structure of an intelligent noise reduction mode switching system provided in an embodiment of this application, as shown below. Figure 2 As shown, the system includes the following steps:

[0118] The acquisition unit is used to acquire real-time environmental data of the user device's current environment after recording is started; the real-time environmental data includes at least: ambient sound data, spatial location data, device usage status, ambient temperature and humidity data, air pressure data, and electromagnetic interference intensity data.

[0119] The analysis unit is used to perform noise analysis on the real-time environmental data through a noise level assessment model to obtain the real-time noise level in the environment where the user equipment is currently located; the real-time noise level is related to the amplitude of noise signal changes in the real-time environmental data.

[0120] The identification unit is used to input the real-time environmental data into the environmental scene identification model to obtain the real-time scene type corresponding to the current environment of the user device; the real-time scene type is determined based on the real-time spatial changes reflected by any one or more of the spatial location data, device usage status, environmental temperature and humidity data, air pressure data, and electromagnetic interference intensity data, as well as the noise signal change trend in the environmental sound data.

[0121] The switching unit is used to automatically switch the user device to the corresponding target noise reduction mode based on the real-time noise level and the real-time scene type to ensure the recording quality of the user device.

[0122] Optionally, it further includes a preference unit for receiving a noise reduction mode adjustment instruction issued by the user device; the noise reduction mode adjustment instruction is constructed based on parameter adjustment information input by the user in the user device; based on the noise reduction mode adjustment instruction and the historical scene type of the user device, a user individual preference library corresponding to the user device is established; wherein, the user individual preference library stores the parameter adjustment strategies corresponding to the user device under each historical scene type.

[0123] Furthermore, after the switching unit automatically switches the user device to the corresponding target noise reduction mode using the noise reduction mode parameters, the preference unit is further configured to select a corresponding noise reduction mode parameter adjustment strategy from the user's individual preference library, and adjust the noise reduction mode parameters in the target noise reduction mode based on the selected parameter adjustment strategy, so that the recording quality matches the user's individual preferences.

[0124] Further optionally, the noise level assessment model includes at least: an extraction layer, an analysis layer, and an output layer;

[0125] The analysis unit performs noise analysis on the real-time environmental data using a noise level assessment model to obtain the real-time noise level of the user equipment's current environment. Specifically, it is used for:

[0126] Environmental audio features are extracted from the real-time environmental data through an extraction layer;

[0127] Through the analysis layer, the environmental audio features are analyzed and processed in real time to obtain the corresponding environmental noise change amplitude features;

[0128] The output layer classifies the environmental noise variation amplitude characteristics obtained from the analysis according to a preset strategy to obtain the real-time noise level.

[0129] Optionally, the analysis unit performs real-time signal amplitude analysis on the environmental audio features through the analysis layer to obtain the corresponding environmental noise change amplitude features, specifically for:

[0130] The environmental audio features are subjected to noise separation processing to obtain multiple environmental noise signals;

[0131] The separated multi-channel environmental noise signals are subjected to channel separation, and the peak points of each independent channel noise signal are identified to determine the peak points of each independent channel noise signal.

[0132] Based on the peak points, the real-time signal amplitude of each independent channel noise signal is determined, and the corresponding environmental noise variation amplitude characteristics are obtained.

[0133] Optionally, the analysis unit performs channel separation on the separated multi-channel environmental noise signals and performs peak identification on each independent channel noise signal to determine the peak points in each independent channel noise signal, specifically for:

[0134] The multiple environmental noise signals are divided into multiple environmental noise signal segments under multiple sliding time windows by a sliding time window.

[0135] According to the channel type, the multi-channel environmental noise signal segments under multiple sliding timing windows are separated and processed to obtain the independent channel noise signal corresponding to each sliding timing window.

[0136] The channel types include at least the left stereo channel, the right stereo channel, and channels under different dry and wet sound levels;

[0137] The quartiles of each independent channel noise signal are calculated to obtain the quartile positions and quartile ranges of each independent channel noise signal.

[0138] Anomaly detection is performed based on the quartile positions and quartile ranges corresponding to the noise signals of each independent channel to obtain the abnormal signal values ​​contained in the noise signals of each independent channel.

[0139] The abnormal signal values ​​in the noise signals of each independent channel are processed to obtain the optimized noise signals of each independent channel. Hilbert transformation and amplitude envelope calculation are then performed on the optimized noise signals of each independent channel to obtain the amplitude envelope of the noise signals of each independent channel.

[0140] Peak identification is performed based on the amplitude envelope of the noise signal in each independent channel to obtain the peak point in the noise signal in each independent channel.

[0141] Further optionally, the environmental scene recognition model includes at least: a preprocessing layer, a feature extraction layer, a construction and fusion layer, a prediction layer, and an output layer;

[0142] The identification unit inputs the real-time environmental data into the environmental scene identification model to obtain the real-time scene type corresponding to the current environment of the user device, specifically for:

[0143] The real-time environmental data is preprocessed through a preprocessing layer;

[0144] Through the feature extraction layer, multi-dimensional real-time environmental features are extracted from the real-time environmental data; the multi-dimensional real-time environmental features include at least: spatial location features, equipment usage status features, environmental temperature and humidity features, air pressure features, electromagnetic interference intensity features, and environmental sound features;

[0145] By constructing a fusion layer, the multidimensional real-time environmental features are projected into the multidimensional analysis space as corresponding feature space points, and the feature point cloud formed by the feature space points is matched with the point cloud type to determine the candidate scene type.

[0146] The prediction layer performs scene adaptability prediction on the candidate scene types to obtain the adaptability probability corresponding to the candidate scene types.

[0147] The real-time scene type is determined by the output layer based on the adaptation probability corresponding to the candidate scene type.

[0148] Further optionally, the switching unit, based on the real-time noise level and the real-time scene type, automatically switches the user equipment to the corresponding target noise reduction mode to ensure the recording quality of the user equipment, specifically for:

[0149] Based on the real-time noise level and the real-time scene type, a pre-configured noise reduction mode selection strategy is used to determine the target noise reduction mode.

[0150] Based on user attributes, the real-time noise level, and the real-time scene type, generate noise reduction mode parameters that match the target noise reduction mode;

[0151] The user device is switched to the corresponding target noise reduction mode using the noise reduction mode parameters.

[0152] Please see Figure 3 , Figure 3 A schematic diagram illustrating an embodiment of the electronic device provided in this application. For example... Figure 3 As shown, this application provides an electronic device 500, including a memory 510, a processor 520, and a computer program 511 stored in the memory 510 and executable on the processor 520. When the processor 520 executes the computer program 511, it implements the aforementioned embodiments.

[0153] Please see Figure 4 , Figure 4 This is a schematic diagram illustrating an embodiment of a computer-readable storage medium provided in this application. For example... Figure 4 As shown, this embodiment provides a computer-readable storage medium 600 on which a computer program 611 is stored, which implements the aforementioned embodiment when executed by a processor.

[0154] It should be noted that the descriptions of each embodiment in the above embodiments have different focuses. For parts that are not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0155] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0156] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0157] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0158] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0159] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.

[0160] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. A method for switching intelligent noise reduction modes, characterized in that, The method includes at least: After recording is started, real-time environmental data of the user device’s current environment is acquired; the real-time environmental data includes at least: ambient sound data, spatial location data, device usage status, ambient temperature and humidity data, air pressure data, and electromagnetic interference intensity data. The real-time environmental data is analyzed using a noise level assessment model to obtain the real-time noise level of the user equipment's current environment; the real-time noise level is correlated with the amplitude of noise signal changes in the real-time environmental data. The real-time environmental data is input into the environmental scene recognition model to obtain the real-time scene type corresponding to the current environment of the user device; the real-time scene type is determined based on the real-time spatial changes reflected by any one or more of the spatial location data, device usage status, environmental temperature and humidity data, air pressure data, and electromagnetic interference intensity data, as well as the noise signal change trend in the environmental sound data. Based on the real-time noise level and the real-time scene type, the user device is automatically switched to the corresponding target noise reduction mode to ensure the recording quality of the user device. The noise level assessment model includes at least: an extraction layer, an analysis layer, and an output layer; The step of performing noise analysis on the real-time environmental data using a noise level assessment model to obtain the real-time noise level of the user equipment's current environment includes: Environmental audio features are extracted from the real-time environmental data through an extraction layer; Through the analysis layer, the environmental audio features are analyzed and processed in real time to obtain the corresponding environmental noise change amplitude features; The output layer classifies the environmental noise variation amplitude characteristics obtained from the analysis according to a preset strategy to obtain the real-time noise level. The step involves performing real-time signal amplitude analysis on the environmental audio features through an analysis layer to obtain the corresponding environmental noise change amplitude features, including: The environmental audio features are subjected to noise separation processing to obtain multiple environmental noise signals; The separated multi-channel environmental noise signals are subjected to channel separation, and the peak points of each independent channel noise signal are identified to determine the peak points of each independent channel noise signal. Based on the peak points, the real-time signal amplitude of each independent channel noise signal is determined, and the corresponding environmental noise change amplitude characteristics are obtained. The process of performing channel separation on the separated multi-channel environmental noise signals and peak identification on each independent channel noise signal to determine the peak point in each independent channel noise signal includes: The multiple environmental noise signals are divided into multiple environmental noise signal segments under multiple sliding time windows by a sliding time window. According to the channel type, the multiple ambient noise signal segments under multiple sliding timing windows are separated and processed to obtain the independent channel noise signal corresponding to each sliding timing window; the channel type includes at least the left stereo channel, the right stereo channel, and channels under different dry and wet sound levels. The quartiles of each independent channel noise signal are calculated to obtain the quartile positions and quartile ranges of each independent channel noise signal. Anomaly detection is performed based on the quartile positions and quartile ranges corresponding to the noise signals of each independent channel to obtain the abnormal signal values ​​contained in the noise signals of each independent channel. The abnormal signal values ​​in the noise signals of each independent channel are processed to obtain the optimized noise signals of each independent channel. Hilbert transformation and amplitude envelope calculation are then performed on the optimized noise signals of each independent channel to obtain the amplitude envelope of the noise signals of each independent channel. Peak identification is performed based on the amplitude envelope of the noise signal in each independent channel to obtain the peak point in the noise signal in each independent channel.

2. The intelligent noise reduction mode switching method according to claim 1, characterized in that, The method further includes: Receive a noise reduction mode adjustment command sent by the user equipment; the noise reduction mode adjustment command is constructed based on the parameter adjustment information input by the user in the user equipment; Based on the noise reduction mode adjustment command and the historical scene type of the user device, a user individual preference library corresponding to the user device is established; wherein, the user individual preference library stores the parameter adjustment strategies corresponding to the user device under each historical scene type; After automatically switching the user device to the corresponding target noise reduction mode using noise reduction mode parameters, the following is also included: Select the corresponding noise reduction mode parameter adjustment strategy from the user's individual preference library, and adjust the noise reduction mode parameters in the target noise reduction mode based on the selected parameter adjustment strategy so that the recording quality matches the user's individual preferences.

3. The intelligent noise reduction mode switching method according to claim 1, characterized in that, The environmental scene recognition model includes at least: a preprocessing layer, a feature extraction layer, a construction and fusion layer, a prediction layer, and an output layer; The step of inputting the real-time environmental data into the environmental scene recognition model to obtain the real-time scene type corresponding to the current environment of the user device includes: The real-time environmental data is preprocessed through a preprocessing layer; Through the feature extraction layer, multi-dimensional real-time environmental features are extracted from the real-time environmental data; the multi-dimensional real-time environmental features include at least: spatial location features, equipment usage status features, environmental temperature and humidity features, air pressure features, electromagnetic interference intensity features, and environmental sound features; By constructing a fusion layer, the multidimensional real-time environmental features are projected into the multidimensional analysis space as corresponding feature space points, and the feature point cloud formed by the feature space points is matched with the point cloud type to determine the candidate scene type. The prediction layer performs scene adaptability prediction on the candidate scene types to obtain the adaptability probability corresponding to the candidate scene types. The real-time scene type is determined by the output layer based on the adaptation probability corresponding to the candidate scene type.

4. The intelligent noise reduction mode switching method according to claim 1, characterized in that, The step of automatically switching the user device to the corresponding target noise reduction mode based on the real-time noise level and the real-time scene type to ensure the recording quality of the user device includes: Based on the real-time noise level and the real-time scene type, a pre-configured noise reduction mode selection strategy is used to determine the target noise reduction mode. Based on user attributes, the real-time noise level, and the real-time scene type, generate noise reduction mode parameters that match the target noise reduction mode; The user device is switched to the corresponding target noise reduction mode using the noise reduction mode parameters.

5. An intelligent noise reduction mode switching system, characterized in that, The system includes at least the following units: The acquisition unit is used to acquire real-time environmental data of the user device's current environment after recording is started; The real-time environmental data includes at least: ambient sound data, spatial location data, equipment usage status, ambient temperature and humidity data, air pressure data, and electromagnetic interference intensity data; The analysis unit is used to perform noise analysis on the real-time environmental data through a noise level assessment model to obtain the real-time noise level in the environment where the user equipment is currently located; the real-time noise level is related to the amplitude of noise signal changes in the real-time environmental data. The identification unit is used to input the real-time environmental data into the environmental scene identification model to obtain the real-time scene type corresponding to the current environment of the user device; the real-time scene type is determined based on the real-time spatial changes reflected by any one or more of the spatial location data, device usage status, environmental temperature and humidity data, air pressure data, and electromagnetic interference intensity data, as well as the noise signal change trend in the environmental sound data. The switching unit is used to automatically switch the user device to the corresponding target noise reduction mode based on the real-time noise level and the real-time scene type to ensure the recording quality of the user device; The noise level assessment model includes at least: an extraction layer, an analysis layer, and an output layer; The step of performing noise analysis on the real-time environmental data using a noise level assessment model to obtain the real-time noise level of the user equipment's current environment includes: Environmental audio features are extracted from the real-time environmental data through an extraction layer; Through the analysis layer, the environmental audio features are analyzed and processed in real time to obtain the corresponding environmental noise change amplitude features; The output layer classifies the environmental noise variation amplitude characteristics obtained from the analysis according to a preset strategy to obtain the real-time noise level. The step involves performing real-time signal amplitude analysis on the environmental audio features through an analysis layer to obtain the corresponding environmental noise change amplitude features, including: The environmental audio features are subjected to noise separation processing to obtain multiple environmental noise signals; The separated multi-channel environmental noise signals are subjected to channel separation, and the peak points of each independent channel noise signal are identified to determine the peak points of each independent channel noise signal. Based on the peak points, the real-time signal amplitude of each independent channel noise signal is determined, and the corresponding environmental noise change amplitude characteristics are obtained. The process of performing channel separation on the separated multi-channel environmental noise signals and peak identification on each independent channel noise signal to determine the peak point in each independent channel noise signal includes: The multiple environmental noise signals are divided into multiple environmental noise signal segments under multiple sliding time windows by a sliding time window. According to the channel type, the multiple ambient noise signal segments under multiple sliding timing windows are separated and processed to obtain the independent channel noise signal corresponding to each sliding timing window; the channel type includes at least the left stereo channel, the right stereo channel, and channels under different dry and wet sound levels. The quartiles of each independent channel noise signal are calculated to obtain the quartile positions and quartile ranges of each independent channel noise signal. Anomaly detection is performed based on the quartile positions and quartile ranges corresponding to the noise signals of each independent channel to obtain the abnormal signal values ​​contained in the noise signals of each independent channel. The abnormal signal values ​​in the noise signals of each independent channel are processed to obtain the optimized noise signals of each independent channel. Hilbert transformation and amplitude envelope calculation are then performed on the optimized noise signals of each independent channel to obtain the amplitude envelope of the noise signals of each independent channel. Peak identification is performed based on the amplitude envelope of the noise signal in each independent channel to obtain the peak point in the noise signal in each independent channel.

Citation Information

Patent Citations

  • Audio processing method and microphone equipment

    CN113825068A