Intelligent noise reduction mode switching method and system
Through the collaborative work of noise level evaluation and environmental scene recognition model, the noise reduction mode is adjusted in real time, which solves the problem that static settings cannot meet users' personalized needs, and achieves efficient recording quality and user experience improvement in complex environments.
Patent Information
- Application Number
- CN202510873606.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-06-27
AI Technical Summary
The static settings of existing noise reduction parameters cannot meet the user's personalized and diverse needs, and cannot switch to real-time modifications according to the scene, resulting in poor noise reduction effect in complex and changing environments, affecting recording quality and user experience.
Through the collaborative work of the noise level evaluation model and the environmental scene recognition model, we obtain environmental data in real time, analyze the noise level and scene type, and automatically switch to the corresponding target noise reduction mode to ensure recording quality.
It achieves the optimal noise reduction effect in complex environments, simplifies user operations, improves recording quality and user experience, and enhances product adaptability and user loyalty.
Smart Images

Figure CN120375871A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of noise reduction, and in particular to an intelligent noise reduction mode switching method and system. Background Art
[0002] Among current numerous devices and systems related to audio processing, the noise reduction function has become a key factor in enhancing the user experience. However, the current situation that most noise reduction parameters are statically set has led to a series of problems that urgently need to be solved, seriously restricting the optimization of noise reduction effects and the improvement of user experience.
[0003] Different scenarios have completely different acoustic environment characteristics. In a quiet library, the ambient noise mainly comes from the slight sound of turning pages, whispering, and the faint sound of equipment operation. At this time, the demand for noise reduction focuses on the precise suppression of extremely subtle noises, while maximizing the retention of the details of the original sound to ensure that sounds such as reading records are clearly distinguishable. However, on a bustling and noisy street, there are the roars of cars, honking, the hubbub of the crowd, and the noises of various means of transportation. These noises not only have high intensity but also a wide and rapidly changing frequency range. In a factory workshop, the high-intensity, regular low-frequency noise generated by continuous machine operation dominates. Facing such diverse scenarios, statically set noise reduction parameters are difficult to meet the specific requirements in different scenarios. Fixed noise reduction parameters cannot effectively adapt to the unique frequency characteristics, intensity changes, and dynamic ranges of noises in different scenarios, resulting in over-noise reduction in some scenarios, severely distorting the sound and losing important information; while in other scenarios, the noise reduction is insufficient, and the noise still significantly interferes with the quality of recording or voice communication.
[0004] The actual environment is often in a dynamic change. For example, when a user takes the subway, entering the noisy carriage interior from a relatively quiet platform, the ambient noise increases significantly instantaneously; or during a walk outdoors, suddenly approaching a construction site from an open square, both the noise intensity and frequency characteristics change sharply. In the case of such rapid scene switching, static noise reduction parameters cannot respond in a timely manner, resulting in a serious deviation in the noise reduction effect at the moment of scene conversion. At the moment when the noise increases, due to the failure of the noise reduction parameters to be adjusted synchronously, the noise will enter the audio acquisition device unhindered, seriously affecting the recording quality or voice clarity; while when the noise suddenly weakens, the originally high-intensity noise reduction setting may overprocess the audio signal, making the sound hollow and unnatural.
[0005] In the related art, the static noise reduction parameter setting cannot meet the personalized and diverse needs of users. There are significant differences in the perception and preference of sound among different users. Some users are extremely sensitive to noise and hope to achieve extreme noise reduction effects in any scenario; while others pay more attention to retaining the atmosphere of the ambient sound in order to better integrate into the surrounding environment. In addition, users have different noise reduction requirements in different usage scenarios, such as learning, working, and entertainment. However, the existing static noise reduction parameter setting cannot be flexibly adjusted according to the individual differences of users and different usage scenarios, making it impossible for users to obtain the best audio experience according to their own needs, and greatly reducing the satisfaction and loyalty of users to the device or system.
[0006] Obviously, the static setting of the existing noise reduction parameters severely limits the adaptability, real-time performance of the noise reduction system in a complex and changing environment, as well as its ability to meet the personalized needs of users. There is an urgent need for a solution that can modify the noise reduction parameters in real time according to the scene switch to improve the quality of audio processing and user experience. Summary of the Invention
[0007] In view of the technical problems existing in the prior art, the present invention provides an intelligent noise reduction mode switching method and system, which are used to analyze environmental data from different angles through the collaborative work of a noise level evaluation model and an environmental scene recognition model, effectively improving the adaptability of the user device to the surrounding environment, assisting the user device to maintain the optimal noise reduction effect, and ensuring that the recording quality is not interfered by environmental changes.
[0008] In a first aspect, an embodiment of the present application provides an intelligent noise reduction mode switching method, including: After the recording is started, obtain the real-time environmental data in the current environment where the user device is located; the real-time environmental data at least includes: environmental sound data, spatial position data, device usage status, environmental temperature and humidity data, air pressure data, electromagnetic interference intensity data; Perform noise analysis on the real-time environmental data through a noise level evaluation model to obtain the real-time noise level in the current environment where the user device is located; the real-time noise level is associated with the change amplitude of the noise signal in the real-time environmental data; Input the real-time environmental data into an environmental scene recognition model to obtain the real-time scene type corresponding to the current environment where the user device is located; the real-time scene type is determined based on any one or more of the real-time spatial changes reflected by the spatial position data, device usage status, environmental temperature and humidity data, air pressure data, electromagnetic interference intensity data, and the change trend of the noise signal in the environmental sound data; Based on the real-time noise level and the real-time scene type, automatically switch the user device to the corresponding target noise reduction mode to ensure the recording quality of the user device.
[0009] In a second aspect, an embodiment of the present application provides an intelligent noise reduction mode switching system, which includes the following units: An acquisition unit, configured to acquire real-time environment data in the current environment of the user device after the recording is started; the real-time environment data at least includes: environmental sound data, spatial position data, device usage status, environmental temperature and humidity data, air pressure data, and electromagnetic interference intensity data; An analysis unit, configured to perform noise analysis on the real-time environment data through a noise level evaluation model to obtain the real-time noise level in the current environment of the user device; the real-time noise level is associated with the change amplitude of the noise signal in the real-time environment data; An identification unit, configured to input the real-time environment data into an environment scene identification model to obtain the real-time scene type corresponding to the current environment of the user device; the real-time scene type is determined based on any one or more of the real-time spatial changes reflected by the spatial position data, device usage status, environmental temperature and humidity data, air pressure data, and electromagnetic interference intensity data, and the change trend of the noise signal in the environmental sound data; A switching unit, configured to automatically switch the user device to the corresponding target noise reduction mode based on the real-time noise level and the real-time scene type to ensure the recording quality of the user device.
[0010] In a third aspect, an embodiment of the present application provides an electronic device, which includes: At least one processor, a memory, and an input / output unit; Wherein, the memory is used to store a computer program, and the processor is used to call the computer program stored in the memory to execute the intelligent noise reduction mode switching method of the first aspect.
[0011] In a fourth aspect, a computer-readable storage medium is provided, which includes instructions that, when the instructions are run on a computer, cause the computer to execute the intelligent noise reduction mode switching method of the first aspect.
[0012] The beneficial effects of the present invention are as follows: It provides an intelligent noise reduction mode switching method and system. In this technical solution, after the recording is turned on, the real-time environmental data in the environment where the user device is currently located is obtained; the real-time environmental data at least includes: environmental sound data, spatial position data, device usage status, environmental temperature and humidity data, air pressure data, and electromagnetic interference intensity data. Furthermore, the real-time environmental data is subjected to noise analysis through a noise level evaluation model to obtain the real-time noise level in the environment where the user device is currently located; the real-time noise level is associated with the change amplitude of the noise signal in the real-time environmental data. Then, the real-time environmental data is input into an environmental scene recognition model to obtain the real-time scene type corresponding to the environment where the user device is currently located; the real-time scene type is determined based on any one or more of the real-time spatial changes reflected by the spatial position data, device usage status, environmental temperature and humidity data, air pressure data, and electromagnetic interference intensity data, and the change trend of the noise signal in the environmental sound data. Finally, based on the real-time noise level and the real-time scene type, the user device is automatically switched to the corresponding target noise reduction mode to ensure the recording quality of the user device.
[0013] In the embodiments of the present application, through the collaborative work of the noise level evaluation model and the environmental scene recognition model, the environmental data is analyzed from different perspectives, effectively improving the adaptability of the user device to the surrounding environment, assisting the user device to maintain the optimal noise reduction effect, and ensuring that the recording quality is not interfered by environmental changes. Description of the Drawings
[0014] Figure 1 is a flowchart of an intelligent noise reduction mode switching method according to an embodiment of the present application; Figure 2 is a schematic structural diagram of an intelligent noise reduction mode switching system according to an embodiment of the present application; Figure 3 is a schematic structural diagram of an electronic device according to an embodiment of the present application; Figure 4 is a schematic structural diagram of a medium device according to an embodiment of the present application. Detailed Embodiments
[0015] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts fall within the protection scope of the present application.
[0016] In the description of the present application, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of the described features. In the description of the present application, "a plurality of" means two or more unless otherwise specifically defined.
[0017] In the description of the present application, the term "for example" is used to mean "serving as an example, illustration, or explanation". Any embodiment described as "for example" in the present application is not necessarily construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use the present invention. In the following description, details are set forth for purposes of explanation. It should be understood that those of ordinary skill in the art can recognize that the present invention can be practiced without these specific details. In other instances, well-known structures and processes are not described in detail to avoid obscuring the description of the present invention with unnecessary details. Therefore, the present invention is not intended to be limited to the embodiments shown, but is to be accorded the widest scope consistent with the principles and features disclosed in the present application.
[0018] An embodiment of the present application provides a method and system for switching intelligent noise reduction modes. In this technical solution, after recording is started, real-time environmental data in the environment where the user device is currently located is obtained; the real-time environmental data at least includes: environmental sound data, spatial position data, device usage status, environmental temperature and humidity data, air pressure data, and electromagnetic interference intensity data. Then, the real-time environmental data is analyzed for noise through a noise level evaluation model to obtain the real-time noise level in the environment where the user device is currently located; the real-time noise level is associated with the change amplitude of the noise signal in the real-time environmental data. Next, the real-time environmental data is input into an environmental scene recognition model to obtain the real-time scene type corresponding to the environment where the user device is currently located; the real-time scene type is determined based on any one or more of the real-time spatial changes reflected by the spatial position data, device usage status, environmental temperature and humidity data, air pressure data, and electromagnetic interference intensity data, and the change trend of the noise signal in the environmental sound data. Finally, based on the real-time noise level and the real-time scene type, the user device is automatically switched to the corresponding target noise reduction mode to ensure the recording quality of the user device.
[0019] In an embodiment of the present application, through the collaborative work of the noise level evaluation model and the environmental scene recognition model, the environmental data is analyzed from different perspectives, effectively improving the adaptability of the user device to the surrounding environment, assisting the user device to maintain the optimal noise reduction effect, and ensuring that the recording quality is not interfered by environmental changes.
[0020] The intelligent noise reduction mode switching solution provided by the embodiments of the present application can also be executed by an electronic device, which can be a server, a server cluster, or a cloud server. The electronic device can also be a terminal device such as a mobile phone, a computer, a tablet computer, a wearable device, or a dedicated device (such as a dedicated terminal device with an intelligent noise reduction mode switching system). The above-mentioned chips introduced in the foregoing embodiments can also be installed in these electronic devices. Alternatively, a service program for executing the intelligent noise reduction mode switching solution can also be installed in these electronic devices.
[0021] Figure 1 It is a schematic diagram of an intelligent noise reduction mode switching method provided by the embodiments of the present application. As Figure 1 shown, the method includes the following steps: 101. After the recording is turned on, obtain the real-time environmental data in the environment where the user device is currently located; 102. Perform noise analysis on the real-time environmental data through a noise level evaluation model to obtain the real-time noise level in the environment where the user device is currently located; the real-time noise level is associated with the change amplitude of the noise signal in the real-time environmental data; 103. Input the real-time environmental data into an environmental scene recognition model to obtain the real-time scene type corresponding to the environment where the user device is currently located; 104. Based on the real-time noise level and the real-time scene type, automatically switch the user device to the corresponding target noise reduction mode to ensure the recording quality of the user device.
[0022] In the embodiments of the present application, the real-time scene type is determined based on any one or more of the spatial position data, device usage status, environmental temperature and humidity data, air pressure data, and electromagnetic interference intensity data, which reflect the real-time spatial changes, and the change trend of the noise signal in the environmental sound data.
[0023] For example, when the user turns on the recording function while traveling by high-speed rail. At this time, the spatial position data shows that the device is in a high-speed moving state, and the change in longitude and latitude presents a trend consistent with the high-speed rail route.
[0024] The device usage status indicates that it is placed on the small table and is in the recording state. The environmental temperature and humidity data is relatively stable, the temperature is maintained at 26 degrees Celsius, the humidity is about 40%, and the air pressure data is also stable at about 101 kPa.
[0025] The electromagnetic interference intensity data shows that there is a certain interference from electronic devices, which is in line with the operation of numerous electronic devices in the high-speed rail carriage. From the environmental sound data, the change trend of the noise signal shows a continuous low-frequency roar, mixed with the regular sound of the train rubbing against the track during driving, and occasionally there are fluctuations in the voices of conductors' announcements and passengers' conversations. The specific analysis method is described below and will not be elaborated here for the time being.
[0026] Based on these real-time environmental data, the system determines that the real-time scene type is inside a high-speed rail carriage. The rapid movement and specific trajectory of the spatial position data, combined with the device placement state, initially point to the scene inside a vehicle. The stability of environmental temperature, humidity, air pressure, and the specific electromagnetic interference intensity further narrow down the scope. The change trend of the noise signal in the environmental sound data, such as continuous low-frequency roar and regular friction sound, becomes the key basis for determining the high-speed rail carriage scene, thus achieving accurate scene recognition and laying a foundation for subsequent matching of appropriate noise reduction modes.
[0027] In the above steps, on the one hand, through the analysis of real-time environmental data, the real-time noise level and scene type can be accurately obtained. Based on this, the system can automatically select and switch to the most suitable target noise reduction mode for different noise conditions and scenes. For example, in a noisy construction site, the system will automatically turn on the strong noise reduction mode to effectively filter out a large amount of high-intensity environmental noise and ensure that the recorded sound is clear and distinguishable; while in a quiet scene like a library, turning off the noise reduction function can avoid audio distortion introduced by the noise reduction algorithm and completely retain every detail of the original sound, thereby comprehensively improving the recording quality and meeting the user's requirements for high-quality recording in various complex environments. Moreover, both the real-time noise level and the environmental scene type are dynamically obtained based on real-time environmental data. This means that as the user's environment changes, such as moving from a quiet indoor environment to a noisy outdoor street, the system can promptly detect the changes in the noise level and scene and quickly adjust the noise reduction mode. Compared with traditional fixed noise reduction modes or noise reduction methods that require manual switching, this method can always maintain the optimal noise reduction effect and ensure that the recording quality is not affected by environmental changes.
[0028] On the other hand, in the embodiments of the present application, during the recording process, there is no need for the user to manually adjust the noise reduction mode, and the entire analysis, judgment, and switching process can be automatically completed. This greatly simplifies the operation process and reduces the user's usage threshold. Especially for those users who are not familiar with audio technology or noise reduction functions, it provides great convenience. The user only needs to focus on the recording content itself without being distracted by dealing with noise reduction-related operations, significantly improving the user's experience and concentration during the recording process. When determining the target noise reduction mode and generating noise reduction mode parameters, in addition to considering the real-time noise level and scene type, user attributes can also be combined for personalized customization. For example, users of different ages and occupations have different sensitivities and requirements for sounds. The system can tailor the most suitable noise reduction mode for each user based on the user's historical usage data and preference settings. This personalized service not only improves the user's satisfaction with the product but also enhances the user's loyalty to the product.
[0029] On the other hand, the embodiments of the present application rely on real-time environmental data for analysis and decision-making, fully reflecting the data-driven intelligent characteristics. Through the collection and analysis of a large amount of environmental data, the noise level assessment model and the environmental scene recognition model can continuously optimize and improve their own accuracy and adaptability. With the increase in the amount of data and the continuous iteration of the model, the system can make more accurate judgments and responses to various complex environments, realizing the transformation from simple mode switching to intelligent environmental perception and adaptive processing.
[0030] Moreover, through the collaborative work of the noise level assessment model and the environmental scene recognition model in the embodiments of the present application, the intelligent advantages of multiple models can be effectively integrated. Analyze the environmental data from different perspectives and then synthesize the results, providing a comprehensive and accurate basis for determining the target noise reduction mode. This model fusion method not only improves the intelligent level of the system but also provides useful reference ideas for the design and development of other intelligent systems.
[0031] It is worth noting that the embodiments of the present application can be widely applied in multiple fields. In the education field, it can be used in scenarios such as classroom recording and lecture recording to ensure that students can clearly record the teacher's teaching content; in the business field, it can meet the needs of meeting recording, interview recording, etc., providing reliable support for business communication and information recording; in the medical field, it can be used in scenarios such as doctor-patient communication recording and case recording, helping to improve the quality and efficiency of medical services. Through applications in different fields, its technical value and social value are further expanded.
[0032] The embodiments of the present application can be combined with various user devices. Whether it is a mobile phone, a voice recorder, a tablet computer or a professional recording device, as long as the intelligent noise reduction mode switching system is integrated, the intelligent upgrade of the noise reduction function can be achieved. At the same time, with the continuous development of technology and the emergence of new environmental data types, the system can easily expand and upgrade its functions by updating the noise level evaluation model and the environmental scene recognition model to adapt to more complex environments and diverse user needs.
[0033] As an alternative embodiment, it is also possible to receive a noise reduction mode adjustment instruction sent by the user device; the noise reduction mode adjustment instruction is constructed based on parameter adjustment information input by the user in the user device. Furthermore, based on the noise reduction mode adjustment instruction and the historical scene type where the user device is located, a user individual preference library corresponding to the user device is established; wherein, the user individual preference library stores parameter adjustment strategies corresponding to the user device in each historical scene type. Then, after automatically switching the user device to the corresponding target noise reduction mode using the noise reduction mode parameter in 104, it is also possible to select a parameter adjustment strategy corresponding to the corresponding noise reduction mode parameter from the user individual preference library, and based on the selected parameter adjustment strategy, adjust the noise reduction mode parameter in the target noise reduction mode to make the recording quality meet the individual preferences of the user.
[0034] In the above embodiments, by collecting the noise reduction mode adjustment instructions actively input by the user, the personalized needs of the user for the noise reduction mode in different scenarios are deeply understood. These instructions are constructed based on the parameter adjustment information input by the user in the device, reflecting the user's preferences for aspects such as noise reduction intensity and sound detail retention. At the same time, in combination with the historical scene type where the user device is located, this information is sorted and summarized into the user individual preference library. In this library, each historical scene type corresponds to a set of specific parameter adjustment strategies, which are the concentrated embodiment of the user's preferences. After the system automatically switches the user device to the target noise reduction mode according to the real-time noise level and scene type, it will further retrieve the parameter adjustment strategy matching the current noise reduction mode from the user individual preference library. Then, according to this strategy, the parameters of the target noise reduction mode are optimized and adjusted.
[0035] Through the above embodiments, traditional noise reduction modes are often generalized and cannot meet the diverse needs of different users. However, the above embodiments can customize the noise reduction mode according to the user's past operations and preferences, so that the recording quality meets the user's unique auditory requirements. Whether the user pursues extreme quietness or hopes to retain more environmental sound details, both can be satisfied. Secondly, the adaptability and competitiveness of the product are enhanced. Among the numerous recording devices on the market, products that can provide personalized services are more likely to stand out. By meeting the personalized needs of users, the satisfaction and loyalty of users to the product are improved. Moreover, it helps with the continuous optimization of the product. By analyzing the data in the user's individual preference library, product developers can discover the general preference trends of users, and thus, in subsequent product upgrades, specifically improve the noise reduction algorithm and function settings, enabling the product to continuously evolve and better adapt to market demands.
[0036] After the recording is started, in 101, real-time environmental data in the environment where the user's device is currently located is obtained.
[0037] In the embodiments of the present application, the real-time environmental data at least includes: environmental sound data, spatial position data, device usage status, environmental temperature and humidity data, air pressure data, and electromagnetic interference intensity data. For example, in an outdoor park scene in the early morning, after the recording device is turned on, real-time environmental data is collected. The environmental sound data includes the chirping of birds, the rustling sound of the breeze blowing through the leaves, and the conversations of the morning exercise people in the distance; the spatial position data shows that the device is near a lawn in the park, and the longitude and latitude coordinates are [specific coordinate values]; the device usage status indicates that it is in the handheld recording state, the screen brightness is 50%, and the volume is the default value; the environmental temperature and humidity data is 20 degrees Celsius for temperature and 60% for relative humidity; the air pressure data is 101.325 kPa and is relatively stable; the electromagnetic interference intensity data shows that there is weak interference from the nearby street lamp circuit, but the overall is within the normal range. These data are collected in real time, providing a comprehensive and rich information basis for subsequent noise level evaluation and environmental scene recognition, facilitating the accurate switching of the intelligent noise reduction mode, ensuring that the recording quality is not overly interfered by environmental factors, making the recorded sound clear and real, and bringing a good recording experience to the user.
[0038] As an alternative embodiment, the noise level evaluation model at least includes: an extraction layer, an analysis layer, and an output layer. Based on this structure, in 102, the real-time environmental data is analyzed for noise through the noise level evaluation model to obtain the real-time noise level in the environment where the user's device is currently located, including: Through an extraction layer, environmental audio features are extracted from the real-time environmental data; through an analysis layer, real-time signal amplitude analysis processing is performed on the environmental audio features to obtain corresponding environmental noise change amplitude features; through an output layer, the environmental noise change amplitude features obtained through analysis are classified according to a preset strategy to obtain the real-time noise level.
[0039] For example, assume that the user is in a bustling market. After turning on the recording device, the real-time environmental data is transmitted into the noise level evaluation model. The extraction layer first filters out the environmental audio features from the real-time environmental data containing various information. Audio information such as the rising and falling cries of vendors, the noisy voices of the bustling crowd, and the honking of vehicles are all accurately extracted.
[0040] Next, the analysis layer starts to work and performs real-time signal amplitude analysis on the extracted environmental audio features. Through complex algorithms, the amplitude changes of each sound signal are calculated in detail. For example, the amplitude difference between a sudden high-decibel cry of a vendor and the relatively gentle voices of people chatting is calculated, so as to obtain the corresponding environmental noise change amplitude features, clearly presenting the fluctuating state of the noise at different times.
[0041] Finally, the output layer classifies the environmental noise change amplitude features obtained by the analysis layer according to the preset strategy. If the preset strategy divides the noise level into three levels: low, medium, and high, after comparison and judgment, since the environmental noise in the market has a large and frequent change amplitude, the output layer determines the real-time noise level of the current environment as "high".
[0042] From the perspective of technical effects, the noise level evaluation model with this structure can analyze environmental noise efficiently and accurately. The extraction layer ensures the acquisition of key audio information and is not interfered by other non-audio environmental data. The analysis layer deeply analyzes the audio features, providing solid data support for accurately judging the noise level. The classification mechanism of the output layer converts the complex analysis results into an intuitive and easy-to-understand noise level grade, facilitating subsequent rapid switching to appropriate noise reduction modes according to different grades, greatly improving the reaction speed and noise reduction effect of the entire intelligent noise reduction system, and ensuring a high-quality recording experience for users in various complex environments.
[0043] As an optional embodiment, in the above steps, through the analysis layer, real-time signal amplitude analysis processing is performed on the environmental audio features to obtain corresponding environmental noise change amplitude features, including: Perform noise separation processing on the environmental audio features to obtain multiple channels of environmental noise signals; perform channel separation on the separated multiple channels of environmental noise signals, and perform peak identification on each separated independent channel noise signal to determine the peak points in each independent channel noise signal; determine the real-time signal amplitude of each independent channel noise signal based on the peak points to obtain the corresponding environmental noise change amplitude feature.
[0044] Specifically, environmental audio is usually a complex signal mixed with multiple noises. Through noise separation processing, using signal processing algorithms and models, according to the characteristics of different noises, such as frequency, timbre, etc., the mixed environmental audio features are decomposed into multiple relatively independent environmental noise signals for more refined analysis later. Perform channel separation on each separated environmental noise signal, and further decompose it into noise signals of each independent channel. Then, through a specific algorithm, perform peak identification on each independent channel noise signal to find the maximum value point in the signal, that is, the peak point. These peak points can reflect the strongest amplitude of the channel noise signal at a specific moment. Based on the found peak points, the real-time signal amplitude of each independent channel noise signal at different moments can be determined, and the variation of these amplitude values constitutes the environmental noise change amplitude feature, thus comprehensively and meticulously depicting the dynamic change of environmental noise.
[0045] In this way, various noise components in environmental audio can be deeply analyzed, and the amplitude variation of each noise can be accurately obtained, providing a detailed and reliable data basis for accurately evaluating the environmental noise level, and making the system's understanding of environmental noise more accurate. It can handle complex and changing environmental noises. Whether it is the mixture of multiple different types of noises or the distribution difference of noises on different channels, it can be effectively processed, enhancing the adaptability and stability of the system for noise analysis in various scenarios. The accurate environmental noise change amplitude feature helps the intelligent noise reduction system to more precisely match and adjust the noise reduction mode, adopt more appropriate noise reduction strategies for noises of different amplitudes, thus significantly improving the noise reduction effect and the recording quality.
[0046] Furthermore, in the above steps, performing channel separation on the separated multiple channels of environmental noise signals, and performing peak identification on each separated independent channel noise signal to determine the peak points in each independent channel noise signal can be implemented as: Perform sliding time window cutting on multi-channel environmental noise signals to obtain multi-channel environmental noise signal segments under multiple sliding time windows; separate the multi-channel environmental noise signal segments under multiple sliding time windows according to channel types to obtain individual independent channel noise signals corresponding to each sliding time window of the multi-channel environmental noise signals; the channel types at least include left stereo channels, right stereo channels, and channels under different dry-wet sound levels; calculate the quartiles of each independent channel noise signal to obtain the quartile positions and quartile ranges corresponding to each independent channel noise signal; perform anomaly detection based on the quartile positions and quartile ranges corresponding to each independent channel noise signal to obtain the anomaly signal values contained in each independent channel noise signal; process the anomaly signal values in each independent channel noise signal to obtain optimized independent channel noise signals for each, and perform Hilbert transform and amplitude envelope calculation on the optimized independent channel noise signals for each to obtain the amplitude envelopes of each independent channel noise signal; perform peak identification based on the amplitude envelopes of each independent channel noise signal to obtain the peak points in each independent channel noise signal.
[0047] In the above steps, the sliding time window is a technique that divides a continuous time series signal into multiple overlapping or non-overlapping segments. For multi-channel environmental noise signals, through the sliding time window, it can be divided into multiple shorter signal segments. The size and sliding step of the sliding window can be adjusted according to specific requirements. For example, if the window size is set to w samples and the sliding step is set to s samples, then for a signal with a length of n, multiple signal segments with a length of w will be generated, and the adjacent segments are separated by s samples. This cutting method helps to transform the long sequence signal into multiple short sequence signals, facilitating subsequent processing and analysis, and can better capture the local characteristics of the signal at different time periods.
[0048] This is because environmental noise signals are usually continuous, and the characteristics of the noise signals may vary in different time segments. Through sliding time window cutting, the continuous signal can be locally processed in the time dimension, making it more convenient for the analysis and processing of local signals.
[0049] Thus, it can effectively reduce the complexity of signal processing, split the long sequence signal into multiple short segments, enabling subsequent processing to be carried out within local time windows, making it easier to implement parallel processing or distributed processing and improving the calculation efficiency. It helps to capture the local characteristics of the signal in different time windows, avoiding the loss of local information that may occur when processing the entire long sequence signal uniformly, and providing a finer-grained data basis for subsequent channel separation, anomaly detection, and other operations.
[0050] Next, the multi-channel environmental noise signal segments under multiple sliding time windows are separated according to the channel type. Specifically, according to different channel types (such as the left stereo channel, the right stereo channel, and the channels under different dry-wet sound levels), the multi-channel environmental noise signal segments under the cut sliding time windows are separated. This is based on the physical or acoustic characteristics of the channels. The noise signals of different channels have different characteristics, and through separation, the noise signals of different channels can be analyzed specifically. For stereo signals, the left and right channels may receive sounds from different directions, and their signals have different characteristics; for the channels under different dry-wet sound levels, the frequency and amplitude characteristics of their signals will also be different. Through this separation, it is convenient for subsequent separate processing of different channel signals.
[0051] Thus, independent analysis of different channels is achieved, enabling the system to more accurately identify the noise characteristics in each channel, avoiding interference between different channel signals, and improving the accuracy of noise analysis. Different subsequent processing strategies can be selected for different channels according to their characteristics. For example, for channels under different dry-wet sound levels, different noise reduction algorithms or parameters may be adopted, providing more detailed information for the final noise reduction mode switching.
[0052] Then, the quartiles are calculated for each independent channel noise signal to obtain the quartile positions and quartile ranges corresponding to each independent channel noise signal. Quartiles are statistical measures that divide the data into four equal parts. The first quartile Q1 indicates that 25% of the data is less than this value, and the third quartile Q3 indicates that 75% of the data is less than this value, and IQR = Q3 - Q1. For each independent channel noise signal, calculating its quartiles can understand the distribution of the data. By sorting the noise signals of each channel, finding Q1 and Q3, and calculating IQR, the central tendency and dispersion degree of the data can be understood. This approach helps detect the distribution range of the data and the potential range of outliers because data outside the range of Q1 - 1.5 * IQR to Q3 + 1.5 * IQR may be outliers.
[0053] Thus, the statistical characteristics of the channel noise signals, including the dispersion degree of the signals and the central position of the data, can be quickly understood through the quartile range, providing a basis for subsequent outlier detection. This helps identify possible outliers, prepares for subsequent processing of abnormal signals, and avoids interference from abnormal signals to the overall signal analysis.
[0054] Furthermore, anomaly detection is performed based on the quartile positions and quartile ranges corresponding to each independent channel noise signal to obtain the anomaly signal values contained in each independent channel noise signal. Specifically, according to the calculated Q1, Q3, and IQR, signal values below Q1 - 1.5 * IQR or above Q3 + 1.5 * IQR are regarded as anomaly signal values. This is a statistical outlier detection method. In noise signals, outliers may be caused by sudden noise interference or equipment failures, etc., which will affect the normal analysis of noise signals. For example, in a normal ambient noise signal, sudden extremely large or extremely small signal amplitudes may be due to transient interference of the equipment, and these can be detected by this method.
[0055] In this way, it is possible to effectively screen out anomaly signals that may interfere with normal signal analysis, improve the accuracy of subsequent analysis of noise signals, and make subsequent processing and analysis more stable and reliable. It provides a basis for the purification of noise signals. By removing or correcting these anomaly signals, the signal quality can be improved.
[0056] Further, the anomaly signal values in each independent channel noise signal are processed to obtain optimized independent channel noise signals. Specifically, for the detected anomaly signal values, different processing methods can be adopted. For example, they can be replaced with Q1 or Q3, or interpolated and replaced with adjacent signal values, or directly deleted. The purpose is to eliminate the influence of anomaly signals on subsequent analysis and make the signal smoother and more normal. For example, if an anomaly signal value is much higher than the normal range, replacing it with Q3 can make the amplitude distribution of the signal more reasonable and avoid the influence of outliers on subsequent analysis and processing.
[0057] In this way, a smoother and more normal signal can be obtained, avoiding the interference of outliers on subsequent signal processing steps, and improving the accuracy and stability of signal processing. It helps to improve the reliability of subsequent analysis and processing results based on these signals, and provides a better data basis for subsequent Hilbert transform and amplitude envelope calculation.
[0058] Furthermore, Hilbert transform and amplitude envelope calculation are performed on each optimized independent channel noise signal to obtain the amplitude envelope of each independent channel noise signal. It can be understood that the Hilbert transform is a method of converting a real signal into an analytic signal. For a real signal x(t), its Hilbert transform H(x(t)) can be calculated through an integral formula. The analytic signal z(t) = x(t) + j * H(x(t)), and the amplitude envelope A(t) can be calculated through A(t) = sqrt(x(t)^2 + H(x(t))^2). The Hilbert transform can extract the instantaneous amplitude information of the signal. For a noise signal, its amplitude envelope reflects the amplitude change trend of the noise signal. For an environmental noise signal, its amplitude envelope can more clearly represent the amplitude change of the noise signal at different time points, and can better reflect the actual energy change of the noise. Especially for non-stationary noise signals, it can better capture the amplitude modulation characteristics of the signal.
[0059] In this way, the amplitude envelope of the noise signal can be extracted, which can more accurately reflect the amplitude change of the noise signal, and provide more intuitive and accurate information for the amplitude analysis of the noise signal. It provides more effective data for subsequent peak identification. Because the amplitude envelope contains the amplitude information of the noise signal, it can better reflect the energy and intensity characteristics of the noise than the original signal, which helps to better understand the characteristics of the noise signal.
[0060] Finally, peak identification is performed based on the amplitude envelopes of the independent channel noise signals to obtain the peak points in each independent channel noise signal. For the amplitude envelope signal of each channel, the peak point is the local maximum in the amplitude envelope signal. By traversing the amplitude envelope signal, find the points that are larger than the adjacent points before and after, which are the peak points. These peak points reflect the maximum amplitude of the channel noise signal within a local time range and are important characteristic points of the intensity of the channel noise signal. For example, for an amplitude envelope [1, 2, 3, 2, 1], 3 is the peak point, which represents the maximum amplitude of the noise signal within this time window. In this way, the peak points of the noise signals of each channel can be accurately found, providing key information for subsequent noise analysis, such as the maximum intensity of the noise and the time position where it appears. It helps to analyze the characteristics of the noise in different channels, provides a basis for noise level evaluation and noise reduction mode switching based on peak points, improves the performance of the entire intelligent noise reduction system, and makes the selection and parameter adjustment of the noise reduction mode more targeted and accurate.
[0061] Through the above series of steps, fine analysis and processing can be performed on multi-channel environmental noise signals, from channel separation, anomaly detection to peak identification, providing comprehensive and accurate information for the intelligent noise reduction system, which helps to achieve more precise noise reduction mode switching and better noise reduction effect.
[0062] As an alternative embodiment, it is assumed that the environmental scene recognition model at least includes: a preprocessing layer, a feature extraction layer, a construction fusion layer, a prediction layer, and an output layer. Based on this, in 103, the real-time environmental data is input into the environmental scene recognition model to obtain the real-time scene type corresponding to the environment where the user device is currently located, including: Through the preprocessing layer, the real-time environmental data is preprocessed; through the feature extraction layer, multi-dimensional real-time environmental features are extracted from the real-time environmental data; the multi-dimensional real-time environmental features at least include: spatial position features, device usage status features, environmental temperature and humidity features, air pressure features, electromagnetic interference intensity features, and environmental sound features; through the construction fusion layer, the multi-dimensional real-time environmental features are projected into a multi-dimensional analysis space as corresponding feature space points, and the point cloud type matching is performed on the feature point cloud formed by the feature space points to determine the candidate scene type; through the prediction layer, the scene adaptability prediction is performed on the candidate scene type to obtain the adaptation probability corresponding to the candidate scene type; through the output layer, the real-time scene type is determined based on the adaptation probability corresponding to the candidate scene type.
[0063] In the above model, the preprocessing layer preprocesses the real-time environmental data, including but not limited to data cleaning, removing missing values and outliers. For example, the values that significantly exceed the reasonable range in the environmental temperature and humidity data are corrected or deleted. Data normalization is performed to unify the data of different dimensions to the same dimension and value range. For example, the air pressure data and the electromagnetic interference intensity data are both normalized to the [0, 1] interval for subsequent model processing. Data encoding may also be performed to convert categorical data such as device usage status into numerical forms for convenient model recognition.
[0064] In this way, the data quality is improved, the data is more suitable for model processing, the interference of abnormal data to the model is reduced, the stability and accuracy of the model are enhanced, and the learning deviation of the model caused by data inconsistency is avoided.
[0065] The feature extraction layer extracts various features from the real-time environmental data. For example, spatial position features such as longitude, latitude, altitude, and whether indoors are extracted from the spatial position data; device usage status features such as whether the device is stationary or moving and whether the screen is on are extracted from the device usage status; environmental temperature and humidity features such as temperature change rate and humidity mean are extracted from the environmental temperature and humidity data; air pressure features such as the fluctuation range of air pressure are extracted from the air pressure data; electromagnetic interference intensity features such as the frequency of interference and the intensity change trend are extracted from the electromagnetic interference intensity data; environmental sound features such as the frequency distribution and energy spectrum of sound are extracted from the environmental sound data.
[0066] In this way, the original environmental data is transformed into a feature representation valuable for scene recognition, highlighting the key information related to the scene in the data, providing a rich information basis for subsequent scene recognition, and enabling the model to better distinguish different scenes based on these features.
[0067] The fusion layer is constructed to project multi-dimensional real-time environmental features into a multi-dimensional analysis space to form feature space points, with each feature corresponding to a dimension in the space. These points form a feature point cloud, and by matching with predefined point cloud templates of different scene types, the candidate scene types are determined. For example, the similarity between the current feature point cloud and predefined point clouds representing "office scene", "outdoor street scene", etc. is calculated to find the scene type with a higher similarity as a candidate.
[0068] By fusing multi-dimensional features and using the method of feature point cloud matching, the possible scene types are initially screened, considering the information from multiple dimensions comprehensively, improving the accuracy and comprehensiveness of scene recognition, and reducing the possibility of misjudgment.
[0069] The prediction layer predicts the scene adaptability of the candidate scene types, using the knowledge learned by the model to evaluate the adaptation probability of each candidate scene type in the current environment. This may involve using machine learning algorithms such as neural networks and decision trees to train the model based on historical data, enabling it to judge the possibility of each candidate scene type according to the input features. Thus, a quantitative adaptation probability is assigned to each candidate scene type, further clarifying the matching degree of each candidate scene with the current environment, providing a more persuasive basis for finally determining the real-time scene type, and enhancing the reliability of scene recognition.
[0070] The output layer determines the real-time scene type based on the adaptation probability corresponding to the candidate scene types. Usually, the candidate scene type with the highest adaptation probability is selected as the final real-time scene type output. If the probabilities of multiple candidate scene types are close, other strategies can also be combined, such as further analyzing the feature details or referring to historical scene data to make a decision. In this way, a clear real-time scene type result is given, completing the entire environmental scene recognition process, providing an accurate basis for subsequent switching of the intelligent noise reduction mode according to the scene type, and achieving the precise adaptation of the intelligent noise reduction system to different environments.
[0071] Exemplarily, the pre-configured noise reduction mode selection strategy is to summarize the most suitable noise reduction mode under different combinations of noise levels and scene types based on a large amount of experimental data and actual usage feedback. For example, in a quiet library scene, even if there is occasional slight rustling of books, the overall noise level is low, and at this time, it is suitable to adopt a mild noise reduction mode to avoid over-noise reduction affecting sound details; while in a noisy factory workshop, the noise level is high and continuous, and a strong noise reduction mode is required to ensure clear recording. By matching the real-time obtained noise level and scene type with the conditions in the strategy, the target noise reduction mode can be determined.
[0072] This method can quickly and accurately match suitable noise reduction modes for different environments, greatly improving the pertinence of noise reduction. It avoids noise reduction in inappropriate modes. For example, using strong noise reduction in a quiet environment may cause sound distortion, or mild noise reduction in a noisy environment cannot effectively eliminate noise, thus ensuring that the recording quality is always at a high level.
[0073] Furthermore, user attributes include factors such as age, hearing preferences, and usage habits. Users of different ages have different sensitivities to sounds. For example, the elderly may be more inclined to retain more sound details; users with different hearing preferences have different requirements for noise reduction intensity and frequency range; usage habits also affect parameter settings. For example, users who often record meetings may hope to highlight the vocal frequency band. Combining the real-time noise level and scene type, these factors are comprehensively considered to generate parameters. For example, for a user who is young and prefers clear vocals in a high-noise-level environment such as an airport waiting hall (scene type), the parameters will focus on enhancing the filtering of high-frequency noise while retaining the clarity of the vocal frequency band.
[0074] Thus, personalized customization of noise reduction mode parameters is achieved. Compared with general parameter settings, the parameters generated considering user attributes can better meet the unique needs of each user, enhancing the user experience. It enables users to obtain recording effects that meet their own expectations in different environments, enhancing the applicability of the product and user satisfaction with the product.
[0075] In the above steps, the audio processing module inside the device will receive the generated noise reduction mode parameters. These parameters will control the working mode of the audio processing algorithm, such as adjusting the cut-off frequency of the filter, gain size, noise reduction intensity, etc. For example, if the parameter specifies that in the strong noise reduction mode, the noise in a specific frequency band should be attenuated by 20 dB, the audio processing module will process the input audio signal according to this parameter to achieve the switch from the current noise reduction state to the target noise reduction mode. In this way, it is ensured that the user device can accurately and efficiently switch to the noise reduction mode that meets the current environment and user needs. The device can quickly adapt to environmental changes, continuously provide users with a high-quality recording environment, ensure that the recording quality is not affected by the device switching process, and maintain the continuity and stability of the recording.
[0076] As an alternative embodiment, based on the real-time noise level and the real-time scene type, the user equipment is automatically switched to the corresponding target noise reduction mode to ensure the recording quality of the user equipment, including: Based on the real-time noise level and the real-time scene type, a pre-configured noise reduction mode selection strategy is used to determine the target noise reduction mode; based on user attributes, the real-time noise level, and the real-time scene type, noise reduction mode parameters matching the target noise reduction mode are generated; the user equipment is switched to the corresponding target noise reduction mode using the noise reduction mode parameters.
[0077] In the above steps, the target noise reduction mode is determined based on the real-time noise level and the real-time scene type. The pre-configured noise reduction mode selection strategy is to summarize the most suitable noise reduction mode under different combinations of noise levels and scene types according to a large amount of experimental data and actual usage feedback. For example, in a quiet library scene, even if there is an occasional slight sound of turning pages, the overall noise level is low, and at this time, it is suitable to use a mild noise reduction mode to avoid excessive noise reduction affecting sound details; while in a noisy factory workshop, the noise level is high and continuous, and a strong noise reduction mode is required to ensure clear recording. By matching the real-time obtained noise level and scene type with the conditions in the strategy, the target noise reduction mode can be determined. This method can quickly and accurately match the appropriate noise reduction mode for different environments, greatly improving the pertinence of noise reduction. It avoids noise reduction in an inappropriate mode, such as using strong noise reduction in a quiet environment resulting in sound distortion, or mild noise reduction in a noisy environment being unable to effectively eliminate noise, thus ensuring that the recording quality is always at a high level.
[0078] Then, noise reduction mode parameters are generated based on user attributes, the real-time noise level, and the real-time scene type. User attributes include factors such as age, hearing preference, and usage habits. Users of different ages have different sensitivities to sounds. For example, the elderly may be more inclined to retain more sound details; users with different hearing preferences have different requirements for noise reduction intensity and frequency range; usage habits also affect parameter settings. For example, users who often record meetings may hope to highlight the human voice frequency band. Combining the real-time noise level and scene type, these factors are comprehensively considered to generate parameters. For example, for a young user with a preference for clear human voices in a high-noise-level environment such as an airport waiting hall (scene type), the parameters will focus on enhancing the filtering of high-frequency noise while retaining the clarity of the human voice frequency band.
[0079] In this way, personalized customization of the noise reduction mode parameters is achieved. Compared with the general parameter settings, the parameters generated considering user attributes can better meet the unique needs of each user, improving the user experience. It enables users to obtain a recording effect that meets their own expectations in different environments, enhancing the applicability of the product and the user's satisfaction with the product.
[0080] Finally, the user device is switched to the target noise reduction mode using the noise reduction mode parameters. The audio processing module inside the device receives the generated noise reduction mode parameters. These parameters control the working mode of the audio processing algorithm, such as adjusting the cut-off frequency of the filter, the gain size, the noise reduction intensity, etc. For example, if the parameter specifies that the noise in a specific frequency band should be attenuated by 20 dB in the strong noise reduction mode, the audio processing module will process the input audio signal according to this parameter to achieve the switch from the current noise reduction state to the target noise reduction mode.
[0081] Thus, it is ensured that the user device can accurately and efficiently switch to the noise reduction mode that meets the current environment and user needs. The device can quickly adapt to environmental changes, continuously provide a high-quality recording environment for the user, ensure that the recording quality is not affected by the device switching process, and maintain the coherence and stability of the recording.
[0082] The following presents an example of an intelligent noise reduction mode switching method in the scenario of a recording pen: Scenario 1: Office scenario After the recording pen is turned on, various built-in sensors start to work. The microphone collects environmental sound data and detects occasional conversations of colleagues, keyboard typing sounds, and the faint humming sound of the air conditioner running. The overall sound is relatively stable and the volume is moderate. The position sensor indicates that the device is in a fixed position in the office, the device usage status is stationary on the desktop, the environmental temperature and humidity sensor feedbacks a temperature of 25 °C and a humidity of 50%, the barometric pressure sensor shows that the barometric pressure is stable near the standard atmospheric pressure, and the electromagnetic interference intensity data shows the existence of weak and stable electromagnetic interference from office equipment.
[0083] By analyzing the environmental sound data through the noise level evaluation model, it is concluded that the noise level is low. At the same time, based on the various data collected, the environmental scene recognition model determines that the current scene is an office.
[0084] According to the pre-set strategy, the low-noise office scenario is suitable for the mild noise reduction mode. The neural network model combines the noise level and the scene type and outputs the corresponding mild noise reduction mode parameters. For example, the update rate of the noise estimation in the spectral subtraction method is set to be slow to avoid overestimating the noise. At the same time, the parameters of the Wiener filter are fine-tuned to have a certain inhibitory effect on the low-frequency background noise without overly affecting the speech frequency band. The recording pen automatically switches to the mild noise reduction mode according to these parameters, reducing the background noise while maximizing the retention of the details of colleagues' conversations and keyboard typing sounds, ensuring that the recording quality is clear and natural and meeting the requirements for voice recording in the office scenario.
[0085] Scenario 2: Street scenario The user walks on the street with a recording pen. The microphone captures a large amount of noisy sounds, including the sounds of cars driving, horns honking, and the noise of the crowd. The sound intensity and frequency vary greatly. The position sensor shows that the device is in a moving state, and the longitude and latitude are constantly changing. The environmental temperature and humidity data indicate that the temperature is 30°C and the humidity is 40%. The air pressure is still close to the standard atmospheric pressure, but the electromagnetic interference intensity increases due to the increase in surrounding electronic devices.
[0086] The noise level assessment model analyzes that the current noise level is high, and the environmental scene recognition model determines it as a street scene.
[0087] For the street scene with high noise, the strategy selects the strong noise reduction mode. The neural network model generates the strong noise reduction mode parameters based on the noise level and scene type. For spectral subtraction, the update speed of noise estimation is accelerated to adapt to the rapidly changing noise environment; for Wiener filtering, the attenuation coefficient for high-frequency noise is increased. The recording pen quickly switches to the strong noise reduction mode according to these parameters, effectively filtering out most of the street noise and highlighting useful audio such as human voices, ensuring that the recorded content can still be clearly recognized in a noisy environment.
[0088] Scenario 3: Adaptive optimization according to user habits During a period of use, the recording pen records the user's operation habits of the noise reduction mode in different scenarios. For example, in the gym scenario, the user often manually adjusts the noise reduction mode to appropriately reduce the music and fitness equipment sounds while retaining the voices of people around. By analyzing these operations, the recording pen learns the user's habit of preferring to retain a certain ambient sound in the gym scenario.
[0089] When it is detected again that the device is in the gym scenario and the noise level is high, in addition to considering the scene and noise level, the neural network model also combines this user habit to generate special noise reduction mode parameters. On the basis of adopting the strong noise reduction mode, the parameters are fine-tuned to reduce the noise reduction intensity of some ambient audio segments, effectively reducing noise interference while meeting the user's need to retain a certain ambient sound and further improving the user experience.
[0090] In the embodiments of this application, through the collaborative work of the noise level assessment model and the environmental scene recognition model, the environmental data is analyzed from different perspectives, effectively improving the adaptability of the user device to the surrounding environment, assisting the user device to maintain the optimal noise reduction effect, and ensuring that the recording quality is not disturbed by environmental changes.
[0091] Figure 2 The following is a schematic structural diagram of an intelligent noise reduction mode switching system provided by the embodiments of this application, as Figure 2 shown. The system includes the following steps: An acquisition unit, configured to acquire real-time environmental data in the current environment of the user device after the recording is started; the real-time environmental data at least includes: environmental sound data, spatial location data, device usage status, environmental temperature and humidity data, air pressure data, and electromagnetic interference intensity data; An analysis unit, configured to perform noise analysis on the real-time environmental data through a noise level evaluation model to obtain the real-time noise level in the current environment of the user device; the real-time noise level is associated with the change amplitude of the noise signal in the real-time environmental data; An identification unit, configured to input the real-time environmental data into an environmental scene identification model to obtain the real-time scene type corresponding to the current environment of the user device; the real-time scene type is determined based on any one or more of the real-time spatial changes reflected by the spatial location data, device usage status, environmental temperature and humidity data, air pressure data, and electromagnetic interference intensity data, and the change trend of the noise signal in the environmental sound data; A switching unit, configured to automatically switch the user device to a corresponding target noise reduction mode based on the real-time noise level and the real-time scene type to ensure the recording quality of the user device.
[0092] Further optionally, it further includes a preference unit, configured to receive a noise reduction mode adjustment instruction issued by the user device; the noise reduction mode adjustment instruction is constructed based on parameter adjustment information input by the user in the user device; based on the noise reduction mode adjustment instruction and the historical scene type where the user device is located, a user individual preference library corresponding to the user device is established; wherein, the user individual preference library stores parameter adjustment strategies corresponding to the user device in each historical scene type.
[0093] Furthermore, after the switching unit automatically switches the user device to the corresponding target noise reduction mode using the noise reduction mode parameters, the preference unit is further configured to select the parameter adjustment strategy corresponding to the corresponding noise reduction mode from the user individual preference library, and adjust the noise reduction mode parameters in the target noise reduction mode based on the selected parameter adjustment strategy to make the recording quality meet the individual preferences of the user.
[0094] Further optionally, the noise level evaluation model at least includes: an extraction layer, an analysis layer, and an output layer; The analysis unit performs noise analysis on the real-time environmental data through the noise level evaluation model to obtain the real-time noise level in the current environment of the user device. Specifically, it is configured to: Extract environmental audio features from the real-time environmental data through the extraction layer; Perform real-time signal amplitude analysis processing on the environmental audio features through the analysis layer to obtain corresponding environmental noise change amplitude features; Through the output layer, classify the environmental noise change amplitude characteristics obtained by analysis according to a preset strategy to obtain the real-time noise level.
[0095] Further optionally, the analysis unit performs real-time signal amplitude analysis processing on the environmental audio characteristics through an analysis layer to obtain corresponding environmental noise change amplitude characteristics, specifically for: Perform noise separation processing on the environmental audio characteristics to obtain multiple channels of environmental noise signals; Perform channel separation on the multiple channels of environmental noise signals obtained by separation, and perform peak identification on the separated independent channel noise signals to determine the peak points in each independent channel noise signal; Based on the peak points, determine the real-time signal amplitude of each independent channel noise signal to obtain corresponding environmental noise change amplitude characteristics.
[0096] Further optionally, the analysis unit performs channel separation on the multiple channels of environmental noise signals obtained by separation, and performs peak identification on the separated independent channel noise signals to determine the peak points in each independent channel noise signal, specifically for: Perform sliding time series window cutting on the multiple channels of environmental noise signals to obtain multiple segments of multiple channels of environmental noise signals under multiple sliding time series windows; Perform separation processing on the multiple segments of multiple channels of environmental noise signals under multiple sliding time series windows according to the channel type to obtain each independent channel noise signal corresponding to each sliding time series window of the multiple channels of environmental noise signals; The channel types at least include the left stereo channel, the right stereo channel, and channels under different dry and wet sound levels; Perform quartile calculation on each independent channel noise signal to obtain the quartile positions and quartile ranges corresponding to each independent channel noise signal; Perform anomaly detection based on the quartile positions and quartile ranges corresponding to each independent channel noise signal to obtain the anomaly signal values included in each independent channel noise signal; Process the anomaly signal values in each independent channel noise signal to obtain optimized each independent channel noise signal, and perform Hilbert transform and amplitude envelope calculation on the optimized each independent channel noise signal to obtain the amplitude envelope of each independent channel noise signal; Perform peak identification according to the amplitude envelope of each independent channel noise signal to obtain the peak points in each independent channel noise signal.
[0097] Further optionally, the environmental scene recognition model at least includes: a preprocessing layer, a feature extraction layer, a construction and fusion layer, a prediction layer, and an output layer; An identification unit inputs the real-time environmental data into an environmental scene identification model to obtain the real-time scene type corresponding to the environment where the user device is currently located, specifically for: Preprocess the real-time environmental data through a preprocessing layer; Extract multi-dimensional real-time environmental features from the real-time environmental data through a feature extraction layer; the multi-dimensional real-time environmental features at least include: spatial position features, device usage status features, environmental temperature and humidity features, air pressure features, electromagnetic interference intensity features, environmental sound features; Through a constructed fusion layer, project the multi-dimensional real-time environmental features into a multi-dimensional analysis space as corresponding feature space points, and perform point cloud type matching on the feature point cloud formed by the feature space points to determine candidate scene types; Through a prediction layer, perform scene adaptability prediction on the candidate scene types to obtain the adaptation probability corresponding to the candidate scene types; Through an output layer, determine the real-time scene type based on the adaptation probability corresponding to the candidate scene types.
[0098] Further optionally, a switching unit automatically switches the user device to the corresponding target noise reduction mode based on the real-time noise level and the real-time scene type to ensure the recording quality of the user device, specifically for: Based on the real-time noise level and the real-time scene type, determine the target noise reduction mode by using a pre-configured noise reduction mode selection strategy; Generate noise reduction mode parameters matching the target noise reduction mode based on user attributes, the real-time noise level, and the real-time scene type; Switch the user device to the corresponding target noise reduction mode by using the noise reduction mode parameters.
[0099] Please refer to Figure 3 , Figure 3 which is a schematic diagram of an embodiment of an electronic device provided by an embodiment of the present application. As Figure 3 shown, an embodiment of the present application provides an electronic device 500, including a memory 510, a processor 520, and a computer program 511 stored on the memory 510 and executable on the processor 520. When the processor 520 executes the computer program 511, the foregoing embodiments are implemented.
[0100] Please refer to Figure 4 , Figure 4 which is a schematic diagram of an embodiment of a computer-readable storage medium provided by an embodiment of the present application. As Figure 4 shown, this embodiment provides a computer-readable storage medium 600, on which a computer program 611 is stored. When the computer program 611 is executed by a processor, the foregoing embodiments are implemented.
[0101] It should be noted that in the above embodiments, the descriptions of each embodiment have their own emphases. For parts not described in detail in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0102] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0103] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded computer, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for realizing the functions specified in Figure 1 one or more of the processes or multiple processes and / or blocks Figure 1 one or more of the blocks.
[0104] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means realizes the functions specified in Figure 1 one or more of the processes or multiple processes and / or blocks Figure 1 one or more of the blocks.
[0105] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for realizing the functions specified in Figure 1 one or more of the processes or multiple processes and / or blocks Figure 1 one or more of the blocks.
[0106] Although the preferred embodiments of the present invention have been described, additional changes and modifications can be made to these embodiments by those skilled in the art once they learn of the basic inventive concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications that fall within the scope of the present invention.
[0107] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.
Claims
1. An intelligent noise reduction mode switching method, characterized in that, The method at least includes: After the recording is turned on, obtaining real-time environment data in the environment where the user device is currently located; the real-time environment data at least includes: environmental sound data, spatial position data, device usage status, environmental temperature and humidity data, air pressure data, electromagnetic interference intensity data; Performing noise analysis on the real-time environment data through a noise level evaluation model to obtain the real-time noise level in the environment where the user device is currently located; the real-time noise level is associated with the change amplitude of the noise signal in the real-time environment data; Inputting the real-time environment data into an environment scene recognition model to obtain the real-time scene type corresponding to the environment where the user device is currently located; the real-time scene type is determined based on any one or more of the real-time spatial changes reflected by the spatial position data, device usage status, environmental temperature and humidity data, air pressure data, electromagnetic interference intensity data, and the change trend of the noise signal in the environmental sound data; Based on the real-time noise level and the real-time scene type, automatically switching the user device to the corresponding target noise reduction mode to ensure the recording quality of the user device.
2. The intelligent noise reduction mode switching method according to claim 1, wherein The method further includes: Receiving a noise reduction mode adjustment instruction issued by the user device; the noise reduction mode adjustment instruction is constructed based on parameter adjustment information input by the user in the user device; Based on the noise reduction mode adjustment instruction and the historical scene type where the user device is located, establishing a user individual preference library corresponding to the user device; wherein, the user individual preference library stores parameter adjustment strategies corresponding to the user device in each historical scene type; After automatically switching the user device to the corresponding target noise reduction mode using the noise reduction mode parameters, it further includes: Selecting the parameter adjustment strategy corresponding to the noise reduction mode from the user individual preference library, and adjusting the noise reduction mode parameters in the target noise reduction mode based on the selected parameter adjustment strategy to make the recording quality meet the individual preferences of the user.
3. The intelligent noise reduction mode switching method according to claim 1, wherein The noise level evaluation model at least includes: an extraction layer, an analysis layer, and an output layer; The performing noise analysis on the real-time environment data through a noise level evaluation model to obtain the real-time noise level in the environment where the user device is currently located includes: Extracting environmental audio features from the real-time environment data through the extraction layer; Performing real-time signal amplitude analysis processing on the environmental audio features through the analysis layer to obtain the corresponding environmental noise change amplitude features; Classifying the analyzed environmental noise change amplitude features according to a preset strategy through the output layer to obtain the real-time noise level.
4. The intelligent noise reduction mode switching method according to claim 3, wherein The performing real-time signal amplitude analysis processing on the environmental audio features through the analysis layer to obtain the corresponding environmental noise change amplitude features includes: Performing noise separation processing on the environmental audio features to obtain multiple channels of environmental noise signals; Performing channel separation on the separated multiple channels of environmental noise signals, and performing peak recognition on the separated independent channel noise signals to determine the peak points in each independent channel noise signal; Determining the real-time signal amplitude of each independent channel noise signal based on the peak points to obtain the corresponding environmental noise change amplitude features.
5. The intelligent noise reduction mode switching method according to claim 4, wherein, Performing channel separation on the separated multi-channel environmental noise signals, and performing peak recognition on each separated independent channel noise signal to determine the peak points in each independent channel noise signal, including: Performing sliding time-series window cutting on the multi-channel environmental noise signals to obtain multi-channel environmental noise signal segments under multiple sliding time-series windows; Performing separation processing on the multi-channel environmental noise signal segments under multiple sliding time-series windows according to the channel types to obtain each independent channel noise signal corresponding to each sliding time-series window of the multi-channel environmental noise signals; the channel types at least include the left stereo channel, the right stereo channel, and the channels under different dry-wet sound levels; Calculating the quartiles of each independent channel noise signal to obtain the quartile positions and quartile ranges corresponding to each independent channel noise signal; Performing outlier detection based on the quartile positions and quartile ranges corresponding to each independent channel noise signal to obtain the outlier signal values included in each independent channel noise signal; Processing the outlier signal values in each independent channel noise signal to obtain the optimized each independent channel noise signal, and performing Hilbert transform and amplitude envelope calculation on the optimized each independent channel noise signal to obtain the amplitude envelope of each independent channel noise signal; Performing peak recognition according to the amplitude envelope of each independent channel noise signal to obtain the peak points in each independent channel noise signal.
6. The intelligent noise reduction mode switching method according to claim 1, wherein, The environmental scene recognition model at least includes: a preprocessing layer, a feature extraction layer, a construction fusion layer, a prediction layer, and an output layer; Inputting the real-time environmental data into the environmental scene recognition model to obtain the real-time scene type corresponding to the environment where the user device is currently located, including: Performing preprocessing on the real-time environmental data through the preprocessing layer; Extracting multi-dimensional real-time environmental features from the real-time environmental data through the feature extraction layer; the multi-dimensional real-time environmental features at least include: spatial position features, device usage status features, environmental temperature and humidity features, air pressure features, electromagnetic interference intensity features, and environmental sound features; Projecting the multi-dimensional real-time environmental features into a multi-dimensional analysis space as corresponding feature space points through the construction fusion layer, and performing point cloud type matching on the feature point cloud formed by the feature space points to determine the candidate scene types; Performing scene adaptability prediction on the candidate scene types through the prediction layer to obtain the adaptation probability corresponding to the candidate scene types; Determining the real-time scene type based on the adaptation probability corresponding to the candidate scene types through the output layer.
7. The intelligent noise reduction mode switching method according to claim 1, wherein Based on the real-time noise level and the real-time scene type, automatically switching the user device to the corresponding target noise reduction mode to ensure the recording quality of the user device, including: Determining the target noise reduction mode based on the real-time noise level and the real-time scene type by using a pre-configured noise reduction mode selection strategy; Generating noise reduction mode parameters matching the target noise reduction mode based on user attributes, the real-time noise level, and the real-time scene type; Switching the user device to the corresponding target noise reduction mode by using the noise reduction mode parameters.
8. An intelligent noise reduction mode switching system, characterized in that, The system at least includes the following units: An acquisition unit, configured to acquire real-time environmental data in the current environment where the user device is located after the recording is turned on; The real-time environmental data at least includes: environmental sound data, spatial location data, device usage status, environmental temperature and humidity data, air pressure data, electromagnetic interference intensity data; An analysis unit, configured to perform noise analysis on the real-time environmental data through a noise level evaluation model to obtain the real-time noise level in the current environment where the user device is located; the real-time noise level is associated with the change amplitude of the noise signal in the real-time environmental data; An identification unit, configured to input the real-time environmental data into an environmental scene identification model to obtain the real-time scene type corresponding to the current environment where the user device is located; the real-time scene type is determined based on any one or more of the real-time spatial changes reflected by the spatial location data, device usage status, environmental temperature and humidity data, air pressure data, electromagnetic interference intensity data, and the change trend of the noise signal in the environmental sound data; A switching unit, configured to automatically switch the user device to the corresponding target noise reduction mode based on the real-time noise level and the real-time scene type to ensure the recording quality of the user device.
Citation Information
Patent Citations
Sound mode switching method and apparatus
CN107147794A
Noise reduction control method and device and computer readable storage medium
CN112004174A
Audio processing method and microphone equipment
CN113825068A
Audio noise reduction method and device, equipment and storage medium
CN116962940A
Environment scene recognition and decision-making system and method based on intelligent equipment
CN117877522A