A noise detection method, electronic device and storage medium

By modeling the speaker cavity using a physical information neural network and detecting speaker noise using a psychoacoustic model, the noise problem when the terminal device plays high-volume audio is solved, achieving efficient noise suppression and audio quality improvement.

CN120431952BActive Publication Date: 2026-04-21HONOR DEVICE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HONOR DEVICE CO LTD
Filing Date
2024-10-30
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately detect and suppress noise generated by the speaker of a terminal device when playing high-volume audio, thus affecting the audio listening experience.

Method used

A speaker cavity model is modeled using a Physical Information Neural Network (PINN), combined with a psychoacoustic model and a sound quality evaluation model based on feature extraction. By detecting the vibration velocity and sound pressure signal of the audio signal, noise can be identified and suppressed.

Benefits of technology

It improves the accuracy and suppression of noise, making the output audio signal imperceptible during playback and enhancing the user's listening experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120431952B_ABST
    Figure CN120431952B_ABST
Patent Text Reader

Abstract

This application provides a noise detection method, electronic device, and storage medium, relating to the field of terminal technology. The method includes: the terminal device first inputs the target audio signal to be played by the user into a classification network to obtain its signal category. Then, the target audio signal is input into a speaker unit model to obtain the corresponding vibration velocity and sound pressure level signals. Next, the vibration velocity and sound pressure level signals are input into a speaker cavity model constructed using a PINN network to obtain the target sound pressure level signal that can be received by the human ear. Furthermore, by utilizing a pre-constructed sound quality evaluation model, combined with the signal category of the audio signal and the sound pressure level signal output by the speaker cavity model, the noise level of the target audio signal can be detected, obtaining a more accurate detection result. Therefore, when the detection result indicates the presence of noise, the noise contained in the target audio signal can be suppressed in a timely manner to improve the user's listening experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of terminal technology, and in particular to a noise detection method, electronic device, and storage medium. Background Technology

[0002] With the continuous development of technology, terminal devices, represented by mobile phones and tablets, are increasingly used in people's lives and work, bringing great convenience. For example, people can use these terminal devices to play music, watch videos, and conduct social communication.

[0003] Currently, when users play audio signals using mobile phones or tablets, the tiny built-in speakers produce noticeable hissing and buzzing noises, especially when playing high-volume audio signals like pianos or similar instruments. These noises disrupt the original timbre of the audio, negatively impacting the overall listening experience. Therefore, accurately detecting and effectively suppressing these noises during audio playback is a pressing issue that needs to be addressed. Summary of the Invention

[0004] To address the aforementioned issues, this application provides a noise detection method, electronic device, and storage medium. The purpose is to detect noise in audio signals that a user intends to play using a mobile phone or tablet computer, and to effectively suppress the noise contained in the audio signal in a timely manner when the detection results indicate the presence of noise. This ensures that the suppressed audio signal will not have perceptible noise issues during playback, thereby improving the user's listening experience.

[0005] Firstly, this application provides a noise detection method, which includes: when a user clicks on a music software on a mobile phone or other terminal device to play music, the terminal device first inputs the target audio signal to be played into a classification network to obtain the signal category to which the target audio signal belongs, such as drum sound, vocals, or piano. Then, the target audio signal is input into a speaker unit model to obtain the vibration velocity and sound pressure level signals corresponding to the target audio signal, thereby mapping the digital audio signal onto the vibration velocity and sound pressure level signals. Next, the vibration velocity and sound pressure level signals are input into a speaker cavity model constructed by a PINN network to obtain the actual propagation (received by the human ear) target sound pressure level signal output by the model. Furthermore, by using a pre-constructed sound quality evaluation model, combined with the signal category to which the audio signal belongs and the sound pressure level signal output by the speaker cavity model, the noise level of the target audio signal can be detected to obtain the detection result; and when the detection result indicates that there is a noise problem, the noise contained in the target audio signal is promptly suppressed.

[0006] As can be seen, in the above noise detection method, this embodiment pre-constructs a speaker cavity model based on the PINN network, which improves the simulation accuracy. It also constructs a sound quality evaluation model by combining psychoacoustic models, feature extraction scoring models, and objective index calculation models. This ensures that the noise detection results output by the model are consistent with subjective listening experience, and can be used as a basis for judging whether the target audio signal that the user wants to play has noise. This results in more accurate detection results. Furthermore, when the detection results indicate that the target audio signal has noise problems, the noise contained in the audio signal can be effectively suppressed in a timely manner, so that the suppressed audio signal will not have perceptible noise problems when played, which helps to improve the user's listening experience.

[0007] In one possible implementation, the loudspeaker unit model is a mathematical model; then the target audio signal is input into the loudspeaker unit model to obtain the vibration velocity and sound pressure signal corresponding to the target audio signal, including: inputting the target audio signal into the loudspeaker unit model to solve for displacement, obtaining the diaphragm displacement of the loudspeaker; performing derivative calculation on the displacement to obtain the vibration velocity and sound pressure signal corresponding to the target audio signal, thereby demonstrating the mapping effect of the loudspeaker unit model on the data.

[0008] In one possible implementation, the loudspeaker cavity model is obtained by modeling using the Physical Information Neural Network (PINN).

[0009] In one possible implementation, the loudspeaker cavity model is constructed as follows: First, sample vibration velocity and sound pressure level (SPL) signals are acquired inside the loudspeaker cavity; the PINN model is trained using these signals and a first loss constraint function to obtain the initially trained PINN model; wherein the first loss constraint function includes a first difference equation loss, a first initial condition loss, a first boundary condition loss, and a first data fitting loss; second, sample vibration velocity and SPL signals are acquired at the location of the loudspeaker cavity's sound outlet; the initially trained PINN model is then trained using the second sample vibration velocity and SPL signals and a second loss constraint function to obtain a single loudspeaker unit model; wherein the second loss constraint function includes a second difference equation loss, a second initial condition loss, a second boundary condition loss, and a second data fitting loss. This improves the simulation effect of the loudspeaker cavity model.

[0010] In one possible implementation, the first initial condition loss and the first boundary condition loss are calculated based on the vibration velocity and sound pressure signal of the first sample; the second initial condition loss and the second boundary condition loss are calculated based on the vibration velocity and sound pressure signal of the second sample.

[0011] In one possible implementation, obtaining the first sample vibration velocity and sound pressure level signals inside the loudspeaker cavity includes: using finite element and / or boundary element analysis to simulate and model the loudspeaker cavity, obtaining the time-varying signals of sound pressure and particle velocity distribution inside the loudspeaker cavity under different types of excitation sources, as the first sample vibration velocity and sound pressure level signals; or, proportionally enlarging the internal structure of the loudspeaker cavity, and measuring the distribution signals of sound pressure and particle velocity inside the cavity, walls, and corresponding sound outlet holes at different times within the enlarged model, as the first sample vibration velocity and sound pressure level signals. This can improve the accuracy of the loudspeaker cavity model construction.

[0012] In one possible implementation, the sound quality evaluation model includes a psychoacoustic perception model and a convolutional neural network (CNN). Using this pre-built sound quality evaluation model, the noise level of the target audio signal is detected based on the signal category of the target sound pressure signal and the target audio signal. The detection results include: inputting the target sound pressure signal into the psychoacoustic model within the sound quality evaluation model for signal type conversion to obtain the corresponding perceptual domain signal; inputting the spectral signal corresponding to the perceptual domain signal into the CNN network for high-dimensional feature extraction; inputting the extracted high-dimensional features into a fully connected layer to predict the perceptual domain score; calculating the total harmonic distortion (THD) and signal-to-noise ratio (SNR) of the target audio signal and the target sound pressure signal; and weighting and summing the perceptual domain score, THD, and SNR based on the signal category of the target audio signal to obtain the noise level score corresponding to the target audio signal, which serves as the detection result after detecting the noise level of the target audio signal. This improves the accuracy of the detection results.

[0013] In one possible implementation, the sound quality evaluation model is constructed as follows: A training sound pressure level (SPL) signal corresponding to the training audio signal output from a pre-built loudspeaker cavity model is obtained; the training SPL signal is input into a psychoacoustic model for filtering to obtain a training receptive domain signal corresponding to the training SPL signal; the training receptive domain signal is weighted using initial filtering weights, and the spectral signal corresponding to the weighted receptive domain signal is input into a CNN network for high-dimensional feature extraction. The extracted high-dimensional training features are then input into a fully connected layer to predict the training receptive domain score; the training THD and training SNR values ​​of the training audio signal and training SPL signal are calculated; based on the signal category of the training audio signal, the training receptive domain score, training THD value, and training SNR value are weighted and summed using floating weights to obtain the noise level score corresponding to the training audio signal; by calculating the difference between the actual noise score and the noise level score corresponding to the training audio signal, the initial filtering weights, floating weights, and other model parameters are continuously updated to train and generate the sound quality evaluation model. This improves the accuracy of the sound quality evaluation model's evaluation results.

[0014] In one possible implementation, when the detection result indicates the presence of noise, noise suppression processing is performed promptly on the target audio signal. This includes: when the detection result indicates the presence of noise, performing a Fourier transform on the target audio signal to obtain the signal spectrum in the frequency domain; and finding the peak value of the signal spectrum in the frequency domain to reduce noise by lowering the peak value, thereby achieving noise suppression. This improves the noise suppression effect on the target audio signal.

[0015] Secondly, this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor is used to call and execute the computer program to implement the noise detection method described in any one of the first aspects above.

[0016] Thirdly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when run by a processor of an electronic device, is used to implement the noise detection method described in any one of the first aspects above.

[0017] Fourthly, this application provides a computer program product that, when run on a computer, causes the computer to perform the noise detection method as described in any one of the first aspects. Attached Figure Description

[0018] Figure 1 This is a schematic diagram of a scenario provided for an embodiment of this application;

[0019] Figure 2This is a schematic diagram illustrating the overall implementation process of noise detection and suppression provided in the embodiments of this application;

[0020] Figure 3 A schematic diagram of the terminal device provided in the embodiments of this application;

[0021] Figure 4 This is a software structure block diagram of a terminal device provided in an embodiment of this application;

[0022] Figure 5 A flowchart of the noise detection method provided in the embodiments of this application;

[0023] Figure 6 This is a schematic diagram of the network structure of the PINN model provided in the embodiments of this application;

[0024] Figure 7 This is a schematic diagram illustrating the construction process of the sound quality evaluation model provided in the embodiments of this application. Detailed Implementation

[0025] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. The terminology used in the following embodiments is for the purpose of describing specific embodiments only and is not intended to be a limitation of this application. As used in the specification and appended claims of this application, the singular expressions "a," "an," "the," "the," "the," and "this" are intended to also include expressions such as "one or more," unless the context clearly indicates otherwise.

[0026] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0027] The "multiple" mentioned in the embodiments of this application refers to two or more. It should be noted that in the description of the embodiments of this application, terms such as "first" and "second" are used only for the purpose of distinguishing descriptions and should not be construed as indicating or implying relative importance, nor should they be construed as indicating or implying order.

[0028] To enable those skilled in the art to better understand the solution of this application, the application scenario of the technical solution of this application will be described first.

[0029] See Figure 1 The illustration shows a scenario diagram provided by an embodiment of this application.

[0030] In this example scenario, when a user plays music using a music app on their phone, the built-in miniature speaker, due to its tiny size, produces noticeable hissing and buzzing noises when playing high-volume audio signals such as pianos or piano-like instruments. This is because pianos or piano-like instruments have strong transients and rich harmonic energy, concentrated in specific frequency bands. This energy easily excites the miniature speaker diaphragm to vibrate at high speeds, causing turbulence in the speaker cavity and impacting the inner wall of the cavity. This results in cavity resonance or friction at the sound outlet, producing severe and unpleasant noise.

[0031] Experiments have shown that the shape of a loudspeaker cavity has a significant impact on airflow noise. Therefore, the common approach to address noise issues is to study the loudspeaker cavity to reduce noise. However, current research on loudspeaker cavities is mainly based on simulation modeling, which is computationally intensive and complex. While offline modeling is possible, it cannot meet the requirements of real-time computation. Specifically, existing schemes for modeling loudspeaker cavities are mostly based on finite element method (FEM) and boundary element method (BEM) analysis. The FEM divides the solution domain into a finite number of elements, constructs local approximate solutions on each element, and then approximates the solution for the entire solution domain by combining these local solutions. Its drawback is the large computational cost, and generating high-quality meshes is difficult for complex geometries (such as loudspeaker cavity shapes). The BEM, on the other hand, divides the boundary of the solution domain into a finite number of boundary elements, constructs discrete equations on the boundaries, solves for the physical quantities on the boundaries, and then solves for the physical quantities within the domain through boundary integration. Its drawback is that the application of the BEM is limited for complex internal structures and nonlinear problems, and the boundary integration calculation is complex for high-dimensional problems.

[0032] Furthermore, the subjective manifestation (i.e., human hearing perception) of noise in current miniature loudspeakers is difficult to correlate with objective indicators. Therefore, it is impossible to judge the noise level solely based on objective indicators. For example, in the case of high total harmonic distortion (THD), it does not necessarily mean that the human ear will subjectively hear obvious noise.

[0033] Therefore, how to accurately detect the noise generated during the playback of audio by the speaker, ensure that the output audio effect matches the subjective hearing of the human ear, and effectively suppress the noise contained in the audio signal to improve the user's listening experience is a technical problem that urgently needs to be solved.

[0034] To overcome the above technical problems, this application provides a noise detection (and suppression) method. For example... Figure 2 As shown, the audio signal to be detected is first input into a classification network to obtain the signal category, such as human voice or piano. Then, the audio signal is input into a speaker unit model to obtain the speaker diaphragm displacement. The diaphragm velocity (vibration velocity), acceleration, and sound pressure level are further calculated by differentiating the displacement, thus mapping the digital audio signal to displacement, acceleration, and sound pressure level signals. Next, the vibration velocity and sound pressure level signals are input into the speaker cavity model to obtain the actual propagating (received by the human ear) sound pressure level signal output by the model. The speaker cavity model is modeled using a Physical-Informed Neural Network (PINN). Subsequent embodiments will describe this model in detail. Compared to the finite element method and boundary element method, the PINN network does not require meshing, thus it can handle the modeling of complex geometries (speaker cavities) and can utilize the high-dimensional characteristics of neural networks to handle high-dimensional problems, improving simulation accuracy. Furthermore, a pre-built sound quality evaluation model can be used, combined with the signal category of the audio signal and the sound pressure level signal output by the speaker cavity model, to evaluate the noise level of the audio signal. The sound quality evaluation model is a comprehensive model combining a psychoacoustic model, a feature extraction scoring model, and an objective index calculation model. It outputs a score for the noise level of the audio signal, serving as the noise detection result to determine whether it meets preset indicators (such as whether the score is higher than a preset score threshold). Subsequent embodiments will describe this model in detail. Thus, when the detection result indicates no noise problem, i.e., the score meets the preset indicators (e.g., the score is higher than the preset score threshold), the audio signal can be directly output. Conversely, when the detection result indicates the presence of noise problem, such as a low noise level score that does not meet the preset indicators (e.g., the score is lower than the preset score threshold), a noise suppression algorithm can be used to process the audio signal under test, ensuring that the output (suppressed) audio signal has no perceptible noise problem, thereby improving the user's listening experience.

[0035] In addition, to further ensure the auditory effect of the output audio signal, such as Figure 2As shown, the audio signal processed by the noise suppression algorithm can be mapped again through the feedback link using the speaker unit model and the speaker cavity model. The noise level can then be scored again based on the mapping result using the sound quality evaluation model. Based on the score result, it can be determined whether noise suppression or other processing steps should be performed, thereby improving the noise detection and suppression effect of the audio signal.

[0036] It should be noted that the noise detection method provided in this application embodiment can be applied to terminal devices (also referred to as electronic devices) such as mobile phones, tablets, personal digital assistants (PDAs), desktop, laptop, and notebook computers, ultra-mobile personal computers (UMPCs), handheld computers, netbooks, and wearable devices.

[0037] To enable those skilled in the art to better understand the noise detection method provided in this application, the hardware architecture and software system architecture of the electronic device implementing the noise detection method will be described in detail below.

[0038] See Figure 3 The diagram shows a schematic of the terminal device provided in the embodiments of this application.

[0039] like Figure 3 As shown, the terminal device 300 may include a processor 310, a mobile communication module 320, a wireless communication module 330, a sensor module 340, a display screen 350, an internal memory 360, a camera 370, an audio module 380, a speaker 380A, a receiver 380B, a microphone 380C, a headphone jack 380D, an antenna group 1, and an antenna group 2.

[0040] It is understood that the structure illustrated in the embodiments of this application does not constitute a specific limitation on the terminal device 300. In other embodiments of this application, the terminal device 300 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0041] The processor 310 may include one or more processing units, such as an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural network processing unit (NPU). These different processing units can be independent devices or integrated into one or more processors. The controller can generate operation control signals based on instruction opcodes and timing signals to control instruction fetching and execution. For example, after acquiring the audio signal the user wants to play using the terminal device 300, and determining the noise level detection result of the target audio signal using a speaker unit model, speaker cavity model, and sound quality evaluation model, the controller can promptly suppress the noise contained in the target audio signal when the detection result indicates the presence of noise, thereby improving the user's listening experience.

[0042] The processor 310 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 310 is a cache memory. This memory can store instructions or data that the processor 310 has just used or that are used repeatedly. If the processor 310 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 310, and thus improves the efficiency of the system.

[0043] In some embodiments, the processor 310 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0044] The sensor module 340 can be used to acquire data signals related to various aspects of the terminal device 300, serving as a basis for implementing corresponding functions. In some embodiments, the sensor module 340 may include, but is not limited to, image sensors, gyroscope sensors, barometric pressure sensors, acoustic pressure sensors, accelerometers, temperature sensors, pressure sensors, etc.

[0045] The display screen 350 is used to display images, videos, etc. The display screen 350 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a miniature LED, a microLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the terminal device 300 may include one or P display screens 350, where P is a positive integer greater than 1.

[0046] The internal memory 360 can be used to store executable program code, including instructions. The internal memory 360 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound capture, image capture, etc.). The data storage area may store data created during the use of the terminal device 300 (such as audio data, image data, etc.). Furthermore, the internal memory 360 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory, universal flash storage (UFS), etc. The processor 310 executes various functional applications and data processing of the terminal device 300 by running instructions stored in the internal memory 360 and / or instructions stored in memory located within the processor.

[0047] In some embodiments, the internal memory 360 stores instructions for performing a noise detection method. The processor 310 can execute the instructions stored in the internal memory 360 to perform the following functions: the terminal device 200 first acquires the audio signal (hereinafter referred to as the target audio signal) to which the user wants to play the noise level to be detected, and classifies the target audio signal to determine the signal category to which the target audio signal belongs, such as drum sound, human voice, or piano. Then, the target audio signal is input into the speaker unit model to obtain the vibration velocity and sound pressure level signals corresponding to the target audio signal. The vibration velocity and sound pressure level signals are then input into the pre-built speaker cavity model to obtain the sound pressure level signal when the speaker plays the target audio signal (hereinafter referred to as the target sound pressure level signal). The speaker cavity model is obtained by PINN modeling. Next, using the pre-built sound quality evaluation model, the noise level of the target audio signal is detected according to the target sound pressure level signal and the signal category to which the target audio signal belongs (such as drum sound, human voice or piano). The detection result is obtained, and when the detection result indicates that there is a noise problem, the noise contained in the target audio signal is suppressed in time, so that the output suppressed target audio signal has no perceptible noise problem, thereby improving the user's listening experience.

[0048] The camera 370 is used to capture still images or videos. For example, after a user uses the handheld terminal device 300, they can use the camera 370 installed on the terminal device 300 to take images or videos. In some embodiments, the terminal device 300 may include one or K cameras 370, where K is a positive integer greater than 1.

[0049] The terminal device 300 implements display functions through a GPU, a display screen 350, and an application processor. The GPU is a microprocessor for image processing, connected to the display screen 350 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. The processor 310 may include one or more GPUs, which execute program instructions to generate or modify display information.

[0050] Terminal device 300 can implement audio functions through audio module 380, speaker 380A, receiver 380B, microphone 380C, headphone jack 380D, and application processor, such as music playback and recording for voice input and output.

[0051] The audio module 380 is used to convert digital audio information into analog audio signals for output, and also to convert analog audio input into digital audio signals. The audio module 380 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 380 may be located in the processor 310, or some functional modules of the audio module 380 may be located in the processor 310.

[0052] The speaker 380A, also known as a "loudspeaker," refers to the "loudspeaker" or "miniature loudspeaker" mentioned in this application, used to convert audio electrical signals into sound signals for audio playback. The terminal device 300 can listen to music or make hands-free calls through the speaker 380A.

[0053] The receiver 380B, also known as the "earpiece," is used to convert audio electrical signals into sound signals. When the terminal device 300 answers a phone call or voice message, the receiver 380B can be brought close to the listener's ear to hear the voice.

[0054] Microphone 380C, also known as a "microphone" or "voice transducer," is used to convert sound signals into electrical signals. When making a phone call or sending a voice message, the user can speak by bringing their mouth close to microphone 380C, inputting the sound signal into microphone 380C. Terminal device 300 may be equipped with at least one microphone 380C. In some embodiments, terminal device 300 may be equipped with two microphones 380C, which, in addition to collecting sound signals, can also perform noise reduction. In other embodiments, electronic device 300 may be equipped with three, four, or more microphones 380C, which can collect sound signals, reduce noise, identify the sound source, and perform directional recording, etc.

[0055] The 380D headphone jack is used to connect wired headphones and does not restrict the standard attributes of the jack.

[0056] It is understood that the interface connection relationships between the modules illustrated in the embodiments of this application are merely illustrative and do not constitute a structural limitation on the terminal device 300.

[0057] The wireless communication function of the terminal device 300 can be implemented through antenna 1, antenna 2, mobile communication module 320, wireless communication module 330, modem processor and baseband processor.

[0058] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in terminal device 300 can be used to cover one or more communication frequency bands. Different antennas can also be multiplexed to improve antenna utilization. For example, antenna 1 can be multiplexed as a diversity antenna for a wireless local area network. In some other embodiments, the antennas can be used in conjunction with a tuning switch.

[0059] The mobile communication module 320 can provide solutions for wireless communication, including 2G / 3G / 4G / 5G, applied to the terminal device 300. The mobile communication module 320 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 320 can receive electromagnetic waves via antenna 1, and perform filtering, amplification, and other processing on the received electromagnetic waves before transmitting them to a modem processor for demodulation. The mobile communication module 320 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via antenna 1. In some embodiments, at least some functional modules of the mobile communication module 320 may be housed in the processor 310. In some embodiments, at least some functional modules of the mobile communication module 320 and at least some modules of the processor 310 may be housed in the same device.

[0060] The wireless communication module 330 can provide solutions for wireless communication applications on the terminal device 300, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies. The wireless communication module 330 can be one or more devices integrating at least one communication processing module. The wireless communication module 330 receives electromagnetic waves via antenna 2, performs frequency modulation and filtering of the electromagnetic wave signals, and sends the processed signal to processor 310. The wireless communication module 330 can also receive signals to be transmitted from processor 310, perform frequency modulation and amplification, and convert them into electromagnetic waves for radiation via antenna 2.

[0061] In addition, the terminal device 300 runs an operating system on top of the aforementioned components. Examples include iOS, Android, and Windows operating systems. Applications can be installed and run on this operating system.

[0062] See Figure 4 It shows a schematic diagram of the software structure of the terminal device provided in the embodiments of this application.

[0063] The software system of terminal device 300 can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture. This application embodiment uses the layered architecture Android system as an example to exemplify the software structure of terminal device 300.

[0064] A layered architecture divides software into several layers, each with a clear role and function. Layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom: the application layer, the application framework layer, the Android runtime and system libraries, and the kernel layer.

[0065] The application layer can include a series of application packages (APPs). For example... Figure 4 As shown, the application package may include applications such as camera, call, navigation, WLAN, Bluetooth, and gallery. When the user's handheld terminal device 300 is taking pictures, it can communicate with camera-related devices in the camera access interface of the frame layer through the camera application to request camera functions and obtain image data, etc.

[0066] The application layer can include a series of application packages (APPs). For example... Figure 4 As shown, the application package can include applications such as call, music, video, WLAN, camera, and calendar apps. When a user uses the terminal device 300 to play music, they can click on the music app on the terminal device 300, select the music they want to play, and then pass the user's click command through the corresponding interface to the noise detection algorithm in the application framework layer for subsequent information processing.

[0067] The application framework layer provides application programming interfaces (APIs) and a programming framework for applications within the application layer. The application framework layer includes predefined functions. For example... Figure 4 As shown, the application framework layer may include a window manager, a phone manager, a resource manager, a noise detection algorithm, etc.

[0068] The window manager is used to manage window applications. It can obtain the screen size, determine if a status bar is present, lock the screen, and capture screenshots, among other things.

[0069] The phone manager is used to provide communication functions for electronic devices 300. For example, it manages call status (including answering and ending video calls).

[0070] The file explorer provides applications with various resources, such as localized strings, icons, images, layout files, video files, and more.

[0071] Noise detection algorithms are used to determine the noise level of the target audio signal by using speaker unit models, speaker cavity models, and sound quality evaluation models. When the detection results show that there is noise, the noise contained in the target audio signal can be suppressed in a timely manner, thereby improving the user's listening experience.

[0072] The Android Runtime consists of core libraries and a virtual machine. The Android runtime is responsible for the scheduling and management of the Android system.

[0073] The core library consists of two parts: one part is the functionalities that need to be called by the Java language, and the other part is the Android core library.

[0074] The application layer and application framework layer run in a virtual machine. The virtual machine executes the Java files of the application layer and application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.

[0075] System libraries can include multiple functional modules. For example: surface manager, media libraries, 3D graphics processing libraries (e.g., OpenGL ES), 2D graphics engines (e.g., SGL), etc.

[0076] The Surface Manager is used to manage the display subsystem and provides the blending of 2D and 3D layers for multiple applications.

[0077] The media library supports playback and recording of various common audio and video formats, as well as still image files. It supports multiple audio and video encoding formats, such as MPEG4, H.264, MP3, AAC, AMR, JPG, and PNG.

[0078] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.

[0079] A 2D graphics engine is a graphics engine for 2D drawing.

[0080] The kernel layer is the layer between hardware and software. The kernel layer contains at least the display driver, camera driver, audio driver, and sensor driver.

[0081] The technical solutions involved in the following embodiments can all be implemented in electronic devices with the above-described hardware and software architectures.

[0082] The following section will detail the specific implementation process of the noise detection method provided in this application:

[0083] like Figure 5 As shown, the specific implementation process of this noise detection method may include the following steps S501-S504:

[0084] S501: Acquire the target audio signal to be detected; classify the target audio signal, and determine the signal category to which the target audio signal belongs.

[0085] In this embodiment, when a user clicks on a music app on a mobile phone or other terminal device to play music, such as Figure 1 As shown, any audio signal that the user wants to play can be defined as the target audio signal that needs to be detected for noise. By executing steps S501-S504 in this application, noise detection is performed on the acquired target audio signal so that when the detection result shows that there is a noise problem, the noise contained in the target audio signal can be suppressed in time, thereby improving the user's hearing effect.

[0086] It should be noted that this embodiment does not limit the language type of the target audio signal. For example, the target audio signal can be audio composed of Chinese or English. At the same time, this embodiment does not limit the length of the target audio signal. For example, the target audio signal can be audio lasting several seconds or audio lasting several minutes.

[0087] Specifically, after acquiring the target audio signal to be detected, the target audio signal can first be classified using existing or future audio classification methods to determine its signal category. For example, the target audio signal can be input into a classification network such as a Convolutional Neural Network (CNN) or a Support Vector Machine (SVM) to classify the target audio signal and output its signal category (which can be a specific category or a percentage of probability of the category), such as drum sound, human voice, piano, etc., for subsequent step S504. It is understood that the range of signal categories to which the target audio signal belongs can be preset based on actual conditions and empirical values, and the specific content is not limited.

[0088] S502: Input the target audio signal into the speaker unit model to obtain the vibration velocity and sound pressure signal corresponding to the target audio signal.

[0089] In this embodiment, after the terminal device obtains the target audio signal for noise level detection in step S501, further, to improve the noise detection effect of the target audio signal, the target audio signal can be input into the speaker unit model to map this digital signal onto displacement, velocity, and sound pressure signals, thus obtaining the speaker diaphragm displacement corresponding to the target audio signal. Then, by differentiating the displacement, the diaphragm velocity, acceleration, and sound pressure signals can be further calculated for subsequent step S503.

[0090] Specifically, one possible implementation is that the speaker unit model can be a mathematical model. In this case, the specific implementation process of step S502 can be as follows: first, input the target audio signal into the speaker unit model to solve for the displacement, and obtain the diaphragm displacement of the speaker. The specific calculation formula is as follows:

[0091]

[0092] in, BL represents the driving force; U represents the voltage of the input target audio signal; Q ms K represents the mechanical quality factor of the horn. ms M represents the stiffness coefficient; ms The total vibrating mass of the speaker is represented by f; the frequency of the input target audio signal is represented by f0; the resonant frequency of the speaker is represented by f0; and the diaphragm displacement of the speaker is represented by x.

[0093] Then, by differentiating the obtained diaphragm displacement x of the loudspeaker, the vibration velocity and sound pressure signal corresponding to the target audio signal can be obtained.

[0094] S503: Input the vibration velocity and sound pressure signals into the pre-built loudspeaker cavity model to obtain the target sound pressure signal when the loudspeaker plays the target audio signal.

[0095] It should be noted that most existing methods for modeling loudspeaker cavities are based on finite element method (FEM) and boundary element method (BEM) analysis. The FEM method suffers from high computational cost, and generating high-quality meshes is difficult for complex geometries (such as loudspeaker cavity shapes). The BEM method, on the other hand, is limited in its application to complex internal structures and nonlinear problems, and its boundary integral calculations are complex for high-dimensional problems. Therefore, existing loudspeaker cavity modeling methods struggle to achieve satisfactory modeling results. To address this, this application proposes a method using a Physical Information Neural Network (PINN) for modeling loudspeaker cavities. This PINN network is a general-purpose function approximator that can embed information about specific physical laws during the learning process. These laws satisfy a given dataset and are described in the form of partial differential equations (also known as difference equations).

[0096] Specifically, the PINN network embeds physical laws (such as partial differential equations and boundary conditions) into the loss function of the neural network, combining partial differential equations and neural networks to effectively model complex structures and solve physical problems using deep learning techniques. The PINN network can handle various physical phenomena, such as fluid dynamics, heat conduction, and electromagnetism. Compared to the finite element method and boundary element method, the PINN network does not require mesh generation, thus it can handle complex geometries (such as the shape of a speaker cavity). Furthermore, due to the high-dimensional nature of neural networks, the PINN network has a significant advantage in handling high-dimensional problems. Moreover, the PINN network can combine observational data and physical models, improving simulation accuracy through data assimilation techniques, and utilize transfer learning techniques to quickly adapt to different geometries (e.g., the shape of the speaker cavity may differ for different mobile phone models).

[0097] Based on this, in this embodiment, after the terminal device inputs the target audio signal into the speaker unit model in step S502 to obtain the vibration velocity and sound pressure signal corresponding to the target audio signal, in order to achieve the ideal noise detection effect, the vibration velocity and sound pressure signal can be further input into the pre-constructed speaker cavity model for simulation to predict the sound pressure signal (here defined as the target sound pressure signal) when the speaker plays the target audio signal (received by the human ear), so as to execute the subsequent step S504 and realize the high real-time and high-precision noise level detection of the target audio signal.

[0098] Next, this embodiment will describe in detail the construction process of the loudspeaker cavity model. The specific construction process may include the following steps A1-A4:

[0099] Step A1: Obtain the first sample vibration velocity and sound pressure signal inside the speaker cavity.

[0100] It should be noted that while this application utilizes a PINN network to model the speaker cavity, which improves the modeling effect, the neural network relies on high-quality training data. Current sensor sizes are far larger than the miniature speakers built into mobile phones and other terminal devices (e.g., the cavity size of a miniature speaker is typically on the order of millimeters), while the sensor size is usually larger than the cavity size. Therefore, it is impossible to directly measure the data inside the speaker cavity using a sensor. However, the data inside the cavity is crucial for network training. Therefore, to construct the speaker cavity model, this application provides the following two methods for obtaining training data:

[0101] One approach is to use finite element and / or boundary element analysis to simulate and model the loudspeaker cavity, obtain the time-varying signals of sound pressure and particle velocity distribution inside the loudspeaker cavity under different types of excitation sources, and define them as the first sample velocity and sound pressure signals as model training data.

[0102] Another approach involves scaling up the internal structure of the loudspeaker cavity proportionally and measuring the distribution signals of sound pressure and particle velocity inside the cavity, walls, and corresponding sound outlet at different times within the scaled-up model. These signals serve as the first sample of velocity and sound pressure signals. The method used for scaling up the internal structure of the loudspeaker cavity is not limited and can be chosen based on practical considerations and experience. For example, the internal structure can be scaled up using 3D printing, and sensors can be attached to specific locations inside the structure during assembly to measure pressure, flow velocity, and other data, which can then be used as training data for the model.

[0103] Step A2: Train the PINN model using the first sample vibration velocity and sound pressure signal and the first loss constraint function to obtain the PINN model after the first training; wherein, the first loss constraint function includes the first difference equation loss, the first initial condition loss, the first boundary condition loss and the first data fitting loss.

[0104] In this embodiment, firstly, a PINN network can be selected as the initial speaker cavity model, and the model parameters can be initialized, for example... Figure 6 The PINN network shown is illustrated. It should be noted that this embodiment does not limit the specific network structure of the model; it can be a PINN network with any composition.

[0105] It should be noted that the initial training process of the PINN network incorporates partial differential equations (i.e., difference equations) constraints into the loss function, compared to that of a typical neural network. The network's input consists of two variables: time variable t and spatial variable x (used for mesh generation). The output has three variables: velocity u in the x-axis direction, velocity v in the y-axis direction, and pressure p.

[0106] Here, we take the Navier-Stokes equations for a two-dimensional, steady, incompressible fluid as an example:

[0107] The mass conservation equation is as follows:

[0108]

[0109] The momentum equation in the x-axis direction is as follows:

[0110]

[0111] The momentum equation in the y-axis direction is as follows:

[0112]

[0113] Where u and v represent the velocities in the x-axis and y-axis directions, respectively; ρ represents the fluid density; p represents the pressure; and μ is the fluid viscosity coefficient, the specific value of which is not limited and can be set according to actual conditions and empirical values. The role of the PINN network is to find the solution to the above equation, as shown below:

[0114]

[0115] Where f represents the activation function in the neural network; W represents the weights of the neural network; and b represents the bias in the neural network. The specific value of b is not limited and can be set according to the actual situation and experience.

[0116] During model training, a first sample of vibration velocity and sound pressure signal can be extracted from the training data sequentially as model input, and multiple rounds of model training can be performed. This allows for adjustments to the model parameters (such as...) based on the value of the first loss constraint function. Figure 6 The parameters (θ, λ, etc.) are updated until a preset condition is met, such as the first loss constraint function (e.g., ...). Figure 6 If the value of L in the model is very small and basically unchanged (or less than the preset threshold), then the update of the model parameters is stopped, so that the model parameters change from random to relatively fixed, completing the training of the PINN model and obtaining the PINN model after the first training.

[0117] It should be noted that the loss constraint function for training the PINN network differs from general loss constraint functions. This loss constraint function mainly consists of two parts: a physical information part and a data part. For example... Figure 6 As shown, L PDE L IC L BC These represent the difference equation loss, initial condition loss, and boundary condition loss, respectively. These three error losses are constrained by physical conditions. data This represents the data fitting loss, which is constrained by the data acquisition. Each error loss has a different weight, which can be denoted as w1, w2, w3, and w4 respectively.

[0118] It should also be noted that L PDE The value of L is determined based on the above-mentioned mass conservation equation, the momentum equation in the x-axis direction, and the momentum equation in the y-axis direction. IC The value of L BC The values ​​are all calculated based on the vibration velocity and sound pressure signal output by the loudspeaker unit model. Specifically, L IC and L BC The calculation formula is as follows:

[0119]

[0120] Where, x s This indicates that the horizontal (x) coordinates of the speaker diaphragm surface need to be specified.

[0121] Specifically, one optional implementation is that during the initial training of the PINN model, since the first sample vibration velocity and sound pressure signal are obtained through the two methods mentioned in step A1, rather than directly measuring the sound hole position at the micro-speaker's outlet, the initial training of the PINN model can be completed using a first loss constraint function, including the first difference equation loss, the first initial condition loss, the first boundary condition loss, and the first data fitting loss (i.e., focusing more on physical condition constraints), and the model's network parameters can be updated. The first initial condition loss and the first boundary condition loss are calculated based on the first sample vibration velocity and sound pressure signal. During training, the weights of each loss can be set to satisfy: w1 + w3 + w4 = m, w2 = 1 - m, where m > th1, and the values ​​of m and th1 are not limited; they can be set to values ​​between 0 and 1 based on actual conditions and experience. For example, m and th1 can be set to 0.4 and 0.1 respectively.

[0122] Step A3: Obtain the second sample vibration velocity and sound pressure signal at the location of the speaker cavity's sound outlet.

[0123] The second sample vibration velocity and sound pressure signal can be data such as sound pressure and particle vibration velocity measured at the sound outlet of the miniature loudspeaker, and the specific acquisition method is not limited.

[0124] Step A4: Using the second sample vibration velocity and sound pressure signal and the second loss constraint function, perform transfer learning training on the PINN model after the first training to obtain the loudspeaker unit model; wherein, the second loss constraint function includes the second difference equation loss, the second initial condition loss, the second boundary condition loss and the second data fitting loss.

[0125] After obtaining the initially trained PINN model through step A2, the model can be further retrained using transfer learning, utilizing the data on the location of the micro-speaker's sound outlet obtained in step A3 (i.e., the vibration velocity and sound pressure signal of the second sample). Then, high-quality data is recorded in an anechoic chamber using a speaker array or binaural recording to further fit the transfer-learned PINN model, thus achieving high-precision model training even in the absence of data from the micro-speaker's internal components.

[0126] Specifically, one possible implementation is that, during the training process of transfer learning for the PINN model, a second loss constraint function (such as...) can be used, which includes the second difference equation loss, the second initial condition loss, the second boundary condition loss, and the second data fitting loss. Figure 6 In L=w1L PDE +w2L data +w3L IC +w4L BC This is used to complete the secondary training of the PINN model (i.e., focusing more on data signal fitting constraints) and update the model's network parameters. The second initial condition loss and the second boundary condition loss are calculated based on the vibration velocity and sound pressure signals of the second sample. During training, data fitting is dominant, and the second data fitting loss L can be set to represent this. data The weights satisfy: w2 > th1. Thus, through two training iterations, not only is the problem of obtaining high-quality data from inside the miniature loudspeaker solved, but the training effect on the loudspeaker cavity model is also improved.

[0127] Furthermore, after the loudspeaker cavity model is trained, if the loudspeaker cavity structure changes, the model can be quickly adjusted using the transfer learning aspects mentioned in the above construction process, thereby enabling it to quickly adapt to different loudspeaker cavity structures.

[0128] Based on this, after obtaining the loudspeaker cavity model using a physical information neural network (PINN), the vibration velocity and sound pressure signals output by the loudspeaker unit model obtained in step S502 can be input into the pre-built loudspeaker cavity model to obtain the sound pressure signal when the loudspeaker plays the target audio signal, and define it as the target sound pressure signal for executing the subsequent step S504.

[0129] S504: Using a pre-built sound quality evaluation model, the noise level of the target audio signal is detected based on the target sound pressure signal and the signal category to which the target audio signal belongs, and the detection results are obtained; when the detection results show that there is a noise problem, the noise contained in the target audio signal is suppressed in a timely manner.

[0130] In this embodiment, after the terminal device determines the signal category of the target audio signal in step S501 and obtains the target sound pressure signal output by the speaker cavity model in step S503, in order to achieve the ideal noise detection effect, it can further utilize a pre-constructed sound quality evaluation model to detect the noise level of the target audio signal based on the target sound pressure signal and the signal category of the target audio signal, and obtain the detection result. If the detection result indicates that there is no noise problem, the target audio signal does not need to be processed and can be directly output; conversely, if the detection result indicates that there is a noise problem, a noise suppression algorithm needs to be used in a timely manner to suppress the noise contained in the target audio signal, so that the output suppressed target audio signal has no perceptible noise problem, thereby improving the user's listening experience.

[0131] Specifically, one possible implementation is that the sound quality evaluation model can include, but is not limited to, psychoacoustic perception models and convolutional neural networks (CNNs). The implementation process of "using the pre-built sound quality evaluation model to detect the noise level of the target audio signal according to the signal category to which the target sound pressure signal and the target audio signal belong, and obtaining the detection result" in step S504 can specifically include the following steps S5041-S5044:

[0132] S5041: Input the target sound pressure signal into the psychoacoustic model in the sound quality evaluation model to perform signal type conversion and obtain the perceptual domain signal corresponding to the target sound pressure signal.

[0133] In this implementation, after the terminal device obtains the target sound pressure signal output by the speaker cavity model in step S503, it can further input the target sound pressure signal into the N (specific values ​​are not limited, and can be positive integers greater than 0) filter groups of the psychoacoustic model in the sound quality evaluation model for filtering processing, converting the time-domain signal of the target sound pressure signal into a perceptual domain signal for execution of subsequent step S5042. The specific composition structure of the psychoacoustic model is not limited in this application and can be selected according to actual conditions and experience; for example, the ERB model, Bark model, etc., can be selected as the psychoacoustic model.

[0134] S5042: Input the spectral signal corresponding to the receptive field signal into the CNN network for high-dimensional feature extraction; and input the extracted high-dimensional features into the fully connected layer to predict the receptive field score.

[0135] In this implementation, after obtaining the perceptual domain signal corresponding to the target sound pressure signal in step S5041, the spectral signal corresponding to the perceptual domain signal (the specific acquisition method of the spectrum is limited, and it can be obtained by using existing or future spectrum extraction methods) can be further input into the CNN network in the sound quality evaluation model for high-dimensional feature extraction. The extracted high-dimensional features are then input into the fully connected layer added at the end of the model to predict the perceptual domain score, denoted as M1, which is used to execute the subsequent step S5044.

[0136] The value of the perception domain score M1 is not limited and can be set according to the actual situation and experience. For example, it can be set to between 1 and 5 points, where 4 and 5 points indicate no noise problem, 3 points indicate the presence of noise problem, and 1 and 2 points indicate the presence of obvious perceptible noise problem.

[0137] S5043: Calculate the total harmonic distortion (THD) and signal-to-noise ratio (SNR) values ​​of the target audio signal and the target sound pressure signal.

[0138] In this implementation, in order to improve the noise detection effect of the target audio signal, objective indicators of the target audio signal and the target sound pressure signal (including but not limited to the total harmonic distortion (THD) value (represented as M2) and the signal-to-noise ratio (SNR) value (represented as M3)) can be calculated as part of the detection basis to perform the subsequent step S5044.

[0139] The values ​​of objective noise evaluation indicators such as THD value M2 and SNR value M3 are not limited and can be set according to actual conditions and experience. For example, they can be converted to the same dimension as the perceptual domain score M1, such as being set to between 1 and 5 points. Here, 4 and 5 points indicate no noise problem, 3 points indicate the presence of noise problem, and 1 and 2 points indicate the presence of obvious perceptible noise problem.

[0140] S5044: Based on the signal category to which the target audio signal belongs, the perceptual domain score, THD value, and SNR value are weighted and summed to obtain the noise level score corresponding to the target audio signal, which is used as the detection result obtained after detecting the noise level of the target audio signal.

[0141] In this implementation, after obtaining the perceptual domain score M1 in step S5042 and the objective indicators of noise evaluation such as THD value M2 and SNR value M3 in step S5043, the objective indicators of noise evaluation such as the perceptual domain score M1, THD value M2, and SNR value M3 can be weighted and summed according to the signal category to which the target audio signal belongs. For example, if the input target audio signal is human dialogue, the weight of SNR value M3 can be increased; if the input target audio signal is a piano signal, the weight of THD value M2 can be increased, thereby obtaining a more accurate noise level score corresponding to the target audio signal. The specific calculation formula is as follows:

[0142] M = β1M1 + β2M2 + ... + β n M n

[0143] Where M represents the more accurate noise level score corresponding to the target audio signal; M1 represents the perceptual domain score; M2 represents the THD value; M3 represents the SNR value; n represents the number of objective indicators for noise evaluation, and the specific value is not limited; M n Indicates the value of the nth objective indicator; β1, β2…β n M1, M2...M n The weights are not limited to any specific values.

[0144] Furthermore, an alternative implementation is that when the more accurate noise level score M (i.e., the detection result) corresponding to the target audio signal is lower than a preset threshold (the specific value is not limited and can be set according to the actual situation and experience), it indicates that there is a noise problem in the target audio signal. At this time, it is necessary to perform a Fourier transform on the target audio signal in a timely manner to obtain the signal spectrum in the frequency domain; and find the peak value of the signal spectrum in the frequency domain in order to reduce the noise by reducing the peak value, thereby achieving the suppression of the noise contained in the target audio signal.

[0145] Next, this embodiment will describe in detail the construction process of the sound quality evaluation model. The specific construction process may include the following steps B1-B6:

[0146] Step B1: Obtain the training sound pressure signal corresponding to the training audio signal output by the pre-built loudspeaker cavity model.

[0147] It should be noted that since noise evaluation often relies on subjective listening, which is time-consuming and labor-intensive, this application proposes a noise evaluation model that combines a psychoacoustic perception model with objective indicators for noise evaluation. When constructing this model, the input data is the training sound pressure level (SPL) signal corresponding to the output training audio signal constructed by the PINN network. This SPL signal is a time-domain signal. After processing such as equivalent rectangular bandwidth filtering and feature extraction, the output is the subjective noise perception score corresponding to the current input training audio signal. This score can be used to guide the noise suppression algorithm to process the noise quality of the current input training audio signal to achieve the goal of subjectively perceiving no noise.

[0148] Therefore, in order to construct a sound quality evaluation model, this application needs to obtain a large number of training audio signals and their training sound pressure signals output through the speaker cavity model in advance as model training data. The specific acquisition method is not limited, and the noise true score (represented by MOS) corresponding to these training audio signals is manually labeled.

[0149] Step B2: Input the training sound pressure signal into the psychoacoustic model for filtering to obtain the training perceptual domain signal corresponding to the training sound pressure signal.

[0150] In this embodiment, firstly, a psychoacoustic model (containing N sets of filters) and a CNN network can be selected as the initial sound quality evaluation model, and the model parameters are initialized, for example... Figure 7 The sound quality evaluation model is shown. It should be noted that this embodiment does not limit the specific network structure of the model.

[0151] Then, during model training, a training sound pressure signal can be extracted from the training data and input into the N filters of the psychoacoustic model for filtering to obtain the training perceptual domain signal corresponding to the training sound pressure signal (i.e., the N filtered signals), which is then used to execute the subsequent step B3.

[0152] Step B3: Using the initial filter weights, the training receptive field signal is weighted, and the spectral signal corresponding to the weighted receptive field signal is input into the CNN network for high-dimensional feature extraction. Then, the extracted high-dimensional training features are input into the fully connected layer to predict the training receptive field score.

[0153] In this embodiment, after obtaining the training perceptual domain signal (i.e., the N filtered signals) corresponding to the training sound pressure signal through step B2, the initial filtering weights of the N filters (as follows) can be further utilized. Figure 7 The α1, α2…α shown N The specific values ​​are not limited. The N sets of filtered signals (i.e., training receptive field signals) are weighted and processed. The spectrum signal corresponding to the weighted receptive field signal (the specific method of obtaining the spectrum is limited, and it can be obtained by existing or future spectrum extraction methods) is input into the CNN network for high-dimensional feature extraction. The extracted training high-dimensional features are then input into the fully connected layer to predict the training receptive field score, which can also be represented by M1, and used to execute the subsequent step B4.

[0154] Step B4: Calculate the training THD and training SNR values ​​of the training audio signal and the training sound pressure signal.

[0155] The calculation methods for the training THD and training SNR values ​​of the training audio signal and training sound pressure signal are not limited, and the training THD value can still be represented by M2 and the training SNR value can still be represented by M3. It is understood that the values ​​of other objective indicators can also be calculated, all of which are used as part of the detection criteria to perform the subsequent step B5.

[0156] Step B5: Based on the signal category to which the training audio signal belongs, use floating weights to perform a weighted summation of the training perceptual domain score, the training THD value, and the training SNR value to obtain the noise level score corresponding to the training audio signal.

[0157] After obtaining the training perceptual domain score M1 in step B3 and the objective indicators of noise evaluation such as the training THD value M2 and the training SNR value M3 in step B4, the objective indicators of noise evaluation, such as the training perceptual domain score M1, the training THD value M2, and the training SNR value M3, can be further weighted and summed according to the signal category to which the training audio signal belongs. For example, if the input training audio signal is human dialogue, the weight of the training SNR value M3 can be increased; if the input training audio signal is a piano signal, the weight of the training THD value M2 can be increased, thereby obtaining the final noise level score corresponding to the training audio signal (still represented by M). Figure 7 As shown, the specific formula for calculating M can still be expressed as: M = β1M1 + β2M2 + ... + β n M n .

[0158] Step B6: By calculating the difference between the actual noise score corresponding to the training audio signal and the noise level score corresponding to the training audio signal, continuously update the values ​​of the initial filter weights, floating weights, and other model parameters to train and generate a sound quality evaluation model.

[0159] After obtaining the noise level score M corresponding to the training audio signal in step B5, the noise level score MOS corresponding to the manually labeled training audio signal can be compared with the noise level score M corresponding to the training audio signal obtained in step B5. Based on the difference between the two (i.e., Loss = MOS - M), the initial filter weights (α1, α2...α...) are adjusted. N ), floating weights (β1, β2…β) n The model updates other model parameters until the preset conditions are met, such as the loss value being very small and basically unchanged. Then the model parameter updates are stopped, the training of the sound quality evaluation model is completed, and a trained sound quality evaluation model is generated.

[0160] It is understandable that the sound quality evaluation model constructed in this application can not only detect the noise level of the target audio signal, but also evaluate other aspects, such as spatial sense, or timbre and pitch. Then, it can increase spatial sense or improve the basic performance of sound such as timbre and pitch through algorithms.

[0161] In this way, after obtaining the target audio signal that the user wants to play using a terminal device (such as a mobile phone), the processing steps S501-S504 above can be executed. By utilizing the combined effects of a classification network, a speaker unit model, a speaker cavity model (pre-built based on PINN), and a sound quality evaluation model (pre-built based on psychoacoustic perception models and CNN networks, etc.), the detection result of the noise level corresponding to the target audio signal can be determined more accurately. Thus, when the detection result shows that there is a noise problem in the target audio signal, the noise contained in the target audio signal can be suppressed in a timely manner, thereby improving the user's listening experience.

[0162] Furthermore, this application also provides an electronic device (i.e., a terminal device). For details regarding the hardware structure and software framework of the electronic device, please refer to [link to relevant documentation]. Figure 3 and Figure 4 The corresponding explanation is as follows: The electronic device includes a memory and a processor. The memory stores a computer program, and the processor calls and executes the computer program to implement the noise detection method provided in the above description.

[0163] This application also provides a computer-readable storage medium storing a computer program thereon, which, when run by the processor of a terminal device, is used to implement the noise detection method described above.

[0164] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A noise detection method, characterized in that, Applied to a terminal device, the method includes: Acquire the target audio signal to be detected; and classify the target audio signal to determine the signal category to which the target audio signal belongs; The target audio signal is input into the speaker unit model to obtain the vibration velocity and sound pressure signal corresponding to the target audio signal; The vibration velocity and sound pressure signals are input into a pre-constructed loudspeaker cavity model to obtain the target sound pressure signal when the loudspeaker plays the target audio signal; Using a pre-built sound quality evaluation model, the noise level of the target audio signal is detected based on the target sound pressure signal and the signal category to which the target audio signal belongs, and the detection result is obtained; when the detection result indicates that there is a noise problem, the noise contained in the target audio signal is suppressed in a timely manner; The loudspeaker cavity model is obtained by modeling the loudspeaker cavity using a Physical Information Neural Network (PINN); the construction method of the loudspeaker cavity model is as follows: Acquire the first sample vibration velocity and sound pressure signal inside the speaker cavity; The PINN model is trained using the first sample vibration velocity and sound pressure signal and the first loss constraint function to obtain the PINN model after the first training; the first loss constraint function includes the first difference equation loss, the first initial condition loss, the first boundary condition loss and the first data fitting loss. Acquire the second sample vibration velocity and sound pressure signal at the location of the speaker cavity's sound outlet; Using the second sample vibration velocity and sound pressure signal, as well as the second loss constraint function, the PINN model after the first training is trained by transfer learning to obtain the loudspeaker unit model; the second loss constraint function includes the second difference equation loss, the second initial condition loss, the second boundary condition loss, and the second data fitting loss.

2. The method according to claim 1, characterized in that, The loudspeaker unit model is a mathematical model; the step of inputting the target audio signal into the loudspeaker unit model to obtain the vibration velocity and sound pressure signal corresponding to the target audio signal includes: The target audio signal is input into the speaker unit model for displacement calculation to obtain the speaker diaphragm displacement; The vibration velocity and sound pressure signal corresponding to the target audio signal are obtained by differentiating the diaphragm displacement.

3. The method according to claim 1, characterized in that, The first initial condition loss and the first boundary condition loss are calculated based on the vibration velocity and sound pressure signal of the first sample; the second initial condition loss and the second boundary condition loss are calculated based on the vibration velocity and sound pressure signal of the second sample.

4. The method according to claim 1, characterized in that, The acquisition of the first sample vibration velocity and sound pressure signal inside the loudspeaker cavity includes: Using the finite element and / or boundary element analysis methods, the loudspeaker cavity is simulated and modeled to obtain the time-varying signals of sound pressure and particle velocity distribution inside the loudspeaker cavity under different types of excitation sources, which are used as the first sample velocity and sound pressure signals. Alternatively, the internal structure of the loudspeaker cavity can be enlarged proportionally, and the distribution signals of sound pressure and particle velocity inside the cavity, walls, and corresponding sound outlet holes at different times can be measured in the enlarged model internal structure as the first sample velocity and sound pressure signals.

5. The method according to claim 1, characterized in that, The sound quality evaluation model includes a psychoacoustic perception model and a convolutional neural network (CNN); the step of using the pre-constructed sound quality evaluation model to detect the noise level of the target audio signal based on the target sound pressure level and the signal category to which the target audio signal belongs, and obtaining the detection results, includes: The target sound pressure signal is input into the psychoacoustic model in the sound quality evaluation model for signal type conversion to obtain the perceptual domain signal corresponding to the target sound pressure signal; The spectral signal corresponding to the receptive domain signal is input into the CNN network for high-dimensional feature extraction; and the extracted high-dimensional features are input into the fully connected layer to predict the receptive domain score. Calculate the total harmonic distortion (THD) and signal-to-noise ratio (SNR) of the target audio signal and the target sound pressure signal; Based on the signal category to which the target audio signal belongs, the perceptual domain score, THD value, and SNR value are weighted and summed to obtain the noise level score corresponding to the target audio signal, which is used as the detection result obtained after detecting the noise level of the target audio signal.

6. The method according to claim 1, characterized in that, The sound quality evaluation model is constructed as follows: Obtain the training sound pressure signal corresponding to the training audio signal output by the pre-built loudspeaker cavity model; The training sound pressure signal is input into the psychoacoustic model for filtering to obtain the training perceptual domain signal corresponding to the training sound pressure signal. Using the initial filter weights, the training receptive field signal is weighted, and the spectrum signal corresponding to the weighted receptive field signal is input into the CNN network for high-dimensional feature extraction. Then, the extracted training high-dimensional features are input into the fully connected layer to predict the training receptive field score. Calculate the training THD and training SNR values ​​of the training audio signal and the training sound pressure signal; Based on the signal category to which the training audio signal belongs, the training perceptual domain score, the training THD value, and the training SNR value are weighted and summed using floating weights to obtain the noise level score corresponding to the training audio signal. By calculating the difference between the actual noise score corresponding to the training audio signal and the noise level score corresponding to the training audio signal, the values ​​of the initial filter weights, floating weights, and other model parameters are continuously updated to train and generate a sound quality evaluation model.

7. The method according to any one of claims 1-6, characterized in that, When the detection result indicates the presence of noise, timely noise suppression processing is performed on the target audio signal, including: When the detection results indicate the presence of noise, the target audio signal is promptly subjected to a Fourier transform to obtain the signal spectrum in the frequency domain; and the peak value of the signal spectrum is located in the frequency domain to reduce noise by lowering the peak value, thereby achieving noise suppression.

8. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor being used to invoke and execute the computer program to implement the method of any one of claims 1-7.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by an electronic device, implements the method of any one of claims 1-7.

Citation Information

Patent Citations

  • Airflow noise elimination method and device, computer equipment and storage medium

    CN111768801A

  • Noise suppression method and electronic equipment

    CN116546126A