Audio signal processing method, device and storage medium

By constructing a dynamic speaker model and performing frequency division and masking filtering, the problem of high speaker power consumption in mobile terminal audio playback scenarios was solved, achieving a balance between sound quality and power consumption, and improving the user experience.

CN120499553BActive Publication Date: 2026-03-27HONOR DEVICE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-10
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

In mobile terminal audio playback scenarios, how can we reduce speaker power consumption while ensuring sound quality to improve user experience?

Method used

By constructing a dynamic loudspeaker model, frequency division and masking filtering are performed based on the loudspeaker's key mechanical parameters and listening position transfer function to eliminate low-frequency audio signals that are imperceptible to the human ear and reduce the loudspeaker's power consumption.

Benefits of technology

While maintaining sound quality, the power consumption of the speakers has been reduced, thus improving the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120499553B_ABST
    Figure CN120499553B_ABST
Patent Text Reader

Abstract

The application provides an audio signal processing method, device and storage medium. The method constructs a dynamic loudspeaker model through a key mechanical parameter obtained in real time, constructs a listening area sound pressure mapping according to a listening position transfer function of the loudspeaker, a listening area parameter and the dynamic loudspeaker model, inputs the listening area sound pressure mapping and low-frequency digital signals in an original audio digital signal to be played into a psychoacoustic model, performs masking filtering on the low-frequency digital signals based on a masking spectrum obtained, thereby performing filtering processing on audio in the audio digital signals that cannot be perceived by a user, and reduces the power of the loudspeaker under the premise of ensuring sound quality, and improves user experience.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of terminal devices, and in particular to an audio signal processing method, device and storage medium. BACKGROUND

[0002] With the increasing powerfulness of mobile terminals, people gradually increase the use of mobile terminals in audio playback scenarios, such as listening to music, watching movies and listening to novels, etc.

[0003] However, in order to obtain better listening experience, the speaker of the mobile terminal doubles the power consumption to meet the performance requirements of the playback. Therefore, how to reduce the power consumption while ensuring the sound quality in the audio playback scenario is a problem to be solved. SUMMARY

[0004] The embodiments of the present application provide an audio signal processing method, device and storage medium, aiming to reduce the power consumption of the speaker while ensuring the sound quality in the audio playback scenario of the mobile terminal, and improve the user experience.

[0005] In a first aspect, the embodiments of the present application provide an audio signal processing method applied to an electronic device, the electronic device comprising a speaker, the method comprising: constructing a dynamic speaker model based on a key mechanical parameter of the speaker acquired in real time; constructing a listening area sound pressure mapping based on a listening position transfer function of the speaker, a listening area parameter and the dynamic speaker model, the listening area parameter being used to represent a listening range of a user, and the listening position transfer function being used to represent a relationship between a first audio digital signal received at different listening positions and an original audio digital signal; performing frequency division processing on the original audio digital signal, inputting a low-frequency digital signal obtained by the frequency division processing and the listening area sound pressure mapping into a preset psychoacoustic model to obtain a masking spectrum; performing masking filter processing on the low-frequency digital signal based on the masking spectrum; and reconstructing and synthesizing the low-frequency digital signal after the masking filter processing and a high-frequency digital signal obtained by the frequency division processing to obtain an audio signal for playing by the speaker.

[0006] Exemplarily, the key mechanical parameters of the speaker include but are not limited to magnetic force factor Bl, stiffness coefficient Kms, mechanical impedance Rms and voice coil displacement x, etc. The linear and nonlinear changes of these mechanical parameters can represent different working states of the speaker. At the same time, these mechanical parameters will dynamically change with the temperature change of the speaker.

[0007] Exemplarily, the listening position transfer function of the speaker can be pre-stored in the memory of the mobile phone, can be acquired in real time through the cloud, or can be pre-stored in the memory of the mobile phone and updated in real time through the cloud. The calculation method can be to first construct a listening area model as shown inFigure 5 The mathematical model is shown, and then the standard microphone is set at point B in Figure 5 , or the artificial ears are set on the left and right sides at point C, the audio sound emitted by the mobile phone loudspeaker is collected, and the listening area position transfer function is obtained by calculating the audio sound and the original audio sound.

[0008] Exemplarily, the listening area parameter can be the listening range in Figure 5 , which can be the head movement range of a person, and the value is used to represent the distance between the head movement range and the head contour as y.

[0009] Exemplarily, in the sound system, since the wavelength of low-frequency sound is longer, more energy is needed to push the air molecules to vibrate to produce functional sound, so more power is needed to play low-frequency sound to produce a large enough sound pressure level. In contrast, high-frequency sound has a higher frequency and a shorter wavelength, and it is relatively easy to push air molecules to vibrate, so it does not need too much power to achieve a certain sound pressure level. Therefore, in the embodiment, only the low-frequency digital signal is subjected to masking filter processing.

[0010] Exemplarily, the preset psychoacoustic model can be a frequency domain masking model in the following embodiment, and the power consumption reduction principle is referred to Figure 7 . By using the frequency domain masking thread, the masked sound that cannot be perceived by the human ear is removed, thereby achieving the purpose of reducing power consumption.

[0011] Exemplarily, before the high-frequency digital signal is reconstructed and synthesized with the low-frequency digital signal, a all-pass filter is used to delay and compensate the high-frequency digital signal.

[0012] Therefore, the embodiment of the application constructs a dynamic loudspeaker model by the key mechanical parameters obtained in real time, constructs a listening area sound pressure mapping according to the listening position transfer function of the loudspeaker, the listening area parameter and the dynamic loudspeaker model, inputs the listening area sound pressure mapping and the low-frequency digital signal in the original audio digital signal to be played into the psychoacoustic model, performs masking filter processing on the low-frequency digital signal based on the obtained masking spectrum, thereby filtering the audio that cannot be perceived by the user in the audio digital signal, reducing the power of the loudspeaker under the premise of ensuring the sound quality, and improving the user experience.

[0013] According to the first aspect, the dynamic loudspeaker model is constructed based on the real-time obtained key mechanical parameters of the loudspeaker, and the method comprises: constructing an impedance model of the loudspeaker based on real-time obtained current signals and voltage signals of the loudspeaker, the impedance model being used to represent the relationship between the current signals, the voltage signals and the key mechanical parameters; calculating the key mechanical parameters of the loudspeaker based on the impedance model, the key mechanical parameters comprising a magnetic force factor, a stiffness coefficient and a mechanical impedance; and constructing a dynamic loudspeaker model based on the key mechanical parameters.

[0014] It can be understood that, by constructing an impedance model of the loudspeaker based on real-time obtained current signals and voltage signals of the loudspeaker, and calculating the key mechanical parameters of the loudspeaker based on the impedance model, a loudspeaker model in a real working scenario can be obtained, so that the acoustic state of the loudspeaker can be controlled more accurately.

[0015] According to the first aspect, or any one of the implementation manners of the first aspect, the dynamic loudspeaker model is constructed based on the real-time obtained key mechanical parameters of the loudspeaker, and the method further comprises: selecting corresponding key mechanical parameters from a preset mechanical parameter model according to a real-time obtained temperature value of the loudspeaker, and constructing a dynamic loudspeaker model.

[0016] It can be understood that, since the key mechanical parameters of the loudspeaker change with temperature, the key mechanical parameter models of the loudspeaker at different temperatures can be preset, so that when the dynamic loudspeaker model is constructed, the key mechanical parameters at the corresponding temperature can be selected, so that the loudspeaker model in a real working scenario can be obtained while reducing the performance requirements of the mobile phone, and the acoustic state of the loudspeaker can be controlled more accurately.

[0017] According to the first aspect, or any one of the implementation manners of the first aspect, the sound field pressure mapping is constructed based on the listening position transfer function of the loudspeaker, the sound field parameter and the dynamic loudspeaker model, and the method comprises: obtaining an edge sound pressure envelope of the sound field based on the obtained sound field parameter and the dynamic loudspeaker model; correcting the obtained listening position transfer function based on the edge sound pressure envelope to obtain a sound pressure model of a plurality of positions in the sound field; and calculating a lower envelope line of the sound pressure model of the plurality of positions to obtain the sound field pressure mapping of the sound field.

[0018] It can be understood that, by calculating the lower envelope line of the edge sound pressure envelope of the sound field to determine the edge sound pressure envelope of the sound field, the best listening effect at different positions in the sound field can be ensured, and the listening experience of the user can be improved.

[0019] According to a first aspect, or any of the implementations of the first aspect, after the sound pressure model of the plurality of positions is obtained based on the edge sound pressure envelope of the listening area, the method further includes filtering the sound pressure model of the plurality of positions based on a preset absolute hearing threshold.

[0020] The absolute hearing threshold is an example of the minimum sound pressure level that can be perceived by the human ear. It is a basic characteristic of the auditory system, and represents the minimum sound pressure that can be heard by the human ear at a specific frequency. Therefore, the embodiments of the present application filter the sound pressure model based on the preset absolute hearing threshold, thereby directly filtering the sound pressure that cannot be perceived by the human ear in the sound pressure model, thereby reducing the subsequent calculation amount and saving the computing resources of the mobile terminal.

[0021] According to the first aspect, or any of the implementations of the first aspect, the method further includes adjusting the size of the lower envelope line based on a preset adjustment factor to obtain the listening area sound pressure mapping.

[0022] The larger the edge sound pressure envelope, the more the power consumption is reduced, and the greater the possibility of sound quality loss. The smaller the edge sound pressure envelope, the less the power consumption is reduced, and the smaller the possibility of sound quality loss. The embodiments of the present application adjust the size of the edge sound pressure envelope by using the adjustment factor, thereby balancing the power consumption and the sound quality, and optimizing the listening experience of the user.

[0023] According to the first aspect, or any of the implementations of the first aspect, the method further includes filtering the original audio digital signal using a frequency division filter to obtain a high-frequency digital signal and a low-frequency digital signal; performing frame division processing on the low-frequency digital signal, and performing Fourier transform on each frame of the low-frequency digital signal to obtain a low-frequency frequency domain signal of each frame; and inputting each frame of the low-frequency frequency domain signal and the listening area sound pressure mapping into the preset psychoacoustic model to obtain the masking spectrum.

[0024] According to the first aspect, or any of the implementations of the first aspect, the method further includes designing a filter based on the masking spectrum, and performing masking filter processing on each frame of the low-frequency frequency domain signal based on the filter.

[0025] According to a first aspect, or any possible implementation mode of the first aspect, the method further comprises: performing inverse Fourier transform on each frame of the low-frequency frequency domain signal after the masking filtering processing to obtain a low-frequency digital signal after the masking filtering processing.

[0026] According to a first aspect, or any possible implementation mode of the first aspect, the method further comprises: performing inverse Fourier transform on each frame of the low-frequency frequency domain signal after the masking filtering processing to obtain a low-frequency digital signal after the masking filtering processing.

[0027] In a second aspect, an electronic device is provided. The electronic device comprises: a memory and a processor coupled with each other; the memory stores program instructions, and the program instructions are executed by the processor to enable the electronic device to perform the method in the first aspect or any possible implementation mode of the first aspect.

[0028] In a third aspect, a computer readable medium is provided for storing a computer program, and the computer program comprises instructions for performing the method in the first aspect or any possible implementation mode of the first aspect.

[0029] In a fourth aspect, a computer program is provided, and the computer program comprises instructions for performing the method in the first aspect or any possible implementation mode of the first aspect.

[0030] In a fifth aspect, a chip is provided, and the chip comprises a processing circuit and a transceiver pin. The transceiver pin and the processing circuit communicate with each other through an internal connection path, and the processing circuit executes the method in the first aspect or any possible implementation mode of the first aspect to control the receiving pin to receive a signal and to control the sending pin to send a signal. BRIEF DESCRIPTION OF DRAWINGS

[0031] Figure 1 A schematic diagram of a hardware structure of an electronic device is exemplarily shown;

[0032] Figure 2 A schematic diagram of a software structure of an electronic device is exemplarily shown;

[0033] Figure 3 A schematic diagram of a user interface is exemplarily shown;

[0034] Figure 4 A schematic diagram of relative positions between a person and a mobile phone in a music outdoor playing scene is exemplarily shown;

[0035] Figure 5 A schematic diagram of a position mathematical model in a music outdoor playing scene is exemplarily shown;

[0036] Figure 6 A schematic diagram of the time domain masking phenomenon is shown for example;

[0037] Figure 7 A schematic diagram of the frequency domain masking phenomenon is shown for example;

[0038] Figure 8 A first sound pressure contrast schematic diagram is shown for example;

[0039] Figure 9 A signal processing flow is shown for example;

[0040] Figure 10 A mechanical parameter model schematic diagram of the loudspeaker at different temperatures is shown for example;

[0041] Figure 11 A loudspeaker model real-time update schematic diagram is shown for example;

[0042] Figure 12 A loudspeaker audio signal acquisition flow schematic diagram is shown for example;

[0043] Figure 13 A listening area sound pressure mapping flow schematic diagram is shown for example;

[0044] Figure 14 A listening area edge envelope calculation schematic diagram is shown for example;

[0045] Figure 15 An absolute hearing threshold schematic diagram is shown for example;

[0046] Figure 16 A second sound pressure contrast schematic diagram is shown for example. DETAILED DESCRIPTION

[0047] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some of the embodiments of the present application, but not all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application.

[0048] The term "and / or" in the present document is only to describe the association relationship of the associated objects, which means that there can be three kinds of relationships, for example, A and / or B can represent the following three cases: A exists alone, A and B exist together, and B exists alone.

[0049] The terms "first" and "second" and the like in the description and claims of the present application are used for distinguishing between similar objects talking different reference numbers and do not imply a specific order or sequence. For example, the first target object and the second target object are used for distinguishing between different target objects, and do not imply a specific order or sequence.

[0050] In the present application, the words "exemplary" and "for example" are used to help clarify the description of the present application. Any embodiment or design scheme described as "exemplary" or "for example" in the present application should not be interpreted as being more preferred or advantageous than other embodiments or design schemes. Rather, the use of "exemplary" or "for example" is intended to present relevant concepts in a concrete manner.

[0051] In the description of the present application, unless otherwise specified, the meaning of "a plurality of" is two or more. For example, a plurality of processing units refers to two or more processing units; a plurality of systems refers to two or more systems.

[0052] In order to better understand the technical solutions provided by the embodiments of the present application, before the technical solutions of the embodiments of the present application are described, first, the hardware structure of the terminal device (such as mobile phone, tablet computer, touchable PC, etc.) applicable to the embodiments of the present application is described in conjunction with the drawings. In order to facilitate the description, Figure 1 Take a mobile phone as an example for description.

[0053] Referring to Figure 1 , the mobile phone 100 can include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charge management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a sensor module 180, a key 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc.

[0054] The processor 110 can include one or more processing units, for example: the processor 110 can include an application processor (AP), a modem, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), and the like, which are not listed one by one here, and the present application does not make any limitation thereto.

[0055] Regarding the controller as the processing unit mentioned above, it can be the nerve center and command center of the mobile phone 100. In actual application, the controller can generate operation control signals according to instruction operation codes and timing signals to complete the control of instruction fetching and instruction execution.

[0056] Regarding the modem as mentioned above, it can include a modulator and a demodulator. The modulator is used to modulate the low-frequency baseband signal to be sent into a medium-high frequency signal. The demodulator is used to demodulate the received electromagnetic wave signal into a low-frequency baseband signal, and transmit the demodulated low-frequency baseband signal to the baseband processor for processing.

[0057] Regarding the baseband processor mentioned above, it is used to process the low-frequency baseband signal transmitted by the modulator, and transmit the processed low-frequency baseband signal to the application processor.

[0058] It should be noted that in some implementations, the baseband processor can be integrated into the modem, i.e., the modem can have the function of the baseband processor.

[0059] Regarding the application processor mentioned above, it is used to output sound signals through audio devices (not limited to the loudspeaker 170A, the receiver 170B, etc.), or display images or videos through the display screen 194.

[0060] Regarding the digital signal processor mentioned above, it is used to process digital signals. Specifically, the digital signal processor can process not only digital image signals, but also other digital signals. For example, when the mobile phone 100 selects a frequency point, the digital signal processor can be used to perform Fourier transform on the frequency point energy, etc. The digital signal processor is also used to process the audio digital signal transmitted by the codec, and output the signal-processed audio digital signal to the loudspeaker 170A through the audio module 170.

[0061] As to the neural network processor mentioned above, by referring to the structure of biological neural network, for example, by referring to the transmission mode between human brain neurons, the input information can be quickly processed, and the self-learning can be continuously performed. Through the neural network processor, the intelligent cognition of the mobile phone 100 can be realized, for example, image recognition, face recognition, voice recognition, text understanding, action generation, etc.

[0062] As to the video codec mentioned above, it is used for compressing or decompressing digital video. For example, the mobile phone 100 can support one or more video codecs. In this way, the mobile phone 100 can play or record videos in multiple encoding formats, for example, moving picture experts group (MPEG) 1, MPEG 2, MPEG 3, MPEG 4, etc.

[0063] As to the ISP mentioned above, it is used for outputting digital image signals to the DSP for processing. Specifically, the ISP is used for processing the data fed back by the camera 193. For example, when taking a photo or recording a video, the shutter is opened, the light is transmitted to the camera photosensitive element through the lens, the light signal is converted into an electric signal, and the camera photosensitive element transmits the electric signal to the ISP for processing, and converts it into an image visible to the naked eye. The ISP can also optimize the algorithm of the noise, brightness and skin color of the image. The ISP can also optimize the exposure, color temperature and other parameters of the shooting scene. In some implementations, the ISP can be arranged in the camera 193.

[0064] As to the DSP mentioned above, it is used for converting the digital image signal into a standard image signal in RGB, YUV, etc.

[0065] In addition, it should be further explained that, as to the processor 110 including the above-mentioned processing units, in some implementations, different processing units can be independent devices. That is, each processing unit can be regarded as a processor. In other implementations, different processing units can be integrated in one or more processors. For example, in some implementations, the modem processor can be an independent device. In other implementations, the modem processor can be independent of the processor 110, and arranged in the same device as the mobile communication module 150 or other functional modules.

[0066] It should be understood that the above description is only an example for better understanding the technical scheme of the embodiment, and is not the only limitation of the embodiment.

[0067] In addition, the processor 110 can further include one or more interfaces. Among them, the interface can include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, and the like, which will not be listed one by one here, and the present application does not make any limitation thereto.

[0068] In addition, the processor 110 can also be provided with a memory for storing instructions and data. In some implementations, the memory in the processor 110 is a cache memory. The memory can save instructions or data that the processor 110 has just used or repeatedly uses. If the processor 110 needs to use the instructions or data again, it can be directly called from the memory. Avoiding repeated access, reducing the waiting time of the processor 110, thus improving the efficiency of the system.

[0069] Continuing to refer to Figure 1 The external memory interface 120 can be used to connect an external storage card, such as a Micro SD card, to expand the storage capacity of the mobile phone 100. The external storage card communicates with the processor 110 through the external memory interface 120 to realize the data storage function. For example, save music, video and other files in the external storage card.

[0070] Continuing to refer to Figure 1The internal memory 121 can be used to store computer executable program codes including instructions. The processor 110 performs various functional applications and data processing of the mobile phone 100 by running the instructions stored in the internal memory 121. The internal memory 121 can include a program storage area and a data storage area. The program storage area can store an operating system, at least one application program required by a function (such as a sound playing function, an image playing function, etc.), and the like. The data storage area can store data created during the use of the mobile phone 100, and the like. In addition, the internal memory 121 can include a high-speed random access memory, and can further include a non-volatile memory such as at least one magnetic disk storage device, a flash memory device, a universal flash storage (UFS), and the like.

[0071] With reference back to Figure 1 The charging management module 140 is configured to receive charging input from a charger. The charger can be a wireless charger or a wired charger. In some implementations of wired charging, the charging management module 140 can receive charging input from a wired charger through the USB interface 130. In some implementations of wireless charging, the charging management module 140 can receive wireless charging input through a wireless charging coil of the mobile phone 100. The charging management module 140 can charge the battery 142 and supply power to the terminal device through the power management module 141.

[0072] With reference back to Figure 1 The power management module 141 is configured to connect the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140, and supplies power to the processor 110, the internal memory 121, the external memory, the display screen 194, the camera 193, and the wireless communication module 160, and the like. The power management module 141 can also be configured to monitor parameters such as battery capacity, battery cycle number, battery health status (leakage, impedance), and the like. In some other implementations, the power management module 141 can also be arranged in the processor 110. In some other implementations, the power management module 141 and the charging management module 140 can also be arranged in the same device.

[0073] With reference back to Figure 1 The wireless communication function of the mobile phone 100 can be implemented through the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, a modem processor, a baseband processor, and the like.

[0074] It should be noted that the antenna 1 and the antenna 2 are used for transmitting and receiving electromagnetic wave signals. Each antenna in the mobile phone 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be multiplexed to improve the utilization of the antennas. For example, the antenna 1 can be multiplexed as a diversity antenna of a wireless local area network. In some other implementations, the antennas can be used in combination with a tuning switch.

[0075] With continued reference to Figure 1 , the mobile communication module 150 can provide a solution for wireless communication including 2G / 3G / 4G / 5G, etc. applied in the mobile phone 100. The mobile communication module 150 can include at least one filter, a switch, a power amplifier, a low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves by the antenna 1, and perform filtering, amplification, etc. on the received electromagnetic waves, and transmit the processed signals to the modem processor for demodulation. The mobile communication module 150 can also amplify the signals modulated by the modem processor, and convert the signals into electromagnetic waves radiated by the antenna 1. In some implementations, at least part of the functional modules of the mobile communication module 150 can be arranged in the processor 110. In some implementations, at least part of the functional modules of the mobile communication module 150 can be arranged in the same device as at least part of the modules of the processor 110.

[0076] With continued reference to Figure 1 , the wireless communication module 160 can provide a solution for wireless communication including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) network), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR) technology, etc. applied in the mobile phone 100. The wireless communication module 160 can be one or more devices integrated with at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via the antenna 2, performs frequency modulation and filtering processing on the electromagnetic wave signals, and transmits the processed signals to the processor 110. The wireless communication module 160 can also receive signals to be transmitted from the processor 110, perform frequency modulation and amplification on the signals, and convert the signals into electromagnetic waves radiated by the antenna 2.

[0077] With continued reference to Figure 1The audio module 170 can include a speaker 170A, a receiver 170B, a microphone 170C, a headphone interface 170D, etc. For example, the mobile phone 100 can implement audio functions through the speaker 170A, the receiver 170B, the microphone 170C, the headphone interface 170D, etc. in the application processor and the audio module 170. For example, music playing, recording, and video recording functions.

[0078] In the process of implementing the audio functions through the application processor and the audio module 170, the audio module 170 can be used to convert digital audio information into an analog audio signal output, and also used to convert an analog audio input into a digital audio signal. The audio module 170 can also be used to encode and decode audio signals. In some implementations, the audio module 170 can be disposed in the processor 110, or some functional modules of the audio module 170 can be disposed in the processor 110.

[0079] Continuing to refer to Figure 1 The speaker 170A, also called a "loudspeaker", is used to convert an audio electrical signal into a sound signal. The mobile phone 100 can play music through the speaker 170A. The audio signal processing method provided in the embodiments of the present application is to process the audio electrical signal before it is output to the speaker 170A, so as to reduce the power consumption of the speaker 170A.

[0080] Continuing to refer to Figure 1 The sensor module 180 can include a pressure sensor, a gyroscope sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a distance sensor, a proximity light sensor, a fingerprint sensor, a temperature sensor, a touch sensor, an ambient light sensor, a bone conduction sensor, etc. which are not listed one by one here, and the present application does not limit this.

[0081] Continuing to refer to Figure 1 The key 190 includes a power-on key, a volume key, etc. The key 190 can be a mechanical key. It can also be a touch key. The mobile phone 100 can receive a key input and generate a key signal input related to user settings and function control of the mobile phone 100.

[0082] Continuing to refer to Figure 1 The motor 191 can generate a vibration prompt. The motor 191 can be used for incoming call vibration prompt, and also can be used for touch vibration feedback.

[0083] Continuing to refer to Figure 1 The indicator 192 can be an indicator light, which can be used to indicate a charging state, a power change, and also can be used to indicate a message, a missed call, a notification, etc.

[0084] Continuing to refer to Figure 1The camera 193 is used to capture still images or videos. The mobile phone 100 can achieve its shooting function through an ISP, camera 193, video codec, GPU, display 194, and application processor. Specifically, an object generates an optical image through a lens and projects it onto a photosensitive element. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then passed to the ISP for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into image signals in standard RGB, YUV, or other formats. In some implementations, the mobile phone 100 may include one or N cameras 193, where N is a positive integer greater than 1.

[0085] See also Figure 1 The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. In some implementations, the mobile phone 100 may include one or N displays 194, where N is a positive integer greater than 1. The mobile phone 100 can implement display functions through a GPU, the display screen 194, and an application processor. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. The processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.

[0086] That concludes the introduction to the hardware structure of the Mobile 100. It should be understood that... Figure 1 The mobile phone 100 shown is just an example. In a specific implementation, the mobile phone 100 may have more or fewer components than shown in the figure, may combine two or more components, or may have different component configurations. Figure 1 The various components shown can be implemented in hardware, software, or a combination of hardware and software, including one or more signal processing and / or application-specific integrated circuits.

[0087] To better understand Figure 1 The software structure of the mobile phone 100 shown is described below. Before describing the software structure of the mobile phone 100, the possible architectures for the software system of the mobile phone 100 will be explained first.

[0088] Specifically, in practical applications, the software system of Mobile 100 can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture.

[0089] In addition, it is understandable that the currently mainstream terminal device uses a software system including but not limited to a Windows system, an Android system and an iOS system. For ease of illustration, an Android system in a layered architecture is taken as an example to exemplarily illustrate the software structure of the mobile phone 100.

[0090] In addition, the audio signal processing scheme provided in the embodiments of the present application is also applicable to other systems in specific implementation.

[0091] Referring to Figure 2 , a software structure block diagram of the mobile phone 100 in the embodiments of the present application is shown.

[0092] As Figure 2 indicated, the layered architecture of the mobile phone 100 divides the software into several layers, each layer has a clear role and division of labor. The layers communicate with each other through a software interface. In some implementations, the Android system is divided into four layers, from top to bottom, an application layer, an application framework layer, an Android runtime and a system library, and a kernel layer.

[0093] The application layer can include a series of application packages. As Figure 2 indicated, the application packages can include application programs such as an application market, music, shopping, permission management, Bluetooth, Wi-Fi, settings, and the like, which are not listed one by one here, and the present application does not limit this.

[0094] The application framework layer provides application programming interfaces (APIs) and programming frameworks for the application programs of the application layer. In some implementations, these programming interfaces and programming frameworks can be described as functions. As Figure 2 indicated, the application framework layer can include functions such as an audio signal processing module, a view system, a content provider, a speech recognition module, a wake-up word / command word detection module, a scene recognition module, a speech data fusion module, a voice activity detection module, a voiceprint verification module, a wake-up word voiceprint template management module, and the like, which are not listed one by one here, and the present application does not limit this.

[0095] Exemplarily, in the embodiment of the present application, when a user uses a mobile phone to play music, the audio signal processing module processes the music signal to be played according to an audio effect algorithm and a reliability protection algorithm, and divides the music signal into a high-frequency part and a low-frequency part through a cross-over filter. The low-frequency part is subjected to a short-time Fourier transform, and the obtained low-frequency frequency domain signal and a sound pressure level (SPL) of a listening area are input into a psychoacoustic model (PAM) to output a masking spectrum of the low-frequency frequency domain signal. A frequency domain filter is designed according to the masking spectrum, the low-frequency frequency domain signal is filtered by using the frequency domain filter, and after a causality constraint, the low-frequency frequency domain signal is subjected to an inverse short-time Fourier transform. The obtained low-frequency time domain signal is synthesized with the high-frequency part and then output, so as to realize the reduction of power consumption of a loudspeaker and the improvement of user experience in the case of ensuring the audio effect quality of the mobile phone in an audio playing scenario.

[0096] It should be understood that the above description is only an example for better understanding the technical solution of the embodiment and is not the only limitation of the embodiment.

[0097] In addition, it can be understood that the division of each functional module is only an example for better understanding the technical solution of the embodiment and is not the only limitation of the embodiment. In actual application, the above functions can also be integrated in one functional module, and the embodiment does not limit this.

[0098] In addition, in actual application, the above functional modules can also be represented as services and frameworks, such as the voice recognition module can be represented as a voice recognition service, a voice recognition framework, etc., and the embodiment does not limit this.

[0099] In addition, it should be further pointed out that the window manager in the application framework layer is used to manage window programs. The window manager can obtain the size of a display screen, judge whether there is a status bar, lock a screen, intercept a screen, etc.

[0100] In addition, it should be further pointed out that the content provider in the application framework layer is used to store and obtain data and make the data accessible by an application program. The data can include videos, images, audios, dialed and received calls, browsing history and bookmarks, a phone book, etc., which are not enumerated one by one here, and the present application does not limit this.

[0101] In addition, it should be further pointed out that the view system in the application framework layer includes visual controls, such as a control for displaying text and a control for displaying pictures, etc. The view system can be used to build an application program. A display interface can be composed of one or more views. For example, a display interface including a short message notification icon can include a view for displaying text and a view for displaying pictures.

[0102] In addition, it should be noted that the phone manager in the application framework layer is used to provide the communication function of the mobile phone 100. For example, the management of the call state (including call connection, call hang-up, etc.).

[0103] The resource manager provides various resources for the application program, such as localized strings, icons, pictures, layout files, video files, etc., which are not listed one by one here, and the present application does not limit this.

[0104] In addition, it should be noted that the notification manager in the application framework layer enables the application program to display notification information in the status bar, which can be used to convey a message of the notification type and can automatically disappear after a short stay without user interaction.

[0105] The Android Runtime includes the core library and the virtual machine. The Android Runtime is responsible for the scheduling and management of the Android system.

[0106] The core library includes two parts: one part is the function function called by the java language, and the other part is the core library of Android.

[0107] The application layer and the application framework layer run in the virtual machine. The virtual machine executes the java file of the application layer and the application framework layer into a binary file. The virtual machine is used to perform the management of the object life cycle, the management of the stack, the management of the thread, the management of the security and the exception, and the garbage collection, etc.

[0108] The system library can include a plurality of functional modules. For example: the surface manager, the media library, the three-dimensional (3D) graphics processing library (for example: OpenGL ES), the two-dimensional (2D) graphics engine (for example: SGL), etc.

[0109] The surface manager is used to manage the display subsystem and provides the fusion of 2D and 3D layers for a plurality of application programs.

[0110] The media library supports a plurality of commonly used audio, video format playing and recording, and static image files, etc. The media library can support a plurality of audio and video coding formats, such as MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc.

[0111] The three-dimensional graphics processing library is used to realize three-dimensional graphics drawing, image rendering, synthesis, and layer processing, etc.

[0112] Understandably, the 2D graphics engine mentioned above is a drawing engine for 2D drawing.

[0113] Further, it is appreciated that the kernel layer in the Android system is a layer between hardware and software. The kernel layer contains at least display driver, camera driver, audio driver, sensor driver, etc. For example, the sensor driver can be used to output the detection signal of a sensor (e.g. touch sensor) to the view system, so that the view system displays the corresponding application interface in response to the detection signal.

[0114] As to the software structure of the mobile phone 100, it is appreciated that, Figure 2 The layers in the illustrated software structure and the components contained in each layer do not constitute a specific limitation to the mobile phone 100. In some other embodiments of the present application, the mobile phone 100 can include more or less layers than illustrated, and each layer can include more or less components, which are not limited in the present application.

[0115] In daily life, users often use mobile terminals to play videos, play audios and play music. In order to get a better listening experience, users will also choose a comfortable sitting or lying posture, hold the mobile terminal and turn on the music external playing function.

[0116] For example, after holding the mobile terminal and turning on the music external playing function, the user will also sway his body along with the rhythm of the music, or even sway along with the rhythm.

[0117] In order to better understand the scenario of the user playing music externally, the following takes the mobile terminal as a mobile phone as an example, and the accompanying drawings are combined to Figure 3 to Figure 5 for illustration.

[0118] In the music external playing scenario, referring to Figure 3 , Figure 3 An example of a mobile phone interface 10a is shown. For example, the interface 10a displays icons of a plurality of application programs, such as icons of music 10a-1, camera, address book, phone, message, clock, calendar, gallery, memo, file management, email, music, calculator, video, recorder, weather, browser, settings, etc.

[0119] It should be noted that in some possible implementations, Figure 3 The interface 10a shown in (1) can be referred to as a main interface. When the user clicks the icon 10a-1 in the interface 10a, the music playing function of the music application can be used.

[0120] Continuing to refer to Figure 3For example, when a user clicks the music application icon 10a-1, the phone responds to the user's action, recognizes the icon corresponding to the user's click as the music application icon, and then calls the corresponding interface in the application framework layer to launch the music application. At this time, the phone displays the music application's playback interface, for example... Figure 3 Interface 10b is shown in (2).

[0121] See Figure 3 (2) Figure 3 (2) Exemplarily illustrates a music playback interface 10b of a music application. This interface 10b displays a music information area, a music playback operation area, and a floating adjustment area. The music information area displays music categories (“My Favorite Music”) and album art; the music playback operation area displays music names (“Demo Music”), artists (“Various Artists”), a favorite button, a comment button, a music playback progress bar, playback mode, previous track, play / pause control 10b-1, next track, and playlist; the floating adjustment area displays volume adjustment controls 10b-2.

[0122] For example, the user clicks control 10b-1 to start playing the current music and uses control 10b-2 to adjust the phone's volume to a suitable level. A diagram illustrating the relative position of the phone and the user at this time can be found in [reference needed]. Figure 4 .

[0123] See Figure 4 , Figure 4 This example illustrates the relative positions of a mobile device and a person's head in a music playback scenario. Figure 4 In this context, there is a certain distance between the mobile phone and the user. For example, this distance could be the distance x from the center point of the mobile phone to the tip of the user's nose. Based on Figure 4 The positional relationship can also be used to mathematically model the positions of the mobile terminal and the human body in the above music playback scenario, thus obtaining a mathematical model of the mobile terminal and the human body, such as... Figure 5 As shown.

[0124] See Figure 5 , Figure 5 An exemplary mathematical model of a mobile terminal and the human body in a music playback scenario is provided. In this model, the user holds a mobile phone, which is positioned at a certain azimuth angle θ and elevation angle φ relative to the human body. The distance between the phone and the tip of the nose (A) is denoted as x. The outline of the human head is abstracted as a sphere, with point B as the center point. The ears are located at point C, above and behind point B. Furthermore, since the user's head may move with the rhythm while listening to music, its range of motion is [missing information]. Figure 5 The range of head movement is shown, and the distance between the range of head movement and the head outline is y.

[0125] Continuing to refer to Figure 5 Since there is a certain distance x between the human head and the mobile phone, when the user is listening to music, the user will hear the sound through the air Figure 3 The control 10b-2 shown in FIG. 2 (2) increases the volume of the mobile phone until the user feels that the volume is appropriate. However, when the user uses a large volume to play out, the power consumption of the speaker of the mobile phone increases, which causes the speaker position of the mobile phone to heat up. The heating phenomenon seriously reduces the user experience, and even makes the user feel that the mobile phone has a safety hazard.

[0126] In an implementation manner, the sound that cannot be heard by the human ear in the audio signal corresponding to the music can be filtered by using a psychoacoustic model, so as to achieve the purpose of reducing power consumption.

[0127] Specifically, when the mobile phone plays audio, the codec in the processor decodes the audio file to obtain an audio digital signal. The digital signal processor (DSP) in the processor filters the audio digital signal. The audio digital signal after the filtering is sequentially subjected to digital-to-analog conversion and power amplification, and then the power amplified audio is played by the speaker. However, in the scenario in which the speaker plays audio, the human ear has weak hearing of some audio analog signals in the played audio. These audio analog signals with weak hearing are actually covered by audio signals with strong hearing. In psychoacoustics, this covering is also called a masking phenomenon. The masking phenomenon includes a frequency domain masking (SM) phenomenon and a time domain masking (TM) phenomenon.

[0128] In the time domain masking scenario, for example, refer to Figure 6 , Figure 6 An exemplary schematic diagram of a time domain masking phenomenon is shown. In Figure 6 , a plurality of sounds are displayed on a time axis, which are audible sound A1, masked sound I1, masking sound M1, masked sound I2, and audible sound A2. However, since the sound pressure level of the masking sound M1 is high, it forms a masking area on the time axis. When there is no sound with a sound pressure level higher than the masking sound M1 in the masking area, it will affect all the sounds in the above-mentioned masking area, so that the masked sound I1 and the masked sound I2 cannot be perceived by the human ear, thereby the time domain masking phenomenon occurs.

[0129] According to the time sequence difference of the masking sound and the masked sound in the time domain, the time domain masking includes forward masking (FM), simultaneous masking (which can be understood as frequency domain masking), and backward masking (BM). The forward masking refers to the phenomenon that the masking sound appears after the time sequence of the masked sound, and the duration is about 5-20 ms; the backward masking refers to the phenomenon that the masking sound appears before the time sequence of the masked sound, and the duration is about 50-200 ms; and the simultaneous masking refers to the phenomenon that the masking sound and the masked sound appear at the same time in the time sequence.

[0130] Therefore, the masked sound I1 and the masked sound I2 that cannot be perceived by the human ear in the time domain masking phenomenon can be removed from the original audio signal according to the characteristic that they cannot be perceived by the human ear, so that the loudspeaker does not play the masked sound I1 and the masked sound I2 when playing the audio, and the purpose of reducing power consumption is finally achieved.

[0131] Regarding the time domain masking model calculation of the audio signal, an excitation model E[k, n] can be constructed first, and the time domain masking model is obtained based on the excitation model, as follows:

[0132] E f [k,n]=a·E f [k,n-1]+(1-a)·E2[k,n];

[0133] E[k,n]=max(E f [k,n],E2[k,n]);

[0134] Wherein, E2[k, n] is an inter-band spreading function, k is a frequency point, n is a frame number, and a is calculated from a time constant τ.

[0135] In the simulation of forward masking, the energy of each frequency group is distributed over time using a first-order low-pass filter, and the time constant τ depends on the center frequency of each frequency group and is calculated according to the following formula:

[0136]

[0137] Therefore, a can be represented as:

[0138]

[0139] The time domain masking model M[k, n] is obtained according to the above excitation model E[k, n], and is represented as:

[0140]

[0141] Wherein, res is the resolution of the pitch model scaling factor, res is a constant.

[0142] Based on the above calculation method, the time domain masking model can be obtained, and then the masked sound in the audio digital signal is removed based on the time domain masking model in the digital signal processor, so as to reduce the power consumption of the loudspeaker.

[0143] In the frequency domain masking scene, for example, see Figure 7 , Figure 7 An exemplary schematic diagram of a frequency domain masking phenomenon is shown. In Figure 7 , the signals of adjacent frequencies have a masking sound M2 and a masked sound I3, and the masking sound M2 forms a masking threshold MT on the hearing threshold T. The hearing threshold T refers to the minimum sound pressure level that can be perceived by the human ear in a quiet environment without any other sound interference. The masking sound M2 forms a masking threshold MT on the hearing threshold T, and any sound (masked sound I3) below this threshold is masked by the masking sound M2, so that it cannot be perceived, resulting in a frequency domain masking phenomenon.

[0144] In the frequency domain masking phenomenon, the masked sound I3 that cannot be perceived by the human ear appears, and according to its characteristic of being unable to be perceived by the human ear, it can be removed from the original audio signal through a psychoacoustic model (such as the frequency domain masking model in the following embodiment), so that the loudspeaker does not play the masked sound I3 when playing audio, and finally realizes the purpose of reducing power consumption.

[0145] Specifically, the frequency domain masking model is based on the original signal spectrum, uses a level expansion function, and erases the pitch pattern P p [k,n] in frequency to obtain the upper and lower slopes of each frequency point. The expansion function is a bilateral exponential function, and the lower slope is fixed at 27dB / Bark, while the upper slope depends on the frequency and energy. The upper slope S u and the lower slope S l are calculated as follows:

[0146]

[0147] S l [k,L[k,n]]=27;

[0148] wherein L[k,n]=10·log 10 (P p [k,n]), P p [k,n] is a pitch model, k is a frequency point, and n is a frame number.

[0149] Based on the above calculation method, the frequency domain masking model can be obtained, so that the masked sound in the audio digital signal is removed based on the frequency domain masking model in the digital signal processor, thereby reducing the power consumption of the loudspeaker.

[0150] For example, the above frequency domain masking model is applied to the above mobile phone music external playing scene, so that Figure 5 In the middle B point, the user's listening scene is set to collect audio analog, and the collected audio is subjected to Fourier transform. The comparison between the original audio signal S and the audio signal S1 processed by the above frequency domain masking model is shown in the comparison diagram Figure 8 In the above comparison diagram, it can be seen that the sound pressure level of the processed audio signal S1 is obviously reduced, so that after the above frequency domain masking model is processed, the power consumption of the loudspeaker is reduced. Figure 8

[0151] However, when the above psychoacoustic model is used to reduce the power consumption of the audio signal, all the masked sounds are removed, and these masked sounds are not completely inaudible to the human ear (weakly perceived after masking), which leads to poor sound quality of the music and affects the listening experience of the human ear. At the same time, in the above music external playing scene, the user's head may swing within a certain range (for example, the head movement range in the above comparison diagram), that is, the user has a listening range. When the above psychoacoustic model is used to reduce the power consumption of the audio signal, the listening range and the dynamic change of the acoustic state of the loudspeaker in the music external playing scene are not considered, so that the above psychoacoustic model damages the audio digital signal more, greatly reducing the listening experience of the user. Figure 5

[0152] Therefore, the embodiments of the present application provide an audio signal processing method. After the music signal to be played is subjected to a conventional sound effect algorithm and a reliability protection algorithm, the music signal is divided into a high frequency part and a low frequency part by a frequency division filter. The low frequency part is transformed into a frequency domain after short-time Fourier transform, and the frequency domain signal and the user's listening area SPL are input into a psychoacoustic model (PAM) to calculate the masking spectrum of the current signal. Then, a frequency domain filter system is designed according to the masking spectrum, and after causal constraint, a short-time inverse Fourier transform is performed to output a processed low frequency time domain signal. The low frequency time domain signal and the high frequency part are reconstructed and synthesized, so that the power of the loudspeaker is reduced without reducing the sound effect of the audio signal, and the listening experience of the user is improved.

[0153] ​​It can be understood that in the sound system, due to the longer wavelength of the low frequency sound, more energy is needed to push the air molecules to produce vibration, thereby generating functional sound, so the playback of the low frequency sound needs more power to generate a large enough sound pressure level. In contrast, the high frequency sound has a higher frequency and a shorter wavelength, and it is relatively easy to push the air molecules to produce vibration, so it does not need too much power to reach a certain sound pressure level. Therefore, in the embodiment of the application, the high frequency part is not processed additionally, but only the time delay of the high frequency part is compensated by a full-pass filter, and then the processed low frequency part is reconstructed and synthesized to obtain the final music signal for playback.

[0154] In order to better understand the processing process of the music signal by the audio signal processing method provided by the embodiment of the application, the following embodiment still takes the mobile terminal as the mobile phone 100 as an example to explain the music signal processing flow in the external playing scene.

[0155] Taking the external playing scene of the user using the music application as an example, combined with the model of the relationship between the user and the mobile phone position in the external playing scene in Figure 5 , see Figure 9 , Figure 9 for the signal processing flow of the embodiment of the application, including:

[0156] Step S101, after the input first audio signal sequentially passes through the sound effect algorithm and the reliability algorithm, the first audio signal is filtered by a frequency division filter to obtain a high frequency digital signal and a low frequency digital signal.

[0157] It can be understood that the first audio signal can be a digital signal output after decoding and processing the audio file in MP3 / MP4 / AVI format by the codec.

[0158] It should be understood that the first audio signal sequentially passes through the sound effect algorithm and the reliability algorithm in order to improve the sound quality, thereby achieving the purpose of optimizing the auditory experience. The sound effect algorithm and the reliability algorithm can be conventional algorithms.

[0159] For example, in the external playing scene, the low frequency sound of the audio signal plays a leading role in the speaker power consumption, so in the embodiment of the application, the high frequency digital signal is not processed, but the time delay is compensated by a full-pass filter, and then the processed low frequency digital signal is reconstructed and synthesized.

[0160] Step S102, frame processing is performed on the low frequency digital signal, and short-time Fourier transform is performed on the processed low frequency digital signal to obtain a low frequency frequency domain signal.

[0161] It can be understood that a large file will be obtained after decoding by the codec, and the digital signal processor will need large memory and computing power to process the file. For example, a 48 kHz audio file with a duration of 3 minutes has 3*60*4800 sampling points after decoding, and the digital signal processor cannot process 3*60*4800 sampling points at a time. Therefore, in the embodiments of the present application, the low-frequency digital signal can also be divided into frames according to a certain step size, and the low-frequency audio digital signal is processed to obtain a frame signal. The compensation can be set according to the performance of the mobile phone, and can be set to 5 ms, 10 ms, 20 ms, etc., which is not limited in the embodiments of the present application.

[0162] At the same time, in order to reduce the complexity of subsequent low-frequency digital signal processing, the low-frequency digital signal can also be subjected to noise filtering processing. For example, a beam algorithm or noise suppression can be used for filtering, which is not limited in the embodiments of the present application.

[0163] In step S103, a dynamic loudspeaker model is constructed according to the real-time key mechanical parameters of the loudspeaker of the electronic device.

[0164] For example, the key mechanical parameters of the loudspeaker include but are not limited to magnetic force factor Bl, stiffness coefficient Kms, mechanical impedance Rms and voice coil displacement x, and the linear and nonlinear changes of these mechanical parameters can represent different working states of the loudspeaker. At the same time, these mechanical parameters will dynamically change with the change of temperature of the loudspeaker. Therefore, when constructing the dynamic loudspeaker model, the embodiments of the present application provide two methods for updating the parameters of the loudspeaker model.

[0165] For example, in some implementations, a method for dynamically setting the parameters of the loudspeaker model following the change of temperature is provided. By using a corresponding loudspeaker analysis tool (for example: Klippel system) before the mobile phone is shipped, the change of each mechanical parameter of the loudspeaker with temperature is analyzed to obtain the mechanical parameter model of each loudspeaker, thereby obtaining the mechanical parameters at different temperatures, such as magnetic force factor Bl(x), stiffness coefficient Kms(x) and mechanical impedance Rms(v).

[0166] For example, referring to Figure 10 , Figure 10 An example of a mechanical parameter model at different temperatures is shown. Figure 10 In the Klippel system, the linear and nonlinear parameters of the loudspeaker are analyzed to obtain the mechanical parameter model at different temperatures.

[0167] It can be understood that the mechanical parameter model described above can be written into the memory of the mobile phone before the mobile phone is shipped, or the mechanical parameter model in the mobile phone can be updated regularly through the cloud platform, and the embodiments of the application do not limit this.

[0168] Exemplarily, in other implementations, a dynamic calculation method of the loudspeaker parameters according to the real-time current signal and the real-time voltage signal of the loudspeaker is provided.

[0169] Exemplarily, referring to Figure 11 , Figure 11 Exemplarily, steps of dynamically updating the loudspeaker model according to the real-time current signal and the real-time voltage signal of the loudspeaker are shown, including:

[0170] Step S1031, the current signal and the voltage signal of the loudspeaker are acquired in real time.

[0171] Step S1032, the real-time impedance model of the loudspeaker is calculated according to the current signal and the voltage signal.

[0172] Step S1033, the mechanical parameters of the loudspeaker are calculated according to the real-time impedance model.

[0173] Step S1034, the static model parameters of the loudspeaker model are updated according to the mechanical parameters.

[0174] Specifically, after the current signal V(t) and the voltage signal I(t) are obtained, Fourier transform is performed on them respectively to obtain frequency domain signals V(f) and I(f):

[0175] V(f) = F{V(t)};

[0176] I(f) = F{I(t)};

[0177] The total impedance Z total (f) of the loudspeaker is calculated according to the frequency domain signals V(f) and I(f):

[0178]

[0179] Meanwhile, in the micro loudspeaker (the loudspeaker in the mobile phone is a micro loudspeaker), therefore, the inductance Le can be ignored, and therefore, Z total (f) can also be written as:

[0180]

[0181] Wherein, R e is the DC impedance of the loudspeaker, M ms is the mechanical mass of the loudspeaker, R ms is the mechanical damping factor, and K msis the stiffness coefficient of the loudspeaker, j is the imaginary unit, and ω is the angular frequency.

[0182] The impedance is decomposed into a real part Re(Z total ) and an imaginary part Im(Z total ):

[0183] Re(Z total ) = R e + R ms ;

[0184]

[0185] Since the DC impedance R e and the mechanical mass M ms of the loudspeaker are considered as known quantities, the mechanical damping factor R ms and the stiffness coefficient K ms can be obtained, and thus the mechanical impedance Z ms (f) can be obtained:

[0186]

[0187] Based on the mechanical impedance Z ms (f), the real-time magnetic force factor Bl can be derived:

[0188]

[0189] where X(f) is the diaphragm displacement of the loudspeaker, which is considered as a known quantity.

[0190] Thus, the corresponding diaphragm velocity v(f) can be obtained:

[0191]

[0192] Since the radiation impedance R rad is:

[0193]

[0194] where a is the radius of the loudspeaker diaphragm, which is considered as a known quantity, ρ0 is the air density (about 1.2 kg / m 3 ), c is the sound speed (about 343 m / s), and k is the wave number, k = 2πf / c.

[0195] Thus, the corresponding sound pressure P(f) can be calculated:

[0196] P(f) = v(f) · R rad ;

[0197]

[0198] where P refis the reference sound pressure, generally 20μPa.

[0199] Therefore, the parameter of the static loudspeaker model is updated according to the current acquired current signal and voltage signal, and the static loudspeaker model is made to complete the correction of the real-time sound pressure. The static loudspeaker model can be constructed based on the factory parameters of the loudspeaker.

[0200] In step S104, the listening area SPL mapping is constructed based on the listening position transfer function of the loudspeaker, the listening area parameter, the dynamic loudspeaker model and the absolute threshold of hearing.

[0201] It can be understood that the listening position transfer function of the loudspeaker can be pre-stored in the memory of the mobile phone, can be acquired in real time through the cloud, or can be pre-stored in the memory of the mobile phone and updated in real time through the cloud, and the embodiments of the present application do not limit the listening position transfer function. The process of obtaining the transfer function can be combined with the process of obtaining the listening area parameter, and the process of obtaining the dynamic loudspeaker model. Figure 5 and Figure 12 are obtained.

[0202] For example, the music outdoor scene is constructed in the Figure 5 For example, the music outdoor scene is constructed in the Figure 5 , the standard microphone is set at point B or the artificial ear is set at the left and right sides of the head of the person at point C, the output signal y(t, θ, φ) of the digital signal x(t) played in the mobile phone to the standard microphone or the artificial ear is collected, where t is time, θ is azimuth angle, and φ is elevation angle.

[0203] In combination with Figure 12 , the music signal x(t) is output to the loudspeaker as an electric signal through the codec of the mobile phone, the loudspeaker is driven to emit a sound signal, the sound signal is transmitted through the air to the standard microphone or the artificial ear, the sound signal is converted into an electric signal by the standard microphone or the artificial ear, and the electric signal is transmitted to the computer end through the amplifier and the sound card for analysis, so that the output signal y(t, θ, φ) at different positions is obtained. The steps include:

[0204] In step S1041, the music signal x(t) and the standard sound source signal sent by the mobile phone in a preset time period are collected at a preset collection point through the standard microphone or the artificial ear based on the current posture information of the mobile phone, and the output signal y(t, θ, φ) is obtained, wherein the posture information of the mobile phone includes distance information of the collection point, azimuth angle of the mobile phone and elevation angle of the mobile phone.

[0205] It can be understood that due to the distance of the mobile phone from the collection point, and the azimuth angle and the elevation angle of the mobile phone, the finally collected output signal is also different, so before signal collection, the attitude information of the mobile phone needs to be set, which can be set according to the statistical listening habits of users, and the attitude information can be obtained by big data analysis. The distance information can be set to different values according to different types of mobile terminals. For example, when the mobile terminal is a mobile phone, the distance is set to 30 centimeters (cm), and when the mobile terminal is a tablet, the distance is set to 50 cm. The azimuth angle can also be set according to requirements, and is usually set to 90°. The elevation angle is set according to the operation habits of the user, and can be set to 0°, 30° and 45°.

[0206] It should be noted that after the mobile phone attitude information is set, a prompt box can also be popped up in the user's listening scene to let the user select the habit attitude information; or the attitude information of the mobile phone is dynamically obtained through the related sensors of the mobile phone, so as to dynamically update the transfer function and update the listening area SPL mapping, and then dynamically filter the music signal, so as to reasonably balance the relationship between sound quality and power consumption.

[0207] It should also be understood that the above standard sound source is used to calibrate the absolute amplitude, and the standard sound source is a single frequency 94dB SPL standard sound source of 1kHz.

[0208] Step S1042, based on the Fourier transformed music signal x(t) and the Fourier transformed output signal y(t, θ, φ), the listening position transfer function is calculated.

[0209] Specifically, the music signal x(t) is Fourier transformed:

[0210]

[0211] Where j is an imaginary number, f is the frequency, and t is the time.

[0212] The output signal y(t, θ, φ) is Fourier transformed:

[0213]

[0214] The listening position transfer function is obtained:

[0215]

[0216] It can be understood that the above listening area parameters include a preset listening area range. That is Figure 5 The head movement range shown in the above formula, and the distance between the head contour is the distance y. The distance y can be set according to requirements, and in general cases, it can be set to 20 cm.

[0217] Exemplarily, after the listening area range is determined, the reference position sound pressure is constructed according to the loudspeaker model parameters obtained above (here, the reference position can be the B point or the C point shown in Figure 5 After the reference position loudspeaker sound pressure model is obtained, the step of establishing the listening area sound pressure mapping is described with reference to Figure 13 , which includes:

[0218] In step S1043, the head movement range is approximated as a spherical surface, and the edge sound pressure envelope of the listening area is obtained according to the spherical wave formula and the reference position loudspeaker sound pressure model.

[0219] It can be understood that in the audio playback scene, the head movement range can be considered as a spherical surface, and thus the edge sound pressure envelope of the listening area can be obtained by the spherical wave formula:

[0220]

[0221] wherein A is the sound pressure at the reference position, r is the distance, ω is the angular frequency, k is the wave number, and k = 2πf / c.

[0222] In step S1044, the transfer function is corrected according to the edge sound pressure envelope, and the sound pressure model of multiple positions in the listening area is obtained.

[0223] It can be understood that the embodiments of the present application process low-frequency signals, and thus the phase change in the listening area is small. Therefore, in combination with the listening position transfer function described above, the transfer function correction formula of the arbitrary position M i can be derived from the transfer function of the reference position M0 as follows:

[0224]

[0225] wherein r0 is the distance from the sound source to the reference position during modeling, and r s is the distance from the sound source to the position M i , that is:

[0226] Specifically, the reference position M0 can be the B point or the C point shown in Figure 5 , and the arbitrary position M i can be the D point position in Figure 5 .

[0227] In step S1045, the lower envelope line of the sound pressure model of multiple positions is calculated, and the edge sound pressure envelope of the listening area is obtained.

[0228] Specifically, after the sound pressures of multiple positions are obtained, the lower envelope line of the sound pressures of multiple positions is set as the edge sound pressure envelope of the listening area, and the schematic is shown in Figure 14.

[0229] It can be understood that, before calculating the lower envelope of the sound pressure of multiple positions, the sound pressure of each position can also be filtered based on the absolute threshold of human ears, so as to reduce the calculation amount of the sound pressure SPL mapping of the listening area. The absolute threshold of human ears refers to the minimum sound pressure level (SPL) that can be perceived by human ears. It is a basic characteristic of the auditory system, indicating the minimum sound pressure that human ears can just hear at a certain frequency. The absolute threshold is usually expressed in decibels (dB SPL) and varies at different frequencies. For details, refer to Figure 15 .

[0230] Step S1046, based on the preset adjustment factor λ, the edge sound pressure envelope is adjusted to obtain the final listening area SPL mapping.

[0231] It can be understood that the size of the edge sound pressure envelope has an impact on power consumption and sound quality; when the edge sound pressure envelope is larger, the power consumption is reduced more, and the possibility of sound quality loss is greater; when the edge sound pressure envelope is smaller, the power consumption is reduced less, and the possibility of sound quality loss is smaller. Therefore, the size of the edge sound pressure envelope can also be adjusted by the above-mentioned adjustment factor λ. The adjustment factor λ can be set according to requirements.

[0232] Step S105, inputting the low-frequency frequency signal and the listening area SPL mapping into a psychoacoustic model to obtain a masking spectrum.

[0233] Specifically, the low-frequency frequency signal and the above-mentioned listening area SPL mapping are input into the above-mentioned psychoacoustic model, so as to calculate the masking spectrum of the low-frequency frequency domain signal.

[0234] Step S106, filtering the low-frequency frequency domain signal based on the masking spectrum, and performing inverse Fourier transform on the filtered low-frequency frequency domain signal to obtain a low-frequency time domain signal after masking processing.

[0235] Specifically, a corresponding filter is designed according to the above-mentioned masking spectrum, and the low-frequency frequency domain signal is filtered based on the filter to obtain a filtered low-frequency frequency domain signal, and the filtered low-frequency frequency domain signal is subjected to inverse Fourier transform to obtain a low-frequency time domain signal after masking processing.

[0236] In a possible implementation, since the frequency domain filter belongs to non-causal filtering, it can also be subjected to inverse Fourier transform after causal constraint. Exemplary causal constraints include introducing a 1-frame delay, so that the filter becomes a semi-causal filter.

[0237] Step S107, reconstructing and synthesizing the high-frequency digital signal and the low-frequency time domain signal, and sending the obtained second audio signal to the loudspeaker after digital-to-analog conversion.

[0238] Specifically, by reconstructing and synthesizing the high-frequency digital signal and the low-frequency time-domain signal, the obtained second audio signal is sent to the speaker of the mobile phone after digital-to-analog conversion, so that the speaker plays the second audio signal.

[0239] For example, when playing the second audio signal, based on the mathematical model of Figure 5 , the audio signal S2 output by the speaker is collected at the B point or the C point, and the collected audio is subjected to Fourier transform, and is compared and analyzed with the original audio signal S and the above-mentioned audio signal S1, and the analysis effect is shown in Figure 16 . As can be seen from Figure 16 , the audio signal S2 processed by the audio signal processing method provided in the present application has lower power consumption than the original audio signal S, and the possible performance of the loss of sound quality is smaller than that of the audio signal S1 processed by directly using the frequency domain masking model, and the sound effect is better.

[0240] Therefore, in the embodiments of the present application, the music signal to be played is processed by a conventional sound effect algorithm and a reliability protection algorithm, and then divided into a high-frequency part and a low-frequency part by a frequency division filter. The low-frequency part is transformed into the frequency domain by short-time Fourier transform, and the frequency domain signal and the user's listening area SPL are input into the psychoacoustic model (PAM), and the masking spectrum of the current signal is calculated. Then, a frequency domain filter system is designed according to the masking spectrum, and after the causality constraint, the processed low-frequency time-domain signal is output by inverse short-time Fourier transform. The low-frequency time-domain signal is reconstructed and synthesized with the high-frequency part, so that the power of the speaker is reduced without reducing the sound effect of the audio signal, and the listening experience of the user is improved.

[0241] In addition, it can be understood that the electronic device includes hardware and / or software modules corresponding to each function to realize the above functions. The algorithm steps of each example described in conjunction with the embodiments disclosed herein can be realized in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed by hardware or computer software driven hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application in conjunction with the embodiments, but such implementation should not be considered beyond the scope of the present application.

[0242] In addition, it should be noted that the audio signal processing method provided by the above-mentioned embodiments realized by the electronic device in the actual application scenario can also be executed by a chip system included in the electronic device, wherein the chip system can include a processor. The chip system can be coupled with a memory, so that the chip system invokes the computer program stored in the memory when running, and realizes the steps executed by the above-mentioned electronic device. The processor in the chip system can be an application processor or a processor other than an application processor.

[0243] In addition, the embodiments of the present application also provide a computer readable storage medium, which stores computer instructions, when the computer instructions run on the electronic device, make the electronic device execute the above-mentioned related method steps to realize the audio signal processing method in the above-mentioned embodiments.

[0244] In addition, the embodiments of the present application also provide a computer program product, when the computer program product runs on the electronic device, make the electronic device execute the above-mentioned related steps to realize the audio signal processing method in the above-mentioned embodiments.

[0245] In addition, the embodiments of the present application also provide a chip (which can also be a component or a module), which can include one or more processing circuits and one or more transceiver pins; wherein the transceiver pins and the processing circuits communicate with each other through internal connection paths, and the processing circuits execute the above-mentioned related method steps to realize the audio signal processing method in the above-mentioned embodiments, to control the receiving pins to receive signals, and to control the sending pins to send signals.

[0246] In addition, as known from the above description, the electronic device, computer readable storage medium, computer program product or chip provided by the embodiments of the present application are all used to execute the corresponding methods provided above, so the beneficial effects that can be achieved are referred to the beneficial effects in the corresponding methods provided above, which will not be repeated here.

[0247] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacements for some technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A method of audio signal processing, characterized by, Applied to an electronic device, the electronic device including a speaker, the method includes: A dynamic speaker model is constructed based on the key mechanical parameters of the speaker obtained in real time. Based on the speaker's listening position transfer function, listening zone parameters, and the dynamic speaker model, a listening zone sound pressure mapping is constructed. The listening zone parameters are used to characterize the user's listening range, and the listening position transfer function is used to characterize the relationship between the first audio digital signal received at different listening positions and the original audio digital signal. The original audio digital signal is subjected to frequency division processing. The low-frequency digital signal obtained by frequency division processing and the sound pressure of the listening area are mapped and input into a preset psychoacoustic model to obtain the masking spectrum. The low-frequency digital signal is masked and filtered based on the masking spectrum. The low-frequency digital signal after masking and filtering and the high-frequency digital signal obtained by frequency division are reconstructed and synthesized to obtain the audio signal for playback by the speaker. The step of constructing a sound pressure mapping of the listening area based on the listening position transfer function of the loudspeaker, the listening area parameters, and the dynamic loudspeaker model includes: Based on the obtained listening area parameters and the dynamic loudspeaker model, the edge sound pressure envelope of the listening area is obtained; The obtained listening position transfer function is corrected based on the edge sound pressure envelope to obtain sound pressure models for multiple locations in the listening area; The lower envelope of the sound pressure model at multiple locations is calculated to obtain the sound pressure mapping of the listening area.

2. The method according to claim 1, characterized in that, The construction of a dynamic loudspeaker model based on the key mechanical parameters of the loudspeaker acquired in real time includes: Based on the real-time acquired current and voltage signals of the loudspeaker, an impedance model of the loudspeaker is constructed. The impedance model is used to characterize the relationship between the current and voltage signals and the key mechanical parameters. The key mechanical parameters of the loudspeaker are calculated based on the impedance model. These key mechanical parameters include the magnetic force factor, stiffness coefficient, and mechanical impedance. A dynamic loudspeaker model was constructed based on the key mechanical parameters.

3. The method according to claim 2, characterized in that, The construction of a dynamic loudspeaker model based on the key mechanical parameters of the loudspeaker obtained in real time also includes: Based on the real-time temperature value of the speaker, the corresponding key mechanical parameters are selected from the preset mechanical parameter model to construct a dynamic model of the speaker.

4. The method according to claim 1, characterized in that, After correcting the obtained listening position transfer function based on the edge sound pressure envelope to obtain sound pressure models for multiple locations in the listening area, the method further includes: The sound pressure model at multiple locations is filtered based on a preset absolute hearing threshold.

5. The method according to claim 1, characterized in that, The step of calculating the lower envelope of the sound pressure model at multiple locations to obtain the sound pressure mapping of the listening area further includes: Based on a preset adjustment factor, the size of the lower envelope is adjusted to obtain the sound pressure mapping of the listening area.

6. The method according to claim 1, characterized in that, The step of performing frequency division processing on the original audio digital signal, and mapping the low-frequency digital signal obtained from the frequency division processing and the sound pressure level of the listening area to a preset psychoacoustic model to obtain a masking spectrum includes: The original audio digital signal is filtered using a frequency division filter to obtain a high-frequency digital signal and a low-frequency digital signal; The low-frequency digital signal is divided into frames, and a Fourier transform is performed on the low-frequency digital signal of each frame to obtain the low-frequency frequency domain signal of each frame. The low-frequency signal and the sound pressure mapping of the listening area in each frame are input into a preset psychoacoustic model to obtain the masking spectrum.

7. The method according to claim 6, characterized in that, The masking filtering process for the low-frequency digital signal based on the masking spectrum includes: A filter is designed based on the masking spectrum, and the low-frequency domain signal of each frame is masked and filtered based on the filter.

8. The method according to claim 7, characterized in that, The step of designing a filter based on the masking spectrum, and performing masking filtering on each frame of the low-frequency domain signal based on the filter, includes: A filter is designed based on the masking spectrum, and the low-frequency domain signal of two adjacent frames is masked and filtered based on the filter.

9. The method according to claim 7 or 8, characterized in that, The method further includes: Each frame of the low-frequency frequency domain signal after masking and filtering is subjected to inverse Fourier transform to obtain the low-frequency digital signal after masking and filtering.

10. An electronic device, characterized in that, The electronic device includes: a memory and a processor, the memory and the processor being coupled; the memory stores program instructions, which, when executed by the processor, cause the electronic device to perform the audio signal processing method as described in any one of claims 1 to 9.

11. A computer-readable storage medium, characterized in that, The method includes a computer program that, when run on an electronic device, causes the electronic device to perform the audio signal processing method as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Micro-speaker audio power reproduction system and method with reduced energy use and thermal protection using micro-speaker electro-acoustic response and human hearing thresholds

    US11153682B1

  • Psychoacoustics for improved audio reproduction and speaker protection

    WO2017222562A1