Voice definition control method and device and electronic equipment

By obtaining the signal-to-noise ratio and formant peak of the voice signal, adjusting the formant peak of the voice signal, and generating adaptive voice signals, the user's need for voice clarity control in different scenarios is solved, and flexible voice clarity adjustment is achieved.

CN120388574APending Publication Date: 2025-07-29HONOR DEVICE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410083188.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-19
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

In different scenarios, users want to adjust voice clarity as needed for smooth communication or privacy protection, but the prior art is difficult to achieve flexible voice clarity control.

Method used

By obtaining the signal-to-noise ratio and formant peak of the speech signal, adjusting the formant peak of the speech signal to generate a speech signal that is adapted to the current scene, including frame-based and window processing, Fourier transform, sound pressure level model training and time-frequency domain masking analysis, and generating adaptive speech signals to adjust clarity.

Benefits of technology

It realizes automatic adjustment of voice signal clarity according to the scene to meet users' communication needs and privacy protection needs in different environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388574A_ABST
    Figure CN120388574A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a voice definition control method and device and electronic equipment, and the method comprises the steps that the electronic equipment obtains a to-be-processed first voice signal, determines the signal-to-noise ratio of the first voice signal, obtains a formant of a first acoustic signal, and obtains the formant of the first acoustic signal according to the signal-to-noise ratio of the first voice signal; determining an adjustment amount of a formant of the first acoustic signal; therefore, the definition of the voice signal heard by the user using the electronic equipment can be adjusted according to the current scene of the electronic equipment. And finally, the electronic equipment generates a second acoustic signal according to the adjustment amount, the upper bound and the lower bound of the formant of the first acoustic signal, the fundamental tone frequency of the first acoustic signal and the signal spectrum corresponding to the second voice signal, obtains a to-be-played third voice signal according to the second acoustic signal, and plays the third voice signal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the technical field of intelligent terminals, and particularly to a method, an apparatus, and an electronic device for controlling speech clarity. Background Art

[0002] In daily voice or video calls, the most core and basic purpose of users is to exchange information. Therefore, speech clarity is always the most important performance indicator in communication. The improvement of speech clarity can enable users to enjoy a smooth communication experience even when they are in different physical spaces. However, in daily use, in some scenarios, users hope that the speech clarity is higher so that they can communicate smoothly with the other party, while in some scenarios, in order to protect their privacy, users do not want people around them to hear their call content clearly. Therefore, a solution that can control the speech clarity of a call according to the current scenario is needed. Summary of the Invention

[0003] The embodiments of the present application provide a method, an apparatus, and an electronic device for controlling speech clarity. The embodiments of the present application also provide a computer-readable storage medium to adjust the clarity of the speech signal heard by the user using the electronic device according to the current scenario where the electronic device is located.

[0004] In a first aspect, the embodiments of the present application provide a method for controlling speech clarity, including: obtaining a first speech signal to be processed; determining the signal-to-noise ratio of the first speech signal and obtaining the formants of a first acoustic signal, where the first acoustic signal is obtained by converting the first speech signal; determining an adjustment amount of the formants of the first acoustic signal according to the signal-to-noise ratio of the first speech signal; generating a second acoustic signal according to the adjustment amount, the upper and lower bounds of the formants of the first acoustic signal, the fundamental frequency of the first acoustic signal, and the signal spectrum corresponding to a second speech signal, where the second speech signal is a speech signal received by the electronic device after the first speech signal; obtaining a third speech signal to be played according to the second acoustic signal; and playing the third speech signal.

[0005] In the above method for controlling speech clarity, after the electronic device obtains the first speech signal to be processed, it determines the signal-to-noise ratio of the first speech signal and obtains the formants of the first acoustic signal. Then, according to the signal-to-noise ratio of the first speech signal, it determines the adjustment amount of the formants of the first acoustic signal. Finally, the electronic device can generate a second acoustic signal based on the adjustment amount, the upper and lower bounds of the formants of the first acoustic signal, the fundamental frequency of the first acoustic signal, and the signal spectrum corresponding to the second speech signal, obtain a third speech signal to be played according to the second acoustic signal, and play the third speech signal. Since the adjustment amount of the formants of the first acoustic signal is determined according to the signal-to-noise ratio of the first speech signal, it is possible to adjust the clarity of the speech signal heard by the user using the electronic device according to the current scene where the electronic device is located.

[0006] In one possible implementation, the determining the signal-to-noise ratio of the first speech signal includes: performing frame segmentation and windowing processing on the first speech signal to obtain a first time-domain signal; and calculating the signal-to-noise ratio of the first speech signal according to the first time-domain signal.

[0007] In one possible implementation, the obtaining the formants of the first acoustic signal includes: performing frame segmentation and windowing processing on the first speech signal to obtain a first time-domain signal; converting the first time-domain signal into a first acoustic signal, where the first acoustic signal is a frequency-domain signal; and performing spectral envelope analysis on the first acoustic signal to obtain the formants of the first acoustic signal.

[0008] In one possible implementation, the generating the second acoustic signal based on the adjustment amount, the upper and lower bounds of the formants of the first acoustic signal, the fundamental frequency of the first acoustic signal, and the signal spectrum corresponding to the first speech signal includes: based on the fundamental frequency of the first acoustic signal, adjusting the peak value of the signal spectrum corresponding to the second speech signal according to the adjustment amount to generate a second acoustic signal, where the upper and lower bounds of the formants of the second acoustic signal are respectively parallel to the upper and lower bounds of the formants of the first acoustic signal; the difference between the upper bound of the formants of the second acoustic signal and the upper bound of the formants of the first acoustic signal is the adjustment amount, and the difference between the lower bound of the formants of the second acoustic signal and the lower bound of the formants of the first acoustic signal is the adjustment amount.

[0009] In one possible implementation, the converting the first time-domain signal into a first acoustic signal includes: performing a Fourier transform on the first time-domain signal to obtain a first frequency-domain signal; and converting the first frequency-domain signal into a first acoustic signal through a pre-trained sound pressure level model.

[0010] In a possible implementation, the pre-trained sound pressure level model includes a sound pressure level model for the left ear and a sound pressure level model for the right ear; before converting the first frequency-domain signal into a first acoustic signal through the pre-trained sound pressure level model, the following steps are further included: converting a digital signal of a predetermined frequency into a third acoustic signal; playing the third acoustic signal; using a standard microphone to receive sound, measuring the sound pressure value of a fourth acoustic signal collected by the standard microphone through a sound level meter, and using the measured sound pressure value as a reference sound pressure; and using an artificial head to perform binaural sound reception to obtain a left-ear acoustic signal and a right-ear acoustic signal respectively; obtaining a sound pressure level model for the left ear according to the digital signal of the predetermined frequency, the left-ear acoustic signal, and the reference sound pressure; and obtaining a sound pressure level model for the right ear according to the digital signal of the predetermined frequency, the right-ear acoustic signal, and the reference sound pressure.

[0011] In a possible implementation, before generating the second acoustic signal, the following steps are further included: performing pitch analysis on the first acoustic signal to obtain the pitch frequency of the first acoustic signal.

[0012] In a possible implementation, before generating the second acoustic signal, the following steps are further included: performing time-frequency domain masking analysis on a second speech signal to obtain a signal spectrum corresponding to the second speech signal; wherein, the signal spectrum corresponding to the second speech signal includes the signal spectrum actually perceived by the human ear for the second speech signal.

[0013] In a possible implementation, performing time-frequency domain masking analysis on the second speech signal to obtain the signal spectrum corresponding to the second speech signal includes: performing short-time Fourier transform on the second speech signal; convolving the signal obtained by the short-time Fourier transform with a masking mask; and performing inverse short-time Fourier transform on the convolved signal to obtain the signal spectrum corresponding to the second speech signal.

[0014] In a possible implementation, before convolving the signal obtained by the short-time Fourier transform with the masking mask, the following steps are further included: calculating the signal-to-noise ratio of the second speech signal; determining a masking threshold according to the signal-to-noise ratio of the second speech signal; and generating the masking mask according to the masking threshold.

[0015] In a possible implementation, before generating the second acoustic signal, the following steps are further included: determining the upper bound and the lower bound of the formant of the first acoustic signal through piecewise interpolation according to the frequency and spectral amplitude of the formant of the first acoustic signal.

[0016] Second aspect, an embodiment of the present application provides a device for controlling speech clarity. This device is included in an electronic device and has the function of implementing the behavior of the electronic device in the first aspect and the possible implementation manners of the first aspect. The function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules or units corresponding to the above functions. For example, an acquisition module, a processing module, a calculation module, a conversion module, an analysis module, a determination module, a generation module, a playback module, etc.

[0017] Third aspect, an embodiment of the present application provides an electronic device, including: one or more processors; a memory; a plurality of applications; and one or more computer programs, wherein the above one or more computer programs are stored in the above memory, and the above one or more computer programs include instructions, when the above instructions are executed by the above electronic device, the above electronic device is caused to execute the method provided in the first aspect.

[0018] It should be understood that the second aspect and the third aspect of the embodiments of the present application are consistent with the technical solutions of the first aspect of the embodiments of the present application, and the beneficial effects obtained by each aspect and the corresponding feasible implementation manners are similar, and will not be elaborated herein.

[0019] Fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored, and when it runs on a computer, the computer is caused to execute the method provided in the first aspect.

[0020] Fifth aspect, an embodiment of the present application provides a computer program, which is used to execute the method provided in the first aspect when the computer program is executed by a computer.

[0021] In a possible design, the program in the fifth aspect can be stored in whole or in part on a storage medium packaged together with the processor, or can be stored in whole or in part on a memory not packaged together with the processor. Description of the Drawings

[0022] Figure 1 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application;

[0023] Figure 2 It is an implementation block diagram of a method for controlling speech clarity provided by an embodiment of the present application;

[0024] Figure 3 It is a flowchart of a method for controlling speech clarity provided by an embodiment of the present application;

[0025] Figure 4 It is a schematic diagram of obtaining an SPL model provided by an embodiment of the present application;

[0026] Figure 5 A schematic diagram of spectrum envelope analysis provided in one embodiment of the present application;

[0027] FIG6( a ) is a schematic diagram of the principle of a PV provided by an embodiment of the present application;

[0028] FIG6( b ) is a schematic diagram of the upper and lower bounds of the resonance peak provided by one embodiment of the present application;

[0029] Figure 7 A schematic diagram of pitch analysis provided in one embodiment of the present application;

[0030] Figure 8 A flowchart of a method for controlling speech clarity provided in another embodiment of the present application;

[0031] Figure 9 A schematic diagram of implementing time-frequency domain masking analysis provided in one embodiment of the present application;

[0032] Figure 10 A schematic diagram of a specific example of time-frequency domain masking analysis provided by one embodiment of the present application. DETAILED DESCRIPTION

[0033] The terms used in the implementation section of this application are only used to explain the specific embodiments of this application and are not intended to limit this application.

[0034] In daily voice or video calls, users often desire higher voice clarity in some scenarios. For example, in low signal-to-noise ratio scenarios, users may want to improve the clarity of their own voice signals to facilitate smooth communication. Alternatively, users may want to reduce the clarity of their own voice signals to protect their privacy while preventing others from hearing their conversations. Therefore, a solution is needed to control the voice clarity of calls based on the current scenario.

[0035] Based on the above problems, an embodiment of the present application provides a method for controlling speech clarity, which can adjust the clarity of the speech signal heard by a user of the electronic device according to the current scenario of the electronic device.

[0036] The speech clarity control method provided in the embodiments of the present application can be applied to electronic devices, wherein the above-mentioned electronic devices can be smart phones, tablet computers, wearable devices, vehicle-mounted devices, augmented reality (AR) / virtual reality (VR) devices, laptop computers, ultra-mobile personal computers (UMPCs), netbooks or personal digital assistants (PDAs), etc.; the embodiments of the present application do not impose any restrictions on the specific type of electronic devices.

[0037] For example, Figure 1 A schematic diagram of the structure of an electronic device provided in one embodiment of the present application is shown in FIG. Figure 1 As shown, the electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, an air pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0038] It should be understood that the structures illustrated in the embodiments of the present application do not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0039] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU). The different processing units may be independent devices or integrated into one or more processors.

[0040] The controller can generate operation control signals according to the instruction operation code and timing signal to complete the control of instruction fetching and execution.

[0041] The processor 110 may also be provided with a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. The memory may store instructions or data that the processor 110 has just used or cyclically used. If the processor 110 needs to use the instruction or data again, it may be directly called from the memory. This avoids repeated access, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.

[0042] In some embodiments, the processor 110 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface.

[0043] The USB interface 130 is an interface that conforms to the USB standard specification. Specifically, it can be a Mini USB interface, a Micro USB interface, a USB Type C interface, etc. The USB interface 130 can be used to connect a charger to charge the electronic device 100, and can also be used to transfer data between the electronic device 100 and peripheral devices. It can also be used to connect headphones to play audio. This interface can also be used to connect other electronic devices, such as AR devices, etc.

[0044] It can be understood that the interface connection relationship between the modules illustrated in the embodiments of the present application is only for illustrative purposes and does not constitute a structural limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 can also adopt different interface connection methods in the above embodiments, or a combination of multiple interface connection methods.

[0045] The charging management module 140 is used to receive a charging input from a charger. Among them, the charger can be a wireless charger or a wired charger. In some embodiments of wired charging, the charging management module 140 can receive the charging input of the wired charger through the USB interface 130. In some embodiments of wireless charging, the charging management module 140 can receive the wireless charging input through the wireless charging coil of the electronic device 100. While charging the battery 142, the charging management module 140 can also supply power to the electronic device 100 through the power management module 141.

[0046] The power management module 141 is used to connect the battery 142, the charging management module 140 and the processor 110. The power management module 141 receives the input from the battery 142 and / or the charging management module 140 and supplies power to the processor 110, the internal memory 121, the display screen 194, the camera 193, the wireless communication module 160, etc. The power management module 141 can also be used to monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage, impedance). In some other embodiments, the power management module 141 can also be provided in the processor 110. In other embodiments, the power management module 141 and the charging management module 140 can also be provided in the same device.

[0047] The wireless communication function of the electronic device 100 can be implemented through antenna 1, antenna 2, the mobile communication module 150, the wireless communication module 160, the modulation and demodulation processor, and the baseband processor, etc.

[0048] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in the electronic device 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be multiplexed to improve the utilization rate of the antennas. For example: Antenna 1 can be multiplexed as the diversity antenna of the wireless local area network. In some other embodiments, the antenna can be used in combination with a tuning switch.

[0049] The mobile communication module 150 can provide solutions for wireless communications including 2G / 3G / 4G / 5G applied to the electronic device 100. The mobile communication module 150 may include at least one filter, a switch, a power amplifier, a low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves from the antenna 1, and filter, amplify, and process the received electromagnetic waves, and transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modulation and demodulation processor, and convert it into electromagnetic waves for radiation through the antenna 1. In some embodiments, at least some of the functional modules of the mobile communication module 150 can be set in the processor 110. In some embodiments, at least some of the functional modules of the mobile communication module 150 can be set in the same device as at least some of the modules of the processor 110.

[0050] The modem processor may include a modulator and a demodulator. The modulator is used to modulate the low-frequency baseband signal to be transmitted into a medium-high frequency signal. The demodulator is used to demodulate the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After being processed by the baseband processor, the low-frequency baseband signal is passed to the application processor. The application processor outputs a sound signal through an audio device (not limited to the speaker 170A, the receiver 170B, etc.) or displays an image or video through the display screen 194. In some embodiments, the modem processor may be an independent device. In other embodiments, the modem processor may be independent of the processor 110 and be set in the same device as the mobile communication module 150 or other functional modules.

[0051] The wireless communication module 160 can provide wireless communication solutions including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR), etc., which are applied to the electronic device 100. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via the antenna 2, frequency modulates and filters the electromagnetic wave signals, and sends the processed signals to the processor 110. The wireless communication module 160 can also receive the signal to be sent from the processor 110, frequency modulate it, amplify it, and convert it into electromagnetic waves for radiation through the antenna 2.

[0052] In some embodiments, antenna 1 of electronic device 100 is coupled to mobile communication module 150, and antenna 2 is coupled to wireless communication module 160, such that electronic device 100 can communicate with a network and other devices through wireless communication technologies. The wireless communication technologies may include global system for mobile communications (GSM), general packet radio service (GPRS), code division multiple access (CDMA), wideband code division multiple access (WCDMA), time-division code division multiple access (TD-SCDMA), long term evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technology, etc. The GNSS may include global positioning system (GPS), global navigation satellite system (GLONASS), beidou navigation satellite system (BDS), quasi-zenith satellite system (QZSS), and / or satellite based augmentation systems (SBAS).

[0053] Electronic device 100 implements a display function through a GPU, display screen 194, and an application processor, etc. The GPU is a microprocessor for image processing, and is connected to display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. Processor 110 may include one or more GPUs, which execute program instructions to generate or change display information.

[0054] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. The display panel can adopt a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Miniled, a MicroLed, a Micro-oLed, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device 100 may include one or N display screens 194, where N is a positive integer greater than 1.

[0055] The electronic device 100 can implement the shooting function through the ISP, the camera 193, the video codec, the GPU, the display screen 194, and the application processor, etc.

[0056] The ISP is used to process the data fed back by the camera 193. For example, when taking a photo, the shutter is opened, and the light passes through the lens and is transmitted to the camera sensor. The optical signal is converted into an electrical signal, and the camera sensor transmits the electrical signal to the ISP for processing and converts it into an image visible to the naked eye. The ISP can also perform algorithm optimization on the noise, brightness, and skin color of the image. The ISP can also optimize parameters such as the exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera 193.

[0057] The camera 193 is used to capture static images or videos. The object generates an optical image through the lens and projects it onto the sensor. The sensor can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The sensor converts the optical signal into an electrical signal, and then transmits the electrical signal to the ISP to convert it into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in standard RGB, YUV, etc. formats. In some embodiments, the electronic device 100 may include one or N cameras 193, where N is a positive integer greater than 1.

[0058] The digital signal processor is used to process digital signals. In addition to processing digital image signals, it can also process other digital signals. For example, when the electronic device 100 selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy, etc.

[0059] The video codec is used to compress or decompress digital videos. The electronic device 100 can support one or more video codecs. In this way, the electronic device 100 can play or record videos in multiple coding formats, such as: Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.

[0060] The NPU is a neural-network (NN) computing processor. By learning from the structure of biological neural networks, such as learning from the transmission pattern between human brain neurons, it can quickly process the input information and can also continuously self-learn. Through the NPU, applications such as intelligent cognition of the electronic device 100 can be realized, such as: image recognition, face recognition, speech recognition, text understanding, etc.

[0061] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external memory interface 120 to achieve the data storage function. For example, files such as music and videos are saved in the external memory card.

[0062] The internal memory 121 can be used to store computer-executable program code, and the executable program code includes instructions. The internal memory 121 can include a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as the sound playback function, the image playback function, etc.). The data storage area can store the data created during the use of the electronic device 100 (such as audio data, phone book, etc.). In addition, the internal memory 121 can include high-speed random access memory and can also include non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc. The processor 110 executes various functional applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 121 and / or the instructions stored in the memory provided in the processor.

[0063] The electronic device 100 can implement audio functions through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone jack 170D, and the application processor, etc. For example, music playback, recording, etc.

[0064] The audio module 170 is used to convert digital audio information into an analog audio signal for output, and is also used to convert an analog audio input into a digital audio signal. The audio module 170 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 170 can be disposed in the processor 110, or some functional modules of the audio module 170 can be disposed in the processor 110.

[0065] The speaker 170A, also known as the "loudspeaker", is used to convert an audio electrical signal into a sound signal. The electronic device 100 can listen to music or hands-free calls through the speaker 170A.

[0066] The receiver 170B, also known as the "earpiece", is used to convert an audio electrical signal into a sound signal. When the electronic device 100 answers a call or a voice message, the user can listen to the voice by bringing the receiver 170B close to the ear.

[0067] The microphone 170C, also known as the "microphone" or "transmitter", is used to convert a sound signal into an electrical signal. When making a call or sending a voice message, the user can speak by bringing the mouth close to the microphone 170C to input the sound signal into the microphone 170C. The electronic device 100 can be provided with at least one microphone 170C. In some other embodiments, the electronic device 100 can be provided with two microphones 170C, which can not only collect sound signals but also implement a noise reduction function. In some other embodiments, the electronic device 100 can also be provided with three, four or more microphones 170C to collect sound signals, reduce noise, identify the sound source, and implement functions such as directional recording.

[0068] The headphone jack 170D is used to connect a wired headphone. The headphone jack 170D can be a USB interface 130, or a 3.5mm open mobile terminal platform (OMTP) standard interface, or a cellular telecommunications industry association of the USA (CTIA) standard interface.

[0069] The keys 190 include a power-on key, volume keys, etc. The keys 190 can be mechanical keys or touch keys. The electronic device 100 can receive key inputs and generate key signal inputs related to the user settings and function controls of the electronic device 100.

[0070] The motor 191 can generate vibration prompts. The motor 191 can be used for incoming call vibration prompts and also for touch vibration feedback. For example, touch operations for different applications (such as taking pictures, audio playing, etc.) can correspond to different vibration feedback effects. For touch operations on different regions of the display screen 194, the motor 191 can also correspond to different vibration feedback effects. Different application scenarios (such as time reminder, receiving messages, alarm clock, games, etc.) can also correspond to different vibration feedback effects. The touch vibration feedback effect can also support customization.

[0071] The indicator 192 can be an indicator light and can be used to indicate the charging state, power change, and can also be used to indicate messages, missed calls, notifications, etc.

[0072] The SIM card interface 195 is used to connect the SIM card. The SIM card can be inserted into or removed from the SIM card interface 195 to achieve contact and separation from the electronic device 100. The electronic device 100 can support 1 or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface 195 can support Nano SIM cards, Micro SIM cards, SIM cards, etc. Multiple cards can be inserted into the same SIM card interface 195 at the same time. The types of the multiple cards can be the same or different. The SIM card interface 195 can also be compatible with different types of SIM cards. The SIM card interface 195 can also be compatible with external memory cards. The electronic device 100 interacts with the network through the SIM card to implement functions such as calls and data communication. In some embodiments, the electronic device 100 uses an eSIM, that is, an embedded SIM card. The eSIM card can be embedded in the electronic device 100 and cannot be separated from the electronic device 100. [[ID=?]] [[ID=?]]

[0073] For ease of understanding, the following embodiments of the present application will be described by taking an electronic device with the [[ID=?]] Figure 1 shown structure as an example, and in combination with the accompanying drawings and application scenarios, the method for controlling voice clarity provided by the embodiments of the present application will be specifically described. [[ID=?]] [[ID=?]]

[0074] [[ID=?]] Figure 2 The implementation block diagram of the method for controlling voice clarity provided by an embodiment of the present application is shown below. The method for controlling voice clarity provided by the embodiments of the present application will be introduced in combination with [[ID=?]] Figure 2 below. [[ID=?]] Figure 3 The flowchart of the method for controlling voice clarity provided by an embodiment of the present application is shown as [[ID=?]] Figure 3 shown. The above method for controlling voice clarity may include: [[ID=?]] [[ID=?]]

[0075] Step 301, the electronic device 100 obtains a first voice signal to be processed. [[ID=?]] [[ID=?]]

[0076] Specifically, the electronic device 100 can obtain a first voice signal to be processed through the microphone 170C. Taking the example that the user uses the electronic device 100 to make a voice call with the other end, after the electronic device 100 receives the voice signal sent by the other end, the electronic device 100 can play the received voice signal through the speaker 170A, and then obtain the first voice signal to be processed through the microphone 170C. That is to say, the first voice signal obtained by the microphone 170C is the voice signal played by the speaker 170A.

[0077] Step 302, the electronic device 100 determines the signal-to-noise ratio of the first voice signal and obtains the formants of the first acoustic signal. Wherein, the above first acoustic signal is obtained by converting the first voice signal.

[0078] Specifically, the electronic device 100 can determine the signal-to-noise ratio of the first voice signal as follows: the electronic device 100 performs frame segmentation and windowing processing on the above first voice signal to obtain a first time-domain signal; and calculates the signal-to-noise ratio of the first voice signal according to the first time-domain signal.

[0079] The electronic device 100 can obtain the formants of the first acoustic signal as follows: the electronic device 100 performs frame segmentation and windowing processing on the first voice signal to obtain a first time-domain signal, and converts the first time-domain signal into a first acoustic signal; wherein, the first acoustic signal is a frequency-domain signal; then, the electronic device 100 performs spectral envelope analysis on the above first acoustic signal to obtain the formants of the first acoustic signal.

[0080] In some examples, the electronic device 100 can convert the first time-domain signal into a first acoustic signal as follows: perform a Fourier transform on the first time-domain signal to obtain a first frequency-domain signal; and convert the first frequency-domain signal into a first acoustic signal through a pre-trained sound pressure level (SPL) model. Wherein, the above Fourier transform can be a fast Fourier transform (FFT), or other forms of Fourier transform, and this embodiment does not limit this.

[0081] In specific implementation, the above-mentioned pre-trained SPL model includes the SPL model of the left ear and the SPL model of the right ear; thus, before converting the first frequency-domain signal into the first acoustic signal through the pre-trained SPL model, a digital signal of a predetermined frequency can also be converted into a third acoustic signal, and the above-mentioned third acoustic signal is played; then, a standard microphone is used for sound collection, and the sound pressure value of the fourth acoustic signal collected by the above-mentioned standard microphone is measured by a sound level meter, and the measured sound pressure value is used as the reference sound pressure; and a binaural sound collection is performed using an artificial head to obtain the left-ear acoustic signal and the right-ear acoustic signal respectively; finally, according to the above-mentioned digital signal of the predetermined frequency, the above-mentioned left-ear acoustic signal and the above-mentioned reference sound pressure, the SPL model of the left ear is obtained; and according to the above-mentioned digital signal of the predetermined frequency, the above-mentioned right-ear acoustic signal and the above-mentioned reference sound pressure, the SPL model of the right ear is obtained. Wherein, the above-mentioned predetermined frequency can be set by itself according to system performance and / or implementation requirements, etc. in specific implementation, and the magnitude of the above-mentioned predetermined frequency is not limited in this embodiment. For example, the above-mentioned predetermined frequency can be 1 kHz.

[0082] Specifically, referring to Figure 4 , Figure 4 is a schematic diagram of obtaining the SPL model provided by an embodiment of the present application. First, a digital-to-analog converter (DAC) can be used to convert the digital signal D_A(t) of 1 kHz into a third acoustic signal, and then, the above-mentioned third acoustic signal is amplified by a power amplifier (PA), and then the amplified acoustic signal is played through a speaker; wherein, the amplitude value of the above-mentioned digital signal D_A(t) can be -3 dBFS. Next, a standard microphone is used for sound collection, and the sound pressure value of the fourth acoustic signal collected by the above-mentioned standard microphone is measured by a sound level meter, and the measured sound pressure value is used as the reference sound pressure ref_gain; then, the artificial head is used to replace the standard microphone and the sound level meter, and the artificial head is used for binaural sound collection to obtain the left-ear acoustic signal SPL_A_L(t) and the right-ear acoustic signal SPL_A_R(t) respectively. Finally, according to the above-mentioned digital signal of the predetermined frequency, the above-mentioned left-ear acoustic signal and the above-mentioned reference sound pressure, the SPL model of the left ear is obtained, as shown in Equation (1); and according to the above-mentioned digital signal of the predetermined frequency, the above-mentioned right-ear acoustic signal and the above-mentioned reference sound pressure, the SPL model of the right ear is obtained, as shown in Equation (2).

[0083] SPL_MODEL_L = tfestimate(D_A(t), SPL_A_L(t)) × ref_gain (1)

[0084] SPL_MODEL_R = tfestimate(D_A(t), SPL_A_R(t)) × ref_gain (2)

[0085] In formula (1), SPL_MODEL_L is the SPL model of the left ear; in formula (2), SPL_MODEL_R is the SPL model of the right ear; in formula (1) and formula (2), tfestimate is a function, and the main function of the tfestimate function is to estimate the signal transfer function through the correlation of the input signal.

[0086] The following describes a process in which the electronic device 100 performs spectrum envelope analysis on the first acoustic signal to obtain the resonance peak of the first acoustic signal.

[0087] A spectrum is a collection of many different frequencies, forming a wide frequency range; different frequencies may have different amplitudes. The curve formed by connecting the highest amplitude points of different frequencies is called the spectrum envelope, and the method of calculating or obtaining the spectrum envelope is called spectrum envelope analysis.

[0088] Formants are the peaks on the envelope curve of the sound spectrum of vowels and sonorant consonants. Formants essentially refer to the resonant frequency of the vocal cavity. During the production of vowels and sonorant consonants, the sound source spectrum is modulated by the vocal cavity. The original harmonic amplitudes no longer decrease in sequence with increasing frequency, but instead strengthen and weaken, forming a new, fluctuating envelope curve. The frequency of the peak of the curve is consistent with the resonant frequency of the vocal cavity. For vowels, the first three formants are qualitatively deterministic of their timbre; the first two are particularly sensitive to the height and position of the tongue. Acoustic vowel diagrams are drawn based on the frequency values of these two formants.

[0089] In a specific implementation, the electronic device 100 can perform spectrum envelope analysis on the first acoustic signal through linear prediction coefficients (LPC) analysis to obtain the resonance peak of the first acoustic signal. The LPC analysis can be implemented using the covariance method. Specifically, the LPC analysis can be performed using the function in the matrix laboratory (matlab), which will not be described in detail here. Figure 5 , Figure 5 A schematic diagram of spectrum envelope analysis provided in one embodiment of the present application is provided. Figure 5 In FIG, the curve with more jagged edges is the FFT spectrum of the first acoustic signal, and the relatively smooth curve is the LPC spectrum, that is, the resonance peak of the first acoustic signal.

[0090] In step 303 , the electronic device 100 determines an adjustment amount of the formant of the first acoustic signal according to the signal-to-noise ratio of the first speech signal.

[0091] Specifically, after calculating the signal-to-noise ratio of the first speech signal, the electronic device 100 can determine the adjustment amount of the formant of the first acoustic signal based on the signal-to-noise ratio of the first speech signal. In this way, the clarity of the speech signal heard by the user of the electronic device 100 can be adjusted according to the current scene of the electronic device 100.

[0092] In some examples, the relationship between the signal-to-noise ratio of the first speech signal and the adjustment amount of the resonance peak of the first acoustic signal, that is, when the signal-to-noise ratio of the first speech signal is how much, the adjustment amount of the resonance peak of the first acoustic signal is how much, this relationship can be set by the application in the electronic device 100 according to the needs of the application itself, and this embodiment does not limit this.

[0093] For example, when a user uses the electronic device 100 to request a private (or encrypted) call, the electronic device 100 can be informed that this call is a private call. At this time, if the signal-to-noise ratio of the first voice signal is greater than or equal to the first threshold, this indicates that the user using the electronic device 100 is currently in a relatively quiet environment. In order to prevent the other party's words from being heard by people around the user, the electronic device 100 can determine that the adjustment amount of the resonance peak of the first acoustic signal is to move down 10dB, thereby reducing the clarity of the voice signal heard by the user, so that people around the user cannot hear the content of the user's call clearly.

[0094] When the user uses the electronic device 100 to request a normal call, the electronic device 100 learns that this call is a normal call, not a private (or encrypted) call. At this time, if the signal-to-noise ratio of the first voice signal is less than or equal to the first threshold, this indicates that the user using the electronic device 100 is currently in an environment with relatively high noise levels. Therefore, it can be determined that the adjustment amount of the resonance peak of the first acoustic signal is increased by 10 dB, thereby improving the clarity of the voice signal heard by the user and allowing the user to communicate smoothly with the other party on the call. If the signal-to-noise ratio of the first voice signal is greater than the first threshold and less than or equal to the second threshold, this indicates that the user using the electronic device 100 is currently in an environment with relatively low noise levels. Therefore, it can be determined that the adjustment amount of the resonance peak of the first acoustic signal is increased by 5 dB.

[0095] Among them, the sizes of the first threshold and the second threshold can be set by yourself during specific implementation according to system performance and / or implementation requirements. This embodiment does not limit the sizes of the above-mentioned first threshold and the second threshold. For example, the above-mentioned first threshold can be 5dB and the second threshold can be 10dB.

[0096] It should be noted that when the user makes a call using the electronic device 100, the user can use the phone application in the electronic device 100 or the instant messaging application installed in the electronic device 100. This embodiment does not limit this. In addition, the above call can be a voice call or a video call.

[0097] The above are only some examples of determining the adjustment amount of the formant of the first acoustic signal according to the signal-to-noise ratio of the first voice signal. This embodiment is not limited thereto. In actual use, the application in the electronic device 100 sets the adjustment scheme of the formant of the first acoustic signal according to the requirements of the application itself. This embodiment does not limit this.

[0098] Step 304, the electronic device 100 generates a second acoustic signal according to the above adjustment amount, the upper and lower bounds of the formant of the first acoustic signal, the fundamental frequency of the first acoustic signal, and the signal spectrum corresponding to the second voice signal.

[0099] Among them, the above second voice signal is the voice signal received by the electronic device 100 after the first voice signal. Specifically, the second voice signal is the voice signal received by the electronic device 100 from the opposite end after the first voice signal. In terms of time, the second voice signal and the first voice signal are continuous voice signals. It should be noted that in this embodiment, the second voice signal has not been played by the speaker 170A.

[0100] In some examples, the electronic device 100 generates a second acoustic signal according to the above adjustment amount, the upper and lower bounds of the formant of the first acoustic signal, the fundamental frequency of the first acoustic signal, and the signal spectrum corresponding to the second voice signal, which can be: the electronic device 100 adjusts the peak value of the signal spectrum corresponding to the second voice signal according to the above adjustment amount based on the fundamental frequency of the first acoustic signal to generate a second acoustic signal; among them, the upper and lower bounds of the formant of the second acoustic signal are parallel to the upper and lower bounds of the formant of the first acoustic signal respectively; the difference between the upper bound of the formant of the second acoustic signal and the upper bound of the formant of the first acoustic signal is the above adjustment amount, and the difference between the lower bound of the formant of the second acoustic signal and the lower bound of the formant of the first acoustic signal is the above adjustment amount.

[0101] In a specific implementation, a phase vocoder (PV) can be used to generate the second acoustic signal. PV is a mature application of short-time Fourier transform (STFT) time-frequency analysis. It can be understood as using a fixed filter bank to extract the frequency response of a time-varying signal, and then using these parameters to control a set of sinusoidal oscillations to complete the synthesis of the target new signal. The principle of PV can be shown in Figure 6(a), which is a schematic diagram of the principle of PV provided in one embodiment of the present application.

[0102] In Figure 6(a), h(n) is a fixed filter bank that can be implemented using an impulse function; the input signal may include an adjustment amount, the upper and lower bounds of the resonance peak of the first acoustic signal, the fundamental frequency of the first acoustic signal, and the signal spectrum corresponding to the first speech signal; the output signal is the second acoustic signal.

[0103] In other examples, before generating the second acoustic signal, the electronic device 100 may also determine the upper and lower bounds of the resonance peak of the first acoustic signal by piecewise interpolation according to the frequency and spectrum amplitude of the resonance peak of the first acoustic signal. This process is Figure 2 In the specific implementation, the upper bound and the lower bound of the resonance peak of the first acoustic signal can be determined by formula (3).

[0104] [up, down] = envelope(lpc_x, lpc_y, 'pchip') (3)

[0105] In formula (3), up represents the upper bound of the resonance peak of the first acoustic signal, down represents the lower bound of the resonance peak of the first acoustic signal, envelope represents the envelope of the curve, lpc_x represents the frequency of the resonance peak of the first acoustic signal, lpc_y represents the spectral amplitude of the resonance peak of the first acoustic signal, and pchip represents piecewise cubic Hermitian interpolation. In this embodiment, the proportional relationship between the various spectral lines can be determined based on the upper and lower bounds of the resonance peak of the first acoustic signal. Since the proportional relationship between the spectral lines determines the timbre, when PV generates the second acoustic signal, it is necessary to control the corresponding adjustment amount, that is, the upper and lower bounds of the resonance peak after adjustment are as parallel as possible to those before adjustment, so that the timbre of the generated second acoustic signal remains unchanged. See Figure 6(b), which is a schematic diagram of the upper and lower bounds of the resonance peak provided by an embodiment of the present application. As shown in Figure 6(b), if the adjustment amount of the resonance peak of the first acoustic signal is increased by 5dB, then after the adjustment, the upper bound of the resonance peak of the second acoustic signal is parallel to the upper bound of the resonance peak of the first acoustic signal. Similarly, the lower bound of the resonance peak of the second acoustic signal is also parallel to the lower bound of the resonance peak of the first acoustic signal.

[0106] In some other examples, before generating the second acoustic signal, the electronic device 100 may also perform pitch analysis on the first acoustic signal to obtain the pitch frequency of the first acoustic signal. Pitch analysis is a commonly used technique in the field of speech signal processing. In this embodiment, the short-time autocorrelation function method may be used to perform pitch analysis, as shown by the white curve in Figure 7 . Figure 7 FIG. is a schematic diagram of pitch analysis provided by an embodiment of the present application. The autocorrelation function method is a method based on signal processing for detecting and analyzing the fundamental frequency of periodic signals. In audio signal processing, the autocorrelation function can be used to detect and analyze the fundamental frequency of audio signals, that is, the pitch in music.

[0107] Step 305, the electronic device 100 obtains a third voice signal to be played according to the second acoustic signal.

[0108] Referring to Figure 2 , after generating the second acoustic signal, the electronic device 100 may perform an inverse Fourier transform on the second acoustic signal, convert the time-domain signal obtained by the inverse Fourier transform into a voice signal through a DAC, and then amplify the converted voice signal through a PA to obtain a third language signal to be played. As shown in Figure 2 , the above inverse Fourier transform may be an inverse fast Fourier transform (iFFT).

[0109] Step 306, the electronic device 100 plays the third voice signal.

[0110] Referring to Figure 2 , the electronic device 100 may play the above third voice signal through the speaker 170A.

[0111] In the above-mentioned method for controlling speech clarity, after the electronic device 100 obtains a first speech signal to be processed, the electronic device 100 performs frame and window processing on the first speech signal to obtain a first time domain signal; and calculates the signal-to-noise ratio of the first speech signal based on the first time domain signal, converts the first time domain signal into a first acoustic signal, performs spectrum envelope analysis on the first acoustic signal to obtain the formant of the first acoustic signal, and then determines an adjustment amount for the formant of the first acoustic signal based on the signal-to-noise ratio of the first speech signal. Finally, the electronic device 100 can generate a second acoustic signal based on the adjustment amount, the upper and lower bounds of the formant of the first acoustic signal, the fundamental frequency of the first acoustic signal, and the signal spectrum corresponding to the second speech signal, obtain a third speech signal to be played based on the second acoustic signal, and play the third speech signal. Since the adjustment amount of the formant of the first acoustic signal is determined based on the signal-to-noise ratio of the first speech signal, it is possible to adjust the clarity of the speech signal heard by the user using the electronic device 100 according to the current scene in which the electronic device 100 is located.

[0112] Figure 8 A flowchart of a method for controlling speech clarity provided in another embodiment of the present application is shown in FIG. Figure 8 As shown, this application Figure 3 In the illustrated embodiment, before step 304, the following steps may also be included:

[0113] In step 801 , the electronic device 100 performs time-frequency domain masking analysis on a second speech signal to obtain a signal spectrum corresponding to the second speech signal.

[0114] The signal spectrum corresponding to the second speech signal may be a signal spectrum of the second speech signal actually perceived by the human ear.

[0115] Time-frequency domain masking is an important analytical method in psychoacoustics. Based on the theory of human auditory perception, the original signal is obtained from the time domain and frequency domain respectively. The signal actually perceived by the human ear after hearing the original signal can then be better analyzed and processed using other algorithms.

[0116] Specifically, see Figure 9 The electronic device 100 performs time-frequency domain masking analysis on the second voice signal to obtain a signal spectrum corresponding to the second voice signal. This can be done by: performing STFT on the second voice signal, convolving the signal obtained by STFT with the masking mask, and performing inverse short-time Fourier transform (inverse STFT, iSTFT) on the signal obtained by convolution to obtain a signal spectrum corresponding to the second voice signal. Figure 9 A schematic diagram of implementing time-frequency domain masking analysis provided in one embodiment of the present application.

[0117] Figure 10Schematic diagram of a specific example of time-frequency domain masking analysis provided by an embodiment of this application. Figure 10 In it, from top to bottom are the original signal, the masking mask, and the masked signal. Combining Figure 9 , Figure 10 After the original signal in Figure 10 is subjected to STFT, the signal obtained by STFT is convolved with the masking mask in

[0118] Then, iSTFT is performed on the signal obtained by convolution, so as to obtain the true signal spectrum perceived by the human ear after being masked.

[0119]

[0120] In formula (4), maskThresh(x) is the masking threshold, x is the signal-to-noise ratio of the second voice signal, and a, b, and c are model coefficients respectively. In this embodiment, the magnitudes of a, b, and c are not limited.

[0121] It can be understood that some or all of the steps or operations in the above embodiments are only examples. Embodiments of this application can also perform other operations or various deformations of the operations. In addition, each step can be executed in a different order presented in the above embodiments, and it is possible not to execute all the operations in the above embodiments.

[0122] It can be understood that in order for the electronic device to implement the above functions, it includes the corresponding hardware and / or software modules for executing each function. Combining the algorithm steps of each example described in the embodiments disclosed in this application, this application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application in combination with the embodiments, but such implementation should not be considered to exceed the scope of this application.

[0123] This embodiment can divide the functional modules of the electronic device according to the above method embodiments. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one module. The above integrated module can be implemented in the form of hardware. It should be noted that the division of modules in this embodiment is illustrative, only a logical function division, and there can be other division methods in actual implementation.

[0124] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. When the computer program runs on a computer, it causes the computer to execute the present application Figures 3 to 10 The method provided by the illustrated embodiment.

[0125] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program runs on a computer, it causes the computer to execute the present application Figures 3 to 10 The method provided by the illustrated embodiment.

[0126] In the embodiments of the present application, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may represent the case where A exists alone, A and B exist simultaneously, or B exists alone. Where A and B may be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one of the following" and its similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, and c may represent: a, b, c, a and b, a and c, b and c, or a and b and c, where a, b, and c may be single or multiple.

[0127] Those of ordinary skill in the art can realize that the various units and algorithm steps described in the embodiments disclosed herein can be implemented by a combination of electronic hardware, computer software, and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0128] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be described herein again.

[0129] In several embodiments provided in the present application, any function can be stored in a computer-readable storage medium if it is implemented in the form of a software functional unit and sold or used as an independent product. Based on this understanding, the technical solution of the present application can be embodied in the form of a software product in essence or in other words, the part that contributes to the prior art or the part of the technical solution. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0130] The above is only a specific implementation of the present application. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. The protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A method for controlling speech intelligibility, characterized in that Including: Obtaining a first voice signal to be processed; Determining the signal-to-noise ratio of the first voice signal and obtaining the formants of a first acoustic signal; wherein, the first acoustic signal is obtained by converting the first voice signal; Determining an adjustment amount of the formants of the first acoustic signal according to the signal-to-noise ratio of the first voice signal; Generating a second acoustic signal according to the adjustment amount, the upper and lower bounds of the formants of the first acoustic signal, the fundamental frequency of the first acoustic signal, and the signal spectrum corresponding to a second voice signal; wherein, the second voice signal is a voice signal received by an electronic device after the first voice signal; Obtaining a third voice signal to be played according to the second acoustic signal; Playing the third voice signal.

2. The method according to claim 1, characterized in that The determining the signal-to-noise ratio of the first voice signal includes: Performing frame segmentation and windowing processing on the first voice signal to obtain a first time-domain signal; Calculating the signal-to-noise ratio of the first voice signal according to the first time-domain signal.

3. The method according to claim 1, characterized in that The obtaining the formants of the first acoustic signal includes: Performing frame segmentation and windowing processing on the first voice signal to obtain a first time-domain signal; Converting the first time-domain signal into a first acoustic signal; wherein, the first acoustic signal is a frequency-domain signal; Performing spectral envelope analysis on the first acoustic signal to obtain the formants of the first acoustic signal.

4. The method according to claim 1, characterized in that: The generating the second acoustic signal according to the adjustment amount, the upper and lower bounds of the formants of the first acoustic signal, the fundamental frequency of the first acoustic signal, and the signal spectrum corresponding to the second voice signal includes: Based on the fundamental frequency of the first acoustic signal, adjusting the peak value of the signal spectrum corresponding to the second voice signal according to the adjustment amount to generate a second acoustic signal; Wherein, the upper and lower bounds of the formants of the second acoustic signal are respectively parallel to the upper and lower bounds of the formants of the first acoustic signal; the difference between the upper bound of the formants of the second acoustic signal and the upper bound of the formants of the first acoustic signal is the adjustment amount, and the difference between the lower bound of the formants of the second acoustic signal and the lower bound of the formants of the first acoustic signal is the adjustment amount.

5. The method according to claim 3, wherein The converting the first time-domain signal into a first acoustic signal includes: Performing a Fourier transform on the first time-domain signal to obtain a first frequency-domain signal; Converting the first frequency-domain signal into a first acoustic signal through a pre-trained sound pressure level model.

6. The method according to claim 5, wherein The pre-trained sound pressure level model includes a left-ear sound pressure level model and a right-ear sound pressure level model; before converting the first frequency-domain signal into a first acoustic signal through the pre-trained sound pressure level model, further including: Converting a digital signal of a predetermined frequency into a third acoustic signal; Playing the third acoustic signal; Using a standard microphone to receive sound, measuring the sound pressure value of a fourth acoustic signal collected by the standard microphone through a sound level meter, and using the measured sound pressure value as a reference sound pressure; and using an artificial head for binaural sound reception to respectively obtain a left-ear acoustic signal and a right-ear acoustic signal; Obtain a sound pressure level model for the left ear based on the digital signal of the predetermined frequency, the left ear acoustic signal, and the reference sound pressure; and obtain a sound pressure level model for the right ear based on the digital signal of the predetermined frequency, the right ear acoustic signal, and the reference sound pressure.

7. The method according to claim 1, characterized in that Before generating the second acoustic signal, it further includes: Perform pitch analysis on the first acoustic signal to obtain the pitch frequency of the first acoustic signal.

8. The method according to claim 1, characterized in that, Before generating the second acoustic signal, it further includes: Perform time-frequency domain masking analysis on the second speech signal to obtain the signal spectrum corresponding to the second speech signal; wherein, the signal spectrum corresponding to the second speech signal includes the signal spectrum actually perceived by the human ear for the second speech signal.

9. The method according to claim 8, characterized in that The performing time-frequency domain masking analysis on the second speech signal to obtain the signal spectrum corresponding to the second speech signal includes: Perform short-time Fourier transform on the second speech signal; Convolve the signal obtained by the short-time Fourier transform with the masking mask; Perform inverse short-time Fourier transform on the signal obtained by the convolution to obtain the signal spectrum corresponding to the second speech signal.

10. The method according to claim 9, wherein Before convolving the signal obtained by the short-time Fourier transform with the masking mask, it further includes: Calculate the signal-to-noise ratio of the second speech signal; Determine the masking threshold according to the signal-to-noise ratio of the second speech signal; Generate the masking mask according to the masking threshold.

11. The method according to claim 1, characterized in that, Before generating the second acoustic signal, it further includes: Determine the upper and lower bounds of the formants of the first acoustic signal by piecewise interpolation according to the frequencies and spectral amplitudes of the formants of the first acoustic signal.

12. An electronic device, characterized in that, It includes: One or more processors; A memory; Multiple applications; And one or more computer programs, wherein the one or more computer programs are stored in the memory, and the one or more computer programs include instructions that, when executed by the electronic device, cause the electronic device to execute the method according to any one of claims 1-11.

13. A computer-readable storage medium, characterized in that A computer program is stored in the computer-readable storage medium, and when it runs on a computer, it causes the computer to execute the method according to any one of claims 1-11.