Audio generation method and related apparatus

By receiving and processing microphone signals in electronic devices and generating target audio files according to set modes, the problems of audio signal processing clarity and timbre adjustment are solved, thus improving the user's voice interaction experience.

CN120431945BActive Publication Date: 2026-05-29HONOR DEVICE CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HONOR DEVICE CO LTD
Filing Date
2024-11-29
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

In existing technologies, there are no effective solutions to problems such as how to process the audio signal received by the microphone according to user needs, such as filling in missing words, generating target audio files with specified timbres, and improving the clarity of voice signals.

Method used

The electronic device receives the audio signal to be processed through the microphone, processes the timbre into the target timbre according to the preset processing mode, generates a noise-free speech signal, and generates the target audio file based on the noise-free speech signal and the environmental noise signal or the preset noise signal.

Benefits of technology

It enables flexible adjustment of the timbre of the user's voice signal, improves the clarity of the voice signal, and enhances the user's voice interaction experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120431945B_ABST
    Figure CN120431945B_ABST
Patent Text Reader

Abstract

The application provides an audio generation method and related device, and relates to the terminal field, and the method comprises the following steps: when an electronic device receives a user input audio signal 1 to be processed through a microphone, the electronic device can process the tone of the audio signal 1 to be processed into a target tone (for example, the tone of the microphone when it is currently receiving sound / the preset tone set by the user in advance) according to the processing mode set by the user in advance, and obtain a noise-free voice signal. Wherein, the audio signal 1 to be processed received by the microphone can include a voice signal 1 emitted by the user when speaking and an environmental noise signal 1 existing in the periphery of the electronic device except the voice signal. Then, the electronic device can generate a target audio file based on the noise-free voice signal and the environmental noise signal 1 (or a preset noise signal) extracted from the audio signal 1 to be processed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of terminals, and more particularly to an audio generation method and related apparatus. Background Technology

[0002] With the development of terminal technology, users are using electronic devices to handle daily tasks more and more frequently. Generally, electronic devices can be equipped with one or more microphones, allowing users to interact with the device via voice. For example, users can use the microphone to make calls, record audio, and send voice messages. Therefore, how to process the audio signals received by the microphone according to user needs has become a pressing issue. Summary of the Invention

[0003] This application provides an audio generation method and related apparatus, relating to the field of terminals, which not only allows for flexible adjustment of the timbre of the user's voice signal and improvement of the clarity of the voice signal, but also enhances the user's voice interaction experience.

[0004] In a first aspect, this application provides an audio processing method, comprising: an electronic device receiving a first audio signal via a microphone. The first audio signal includes a user's speech signal and a first ambient noise signal. The electronic device identifies text to be synthesized from the first audio signal. When the electronic device is in a first mode, it generates a first noise-free speech signal based on the text to be synthesized. The timbre of the first noise-free speech signal is the timbre of the speech signal. When the electronic device determines that a preset noise signal should not be used, it extracts the first ambient noise signal from the first audio signal. The electronic device generates a first target audio file based on the first noise-free speech signal and the first ambient noise signal. When the electronic device determines that a preset noise signal should be used, it generates a second target audio file based on the first noise-free speech signal and the preset noise signal.

[0005] In one possible implementation, the method further includes: when the electronic device is in a second mode, the electronic device generates a second noiseless speech signal based on the text to be synthesized. The timbre of the second noiseless speech signal is a preset timbre. When the electronic device determines that a preset noise signal is not to be used, the electronic device extracts a first ambient noise signal from the first audio signal. The electronic device generates a third target audio file based on the second noiseless speech signal and the first ambient noise signal. When the electronic device determines that a preset noise signal is to be used, the electronic device generates a fourth target audio file based on the second noiseless speech signal and the preset noise signal.

[0006] In one possible implementation, the method further includes: the electronic device acquiring first text information; the electronic device generating a third noise-free speech signal based on the first text information; wherein the timbre of the third noise-free speech signal is a preset timbre; when the electronic device determines that it will not use the preset noise signal, the electronic device acquiring a second ambient noise signal through the microphone; the electronic device generating a fifth target audio file based on the third noise-free speech signal and the second ambient noise signal; and when the electronic device determines that it will use the preset noise signal, the electronic device generating a sixth target audio file based on the third noise-free speech signal and the preset noise signal.

[0007] In one possible implementation, the method further includes: the electronic device acquiring first text information; the electronic device acquiring a second audio signal via the microphone; the electronic device extracting a first acoustic feature from the second audio signal; the electronic device generating a fourth noise-free speech signal based on the first text information and the first acoustic feature, wherein the timbre of the fourth noise-free speech signal is the user's current timbre; when the electronic device determines not to use a preset noise signal, the electronic device extracting a third ambient noise signal from the second audio signal; the electronic device generating a seventh target audio file based on the fourth noise-free speech signal and the third ambient noise signal; and when the electronic device determines to use a preset noise signal, the electronic device generating an eighth target audio file based on the fourth noise-free speech signal and the preset noise signal.

[0008] In one possible implementation, when the electronic device is in a first mode, it generates a first noiseless speech signal based on the text to be synthesized. Specifically, this includes: the electronic device extracting a second acoustic feature from the first audio signal. The second acoustic feature is the acoustic feature of the speech signal. The electronic device generates the first noiseless speech signal based on the first acoustic feature and the text to be synthesized.

[0009] In one possible implementation, the preset tone is the tone that the user is in a specified state.

[0010] In one possible implementation, the electronic device generates a first noiseless speech signal based on the second acoustic feature and the text to be synthesized, specifically including: the electronic device generates the first noiseless speech signal based on the second acoustic feature and the text to be synthesized through an audio generation model.

[0011] In one possible implementation, when the electronic device is in the second mode, it generates a second noiseless speech signal based on the text to be synthesized. Specifically, when the electronic device is in the second mode, it generates the second noiseless speech signal based on the text to be synthesized using a preset timbre generation model. The preset timbre generation model includes a third acoustic feature, which is an acoustic feature corresponding to the preset timbre.

[0012] In a second aspect, this application provides an electronic device comprising: one or more processors and a memory. The memory is coupled to the one or more processors and is used to store computer program code, the computer program code including computer instructions, which the one or more processors invoke to cause the electronic device to perform a method as described in any of the possible implementations of the first aspect above.

[0013] Thirdly, this application provides a chip system applied to an electronic device, the chip system including one or more processors, the processors being configured to invoke computer instructions to cause the electronic device to perform a method as described in any of the possible implementations of the first aspect above.

[0014] Fourthly, this application provides a computer-readable storage medium including instructions that, when executed on an electronic device, cause the electronic device to perform a method as described in any of the possible implementations of the first aspect above.

[0015] Fifthly, this application provides a computer program product, including a computer program that, when executed by a processor, causes the electronic device to perform the method as described in any of the possible implementations of the first aspect above. Attached Figure Description

[0016] Figure 1 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application;

[0017] Figures 2A-2S A set of user interface diagrams provided for embodiments of this application;

[0018] Figure 3A This is a schematic flowchart illustrating a specific audio processing method provided in an embodiment of this application.

[0019] Figure 3B A schematic diagram of an apparatus framework for an audio processing method provided in an embodiment of this application;

[0020] Figure 4A A detailed flowchart illustrating another audio processing method provided in this application embodiment;

[0021] Figure 4B A schematic diagram of the apparatus framework for another audio processing method provided in an embodiment of this application;

[0022] Figure 5A A detailed flowchart illustrating another audio processing method provided in this application embodiment;

[0023] Figure 5B This is a schematic diagram of the device framework for another audio processing method provided in the embodiments of this application;

[0024] Figure 6A A detailed flowchart illustrating another audio processing method provided in this application embodiment;

[0025] Figure 6B This is a schematic diagram of the apparatus framework for another audio processing method provided in the embodiments of this application. Detailed Implementation

[0026] The technical solutions in the embodiments of this application will be clearly and thoroughly described below with reference to the accompanying drawings. In the description of the embodiments of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B; the word "and / or" in the text is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Furthermore, in the description of the embodiments of this application, "multiple" refers to two or more than two.

[0027] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as implying or suggesting relative importance or implicitly indicating the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature, and in the description of the embodiments of this application, unless otherwise stated, "multiple" means two or more.

[0028] Figure 1 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application.

[0029] In the embodiments of this application, Figure 1 The electronic device 100 shown is the electronic device in the embodiments of this application.

[0030] Electronic device 100 may be a mobile phone, tablet computer, desktop computer, laptop computer, handheld computer, notebook computer, ultra-mobile personal computer (UMPC), netbook, cellular phone, personal digital assistant (PDA), augmented reality (AR) device, virtual reality (VR) device, artificial intelligence (AI) device, wearable device, in-vehicle device, smart home device and / or smart city device. The embodiments of this application do not impose any special restrictions on the specific type of electronic device 100.

[0031] like Figure 1 As shown, the electronic device 100 may include a processor 101, a memory 102, a wireless communication module 103 (optional), a display screen 104, a sensor module 105, an audio module 106, a speaker 107, and a microphone 108.

[0032] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may also include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0033] Processor 101 may include one or more processor units, such as an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural network processing unit (NPU). Different processing units may be independent devices or integrated into one or more processors.

[0034] The controller can generate operation control signals based on the instruction opcode and timing signals to complete the control of instruction fetching and execution.

[0035] The processor 101 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 101 is a cache memory. This memory can store instructions or data that the processor 101 has just used or that are used repeatedly. If the processor 101 needs to use the instruction or data again, it can directly retrieve it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 101, and thus improves the efficiency of the system.

[0036] In some embodiments, the processor 101 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a USB interface, etc.

[0037] The memory 102 is coupled to the processor 101 and is used to store various software programs and / or multiple sets of instructions. In specific implementations, the memory 102 may include volatile memory, such as random access memory (RAM); it may also include non-volatile memory, such as ROM, flash memory, hard disk drive (HDD), or solid-state drive (SSD); the memory 102 may also include combinations of the above types of memory. The memory 102 may also store some program code so that the processor 101 can call the program code stored in the memory 102 to implement the implementation method of the present application embodiment in the electronic device 100. The memory 102 may store an operating system, such as uCOS, VxWorks, RTLinux, or other embedded operating systems.

[0038] The wireless communication module 103 can provide solutions for wireless communication applications on the electronic device 100, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies. The wireless communication module 103 can be one or more devices integrating at least one communication processing module. The wireless communication module 103 receives electromagnetic waves via an antenna, performs frequency modulation and filtering of the electromagnetic wave signals, and sends the processed signal to the processor 101. The wireless communication module 103 can also receive signals to be transmitted from the processor 101, perform frequency modulation and amplification, and convert them into electromagnetic waves for radiation via the antenna. In some embodiments, the electronic device 100 can also communicate via the Bluetooth module in the wireless communication module 103 (… Figure 1 (not shown), WLAN module ( Figure 1 (Not shown) The device transmits signals to detect or scan devices near electronic device 100 and establishes wireless communication connections with those devices to transmit data. The Bluetooth module can provide solutions for one or more Bluetooth communication methods, including basic rate / enhanced data rate (BR / EDR) or Bluetooth Low Energy (BLE), and the WLAN module can provide solutions for one or more WLAN communication methods, including Wi-Fi direct, Wi-Fi LAN, or Wi-Fi softAP.

[0039] The display screen 104 can be used to display images, videos, etc. The display screen 104 may include a display panel. The display panel may be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Mini LED, a MicroLED, a Micro-OLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device 100 may include one or N display screens 104, where N is a positive integer greater than 1.

[0040] The sensor module 105 may include a touch sensor 105A. The touch sensor 105A may also be referred to as a "touch device". The touch sensor 105A may be disposed on the display screen 104, and the touch sensor 105A and the display screen 104 together form a touch screen, also known as a "touchscreen". The touch sensor 105A can be used to detect touch operations applied to or near it.

[0041] The audio module 106 can be used to convert digital audio information into analog audio signal output, and can also be used to convert analog audio input into digital audio signal. The audio module 106 can also be used to encode and decode audio signals. In some embodiments, the audio module 106 can also be disposed in the processor 101, or some functional modules of the audio module 106 can be disposed in the processor 101.

[0042] The speaker 107, also known as a "loudspeaker," is used to convert audio electrical signals into sound signals. The electronic device 100 can listen to music or make hands-free calls through the speaker 107.

[0043] Microphone 108, also known as a "microphone" or "voice transducer," is used to collect sound signals from the environment surrounding the electronic device. It then converts these sound signals into electrical signals, processes them (e.g., analog-to-digital conversion), and obtains a digital audio signal that can be processed by the processor 101 of the electronic device. When making a phone call or sending a voice message, the user can speak by bringing their mouth close to the microphone 108, inputting the sound signal into the microphone 108. The electronic device 100 may have at least one microphone 108. In some embodiments, the electronic device 100 may have two microphones 108, which, in addition to collecting sound signals, can also perform noise reduction. In other embodiments, the electronic device 100 may have three or more microphones 108, enabling sound signal collection, noise reduction, sound source identification, and directional recording, among other functions.

[0044] It should be noted that, Figure 1 The electronic device 100 shown is merely an illustrative explanation of the hardware structure of the electronic device provided in this application and does not constitute a specific limitation on this application.

[0045] Generally, electronic device 100 can be equipped with one or more microphones, allowing users to interact with the device via voice. For example, users can use the microphones on electronic device 100 for calls, voice recording, and sending voice messages. Therefore, how to process the audio signal received by the microphone according to user needs—for example, by completing missing words in the audio signal, generating a target audio file with a specified timbre based on the received audio signal, and improving the clarity of the user's voice signal—has become a pressing issue.

[0046] Therefore, this application provides an audio generation method. In this method, when an electronic device 100 receives a user-input audio signal 1 to be processed via a microphone, the electronic device 100 can process the timbre in the audio signal 1 to a target timbre (e.g., the timbre of the microphone at the current time of recording (also known as the timbre of the user in the current time and space, i.e., the user's current timbre) / the user's preset timbre) according to a user-preset processing mode, thereby obtaining a noise-free speech signal. The audio signal 1 to be processed received by the microphone may include: the speech signal 1 emitted by the user when speaking and environmental noise signal 1 existing around the electronic device 100, excluding the speech signal. Then, the electronic device 100 can generate a target audio file based on the noise-free speech signal and the environmental noise signal 1 (or preset noise signal) extracted from the audio signal 1 to be processed. This not only allows for flexible adjustment of the timbre of the user's speech signal and improves the clarity of the speech signal, but also enhances the user's voice interaction experience.

[0047] In this embodiment of the application, timbre generally refers to the timbre of the speech signal emitted by the user when speaking. The timbre of the audio signal 1 to be processed is also the timbre of the speech signal 1.

[0048] Figures 2A-2S A set of user interface diagrams provided for embodiments of this application.

[0049] like Figure 2A As shown, the electronic device 100 can display a desktop 210. This desktop 210 can display a page containing application icons, including icons for multiple applications (e.g., weather app icon, stock app icon, calculator app icon, settings app icon 211, email app icon, theme app icon, calendar app icon, video app icon, etc.). A status bar is displayed in the upper portion of the desktop 210. This status bar can include one or more indicators, such as one or more mobile communication signal (also known as cellular signal) strength indicators, battery status indicators, time indicators, Wi-Fi signal indicators, etc.

[0050] The electronic device 100 can receive a touch operation (e.g., a tap) on the icon 211 of the settings application. In response to the touch operation, the electronic device 100 can display the settings interface 220.

[0051] like Figure 2B As shown, the settings interface 220 can display one or more page options, such as application page options, battery page options, storage page options, security page options, privacy page options, health usage page options, smart assistant page options 221, and accessibility page options, etc.

[0052] Electronic device 100 can receive a touch operation (e.g., a click) on the smart assistant page option 221. In response to the touch operation, electronic device 100 can display the smart assistant page 230.

[0053] like Figure 2C As shown, the smart assistant page 230 may include one or more function options, such as personalized voice function option 231, auxiliary vision function option, smart screen recognition function option, contextual intelligence function option, smart search function option, and YOYO suggestion function option, etc.

[0054] Electronic device 100 can receive a touch operation (e.g., a click) on the personalized voice function option 231. In response to the touch operation, electronic device 100 can display the personalized voice page 240.

[0055] like Figure 2DAs shown, the personalized voice page 240 may include one or more functional options, such as a "Maintain Current Tone" option, including the text prompt "When turned on, the generated voice will retain the tone from when it was picked up" and a control 241. The control 241 can receive touch operations (e.g., clicks) applied to it. In response to the touch operation, when the indicator on the control 241 changes from "off" to "on", the electronic device 100 can maintain the tone of the audio signal 1 to be processed received by the microphone as the tone from when the microphone picked up the sound (i.e., the user's current tone) and improve the clarity of the voice signal 1.

[0056] The "Use Preset Tone" option includes the text prompt "When enabled, the generated speech will closely resemble the user-set preset tone" and control 242. This control 242 can receive touch operations (e.g., clicks) applied to it. In response to this touch operation, the electronic device 100 can process the tone of the audio signal 1 received by the microphone to be processed into the user-set preset tone and improve the clarity of the speech signal 1. The preset tone can be the tone of the user in a specific state (e.g., a sick state / normal physiological state), and this application does not limit the specific type of preset tone.

[0057] The "Use Preset Noise" option includes the text message "When turned on, an audio file will be generated based on the preset noise" and control 243. Control 243 can receive touch operations (e.g., clicks) applied to it. In response to this touch operation, when the indicator on control 243 changes from "off" to "on", the electronic device 100 can use the selected preset noise signal superimposed on the noise-free speech signal to generate the target audio file.

[0058] The "Update your preset tone" option includes the text prompt "Clicking this will re-encode your tone and update your preset tone" and control 244. This control 244 can receive touch operations (e.g., clicks) applied to it. In response to this touch operation, the electronic device 100 can receive the user-input training audio signal 1 via the microphone in the voice recording interface (e.g., the voice recording interface 250 described later), and / or, the electronic device 100 can acquire the specified application (e.g., gallery application, recorder application, etc.). The electronic device 100 uses training voice message 1 and / or training audio file 1 in the specified application to train a preset timbre generation model. The electronic device 100 can process the audio signal 1 to be processed into a preset timbre based on the preset timbre generation model, thereby improving the clarity of the voice signal. The training audio signal 1, training voice message 1, and training audio file 1 include: the voice signal of a user speaking in a specified state (e.g., a sick state / normal physiological state, etc.) and the surrounding environmental noise signal.

[0059] like Figure 2E As shown, the electronic device 100 can receive a touch operation (e.g., a click) applied to the control 242. In response to the touch operation, the indicator on the control 242 is displayed as "on," and the electronic device 100 can process the timbre of the audio signal 1 to be processed to a preset timbre. In one possible implementation, the electronic device 100 can receive a touch operation (e.g., a click) applied to the control 241. In response to the touch operation, the indicator on the control 241 is displayed as "on," and the electronic device 100 can maintain the timbre of the audio signal 1 to be processed as the timbre picked up by the microphone.

[0060] In one possible implementation, if the electronic device 100 first receives a touch operation on the control 242 and the identifier on the control 242 is displayed as "on", or if the electronic device 100 receives a touch operation on the control 244 to update the preset tone, the electronic device 100 may display the voice recording interface 250 in response to the aforementioned touch operation.

[0061] like Figure 2F As shown, the voice recording interface 250 may include the text prompt "Please record three or more voice messages", control 251, and one or more application icons (e.g., gallery application icon, recorder application icon 252, etc.). App icon Application icons, etc., and controls 253, etc. Control 251 can be used to receive touch operations (e.g., clicks) applied to it. In response to the touch operation, the electronic device 100 can activate the microphone and receive the training audio signal 1 through the microphone in the voice recording interface 250. An application icon of a specific application can be used to receive touch operations (e.g., clicks) applied to it. In response to the touch operation, the electronic device 100 can display the application interface of the specified application, allowing the user to select the training voice message 1 and / or training audio file 1 in the specified application. Control 253 can be used to receive touch operations (e.g., clicks) applied to it. In response to the touch operation, the electronic device 100 can train a preset timbre generation model based on the training audio signal 1, and / or training voice message 1, and / or training audio file 1, etc.

[0062] like Figure 2G As shown, when the electronic device 100 acquires training audio signal 1, and / or training voice message 1, and / or training audio file 1, and receives a touch operation applied to control 253, in response to the touch operation, the electronic device 100 can train a preset timbre generation model based on training audio signal 1, and / or training voice message 1, and / or training audio file 1. At this time, as... Figure 2H As shown, the electronic device 100 can display the text prompt "Generating your preset tone..." on the voice recording interface 250.

[0063] like Figure 2I As shown, when the preset timbre generation model is trained, the electronic device 100 can display the text prompt "Your preset timbre has been generated" on the voice recording interface 250. Then, the electronic device can process the timbre of the audio signal to be processed into the preset timbre based on the preset timbre generation model.

[0064] In one possible implementation, the electronic device 100 can receive a touch operation (e.g., click) on an application icon in the voice recording interface. In response to the touch operation, the electronic device 100 can display the application interface of the specified application, allowing the user to select training voice message 1 and / or training audio file 1 in the specified application to train a preset timbre generation model.

[0065] like Figure 2J As shown, the electronic device 100 can receive a touch operation on the recorder application icon 252 on the voice recording interface 250. In response to the touch operation, the electronic device 100 can display the recorder application interface 260.

[0066] like Figure 2KAs shown, the recorder application interface 260 may include one or more training audio files 1 (e.g., training audio file 261, etc.) and corresponding selection controls (e.g., selection control 262 corresponding to training audio file 261, etc.). When the selection control displays a black dot, it indicates that the training audio file 1 corresponding to the selection control has been selected. When the electronic device 100 selects a training audio file 1 and receives a touch operation on the selection control 263, the electronic device 100 can use the specified training audio file 1 to train a preset timbre generation model.

[0067] like Figure 2L As shown, the electronic device 100 can receive a touch operation applied to the selected control 262. In response to the touch operation, the selected control 262 can display a black dot indicator to indicate that the training audio file 261 has been selected.

[0068] Then, as Figure 2M As shown, the electronic device 100 can receive a touch operation applied to the determination control 263. In response to the touch operation, the electronic device 100 can use the training audio file 261 to train a preset timbre generation model.

[0069] In one possible implementation, such as Figure 2N As shown, the electronic device 100 can receive a touch operation on the control 243 within the personalized voice page 240, at which point the indicator on the control 243 changes from "off" to "on". In response to this touch operation, such as... Figure 2O As shown, the electronic device 100 can display a preset noise interface 270. The preset noise interface 270 may include one or more preset noise signal options (e.g., preset noise signal 1 option 271) and corresponding selection controls (e.g., selection control 272 corresponding to preset noise signal 1 option 271), and an add control 273. When a selected control displays a black dot, it indicates that the preset noise signal corresponding to that selected control has been selected for generating the target audio file. For example, when selection control 272 displays a black dot, it indicates that preset noise signal 1 has been selected for generating the target audio file.

[0070] In one possible implementation, such as Figure 2O As shown, when the electronic device 100 receives a touch operation on the addition control 273, in response to the touch operation, the electronic device 100 can display a preset noise signal addition interface 280.

[0071] like Figure 2PAs shown, the preset noise signal adding interface 280 may include the text prompt "Please record ambient sound", control 281, and one or more application icons (e.g., gallery application icon, recorder application icon, etc.). App icon Application icons, etc., and controls 282, etc. Control 281 can be used to receive touch operations (e.g., clicks) applied to it. In response to the touch operation, the electronic device 100 can receive ambient noise signals through a microphone in the preset noise signal addition interface 280. An application icon of a specific application can be used to receive touch operations (e.g., clicks) applied to it by the user. In response to the touch operation, the electronic device 100 can display the application interface of the specified application, allowing the user to select voice messages and / or audio files within the specified application. Control 282 can be used to receive touch operations (e.g., clicks) applied to it by the user. In response to the touch operation, the electronic device 100 can determine a preset noise signal based on the received ambient noise signal and / or the selected voice message and / or audio file in the specified application. For example, the electronic device 100 determines the received ambient noise signal as the preset noise signal, or the electronic device 100 extracts noise signals from the selected voice message and / or audio file in the specified application and determines them as the preset noise signal.

[0072] Electronic device 100 can receive touch operations applied to control 282, and respond to the touch operation, such as Figure 2Q As shown, the electronic device 100 can determine a preset noise signal based on the received ambient noise signal and / or a selected voice canceller and / or audio file in a specified application, and display a text prompt message "Generating your preset noise..." in the preset noise signal adding interface 280. When the preset noise signal is successfully extracted, such as... Figure 2R As shown, the electronic device 100 can display the text prompt "Your preset noise has been generated" and the confirmation control 283 in the preset noise signal addition interface 280.

[0073] Electronic device 100 can receive a touch operation (e.g., a click) applied to a determination control 283, and in response to the touch operation, electronic device 100 can display a preset noise interface 270.

[0074] like Figure 2S As shown, the preset noise interface 270 can display options for newly extracted preset noise signals (e.g., option 274 for preset noise signal 2) and corresponding selection controls (e.g., control 275 corresponding to option 274 for preset noise signal 2), the description of which can be found in the foregoing. Figure 2O The explanation shown.

[0075] In the embodiments of this application, Figures 2A-2S This application is provided as an example and does not constitute any limitation.

[0076] Figure 3A This is a schematic diagram illustrating the specific process of an audio processing method provided in an embodiment of this application.

[0077] like Figure 3A As shown, this audio processing method aims to maintain the timbre of the audio signal 1 to be processed as it was when picked up by the microphone. The specific process may include:

[0078] S301. When the electronic device 100 is in the first mode, the electronic device 100 receives the audio signal 1 to be processed through the microphone.

[0079] Specifically, when the electronic device 100 receives an input to activate the first mode, the electronic device 100 determines that it is in the first mode. The input to activate the first mode can, for example, act on... Figure 2D When a touch operation is received on the control 241, the indicator on the control 241 can change from "off" to "on".

[0080] The first mode can be: maintaining the timbre of the audio signal 1 to be processed as it was when the microphone picked it up, or generating a target audio file with the user's current timbre based on text information. Conversely, the second mode can be: processing the timbre of the audio signal 1 to be processed to the user's preset timbre, or generating a target audio file with a preset timbre based on text information.

[0081] Specifically, the audio signal to be processed 1 can be an audio signal input by a user when using electronic device 100 for calls, live streaming, recording, or entering voice notes. This audio signal to be processed 1 can include: a voice signal 1 emitted by the user when speaking and an ambient noise signal 1 existing around electronic device 100, excluding the voice signal. Electronic device 100 can acquire the audio signal to be processed 1 through a microphone mounted on electronic device 100. The sound source of the audio signal to be processed 1 can be located in the surrounding environment of electronic device 100. The audio signal to be processed 1 can be acquired through one microphone on electronic device 100, or through multiple microphones on electronic device 100. Alternatively, the audio signal to be processed 1 can also be a voice signal sent to electronic device 100 by other electronic devices. That is to say, this application does not limit the source of the audio signal to be processed 1 acquired by electronic device 100.

[0082] S302. Electronic device 100 determines whether to use a preset noise signal.

[0083] Specifically, when the electronic device 100 receives an input to enable the function of superimposing a preset noise signal, the electronic device 100 determines to use the preset noise signal. Otherwise, the electronic device 100 does not use the preset noise signal, but instead uses the ambient noise signal 1 in the audio signal 1 to be processed for superposition.

[0084] For example, the input to enable the superimposed preset noise signal function can be as follows: Figure 2N As shown, the electronic device 100 can receive a touch operation on the control 243 in the personalized voice page 240, at which time the label on the control 243 changes from "off" to "on".

[0085] Scenario 1: The electronic device does not use the preset noise signal, but instead uses the ambient noise signal from when the radio was being recorded.

[0086] S303. Electronic device 100 extracts the text to be synthesized and acoustic features 1 from the audio signal 1 to be processed through a large speech recognition model.

[0087] Among them, acoustic features can refer to features related to the anatomical structure of the human vocal mechanism, and may include one or more of the following: vocal tone, pitch, timbre, formants, fundamental frequency, and speech spectrum envelope.

[0088] The large speech recognition model (also known as the sound field recording algorithm model) can be a large model based on recurrent neural network (RNN), convolutional neural network (CNN), Transformer model, or end-to-end model. This application does not impose any restrictions on this.

[0089] In one possible implementation, the large-scale speech recognition model may include a noise signal extraction module, an acoustic feature extraction module, and a text recognition module. The electronic device 100 can extract environmental noise signal 1 through the noise signal extraction module, extract acoustic features 1 through the acoustic feature extraction module, and extract the text to be synthesized through the text recognition module. The text recognition module may integrate modules such as a dialect recognition module, a missing word recognition module, and a semantic confusion correction module. Specifically, the dialect recognition module can be used to recognize text in a user's speech signal with a dialect accent; the missing word recognition module can identify and complete missing words in the text included in the speech signal; and the semantic confusion correction module can correct misidentified or semantically confused words to correct words.

[0090] The text to be synthesized can be generated by a large speech recognition model based on the text included in the speech signal 1 of the audio signal to be processed. For example:

[0091] If the electronic device 100 collects the audio signal 1 to be processed through the microphone, and the text included in the speech signal 1 of the audio signal 1 to be processed is "I climbed the beautiful Huangshan", then the text to be synthesized can be "I climbed the beautiful Huangshan".

[0092] If the electronic device 100 collects the audio signal 1 to be processed through the microphone, and the text included in the speech signal 1 of the audio signal 1 to be processed is "I climbed the beautiful Huangshan", then the speech recognition big model can identify the missing character in the text through the missing character recognition module, and fill in the missing character "beauty" to generate the text to be synthesized "I climbed the beautiful Huangshan".

[0093] If the electronic device 100 collects the audio signal 1 to be processed through the microphone, and the text included in the speech signal 1 of the audio signal 1 to be processed is "I climbed the beautiful Huangshan", the speech recognition big model recognizes the text as "I climbed the charming Huangshan" based on the audio signal 1 to be processed. Then, the speech recognition big model, based on the semantic confusion correction module, recognizes that the word "charm" has semantic confusion and corrects it to the correct word "beautiful", generating the text to be synthesized as "I climbed the beautiful Huangshan".

[0094] S304. Electronic device 100 generates a noiseless speech signal 1 based on acoustic feature 1 and the text to be synthesized using an audio generation model.

[0095] Specifically, the audio generation model can be: Tacotron2, an end-to-end audio generation model based on WaveNet and Tacotron combined with a spectrogram prediction network and a vocoder; Transformer-TTS, an end-to-end audio generation model based on Transformer and a TTS system; FastSpeech2, a non-autoregressive end-to-end audio generation model based on the Transformer-TTS architecture; or DeepVoice3, an audio generation model based on fully convolutional sequence-to-sequence (which can convert the text features of the text to be synthesized into vocoder parameters through fully parallel computation and use them as input to the waveform synthesis model to generate noiseless speech signals), etc. In other words, this application does not limit the specific type of audio generation model.

[0096] Since the audio generation model uses acoustic feature 1 to generate noiseless speech signal 1 corresponding to the text to be synthesized, and acoustic feature 1 is the acoustic feature extracted from speech signal 1 in the audio signal 1 to be processed, the timbre of the generated noiseless speech signal 1 can be close to the timbre of speech signal 1 in the audio signal 1 to be processed when the electronic device 100 picks up the sound through the microphone, that is, close to the user's current timbre.

[0097] S305. Electronic device 100 extracts ambient noise signal 1 from audio signal 1 to be processed.

[0098] Specifically, since the spectral characteristics, temporal characteristics, and other feature parameters of environmental noise signals and human-generated speech signals are different, electronic device 100 can extract the environmental noise signal 1 from the audio signal 1 to be processed through the noise signal extraction module in the speech recognition large model. Electronic device 100 can extract the environmental noise signal 1 from the audio signal 1 to be processed based on artificial intelligence algorithms such as RNN, CNN, and deep neural (machine learning, ML) networks.

[0099] It is understandable that step S305 can be incorporated into step S303, that is, in step S303, electronic device 100 can extract the text to be synthesized, acoustic features 1 and environmental noise signal 1 from the audio signal 1 to be processed through the speech recognition big model.

[0100] S306. Electronic device 100 generates target audio file 1 based on noiseless speech signal 1 and ambient noise signal 1.

[0101] Specifically, the electronic device 100 can use an environmental sound field fusion algorithm model (also known as a mixing big model) to superimpose one or more features of the environmental noise signal 1, such as the amplitude and power of the environmental noise signal 1, the signal-to-noise ratio between the user's speech signal 1 and the environmental noise signal 1, and noise frequency masking, onto the noiseless speech signal 1 to generate the target audio file 1.

[0102] The electronic device 100 can pre-train and store an environmental sound field fusion algorithm model. For example, when the electronic device 100 first receives input indicating the activation of a first / second mode, it can train and store the environmental sound field fusion algorithm model. In one possible implementation, the electronic device 100 can acquire multiple training audio signals 2 input by the user via a microphone, and / or obtain training voice messages 2 and / or training audio files 2 stored in a specified application (e.g., a recorder application, a gallery application, etc.). The training voice messages 2, training audio signals 2, and training audio files 2 are used to train the environmental sound field fusion algorithm model. These include: the voice signal 2 emitted by the user when speaking and the ambient noise signal 2 surrounding the user when speaking. The electronic device 100 can train the environmental sound field fusion algorithm model based on the training voice messages 2 and / or training audio signals 2 and / or training audio files 2, using artificial intelligence algorithms such as RNN, CNN, and ML networks. This environmental sound field fusion algorithm model can be used to improve the clarity of user voice signals while preserving the loudness of user voice signals under the influence of environmental noise signals and the environmental masking characteristics.

[0103] In one possible implementation, the aforementioned training audio signal 1 can be used as training audio signal 2, training speech message 1 can be used as training speech message 2, and training audio file 1 can be used as training audio file 2. When training the preset timbre generation model, the electronic device 100 can also train the environmental sound field fusion algorithm model.

[0104] In another possible implementation, the electronic device 100 can update the environmental sound field fusion algorithm model based on the audio signal 1 to be processed. When the electronic device 100 acquires a specified number of audio signals 1 to be processed, it can train and update the environmental sound field fusion algorithm model once based on the specified number of audio signals 1 to be processed.

[0105] In one possible implementation, the noise signal extraction module may include an effect feature recognition module, through which the speech recognition big data model in the electronic device 100 can identify the feature parameters in the environmental noise signal.

[0106] Scenario 2: Electronic device 100 uses a preset noise signal:

[0107] S307. Electronic device 100 extracts the text to be synthesized and acoustic features 1 from the audio signal 1 to be processed through a large speech recognition model.

[0108] The explanation of this step can be found in step S303 above.

[0109] S308. Electronic device 100 generates a noiseless speech signal 1 based on acoustic feature 1 and the text to be synthesized using an audio generation model.

[0110] The explanation of this step can be found in step S304 above.

[0111] S309. Electronic device 100 generates target audio file 2 based on noiseless speech signal 1 and preset noise signal.

[0112] Specifically, the electronic device 100 can use an environmental sound field fusion algorithm model (also known as a mixing model) to superimpose one or more features of a preset noise signal, such as the amplitude, power, and noise frequency masking of the preset noise signal, onto the noiseless speech signal 1 to generate a target audio file 2. For an explanation of the environmental sound field fusion algorithm model, please refer to the description in S306; it will not be repeated here.

[0113] Figure 3B This is a schematic diagram of the device framework for an audio processing method provided in an embodiment of this application.

[0114] like Figure 3B As shown, the audio processing method maintains the timbre of the audio signal 1 to be processed as it is when picked up by the microphone. The related device framework diagram may include: an environmental noise database in memory, an audio recording module, a preset noise signal module, and a processing mode determination module in the application layer; a sample rate converter (SRC), an audio generation model, and an environmental sound field fusion algorithm model in the application framework layer; a large speech recognition model in the hardware abstraction layer; a signal processing module (ADCB COPP) in the digital audio signal processor (ADSP); and a microphone, wherein:

[0115] S0. The environmental noise database can be used to train an environmental sound field fusion algorithm model based on training audio signal 2, and / or training voice message 2, and / or training audio file 2, through artificial intelligence algorithms such as RNN, CNN, and ML networks.

[0116] S1. When the processing mode determination module determines that the current electronic device 100 is in the first mode, it can send the first mode indication information to the speech recognition big model to indicate that the current electronic device 100 is in the first mode, that is, to maintain the timbre of the audio signal 1 to be processed as the timbre when the microphone picks up the sound.

[0117] The preset noise signal module determines whether the electronic device 100 uses preset noise and sends noise superposition indication information to the speech recognition big data model. When the preset noise signal module determines that the electronic device 100 does not use the preset noise signal, the value of the noise superposition indication information is the first value, indicating to the speech recognition big data model that the preset noise signal is not used; when the preset noise signal module determines that the electronic device 100 uses the preset noise signal, the value of the noise superposition indication information is the second value, indicating to the speech recognition big data model that the preset noise signal is used.

[0118] S2. The microphone acquires the audio signal to be processed 1 and sends it to the signal processing module. The signal processing module then sends the audio signal to be processed 1 to the speech recognition model.

[0119] S3. The large-scale speech recognition model extracts the text to be synthesized and acoustic features 1 from the audio signal 1 to be processed, and sends the text to be synthesized and acoustic features 1 to the audio generation model. The audio generation model generates a noiseless speech signal 1 based on the text to be synthesized and acoustic features 1.

[0120] S4A. When the speech recognition big model determines that the electronic device 100 does not use the preset noise signal based on the value of the noise superposition indication information, the speech recognition big model extracts the environmental noise signal 1 from the audio signal 1 to be processed and sends the environmental noise signal 1 to the environmental sound field fusion algorithm model.

[0121] Alternatively, the speech recognition big model can extract the environmental noise signal 1, the text to be synthesized, and the acoustic feature 1 from the audio signal to be processed 1. When the speech recognition big model determines that the electronic device 100 does not use the preset noise signal based on the value of the noise superposition indication information, the speech recognition big model sends the environmental noise signal 1 to the environmental sound field fusion algorithm model.

[0122] S4B. When the value of the noise superposition indication information is the second value, the speech recognition big model will not extract the environmental noise signal 1 from the audio signal to be processed 1. At this time, the preset noise signal module sends the preset noise signal to the environmental sound field fusion algorithm model.

[0123] Alternatively, the speech recognition big model can extract the environmental noise signal 1, the text to be synthesized, and the acoustic feature 1 from the audio signal to be processed 1. When the value of the noise superposition indication information is the second value, and the speech recognition big model determines to use the preset noise signal, the speech recognition big model will not send the environmental noise signal 1 to the environmental sound field fusion algorithm model. At this time, the preset noise signal module sends the preset noise signal to the environmental sound field fusion algorithm model.

[0124] S5. The audio generation model sends a noiseless speech signal to the environmental sound field fusion algorithm model.

[0125] S6. When the electronic device 100 does not use the preset noise signal, the environmental sound field fusion algorithm model generates the target audio file 1 based on the noiseless speech signal 1 and the environmental noise signal 1; when the electronic device 100 uses the preset noise signal, the environmental sound field fusion algorithm model generates the target audio file 2 based on the noiseless speech signal 1 and the preset noise signal.

[0126] Then, the ambient sound field fusion algorithm model sends the target audio file 1 / target audio file 2 to the audio recording module via a sample rate converter (SRC) so that the user can play the target audio signal 1 / target audio file 2 through the audio recording module.

[0127] In one possible implementation:

[0128] S7. When the electronic device 100 neither activates the first mode nor the second mode, the electronic device 100 does not perform any processing on the audio signal 1 to be processed. The processing mode determination module sends a third mode indication message to the speech recognition big model to instruct the speech recognition big model not to perform any processing on the audio signal 1 to be processed.

[0129] S8. When the speech recognition big model receives the third mode indication information and receives the audio signal to be processed 1, the speech recognition big model sends the audio signal to be processed 1 to the audio recording (AudioRecord) module through the sampling rate converter (SRC).

[0130] Figure 4A This is a schematic diagram illustrating the specific process of another audio processing method provided in an embodiment of this application.

[0131] like Figure 4A As shown, this audio processing method processes the timbre of the audio signal 1 to a preset timbre, and its specific process may include:

[0132] S401. When the electronic device 100 is in the second mode, the electronic device 100 receives the audio signal 1 to be processed through the microphone.

[0133] When the electronic device 100 receives an input to activate the second mode, the electronic device 100 determines that it is in the second mode. The input to activate the second mode can, for example, act on... Figure 2D When a touch operation is received on the control 242, the indicator on the control 242 can change from "off" to "on".

[0134] Further details regarding this step can be found in the description of S301, and will not be repeated here.

[0135] S402. Electronic device 100 determines whether to use a preset noise signal.

[0136] Further details regarding this step can be found in the description of S302, and will not be repeated here.

[0137] Scenario 1: The electronic device does not use the preset noise signal, but instead uses the ambient noise signal from when the radio was being recorded.

[0138] S403. Electronic device 100 extracts the text to be synthesized from the audio signal 1 to be processed using a large speech recognition model.

[0139] Further details regarding this step can be found in the description of S303, and will not be repeated here.

[0140] S404. Electronic device 100 generates a noiseless speech signal 2 corresponding to the text to be synthesized through a preset timbre generation model. The timbre of the noiseless speech signal 2 is a preset timbre.

[0141] The electronic device 100 can pre-train and store a preset timbre generation model. The electronic device 100 can also be used in a voice recording interface (e.g., Figure 2F The voice recording interface 250 shown receives the user-input training audio signal 1 via the microphone, and / or obtains the specified application (e.g., gallery application, recorder application, etc.). The training audio signal 1 and / or training audio file 1 in the specified application are used to train a preset timbre generation model. The timbre of the audio signal 1 in the training audio signal 1, training audio message 1, and training audio file 1 is the preset timbre.

[0142] Specifically, the electronic device 100 can extract acoustic features 2 from the training audio signal 1, and / or training voice message 1, and / or training audio file 1. These acoustic features 2 are the acoustic features corresponding to the preset timbre. Then, the electronic device 100 can train a preset timbre generation model based on artificial intelligence algorithms such as RNN, CNN, and ML networks. This preset timbre generation model includes acoustic features 2. Next, the electronic device 100 can input the text to be synthesized extracted from the audio signal 1 into the preset timbre generation model. Through the preset timbre generation model, a noiseless speech signal 2 corresponding to the text to be synthesized is generated. Since the electronic device 100 uses the preset timbre generation model to generate the noiseless speech signal 2 corresponding to the text to be synthesized, the acoustic feature of the noiseless speech signal 2 is acoustic feature 2, and acoustic feature 2 is the acoustic feature corresponding to the preset timbre. Therefore, the timbre of the generated noiseless speech signal 2 is the user's preset timbre.

[0143] S405. Electronic device 100 extracts ambient noise signal 1 from audio signal 1 to be processed.

[0144] The explanation of this step can be found in the description of S305, and will not be repeated here.

[0145] S406. Electronic device 100 generates target audio file 3 based on noiseless speech signal 2 and ambient noise signal 1.

[0146] Specifically, the electronic device 100 can use an environmental sound field fusion algorithm model (also known as a mixing large model) to superimpose the feature parameters of the environmental noise signal 1 onto the noiseless speech signal 2 to generate the target audio file 3.

[0147] The description of the environmental sound field fusion algorithm model can be found in the aforementioned S306, and will not be repeated here.

[0148] Scenario 2: Electronic device 100 uses a preset noise signal:

[0149] S407. Electronic device 100 extracts the text to be synthesized from the audio signal 1 to be processed using a large speech recognition model.

[0150] Further details regarding this step can be found in the description of S403, and will not be repeated here.

[0151] S408. Electronic device 100 generates a noiseless speech signal 2 corresponding to the text to be synthesized through a preset timbre generation model. The timbre of the noiseless speech signal 2 is a preset timbre.

[0152] Further details regarding this step can be found in the description of S404, and will not be repeated here.

[0153] S409. Electronic device 100 generates target audio file 4 based on noiseless speech signal 2 and preset noise signal.

[0154] Specifically, the electronic device 100 can use an environmental sound field fusion algorithm model (also known as a mixing model) to superimpose the characteristic parameters of a preset noise signal (such as the amplitude, power, noise frequency masking, or one or more other features of the preset noise signal) onto the noiseless speech signal 2 to generate a target audio file 4. For an explanation of the environmental sound field fusion algorithm model, please refer to the description in S306; it will not be repeated here.

[0155] Figure 4B A schematic diagram of the apparatus framework for another audio processing method provided in an embodiment of this application.

[0156] like Figure 4BAs shown, the audio processing method processes the timbre of the audio signal 1 to be processed into a preset timbre. The related device framework diagram may include: an environmental noise database and a training audio database in memory; an audio recording module, a preset noise signal module, and a processing mode determination module in the application layer; a sample rate converter (SRC), a preset timbre generation model, and an environmental sound field fusion algorithm model in the application framework layer; a large speech recognition model in the hardware abstraction layer; a signal processing module (ADCB COPP) in the digital audio signal processor (ADSP); and a microphone, wherein:

[0157] S0. The environmental noise database can be used to train an environmental sound field fusion algorithm model based on training audio signal 2, and / or training speech message 2, and / or training audio file 2, using artificial intelligence algorithms such as RNN, CNN, and ML networks. The training audio database can be used to train a preset timbre generation model based on acoustic features 2 in training audio signal 1, and / or training speech message 1, and / or training audio file 1, using artificial intelligence algorithms such as RNN, CNN, and ML networks.

[0158] S1. When the processing mode determination module determines that the current electronic device 100 is in the second mode, it can send the second mode indication information to the speech recognition big model to indicate that the current electronic device 100 is in the second mode, that is, to process the timbre of the audio signal 1 to be processed into the preset timbre.

[0159] The preset noise signal module determines whether the electronic device 100 uses preset noise and sends noise superposition indication information to the speech recognition big data model. When the preset noise signal module determines that the electronic device 100 does not use the preset noise signal, the value of the noise superposition indication information is the first value, indicating to the speech recognition big data model that the preset noise signal is not used; when the preset noise signal module determines that the electronic device 100 uses the preset noise signal, the value of the noise superposition indication information is the second value, indicating to the speech recognition big data model that the preset noise signal is used.

[0160] S2. The microphone acquires the audio signal to be processed 1 and sends it to the signal processing module (ADCB COPP). The signal processing module (ADCB COPP) then sends the audio signal to be processed 1 to the speech recognition model.

[0161] S3. The large-scale speech recognition model extracts the text to be synthesized from the audio signal 1 to be processed, and sends the text to be synthesized to the preset timbre generation model. The preset timbre generation model generates a noiseless speech signal 2 based on the text to be synthesized.

[0162] S4A. When the speech recognition big model determines that the electronic device 100 does not use the preset noise signal based on the value of the noise superposition indication information, the speech recognition big model extracts the environmental noise signal 1 from the audio signal 1 to be processed and sends the environmental noise signal 1 to the environmental sound field fusion algorithm model.

[0163] Alternatively, the speech recognition big model can extract the environmental noise signal 1 and the text to be synthesized from the audio signal 1 to be processed. When the speech recognition big model determines that the electronic device 100 does not use the preset noise signal based on the value of the noise superposition indication information, the speech recognition big model sends the environmental noise signal 1 to the environmental sound field fusion algorithm model.

[0164] S4B. When the value of the noise superposition indication information is the second value, the speech recognition big model will not extract the environmental noise signal 1 from the audio signal to be processed 1. At this time, the preset noise signal module sends the preset noise signal to the environmental sound field fusion algorithm model.

[0165] Alternatively, the speech recognition big model can extract the environmental noise signal 1 and the text to be synthesized from the audio signal 1 to be processed. When the value of the noise superposition indicator information is the second value, and the speech recognition big model determines to use the preset noise signal, the speech recognition big model will not send the environmental noise signal 1 to the environmental sound field fusion algorithm model. At this time, the preset noise signal module sends the preset noise signal to the environmental sound field fusion algorithm model.

[0166] S5. The preset timbre generation model sends a noiseless speech signal to the environmental sound field fusion algorithm model.

[0167] S6. When the electronic device 100 does not use the preset noise signal, the environmental sound field fusion algorithm model generates the target audio file 3 based on the noiseless speech signal 2 and the environmental noise signal 1; when the electronic device 100 uses the preset noise signal, the environmental sound field fusion algorithm model generates the target audio file 4 based on the noiseless speech signal 2 and the preset noise signal.

[0168] Then, the ambient sound field fusion algorithm model sends the target audio file 3 / target audio file 4 to the audio recording module via a sample rate converter (SRC) so that the user can play the target audio signal 3 / target audio file 4 through the audio recording module.

[0169] Figure 5A This is a schematic diagram illustrating the specific process of another audio processing method provided in an embodiment of this application.

[0170] like Figure 5A As shown, this audio processing method involves an electronic device 100 generating a target audio file with a preset timbre based on received text information 1. The specific process may include:

[0171] S501. When the electronic device 100 is in the second mode, the electronic device 100 receives text information 1.

[0172] The text information 1 can be a piece of text input by the user, or it can be a piece of text obtained by the electronic device 100 through a specified application, or it can be a piece of text captured by the electronic device 100 through a camera. That is to say, this application does not limit the specific type and source of the text information 1.

[0173] S502. Electronic device 100 determines whether to use a preset noise signal.

[0174] Further details regarding this step can be found in the description of S302, and will not be repeated here.

[0175] Scenario 1: The electronic device does not use the preset noise signal, but instead uses the current ambient noise signal:

[0176] S503. Electronic device 100 generates noiseless speech signal 3 based on text information 1 by using a preset timbre generation model.

[0177] The electronic device 100 can pre-train and store a preset timbre generation model. The electronic device 100 can also be used in a voice recording interface (e.g., Figure 2F The voice recording interface 250 shown receives the user-input training audio signal 1 via the microphone, and / or obtains the signal from a specified application (e.g., a gallery application, a voice recorder application, etc.). The training audio signal 1 and / or training audio file 1 in the specified application are used to train a preset timbre generation model. The timbre of the audio signal 1, training audio message 1, and training audio file 1 used to train the preset timbre generation model is the preset timbre.

[0178] Specifically, the electronic device 100 can extract acoustic features 2 from the training audio signal 1, and / or training voice message 1, and / or training audio file 1. These acoustic features 2 are the acoustic features corresponding to the preset timbre. Then, the electronic device 100 can train a preset timbre generation model based on artificial intelligence algorithms such as RNN, CNN, and ML networks. Next, the electronic device 100 can input text information 1 into the preset timbre generation model, and generate a noiseless speech signal 3 corresponding to the text information 1. Since the electronic device 100 uses the preset timbre generation model to generate the noiseless speech signal 3 corresponding to the text information 1, the acoustic feature of the noiseless speech signal 3 is acoustic feature 2, and acoustic feature 2 is the acoustic feature corresponding to the preset timbre. Therefore, the timbre of the generated noiseless speech signal 3 is the user's preset timbre.

[0179] S504. Electronic device 100 collects ambient noise signal 2 through microphone.

[0180] Specifically, when the electronic device 100 determines that it will not use the preset noise signal, the electronic device 100 can turn on the microphone and collect the ambient noise signal 2 of the environment around the electronic device 100 through the microphone.

[0181] S505. Electronic device 100 generates target audio file 5 based on noiseless speech signal 3 and ambient noise signal 2.

[0182] Specifically, the electronic device 100 can use an environmental sound field fusion algorithm model (also known as a mixing big model) to superimpose one or more features of the environmental noise signal 2, such as the amplitude, power, and noise frequency masking of the environmental noise signal 2, onto the noiseless speech signal 3 to generate a target audio file 5.

[0183] The description of the environmental sound field fusion algorithm model can be found in the aforementioned S306, and will not be repeated here.

[0184] Scenario 2: Electronic device 100 uses a preset noise signal:

[0185] S506. Electronic device 100 generates noiseless speech signal 3 based on text information 1 using a preset timbre generation model.

[0186] Further details regarding this step can be found in the description of S503, and will not be repeated here.

[0187] S507. Electronic device 100 generates target audio file 6 based on noiseless speech signal 3 and preset noise signal.

[0188] Specifically, the electronic device 100 can use an environmental sound field fusion algorithm model (also known as a mixing model) to superimpose one or more features of a preset noise signal, such as the amplitude, power, and noise frequency masking of the preset noise signal, onto the noiseless speech signal 3 to generate a target audio file 6. For an explanation of the environmental sound field fusion algorithm model, please refer to the description in S306; it will not be repeated here.

[0189] In this embodiment of the application, text information 1 can also be replaced with image information, that is to say, electronic device 100 can... Figure 5A The illustrated embodiment generates a target audio file based on image information. Implementation details are provided below. Figure 5A The process shown is not detailed here.

[0190] Figure 5B This is a schematic diagram of the apparatus framework for another audio processing method provided in the embodiments of this application.

[0191] like Figure 5B As shown, this audio processing method generates a target audio file with a preset timbre based on text information 1. The related device framework diagram may include: an environmental noise database and a training audio database in memory; an audio tracking module, a processing mode determination module, and a preset noise signal module in the application layer; a sample rate converter & volume (SRC & Volume) module, a preset timbre generation model, and an environmental sound field fusion algorithm model in the application framework layer; an audio file output (Outwrite) module in the hardware abstraction layer; a smart power amplifier module (SmartPA COPP) in the digital audio signal processor (ADSP); a microphone; and a speaker, wherein:

[0192] S0. The environmental noise database can be used to train an environmental sound field fusion algorithm model based on training audio signal 2, and / or training speech message 2, and / or training audio file 2, using artificial intelligence algorithms such as RNN, CNN, and ML networks. The training audio database can be used to train a preset timbre generation model based on acoustic features 2 in training audio signal 1, and / or training speech message 1, and / or training audio file 1, using artificial intelligence algorithms such as RNN, CNN, and ML networks.

[0193] When the processing mode determination module determines that the current electronic device 100 is in the second mode, it can send the second mode indication information to the sampling rate converter & volume (SRC & Volume) module to indicate that the current electronic device 100 is in the second mode, that is, to generate a target audio file with a preset timbre based on the text information.

[0194] S1. The Audio Track module receives text message 1. When the Sample Rate Converter & Volume (SRC & Volume) module receives the second mode indication information, the SRC & Volume module sends text message 1 to the preset timbre generation model.

[0195] S2. The preset timbre generation model generates a noiseless speech signal 3 corresponding to the text information 1, and sends the noiseless speech signal 3 to the environmental sound field fusion algorithm model.

[0196] S3A. The preset noise signal module determines whether the electronic device 100 uses preset noise. When the preset noise signal module determines that the electronic device 100 uses preset noise, the preset noise signal module sends the preset noise signal to the ambient sound field fusion algorithm model.

[0197] S3B. When the preset noise signal module determines that the electronic device 100 does not use the preset noise signal, the preset noise signal module sends a microphone start command to the microphone, triggering the microphone to collect the ambient noise signal 2. Then, the microphone sends the ambient noise signal 2 to the ambient sound field fusion algorithm model.

[0198] S4. The ambient sound field fusion algorithm model generates a target audio file 5 based on the noiseless speech signal 3 and the ambient noise signal 2, or generates a target audio file 6 based on the noiseless speech signal 3 and a preset noise signal. Then, the ambient sound field fusion algorithm model sends the target audio file 5 / target audio file 6 to the sampling rate converter & volume (SRC & Volume) module.

[0199] The S5. Sample Rate Converter & Volume (SRC & Volume) module sends the target audio file 5 / target audio file 6 to the speaker through the audio file output (Outwrite) module and the SmartPA COPP module, enabling the speaker to play the target audio file 5 / target audio file 6.

[0200] Figure 6A This is a schematic diagram illustrating the specific process of another audio processing method provided in an embodiment of this application.

[0201] like Figure 6A As shown, this audio processing method involves an electronic device 100 generating a target audio file with the user's current voice timbre based on received text information 1. The specific process may include:

[0202] S601. When the electronic device 100 is in the first mode, the electronic device 100 receives text information 1.

[0203] For an explanation of text information 1, please refer to the explanation of S501 above.

[0204] S602. Electronic device 100 acquires audio signal 2 via microphone.

[0205] Specifically, audio signal 2 can be the audio signal input by the user when using electronic device 100 to make a call, live stream, record, or enter voice notes.

[0206] The audio signal 2 may include: a voice signal emitted by the user when speaking and an ambient noise signal 3 existing around the electronic device 100, excluding the voice signal. The electronic device 100 can acquire the audio signal 2 through a microphone mounted on the electronic device 100. The sound source of the audio signal 2 may be located in the surrounding environment of the electronic device 100. The audio signal 2 may be acquired through one microphone on the electronic device 100, or through multiple microphones on the electronic device 100. Alternatively, the audio signal 2 may also be a voice signal sent to the electronic device 100 by other electronic devices. That is to say, this application does not limit the source of the audio signal 2 acquired by the electronic device 100.

[0207] S603. Electronic device 100 extracts acoustic features 3 from audio signal 2 using a large speech recognition model. Acoustic feature 3 is the acoustic feature corresponding to the user's current voice timbre.

[0208] For details regarding acoustic feature extraction and related explanations, please refer to step S303 above.

[0209] In one possible implementation, the electronic device 100 can also obtain the audio file 2 in the specified application, and then the electronic device 100 extracts the acoustic features 3 in the audio file 2 through a large speech recognition model.

[0210] S604. Electronic device 100 determines whether to use a preset noise signal.

[0211] Further details regarding this step can be found in the description of S302, and will not be repeated here.

[0212] Scenario 1: The electronic device does not use the preset noise signal, but instead uses the current ambient noise signal:

[0213] S605. Electronic device 100 generates noiseless speech signal 4 based on text information 1 and acoustic features 3 using an audio generation model.

[0214] Since the audio generation model uses acoustic feature 3 to generate noiseless speech signal 4 corresponding to text information 1, and acoustic feature 3 is an acoustic feature extracted from audio signal 2, the timbre of the generated noiseless speech signal 4 can be close to the timbre of audio signal 2 when the electronic device 100 picks up sound through the microphone, that is, close to the user's current timbre.

[0215] S606. Electronic device 100 extracts ambient noise signal 3 from audio signal 2.

[0216] For instructions on this step, please refer to the description in S305.

[0217] S607. Electronic device 100 generates a target audio file 7 based on noiseless speech signal 4 and ambient noise signal 3. Specifically, electronic device 100 can use an ambient sound field fusion algorithm model (also known as a mixing model) to superimpose one or more features of the ambient noise signal 3, such as amplitude, power, signal-to-noise ratio between the user's speech signal in audio signal 2 and ambient noise signal 3, noise frequency masking, etc., onto the noiseless speech signal 4 to generate the target audio file 7.

[0218] The description of the environmental sound field fusion algorithm model can be found in the aforementioned S306, and will not be repeated here.

[0219] Scenario 2: Electronic device 100 uses a preset noise signal:

[0220] S608. Electronic device 100 generates noiseless speech signal 4 based on text information 1 and acoustic features 3 using an audio generation model.

[0221] Further details regarding this step can be found in the description of S605, and will not be repeated here.

[0222] S609. Electronic device 100 generates target audio file 8 based on noiseless speech signal 4 and preset noise signal.

[0223] Specifically, the electronic device 100 can use an environmental sound field fusion algorithm model (also known as a mixing model) to superimpose one or more features of a preset noise signal, such as the amplitude, power, and noise frequency masking of the preset noise signal, onto the noiseless speech signal 4 to generate a target audio file 8. For an explanation of the environmental sound field fusion algorithm model, please refer to the description in S306; it will not be repeated here.

[0224] In this embodiment of the application, text information 1 can also be replaced with image information, that is to say, electronic device 100 can... Figure 6A The illustrated embodiment generates a target audio file based on image information. Implementation details are provided below. Figure 6AThe process shown is not detailed here.

[0225] Figure 6B This is a schematic diagram of the apparatus framework for another audio processing method provided in the embodiments of this application.

[0226] like Figure 6B As shown, this audio processing method generates a target audio file based on the user's current timbre according to text information 1. Its related device framework diagram may include: an environmental noise database in memory; an audio tracking module, a processing mode determination module, and a preset noise signal module in the application layer; a sample rate converter & volume (SRC & Volume) module, an audio generation model, and an environmental sound field fusion algorithm model in the application framework layer; an audio file output (Outwrite) module and a large speech recognition model in the hardware abstraction layer; a smart power amplifier module (SmartPA COPP) and a signal processing module (ADCB COPP) in the digital audio signal processor (ADSP); a microphone; and a speaker, wherein:

[0227] S0. The environmental noise database can be used to train an environmental sound field fusion algorithm model based on training audio signal 2, and / or training voice message 2, and / or training audio file 2, through artificial intelligence algorithms such as RNN, CNN, and ML networks.

[0228] S1. When the processing mode determination module determines that the current electronic device 100 is in the first mode, it can send the first mode indication information to the sampling rate converter & volume (SRC & Volume) module to indicate that the current electronic device 100 is in the first mode, that is, to generate the target audio file of the user's current timbre based on the text information.

[0229] The processing mode determination module sends an audio signal acquisition command to the microphone.

[0230] The preset noise signal module determines whether the electronic device 100 uses preset noise and sends noise superposition indication information to the speech recognition big data model. When the preset noise signal module determines that the electronic device 100 does not use the preset noise signal, the value of the noise superposition indication information is the first value, indicating to the speech recognition big data model that the preset noise signal is not used; when the preset noise signal module determines that the electronic device 100 uses the preset noise signal, the value of the noise superposition indication information is the second value, indicating to the speech recognition big data model that the preset noise signal is used.

[0231] S2. The Audio Track module receives text message 1. When the Sample Rate Converter & Volume (SRC & Volume) module receives the first mode indication information, it sends text message 1 to the audio generation model.

[0232] In response to the audio signal acquisition command, the microphone acquires audio signal 2 and sends audio signal 2 to the speech recognition big model through the signal processing module (ADCB COPP).

[0233] S3. The large speech recognition model extracts acoustic features 3 from audio signal 2 and sends acoustic features 3 to the audio generation model.

[0234] S4. The audio generation model generates a noiseless speech signal 4 based on acoustic feature 3 and text information 1, and sends the noiseless speech signal 4 to the environmental sound field fusion algorithm model.

[0235] S5A. When the speech recognition big model determines that the electronic device 100 does not use the preset noise signal based on the value of the noise superposition indication information, the speech recognition big model extracts the environmental noise signal 3 from the audio signal 2 and sends the environmental noise signal 3 to the environmental sound field fusion algorithm model.

[0236] Alternatively, the speech recognition big model can extract the environmental noise signal 3 and acoustic features 3 from the audio signal 2. When the speech recognition big model determines that the electronic device 100 does not use the preset noise signal based on the value of the noise superposition indication information, the speech recognition big model sends the environmental noise signal 3 to the environmental sound field fusion algorithm model.

[0237] S5B. When the speech recognition big model determines that the electronic device 100 uses a preset noise signal based on the value of the noise superposition indication information, the preset noise signal module sends the preset noise signal to the environmental sound field fusion algorithm model.

[0238] Alternatively, the speech recognition big model can extract the environmental noise signal 3 and acoustic feature 3 from the audio signal 2. When the speech recognition big model determines that the electronic device 100 uses a preset noise signal based on the value of the noise superposition indication information, the speech recognition big model will not send the environmental noise signal 3 to the environmental sound field fusion algorithm model. At this time, the preset noise signal module sends the preset noise signal to the environmental sound field fusion algorithm model.

[0239] S6. The ambient sound field fusion algorithm model generates a target audio file 7 based on the noiseless speech signal 4 and the ambient noise signal 3, or generates a target audio file 8 based on the noiseless speech signal 4 and a preset noise signal. Then, the ambient sound field fusion algorithm model sends the target audio file 7 / target audio file 8 to the sampling rate converter & volume (SRC & Volume) module.

[0240] The S7. Sample Rate Converter & Volume (SRC & Volume) module sends the target audio file 7 / target audio file 8 to the speaker through the audio file output (Outwrite) module and the SmartPA COPP module, enabling the speaker to play the target audio file 7 / target audio file 8.

[0241] In this embodiment, the audio signal to be processed 1 can be called the first audio signal, the audio signal 2 can be called the second audio signal, the environmental noise signal 1 can be called the first environmental noise signal, the environmental noise signal 2 can be called the second environmental noise signal, the environmental noise signal 3 can be called the third environmental noise signal, the noiseless speech signal 1 can be called the first noiseless speech signal, the noiseless speech signal 2 can be called the second noiseless speech signal, the noiseless speech signal 3 can be called the third noiseless speech signal, the noiseless speech signal 4 can be called the fourth noiseless speech signal, the target audio file 1 can be called the first target audio file, the target audio file 2 can be called the second target audio file, the target audio file 3 can be called the third target audio file, the target audio file 4 can be called the fourth target audio file, the target audio file 5 can be called the fifth target audio file, the target audio file 6 can be called the sixth target audio file, the target audio file 7 can be called the seventh target audio file, the target audio file 8 can be called the eighth target audio file, the text information 1 can be called the first text information, the acoustic feature 3 can be called the first acoustic feature, the acoustic feature 1 can be called the second acoustic feature, and the acoustic feature 2 can be called the third acoustic feature.

[0242] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, can implement the steps in the above-described method embodiments.

[0243] This application also provides a computer program product, including a computer program that, when run on a processor, can implement the steps executed by the electronic device in the above-described method embodiments.

[0244] This application also provides a chip system, which includes a processing circuit interface circuit. The interface circuit receives instructions and transmits them to the processing circuit, which executes the instructions to cause the chip system to perform the steps executed by the electronic device in any of the method embodiments of this application. The chip system can be a single chip or a chip module composed of multiple chips.

[0245] The term "user interface (UI)" used in the specification and accompanying drawings of this application refers to the medium through which an application or operating system interacts and exchanges information with the user. It converts the internal form of information into a form acceptable to the user. The user interface of an application is source code written in a specific computer language such as Java or Extensible Markup Language (XML). This source code is parsed and rendered on the terminal device, ultimately presenting user-recognizable content, such as images, text, buttons, and other controls. Controls, also known as widgets, are the basic elements of the user interface. Typical controls include toolbars, menu bars, text boxes, buttons, scroll bars, images, and text. The attributes and content of controls in the interface are defined through tags or nodes, such as XML. <textview> 、 <imgview> 、 <videoview>Nodes define the controls contained in the interface. A node corresponds to a control or property in the interface, and after parsing and rendering, the node is presented as the content visible to the user. In addition, many applications, such as hybrid applications, often contain web pages within their interfaces. A web page, also known as a webpage, can be understood as a special control embedded in the application interface. Web pages are source code written in a specific computer language, such as Hypertext Markup Language (HTML), Cascading Style Sheets (CSS), JavaScript (JS), etc. Web page source code can be loaded and displayed as user-readable content by a browser or a web page display component with browser-like functionality. The specific content contained in a webpage is also defined through tags or nodes in the webpage source code; for example, HTML uses tags or nodes to define the content. 、 、 <video> 、 <canvas>Used to define the elements and attributes of a webpage.

[0246] The most common form of user interface is the graphical user interface (GUI), which refers to a user interface related to computer operation displayed graphically. It can be an icon, window, control, or other interface element displayed on the screen of an electronic device. Controls can include visual interface elements such as icons, buttons, menus, tabs, text boxes, dialog boxes, status bars, navigation bars, and widgets.

[0247] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive), etc.

[0248] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.

[0249] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.< / canvas> < / video> < / videoview> < / imgview> < / textview>

Claims

1. An audio processing method, characterized in that, include: The electronic device receives a first audio signal via a microphone; wherein the first audio signal includes the user's voice signal and a first ambient noise signal; The electronic device identifies the text to be synthesized from the first audio signal; When the electronic device is in the first mode, the electronic device generates a first noiseless speech signal based on the text to be synthesized; wherein, the timbre of the first noiseless speech signal is the timbre of the speech signal; When the electronic device determines that it does not use the preset noise signal, the electronic device extracts the first ambient noise signal from the first audio signal; The electronic device generates a first target audio file based on the first noiseless speech signal and the first ambient noise signal; When the electronic device determines that a preset noise signal is to be used, the electronic device generates a second target audio file based on the first noiseless speech signal and the preset noise signal.

2. The method according to claim 1, characterized in that, The method further includes: When the electronic device is in the second mode, the electronic device generates a second noiseless speech signal based on the text to be synthesized; wherein the timbre of the second noiseless speech signal is a preset timbre; When the electronic device determines that it does not use the preset noise signal, the electronic device extracts the first ambient noise signal from the first audio signal; The electronic device generates a third target audio file based on the second noiseless speech signal and the first environmental noise signal; When the electronic device determines that a preset noise signal is to be used, the electronic device generates a fourth target audio file based on the second noiseless speech signal and the preset noise signal.

3. The method according to claim 1 or 2, characterized in that, The method further includes: The electronic device acquires the first text information; The electronic device generates a third noise-free speech signal based on the first text information; wherein the timbre of the third noise-free speech signal is a preset timbre; When the electronic device determines that it will not use the preset noise signal, the electronic device collects a second ambient noise signal through the microphone; The electronic device generates a fifth target audio file based on the third noiseless speech signal and the second environmental noise signal; When the electronic device determines to use a preset noise signal, the electronic device generates a sixth target audio file based on the third noiseless speech signal and the preset noise signal.

4. The method according to claim 1 or 2, characterized in that, The method further includes: The electronic device acquires the first text information; The electronic device acquires a second audio signal through the microphone; The electronic device extracts a first acoustic feature from the second audio signal; The electronic device generates a fourth noiseless speech signal based on the first text information and the first acoustic feature; wherein the timbre of the fourth noiseless speech signal is the user's current timbre. When the electronic device determines that it does not use the preset noise signal, the electronic device extracts the third ambient noise signal from the second audio signal; The electronic device generates a seventh target audio file based on the fourth noiseless speech signal and the third environmental noise signal; When the electronic device determines that a preset noise signal is to be used, the electronic device generates an eighth target audio file based on the fourth noiseless speech signal and the preset noise signal.

5. The method according to claim 1, characterized in that, When the electronic device is in the first mode, the electronic device generates a first noiseless speech signal based on the text to be synthesized, specifically including: The electronic device extracts a second acoustic feature from the first audio signal; wherein the second acoustic feature is the acoustic feature of the speech signal; The electronic device generates the first noiseless speech signal based on the second acoustic feature and the text to be synthesized.

6. The method according to claim 2, characterized in that, The preset timbre is the timbre that the user is in when in a specified state.

7. The method according to claim 5, characterized in that, The electronic device generates a first noiseless speech signal based on the second acoustic feature and the text to be synthesized, specifically including: The electronic device generates the first noiseless speech signal based on the second acoustic features and the text to be synthesized using an audio generation model.

8. The method according to claim 2, characterized in that, When the electronic device is in the second mode, the electronic device generates a second noiseless speech signal based on the text to be synthesized, specifically including: When the electronic device is in the second mode, the electronic device generates a second noiseless speech signal based on the text to be synthesized using a preset timbre generation model; wherein, the preset timbre generation model includes a third acoustic feature, and the third acoustic feature is the acoustic feature corresponding to the preset timbre.

9. An electronic device, characterized in that, The electronic device includes: one or more processors and a memory; the memory is coupled to the one or more processors, the memory being used to store computer program code, the computer program code including computer instructions, and the one or more processors invoking the computer instructions to cause the electronic device to perform the method as described in any one of claims 1-8.

10. A chip system, characterized in that, The chip system is applied to an electronic device, the chip system including one or more processors, the processors being configured to invoke computer instructions to cause the electronic device to perform the method as described in any one of claims 1-8.

11. A computer-readable storage medium comprising instructions, characterized in that, When the instructions are executed on an electronic device, the electronic device causes the electronic device to perform the method as described in any one of claims 1-8.

12. A computer program product, characterized in that, Includes a computer program, which, when executed by a processor, causes the electronic device to perform the method as described in any one of claims 1-8.