Age-based sound generation method and apparatus

By decoupling timbre from age characteristics, a reconstructed timbre feature matching the target age is generated, solving the problems of sound library dependence and sound distortion in existing technologies, and achieving accurate reproduction of the user's timbre.

CN116030787BActive Publication Date: 2026-03-31HUAWEI TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-26
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing voice age control technology relies on high-quality training voice libraries, and it is difficult to collect voice libraries for different age groups of the same person, resulting in a large degree of distortion in the generated voices and a lack of quantitative dimensional definition.

Method used

By obtaining reference voices for users at a set age, the timbre and age features are decoupled. A pre-trained model is used to generate reconstructed timbre features that match the target age, enabling timbre reproduction for users of any age and reducing dependence on the voice library.

Benefits of technology

It achieves accurate reproduction of timbre at different ages and greatly reduces the distortion of the user's personal timbre. It does not require collecting a multi-age voice library for each individual and preserves the user's personal timbre characteristics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116030787B_ABST
    Figure CN116030787B_ABST
Patent Text Reader

Abstract

This application provides an age-based voice generation method and apparatus. The age-based voice generation method includes: acquiring reference speech, which is recorded at a user's set age; acquiring a first operation instruction generated by an operation performed on an interactive interface, the first operation instruction indicating a target age; acquiring target content; and generating target speech based on the target content, the reference speech, and the first operation instruction, wherein the target speech presents the target content using the user's timbre at the target age. This application reduces reliance on a voice library and significantly reduces the distortion between the timbre reproduction at different ages and the user's individual timbre.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio processing, and more particularly to a method and apparatus for generating sound based on age. Background Technology

[0002] With the rapid development of the internet, diverse entertainment options and fragmented mobile browsing occupy most of people's leisure time. Online entertainment methods are also constantly evolving, gradually entering the era of artificial intelligence (AI). Currently, AI-based face-swapping technology has demonstrated enormous commercial value and prospects in entertainment scenarios, especially in applications that present users with different ages and short video effects, which have become phenomenal products.

[0003] As mentioned above, age control technology is relatively mature in the facial field. Because the face has relatively intuitive facial features and muscle distribution, aging changes can often be simulated by modeling facial muscle tightness, skin texture, the number of wrinkles, and the degree of gray hair. However, voice, as another important means of expressing human image, lacks similar quantifiable dimensional definitions, making research on voice age control technology more difficult.

[0004] Existing limited research often relies on building training voice libraries for multiple individuals and for the same individual across multiple age groups, using age as the sole variable to perform blind-box modeling through neural networks. However, this type of research is overly dependent on comprehensive and high-quality training voice libraries, and collecting voice libraries for the same individual across different age groups is extremely difficult. Furthermore, voice libraries based on the target age of other individuals can lead to significant distortion in the generated sounds. Summary of the Invention

[0005] This application provides an age-based voice generation method and apparatus to reduce reliance on sound libraries and significantly reduce the distortion of timbre reproduction at different ages and the user's personal timbre.

[0006] In a first aspect, embodiments of this application provide an age-based voice generation method, comprising: acquiring reference speech, the reference speech being recorded at a user's set age; acquiring a first operation instruction generated by an operation performed on an interactive interface, the first operation instruction indicating a target age; acquiring target content; generating target speech based on the target content, the reference speech, and the first operation instruction, the target speech presenting the target content using the user's timbre at the target age.

[0007] This application's embodiments decouple the timbre and age of the reference speech to obtain age-independent timbre features. Then, based on these age-independent timbre features and the target age, reconstructed timbre features are obtained. These reconstructed timbre features can then be used to express arbitrary target content, thereby enabling the reproduction of a user's timbre at any age. This method does not require collecting a multi-age voice database for an individual, reducing reliance on such databases, and preserves the user's individual timbre, significantly reducing distortion in the reproduction of timbre at different ages compared to the user's individual timbre.

[0008] In one possible implementation, obtaining the reference speech includes: obtaining a second operation instruction generated by an operation performed on the interactive interface, the second operation instruction being used to instruct the recording function to be started; and starting the recording function to record the user's voice to obtain the reference speech.

[0009] In one possible implementation, obtaining the reference speech includes: obtaining a third operation instruction generated by an operation performed on the interactive interface, the third operation instruction being used to instruct the display of a pre-recorded speech file; and obtaining the reference speech based on the user's selection operation on the pre-recorded speech file.

[0010] The set age can refer to the user's current age (i.e., the age at the moment the user generates the speech using the method provided in this application embodiment), or it can be the user's age at the time the selected historical speech was recorded. For example, if the user is currently 30 years old, and the newly recorded speech is used as the reference speech, then the set age corresponding to that reference speech is 30 years old; if the historical speech is used as the reference speech, then the set age is the user's age when the historical speech was recorded (e.g., 20 years old).

[0011] The interactive interface provides controls that users can interact with. User actions on these controls generate commands within the electronic device, which then performs corresponding processing based on these commands. Therefore, users can select a target age group by manipulating the controls on the interactive interface.

[0012] In one possible implementation, the operation on the interactive interface includes clicking on a control on the interactive interface; the control includes an age increase control and an age decrease control, wherein the age increase control corresponds to a preset age range greater than the set age, and the age decrease control corresponds to a preset age range less than the set age; or, the control includes an age increase control and an age decrease control, wherein the age increase control corresponds to a step size for increasing age, and the age decrease control corresponds to a step size for decreasing age.

[0013] In one possible implementation, the operation on the interactive interface includes dragging a slider on the interactive interface; the position of the slider corresponds to the target age.

[0014] The target content is the information conveyed by the target speech.

[0015] In one possible implementation, obtaining the target content includes: obtaining a fourth operation instruction generated by an operation performed on the interactive interface, the fourth operation instruction including the text content input by the user; and using the text content as the target content.

[0016] In one possible implementation, obtaining the target content includes: obtaining a second operation instruction generated by an operation performed on the interactive interface, the second operation instruction being used to instruct the activation of a recording function; activating the recording function to record the user's voice to obtain content audio; and obtaining the target content based on the content audio.

[0017] In one possible implementation, obtaining the target content includes: obtaining the target content based on the reference speech.

[0018] In one possible implementation, when the target age is a single age value, the target voice reflects the user's timbre at that target age; when the target age is an age range, the target voice reflects the user's timbre variations within that target age range.

[0019] In one possible implementation, generating the target speech based on the target content, the reference speech, and the first operation instruction includes: obtaining age-independent timbre features based on the reference speech; obtaining reconstructed timbre features based on the age-independent timbre features and the target age, wherein the reconstructed timbre features are associated with the target age; and generating the target speech based on the target content and the reconstructed timbre features.

[0020] In one possible implementation, obtaining age-independent timbre features based on the reference speech includes: inputting the reference speech into a pre-trained age extraction model to obtain an estimated age; inputting the reference speech into a pre-trained timbre extraction model to obtain estimated timbre features; and decoupling the estimated timbre features and the estimated age to obtain the age-independent timbre features.

[0021] In one possible implementation, obtaining the reconstructed timbre features based on the age-independent timbre features and the target age includes: inputting the age-independent timbre features and the target age into a pre-trained timbre generation model to obtain the reconstructed timbre features.

[0022] In one possible implementation, generating the target speech based on the target content and the reconstructed timbre features includes: inputting the target content and the reconstructed timbre features into a pre-trained speech generation model to obtain the target speech.

[0023] Secondly, embodiments of this application provide an age-based voice generation device, comprising: an acquisition module for acquiring reference speech, the reference speech being recorded at a user's set age; acquiring a first operation instruction generated by an operation performed on an interactive interface, the first operation instruction indicating a target age; acquiring target content; and a processing module for generating target speech based on the target content, the reference speech, and the first operation instruction, the target speech presenting the target content using the user's timbre at the target age.

[0024] In one possible implementation, the operation on the interactive interface includes clicking on a control on the interactive interface; the control includes an age increase control and an age decrease control, wherein the age increase control corresponds to a preset age range greater than the set age, and the age decrease control corresponds to a preset age range less than the set age; or, the control includes an age increase control and an age decrease control, wherein the age increase control corresponds to a step size for increasing age, and the age decrease control corresponds to a step size for decreasing age.

[0025] In one possible implementation, the operation on the interactive interface includes dragging a slider on the interactive interface; the position of the slider corresponds to the target age.

[0026] In one possible implementation, the acquisition module is specifically used to acquire a second operation instruction generated by an operation performed on the interactive interface, the second operation instruction being used to instruct the recording function to be started; the recording function is then activated to record the user's voice to obtain the reference speech.

[0027] In one possible implementation, the acquisition module is specifically used to acquire a third operation instruction generated by an operation performed on the interactive interface, the third operation instruction being used to instruct the display of a pre-recorded voice file; and to acquire the reference voice based on the user's selection operation of the pre-recorded voice file.

[0028] In one possible implementation, the acquisition module is specifically used to acquire a fourth operation instruction generated by an operation performed on the interactive interface, the fourth operation instruction including the text content input by the user; and to use the text content as the target content.

[0029] In one possible implementation, the acquisition module is specifically used to acquire a second operation instruction generated by an operation performed on the interactive interface, the second operation instruction being used to instruct the recording function to be started; to start the recording function to record the user's voice to obtain content speech; and to acquire the target content based on the content speech.

[0030] In one possible implementation, the acquisition module is specifically used to acquire the target content based on the reference speech.

[0031] In one possible implementation, when the target age is a single age value, the target voice reflects the user's timbre at that target age; when the target age is an age range, the target voice reflects the user's timbre variations within that target age range.

[0032] In one possible implementation, the processing module is specifically configured to: acquire age-independent timbre features based on the reference speech; acquire reconstructed timbre features based on the age-independent timbre features and the target age, wherein the reconstructed timbre features are associated with the target age; and generate the target speech based on the target content and the reconstructed timbre features.

[0033] In one possible implementation, the processing module is specifically used to obtain an estimated age by using a pre-trained age extraction model of the reference speech input; to obtain an estimated timbre feature by using a pre-trained timbre extraction model of the reference speech input; and to decouple the estimated timbre feature from the estimated age to obtain the age-independent timbre feature.

[0034] In one possible implementation, the processing module is specifically used to input the age-independent timbre features and the target age into a pre-trained timbre generation model to obtain the reconstructed timbre features.

[0035] In one possible implementation, the processing module is specifically used to input the target content and the reconstructed timbre features into a pre-trained speech generation model to obtain the target speech.

[0036] Thirdly, embodiments of this application provide an electronic device, including: one or more processors; a memory for storing one or more programs; and when the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of the first aspects above.

[0037] Fourthly, embodiments of this application provide a computer-readable storage medium including a computer program, which, when executed on a computer, causes the computer to perform the method described in any one of the first aspects above.

[0038] Fifthly, embodiments of this application provide a computer program product, the computer program product including computer program code, which, when run on a computer, is used to perform the method described in any one of the first aspects above. Attached Figure Description

[0039] Figure 1 A schematic diagram of the structure of the electronic device 100 is shown;

[0040] Figure 2 This is a software structure block diagram of the electronic device 100 according to an embodiment of this application;

[0041] Figure 3 A framework diagram of the voice system provided in the embodiments of this application;

[0042] Figure 4 A flowchart of process 400 of the age-based voice generation method provided in this application embodiment;

[0043] Figure 5a and Figure 5b An example interface for recording voice;

[0044] Figure 5c An example interface for selecting historical audio;

[0045] Figure 5d An example interface for selecting the target age;

[0046] Figure 5e An example interface for selecting the target age;

[0047] Figure 5f An example interface for selecting the target age;

[0048] Figure 5g and Figure 5h An exemplary interface for playing the target audio;

[0049] Figures 6a-6c This is a schematic diagram illustrating the training of the model;

[0050] Figure 7 A schematic block diagram of an apparatus 700 according to an embodiment of this application is shown. Detailed Implementation

[0051] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0052] In this article, the term "and / or" is merely a description of the relationship between related objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone.

[0053] The terms "first" and "second," etc., used in the specification and claims of this application are used to distinguish different objects, not to describe a specific order of objects. For example, "first target object" and "second target object," etc., are used to distinguish different target objects, not to describe a specific order of target objects.

[0054] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.

[0055] In the description of the embodiments in this application, unless otherwise stated, "multiple" means two or more. For example, multiple processing units means two or more processing units; multiple systems means two or more systems.

[0056] Before describing the technical solutions of the embodiments of this application, the electronic devices of the embodiments of this application will first be described in conjunction with the accompanying drawings. Figure 1 A schematic diagram of the structure of electronic device 100 is shown. Electronic device 100 can be any one or more of the following: mobile phone, tablet computer, laptop computer, smart screen, wearable device, augmented reality (AR) / virtual reality (VR) device, etc.

[0057] It should be understood that, Figure 1 The electronic device 100 shown is merely an example of an electronic device, and the electronic device 100 may have more or fewer components than those shown in the figure, may combine two or more components, or may have different component configurations. Figure 1 The various components shown can be implemented in hardware, software, or a combination of hardware and software, including one or more signal processing and / or application-specific integrated circuits.

[0058] Electronic device 100 may include: processor 110, external memory interface 120, internal memory 121, universal serial bus (USB) interface 130, charging management module 140, power management module 141, battery 142, antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, sensor module 180, button 190, motor 191, indicator 192, camera 193, display screen 194, and subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an accelerometer sensor 180E, a distance sensor 180F, a proximity sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0059] Processor 110 may include one or more processing units, such as: application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, memory, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU), etc. Different processing units may be independent devices or integrated into one or more processors.

[0060] The controller can be the nerve center and command center of the electronic device 100. The controller can generate operation control signals according to the instruction opcode and timing signals to complete the control of instruction fetching and execution.

[0061] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.

[0062] In some embodiments, the processor 110 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0063] The I2C interface is a bidirectional synchronous serial bus, including a serial data line (SDA) and a serial clock line (SCL). In some embodiments, the processor 110 may include multiple I2C buses. The processor 110 can couple to the touch sensor 180K, charger, flash, camera 193, etc., through different I2C bus interfaces. For example, the processor 110 can couple to the touch sensor 180K through the I2C interface, enabling the processor 110 and the touch sensor 180K to communicate through the I2C bus interface, thereby realizing the touch function of the electronic device 100.

[0064] The I2S interface can be used for audio communication. In some embodiments, the processor 110 may include multiple I2S buses. The processor 110 can be coupled to the audio module 170 via the I2S bus to enable communication between the processor 110 and the audio module 170. In some embodiments, the audio module 170 can transmit audio signals to the wireless communication module 160 via the I2S interface to enable the function of answering phone calls through a Bluetooth headset.

[0065] The PCM interface can also be used for audio communication, sampling, quantizing, and encoding analog signals. In some embodiments, the audio module 170 and the wireless communication module 160 can be coupled via the PCM bus interface. In some embodiments, the audio module 170 can also transmit audio signals to the wireless communication module 160 via the PCM interface, enabling the function of answering phone calls through a Bluetooth headset. Both the I2S interface and the PCM interface can be used for audio communication.

[0066] The UART interface is a universal serial data bus used for asynchronous communication. This bus can be a bidirectional communication bus. It converts the data to be transmitted between serial and parallel communication. In some embodiments, the UART interface is typically used to connect the processor 110 and the wireless communication module 160. For example, the processor 110 communicates with the Bluetooth module in the wireless communication module 160 via the UART interface to implement Bluetooth functionality. In some embodiments, the audio module 170 can transmit audio signals to the wireless communication module 160 via the UART interface to enable music playback through Bluetooth headphones.

[0067] The MIPI interface can be used to connect the processor 110 to peripheral devices such as the display screen 194 and the camera 193. The MIPI interface includes a camera serial interface (CSI) and a display serial interface (DSI). In some embodiments, the processor 110 and the camera 193 communicate via the CSI interface to enable the electronic device 100 to capture images. The processor 110 and the display screen 194 communicate via the DSI interface to enable the electronic device 100 to display images.

[0068] The GPIO interface can be configured via software. It can be configured as a control signal or a data signal. In some embodiments, the GPIO interface can be used to connect the processor 110 to a camera 193, a display screen 194, a wireless communication module 160, an audio module 170, a sensor module 180, etc. The GPIO interface can also be configured as an I2C interface, an I2S interface, a UART interface, a MIPI interface, etc.

[0069] USB port 130 is a USB standard compliant interface, specifically a Mini USB port, Micro USB port, USB Type-C port, etc. USB port 130 can be used to connect a charger to charge electronic device 100, and can also be used for data transfer between electronic device 100 and peripheral devices. It can also be used to connect headphones for audio playback. This interface can also be used to connect other electronic devices, such as AR devices.

[0070] It is understood that the interface connection relationships between the modules illustrated in the embodiments of this application are merely illustrative and do not constitute a structural limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may also employ different interface connection methods or combinations of multiple interface connection methods as described in the above embodiments.

[0071] The charging management module 140 receives charging input from a charger. The charger can be a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 140 receives charging input from the wired charger via the USB interface 130. In some wireless charging embodiments, the charging management module 140 receives wireless charging input via the wireless charging coil of the electronic device 100. While charging the battery 142, the charging management module 140 can also supply power to the electronic device via the power management module 141.

[0072] The power management module 141 connects the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140, providing power to the processor 110, internal memory 121, external memory, display screen 194, camera 193, and wireless communication module 160, etc. The power management module 141 can also monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage current, impedance). In some other embodiments, the power management module 141 may also be located within the processor 110. In other embodiments, the power management module 141 and the charging management module 140 may be located in the same device.

[0073] The wireless communication function of electronic device 100 can be realized through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor and baseband processor, etc.

[0074] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 100 can be used to cover one or more communication frequency bands. Different antennas can also be multiplexed to improve antenna utilization. For example, antenna 1 can be multiplexed as a diversity antenna for a wireless local area network. In some other embodiments, the antennas can be used in conjunction with tuning switches.

[0075] The mobile communication module 150 can provide solutions for wireless communication, including 2G / 3G / 4G / 5G, applied to the electronic device 100. The mobile communication module 150 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves via antenna 1, and perform filtering, amplification, and other processing on the received electromagnetic waves before transmitting them to a modem processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via antenna 1. In some embodiments, at least some functional modules of the mobile communication module 150 may be housed in the processor 110. In some embodiments, at least some functional modules of the mobile communication module 150 and at least some modules of the processor 110 may be housed in the same device.

[0076] The modem processor may include a modulator and a demodulator. The modulator modulates the low-frequency baseband signal to be transmitted into a mid-to-high frequency signal. The demodulator demodulates the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After processing by the baseband processor, the low-frequency baseband signal is transmitted to the application processor. The application processor outputs sound signals through an audio device (not limited to speaker 170A, receiver 170B, etc.) or displays images or videos through the display screen 194. In some embodiments, the modem processor may be a separate device. In other embodiments, the modem processor may be independent of the processor 110 and may be housed in the same device as the mobile communication module 150 or other functional modules.

[0077] The wireless communication module 160 can provide solutions for wireless communication applications on the electronic device 100, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via antenna 2, performs frequency modulation and filtering of the electromagnetic wave signals, and sends the processed signal to processor 110. The wireless communication module 160 can also receive signals to be transmitted from processor 110, perform frequency modulation and amplification, and convert them into electromagnetic waves for radiation via antenna 2.

[0078] In some embodiments, antenna 1 of electronic device 100 is coupled to mobile communication module 150, and antenna 2 is coupled to wireless communication module 160, enabling electronic device 100 to communicate with networks and other devices via wireless communication technology. The wireless communication technology may include Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Time Division Code Division Multiple Access (TD-SCDMA), Long Term Evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technologies, etc. The GNSS may include the Global Positioning System (GPS), the Global Navigation Satellite System (GLONASS), the BeiDou Navigation Satellite System (BDS), the Quasi-Zenith Satellite System (QZSS), and / or satellite-based augmentation systems (SBAS).

[0079] Electronic device 100 implements display functions through a GPU, a display screen 194, and an application processor. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. Processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.

[0080] Display screen 194 is used to display images, videos, etc. Display screen 194 includes a display panel. The display panel may be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a miniature LED, a microLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, electronic device 100 may include one or N displays 194, where N is a positive integer greater than 1.

[0081] Electronic device 100 can perform shooting functions through ISP, camera 193, video codec, GPU, display 194 and application processor.

[0082] The ISP (Image Signal Processor) is used to process data fed back from the camera 193. For example, when taking a picture, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing, transforming it into an image visible to the naked eye. The ISP can also perform algorithmic optimization of image noise, brightness, and skin tone. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera 193.

[0083] Camera 193 is used to capture still images or videos. An object is projected onto a photosensitive element by generating an optical image through the lens. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then passed to an ISP for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into image signals in standard RGB, YUV, or other formats. In some embodiments, the electronic device 100 may include one or N cameras 193, where N is a positive integer greater than 1.

[0084] Digital signal processors (DSPs) are used to process digital signals. Besides digital image signals, they can also process other digital signals. For example, when electronic device 100 selects a frequency, the DSP can perform Fourier transforms on the frequency energy.

[0085] Video codecs are used to compress or decompress digital video. Electronic device 100 may support one or more video codecs. Thus, electronic device 100 can play or record videos in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.

[0086] An NPU (Neural Processing Unit) is a computational processor for neural networks (NNs). By borrowing the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it can rapidly process input information and continuously learn on its own. NPUs enable intelligent cognitive applications in electronic devices, such as image recognition, facial recognition, speech recognition, and text understanding.

[0087] The external storage interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external storage interface 120 to perform data storage functions. For example, music, video, and other files can be saved on the external memory card.

[0088] Internal memory 121 can be used to store computer executable program code, which includes instructions. Processor 110 executes various functional applications and data processing of electronic device 100 by running the instructions stored in internal memory 121. Internal memory 121 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback, image playback, etc.), etc. The data storage area may store data created during the use of electronic device 100 (such as audio data, phonebook, etc.). Furthermore, internal memory 121 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc.

[0089] Electronic device 100 can implement audio functions, such as music playback and recording, through audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, and application processor.

[0090] The audio module 170 is used to convert digital audio information into analog audio signals for output, and also to convert analog audio input into digital audio signals. The audio module 170 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 170 may be located in the processor 110, or some functional modules of the audio module 170 may be located in the processor 110.

[0091] The speaker 170A, also known as a "loudspeaker," is used to convert audio electrical signals into sound signals. The electronic device 100 can listen to music or make hands-free calls through the speaker 170A.

[0092] The receiver 170B, also known as the "earpiece," is used to convert audio electrical signals into sound signals. When the electronic device 100 answers a telephone call or voice message, the receiver 170B can be brought close to the ear to listen to the voice.

[0093] Microphone 170C, also known as a "microphone" or "voice transducer," is used to convert sound signals into electrical signals. When making a phone call or sending a voice message, the user can speak by bringing their mouth close to microphone 170C, inputting the sound signal into microphone 170C. Electronic device 100 may have at least one microphone 170C. In some embodiments, electronic device 100 may have two microphones 170C, which, in addition to collecting sound signals, can also perform noise reduction. In other embodiments, electronic device 100 may also have three, four, or more microphones 170C, which can collect sound signals, reduce noise, identify the sound source, and perform directional recording, etc.

[0094] The 170D headphone jack is used to connect wired headphones. The 170D headphone jack can be a USB 130 interface or a 3.5mm Open Mobile Terminal Platform (OMTP) standard interface, a CTIA (Cellular Telecommunications Industry Association of the USA) standard interface.

[0095] Pressure sensor 180A is used to sense pressure signals and convert them into electrical signals. In some embodiments, pressure sensor 180A can be disposed on display screen 194. There are many types of pressure sensors 180A, such as resistive pressure sensors, inductive pressure sensors, and capacitive pressure sensors. A capacitive pressure sensor may include at least two parallel plates with conductive material. When force is applied to pressure sensor 180A, the capacitance between the electrodes changes. Electronic device 100 determines the pressure intensity based on the change in capacitance. When a touch operation is applied to display screen 194, electronic device 100 detects the intensity of the touch operation based on pressure sensor 180A. Electronic device 100 can also calculate the touch position based on the detection signal from pressure sensor 180A. In some embodiments, touch operations applied to the same touch position but with different touch operation intensities can correspond to different operation commands. For example, when a touch operation with an intensity less than a first pressure threshold is applied to the SMS application icon, a command to view an SMS is executed. When a touch operation with an intensity greater than or equal to the first pressure threshold is applied to the SMS application icon, a command to create a new SMS is executed.

[0096] The gyroscope sensor 180B can be used to determine the motion attitude of the electronic device 100. In some embodiments, the gyroscope sensor 180B can determine the angular velocity of the electronic device 100 about three axes (i.e., the x, y, and z axes). The gyroscope sensor 180B can be used for image stabilization. For example, when the shutter is pressed, the gyroscope sensor 180B detects the angle of the shake of the electronic device 100, calculates the distance that the lens module needs to compensate based on the angle, and allows the lens to counteract the shake of the electronic device 100 by moving in the opposite direction, thus achieving image stabilization. The gyroscope sensor 180B can also be used in navigation and motion-sensing game scenarios.

[0097] The barometric pressure sensor 180C is used to measure air pressure. In some embodiments, the electronic device 100 calculates altitude using the air pressure value measured by the barometric pressure sensor 180C to assist in positioning and navigation.

[0098] The magnetic sensor 180D includes a Hall sensor. The electronic device 100 can use the magnetic sensor 180D to detect the opening and closing of the flip cover. In some embodiments, when the electronic device 100 is a flip phone, the electronic device 100 can detect the opening and closing of the flip cover using the magnetic sensor 180D. Then, based on the detected opening and closing state of the cover or the flip cover, features such as automatic flip unlocking can be set.

[0099] The 180E accelerometer can detect the magnitude of acceleration of electronic device 100 in various directions (typically three axes). When electronic device 100 is stationary, it can detect the magnitude and direction of gravity. It can also be used to identify the posture of electronic devices and applied to applications such as screen orientation switching and pedometers.

[0100] A distance sensor 180F is used to measure distance. Electronic device 100 can measure distance via infrared or laser. In some embodiments, during a shooting scene, electronic device 100 can utilize the distance sensor 180F to measure distance for rapid focusing.

[0101] The proximity sensor 180G may include, for example, a light-emitting diode (LED) and a light detector, such as a photodiode. The LED may be an infrared LED. The electronic device 100 emits infrared light outward through the LED. The electronic device 100 uses the photodiode to detect infrared reflected light from nearby objects. When sufficient reflected light is detected, it can be determined that there is an object near the electronic device 100. When insufficient reflected light is detected, the electronic device 100 can determine that there is no object near the electronic device 100. The electronic device 100 may use the proximity sensor 180G to detect when a user holds the electronic device 100 close to their ear for a call, so as to automatically turn off the screen to save power. The proximity sensor 180G can also be used in holster mode and pocket mode for automatic unlocking and locking of the screen.

[0102] The ambient light sensor 180L is used to sense the brightness of ambient light. The electronic device 100 can adaptively adjust the brightness of the display screen 194 based on the sensed ambient light brightness. The ambient light sensor 180L can also be used to automatically adjust the white balance when taking pictures. The ambient light sensor 180L can also work with the proximity sensor 180G to detect whether the electronic device 100 is in a pocket to prevent accidental touches.

[0103] The fingerprint sensor 180H is used to collect fingerprints. The electronic device 100 can utilize the characteristics of the collected fingerprints to achieve fingerprint unlocking, accessing application locks, taking photos with fingerprints, answering calls with fingerprints, etc.

[0104] Temperature sensor 180J is used to detect temperature. In some embodiments, electronic device 100 uses the temperature detected by temperature sensor 180J to execute a temperature handling strategy. For example, when the temperature reported by temperature sensor 180J exceeds a threshold, electronic device 100 performs thermal protection by reducing the performance of a processor located near temperature sensor 180J to reduce power consumption. In other embodiments, when the temperature is below another threshold, electronic device 100 heats battery 142 to prevent abnormal shutdown of electronic device 100 due to low temperature. In still other embodiments, when the temperature is below yet another threshold, electronic device 100 boosts the output voltage of battery 142 to prevent abnormal shutdown due to low temperature.

[0105] Touch sensor 180K, also known as a "touch panel," can be located on display screen 194. The touch sensor 180K and display screen 194 together form a touchscreen, also known as a "touch screen." Touch sensor 180K detects touch operations applied to or near it. The touch sensor can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through display screen 194. In other embodiments, touch sensor 180K may also be located on the surface of electronic device 100, in a different position than display screen 194.

[0106] The bone conduction sensor 180M can acquire vibration signals. In some embodiments, the bone conduction sensor 180M can acquire vibration signals from the vibrating bone segments of the human vocal cords. The bone conduction sensor 180M can also contact the human pulse to receive blood pressure signals. In some embodiments, the bone conduction sensor 180M can also be incorporated into headphones to form bone conduction headphones. The audio module 170 can parse the voice signals from the vibrating bone segments of the vocal cords acquired by the bone conduction sensor 180M to realize voice functionality. The application processor can parse heart rate information from the blood pressure signals acquired by the bone conduction sensor 180M to realize heart rate detection functionality.

[0107] Buttons 190 include a power button, volume buttons, etc. Buttons 190 can be mechanical buttons or touch-sensitive buttons. Electronic device 100 can receive button input and generate key signal inputs related to user settings and function control of electronic device 100.

[0108] Motor 191 can generate vibration alerts. Motor 191 can be used for incoming call vibration alerts or for touch vibration feedback. For example, different vibration feedback effects can correspond to touch operations performed on different applications (such as taking photos, playing audio, etc.). Motor 191 can also correspond to different vibration feedback effects for touch operations performed on different areas of the display screen 194. Different application scenarios (such as time reminders, receiving messages, alarm clocks, games, etc.) can also correspond to different vibration feedback effects. The touch vibration feedback effect can also be customized.

[0109] Indicator 192 can be an indicator light, used to indicate charging status, power changes, or to indicate messages, missed calls, notifications, etc.

[0110] The SIM card interface 195 is used to connect a SIM card. The SIM card can be inserted into or removed from the SIM card interface 195 to make contact with and separate from the electronic device 100. The electronic device 100 can support one or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface 195 can support Nano SIM cards, Micro SIM cards, SIM cards, etc. Multiple cards can be inserted into the same SIM card interface 195 simultaneously. The multiple cards can be of the same or different types. The SIM card interface 195 is also compatible with different types of SIM cards. The SIM card interface 195 is also compatible with external memory cards. The electronic device 100 interacts with the network through the SIM card to realize functions such as calls and data communication. In some embodiments, the electronic device 100 uses an eSIM, i.e., an embedded SIM card. The eSIM card can be embedded in the electronic device 100 and cannot be separated from the electronic device 100.

[0111] The software system of electronic device 100 can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture. This application embodiment uses the layered architecture Android system as an example to exemplify the software structure of electronic device 100.

[0112] Figure 2 This is a software structure block diagram of the electronic device 100 according to an embodiment of this application.

[0113] The layered architecture of the electronic device 100 divides the software into several layers, each with a clear role and division of labor. Layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom: the application layer, the application framework layer, the Android runtime and system libraries, and the kernel layer.

[0114] The application layer can include a series of application packages.

[0115] like Figure 2 As shown, the application package may include applications such as camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, and SMS.

[0116] The application framework layer provides application programming interfaces (APIs) and a programming framework for applications in the application layer. The application framework layer includes some predefined functions.

[0117] like Figure 2 As shown, the application framework layer may include a window manager, content provider, view system, phone manager, resource manager, notification manager, etc.

[0118] The window manager is used to manage windowed applications. It can retrieve screen size, determine the presence of a status bar, lock the screen, and capture screenshots, among other things.

[0119] Content providers store and retrieve data, making that data accessible to applications. This data may include videos, images, audio, made and received phone calls, browsing history and bookmarks, phone books, etc.

[0120] A view system includes visual controls, such as controls for displaying text and controls for displaying images. View systems can be used to build applications. A display interface can consist of one or more views. For example, a display interface including a text notification icon could include views for displaying text and views for displaying images.

[0121] The phone manager is used to provide communication functions for electronic device 100. For example, it manages call status (including connection and disconnection).

[0122] The file explorer provides applications with various resources, such as localized strings, icons, images, layout files, video files, and more.

[0123] The notification manager allows applications to display notifications in the status bar. These notifications can be used to deliver informational messages and can disappear automatically after a short pause, requiring no user interaction. For example, the notification manager can be used to notify users of completed downloads or message alerts. The notification manager can also display notifications as icons or scrolling text in the top status bar, such as notifications from background applications, or as dialog boxes on the screen. Examples include displaying text messages in the status bar, emitting sounds, vibrating electronic devices, and flashing indicator lights.

[0124] The Android Runtime consists of core libraries and a virtual machine. The Android runtime is responsible for scheduling and managing the Android system.

[0125] The core library consists of two parts: one part is the functionalities that need to be called by the Java language, and the other part is the Android core library.

[0126] The application layer and application framework layer run in a virtual machine. The virtual machine executes the Java files of the application layer and application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.

[0127] System libraries can include multiple functional modules. For example: surface manager, media libraries, 3D graphics processing libraries (e.g., OpenGL ES), 2D graphics engines (e.g., SGL), etc.

[0128] The Surface Manager is used to manage the display subsystem and provides the blending of 2D and 3D layers for multiple applications.

[0129] The media library supports playback and recording of various common audio and video formats, as well as still image files. It supports multiple audio and video encoding formats, such as MPEG4, H.264, MP3, AAC, AMR, JPG, and PNG.

[0130] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.

[0131] A 2D graphics engine is a graphics engine for 2D drawing.

[0132] The kernel layer is the layer between hardware and software. The kernel layer contains at least the display driver, camera driver, audio driver, and sensor driver.

[0133] It should be understood that, Figure 2 The components included in the illustrated software structure do not constitute a limitation on the software structure of the electronic device 100. In other embodiments of this application, the software structure of the electronic device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements.

[0134] This application provides an age-based voice generation method, which can be applied to the aforementioned electronic device 100 in the form of applications, mini-programs, etc., such as mobile phones, tablets, etc., to provide users with an audio entertainment experience.

[0135] Figure 3 A framework diagram of the voice system provided in the embodiments of this application, such as Figure 3 As shown, the speech system includes a timbre generation module, a content encoding module, and a speech generation module, wherein...

[0136] The timbre generation module takes a reference speech and a target age as input and outputs the reconstructed timbre features of the speaker at the target age. The timbre generation module includes a timbre extraction module, an age extraction module, an age decoupling module, and a timbre reconstruction module.

[0137] The content encoding module takes content as input, which can be text or speech, and outputs the target content, which can be represented in the form of content features. The content encoding module includes an audio encoding module and a text encoding module.

[0138] The input to the speech generation module is the output of the timbre generation module and the content encoding module, that is, to reconstruct the timbre features and the target content, and the output is the target speech, which presents the target content using the speaker's timbre at the target age.

[0139] The functions of the aforementioned modules will be explained in the following description in conjunction with the method implementation.

[0140] Figure 4 A flowchart of process 400 of the age-based voice generation method provided in this application embodiment. Process 400 can be executed by the aforementioned electronic device 100. Process 400 is described as a series of steps or operations, and it should be understood that process 400 can be executed in various orders and / or occur simultaneously, and is not limited to... Figure 4 The execution order is shown. Process 400 may include:

[0141] Step 401: Obtain reference audio. The reference audio is recorded when the user sets their age.

[0142] In one possible implementation, the electronic device can acquire a second operation instruction generated by an action performed on the interactive interface. This second operation instruction instructs the activation of a recording function to record the user's voice and obtain reference speech. The reference speech can be newly recorded speech by the user.

[0143] Figure 5a and Figure 5b An exemplary interface for recording voice, such as Figure 5a As shown, after opening the application, the user enters the voice selection interface, which includes two controls: "New Voice" and "Historical Voice". When the user clicks the "New Voice" control, it means that the user selects to record a new voice as a reference. Figure 5bAs shown, after clicking the "New Voice" control, the user enters the recording interface, which includes a record button 501, a play button 502, and an OK button 503. When the user presses and holds the record button 501, the electronic device starts recording. The user can speak directly into the microphone, connect a headset to the electronic device and speak into it, or connect a microphone to the electronic device and speak into it, and so on. After the user releases the record button 501, the electronic device ends recording and generates a voice clip from the previously recorded sound. After recording, the user can click the play button 502 to play back the previously recorded voice clip. If the user is not satisfied with the voice clip, they can press and hold the record button 501 again to record again; if the user is satisfied with the voice clip, they can click the OK button 503, and the voice clip will be saved and used as a reference voice clip. It can be seen that the newly recorded voice clip reflects the user's voice at their current age.

[0144] In one possible implementation, the electronic device can acquire a third operation instruction generated by an action performed on the interactive interface. This third operation instruction is used to instruct the display of a pre-recorded voice file, and a reference voice is obtained based on the user's selection of the pre-recorded voice file. The reference voice may be historical voice recordings previously made by the user.

[0145] When the user clicks Figure 5a When the "Historical Voice" control is displayed on the interface, it indicates that the user has selected a segment of historical voice as a reference. Figure 5c An example interface for selecting historical audio, such as Figure 5c As shown, upon entering the voice library interface, multiple voice files are displayed. These voices were all recorded by the user previously, possibly recently or at a different age. Therefore, historical voice recordings do not necessarily reflect the user's voice at their current age. Users can click on a voice file in the voice library, and the voice displayed in the clicked file will serve as a reference.

[0146] As mentioned above, the set age can refer to the user's current age (i.e., the age at the moment the user generates the speech using the method provided in this application embodiment), or it can be the user's age at the time the selected historical speech was recorded. For example, if the user is currently 30 years old, and the newly recorded speech is used as the reference speech, then the set age corresponding to that reference speech is 30 years old; if the historical speech is used as the reference speech, then the set age is the user's age when the historical speech was recorded (e.g., 20 years old).

[0147] Step 402: Obtain the first operation instruction generated by the operation applied to the interactive interface, wherein the first operation instruction indicates the target age.

[0148] The interactive interface provides controls that users can interact with. User actions on these controls generate commands within the electronic device, which then performs corresponding processing based on these commands. Therefore, users can select a target age group by manipulating the controls on the interactive interface.

[0149] In one possible implementation, user actions on the interactive interface include clicking on controls. These controls include an "age increase" control and an "age decrease" control, where the "age increase" control corresponds to a preset age range greater than a set age, and the "age decrease" control corresponds to a preset age range less than a set age.

[0150] Figure 5d An example interface for selecting a target age, such as Figure 5d As shown, the method of this application embodiment can provide an entertainment experience of "My Life", which includes an age increase control 504 and an age decrease control 505 on the interactive interface of the experience.

[0151] When a user clicks the "Age Increase" control 504, it corresponds to a preset age range greater than the set age. Therefore, the user's action indicates they wish to hear their voice at a future age, which can be a preset age range. For example, if the user is currently 30 years old and the preset age range is 30 years (30-60 years old), clicking the "Age Increase" control 504 will generate a target voice reflecting the vocal changes from 30 to 60 years old. Similarly, when a user clicks the "Age Decrease" control 505, it corresponds to a preset age range less than the set age. Therefore, the user's action indicates they wish to hear their voice at a previous age, which can also be a preset age range. For example, if the user is currently 30 years old and the preset age range is 30 years (0-30 years old), clicking the "Age Decrease" control 505 will generate a target voice reflecting the vocal changes from 0 to 30 years old.

[0152] In one possible implementation, user actions on the interactive interface include clicking on controls. These controls include an age increase control and an age decrease control, where the age increase control corresponds to a step size for increasing age, and the age decrease control corresponds to a step size for decreasing age.

[0153] Figure 5e An example interface for selecting a target age, such as Figure 5e As shown, the method of this application embodiment can provide a "voice-customized" entertainment experience, which includes an age increase control 506 and an age decrease control 507 on the interactive interface of the experience.

[0154] When a user clicks the age increment control 506, which corresponds to a step size for increasing age, the user's action indicates that they want to hear their voice at a specific age or several ages in the future, with each age being an integer multiple (e.g., 1 to n) of the aforementioned step size from the user's current age. For example, if the user is currently 30 years old and the age increment step size is set to 5 years, clicking the age increment control 506 once will generate a target voice reflecting their 35-year-old voice; clicking it again will generate a target voice reflecting their 40-year-old voice, and so on. Similarly, clicking the age decrement control 507 corresponds to a step size for decreasing age, indicating that the user wants to hear their voice at a specific age or several ages in the past, with each age being an integer multiple (e.g., 1 to n) of the aforementioned step size from the user's current age. For example, if a user is currently 30 years old and the age reduction step is set to 5 years, clicking the age reduction control 507 once will generate a target voice that reflects the user's 25-year-old voice. Clicking the age reduction control 507 again will generate a target voice that reflects the user's 20-year-old voice, and so on.

[0155] In one possible implementation, user actions on the interface include dragging a slider. The position of the slider corresponds to the target age.

[0156] Figure 5f An example interface for selecting a target age, such as Figure 5f As shown, the method of this application embodiment can provide a "voice-customized" entertainment experience, which includes a slider 508 on the interactive interface of the experience, and the position of the slider 508 corresponds to the target age.

[0157] Users can drag slider 508 within a set age range to set the target age to any value within that range. For example, if the set age range is 0-80 years old, dragging slider 508 to 30 years old will generate a target voice reflecting the user's 30-year-old voice; dragging slider 508 to 60 years old will generate a target voice reflecting the user's 60-year-old voice. Similarly, if the set age range is 0-80 years old, dragging slider 508 to 30 years old and marking it, then dragging slider 508 to 60 years old and marking it again, will generate a target voice reflecting the changes in voice reflecting the user's 30-60 years old voice.

[0158] It should be noted that, in addition to the methods mentioned above, other methods can also be used to obtain the target age in the embodiments of this application, and no specific limitations are made thereto.

[0159] Step 403: Obtain the target content.

[0160] The target content is the information conveyed by the target speech.

[0161] In one possible implementation, the electronic device can acquire a fourth operation instruction generated by an operation performed on the interactive interface, the fourth operation instruction including text content input by the user, and using the text content as the target content.

[0162] like Figure 5e and Figure 5f As shown, the method in this application embodiment can provide a "voice customization" entertainment experience, and the interactive interface of this experience also includes an "input text" control. After the user clicks the "input text" control, a simulated keyboard pops up, and the user can use the simulated keyboard to input text. This text will then be the content presented in the target voice. For example, if the user inputs the three words "I love you", the target voice will read "I love you" using the user's voice at the target age.

[0163] In one possible implementation, the electronic device can acquire a second operation instruction generated by an operation performed on the interactive interface. This second operation instruction is used to instruct the recording function to be started, the recording function to record the user's voice to obtain content speech, and the target content to be obtained based on the content speech.

[0164] like Figure 5e and Figure 5f As shown, the method in this application embodiment can provide a "voice-customized" entertainment experience, and the interactive interface of this experience also includes an "input voice" control. After the user clicks the "input voice" control, the recording function is started, and the following can be displayed: Figure 5b The interface shown allows the user to record a new voice message (i.e., content voice). The electronic device performs speech recognition on this message to obtain the textual meaning within the content voice, which is the content to be presented in the target voice. For example, if the user records the content voice message as "I love you," the electronic device will perform speech recognition on this content voice message to obtain the meaning of "I love you," and the target voice will read "I love you" in the tone appropriate for the user at a target age.

[0165] Optionally, after clicking the "Enter Voice" control, users can be allowed to input their voice through various means, such as... Figure 5c The interface shown selects a historical voice as the content voice. This historical voice can be the user's own voice or the voice of another user. The key is to identify the target content from the content voice.

[0166] In one possible implementation, the electronic device can acquire the target content based on reference speech.

[0167] The content speech can also be the reference speech obtained in step 401, that is, the content presented by the reference speech is presented using the user's voice at the target age. Therefore, the electronic device also needs to perform speech recognition on the reference speech to obtain the content expressed by the reference speech.

[0168] It should be understood that the embodiments of this application may also use other methods to obtain the target content, and the carrier may be text, audio, video, etc., without any specific limitation.

[0169] This application embodiment can automatically select the activation method based on the different carriers of the content. Figure 3 The audio encoding module or text encoding module in the framework diagram encodes the content into target content that can be used in subsequent steps.

[0170] Step 404: Generate target speech based on target content, reference speech and first operation command. The target speech presents the target content using the user's voice at the target age.

[0171] In this embodiment of the application, age-independent timbre features can be obtained from the reference speech obtained in step 401, and then reconstructed timbre features can be obtained from the age-independent timbre features and the target age obtained in step 402. The reconstructed timbre features are associated with the target age. Then, the target speech can be generated from the target content obtained in step 403 and the reconstructed timbre features.

[0172] In the above process, a pre-trained age extraction model based on the reference speech input can be used to obtain an estimated age, and a pre-trained timbre extraction model based on the reference speech input can be used to obtain estimated timbre features. The estimated age and estimated timbre features are then decoupled to obtain age-independent timbre features. These age-independent timbre features and the target age are then input into a pre-trained timbre generation model to obtain reconstructed timbre features. These reconstructed timbre features are correlated with the target age and reflect the user's timbre at that target age.

[0173] The reconstructed timbre features and target content are input into a pre-trained speech generation model to obtain the target speech. The user's timbre at the target age plus the target content equals the target speech. As mentioned above, the target speech uses the user's timbre at the target age to present the target content. For example, if the user is currently 30 years old, and the newly recorded reference speech uses the timbre of a 30-year-old to read "I love you," and the target age is 80 years old, while the target content is still "I love you," then the processed target speech will use the timbre of an 80-year-old to say "I love you." As another example, if the user is currently 30 years old, and the newly recorded reference speech uses the timbre of a 30-year-old to read "I love you," and the target age is 0-60 years old, while the target content is "The Ballad of Mulan," then the processed target speech will use the timbre of a 0-60-year-old to read "The Ballad of Mulan," reflecting the timbre changes from 0 to 60 years old during the reading.

[0174] Age extraction models can correspond Figure 3 The age extraction module and timbre extraction model in the framework diagram shown can correspond to... Figure 3 The timbre extraction model in the framework diagram shown can be decoupled accordingly. Figure 3 The age decoupling module in the framework diagram shown can correspond to the timbre generation model. Figure 3 The timbre reconstruction module in the framework diagram shown corresponds to the speech generation model. Figure 3 The speech generation module is shown in the framework diagram. The training scheme for the model in this embodiment includes using a large amount of speech-text-age pair data to train an age extraction model, a timbre extraction model, a timbre generation model, and a speech generation model. After training, the parameters of these models are saved for use in step 404.

[0175] It should be understood that the process of obtaining reconstructed timbre features can be performed after step 402. That is, after obtaining the reference speech and the target age, the reconstructed timbre features can be obtained first, and then after obtaining the target content, the aforementioned reconstructed timbre features can be combined with the target content to generate the target speech. Alternatively, the reconstructed timbre features can be obtained after performing steps 401 to 403, and then the target speech can be generated based on the reconstructed timbre features and the target content. This application embodiment does not specifically limit this.

[0176] like Figures 5d to 5f The interactive interface also includes a download button 509 and a share button 510. After obtaining the target audio, the user can click the download button 509 to save the target audio to their local device, or click the share button 510 to share the target audio to a social media platform.

[0177] Figures 6a-6c This is a diagram illustrating the training of the model, such as... Figure 6a As shown, the above-mentioned models can be trained using the following methods in the embodiments of this application:

[0178] 1. Construct a multi-speaker text-to-speech (TTS) audio library, where speakers cover different age groups, although it is not required that the same speaker cover different age groups. The audio library includes speech, text, and age tags.

[0179] 2. Feature preparation stage:

[0180] a) Extract acoustic features from audio.

[0181] b) Use a pre-trained automatic speech recognition (ASR) model as a feature extraction tool to extract phonetic posterior probabilities (PPGs) from the audio as content features.

[0182] 3. Text encoding model pre-training stage:

[0183] The text encoding model accepts text as input and uses PPG extracted from the corresponding audio as target features, and is pre-trained separately.

[0184] 4. Audio coding model pre-training stage:

[0185] The audio coding model takes the acoustic features of the speech signal as input, the extracted PPG as the target output, and learns the PPG modeling ability through distillation.

[0186] 5. Age extraction model training phase:

[0187] The acoustic features of the speech signal are taken as input, and the corresponding age label is used as the prediction target. These are trained separately to obtain an age extraction model. One possible network structure is as follows: Figure 6b As shown, the acoustic features of the input speech, with dimensions [B, T, C1], are processed by a convolutional neural network to obtain hidden features with dimensions [B, T, C2]. Then, pooling is performed through a gated recurrent unit (GRU), and the final hidden state is taken as the output to obtain age features with dimensions [B, C3]. This is then passed through a linear prediction layer to obtain an age estimate with dimensions [B, 1]. The mean squared error (MSE) loss between the age estimate and the target age is calculated to obtain the age prediction loss. Gradient descent is then used to update the model parameters, resulting in the final age extraction model. In the figure, B represents the batch size, i.e., the size of the batch of data fed into training each time; T represents the time frame length of the input data; C1 is the dimension of the input acoustic features; C2 is the dimension of the hidden features; and C3 is the dimension of the age features.

[0188] 6. In the joint training phase, the entire model, except for the age extraction model, is jointly trained. The detailed steps are as follows:

[0189] a) Input speech is fed into the timbre extraction model, and the output is an estimated timbre feature. One possible network structure is similar to the age extraction model, except that the dimension of the final linear layer output is changed to [B, S], where S represents the total number of speakers. The output of the linear layer serves as the predicted probability for timbre classification, and cross-entropy loss is calculated with the timbre label. Gradient descent is then used to update the network parameters. The timbre features output by the GRU layer are used as the estimated timbre features of the model output.

[0190] b) Timbre feature estimation and target age are sequentially input into the age decoupling model and the timbre reconstruction model, respectively, and the output is the reconstructed timbre features. One possible network structure is as follows: Figure 6c As shown, timbre features are input into an age decoupling model constructed from convolutional networks to obtain age-independent timbre features. These features are then passed through a gradient inversion (GRL) layer and fed into a linear prediction layer to obtain an age estimate. The MSE loss is calculated between this estimate and the target age to obtain the age prediction loss. Simultaneously, the age-independent timbre features and the target age are concatenated and input into a timbre reconstruction model constructed from convolutional networks to output reconstructed timbre features. The MSE reconstruction loss is calculated between the reconstructed timbre features and the original input timbre features to obtain the timbre reconstruction loss.

[0191] c) Randomly select one from the input speech and input text, input the corresponding encoding model, and output the target content.

[0192] d) The target content obtained in step c) and the timbre feature estimation obtained in step a) are input into the speech generation model, and the predicted speech 1 is output.

[0193] e) The target content obtained in step c) and the reconstructed timbre features obtained in step b) are input into the speech generation model with weights shared with d), and the predicted speech 2 is output.

[0194] f) In the loss calculation stage, the MSE loss (including audio reconstruction loss T1, audio reconstruction loss T2, and audio reconstruction loss T2) is calculated pairwise for predicted speech 1, predicted speech 2, and the original input speech. The sum of the three is used as the audio reconstruction loss. The audio reconstruction loss is added to the timbre reconstruction loss and age prediction loss obtained in step b) to obtain the final loss of the model. The model parameters are then updated using gradient descent.

[0195] The above training process can achieve arbitrary combinations of timbre and age by using an explicit modeling method of age-timbre features; by using PPG as an intermediate feature for pre-training the content encoding module and a training strategy of randomly selecting input speech and text, the speech and text input can be shared for the speech generation module; by combining timbre reconstruction loss and audio reconstruction loss, high-quality output speech can be generated while ensuring timbre stability.

[0196] In step 402, if the user clicks as Figure 5d The age-increasing control 504 shown indicates that after the electronic device generates target speech corresponding to a preset age range that is older than the set age, it can display something like this. Figure 5g The interface shown. If the user clicks as shown... Figure 5d The age reduction control 505 shown indicates that after the electronic device generates target speech corresponding to a preset age range that is younger than the set age, it can display the following: Figure 5h The interface shown is described above. The "My Life" entertainment experience includes a voice display that gradually transitions from the user's current age's voice to either younger or older. Starting with the user's current age's voice, it allows the user to intuitively perceive the change in voice. When the interface displays an infant icon, it corresponds to the experience of the voice gradually decreasing in age; when the interface displays an elderly icon, it corresponds to the experience of the voice gradually decreasing in age. Users can switch between the two modes by clicking the "age increases" control 504 or the "age decreases" control 505 on the interface.

[0197] This application's embodiments decouple the timbre and age of the reference speech to obtain age-independent timbre features. Then, based on these age-independent timbre features and the target age, reconstructed timbre features are obtained. These reconstructed timbre features can then be used to express arbitrary target content, thereby enabling the reproduction of a user's timbre at any age. This method does not require collecting a multi-age voice database for an individual, reducing reliance on such databases, and preserves the user's individual timbre, significantly reducing distortion in the reproduction of timbre at different ages compared to the user's individual timbre.

[0198] It is understood that, in order to achieve the above-mentioned functions, electronic devices include hardware and / or software modules that perform the respective functions. Based on the algorithmic steps of the examples described in the embodiments disclosed herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed by hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application in conjunction with the embodiments, but such implementation should not be considered beyond the scope of this application.

[0199] In one example, Figure 7The diagram shows a schematic block diagram of an apparatus 700 according to an embodiment of the present application. The apparatus 700 may include a processor 701 and a transceiver / transceiver pin 702, and optionally, a memory 703.

[0200] The various components of device 700 are coupled together via bus 704, which includes a data bus, a power bus, a control bus, and a status signal bus. However, for clarity, all buses are referred to as bus 704 in the figure.

[0201] Optionally, the memory 703 can be used for the instructions in the foregoing method embodiments. The processor 701 can be used to execute the instructions in the memory 703, control the receive pin to receive signals, and control the transmit pin to transmit signals.

[0202] The device 700 may be an electronic device or a chip of an electronic device in the above method embodiments.

[0203] All relevant content of each step involved in the above method embodiments can be referenced from the functional description of the corresponding functional module, and will not be repeated here.

[0204] This embodiment also provides a computer storage medium storing computer instructions. When the computer instructions are executed on an electronic device, the electronic device performs the aforementioned method steps to implement the dual Wi-Fi connection method in the above embodiment.

[0205] This embodiment also provides a computer program product that, when run on a computer, causes the computer to perform the aforementioned steps to implement the dual Wi-Fi connection method in the above embodiment.

[0206] In addition, embodiments of this application also provide an apparatus, which may specifically be a chip, component or module. The apparatus may include a connected processor and a memory; wherein the memory is used to store computer execution instructions, and when the apparatus is running, the processor may execute the computer execution instructions stored in the memory to cause the chip to execute the dual Wi-Fi connection method in the above method embodiments.

[0207] In this embodiment, the electronic device, computer storage medium, computer program product or chip are all used to execute the corresponding method provided above. Therefore, the beneficial effects that can be achieved can be referred to the beneficial effects of the corresponding method provided above, and will not be repeated here.

[0208] Through the above description of the embodiments, those skilled in the art will understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0209] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another apparatus, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0210] The units described as separate components may or may not be physically separate. A component shown as a unit can be one or more physical units; that is, it can be located in one place or distributed in multiple different locations. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0211] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0212] Any content in the various embodiments of this application, as well as any content in the same embodiment, can be freely combined. Any combination of the above content is within the scope of this application.

[0213] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, in essence, or the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0214] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

[0215] The steps of the methods or algorithms described in conjunction with the embodiments of this application can be implemented in hardware or by a processor executing software instructions. The software instructions can consist of corresponding software modules, which can be stored in random access memory (RAM), flash memory, read-only memory (ROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, portable hard disks, CD-ROMs, or any other form of storage medium well known in the art. An exemplary storage medium is coupled to a processor, enabling the processor to read information from and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and storage medium can reside in an ASIC. Additionally, the ASIC can reside in a network device. Alternatively, the processor and storage medium can exist as discrete components in the network device.

[0216] Those skilled in the art will recognize that the functions described in the embodiments of this application in one or more of the above examples can be implemented using hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or code on a computer-readable medium. Computer-readable media include computer storage media and communication media, wherein communication media include any medium that facilitates the transfer of a computer program from one place to another. Storage media can be any available medium that can be accessed by a general-purpose or special-purpose computer.

[0217] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. An age-based sound generation method, characterized by, The method comprises: obtaining a reference voice, the reference voice being recorded when a user is at a set age; obtaining a first operation instruction generated by an operation on an interactive interface, the first operation instruction indicating a target age; obtaining target content; generating a target voice according to the target content, the reference voice, and the first operation instruction, the target voice presenting the target content in a timbre of the user at the target age; wherein the generating of the target voice according to the target content, the reference voice, and the first operation instruction comprises: inputting the reference voice into a pre-trained age extraction model to obtain an estimated age; inputting the reference voice into a pre-trained timbre extraction model to obtain estimated timbre features; decoupling the estimated timbre features and the estimated age to obtain age-independent timbre features; obtaining reconstructed timbre features according to the age-independent timbre features and the target age, the reconstructed timbre features being associated with the target age; generating the target voice according to the target content and the reconstructed timbre features.

2. The method of claim 1, wherein, The operation on the interactive interface comprises a click operation on a control on the interactive interface; The control comprises an age-increasing control and an age-decreasing control, wherein the age-increasing control corresponds to a preset age interval greater than the set age, and the age-decreasing control corresponds to a preset age interval less than the set age. Alternatively, the control comprises an age-increasing control and an age-decreasing control, wherein the age-increasing control corresponds to an age-increasing step, and the age-decreasing control corresponds to an age-decreasing step.

3. The method of claim 1, wherein, The operation on the interactive interface comprises a drag operation on a slider on the interactive interface; The position of the slider corresponds to the target age.

4. The method according to any one of claims 1 to 3, characterized in that, The obtaining of the reference voice comprises: obtaining a second operation instruction generated by an operation on the interactive interface, the second operation instruction being used to indicate starting of a recording function; starting the recording function to record a voice of the user to obtain the reference voice.

5. The method according to any one of claims 1-3, characterized in that, The obtaining of the reference voice comprises: obtaining a third operation instruction generated by an operation on the interactive interface, the third operation instruction being used to indicate display of a pre-recorded voice file; obtaining the reference voice based on a selection operation of the user on the pre-recorded voice file.

6. The method according to any one of claims 1-3, characterized in that, The obtaining of the target content comprises: obtaining a fourth operation instruction generated by an operation on the interactive interface, the fourth operation instruction comprising text content input by the user; taking the text content as the target content.

7. The method according to any one of claims 1-3, characterized in that, The obtaining of the target content comprises: obtaining a second operation instruction generated by an operation on the interactive interface, the second operation instruction being used to indicate starting of a recording function; starting the recording function to record a voice of the user to obtain a content voice; obtaining the target content according to the content voice.

8. The method of any one of claims 1-3, wherein, The obtaining of the target content comprises: obtaining the target content according to the reference voice.

9. The method of any one of claims 1-3, wherein, When the target age is an age value, the target voice embodies a timbre of the user at the target age; When the target age is an age interval, the target voice embodies a timbre change of the user in the target age interval.

10. The method of claim 1, wherein, The obtaining of the reconstructed timbre feature according to the age-independent timbre feature and the target age comprises: inputting the age-independent timbre feature and the target age into a pre-trained timbre generation model to obtain the reconstructed timbre feature.

11. The method according to any one of claims 1-3, 10, characterized in that, The generating of the target voice according to the target content and the reconstructed timbre feature comprises: inputting the target content and the reconstructed timbre feature into a pre-trained voice generation model to obtain the target voice.

12. An age-based sound generating device, characterized by, comprise: The obtaining module is configured to obtain a reference voice, the reference voice being recorded at a set age of a user. The obtaining module is configured to obtain a first operation instruction generated by an operation on an interactive interface, the first operation instruction indicating a target age. The obtaining module is configured to obtain target content. The processing module is configured to generate a target voice according to the target content, the reference voice and the first operation instruction, the target voice presenting the target content with a timbre of the user at the target age. The processing module is specifically configured to input the reference voice into a pre-trained age extraction model to obtain an estimated age, input the reference voice into a pre-trained timbre extraction model to obtain an estimated timbre feature, and decouple the estimated timbre feature and the estimated age to obtain an age-independent timbre feature. The obtaining of the reconstructed timbre feature according to the age-independent timbre feature and the target age comprises:

13. The apparatus of claim 12, wherein, The operation on the interactive interface comprises a click operation on a control on the interactive interface; the control comprises an age-increasing control and an age-decreasing control, wherein the age-increasing control corresponds to a preset age interval greater than the set age, and the age-decreasing control corresponds to a preset age interval smaller than the set age; or the control comprises an age-increasing control and an age-decreasing control, wherein the age-increasing control corresponds to an age-increasing step, and the age-decreasing control corresponds to an age-decreasing step.

14. The apparatus of claim 12, wherein, The operation on the interactive interface comprises a drag operation on a slider on the interactive interface; a position of the slider corresponds to the target age.

15. The apparatus of any one of claims 12-14, wherein, The obtaining module is specifically configured to obtain a second operation instruction generated by an operation on the interactive interface, the second operation instruction being used to instruct to start a recording function; and the starting of the recording function is used to record a sound of the user to obtain the reference voice.

16. The apparatus of any one of claims 12-14, wherein, The obtaining module is specifically configured to obtain a third operation instruction generated by an operation on the interactive interface, the third operation instruction being used to instruct to display a pre-recorded voice file; and the obtaining of the reference voice is based on a selection operation of the user on the pre-recorded voice file.

17. The apparatus of any one of claims 12-14, wherein, The acquisition module is specifically configured to acquire a fourth operation instruction generated by an operation on the interactive interface, the fourth operation instruction comprising text content input by the user; and take the text content as the target content.

18. The apparatus of any one of claims 12-14, wherein, The acquisition module is specifically configured to acquire a second operation instruction generated by an operation on the interactive interface, the second operation instruction being used to instruct to start a recording function; start the recording function to record a voice of the user to obtain content voice; and acquire the target content according to the content voice.

19. The apparatus of any one of claims 12-14, wherein, The acquisition module is specifically configured to acquire the target content according to the reference voice.

20. The apparatus of any one of claims 12-14, wherein, When the target age is an age value, the target voice embodies a timbre of the user at the target age; and when the target age is an age interval, the target voice embodies a timbre change of the user in the target age interval.

21. The apparatus of claim 12, wherein, The processing module is specifically configured to input the age-independent timbre feature and the target age into a pre-trained timbre generation model to obtain the reconstructed timbre feature.

22. The apparatus of any one of claims 12-14, 21, wherein, The processing module is specifically configured to input the target content and the reconstructed timbre feature into a pre-trained voice generation model to obtain the target voice.

23. An electronic device, comprising: The computer program product comprises computer program code for performing the method of any one of claims 1-11 when the computer program code runs on a computer. The computer program product comprises computer program code for performing the method of any one of claims 1-11 when the computer program code runs on a computer. The computer program product comprises computer program code for performing the method of any one of claims 1-11 when the computer program code runs on a computer. ​ 24. A computer-readable storage medium, characterized in that, ​ 25. A computer program product, characterised in that, ​

Citation Information

Patent Citations

  • User tone based method and device for voice synthesis

    CN108847215A

  • Method, device, equipment and medium for generating audio

    CN112652292A

  • Audio processing method and device and readable storage medium

    CN113113033A