Emotional processing method of input information and electronic device

CN116705072BActive Publication Date: 2026-09-22HONOR DEVICE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310257753.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-08-31
Publication Date
2026-09-22
Estimated Expiration
2041-08-31

AI Technical Summary

Technical Problem

可以看出,目前的语音处理技术仅能够改变语音的音色,造成语音在情感表达方面较为单一,导致语音情感丰富性较差

Benefits of technology

[0047]可以理解地,上述提供的第二方面及其任一种可能的设计方式所述的电子设备,第三方面所述的计算机存储介质,第四方面所述的计算机程序产品所能达到的有益效果,可参考第一方面及其任一种可能的设计方式中的有益效果,此处不再赘述。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116705072B_ABST
    Figure CN116705072B_ABST
Patent Text Reader

Abstract

The application discloses an emotional processing method of input information and an electronic device, relates to the technical field of terminals, and can make the same input information expressed in different emotions and produce different hearing effects, thereby enriching the emotional expression of voice. The method comprises the following steps: an electronic device displays an input interface; the input interface is used for receiving input information input by a user, the input information comprises voice information or text information, and the emotion of the input information comprises a first emotion; the electronic device determines a target emotion type; and the electronic device performs emotional processing on the input information according to the target emotion type, so as to obtain target voice; the emotion of the target voice comprises a second emotion, the content of the target voice is the same as that of the input information, and the first emotion is different from the second emotion.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This is a divisional application. The original application number is 202111017308.3, and the original application date is August 31, 2021. The entire contents of the original application are incorporated herein by reference. Technical Field

[0002] This application relates to the field of terminal technology, and in particular to an emotion processing method and electronic device for input information. Background Technology

[0003] With the development of smart electronic devices, voice processing technology has also made great strides and has been widely applied in users' lives. For example, voice processing technology can be widely used in various scenarios such as short video shooting, live streaming, and text-to-speech. In these scenarios, voice processing technology can enhance the user's voice to improve the enjoyment of their life and meet their diverse needs.

[0004] However, current voice enhancement functions mainly focus on timbre equalization and speech noise reduction. Timbre equalization alters the energy distribution of speech frequencies, producing a duller or brighter sound. It can be seen that current speech processing technology can only change the timbre of the speech, resulting in a relatively simplistic expression of emotion and a lack of emotional richness. Summary of the Invention

[0005] This application provides an emotion processing method and electronic device for input information, which enables the same input information to be expressed with different emotions, producing different auditory effects, thereby enriching the emotional expression of speech.

[0006] In a first aspect, embodiments of this application provide a method for processing the emotion of input information. The method includes: an electronic device displaying an input interface; the input interface receiving input information from a user, the input information including voice information or text information; the emotion of the input information including a first emotion; the electronic device determining a target emotion type; the electronic device performing emotion processing on the input information according to the target emotion type to obtain target speech; the emotion of the target speech including a second emotion, the content of the target speech being the same as the content of the input information, and the first emotion being different from the second emotion.

[0007] Using this scheme, electronic devices can process the input information according to the determined target emotion type to obtain the target speech. Since the emotion of the input information includes the first emotion and the emotion of the target speech includes the second emotion, and the content of the input information is the same as the content of the target speech, but the first emotion and the second emotion are different, the same input information can be expressed with different emotions, producing different auditory effects, thereby enriching the emotional expression of speech.

[0008] In one possible design of the first aspect, the electronic device displays an input interface, including: the electronic device displaying an input interface in response to a user activating the camera; the input interface being a preview interface before the electronic device records video; or, the input interface being an interface during video recording; or, the input interface being an interface after the electronic device has finished recording video.

[0009] In this design approach, the input interface serves as the interface for electronic devices to record videos. This means that the electronic devices can enhance the emotional expression of the recorded videos through voice, allowing users to hear different emotional voices when the videos are played back, thus enriching the emotional expression of the video recordings.

[0010] In one possible design of the first aspect, the electronic device includes a gallery application, a recording application, and a notepad application; the electronic device displays an input interface, including: the electronic device displays an input interface in response to a user opening any video file in the gallery application; or, the electronic device displays an input interface in response to a user opening any audio file in the recording application; or, the electronic device displays an input interface in response to a user opening any text file in the notepad application.

[0011] In this design approach, electronic devices can enhance the emotional tone of video files played in a gallery application; or, electronic devices can enhance the emotional tone of audio files played in a recording application; or, electronic devices can imbue text files in a notepad application with emotion during the conversion process, making the converted audio more emotionally charged and further enriching the expression of emotional voice effects.

[0012] In one possible design of the first aspect, the input interface includes multiple voice emotion controls, each voice emotion control corresponding to an emotion type; determining the target emotion type includes: the electronic device determining the target emotion type in response to the user's operation on at least one of the multiple voice emotion controls.

[0013] In this design approach, users can select at least one voice emotion control from multiple voice emotion controls, allowing them to customize the emotion type they need for the input information. When the electronic device plays the target voice, the voice the user hears is the voice with the customized emotion type, thus improving the user experience.

[0014] In one possible design of the first aspect, the input interface includes a captured image; determining the target emotion type includes: the electronic device identifying the style of the captured image and automatically matching the target emotion type corresponding to the style of the captured image.

[0015] In this design approach, electronic devices can also recognize the style of the captured image and automatically match the target emotional type that corresponds to the style of the captured image, further improving the user experience.

[0016] In one possible design approach of the first aspect, the input information is voice information; determining the target emotion type includes: the electronic device recognizing the emotion in the voice information and automatically matching the emotion type corresponding to the emotion in the voice information to determine the target emotion type.

[0017] In this design approach, when the input information is voice information, the electronic device can also recognize the emotions in the voice information and automatically match the emotion type corresponding to the emotion in the voice information to determine the target emotion type, thereby further improving the user experience.

[0018] In one possible design approach of the first aspect, the input information is text information; determining the target sentiment type includes: the electronic device recognizing the semantics of the text information and automatically matching the sentiment type corresponding to the semantics of the text information to determine the target sentiment type.

[0019] In this design approach, when the input information is text, the electronic device can also recognize the semantics of the text and automatically match the sentiment type corresponding to the semantics of the text to determine the target sentiment type, thereby further improving the user experience.

[0020] In one possible design of the first aspect, the electronic device performs emotion processing on the input information according to the target emotion type to obtain the target speech, including: the electronic device inputs the input information into a speech emotion model to obtain the target speech; the speech emotion model is used to modify the emotion of the input information according to the target emotion type.

[0021] In this design approach, the electronic device inputs input information into a speech emotion model to obtain the target speech. Since the speech emotion model is used to modify the emotion of the input information according to the target emotion type, the emotion of the target speech output by the electronic device is different from the emotion of the input information, thus enriching the emotional expression of the speech.

[0022] In one possible design of the first aspect, the electronic device inputs the input information into a speech emotion model to obtain target speech, including: the electronic device encoding the input information to obtain time-frequency features of the input information; the encoding process includes frame segmentation and Fourier transform; the frame segmentation process is used to divide the input information into multiple speech frames, and the time-frequency features are used to describe the relationship between the frequency and amplitude of each speech frame over time; the electronic device inputs the time-frequency features of the input information into the speech emotion model to obtain time-frequency features of the target speech; the electronic device decodes the time-frequency features of the target speech and performs speech synthesis processing to obtain target speech; the decoding process includes inverse Fourier transform and time-domain waveform superposition.

[0023] In this design approach, the electronic device first encodes the input information, then inputs the processed information into the speech emotion model to obtain the time-frequency features of the target speech. Finally, the time-frequency features of the target speech are decrypted and processed by speech synthesis to obtain the target speech, thereby enabling more accurate expression of emotion in the target speech.

[0024] In one possible design of the first aspect, the method further includes: an electronic device acquiring a speech emotion dataset; the speech emotion dataset includes multiple emotional speech pieces; each of the multiple emotional speech pieces corresponds to a different emotion type; for each emotional speech piece in the speech emotion dataset, the electronic device performs feature extraction processing on the emotional speech piece to obtain the time-frequency features of the emotional speech piece; the electronic device inputs the time-frequency features of the emotional speech piece into a neural network model for emotion training to obtain a speech emotion model.

[0025] In this design approach, electronic devices can train neural network models on language emotion datasets to improve the maturity of the trained speech emotion model, which is beneficial for improving the accuracy of modifying the emotion of input information.

[0026] In one possible design approach of the first aspect, the speech emotion model includes a first model and a second model; the first model is used to indicate the mapping relationship between emotional speech and emotion type; the second model is used to modify the emotion of the input information.

[0027] In this design approach, since the voice emotion model includes a first model and a second model; the first model is used to indicate the mapping relationship between emotional speech and emotion type; and the second model is used to modify the emotion of the input information, after the electronic device inputs the input information into the voice emotion model, the first model first determines the emotion of the input information, and then the second model converts the emotion of the input information, thereby improving the accuracy of modifying the emotion of the input information.

[0028] In one possible design of the first aspect, when the electronic device plays an audio video, the electronic device outputs the target speech; the audio video is a video video; or, the audio video is an audio file.

[0029] In this design approach, when an electronic device plays audio, it outputs the target speech, thereby creating different auditory experiences for the user and enriching the emotional expression of the speech.

[0030] In one possible design of the first aspect, when the electronic device plays audio, the interface of the electronic device displays instruction information; the instruction information is used to indicate the emotion type corresponding to the target speech.

[0031] In this design approach, when the electronic device plays audio, the interface also displays instruction information. Since the instruction information is used to indicate the emotional type of the target speech, the user can see the emotional type of the target speech at this time from the interface of the electronic device, thus improving the user experience.

[0032] Secondly, this application provides an electronic device including a memory, a display screen, one or more cameras, and one or more processors. The memory, display screen, cameras, and processors are coupled together. The cameras are used to capture images, the display screen is used to display images captured by the cameras or images generated by the processor, and the memory stores computer program code including computer instructions. When the computer instructions are executed by the processor, the electronic device performs the following steps: the electronic device displays an input interface; the input interface is used to receive input information from a user, including voice information or text information; the emotion of the input information includes a first emotion; the electronic device determines a target emotion type; the electronic device performs emotion processing on the input information according to the target emotion type to obtain target speech; the emotion of the target speech includes a second emotion, the content of the target speech is the same as the content of the input information, and the first emotion and the second emotion are different.

[0033] In one possible design of the second aspect, when the computer instruction is executed by the processor, the electronic device specifically performs the following steps: the electronic device displays an input interface in response to the user's operation of activating the camera; the input interface is a preview interface before the electronic device records video; or, the input interface is an interface during the recording of video by the electronic device; or, the input interface is an interface after the electronic device has finished recording video.

[0034] In one possible design of the second aspect, the electronic device includes a gallery application, a recording application, and a notepad application; when the computer instruction is executed by the processor, the electronic device specifically performs the following steps: the electronic device displays an input interface in response to a user opening any video file in the gallery application; or, the electronic device displays an input interface in response to a user opening any audio file in the recording application; or, the electronic device displays an input interface in response to a user opening any text file in the notepad application.

[0035] In one possible design of the second aspect, the input interface includes multiple voice emotion controls, each voice emotion control corresponding to an emotion type; when the computer instruction is executed by the processor, the electronic device specifically performs the following steps: the electronic device, in response to the user's operation on at least one of the multiple voice emotion controls, determines the target emotion type.

[0036] In one possible design of the second aspect, the input interface includes a captured image; when the computer instruction is executed by the processor, the electronic device specifically performs the following steps: the electronic device identifies the style of the captured image and automatically matches a target emotion type corresponding to the style of the captured image.

[0037] In one possible design of the second aspect, the input information is voice information; when the computer instruction is executed by the processor, the electronic device specifically performs the following steps: the electronic device identifies the emotion in the voice information and automatically matches the emotion type corresponding to the emotion in the voice information to determine the target emotion type.

[0038] In one possible design approach of the second aspect, the input information is text information; when the computer instruction is executed by the processor, the electronic device specifically performs the following steps: the electronic device recognizes the semantics of the text information and automatically matches the sentiment type corresponding to the semantics of the text information to determine the target sentiment type.

[0039] In one possible design of the second aspect, when the computer instruction is executed by the processor, the electronic device specifically performs the following steps: the electronic device inputs input information into a speech emotion model to obtain target speech; the speech emotion model is used to modify the emotion of the input information according to the target emotion type.

[0040] In one possible design of the second aspect, when the computer instruction is executed by the processor, the electronic device specifically performs the following steps: the electronic device encodes the input information to obtain the time-frequency features of the input information; the encoding process includes framing and Fourier transform; the framing process is used to divide the input information into multiple speech frames, and the time-frequency features are used to describe the relationship between the frequency and amplitude of each speech frame over time; the electronic device inputs the time-frequency features of the input information into a speech emotion model to obtain the time-frequency features of the target speech; the electronic device decodes and synthesizes the time-frequency features of the target speech to obtain the target speech; the decoding process includes inverse Fourier transform and time-domain waveform superposition.

[0041] In one possible design of the second aspect, when the computer instruction is executed by the processor, the electronic device further performs the following steps: the electronic device acquires a speech emotion dataset; the speech emotion dataset includes multiple emotional speech pieces; each of the multiple emotional speech pieces corresponds to a different emotion type; for each emotional speech piece in the speech emotion dataset, the electronic device performs feature extraction processing on the emotional speech piece to obtain the time-frequency features of the emotional speech piece; the electronic device inputs the time-frequency features of the emotional speech piece into a neural network model for emotion training to obtain a speech emotion model.

[0042] In one possible design approach of the second aspect, the speech emotion model includes a first model and a second model; the first model is used to indicate the mapping relationship between emotional speech and emotion type; the second model is used to modify the emotion of the input information.

[0043] In one possible design of the second aspect, when the computer instruction is executed by the processor, the electronic device also performs the following steps: when the electronic device plays an audio picture, the electronic device outputs the target speech; the audio picture is a video picture; or, the audio picture is an audio file.

[0044] In one possible design of the second aspect, when the computer instruction is executed by the processor, the electronic device also performs the following steps: when the electronic device plays audio, the interface of the electronic device displays instruction information; the instruction information is used to indicate the emotion type corresponding to the target speech.

[0045] Thirdly, this application provides a computer-readable storage medium including computer instructions that, when executed on an electronic device, cause the electronic device to perform the method described in the first aspect and any possible design thereof.

[0046] Fourthly, this application provides a computer program product that, when run on a computer, causes the computer to perform the methods described in the first aspect and any possible design. The computer may be the aforementioned electronic device.

[0047] It is understood that the beneficial effects achieved by the electronic device described in the second aspect and any possible design thereof, the computer storage medium described in the third aspect, and the computer program product described in the fourth aspect can be referred to in light of the beneficial effects in the first aspect and any possible design thereof, and will not be repeated here. Attached Figure Description

[0048] Figure 1 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application;

[0049] Figure 2 A schematic diagram of the software structure of an electronic device provided in an embodiment of this application;

[0050] Figure 3a A schematic diagram of an input interface provided in an embodiment of this application. Figure 1 ;

[0051] Figure 3b A schematic diagram of an input interface provided in an embodiment of this application. Figure 2 ;

[0052] Figure 3c Schematic diagram three of an input interface provided in an embodiment of this application;

[0053] Figure 4a Schematic diagram four of an input interface provided for an embodiment of this application;

[0054] Figure 4b Schematic diagram five of an input interface provided for an embodiment of this application;

[0055] Figure 4c A schematic diagram of the interface of an electronic device for playing video, provided as an embodiment of this application;

[0056] Figure 5a A schematic diagram of an input interface provided in an embodiment of this application. Figure 6 ;

[0057] Figure 5b A schematic diagram of an input interface provided in an embodiment of this application. Figure 7 ;

[0058] Figure 5c A schematic diagram of an input interface provided in an embodiment of this application. Figure 8 ;

[0059] Figure 6 A flowchart illustrating an emotion processing method for input information provided in this application embodiment. Figure 1 ;

[0060] Figure 7 A flowchart illustrating an emotion processing method for input information provided in this application embodiment. Figure 2 ;

[0061] Figure 8 A schematic diagram of a real-time spectrum provided in an embodiment of this application;

[0062] Figure 9 A schematic diagram illustrating a process for training a speech emotion model, provided in an embodiment of this application;

[0063] Figure 10 A schematic diagram of a voice emotion dataset provided in an embodiment of this application;

[0064] Figure 11 This is a schematic diagram of a chip system provided in an embodiment of this application. Detailed Implementation

[0065] The technical solutions of the embodiments of this application will be described below with reference to the accompanying drawings. In the description of the embodiments of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B; the term "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone.

[0066] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of embodiments of this application, unless otherwise stated, "a plurality of" means two or more.

[0067] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.

[0068] To address the problems in the background art, embodiments of this application provide an emotion processing method for input information, applied in an electronic device. Through this method, the electronic device can perform emotion processing on input information, enabling the same input information to be expressed with different emotions, producing different auditory effects, thereby enriching the emotional expression of speech. The input information may include speech or text.

[0069] Specifically, the electronic device inputs input information into a pre-trained neural network model, which then performs emotion processing on the input information to obtain target speech. The emotion type of the target speech differs from that of the input information; alternatively, the input information may not contain emotion, while the target speech may contain emotion. Emotion types may include, for example, neutral, angry, disgusted, fearful, joyful, sad, and surprised. In some embodiments, when the input information is speech to be processed, the electronic device inputs the speech to be processed into a pre-trained neural network model, which then performs emotion transformation on the speech to obtain target speech. For example, the emotion type of the speech to be processed may be angry, while the emotion type of the target speech may be joyful. In other embodiments, when the input information is text to be processed, the electronic device inputs the text to be processed into a pre-trained neural network model, which then performs emotion processing on the text to obtain target speech. For example, the text to be processed may not possess emotion, while the emotion type of the obtained target speech may be joyful.

[0070] The emotion processing method for input information provided in this application can be applied to electronic devices including intelligent voice devices, such as voice assistants, smart speakers, smartphones, tablets, computers, wearable electronic devices, and intelligent robots. In these devices, the desired voice emotion can be output. Several possible application scenarios for emotion processing of input information are described below.

[0071] Application Scenario 1: Text-to-Speech

[0072] In text-to-speech applications, the text content can be combined to make the speech heard by the user more emotional, thus enriching the sound quality while ensuring a high accuracy rate.

[0073] Application Scenario 2: Smartphone Voice Interaction

[0074] In smartphone voice interaction scenarios, the voice of a smartphone's voice assistant is no longer a simple machine voice, but a user-customized voice with emotions. For example, a user can customize the voice assistant's voice to be cheerful, so the user will hear a cheerful and emotional voice when communicating with the voice assistant.

[0075] Application Scenario 3: Short Video Shooting

[0076] In short video shooting scenarios, for example, users' voices can be customized into personalized voices with specific emotions such as anger or joy, thereby beautifying the user's voice during the video shooting process and adapting it to the visual style of different shooting themes.

[0077] To better understand the solutions of the embodiments of this application, the implementation methods of the embodiments of this application will be described in detail below with reference to the accompanying drawings.

[0078] Please refer to Figure 1 This is a schematic diagram of the structure of an electronic device 100 provided in an embodiment of this application. Figure 1 As shown, the electronic device 100 may include: a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, buttons 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc.

[0079] The aforementioned sensor module 180 may include sensors such as a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an accelerometer sensor 180E, a distance sensor 180F, a proximity sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, and a bone conduction sensor 180M.

[0080] It is understood that the structure illustrated in this embodiment does not constitute a specific limitation on the electronic device 100. In other embodiments, the electronic device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0081] Processor 110 may include one or more processing units, such as: application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, memory, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU), etc. Different processing units may be independent devices or integrated into one or more processors.

[0082] The controller can be the nerve center and command center of the electronic device 100. The controller can generate operation control signals according to the instruction opcode and timing signals to complete the control of fetching and executing instructions.

[0083] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.

[0084] In some embodiments, the processor 110 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0085] It is understood that the interface connection relationships between the modules illustrated in this embodiment are merely illustrative and do not constitute a structural limitation on the electronic device 100. In other embodiments, the electronic device 100 may also employ different interface connection methods or combinations of multiple interface connection methods as described in the above embodiments.

[0086] Electronic device 100 implements display functions through a GPU, a display screen 194, and an application processor. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. Processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.

[0087] Display screen 194 is used to display images, videos, etc. Display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Mini-LED, a Micro-LED, a Micro-OLED, a quantum dot light-emitting diode (QLED), etc.

[0088] Electronic device 100 can perform shooting functions through ISP, camera 193, video codec, GPU, display 194 and application processor.

[0089] The ISP (Image Signal Processor) is used to process data fed back from the camera 193. For example, when taking a picture, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing, transforming it into an image visible to the naked eye. The ISP can also perform algorithmic optimization of image noise, brightness, and skin tone. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera 193.

[0090] Camera 193 is used to capture still images or videos. An object is projected onto a photosensitive element by generating an optical image through the lens. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then passed to an ISP for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into image signals in standard RGB, YUV, or other formats. In some embodiments, the electronic device 100 may include one or N cameras 193, where N is a positive integer greater than 1.

[0091] Digital signal processors (DSPs) are used to process digital signals. Besides digital image signals, they can also process other digital signals. For example, when electronic device 100 selects a frequency, the DSP can perform Fourier transforms on the frequency energy.

[0092] Video codecs are used to compress or decompress digital video. Electronic device 100 may support one or more video codecs. Thus, electronic device 100 can play or record videos in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.

[0093] An NPU (Neural Processing Unit) is a computational processor for neural networks (NNs). By borrowing the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it can rapidly process input information and continuously learn on its own. NPUs enable intelligent cognitive applications in electronic devices, such as image recognition, facial recognition, speech recognition, and text understanding.

[0094] The external storage interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external storage interface 120 to perform data storage functions. For example, music, video, and other files can be saved on the external memory card.

[0095] Internal memory 121 can be used to store computer executable program code, which includes instructions. Processor 110 executes various functional applications and data processing of electronic device 100 by running the instructions stored in internal memory 121. For example, in this embodiment, processor 110 can execute instructions stored in internal memory 121, which may include a program storage area and a data storage area.

[0096] The program storage area can store the operating system, at least one application program required for a function (such as sound playback, image playback, etc.). The data storage area can store data created during the use of the electronic device 100 (such as audio data, phonebook, etc.). Furthermore, the internal memory 121 can include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc.

[0097] Electronic device 100 can implement audio functions, such as music playback and recording, through audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, and application processor.

[0098] Audio module 170 is used to convert digital audio information into analog audio signal output, and also to convert analog audio input into digital audio signal. Audio module 170 can also be used for encoding and decoding audio signals. In some embodiments, audio module 170 may be located in processor 110, or some functional modules of audio module 170 may be located in processor 110. Speaker 170A, also called a "loudspeaker," is used to convert audio electrical signals into sound signals. Receiver 170B, also called a "handset," is used to convert audio electrical signals into sound signals. Microphone 170C, also called a "microphone" or "microphone," is used to convert sound signals into electrical signals.

[0099] The 170D headphone jack is used to connect wired headphones. The 170D headphone jack can be a USB 130 interface or a 3.5mm Open Mobile Terminal Platform (OMTP) standard interface, a CTIA (Cellular Telecommunications Industry Association of the USA) standard interface.

[0100] Buttons 190 include a power button, volume buttons, etc. Buttons 190 can be mechanical buttons or touch buttons. Motor 191 can generate vibration alerts. Motor 191 can be used for incoming call vibration alerts or for touch vibration feedback. Indicator 192 can be an indicator light, used to indicate charging status, battery level changes, messages, missed calls, notifications, etc. SIM card interface 195 is used to connect a SIM card. The SIM card can be inserted into or removed from the SIM card interface 195 to achieve contact and separation with the electronic device 100. The electronic device 100 can support one or N SIM card interfaces, where N is a positive integer greater than 1. SIM card interface 195 can support Nano SIM cards, Micro SIM cards, SIM cards, etc.

[0101] The emotion processing method for input information provided in this application embodiment can be executed, for example, by the processor 110 included in the aforementioned electronic device 100. Exemplarily, it can be implemented by a neural network processor within the processor 110. This neural network processor can carry a neural network model to implement the method described in this application embodiment.

[0102] The software system of electronic device 100 can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture. This embodiment of the invention uses the layered architecture Android system as an example to exemplify the software architecture of electronic device 100.

[0103] Figure 2 This is a software structure block diagram of an electronic device 100 according to an embodiment of this application.

[0104] A layered architecture divides software into several layers, each with a clear role and function. Layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom: the application layer (or application layer), the application framework layer (or framework layer), the Android runtime and system libraries, and the kernel layer.

[0105] The application layer can include a series of application packages.

[0106] like Figure 2 As shown, the application package may include applications such as camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, SMS and voice assistant.

[0107] The application framework layer provides application programming interfaces (APIs) and a programming framework for applications in the application layer. The application framework layer includes some predefined functions.

[0108] like Figure 2 As shown, the application framework layer may include a window manager, content manager, view system, phone manager, resource manager, notification manager, etc.

[0109] The window manager is used to manage windowed applications. It can retrieve screen size, determine the presence of a status bar, lock the screen, and capture screenshots, among other things.

[0110] Content providers store and retrieve data, making that data accessible to applications. This data may include videos, images, audio, call logs, browsing history and bookmarks, phone books, etc.

[0111] A view system includes visual controls, such as controls for displaying text and controls for displaying images. View systems can be used to build applications. A display interface can consist of one or more views. For example, a display interface including a text notification icon could include views for displaying text and views for displaying images.

[0112] The phone manager is used to provide communication functions for electronic device 100. For example, it manages call status (including connection and disconnection).

[0113] The file explorer provides applications with various resources, such as localized characters, icons, images, layout files, video files, etc.

[0114] The notification manager allows applications to display notifications in the status bar. These notifications can be used to deliver informational messages and can disappear automatically after a short pause, requiring no user interaction. For example, the notification manager can be used to notify users of completed downloads or message alerts. The notification manager can also display notifications as icons or scrolling text in the top status bar, such as notifications from background applications, or as dialog boxes on the screen. Examples include displaying text messages in the status bar, emitting sounds, vibrating electronic devices, and flashing indicator lights.

[0115] The Android Runtime consists of core libraries and a virtual machine. The Android Runtime is responsible for scheduling and managing the Android system.

[0116] The core library consists of two parts: one part is the functionalities that need to be called by the Java language, and the other part is the Android core library.

[0117] The application layer and application framework layer run in a virtual machine. The virtual machine executes the Java files of the application layer and application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.

[0118] System libraries can include multiple functional modules. For example: surface manager, media libraries, 3D graphics processing libraries (e.g., OpenGL ES), 2D graphics engines (e.g., SGL), etc.

[0119] The Surface Manager is used to manage the display subsystem and provides the blending of 2D and 3D layers for multiple applications.

[0120] The media library supports playback and recording of various common audio and video formats, as well as still image files. It supports multiple audio encoding formats, such as MPEG4, H.264, MP3, AAC, AMR, JPG, and PNG.

[0121] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.

[0122] A 2D graphics engine is a graphics engine for 2D drawing.

[0123] The kernel layer is the layer between hardware and software. The kernel layer contains at least the display driver, camera driver, audio driver, and sensor driver.

[0124] Using the aforementioned electronic device 100 as a mobile phone, the emotion processing method for input information provided in this application embodiment will be described in detail. It should be understood that the methods in the following embodiments can all be implemented in an electronic device having the above-described hardware and software structures.

[0125] The mobile phone displays an input interface to receive user input, which may include voice or text. The input includes a primary emotion. The phone then responds to the user's actions on the input interface and determines the target emotion type. Finally, the phone processes the input based on the target emotion type (e.g., changing the emotion of the input from the primary emotion to the secondary emotion) to obtain the target speech. The target speech includes the secondary emotion, and the content of the target speech is the same as the content of the input, but the primary and secondary emotions differ.

[0126] The following description, in conjunction with the accompanying drawings, details the emotion processing method for input information provided in the embodiments of this application, based on different scenarios.

[0127] In some embodiments, the emotion processing method for input information provided in this application can be applied to short video shooting scenarios. For example, a short video shooting scenario can be a scenario where a user records video using a camera application on their mobile phone. For instance, a user can record video using the camera application's video recording mode. Another example is that a user can record video using the camera application's professional mode. Yet another example is that a user can record video using the camera application's movie mode.

[0128] It should be noted that users can also use other short video shooting applications on their mobile phones to record videos. These short video shooting applications can be system applications on the phone or third-party applications (such as applications downloaded from the phone's app store or app market). This application embodiment does not impose any restrictions on this.

[0129] For example, users can input voice during video recording. When recording a food video, a user can provide an introduction or evaluation of the food. During video recording, the phone can adjust the emotion of the user's input voice to produce different emotional expressions for the food description or evaluation. In other words, the emotional type of the voice output by the phone differs from the emotional type of the user's input, resulting in different auditory effects and enriching the emotional expression of the voice.

[0130] In one possible implementation, users can choose different voice emotion types based on their subjective consciousness or personal preferences, so that different videos will have different voice emotion types, thereby satisfying users' needs for beautifying voice effects.

[0131] In some embodiments, such as Figures 3a-3c As shown, the mobile phone displays an input interface 201, which can be a preview interface before the phone takes a picture; or it can be an interface during the shooting process; or it can be an interface after the shooting is completed. For example, the input interface 201 is the interface for the phone to record video in video recording mode.

[0132] Taking the input interface 201 as an example, which is the preview interface before the mobile phone takes a picture, for reference... Figure 3a As shown, the input interface 201 includes a voice emotion template 202; the voice emotion template 202 includes multiple different voice emotion types. For example, the voice emotion template 202 includes emotion types such as neutral, angry, disgusted, fearful, cheerful, sad, and surprised. Based on this, users can select the emotion type from the voice emotion template 202 before recording a video. After the phone completes the recording, the recorded audio is obtained. The voice in this audio is the emotional voice corresponding to the emotion type selected by the user.

[0133] For example, before recording a video, the user can select one of the voice emotion templates 202 based on the style of the current preview screen, making it match the style of the current preview screen. For instance, if the preview screen's style is minimalist, the user can choose a neutral voice emotion type to match the minimalist style. Or, if the preview screen's style is exaggerated, the user can choose a cheerful voice emotion type to match the style of the current preview screen.

[0134] Taking the input interface 201 as an example of the interface used in mobile phone photography, for instance, refer to... Figure 3b As shown, the input interface 201 includes a voice emotion template 202; wherein, the voice emotion template 202 includes multiple different voice emotion types. For example, the voice emotion template 202 includes emotion types such as neutral, angry, disgusted, fearful, cheerful, sad, and surprised.

[0135] For example, a user can select one of the voice emotion templates 202 based on the style of the captured scene to match its style. For instance, if the current scene's style is minimalist, the user can choose a neutral voice emotion to match it. Or, if the scene's style is exaggerated, the user can choose a cheerful voice emotion to match it. Alternatively, if the style of the scene changes during video recording, the user can change the voice emotion accordingly. For example, if the scene changes from minimalist to exaggerated, the user can first choose a neutral voice emotion and then a cheerful voice emotion to include two different voice emotion types in the same video.

[0136] Taking input interface 201 as an example, which is the interface after the phone has finished taking a picture. See also: Figure 3c As shown, the first interface 201 includes a voice emotion template 202; wherein, the voice emotion template 202 includes multiple different voice emotion types. For example, the voice emotion template 202 includes emotion types such as neutral, angry, disgusted, fearful, cheerful, sad, and surprised.

[0137] For example, a user can select one of the voice emotion templates 202 based on the overall style of the video being filmed, making it match the overall style of the video. In some embodiments, the user can determine the style of the filmed video based on their subjective perception. For example, when the user determines that the overall style of the filmed video leans towards a minimalist style, the user can select a neutral voice emotion type to match the overall style of the filmed video. As another example, when the user determines that the overall style of the filmed video leans towards an exaggerated style, the user can select a cheerful voice emotion type to match the overall style of the filmed video. Then, the mobile phone can apply emotion transformation to the voice in the filmed video based on the voice emotion type selected by the user, and save the filmed video. The voice in the video is the emotional voice corresponding to the voice emotion type selected by the user.

[0138] In other embodiments, such as Figure 4a and Figure 4b As shown, the mobile phone displays an input interface 203, which can be a preview interface before the phone takes a picture; or, the input interface 203 can be an interface during the shooting process; or, the input interface 203 can be an interface after the shooting is completed. For example, the input interface is the interface for the mobile phone to record video in video recording mode.

[0139] For example, the input interface 203 includes a voice emotion setting item 204, which may include, for example, a "neutral" setting, an "angry" setting, a "disgusted" setting, and a "fear" setting. In this embodiment, the voice emotion setting item 204 is... Figure 4a and Figure 4b Using the scroll bar shown as an example, the method of this application embodiment is introduced.

[0140] It should be understood that the voice emotion template 202 in the above embodiments and the voice emotion setting item 204 in this embodiment are only different ways of expression, while the method for the user to select an emotion type that matches the style of the screen of the input interface 203 is the same. For example, the user can select an emotion type according to the style of the preview screen of the preview interface; or, the user can select an emotion type according to the style of the current shooting screen; or, the user can select an emotion type according to the overall style of the video after shooting. Specific examples can be found in the above embodiments, and will not be repeated here.

[0141] For example, a user can slide the scroll bar of one of the voice emotion settings 204 to select a corresponding voice emotion type that matches the style of the currently captured scene. The scroll bar's length is 0-1; for instance, when the scroll bar is at position 0, it indicates that the user has not selected that setting; correspondingly, when the scroll bar is at position 1, it indicates that the user has selected that setting. (Reference) Figure 4a As shown, for example, the scroll bar for the "neutral" setting included in the voice emotion setting 204 is moved to position 1, while the scroll bars for other settings (such as the "angry", "disgusted", and "fearful" settings) are moved to position 0, meaning that the user selects a neutral emotion type to match the style of the currently captured scene.

[0142] In some embodiments, the scroll bar 0-1 is also used to represent the intensity of the emotional tone in the voice. For example, the closer the scroll bar for a setting is to a position of 1, the stronger the emotional intensity of the corresponding emotional type; conversely, the closer the scroll bar is to a position of 0, the weaker the emotional intensity of the corresponding emotional type. This allows users to position the scroll bar at different locations between 0 and 1, enabling them to not only assign different emotional types to different voices but also to assign different emotional intensities to different emotional types, thereby further satisfying users' desire for enhanced voice effects.

[0143] Building upon this, users can further swipe through each setting in the voice emotion settings 204. The scroll bar positions for each setting are not entirely the same, meaning the intensity of the voice emotion corresponding to each setting is not entirely the same. This allows the phone's output voice emotion to include various emotion types with different intensities, further enriching the emotional expression of the voice. For example... Figure 4b As shown, voice emotion settings 204 include "Neutral," "Angry," "Disgusted," and "Fearful" settings. The scroll bar positions for each of these settings are different. After video recording, the audio will include complex emotions such as neutral, angry, disgusted, and fear at varying intensities.

[0144] In another possible implementation, the phone can automatically match the corresponding voice emotion type based on the style of different shooting scenes. This satisfies users' needs for enhancing the recording effect. The shooting scene style can be: minimalist, exaggerated, gradient, ink painting, elegant, and artistic, among others. Of course, other styles are also possible, which will not be listed here.

[0145] For example, when a phone identifies the style of the current scene as minimalist, it can automatically match a neutral emotional tone in the voice to the style of the scene. Or, when the phone identifies the style of the current scene as exaggerated, it can automatically match a cheerful emotional tone in the voice to the style of the scene. Alternatively, the phone can identify the emotion in the voice within a video and match it to the corresponding emotion type in the language. For example, the phone can detect keywords in the user's speech in real time and automatically match the corresponding emotional tone in the voice. For instance, if the phone detects keywords such as "sad" or "wronged" during recording, it might match the corresponding emotional tone in the voice as "grief."

[0146] In conjunction with any of the above embodiments, in some embodiments, after the mobile phone has recorded a video, when the user plays the video, the emotional type of the voice the user hears is the emotional type previously selected by the user. In other embodiments, when the user plays the video, indication information is displayed on the mobile phone's playback interface. This indication information is used to indicate the emotional type of the voice in the currently playing video. For example, such as... Figure 4c As shown, the indication information 205 is: the current voice emotion type is cheerful.

[0147] It should be noted that, in the embodiments of this application, Figures 3a-4c The images shown in the mobile phone video recording examples all depict people. It should be understood that mobile phone video recordings can also include landscapes, food, and other scenes; the actual recording will prevail, and these will not be listed here.

[0148] In some embodiments, the emotion processing method for input information provided in this application can also be applied to scenarios involving audio recording (hereinafter referred to as recording). The recording scenario can be a scenario where a user records audio using a recording application on a mobile phone. For example, the mobile phone displays an input interface; this input interface can be the interface before recording; or, the input interface can be the interface during the recording process; or, the input interface can also be the interface when playing back the recording.

[0149] The following example illustrates the interface used when playing a recording (also known as an audio file). For instance, as shown... Figure 5aAs shown, the input interface 206 includes a voice emotion template 207; wherein, the voice emotion template 207 includes multiple different voice emotion types. For example, the voice emotion template 207 includes emotion types such as neutral, angry, disgusted, fearful, cheerful, sad, and surprised. Based on this, the user can select the emotion type in the voice emotion template 207 before playing the recording to assign emotion to the recording; or, when the user plays the recording, the mobile phone can identify the emotion in the recording and match the emotion type corresponding to the emotion in the recording. For example, the mobile phone can detect keywords in the recording in real time and match the emotion type corresponding to the keywords to assign emotion to the recording. For example, when the keywords identified by the mobile phone include "sad" and "cry," the emotion type matched by the mobile phone for the recording can be "sad." It should be understood that when the mobile phone plays the recording, the emotion of the recording heard by the user corresponds to the emotion type of the recording.

[0150] In some embodiments, the emotion processing method for input information provided in this application can also be applied to scenarios involving playing videos (also known as video files). For example, a user opens a photo gallery app on their phone and plays a video file from that app. An example is provided where the input interface is the interface for playing the video file. Figure 5b As shown, the input interface 208 includes a voice emotion template 209; wherein, the voice emotion template 209 includes multiple different voice emotion types. For example, the voice emotion template 209 includes emotion types such as neutral, angry, disgusted, fearful, cheerful, sad, and surprised. Based on this, the user can select an emotion type from the emotion template 209 before playing the video file to assign emotion to the voice in the video file; or, the mobile phone can recognize the emotion in the video file and match the emotion type corresponding to the emotion in the video file. For example, the mobile phone can recognize keywords in the video file and match the emotion type corresponding to the keywords in the video file to assign emotion to the video segment. For example, when the keywords recognized by the mobile phone include "sad" and "cry," the emotion type matched by the mobile phone for the video segment can be "sad." It should be understood that when the mobile phone plays the video file, the emotion of the voice in the video file heard by the user is the voice corresponding to the emotion type.

[0151] In some embodiments, the emotion processing method for input information provided in this application can also be applied to text-to-speech scenarios. For example, the text (also called a text file) can be a piece of text in a mobile phone's notepad application. An example is provided using a text-to-speech input interface. For example, as... Figure 5cAs shown, the input interface 210 includes a voice emotion template 211; wherein, the voice emotion template 211 includes multiple different voice emotion types. For example, the voice emotion template 211 includes emotion types such as neutral, angry, disgusted, fearful, cheerful, sad, and surprised. Based on this, when a user wants to convert the text into speech, the user can first select an emotion type in the voice emotion template 211 to assign emotion to the converted speech; or, the mobile phone can automatically recognize the semantics of the text and match the corresponding emotion type to the text based on the semantics. After the mobile phone converts the text into speech, the emotion of the speech heard by the user is the emotion voice corresponding to the emotion type selected by the user.

[0152] It should be noted that in the above embodiments Figures 5a-5c The input interface shown can include voice emotion templates, which can also be voice emotion settings. Examples of voice emotion settings can be found in the above embodiments and will not be repeated here. This application provides a method for processing the emotion of input information, which can be applied to electronic devices. For example... Figure 6 As shown, the method may include S301-S304.

[0153] S301, Electronic devices acquire input information.

[0154] The input information includes either voice or text. Voice information can be, for example, a spoken audio recording, while text information can be, for example, a written text. The emotion expressed in the input information is designated as the primary emotion; the primary emotion corresponds to the primary emotion type. For example, the primary emotion type can include neutral, angry, disgusted, fearful, joyful, sad, and surprised emotion types. A neutral emotion type refers to an emotion type that does not possess any particular emotional connotation. For instance, when the input information is text, the primary emotion type can be neutral.

[0155] S302. The electronic device processes the input information to obtain the time-frequency characteristics of the input information.

[0156] In conjunction with the above embodiments, the input information includes voice or text. The electronic device processes the input information by first converting the input information into a voice signal, and then encoding the voice signal to obtain the time-frequency characteristics corresponding to the voice signal. Here, the voice signal refers to the signal of a voice wave (i.e., a voice waveform), and the voice signal is the information carrier of the wavelength and intensity of the voice wave.

[0157] Taking voice as the input information as an example, an electronic device can recognize the audio of the voice to obtain the voice signal corresponding to the audio. Taking text as the input information as an example, the electronic device first converts the text into voice, and then recognizes the audio of the converted voice to obtain the voice signal corresponding to the text.

[0158] It should be noted that for a detailed description of the speech signal obtained by recognizing the audio of the speech, please refer to the speech recognition technology in relevant fields, which will not be elaborated here.

[0159] In some embodiments, such as Figure 7 As shown, the encoding process may include, for example, framing and Fourier transform. For instance, framing refers to dividing the speech signal into multiple speech frames according to a pre-set frame length and frame shift, thereby obtaining a speech frame sequence corresponding to the speech signal. Each speech frame can be a speech segment, and the speech signal can be processed frame by frame.

[0160] For example, the frame length can be used to represent the duration of each speech frame, and the frame shift can be used to represent the overlap between adjacent speech frames. For instance, when the frame length is 25ms and the frame shift is 15ms, the first speech frame lasts 0–25ms, the second speech frame lasts 15–40ms, and so on, thus enabling frame-based processing of the speech signal. It should be understood that specific frame lengths and frame shifts can be set according to actual circumstances, and this application embodiment does not impose any limitations on this.

[0161] Then, the electronic device sequentially performs Fourier transform processing on each speech frame in the speech frame sequence to obtain the time-frequency characteristics of each speech frame. These time-frequency characteristics describe the relationship between the frequency and amplitude of each speech frame over time. In some embodiments, the time-frequency characteristics can be represented using a real-time spectrogram (referred to as the time-spectrum). This real-time spectrogram can be a three-primary-color pixel image (i.e., an RGB pixel image); or, it can be a time-frequency waveform. Taking an RGB pixel image as an example, such as... Figure 8 As shown, in this RGB pixel image, the horizontal axis (i.e., the X-axis) represents time, the vertical axis (i.e., the Y-axis) represents frequency, and the color depth represents amplitude.

[0162] It should be noted that the Fourier transform described in the above embodiments can be, for example, a short-time Fourier transform; or, a fast Fourier transform, and the embodiments of this application do not limit this.

[0163] S303. The electronic device inputs the time-frequency features of the input information into a pre-trained speech emotion model to obtain the time-frequency features of the target speech.

[0164] The speech emotion model is used to transform the emotion of the input information. For example, the speech emotion model can be a model built on a neural network, meaning it is a model trained on a neural network. The neural network here includes, but is not limited to, combinations, superpositions, or nesting of at least one of the following networks: convolutional neural network (CNN), recurrent neural network (RNN), long-short term memory (LSTM), bidirectional long-short term memory (BLSTM), deep convolutional neural network (DCNN), etc.

[0165] Neural networks can be composed of neural units. Taking deep neural networks as an example, specifically, deep neural networks can also be called multi-layer neural networks, which can be understood as neural networks with multiple hidden layers. Deep neural networks can be divided according to the position of different layers. For example, the internal neural network of a deep neural network can be divided into three categories: input layer, hidden layer, and output layer. Generally, the first layer of a deep neural network is the input layer, the last layer is the output layer, and the layers in between are hidden layers. The layers are fully connected, meaning that any neuron in the i-th layer is connected to any neuron in the (i+1)-th layer.

[0166] In some embodiments, the speech emotion model can be trained based on a neural network using forward propagation and backpropagation algorithms. For example, the forward propagation algorithm includes convolutional layers and pooling layers; the backpropagation algorithm includes deconvolution and depooling. In this embodiment, the backpropagation algorithm can correct the parameter data in the initial neural network model during training, thus reducing the reconstruction error loss of the neural network model. Specifically, the forward propagation algorithm generates error loss, and the parameters in the initial neural network model are updated by backpropagating the error loss information, thereby converging the error loss. The backpropagation algorithm is a backpropagation motion dominated by error loss, aiming to obtain the optimal parameters of the neural network model, such as the weight matrix. In other words, the backpropagation algorithm can optimize the weights of the neural network model.

[0167] In some embodiments, the speech emotion model includes a first model and a second model. The first model indicates the mapping relationship between emotional speech and emotion type; the second model modifies the emotion of the input information. For example, the electronic device inputs the time-frequency features of the input information into the first model to obtain the time-frequency features of the input information and the first emotion type corresponding to those time-frequency features. Then, the electronic device uses the time-frequency features of the input information and the corresponding first emotion type as input to the second model, causing the second model to output the time-frequency features of the target speech. The time-frequency features of the target speech differ from the time-frequency features of the input information; the time-frequency features of the target speech correspond to the second emotion type, while the time-frequency features of the input information correspond to the first emotion type.

[0168] For example, taking anger as the first emotion type, the electronic device inputs the input information into the first model and obtains the corresponding emotion type. That is, the electronic device knows that the emotion type of the input information is the first emotion type (anger). It should be noted that the input information to the first model is represented in the form of time-frequency features. Then, the electronic device inputs the time-frequency features of the input information into the second model. The second model can modify the time-frequency features of the input information to obtain the time-frequency features of the target speech with the second emotion type.

[0169] Taking the representation of time-frequency features using a real-time spectrogram as an example, the second model modifies the time-frequency features of the input information by modifying the energy, frequency, and time distribution of the real-time spectrogram.

[0170] It should be noted that examples illustrating the time-frequency characteristics of the target speech can be found in the above embodiments illustrating the time-frequency characteristics of the input information, and will not be repeated here. It should be understood that the time-frequency characteristics of the target speech describe the relationship between the frequency and amplitude of the speech frames over time.

[0171] S304. The electronic device processes the time-frequency characteristics of the target speech to obtain the target speech.

[0172] Among them, the emotion of the target speech is the second emotion; the second emotion corresponds to the second emotion type; the content of the input information is the same as the content of the target speech, but the first emotion and the second emotion are different, that is, the first emotion type and the second emotion type are different.

[0173] In some embodiments, the electronic device can perform decoding and speech synthesis processing on the time-frequency features of the target speech. The decoding process includes inverse Fourier transform and time-domain waveform superposition. For example, the electronic device can perform inverse Fourier transform on the frequencies, amplitudes at each frequency, and phases at each frequency in the time-frequency features of the target speech to obtain the time-domain signal (i.e., time-domain waveform) of the target speech.

[0174] As can be seen from the above embodiments, since the speech signal is segmented into frames during the encoding process, the time-domain signal of the target speech obtained in this embodiment is the time-domain signal of each speech frame. Based on this, the electronic device should also perform time-domain waveform superposition on the time-domain signals of each speech frame to obtain the target speech signal (i.e., the target speech waveform). Accordingly, the electronic device can use speech synthesis technology to synthesize the target speech signal to obtain the target speech, which is the speech finally output by the electronic device.

[0175] It should be noted that the speech synthesis technology can refer to the speech synthesis technology of related technologies, which will not be elaborated here.

[0176] In summary, in the embodiments of this application, the electronic device can perform emotional processing on the input information, so that the same input information can be expressed with different emotions, producing different auditory effects, thereby enriching the emotional expression of speech.

[0177] Taking voice input as an example, let's say a speaker utters a sentence (such as "The scenery here is beautiful"). The emotion in this sentence is neutral, meaning the speaker hasn't added any emotional color to it. After receiving the speaker's voice, the electronic device modifies its emotion and outputs a more emotionally charged version. For example, the output might be of a cheerful emotion. In other words, the speaker's voice is neutral, but after emotional processing by the electronic device, the output is cheerful, meaning the listener hears a cheerful voice.

[0178] This application also provides a method for training a model, which is used to train the speech emotion model described in the above embodiments. Figure 9 As shown, the method includes:

[0179] S401. Electronic device acquires voice emotion dataset.

[0180] The voice emotion dataset includes multiple emotional voice recordings.

[0181] In some embodiments, the electronic device can collect a user's voice to construct a voice emotion dataset. The user can be a professional voice actor. For example, the user can express different emotion types in the text included in the corpus. The text included in the corpus can be standard Mandarin Chinese text. For example, the corpus includes standard Mandarin Chinese texts such as "Have you eaten?" and "The weather is so nice today!"

[0182] Taking the text "Have you eaten?" from the corpus as an example, suppose the user expresses the text using five different emotional types. These five different emotional types include: fear, joy, surprise, sadness, and neutral.

[0183] For example, such as Figure 10 As shown, the user expresses fear in the text, resulting in the first emotional speech; joy in the text, the second emotional speech; surprise in the text, the third emotional speech; sadness in the text, the fourth emotional speech; and neutral emotion in the text, the fifth emotional speech. Correspondingly, the same method can be used for other texts in the corpus to obtain multiple emotional speeches with different emotions, thus enabling the collection of user speech data from a corpus and the construction of a speech emotion dataset.

[0184] Considering the richness of user emotional expression, users may express other emotions (such as neutral, angry, or disgust) in addition to the customized emotion (e.g., fear) when conveying a particular text; furthermore, the intensity of each emotion may vary. Therefore, after obtaining each emotional speech segment corresponding to different texts, users can evaluate the emotional intensity of each segment. For example, they can rate the emotional intensity of each emotional speech segment. This rate refers to the percentage (%) of each emotion included in the emotional speech segment. It should be noted that the rating of the emotional intensity of each emotional speech segment can be based on the user's subjective perception, for example.

[0185] Still Figure 10As shown, for example, in the first emotional voice message, the emotional intensity of "neutral" emotion accounts for 5%, "anger" emotion accounts for 10%, "disgust" emotion accounts for 15%, "fear" emotion accounts for 70%, and the emotional intensity of other emotions (such as joy, sadness, and surprise) accounts for 0%. In the second emotional voice message, the emotional intensity of "neutral" emotion accounts for 5%, "joy" emotion accounts for 70%, "sadness" emotion accounts for 10%, "surprise" emotion accounts for 15%, and the emotional intensity of other emotions (such as anger, disgust, and fear) accounts for 0%. In the third emotional voice message, the emotional intensity of "anger" emotion accounts for 5%, "neutral" emotion accounts for 10%, "joy" emotion accounts for 15%, "surprise" emotion accounts for 75%, and the emotional intensity of other emotions (such as disgust, sadness, and fear) accounts for 0%. In the fourth category of emotional voice recordings, the percentage of "neutral" emotion was 5%, "sadness" was 70%, "disgust" was 10%, "fear" was 15%, and other emotions (such as anger, joy, and surprise) accounted for 0%. In the fifth category of emotional voice recordings, the percentage of "neutral" emotion was 80%, "joy" was 5%, "sadness" was 10%, "surprise" was 5%, and other emotions (such as anger, disgust, and fear) accounted for 0%.

[0186] S402. For each emotional speech in the speech emotion dataset, the electronic device performs feature extraction processing on the emotional speech to obtain the time-frequency features of the emotional speech.

[0187] In some embodiments, feature extraction processing may include, for example, frame segmentation and Fourier transform. Furthermore, after feature extraction, the resulting time-frequency features of the emotional speech are used to describe the relationship between the frequency and amplitude of each speech frame in the emotional speech over time.

[0188] It should be noted that the examples of frame processing, Fourier transform, and time-frequency characteristics can be found in the above embodiments, and will not be repeated here.

[0189] S403. The electronic device inputs the time-frequency features of the emotional speech into the neural network model for emotion training to obtain the speech emotion model.

[0190] For example, an electronic device inputs the time-frequency features of each emotional speech into a neural network model until the model fully converges, thus obtaining a relatively mature speech emotion model. For instance, the speech emotion model is used to convert the emotion type of the input information.

[0191] It should be noted that the examples of neural network models can be found in the examples of neural networks in the above embodiments, and will not be repeated here.

[0192] In some embodiments, the speech emotion model includes a first model and a second model. The first model indicates the mapping relationship between the time-frequency features of the emotional speech and the emotion type; the second model is used to convert the emotion type of the emotional speech.

[0193] In summary, in the embodiments of this application, the electronic device can train a neural network model based on a speech emotion dataset to obtain a speech emotion model. This allows the electronic device to transform the emotion type of the input information according to the speech emotion model to obtain the target speech, that is, the speech corresponding to the emotion desired by the user, thereby enriching the emotional expression of speech and improving the user experience.

[0194] This application provides an electronic device that may include a display screen (such as a touchscreen), a camera, a memory, and one or more processors. The display screen, camera, memory, and processors are coupled. The display screen is used to display images captured by the camera or images generated by the processor, and the memory is used to store computer program code, including computer instructions. When the processor executes the computer instructions, the electronic device can perform various functions or steps performed by the mobile phone in the above method embodiments. The structure of this electronic device can be referred to... Figure 1 The structure of the electronic device 100 shown.

[0195] This application also provides a chip system, such as... Figure 11 As shown, the chip system 1800 includes at least one processor 1801 and at least one interface circuit 1802.

[0196] The processor 1801 and interface circuit 1802 described above can be interconnected via lines. For example, interface circuit 1802 can be used to receive signals from other devices (e.g., the memory of an electronic device). As another example, interface circuit 1802 can be used to send signals to other devices (e.g., processor 1801). Exemplarily, interface circuit 1802 can read instructions stored in memory and send those instructions to processor 1801. When the instructions are executed by processor 1801, the electronic device can perform the various steps performed by mobile phone 180 in the above embodiments. Of course, this chip system may also include other discrete components, and this application embodiment does not specifically limit this.

[0197] This application also provides a computer storage medium that includes computer instructions. When the computer instructions are executed on an electronic device, the electronic device causes the electronic device to perform various functions or steps performed by the mobile phone in the above method embodiments.

[0198] This application also provides a computer program product that, when run on a computer, causes the computer to perform the various functions or steps performed by the mobile phone in the above method embodiments.

[0199] Through the above description of the embodiments, those skilled in the art can clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0200] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0201] The units described as separate components may or may not be physically separate. A component shown as a unit can be one or more physical units; that is, it can be located in one place or distributed in multiple different locations. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0202] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0203] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, essentially or in other words, the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0204] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for processing the emotion of input information, characterized in that, The method includes: The electronic device receives input information from the user regarding the first emotion; the input information includes voice information or text information. The electronic device displays an input interface, which includes one or more emotion controls. Each emotion control corresponds to an emotion type, and each emotion control is used to adjust the emotion intensity of the corresponding emotion type. The electronic device, in response to a user's operation on one or more emotion controls, determines a target emotion type and emotion intensity; wherein the target emotion type includes one or more emotion types; The electronic device inputs the first emotion input information into the speech emotion model and outputs target speech including the second emotion; wherein, the speech emotion model includes a first model and a second model, the first model is used to determine the first emotion type corresponding to the input information based on the input information, and the second model is used to modify the emotion of the input information based on the first emotion type, the target emotion type and the emotion intensity; the content of the input information is the same as the content of the target speech, and the first emotion and the second emotion are different.

2. The method according to claim 1, characterized in that, The electronic device displays an input interface, including: The electronic device displays the input interface in response to the user's operation of activating the camera; the input interface is a preview interface before the electronic device shoots video; or, the input interface is an interface during the video shooting process; or, the input interface is an interface after the electronic device finishes shooting video.

3. The method according to claim 1, characterized in that, The electronic device includes a gallery application, a recording application, and a notepad application; the electronic device displays an input interface, including: The electronic device displays the input interface in response to a user opening any video file in the gallery application; or... The electronic device displays the input interface in response to a user opening any recording file in the recording application; or... The electronic device displays the input interface in response to a user opening any text file in the Notepad application.

4. The method according to any one of claims 1-3, characterized in that, The input interface includes the captured image; the method further includes: The electronic device identifies the style of the captured image and automatically matches the emotional type corresponding to the style of the captured image to determine the target emotional type.

5. The method according to claim 2 or 3, characterized in that, The input information is voice information; the method further includes: The electronic device identifies the emotions in the voice information and automatically matches the emotion type corresponding to the emotions in the voice information to determine the target emotion type.

6. The method according to claim 3, characterized in that, The input information is text information; the method further includes: The electronic device identifies the semantics of the text information and automatically matches the sentiment type corresponding to the semantics of the text information to determine the target sentiment type.

7. The method according to any one of claims 1-3 or 6, characterized in that, The method further includes: The electronic device outputs target speech including a third emotion based on the input information of the first emotion and the target emotion type; The third emotion is different from the second emotion, and the intensity of the third emotion is different from the intensity of the second emotion.

8. The method according to any one of claims 1-3 or 6, characterized in that, The target emotion types include one or more of the following: neutral, angry, disgusted, fearful, joyful, sad, and surprised.

9. The method according to any one of claims 1-3 or 6, characterized in that, The output includes the target speech with a second emotion, including: When the electronic device plays audio, the electronic device outputs target speech including the second emotion; The audio feed is either a video feed or an audio file.

10. The method according to claim 9, characterized in that, When the electronic device plays audio, the interface of the electronic device displays instruction information; the instruction information is used to indicate the emotion type corresponding to the target voice.

11. The method according to claim 1, characterized in that, The electronic device inputs the input information of the first emotion into the speech emotion model and outputs target speech including the second emotion, including: The electronic device encodes the input information to obtain the time-frequency characteristics of the input information; the encoding process includes frame segmentation and Fourier transform; the frame segmentation is used to divide the input information into multiple speech frames, and the time-frequency characteristics are used to describe the relationship between the frequency and amplitude of each speech frame over time. The electronic device inputs the time-frequency features of the input information into the speech emotion model to obtain the time-frequency features of the target speech; The electronic device performs decoding and speech synthesis processing on the time-frequency features of the target speech to obtain the target speech; the decoding process includes inverse Fourier transform and time-domain waveform superposition.

12. The method according to claim 11, characterized in that, The method further includes: The electronic device acquires a voice emotion dataset; the voice emotion dataset includes multiple emotional voices, each of which corresponds to multiple different emotions, and each of the multiple different emotions has a different emotional intensity ratio in the emotional voice; For each emotional speech in the emotional speech dataset, the electronic device performs feature extraction processing on the emotional speech to obtain the time-frequency features of the emotional speech; The electronic device inputs the time-frequency features of the emotional speech into a neural network model for emotion training, thereby obtaining the speech emotion model.

13. An electronic device, characterized in that, The electronic device includes a memory, a display screen, one or more cameras, and one or more processors; the display screen is used to display images captured by the cameras or images generated by the processors, and the memory stores computer program code, the computer program code including computer instructions, which, when executed by the processor, cause the electronic device to perform the method as described in any one of claims 1-12.

14. A computer-readable storage medium, characterized in that, Includes computer instructions that, when executed on an electronic device, cause the electronic device to perform the method as described in any one of claims 1-12.

15. A computer program product, characterized in that, The computer program product includes computer instructions that, when executed on an electronic device, cause the electronic device to perform the method as described in any one of claims 1-12.

Citation Information

Patent Citations

  • Speech emotion migration method

    CN107221344A

  • Video dubbing method and device based on voice synthesis, computer equipment and medium

    CN111031386A