Processing method, intelligent terminal and storage medium
By analyzing the differences between the learning audio and the target speech style, the problem of users being unaware of speech differences on smart terminals was solved, thus improving the personalized learning experience.
Patent Information
- Application Number
- CN202511643292.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-10
- Publication Date
- 2026-02-06
AI Technical Summary
Existing smart terminals cannot effectively analyze the differences between user voice and target voice during the voice learning process, which makes it difficult for users to clearly understand the differences and affects the learning experience.
By acquiring information about the differences between the speech style of the learning audio and the target speech style, including generating and displaying prosodic curves, virtual avatars, and text differences, it helps users correct their speech style.
Users can clearly understand the differences in voice styles, enabling a personalized learning experience and improving learning outcomes and user experience.
Smart Images

Figure CN121483296A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent interaction technology, specifically to a processing method, an intelligent terminal, and a storage medium. Background Technology
[0002] With the development of internet technology, students and other users are not only learning foreign languages or imitating specific speaking styles in the classroom, but also conducting online learning through smartphones and other smart devices.
[0003] In conceiving and implementing this application, the inventors discovered at least the following problems: Existing smart terminal processing methods mainly involve providing learning videos or audio, the speech contained in which is called the target speech. Users watch the video or audio to imitate and learn the target speech. During the imitation learning process, users repeatedly play a segment of the video or audio to correct their pronunciation. However, this method does not analyze the differences between the user's speech and the target speech, making it difficult for users to clearly perceive the differences. Although some smart terminals also collect user speech, they only analyze text errors and pronunciation errors in the user's speech, i.e., mispronounced characters and sounds, without involving analysis of differences in speech style, resulting in a poor user experience.
[0004] The preceding description is intended to provide general background information and does not necessarily constitute prior art. Summary of the Invention
[0005] To address the aforementioned technical problems, embodiments of this application provide a processing method, a smart terminal, and a storage medium that can analyze the differences between the speech style of a user's learning audio and the target speech style, enabling the user to clearly understand the differences and correct their own speech style accordingly, thereby improving the user's learning experience.
[0006] This application provides a processing method, including the following steps: S11. Obtain the learning audio; S12. Output the difference information between the speech style of the learning audio and the target speech style.
[0007] Optionally, it further includes: determining the target speech style; step S11 includes: in response to determining the target speech style, acquiring learning audio based on the target speech style.
[0008] Optionally, the method of determining the target speech style includes: selecting a first target object and using the speech style corresponding to the first target object as the target speech style.
[0009] Optionally, the method of determining the target speech style includes: analyzing the speech style of the target audio as the target speech style.
[0010] Optional methods for determining the target speech style include: Determine the language information corresponding to the first target object and / or the target audio; Based on the language information, the target audio and / or the first target object are processed to generate the target audio; In addition, the style of the target audio is taken as the target speech style.
[0011] Optionally, the method of obtaining the target audio includes: Collect the audio of the current environment as the target audio.
[0012] Optionally, the method of obtaining the target audio includes: in response to selecting a target language, using an audio example corresponding to the target language as the target audio.
[0013] Optionally, the method of obtaining the target audio includes: in response to selecting a target scene, using the audio example corresponding to the target scene as the target audio.
[0014] Optionally, the method of obtaining the target audio includes: in response to selecting a second target object, using the audio example corresponding to the second target object as the target audio.
[0015] Optionally, the method further includes: Determine the identification rules for the difference information; In response to the identification rule, step S12 is executed.
[0016] Optionally, the identification rules include: Generate a first prosodic curve corresponding to the learning audio and a second prosodic curve when the target audio is pronounced according to the target speech style, and display the difference information between the speech style of the learning audio and the target speech style through the first prosodic curve and the second prosodic curve.
[0017] Optionally, the identification rules include: generating a first virtual image corresponding to the learning audio and a second virtual image when the target audio is spoken according to the target speech style, and displaying the difference information between the speech style of the learning audio and the target speech style through the first virtual image and the second virtual image.
[0018] Optionally, the identification rules include: generating a first text corresponding to the learning audio and a second text corresponding to the target audio, and displaying the difference information between the speech style of the learning audio and the target speech style through the first text and the second text.
[0019] Optionally, the difference information includes text errors and / or pronunciation errors; The step of displaying the difference information between the speech style of the learning audio and the target speech style through the first text and the second text includes: Text errors and / or pronunciation errors corresponding to the learning audio are highlighted.
[0020] Optionally, the method further includes: Preset prompt symbols are marked in the difference information.
[0021] Optionally, marking the difference information with a preset prompt symbol includes: In the second text, emphasize the phonetic symbols.
[0022] Optionally, marking the difference information with a preset prompt symbol includes: The second prosodic curve is marked with a phonological symbol.
[0023] Optionally, the method further includes: In response to a first preset operation on the first rhythm curve or the second rhythm curve, a first position selected on the corresponding rhythm curve by the first preset operation is determined. Displays the lip movements of the second virtual character corresponding to the first position.
[0024] Optionally, the method further includes: In response to a second preset operation on the first text or the second text, a second position selected on the corresponding text by the second preset operation is determined; Displays the lip movements of the second virtual character corresponding to the second position.
[0025] This application also provides a smart terminal, including a memory and a processor, wherein the memory stores a processing program, which, when executed by the processor, implements the steps of any of the above processing methods.
[0026] This application also provides a storage medium storing a computer program, which, when executed by a processor, implements the steps of any of the processing methods described above.
[0027] As described above, the technical solution of this application includes: S11, acquiring learning audio; S12, outputting difference information between the speech style of the learning audio and the target speech style. Therefore, this application allows users to emit audio according to the target speech style, i.e., the learning audio, and compares the speech style of the learning audio with the target speech style to obtain difference information between the two styles. This allows users to clearly understand the differences and correct their lip movements, facial expressions, and other pronunciation methods related to speech style. This enables processing methods for new languages, different languages, or imitating specific speaking styles, and provides a personalized learning experience for both tutoring and free practice, thereby improving the user's experience and learning experience. Attached Figure Description
[0028] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, those skilled in the art can obtain other drawings based on these drawings without any creative effort.
[0029] Figure 1 A schematic diagram of the hardware structure of a mobile terminal to implement the various embodiments of this application; Figure 2 A communication network system architecture diagram provided for an embodiment of this application; Figure 3 A flowchart illustrating a processing method provided in the first embodiment of this application; Figure 4a A schematic diagram of an interface used to determine the target speech style for this application; Figure 4b This is a schematic diagram of an interface for obtaining target audio based on the target language in this application; Figure 5a A flowchart illustrating another processing method provided in the first embodiment of this application; Figure 5b A flowchart illustrating a processing method provided in the second embodiment of this application; Figure 6a A schematic diagram of an interface showing the first prosodic curve and the second prosodic curve for the purposes of this application; Figure 6b This application is based on Figure 6a The diagram shows the difference between the first and second prosodic curves at a given moment, with custom selections made. Figure 7 A schematic diagram of an interface for displaying difference information provided in an embodiment of this application; The realization of the objectives, functional features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. The accompanying drawings have illustrated specific embodiments of this application, which will be described in more detail below. These drawings and textual descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concepts of this application to those skilled in the art through reference to specific embodiments. Detailed Implementation
[0030] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0031] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, components, features, and elements with the same names in different embodiments of this application may have the same meaning or different meanings, the specific meaning of which must be determined by its interpretation in that specific embodiment or further in conjunction with the context of that specific embodiment.
[0032] It should be understood that although the terms first, second, third, etc., may be used herein to describe various information, such information should not be limited to these terms. These terms are used only to distinguish information of the same type from one another. For example, without departing from the scope of this document, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if," as used herein, may be interpreted as "when," "when," or "in response to determination." Furthermore, as used herein, the singular forms "a," "an," and "the" are intended to also include the plural forms unless the context indicates otherwise. It should be further understood that the terms "comprising," "including," indicate the presence of the stated feature, step, operation, element, component, item, kind, and / or group, but do not exclude the presence, occurrence, or addition of one or more other features, steps, operations, elements, components, items, kinds, and / or groups. The terms "or," "and / or," "including at least one of the following," etc., as used in this application, may be interpreted as inclusive, or mean any one or any combination thereof. For example, "including at least one of the following: A, B, C" means "any one of the following: A; B; C; A and B; A and C; B and C; A and B and C." Similarly, "A, B, or C" or "A, B, and / or C" means "any one of the following: A; B; C; A and B; A and C; B and C; A and B and C." Exceptions to this definition only occur when the combination of elements, functions, steps, or operations is inherently mutually exclusive in some way.
[0033] It should be understood that although the steps in the flowcharts of this application's embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least a portion of the sub-steps or stages of other steps.
[0034] Depending on the context, the words “if” or “suppose” as used here can be interpreted as “when” or “in response to determination” or “in response to detection.” Similarly, depending on the context, the phrases “if determination” or “if detection (of the stated condition or event)” can be interpreted as “when determination” or “in response to determination” or “when detection (of the stated condition or event)” or “in response to detection (of the stated condition or event).”
[0035] It should be noted that step designations such as S21 and S22 are used in this document for the purpose of more clearly and concisely describing the corresponding content, and do not constitute a substantial limitation on the order. In specific implementation, those skilled in the art may execute S21 first and then S22, etc., but these should all be within the protection scope of this application.
[0036] It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit this application.
[0037] In the following description, the use of suffixes such as "module," "part," or "unit" to denote elements is solely for the purpose of illustrative purposes and has no specific meaning in itself. Therefore, "module," "part," or "unit" may be used interchangeably.
[0038] Smart terminals can be implemented in various forms. For example, the smart terminals described in this application may include smart terminals such as mobile phones, tablets, laptops, handheld computers, personal digital assistants (PDAs), portable media players (PMPs), navigation devices, wearable devices, smart bracelets, pedometers, etc., as well as fixed terminals such as digital TVs and desktop computers.
[0039] The following description will use a mobile terminal as an example. Those skilled in the art will understand that, apart from elements specifically designed for mobile purposes, the construction according to the embodiments of this application can also be applied to fixed-type terminals.
[0040] Please see Figure 1 This is a schematic diagram of the hardware structure of a mobile terminal implementing various embodiments of this application. The mobile terminal 100 may include: an RF (Radio Frequency) unit 101, a WiFi module 102, an audio output unit 103, an A / V (Audio / Video) input unit 104, a sensor 105, a display unit 106, a user input unit 107, an interface unit 108, a memory 109, a processor 110, and a power supply 111, etc. Those skilled in the art will understand that... Figure 1 The mobile terminal structure shown does not constitute a limitation on the mobile terminal. The mobile terminal may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0041] The following is combined with Figure 1 A detailed introduction to each component of the mobile terminal: The radio frequency unit 101 can be used for receiving and transmitting signals during information transmission or calls. Specifically, it receives downlink information from the base station and processes it with the processor 110; additionally, it transmits uplink data to the base station. Typically, the radio frequency unit 101 includes, but is not limited to, an antenna, at least one amplifier, a transceiver, a coupler, a low-noise amplifier, and a duplexer. Furthermore, the radio frequency unit 101 can also communicate wirelessly with networks and other devices. The aforementioned wireless communications may use any communication standard or protocol, including but not limited to GSM (Global System of Mobile communication), GPRS (General Packet Radio Service), CDMA2000 (Code Division Multiple Access 2000), WCDMA (Wideband Code Division Multiple Access), TD-SCDMA (Time Division-Synchronous Code Division Multiple Access), FDD-LTE (Frequency Division Duplexing-Long Term Evolution), TDD-LTE (Time Division Duplexing-Long Term Evolution), 5G, and 6G.
[0042] WiFi is a short-range wireless transmission technology. Mobile terminals, through the WiFi module 102, can help users send and receive emails, browse web pages, and access streaming media, providing users with wireless broadband internet access. Although Figure 1 WiFi module 102 is shown, but it is understood that it is not a necessary component of a mobile terminal and can be omitted as needed without changing the nature of the invention.
[0043] The audio output unit 103 can convert audio data received by the radio frequency unit 101 or the WiFi module 102 or stored in the memory 109 into audio signals and output them as sound when the mobile terminal 100 is in call signal receiving mode, call mode, recording mode, voice recognition mode, broadcast receiving mode, etc. Furthermore, the audio output unit 103 can also provide audio output related to specific functions performed by the mobile terminal 100 (e.g., call signal receiving sound, message receiving sound, etc.). The audio output unit 103 may include a speaker, a buzzer, etc.
[0044] The A / V input unit 104 is used to receive audio or video signals. The A / V input unit 104 may include a graphics processing unit (GPU) 1041 and a microphone 1042. The GPU 1041 processes image data of still images or videos acquired by an image capture device (such as a camera) in video capture mode or image capture mode. The processed image frames can be displayed on the display unit 106. The image frames processed by the GPU 1041 can be stored in the memory 109 (or other storage media) or transmitted via the radio frequency unit 101 or the WiFi module 102. The microphone 1042 can receive sound (audio data) in operating modes such as telephone call mode, recording mode, and voice recognition mode, and can process such sound into audio data. The processed audio (voice) data can be converted into a format that can be transmitted to a mobile communication base station via the radio frequency unit 101 in telephone call mode. The microphone 1042 can implement various types of noise cancellation (or suppression) algorithms to eliminate (or suppress) noise or interference generated during the reception and transmission of audio signals.
[0045] The mobile terminal 100 also includes at least one sensor 105, such as a light sensor, a motion sensor, and other sensors. Optionally, the light sensor includes an ambient light sensor and a proximity sensor. Optionally, the ambient light sensor can adjust the brightness of the display panel 1061 according to the ambient light level, and the proximity sensor can turn off the display panel 1061 and / or backlight when the mobile terminal 100 is moved to the ear. As a type of motion sensor, the accelerometer sensor can detect the magnitude of acceleration in various directions (generally three axes), and can detect the magnitude and direction of gravity when stationary. It can be used for applications that recognize the phone's posture (such as landscape / portrait switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc. Other sensors that may be configured in the phone, such as fingerprint sensors, pressure sensors, iris sensors, molecular sensors, gyroscopes, barometers, hygrometers, thermometers, and infrared sensors, will not be described in detail here.
[0046] The display unit 106 is used to display information input by the user or information provided to the user. The display unit 106 may include a display panel 1061, which may be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), or the like.
[0047] User input unit 107 can be used to receive input numerical or character information, and generate key signal inputs related to user settings and function control of the mobile terminal. Optionally, user input unit 107 may include touch panel 1071 and other input devices 1072. Touch panel 1071, also known as touch screen, can collect touch operations on or near the user (such as operations performed by the user using a finger, stylus, or any suitable object or accessory on or near touch panel 1071), and drive corresponding connection devices according to a pre-set program. Touch panel 1071 may include two parts: a touch detection device and a touch controller. Optionally, the touch detection device detects the user's touch position and the signal generated by the touch operation, and transmits the signal to the touch controller; the touch controller receives touch information from the touch detection device, converts it into touch point coordinates, sends it to processor 110, and can receive and execute commands sent by processor 110. In addition, touch panel 1071 can be implemented using various types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch panel 1071, the user input unit 107 may also include other input devices 1072. Optionally, other input devices 1072 may include, but are not limited to, one or more of the following: physical keyboard, function keys (such as volume control buttons, power buttons, etc.), trackball, mouse, joystick, etc., without being specifically limited here.
[0048] Optionally, the touch panel 1071 may cover the display panel 1061. When the touch panel 1071 detects a touch operation on or near it, it transmits the information to the processor 110 to determine the type of touch event. Subsequently, the processor 110 provides corresponding visual output on the display panel 1061 based on the type of touch event. Although in Figure 1 In this embodiment, the touch panel 1071 and the display panel 1061 are two independent components to realize the input and output functions of the mobile terminal. However, in some embodiments, the touch panel 1071 and the display panel 1061 can be integrated to realize the input and output functions of the mobile terminal. The specific implementation is not limited here.
[0049] Interface unit 108 serves as an interface through which at least one external device can connect to mobile terminal 100. For example, the external device may include a wired or wireless headset port, an external power supply (or battery charger) port, a wired or wireless data port, a memory card port, a port for connecting a device with an identification module, an audio input / output (I / O) port, a video I / O port, a headphone port, and so on. Interface unit 108 may be used to receive input (e.g., data, power, etc.) from the external device and transmit the received input to one or more elements within mobile terminal 100, or it may be used to transmit data between mobile terminal 100 and the external device.
[0050] The memory 109 can be used to store software programs and various data. The memory 109 may primarily include a program storage area and a data storage area. Optionally, the program storage area may store the operating system, applications required for at least one function (such as sound playback, image playback, etc.), etc.; the data storage area may store data created based on the use of the mobile phone (such as audio data, phonebook, etc.). Furthermore, the memory 109 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0051] The processor 110 is the control center of the mobile terminal. It connects various parts of the mobile terminal via various interfaces and lines. By running or executing software programs and / or modules stored in the memory 109, and by calling data stored in the memory 109, it performs various functions and processes data of the mobile terminal, thereby providing overall monitoring of the mobile terminal. The processor 110 may include one or more processing units; preferably, the processor 110 may integrate an application processor and a modem processor. Optionally, the application processor mainly handles the operating system, user interface, and applications, while the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 110.
[0052] The mobile terminal 100 may also include a power supply 111 (such as a battery) that supplies power to various components. Preferably, the power supply 111 can be logically connected to the processor 110 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system.
[0053] although Figure 1 As not shown, the mobile terminal 100 may also include a Bluetooth module, etc., which will not be described in detail here.
[0054] To facilitate understanding of the embodiments of this application, the communication network system on which the mobile terminal of this application is based is described below.
[0055] Please see Figure 2 , Figure 2 This application provides a communication network system architecture diagram. The communication network system is an LTE system based on the universal mobile communication technology. The LTE system includes a UE (User Equipment) 201, an E-UTRAN (Evolved UMTS Terrestrial Radio Access Network) 202, an EPC (Evolved Packet Core) 203, and the operator's IP services 204, which are connected in sequence.
[0056] Optionally, UE201 can be the aforementioned mobile terminal 100, which will not be described in detail here.
[0057] E-UTRAN202 includes eNodeB2021 and other eNodeB2022s. Optionally, eNodeB2021 can connect to other eNodeB2022s via backhaul (e.g., X2 interface). eNodeB2021 connects to EPC203 and can provide UE201 with access to EPC203.
[0058] EPC203 may include an MME (Mobility Management Entity) 2031, an HSS (Home Subscriber Server) 2032, other MMEs 2033, an SGW (Serving Gateway) 2034, a PGW (Packet Data Network Gateway) 2035, and a PCRF (Policy and Charging Rules Function) 2036, etc. Optionally, MME2031 is the control node that handles signaling between UE201 and EPC203, providing bearer and connection management. HSS12032 is used to provide registers to manage functions such as the Home Location Register (not shown in the figure) and stores user-specific information such as service characteristics and data rates. All user data can be sent through SGW2034. PGW2035 can provide UE 201 IP address allocation and other functions. PCRF2036 is the policy and charging control decision point for service data flow and IP bearer resources. It selects and provides available policy and charging control decisions for the policy and charging enforcement function unit (not shown in the figure).
[0059] IP services 204 may include the Internet, intranet, IMS (IP Multimedia Subsystem), or other IP services.
[0060] Although the above description uses the LTE system as an example, those skilled in the art should know that this application is not only applicable to the LTE system, but also to other wireless communication systems, such as GSM, CDMA2000, WCDMA, TD-SCDMA, 5G and future new network systems (such as 6G), etc., without limitation.
[0061] Based on the above-described mobile terminal hardware structure and communication network system, various embodiments of this application are proposed.
[0062] First Embodiment Figure 3 This application provides a processing method as a first embodiment. The executing entity of this processing method can be at least one of a mobile phone or other smart terminal, a wearable device, or a processor or storage medium with processing capabilities. Please refer to [link to relevant documentation]. Figure 3 The processing method includes the following steps: S11. Obtain the learning audio; S12. Output information on the differences between the speech style of the learning audio and the target speech style.
[0063] In step S11, the method for obtaining the learning audio can be: in response to determining the target speech style, obtaining learning audio based on the target speech style. That is, after determining the target speech style, the smart terminal obtains the corresponding learning audio through the audio emitted by the user based on the target speech style. The learning audio is the audio emitted by the user according to the target speech style, which can be regarded as the audio obtained through imitation. Based on this, the method further includes: determining the target speech style.
[0064] Voice style refers to a unique vocal expression formed by the combination of elements such as timbre, intonation, rhythm, and stress. It encompasses multiple dimensions, including emotional expression, scene adaptation, and personalized characteristics. The target voice style can be regarded as a model for users to refer to when learning.
[0065] In scenarios where users learn through terminals or devices, the target speech style can be represented by one or more audio or video segments with sound, referred to as audio for ease of description below. When the terminal plays the audio, the user can obtain the target speech style by listening to the sound. Alternatively, the target speech style can be represented by a target object that has a relationship with the target speech style, also known as the "first target object," which includes, but is not limited to, images with a specific language style such as real person portraits, cartoon characters, and anthropomorphic characters. When the terminal displays the first target object, the user can intuitively know the corresponding speech style, which is the target speech style.
[0066] For example, if the primary target is the real image or cartoon character of a well-known comedian, users can intuitively identify classic comedic scenes associated with that character and use the corresponding target speech style. As another example, if the primary target is a well-known animal cartoon character specific to a particular province in China, users can intuitively identify the dialect style of that province and use that as the target speech style. Yet another example is if the primary target is a mascot of a country or region, users can intuitively identify the language spoken by that mascot, such as British English, and use the corresponding speech style as the target speech style.
[0067] Based on the above manifestations, the target speech style can be determined by at least one of the following methods 1 to 4: Method 1: Select the first target object and use the speech style corresponding to the first target object as the target speech style.
[0068] Method 2: Determine the target audio and analyze its speech style to determine the target speech style.
[0069] In one example, the method can be applied to an app on a smart terminal, named, for example, "XX Processing Method" program, in combination with... Figure 4a As shown, the smart terminal can display a variety of different "attribute" menu items on the running interface of the APP. Each "attribute" menu item corresponds to a form of expression, including but not limited to the first attribute "real person's head" menu item, the second attribute "anthropomorphic image" menu item, and the third attribute "audio" menu item shown in the figure. Each attribute menu item includes at least one object.
[0070] In response to the user selecting the first attribute menu item, the smart terminal switches to the APP running interface displaying multiple real person portraits. When the user clicks to select one of the real person portraits A1x, the real person portrait A1x is taken as the first target object, and the corresponding voice style is taken as the target voice style.
[0071] In response to the user selecting the third attribute menu item, the smart terminal switches to the APP running interface displaying multiple audio files. The user clicks on an audio file labeled A3x, and the smart terminal plays that audio file A3x. If the user confirms after listening, and then clicks again to select audio file A3x as the target audio file, the corresponding voice style of audio file A3x will be used as the target voice style. Alternatively, these audio files can be labeled A31~A3n in a preset order or display corresponding names and other information. The user clicks on one of the audio files A3x, and that audio file A3x will be used as the target audio file, and its corresponding voice style will be used as the target voice style.
[0072] In real-world scenarios, such as Figure 4a As shown, the APP's running interface can provide custom menu items and / or switch menu items. The switch menu items are used to display the words "on" and "off" to allow users to turn the custom menu items on or off accordingly. The custom menu items allow users to edit attribute menu items, such as adding a fourth attribute menu item or deleting the aforementioned third attribute menu item.
[0073] Method 3: For scenarios where either the first target object or the target audio is obtained, determine the language information corresponding to the first target object and / or the target audio, process the target audio and / or the first target object based on the language information, generate the target audio, and then use the style of the target audio as the target speech style.
[0074] Method 4: In one example, after determining the language information corresponding to the first target object and / or the target audio, this application may, in response to the language information being different from the language information corresponding to the learning audio, translate the target audio into audio using the language information corresponding to the learning audio; and take the style corresponding to the translated audio as the target speech style.
[0075] When the language of the learning audio sent by the user differs from that of the target audio, the current technology typically translates the user-sent learning audio into the language of the target audio. However, this results in the user's imitated learning audio having a stiff and mechanical pronunciation style, significantly different from the target audio's pronunciation style, leading to poor learning effectiveness and experience. For example, when the learning audio is Chinese and the target audio is English, the user's imitated learning audio usually sounds like English with a Chinese accent. While this may be sufficient for communication with native speakers, it fails to capture the essence of English pronunciation. Furthermore, in scenarios where the language information is a dialect, existing technologies are completely incapable of achieving pronunciation style imitation when the learning audio is in the first dialect of the same language and the target audio is in the second dialect of the same language.
[0076] To address this technical issue, Method 4 translates the target audio into a language familiar to the user when the language of the learning audio sent by the user is different from that of the target audio. This preserves the target speech style, allowing the user to understand both the character information contained in the target audio and the target speech style, resulting in a better learning outcome and experience.
[0077] Taking the example of learning audio in Chinese and target audio in English, after translating the target audio into Chinese, the pronunciation of the Chinese audio still retains the English accent, intonation, and rhythm. Chinese users can not only understand the target audio, but also learn and imitate the pronunciation.
[0078] For any one or more audio files displayed on the app interface of a smart terminal using any of the aforementioned methods, they can be obtained through at least one of the following methods 21 to 24: Method 21: Collect the audio of the current environment as the target audio.
[0079] Method 22: In response to the selected target language, use the audio example corresponding to the target language as the target audio. For example, combine... Figure 4a and Figure 4b As shown, in response to a user's operation of selecting the third attribute menu item, the smart terminal and APP switch to an APP running interface displaying multiple languages. These languages include, but are not limited to, international languages such as English and Chinese, as well as regional languages such as local dialects. Then, in response to selecting a certain language, i.e., as the target language, the system switches to an APP running interface displaying at least one audio. The user can select one of the audios in the aforementioned manner to use it as the target audio.
[0080] The audio samples described in this application can be pre-stored in smart terminals and apps, allowing users to record, delete, or edit them at any time to adjust the audio samples.
[0081] Method 23: In response to a selected target scene, use the audio example corresponding to the target scene as the target audio. For example, combine... Figure 4a As shown, and see also Figure 4b The interface display and switching method shown is as follows: In response to a user's operation of selecting the third attribute menu item, the smart terminal and APP switch to the APP running interface displaying multiple scenes. These scenes include, but are not limited to, visually identifiable regional images. For example, if a famous building image is displayed, it is identified as India. Then, in response to selecting a specific regional image, it becomes the target scene, and the interface switches to displaying images such as... Figure 4b The right image shows the app's interface for at least one audio file. Users can select one of the audio files as the target audio file using the aforementioned method.
[0082] Method 24: In response to selecting a second target object, use the audio example corresponding to the second target object as the target audio. For example, combine... Figure 4a As shown, and see also Figure 4b The interface display and switching method shown, in response to a user operation of selecting the third attribute menu item, switches the smart terminal and APP to an APP running interface displaying multiple objects. The representation of these objects includes, but is not limited to, the representation of the aforementioned first target object. Then, in response to selecting a specific object, i.e., as the second target object, the interface switches to displaying, for example... Figure 4b The right image shows the app's interface for at least one audio file. Users can select one of the audio files as the target audio file using the aforementioned method.
[0083] Therefore, this application may obtain the target audio through at least one of methods 21 to 24.
[0084] It should be noted that any embodiment provided throughout this application includes multiple scenarios and feasible implementation methods for a particular technical feature. Unless otherwise specified, it means that the corresponding technical feature can be implemented by combining any of these methods. For example, the target speech style can be determined by jointly identifying a first target object and the target audio. Another example is that the target audio can be determined by jointly identifying the target language, the target scene, and a second target object. By combining these methods, the corresponding technical feature can be implemented more accurately and / or intelligently, thereby improving the accuracy of the technical feature implementation and the user experience.
[0085] In step S12, "outputting difference information" can be understood as displaying the difference information on the screen of the smart terminal. Through this application embodiment, the user is allowed to emit audio according to a target speech style, i.e., the learning audio, and the speech style of the learning audio is compared with the target speech style to obtain the difference information between the two styles. This allows the user to clearly understand the differences and correct their lip movements, facial expressions, and other pronunciation methods related to the speech style. This achieves processing methods for new languages, different languages, or imitating specific speaking styles, and enables personalized learning experiences such as tutoring and free practice, thereby improving the user's experience and learning experience.
[0086] The form of the difference information can be determined according to the adaptability of the actual scenario, including but not limited to at least one of the following: rhythm curve difference, preset information difference of virtual image, and text difference. This application does not limit this. Therefore, this application can first determine the identification rules for the difference information. The identification rules can be considered as the form of expression, output, or display of the difference information, such as... Figure 5a As shown, the method may further include: S11. Obtain the learning audio; S120. Determine the identification rules for the difference information; S12. In response to the identification rule, output the difference information between the speech style of the learning audio and the target speech style.
[0087] The following section uses the example of differences in rhythm curves, differences in preset information of virtual avatars, and differences in text to describe in detail the implementation scheme of the processing method of this application.
[0088] Second Embodiment Figure 5b This is a flowchart illustrating the processing method of the second embodiment of this application. Figure 5b As shown, the processing method includes at least the following steps S20 to S222.
[0089] S20. Obtain the target audio; S21. Obtain the learning audio; S221. Generate the first prosodic curve corresponding to the learning audio and the second prosodic curve when the target audio is pronounced according to the target speech style; S222, Display the difference information between the speech style of the learning audio and the target speech style through the first prosodic curve and the second prosodic curve.
[0090] In the first embodiment described above, if step S21 obtains the target audio when determining the target speech style, that is, step S21 already includes step S20, then this second embodiment does not need to execute step S20 again.
[0091] The same features as the first embodiment described above can be found in the foregoing, such as step S21, which corresponds to the same features as step S11 described above, and will not be repeated here. The difference is that the third embodiment can be regarded as decomposing step S12 of the first embodiment into steps S221 and S222.
[0092] The generation method of the rhythm curve can be found in existing technologies in this field, such as extracting relevant pitch, timbre, and intensity factors from the corresponding audio, which will not be elaborated here. Figure 6a As shown, the app interface of the smart terminal can display a first prosodic curve and a second prosodic curve. The horizontal axis represents time information, such as the sequence from 0 seconds to n seconds (n seconds represents the duration of the audio). The two prosodic curves can be different colors to facilitate quick and intuitive differentiation for users.
[0093] Optionally, the first and second prosodic curves are aligned sequentially, allowing users to visually compare the curve shapes at any given moment to intuitively determine whether the pronunciation style of the learning audio at that moment matches the target pronunciation style, thus clearly identifying the differences. For example, if the first and second prosodic curves differ at positions Z1 and Z2 on the graph, users can adjust their pronunciation accordingly in a timely and accurate manner.
[0094] This example demonstrates the differences between the user's learned speech style and the target speech style by comparing the prosody curves. This difference is presented in a visually clear form, allowing users to intuitively understand their shortcomings and thus gain a better experience with the processing methods.
[0095] In one example, after step S222, the method may further include: in response to a user operation on the first prosodic curve or the second prosodic curve, determining a position selected on the corresponding prosodic curve by the user operation, thereby displaying difference information corresponding to the position.
[0096] Combined Figure 6a and Figure 6b As shown, for example, if the user clicks on position Z1 of the second prosodic curve, the smart terminal recognizes the time t1 corresponding to position Z1. A vertical line can be displayed on the APP interface, which can be called the time line corresponding to time t1. This time line intersects with both the first and second prosodic curves that are aligned vertically. The user can then intuitively understand the difference in speech style at time t1. For example, the difference in speech style can be directly displayed in text form on the APP interface (not shown in the figure).
[0097] Or, combine Figure 7 As shown, the first display area A of the APP running interface can display the text of the target audio. The first rhythm curve and the second rhythm curve are both displayed vertically aligned in the second display area B, which can be located below the first display area A. When the user selects the position, the character corresponding to that position is highlighted in the first display area A. For example, at least one of the character's font color, font size, background color, and background pattern is displayed as a different character from the text.
[0098] Please continue reading. Figure 5b The processing method may further include the following steps S223 to S224.
[0099] S223. Generate a first virtual image corresponding to the learning audio and a second virtual image when the target audio is spoken according to the target speech style; S224. Display the difference information between the speech style of the learning audio and the target speech style through the first virtual image and the second virtual image; This example focuses on realistic lip-syncing animation and view display for a virtual avatar. In one example, the virtual avatar may include a front view and a side view of the virtual character; see [link / reference]. Figure 7As shown, the APP interface can also include a third display area, which may include four separately set sub-areas, namely the first to fourth sub-areas from left to right in the figure, and labeled C, D, E, and F respectively. The first sub-area C is used to display a standard frontal view of the second virtual avatar, the second sub-area D is used to display a standard frontal view of the first virtual avatar generated based on the user's upper body head image, the third sub-area E is used to display a standard side view of the second virtual avatar, and the fourth sub-area F is used to display a standard side view of the first virtual avatar generated based on the user's upper body head image. The first sub-area C and the second sub-area D can be set adjacent to each other, and the third sub-area E and the fourth sub-area F can be set adjacent to each other to facilitate comparison by the user. Optionally, once the user determines the target audio for this imitation, the second virtual avatar displayed in the first sub-area C and the third sub-area E can remain unchanged during the user's imitation learning process. However, the pronunciation features (such as mouth shape and facial features) of these two views will change in real time according to the user's reading of the text. At the same time, the first virtual avatar displayed in the second sub-area D and the fourth sub-area F will also be adjusted in real time according to the user's pronunciation features.
[0100] This example demonstrates the differences between the pronunciation features of a virtual avatar and the target pronunciation style by comparing them. This difference is presented in a visually perceptible way, allowing users to intuitively understand their shortcomings and thus gain a better experience with the processing methods.
[0101] In one example, in response to user actions on the first and second virtual avatars, the app's interface can highlight specific features of the corresponding virtual avatar. For example, a user can click on features such as... Figure 7 If the first sub-region C shown is identified as the mouth of the second virtual character, the mouth features of the second virtual character and the mouth features of the first virtual character in the second sub-region D can be enlarged and displayed, so that users can clearly see the mouth features when pronouncing the target speech style and compare them with their own mouth features when pronouncing, thus achieving rapid correction.
[0102] Please continue reading. Figure 5b The processing method may further include the following steps S225 to S226.
[0103] S225. Generate the first text corresponding to the learning audio and the second text corresponding to the target audio; S226. Display the difference information between the speech style of the learning audio and the target speech style through the first text and the second text.
[0104] Combination Figure 7As shown, the first display area A of the APP's running interface can display the text of the target audio, which is the second text. In other examples, the first display area A can also display the text corresponding to the learning audio, which is the first text. Refer to the aforementioned alignment of the first and second prosodic curves. The first text and the second text can also be aligned vertically within the first display area A.
[0105] As the user progresses in pronunciation, corresponding characters in the first and second texts are highlighted. For differences between the learning audio's speech style and the target speech style, characters corresponding to these differences are highlighted. The first and second highlights differ, allowing the user to intuitively perceive text differences. Either highlight can be manifested by displaying the character's font color, font size, background color, and background pattern as a different character from the text. For example, as the user progresses in pronunciation, the currently pronounced character is displayed in blue in both the first and second texts, other characters that have been pronounced are displayed in black, and unpronounced characters are displayed in gray. The first highlight would then be displayed in blue, and the second highlight could be displayed in red to indicate that the user has mispronounced the red character.
[0106] Third Embodiment Regarding the specific manifestation of the difference information provided in the second embodiment, this embodiment may continue to perform the following steps after step S222: S322. In response to a first preset operation on the first or second rhythmic curve, determine a first position selected on the corresponding rhythmic curve by the first preset operation; and S323. Display the pronunciation lip movements of the second virtual image corresponding to the first position.
[0107] Combined Figure 6b and Figure 7 As shown, the first preset operation can be an operation whereby the user clicks and selects a position on the first or second rhythm curve, and this position is the first position. For example, combined with Figure 6b As shown, the user clicks on position Z1 of the second rhythm curve. The smart terminal identifies the time t1 corresponding to position Z1, and the second virtual image's mouth shape at that time t1 is displayed in the third display area C. The user can then intuitively understand the correct speech style at that time t1.
[0108] Similarly, in this embodiment, after step S226, the following steps can be performed: S326. In response to a second preset operation on the first text or the second text, determine a second position selected on the corresponding text by the second preset operation; and S327. Display the pronunciation lip movements of the second virtual image corresponding to the second position.
[0109] Combined Figure 6b and Figure 7 As shown, the second preset operation can also be an operation where the user clicks and selects a certain position on the first or second text, and that certain position is the second position. For example, if the user clicks position Z3 on the second text, the smart terminal recognizes the time t3 corresponding to position Z3, and the mouth shape of the second virtual image at that time t3 can be displayed in the third display area C, so that the user can intuitively know the correct speech style at that time t3.
[0110] Based on any of the foregoing examples, the method of this application may further include: marking the difference information with preset prompt symbols. This step can be implemented after the aforementioned step S12. The preset prompt symbols are prompts related to speech style, including but not limited to accent marks, to remind the user.
[0111] In the example of the difference information identified by the aforementioned text and / or prosodic curve, marking the difference information with a preset prompt symbol can be manifested as: marking the second text with an emphasis symbol; and / or, marking the second prosodic curve with an emphasis symbol.
[0112] This application also provides a smart terminal, including a memory and a processor. The memory stores a processing program, which, when executed by the processor, implements the steps of the processing method in any of the above embodiments.
[0113] This application also provides a storage medium storing a processing program, which, when executed by a processor, implements the steps of the processing method in any of the above embodiments.
[0114] The embodiments of the smart terminal and storage medium provided in this application may include all the technical features of any of the above-described processing method embodiments, and thus have corresponding beneficial effects. The extended and explanatory content of the specification is basically the same as the various embodiments of the above methods, and will not be repeated here.
[0115] This application also provides a computer program product, which includes computer program code. When the computer program code is run on a computer, it causes the computer to perform the methods described in the various possible implementations above.
[0116] This application also provides a chip, including a memory and a processor. The memory is used to store a computer program, and the processor is used to call and run the computer program from the memory, so that a device with the chip installed performs the methods described in the various possible implementations above.
[0117] It is understood that the above scenarios are merely examples and do not constitute a limitation on the application scenarios of the technical solutions provided in the embodiments of this application. The technical solutions of this application can also be applied to other scenarios. For example, as those skilled in the art will know, with the evolution of system architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0118] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0119] The steps in the method of this application embodiment can be adjusted, combined, or deleted according to actual needs.
[0120] The units in the device of this application embodiment can be merged, divided, and deleted according to actual needs.
[0121] In this application, the same or similar terms, concepts, technical solutions and / or application scenario descriptions are generally described in detail only when they appear for the first time. When they appear again, they are generally not repeated for the sake of brevity. When understanding the technical solutions and other contents of this application, the same or similar terms, concepts, technical solutions and / or application scenario descriptions that are not described in detail later can be referred to their previous relevant detailed descriptions.
[0122] In this application, the descriptions of the various embodiments have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0123] The technical features of the present application can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of the present application.
[0124] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, controlled terminal, or network device, etc.) to execute the methods of each embodiment of this application.
[0125] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this application is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a storage medium or transmitted from one storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, storage disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state disk (SSD)).
[0126] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
Claims
1. A processing method, characterized in that, Including the following steps: S11. Obtain the learning audio; S12. Output the difference information between the speech style of the learning audio and the target speech style.
2. The method according to claim 1, characterized in that, Also includes: Determine the target speech style; the method for determining the target speech style includes at least one of the following: Select a first target object and use the speech style corresponding to the first target object as the target speech style; Analyze the speech style of the target audio to determine the target speech style.
3. The method according to claim 2, characterized in that, The method for determining the target speech style includes: Determine the language information corresponding to the first target object and / or the target audio; Based on the language information, the target audio and / or the first target object are processed to generate the target audio; The style of the target audio is taken as the target speech style.
4. The method according to claim 2, characterized in that, The method of obtaining the target audio includes at least one of the following: Collect the audio of the current environment as the target audio; In response to the selection of a target language, the audio example corresponding to the target language is used as the target audio; In response to the selection of a target scene, the audio example corresponding to the target scene is used as the target audio; In response to selecting a second target object, the audio example corresponding to the second target object is used as the target audio.
5. The method according to any one of claims 1 to 4, characterized in that, Also includes: Determine the identification rules for the difference information; The difference information is output according to the identification rules.
6. The method according to claim 5, characterized in that, The step of outputting the difference information according to the identification rule includes at least one of the following: Generate a first prosodic curve corresponding to the learning audio and a second prosodic curve when the target audio is pronounced according to the target speech style, and display the difference information between the speech style of the learning audio and the target speech style through the first prosodic curve and the second prosodic curve; Generate a first virtual image corresponding to the learning audio and a second virtual image when the target audio is spoken according to the target speech style, and display the difference information between the speech style of the learning audio and the target speech style through the first virtual image and the second virtual image; Generate a first text corresponding to the learning audio and a second text corresponding to the target audio, and display the difference information between the speech style of the learning audio and the speech style of the target audio through the first text and the second text.
7. The method according to claim 6, characterized in that, Also includes: In response to a first preset operation on the first rhythm curve or the second rhythm curve, a first position selected on the corresponding rhythm curve by the first preset operation is determined. Display the lip movements of the second virtual character corresponding to the first position; or, In response to a second preset operation on the first text or the second text, a second position selected on the corresponding text by the second preset operation is determined; Displays the lip movements of the second virtual character corresponding to the second position.
8. The method according to any one of claims 1 to 4, characterized in that, Also includes: Preset prompt symbols are marked in the difference information.
9. A smart terminal, characterized in that, include: A memory and a processor, wherein the memory stores a processing program that, when executed by the processor, implements the steps of the processing method as described in any one of claims 1 to 8.
10. A storage medium, characterized in that, The device contains a computer program that, when executed by a processor, implements the steps of the processing method as described in any one of claims 1 to 8.