Method for converting speech to text and terminal device
Patent Information
- Application Number
- PCT/CN2025/138893
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-14
- Filing Date
- 2025-12-01
- Publication Date
- 2026-09-17
Smart Images

Figure CN2025138893_17092026_PF_FP_ABST
Abstract
Description
A method and terminal device for speech-to-text conversion
[0001] This application claims priority to Chinese Patent Application No. 202510318315.9, filed on March 14, 2025, entitled "A Method and Terminal Device for Speech-to-Text Conversion", the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of terminal technology, and in particular to a method and terminal device for speech-to-text conversion. Background Technology
[0003] With the widespread use of terminal devices in daily life, information interaction within these devices is frequent. In traditional information interaction, text input is typically the core method. Users in various applications rely on the traditional keyboard provided by the input method to input text via handwriting or pinyin. However, text input using a traditional keyboard is inefficient and inconvenient. For example, when a user only has one free hand for inputting text, it is difficult to accurately and quickly complete the input using only the keyboard.
[0004] Therefore, to improve users' text input efficiency, various application software input methods have successively provided speech-to-text functionality. This means that by converting the user's voice input into text, it assists the user in text input. When a user needs speech-to-text, they need to enable the function through the corresponding component in the input method, displaying a voice input control. Then, using the voice input control, the speech-to-text input is achieved. It is evident that the above speech-to-text input process is relatively complex and cumbersome, resulting in a poor user experience. Summary of the Invention
[0005] This application provides a method and terminal device for speech-to-text conversion, which allows users to realize the speech-to-text function through various triggering methods, making the triggering methods of speech-to-text function more diversified. This can reduce the probability of user fatigue and improve the user experience.
[0006] To achieve the above objectives, the embodiments of this application adopt the following technical solutions:
[0007] Firstly, a speech-to-text method is provided, applied to a terminal device. First, a first interactive interface is displayed on the terminal device. Second, in response to an input operation performed on the first interactive interface, the terminal device displays a first microphone interface on the first interactive interface. Then, in response to a first long press operation on the first microphone control, the terminal device continuously acquires the user's voice information during the execution of the first long press operation. Alternatively, in response to a first click operation on the first microphone control, the terminal device continuously acquires the user's voice information. Finally, the terminal device displays text information on the first interactive interface.
[0008] This allows users to choose the appropriate operation method (long press or tap) based on their own situation when they need to convert speech to text, triggering the continuous acquisition of voice information to achieve the speech-to-text function. Furthermore, it diversifies the triggering methods for the speech-to-text function, thereby reducing the probability of user fatigue, improving the convenience of the speech-to-text function, and enhancing the user experience.
[0009] In one possible implementation of the first aspect, while the terminal device is continuously acquiring voice information through a first click operation, it responds to an information modification instruction by modifying the first information already displayed on the first interactive interface and then displays text information, including the modified first information, on the first interactive interface. This allows the user to edit the converted text information without stopping voice input, reducing the complexity of the text modification process and decreasing the response frequency of the terminal device.
[0010] In one possible implementation of the first aspect, the first information on the first interactive interface is highlighted before it is modified. This prompts the user that the first information will soon be modified, assisting the user in verifying the accuracy of the information to be modified and reducing the probability of incorrect modification.
[0011] In one possible implementation of the first aspect, after receiving an information modification instruction, the terminal device displays the content of the instruction in text form on the first audio interface. This facilitates the user's verification of the issued information modification instruction, allowing the user to promptly adjust the instruction in case of an incorrect issuance, reducing the probability of the user repeatedly modifying the displayed text information, and also reducing the phone's response frequency.
[0012] In one possible implementation of the first aspect, during the process of continuously acquiring voice information in response to a first click operation on the first microphone control, the terminal device changes the display mode of the first microphone control from a first mode to a second mode. In this way, by changing the display mode of the first microphone control, the user is notified that the device is currently in the process of continuously acquiring voice information, thereby reducing the probability that the user unknowingly inputs invalid voice information into the phone.
[0013] In one possible implementation of the first aspect, the first radio interface includes a first switching control. During the process of the terminal device continuously acquiring voice information in response to a first click operation on the first radio control, the terminal device can first respond to a trigger operation on the first switching control, display a text input keyboard in the first interactive interface, and then stop acquiring voice information. The text input keyboard further includes a second switching control. Secondly, the terminal device responds to a trigger operation on the second switching control, displaying the first radio interface in the first interactive interface. This allows the user to switch between the first radio interface and the text input keyboard during the speech-to-text function.
[0014] In one possible implementation of the first aspect, firstly, in response to a first long press operation on the first microphone control, the terminal device displays a second microphone interface in the first interactive interface. Secondly, after displaying the second microphone interface, if the first long press operation has not ended, the terminal device continues to acquire voice information. In this way, while facilitating information interaction with other users (or sending "short voice messages"), users can still maintain their previous operating habits and quickly achieve voice-to-text conversion by performing a single long press operation to conveniently send text information.
[0015] In one possible implementation of the first aspect, after displaying the second audio interface, the terminal device, in response to the end of the first long press operation, displays the first audio interface in the first interactive interface and ends the acquisition of voice information. Thus, after the user completes the use of the speech-to-text function, the phone can automatically switch back to the first audio interface without requiring any additional user action. This allows the user to reselect an appropriate trigger operation (click operation, long press operation, etc.) during subsequent information interaction to re-enable the speech-to-text function.
[0016] In one possible implementation of the first aspect, the second audio interface includes a voice editing control. After displaying the second audio interface, the terminal device responds to the first long press operation by changing it to a first swipe operation, and the first swipe operation ends at the voice editing control, thus obtaining an information modification command. In this way, the user can actively trigger the modification of text information during the speech-to-text process using the voice editing control, reducing the probability that the terminal device might mistakenly identify interactive information as an information modification command when modifying text information by recognizing the modification command itself.
[0017] In one possible implementation of the first aspect, firstly, the terminal device responds to a first long press operation 1 on the first microphone control and displays a second microphone interface. Secondly, the phone responds to the first long press operation on the second microphone interface changing to a second swipe operation. After the second swipe operation ends, a third microphone interface is displayed, and voice information is continuously acquired. This allows users to adjust the triggering method for the voice-to-text function at any time according to their own situation during the voice-to-text conversion process, making the triggering methods of the voice-to-text function more diverse. This reduces the probability of user fatigue and improves the user experience. Furthermore, the size of the third microphone interface is smaller than the size of the second microphone interface, thus reducing the obstruction of the first interactive interface by the microphone interface during continuous acquisition of voice information.
[0018] In one possible implementation of the first aspect, the third audio interface includes a voice editing control. When the terminal device responds to a change in the first long-press operation to a second swipe operation within the second audio interface, and continues to acquire voice information, it responds to a trigger operation on the voice editing control to obtain information modification instructions. In this way, the user can actively trigger the modification of text information during the speech-to-text process using the voice editing control, thereby reducing the probability that the phone mistakenly identifies interactive information as information modification instructions when modifying text information by recognizing modification instructions.
[0019] In one possible implementation of the first aspect, the input operation includes triggering a text input operation on a first interactive interface. Alternatively, the text input operation may be triggered in an input box displayed on the first interactive interface.
[0020] Secondly, an alternative speech-to-text method is provided for application on terminal devices. First, a second interactive interface is displayed on the terminal device. Second, in response to a second long press on the first key of the text input keyboard, a fourth audio interface is displayed on the second interactive interface. Then, in response to the second long press on the fourth audio interface changing to a third swipe operation, the terminal device continues to acquire voice information after the third swipe operation ends. Finally, the terminal device displays text information on the second interactive interface. In this way, users can adjust the triggering method for the speech-to-text function at any time according to their own needs, making the triggering methods more diverse. This reduces the probability of user fatigue and improves the user experience.
[0021] In one possible implementation of the second aspect, during the continuous acquisition of voice information, the terminal device responds to a received information modification instruction by modifying the second information already displayed on the second interactive interface, and then displays text information, including the modified second information, on the second interactive interface. This allows the user to edit the converted text information without stopping voice input, reducing the complexity of the operation and decreasing the phone's response frequency.
[0022] In one possible implementation of the second aspect, the second information on the second interactive interface is highlighted before it is modified. This prompts the user that the second information will soon be modified, helping the user verify the accuracy of the first information to be modified and reducing the probability of incorrect modification.
[0023] In one possible implementation of the second aspect, after receiving an information modification instruction, the terminal device displays the content of the instruction in text form on the fourth audio interface. This facilitates the user's verification of the issued information modification instruction, allowing the user to promptly adjust the instruction in case of an erroneous issuance, reducing the probability of the user repeatedly modifying already displayed text information, and also reducing the phone's response frequency.
[0024] In one possible implementation of the second aspect, the fourth radio interface includes a second radio control; the third sliding operation ends at the second radio control.
[0025] In one possible implementation of the second aspect, while the terminal device is continuously acquiring voice information, the terminal device responds to a trigger operation on the second microphone control, pauses the acquisition of voice information, and displays the first prompt information on the fourth microphone interface. This allows for timely prompting to the user when the terminal device is in a paused state, reducing the probability of the user outputting invalid voice information during the paused state.
[0026] In one possible implementation of the second aspect, firstly, the terminal device responds to a second long press operation on the fourth audio interface, continuously acquiring voice information during the execution of the second long press operation. Secondly, in response to the end of the second long press operation, a text input keyboard is displayed on the second interactive interface, and the acquisition of voice information ceases. Thus, after the user completes the use of the speech-to-text function, the terminal device can automatically switch back to the first audio interface, allowing the user to reselect a suitable text input method in subsequent information interactions to again implement the speech-to-text function.
[0027] In one possible implementation of the second aspect, if the terminal device fails to acquire voice information within a preset time period during the continuous acquisition of voice information, a second prompt message is displayed on the fourth audio interface. This allows the user to be prompted with the currently available voice input information even before they are familiar with the voice input rules, thereby assisting the user in using the voice-to-text function and improving the user experience.
[0028] Thirdly, a terminal device is provided, the terminal device including a memory and one or more processors; the memory is coupled to the processors; wherein the memory stores computer program code, the computer program code including computer instructions, and when the computer instructions are executed by the processor, the terminal device performs a speech-to-text method as in the first aspect and any implementation thereof, or a speech-to-text method as in the second aspect and any implementation thereof.
[0029] Fourthly, a computer-readable storage medium is provided, including computer instructions that, when executed on a terminal device, cause the terminal device to perform a speech-to-text method as described in the first aspect and any implementation thereof, or a speech-to-text method as described in the second aspect and any implementation thereof.
[0030] Fifthly, a computer program product is provided, which, when run on a terminal device, causes the terminal device to execute the speech-to-text method as described in the first aspect and any implementation thereof, or the speech-to-text method as described in the second aspect and any implementation thereof.
[0031] The beneficial effects that can be achieved by the terminal equipment provided in the third aspect, the computer-readable storage medium provided in the fourth aspect, and the computer program product provided in the fifth aspect can be referred to as the beneficial effects that can be achieved by the first aspect and any of its implementations, or the beneficial effects that can be achieved by the second aspect and any of its implementations, which will not be repeated here. Attached Figure Description
[0032] Figure 1 shows a schematic diagram of a speech-to-text trigger provided in an embodiment of this application;
[0033] Figure 2 shows a schematic diagram of the hardware structure of a terminal device provided in an embodiment of this application;
[0034] Figure 3 shows a schematic diagram of the software structure of a terminal device provided in an embodiment of this application;
[0035] Figure 4 shows a flowchart of a speech-to-text method provided in an embodiment of this application;
[0036] Figure 5 shows one of the schematic diagrams of a radio interface provided in an embodiment of this application;
[0037] Figure 6 shows a schematic diagram of a radio interface trigger provided in an embodiment of this application;
[0038] Figure 7 shows a second schematic diagram of a radio interface display provided in an embodiment of this application;
[0039] Figure 8 shows a third schematic diagram of a radio interface provided in an embodiment of this application;
[0040] Figure 9 shows a fourth schematic diagram of a radio interface display provided in an embodiment of this application;
[0041] Figure 10 shows a fifth schematic diagram of a radio interface display provided in an embodiment of this application;
[0042] Figure 11 shows a sixth schematic diagram of a radio interface display provided in an embodiment of this application;
[0043] Figure 12 shows one of the interface switching diagrams provided in an embodiment of this application;
[0044] Figure 13 shows a seventh schematic diagram of a radio interface display provided in an embodiment of this application;
[0045] Figure 14 shows an eighth schematic diagram of a radio interface display provided in an embodiment of this application;
[0046] Figure 15 shows a second schematic diagram of an interface switching provided in an embodiment of this application;
[0047] Figure 16 shows a ninth schematic diagram of a radio interface display provided in an embodiment of this application;
[0048] Figure 17 shows a schematic diagram of a radio interface provided in an embodiment of this application;
[0049] Figure 18 shows a third schematic diagram of an interface switching embodiment provided in this application;
[0050] Figure 19 shows an eleventh schematic diagram of a radio interface display provided in an embodiment of this application;
[0051] Figure 20 shows a schematic diagram of a radio interface provided in an embodiment of this application;
[0052] Figure 21 shows a schematic diagram of a radio interface provided in an embodiment of this application;
[0053] Figure 22 shows a flowchart of another speech-to-text method provided in an embodiment of this application;
[0054] Figure 23 shows fourteenth of a schematic diagram of a radio interface provided in an embodiment of this application;
[0055] Figure 24 shows a fourth schematic diagram of an interface switching embodiment provided in this application;
[0056] Figure 25 shows a fifteenth schematic diagram of a radio interface display provided in an embodiment of this application;
[0057] Figure 26 shows a sixteenth schematic diagram of a radio interface display provided in an embodiment of this application;
[0058] Figure 27 illustrates a fifth schematic diagram of an interface switching method provided in an embodiment of this application;
[0059] Figure 28 shows a schematic diagram of a radio interface provided in an embodiment of this application;
[0060] Figure 29 shows an eighteenth schematic diagram of a radio interface display provided in an embodiment of this application;
[0061] Figure 30 shows a nineteenth schematic diagram of a radio interface display provided in an embodiment of this application;
[0062] Figure 31 shows a schematic diagram of a radio interface display provided in an embodiment of this application;
[0063] Figure 32 shows a schematic diagram of the hardware structure of another terminal device provided in an embodiment of this application. Detailed Implementation
[0064] The technical solutions of the embodiments of this application will be described below with reference to the accompanying drawings. In the description of this application, unless otherwise stated, " / " indicates that the objects before and after are in an "or" relationship. For example, A / B can represent A or B. "And / or" in this application is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone, where A and B can be singular or plural. Furthermore, in the description of this application, unless otherwise stated, "multiple" refers to two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple. Furthermore, to facilitate a clear description of the technical solutions in the embodiments of this application, the terms "first" and "second" are used in the embodiments of this application to distinguish identical or similar items with substantially the same function and effect. Those skilled in the art will understand that the terms "first" and "second" do not limit the quantity or execution order, and that "first" and "second" are not necessarily different. Meanwhile, in the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is being used as an example, illustration, or description. Any embodiment or design scheme described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design schemes. Specifically, the use of terms such as "exemplary" or "for example" is intended to present related concepts in a concrete manner for ease of understanding.
[0065] Furthermore, the business scenarios described in the embodiments of this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the emergence of new business scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
[0066] In today's world of ubiquitous mobile devices, users frequently exchange information. Traditionally, text input is the core method of interaction. Therefore, the efficiency and convenience of text input are crucial for information exchange. Traditional text input relies on a keyboard, where users manually type to input text. However, in some interaction scenarios, users may encounter inconveniences, such as when both hands are busy, making it difficult to input text on the device promptly. Even if a user has one hand free for single-handed input, this still impacts input efficiency.
[0067] Therefore, to improve users' text input efficiency, various application software input methods have successively provided speech-to-text functions. These functions acquire the user's voice information (or voice signal) through a microphone, convert it into text using a speech recognition engine, and assist users in text input. With its advantage of eliminating the need for manual input, speech-to-text functionality can help users quickly complete text input in different scenarios, becoming an important solution for improving input efficiency.
[0068] In some examples, when a user needs to convert speech to text, the speech-to-text function needs to be enabled through the corresponding component in the input method, displaying a voice input control. Then, the voice input is achieved using the voice input control. For instance, during information interaction, if a user needs to convert speech to text, they first need to access the function interface through the "+" component. Secondly, in the function interface, the voice input control is triggered to enable the speech-to-text function, and the voice input interface is displayed in the interactive interface. Finally, the user can long-press the voice input control in the voice input interface to achieve speech-to-text input, as shown in Figure 1(a). It is evident that the above speech-to-text input process is relatively complex and cumbersome. Users cannot perceive the existence of the speech-to-text function during interaction, affecting their ability to use the speech-to-text function. Furthermore, in the implementation of the aforementioned voice-to-text function, users can only input voice by "long-pressing" the microphone control. It is evident that this method is simplistic, easily causes user fatigue, is not conducive to long-term voice input, and results in a poor user experience for the voice-to-text function.
[0069] In other examples, the speech-to-text function is implemented using the microphone button on the existing text input keyboard. However, during the speech-to-text process, the display of the text input keyboard unnecessarily obstructs the content displayed on the interactive interface. Therefore, for some terminal devices with small screens (e.g., the Verde external screen), or in situations where high screen integrity is required (e.g., during gameplay), excessive occupation of the interactive interface can negatively impact the user's interaction. For example, during gameplay, a user needs to use speech-to-text to communicate with teammates. If the user clicks the microphone button on the text input keyboard to input speech-to-text, the text input keyboard will continuously obstruct the interactive interface during the speech-to-text process, as shown in Figure 1(b), which will also affect the user's game progress and, consequently, the user's experience with the speech-to-text function.
[0070] Based on the above, this application provides a speech-to-text method. When a user needs speech-to-text conversion, firstly, the user can perform an input operation on a first interactive interface to trigger the display of a first audio interface including a first audio control. It is evident that, to improve the user's perception of the speech-to-text function during information interaction, a separate first audio interface displaying the first audio control is provided, enabling the user to quickly input text using the first audio control. Secondly, the user can not only convert continuously acquired user speech information into text information by long-pressing the first audio control, but also by clicking the first audio control. This allows users to choose an appropriate triggering method to achieve the speech-to-text function based on their own circumstances, making the triggering methods of the speech-to-text function more diverse. This reduces the probability of user fatigue and improves the user experience.
[0071] This application also provides another method for speech-to-text conversion. When a user needs speech-to-text conversion, firstly, the user can trigger the display of the fourth audio interface using the text input keyboard already displayed in the second interactive interface. Secondly, the user can change the second long-press operation to a third swipe operation in the fourth audio interface. After the third swipe operation ends, voice information is continuously acquired without requiring the user to continuously perform a long-press operation, and the continuously acquired user voice information is converted into text information to achieve the speech-to-text function. It is evident that in this method, users can adjust the triggering method for the speech-to-text function at any time according to their own circumstances, making the triggering method of the speech-to-text function more diverse. This can reduce the probability of user fatigue and improve the user experience.
[0072] The terminal device 100 involved in the speech-to-text method provided in this application embodiment can be seen in Figure 2. The terminal device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a microphone 170A, a display screen 180, etc.
[0073] Processor 110 may include one or more processing units, such as: application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, memory, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU), etc. Different processing units may be independent devices or integrated into one or more processors.
[0074] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0075] The charging management module 140 receives charging input from a charger, which can be either a wireless charger or a wired charger. While charging the battery 142, the charging management module 140 can also supply power to the terminal device via the power management module 141.
[0076] The power management module 141 is used to connect the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140 to power the processor 110, internal memory 121, external memory (not shown in FIG. 2), display screen 180, and wireless communication module 160, etc.
[0077] The wireless communication function of the terminal device 100 can be implemented through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor, and baseband processor.
[0078] Antennas 1 and 2 are used to transmit and receive electromagnetic wave signals. Each antenna in terminal device 100 can be used to cover one or more communication frequency bands. Different antennas can also be multiplexed to improve antenna utilization. For example, antenna 1 can be multiplexed as a diversity antenna for a wireless local area network. In some other embodiments, the antennas can be used in conjunction with a tuning switch.
[0079] The mobile communication module 150 can provide solutions for wireless communication, including 2G / 3G / 4G / 5G, applied to the terminal device 100. The mobile communication module 150 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc.
[0080] The wireless communication module 160 can provide solutions for wireless communication applications on the terminal device 100, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR) technology, etc.
[0081] In some embodiments, the antenna 1 of the terminal device 100 is coupled to the mobile communication module 150, and the antenna 2 is coupled to the wireless communication module 160, so that the terminal device 100 can communicate with the network and other devices through wireless communication technology.
[0082] Terminal device 100 implements display functions through a GPU, display screen 180, etc. The GPU is a microprocessor for image processing, connected to the display screen 180 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. Processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.
[0083] The display screen 180 is used to display images, videos, a first radio interface, text information, etc. The display screen 180 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Mini LED, a MicroLED, a Micro-OLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the terminal device 100 may include one or N display screens 180, where N is a positive integer greater than 1.
[0084] The external storage interface 120 can be used to connect an external storage card, such as a Micro SD card, to expand the storage capacity of the terminal device 100. The external storage card communicates with the processor 110 through the external storage interface 120 to perform data storage functions. For example, it can save acquired voice information, text information converted from voice information, video files, etc., on the external storage card.
[0085] Internal memory 121 can be used to store computer executable program code, which includes instructions. Processor 110 executes various functional applications (e.g., speech-to-text) and data processing of terminal device 100 by running the instructions stored in internal memory 121. Internal memory 121 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (e.g., speech-to-text function, sound playback function, image playback function, etc.). The data storage area may store data created during the use of terminal device 100 (e.g., acquired voice information, text information converted from voice information, audio data, phonebook, etc.). Furthermore, internal memory 121 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc.
[0086] Terminal device 100 can implement audio functions through audio module 170, microphone 170A, and application processor, such as acquiring user voice information, playing music, and recording.
[0087] The audio module 170 is used to convert digital audio information into analog audio signals for output, and also to convert analog audio input into digital audio signals. The audio module 170 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 170 may be located in the processor 110, or some functional modules of the audio module 170 may be located in the processor 110.
[0088] Microphone 170A, also known as a "microphone" or "voice transducer," is used to convert sound signals into electrical signals. When making phone calls, sending voice messages, or performing speech-to-text conversion, the user can speak by bringing their mouth close to microphone 170A, inputting the sound signal into microphone 170A. Terminal device 100 may be equipped with at least one microphone 170A. In other embodiments, terminal device 100 may also be equipped with two, three, four, or more microphones 170A to achieve functions such as collecting user voice information, noise reduction, speech-to-text conversion, sound source identification, and directional recording.
[0089] The terminal device provided in this application embodiment can run an operating system (OS). This operating system can be various operating systems used in the industry, such as an operating system based on OpenHarmony, like HarmonyOS; or other operating systems such as Android. TM An operating system can refer to the iOS mobile operating system; it can also refer to various open-source operating systems or their derivatives, such as Linux OS and other embedded operating systems; or it can refer to future new operating systems, such as AI operating systems based on artificial intelligence. An operating system is a set of interconnected system software programs that manage and control the operation of terminal devices, utilize and run hardware and software resources, and provide public services to organize user interactions. In a terminal device, the operating system connects downwards to the physical hardware layer and upwards to provide a runtime environment for application software.
[0090] An operating system typically includes a kernel layer, a middleware layer, and an application layer. The application layer includes applications, which can include system applications and third-party applications, as shown in Figure 3. The middleware layer includes a series of software providing various services to application developers, or frameworks providing services such as databases, multimedia, and graphics, or capabilities such as distributed scheduling and system expansion. For example, the middleware layer may include a framework layer and / or a system service layer. The framework layer provides application programming interfaces (APIs) and programming frameworks for applications in the application layer. The system service layer includes the core capabilities of the system, providing services to applications through the framework layer. The kernel layer is the layer between hardware and software. The kernel layer may include hardware drivers and the operating system kernel. In addition to providing hardware drivers, the kernel layer also supports functions such as memory management and system process management.
[0091] The types and forms of terminal devices we use in our daily lives vary greatly, and the application scenarios for these devices are also very wide. Therefore, based on different forms and functions of terminal devices, different application scenarios, and different user needs, the operating systems used on these terminal devices may also differ. The basic functions implemented by the terminal device provided in this application can be implemented using a general-purpose operating system or a dedicated operating system. To more clearly illustrate the implementation of the embodiments of this application under a specific operating system, the architecture of HarmonyOS is shown below. Those skilled in the art can deduce the implementation of the embodiments of this application under other specific operating systems, such as Android. TM Implementation under operating systems, etc.
[0092] The software architecture of a terminal device can be divided into several layers. In some embodiments, from bottom to top, these layers are: kernel layer, system service layer, framework layer, and application layer. Layers communicate with each other through software interfaces. System functions can be tailored, added, or combined at the subsystem level depending on the deployment scenario of different device forms. Each subsystem can also be tailored, added, or combined at the functional level.
[0093] The kernel layer includes, but is not limited to: the kernel abstract layer (KAL), the kernel subsystem, and the driver subsystem.
[0094] Kernel Abstraction Layer: By shielding the differences between multiple kernels, it provides basic kernel capabilities to the upper layers, including but not limited to process / thread management, memory management, file system, network management, and peripheral device management.
[0095] Kernel Subsystem: Supports the selection of a suitable OS kernel for different resource-constrained devices, including but not limited to Linux kernel, HarmonyOS kernel, LiteOS (lite operating system), etc.
[0096] Driver Subsystem: The driver framework is the foundation for the open system hardware ecosystem, providing unified peripheral access capabilities and a framework for driver development and management. The driver framework includes: display drivers, camera drivers, audio drivers, Bluetooth drivers, sensor drivers, etc.
[0097] The system service layer comprises the core capabilities of the system, providing services to applications through the framework layer. This layer involves multiple subsystem sets, including but not limited to: basic system capability subsystem set, basic software service subsystem set, enhanced software service subsystem set, and hardware service subsystem set.
[0098] The system's basic capability subsystem set provides fundamental capabilities for the operation, scheduling, and migration of distributed applications across multiple devices. This set may include distributed soft bus, distributed data management, distributed task scheduling, and Ark multi-language runtime; it may also include multi-modal input subsystem, graphics subsystem, security subsystem, and AI business subsystem.
[0099] Basic software service subsystem set: provides public and general software services; the basic software service subsystem set may include event notification subsystem, telephone service subsystem, multimedia subsystem, etc.
[0100] Enhanced software service subsystem suite: Provides differentiated enhanced software services for different devices; the enhanced software service subsystem suite may include smart screen proprietary business subsystem, wearable proprietary business subsystem, IoT proprietary business subsystem, etc.
[0101] Hardware service subsystem set: Provides hardware services; the hardware service subsystem set may include location service subsystem, user IAM (Identity and Access Management) subsystem, wearable proprietary hardware service subsystem, biometric identification, IoT proprietary hardware service subsystem, etc.
[0102] Distributed task scheduling enables distributed service management (discovery, synchronization, registration, and invocation), supporting remote startup, remote invocation, remote connection, and migration of applications across devices.
[0103] Distributed data management enables data synchronization, data storage, data sharing, and data access across all scenarios and devices.
[0104] The distributed soft bus provides communication-related capabilities for seamless interconnection between multiple devices, including: WLAN service capabilities, Bluetooth service capabilities, soft bus, inter-process communication RPC (remote procedure call), and StarFlash communication capabilities.
[0105] Ark Multilingual Runtime is a unified compilation runtime platform designed to support the joint compilation and execution of multiple programming languages and multiple chip platforms.
[0106] The framework layer provides application programming interfaces (APIs) and programming frameworks for applications in the application layer. The framework layer includes: the ArkUI framework (which provides a complete infrastructure for UI development of system applications, including UI functions such as components, layouts, animations, and interactive events, as well as a real-time interface preview tool), the user application framework, and the Ability framework (an Ability is a lightweight application; the Ability framework schedules and manages the operation and lifecycle of Abilities). Different devices may have different operating systems, and the APIs they support may also differ.
[0107] The HarmonyOS API is a series of open capabilities provided to support HarmonyOS application development. The HarmonyOS API can be set at the framework layer or independently of the framework layer. The HarmonyOS API includes the Audio API (audio service), Push API (push service), and Account API (account service), among others.
[0108] Applications can include system apps and extended / third-party apps. System apps can include the desktop, control bar, settings, contacts, phone, camera, etc., while extended / third-party apps can include social apps, travel apps, etc.
[0109] The speech-to-text method provided in this application can be applied to smart devices such as mobile phones, tablets, laptops, and desktop computers. Alternatively, the terminal device can also include wearable devices such as head-mounted display devices (e.g., virtual reality devices, augmented reality devices, mixed reality devices), smartwatches, smart bracelets, and smart glasses. Furthermore, the terminal device can also include smart home devices (such as televisions) and in-vehicle systems (such as vehicle terminals). The mobile phone can be a foldable screen phone or a non-foldable screen phone.
[0110] The following description uses a mobile phone as an example to illustrate a speech-to-text method provided in this application. As shown in Figure 4, the method may include the following steps S401 to S405.
[0111] S401, The first interactive interface is displayed on the mobile phone.
[0112] The first interactive interface can include the chat interface of a social application, the game interface during gameplay, or the recording interface of a memo or other note-taking application. The first interactive interface can be displayed on the phone's home screen or on the phone's external screen.
[0113] S402, The mobile phone responds to the input operation performed on the first interactive interface and displays the first radio interface on the first interactive interface.
[0114] The first radio interface can be displayed on the first interactive interface in several ways. For example, the first radio interface can be displayed in an "embedded" form within the first interactive interface, as shown in Figure 5(a). Another example is that the first radio interface can also be displayed in a "floating" form within the first interactive interface, as shown in Figure 5(b). The first radio interface differs from a traditional text input keyboard; it may only include the first radio control and may not involve or include other keys, such as number keys, letter keys, symbol keys, or function keys (e.g., delete key, send key, etc.). This not only improves the user's perception of the first radio control but also reduces the amount of data the phone needs to process when rendering the first radio interface.
[0115] In some implementations, the input operation described above may include triggering a text input operation in the first interactive interface. For example, taking a memo recording interface as the first interactive interface, the input operation refers to the user clicking on a "blank" area in the recording interface to trigger a text input operation, as shown in Figure 6(a).
[0116] Alternatively, the input operation can also include triggering the input of content in the input box of the first interactive interface. For example, taking the first interactive interface as the dialogue interface of a social application, the input operation refers to the user clicking the "input box" in the dialogue interface to trigger the input of content, as shown in Figure 6(b).
[0117] In daily information interaction, users can continuously obtain voice information in at least two ways: Method 1, long press; in this case, the user needs to press and hold the first microphone control for a long time to continuously obtain voice information. Method 2, tap; in this case, the user only needs to tap the first microphone control to continuously obtain voice information.
[0118] When a user continuously acquires voice information through a long press operation, in some implementations, S403, the mobile phone responds to the long press operation 1 on the first microphone control, and continuously acquires the user's voice information during the execution of the long press operation 1. In this way, while conveniently interacting with other users (or sending "short voice" messages), the user can still continue their previous operating habits and quickly realize the voice-to-text function by performing a long press operation to conveniently complete the sending of text information.
[0119] The long press operation 1 includes a first long press operation. Long press operation 1 can refer to a touch operation performed by the user using the touch device on the mobile phone. Alternatively, long press operation 1 can also refer to a trigger operation performed by the user using an additional input device (e.g., a game controller, mouse, keyboard, etc.).
[0120] In some embodiments, when a user continuously acquires voice information through a click operation, in S404, the mobile phone responds to the click operation 1 on the first microphone control and continuously acquires the user's voice information. Thus, when the user's hands are busy or other situations make information interaction inconvenient, the voice-to-text function can be quickly and conveniently implemented through a click operation.
[0121] The click operation 1 includes a first click operation. Click operation 1 can refer to a touch operation performed by the user using the touch device on the mobile phone. Alternatively, click operation 1 can also refer to a trigger operation performed by the user using an additional input device (e.g., a game controller, mouse, keyboard, etc.).
[0122] As can be seen, when users need to convert speech to text, they can choose a suitable operation method (long press or tap) based on their own situation to trigger the continuous acquisition of voice information and realize the speech-to-text function. This makes the triggering methods for the speech-to-text function more diverse, thereby reducing the probability of user fatigue, improving the convenience of the speech-to-text function, and enhancing the user experience.
[0123] S405: The mobile phone displays text information on the first interactive interface.
[0124] The text information is converted from the speech information. In some examples, during the process of converting speech information into text, the mobile phone first collects speech information (or speech signals) through the microphone, and then the "speech recognition engine" built into the mobile phone system converts the speech information into text information for output. In some examples, the speech recognition engine can use machine learning algorithms such as deep learning algorithms to train the model with a large amount of sample speech information to improve the model's ability to recognize and convert common speech information.
[0125] Here, based on the above content, it can be seen that the first radio interface provided in this application embodiment includes at least two display forms in the first interactive interface: one is an "embedded" display; the other is a "floating" display. In some embodiments, under different display forms, after the user performs different interactive operations (including long press operation 1 and click operation 1), the interface changes presented in the first radio interface (including changes in the display effect of controls in the first radio interface) are different.
[0126] Therefore, at least four implementation methods can be adopted in this application embodiment to realize the speech-to-text function. Implementation method 1: The first radio interface is displayed in the first interactive interface in an "embedded" form, and the user's interactive operation on the first radio control is click operation 1; Implementation method 2: The first radio interface is displayed in the first interactive interface in an "embedded" form, and the user's interactive operation on the first radio control is long press operation 1; Implementation method 3: The first radio interface is displayed in the first interactive interface in a "floating" form, and the user's interactive operation on the first radio control is click operation 1; Implementation method 4: The first radio interface is displayed in the first interactive interface in a "floating" form, and the user's interactive operation on the first radio control is long press operation 1.
[0127] Implementation method 1:
[0128] In some examples, while the phone is continuously acquiring voice information in response to a click operation 1 on the first microphone control, it changes the display mode of the first microphone control from a first mode to a second mode. This way, by changing the display mode of the first microphone control, the user is notified that the phone is continuously acquiring voice information, thus reducing the probability that the user unknowingly inputs invalid voice information into the phone.
[0129] The display effect presented when the first radio control is displayed in the first mode is different from the display effect presented when the first radio control is displayed in the second mode. The display effect may include at least one of the following: display icon, display color, display position, and display size.
[0130] For example, taking a social application's dialogue interface as the first interactive interface, an input box is displayed in the dialogue interface. The user performs an input operation in the input box (e.g., clicks the input box), and the first audio interface is displayed on the dialogue interface in a "embedded" form within the input method. The first audio interface includes a first audio control, which is displayed as a first icon in the first audio interface. In response to the user's click operation 1 on the first audio control, while continuously acquiring voice information, the first audio control is displayed as a second icon in the first audio interface, as shown in Figure 7.
[0131] In some examples, before the user performs click operation 1 on the first audio interface, the first audio interface can also display prompt message 3 to prompt the user to enter the state of continuously acquiring voice information. That is to say, the first audio interface can also prompt the user to use the speech-to-text function by displaying prompt message 3. For example, prompt message 3 can be "Click or long press to convert speech to text", as shown in Figure 7.
[0132] In other examples, while the phone is continuously acquiring voice information, the first audio interface can also display prompt message 4 to guide the user on how to stop continuously acquiring voice information. For example, prompt message 4 could be "Click to stop audio acquisition," as shown in Figure 7.
[0133] In some implementations, the first radio interface includes a first switching control. While the phone continuously acquires voice information in response to a click operation 1 on the first radio control, firstly, the phone can also respond to a trigger operation on the first switching control, displaying a text input keyboard in the first interactive interface and ending the acquisition of voice information. Secondly, the text input keyboard also includes a second switching control; the phone responds to a trigger operation on the second switching control, displaying the first radio interface in the first interactive interface. This allows the user to switch between the first radio interface and the text input keyboard while performing the speech-to-text function.
[0134] The triggering operation can refer to a touch operation performed by a user using the touch device on a mobile phone, which can include single-point touch operations (e.g., a click operation) and multi-point touch operations, such as clicking the first toggle control. Alternatively, the triggering operation can also refer to an operation performed by a user using an additional input device (e.g., a game controller, mouse, keyboard, etc.).
[0135] Here, when the display size of the first interactive interface is large (e.g., the first interactive interface displayed on a tablet computer), the first radio interface and the text input keyboard can be displayed simultaneously on the first interactive interface. Conversely, when the display size of the first interactive interface is small (e.g., the first interactive interface displayed on a mobile phone), in order to reduce unnecessary obstruction of the first interactive interface, the first radio interface and the text input keyboard may not be displayed simultaneously on the first interactive interface. In this case, the first interactive interface and the text input keyboard are displayed in a "switching" manner.
[0136] In some examples, taking a social application's chat interface as the first interactive interface, the chat interface displays a first audio interface, which includes a first audio control and a first switching control. While the phone is continuously acquiring voice information in response to a click on the first audio control, the phone also responds to a click on the first switching control, switching the first audio interface to a text input keyboard within the chat interface and ending the acquisition of voice information.
[0137] In this case, the text input keyboard also includes a second switching control. In subsequent information interaction, when the user has a need for voice-to-text conversion again, the mobile phone can also respond to the click operation performed on the second switching control and switch the text input keyboard to the first voice interface in the dialogue interface, as shown in Figure 8.
[0138] In some implementations, while the mobile phone is continuously acquiring voice information through click operation 1, it responds to a received information modification instruction by modifying the first information already displayed on the first interactive interface, and then displays text information, including the modified first information, on the first interactive interface. This allows the user to edit the converted text information without stopping voice input, reducing the complexity of the text modification process and decreasing the phone's response frequency.
[0139] There are several ways to obtain information modification instructions. For example, an information modification instruction can be issued by the user via voice. Another example is that the user can issue the instruction by triggering a modification control (including a delete control) on the first audio interface. Specifically, when a user issues an information modification instruction via voice, the user can input the instruction in voice form into the phone through the microphone. When a user issues an information modification instruction by triggering a modification control, the user can send the instruction to the phone by clicking the modification control.
[0140] The first information is text information related to the information modification instruction. That is, the first information is the text information that the information modification instruction requests to be modified. The first information can be displayed in multiple locations on the first interactive interface. For example, the first information can be displayed in the input box of the first interactive interface. As another example, the first information can also be displayed in a blank space on the first interactive interface.
[0141] For example, the text message already displayed on the first interactive interface is "Happy New Year," and the message modification instruction is "Change 'Happy' to 'Joyful'." In this case, the phone responds to the message modification instruction, recognizing "Happy" as the first message and "Joyful" as the modified first message.
[0142] In some examples, the first piece of information on the initial interactive interface is highlighted before it is modified. This prompts the user that the first piece of information will be modified soon, helping the user to verify the accuracy of the information to be modified and reducing the probability of incorrect modification.
[0143] In some examples, highlighting the first piece of information can include several highlighting methods. These methods can include at least one of the following: changing the background color of the first piece of information, changing the font size of the first piece of information, and changing the font color of the first piece of information.
[0144] If the message modification instruction is issued by the user via voice message, that is, when the voice message includes the message modification instruction, in some implementations, the mobile phone displays the content of the message modification instruction in text form on the first voice interface after receiving the instruction. This facilitates the user's verification of the issued message modification instruction, allowing the user to promptly adjust the instruction in case of an error, reducing the probability of the user repeatedly modifying the displayed text information, and also reducing the phone's response frequency.
[0145] For example, taking a social application's dialogue interface as the first interactive interface, the dialogue interface displays a first audio interface. The first audio interface includes a first audio control. When the phone responds to the user's click operation 1 on the first audio control, and while continuously acquiring voice information, the phone displays the acquired voice message: "Happy New Year" in the input box. During the continuous acquisition of voice information, the phone receives a message modification instruction: "Change 'happy' to 'joyful'." In this case, the phone displays the message modification instruction in text form on the first audio interface. Furthermore, it highlights the first message (i.e., "joyful") by changing the background color and font color of the first message in the input box, as shown in Figure 9.
[0146] In some examples, the display size of the first radio interface is limited. When the information modification instruction involves a lot of content, in order to reduce the number of displayed information modification instructions and prevent obstruction of other displayed content (such as the first radio control) in the first radio interface, part of the information modification instruction can be displayed in the first listening interface. For example, only the first part (or the second part) of the information modification instruction can be displayed.
[0147] In other examples, to assist users in verifying issued information modification instructions when the instructions are lengthy, the mobile phone, upon receiving the instruction, analyzes it (e.g., through AI recognition and analysis) to obtain both the initial information and the modified initial information. By combining the initial and modified information, the content indicated by the information modification instruction is obtained and displayed on the first radio interface. This allows for a simplified display of the information modification instruction content.
[0148] For example, during the process of continuously acquiring voice information by clicking operation 1, the phone receives a message modification instruction: "Change 'happy' in 'Happy New Year' to 'joyful'." After analyzing the message modification instruction, the phone identifies "happy" as the first message and "joyful" as the modified first message. By combining "happy" and "joyful," the content indicated by the message modification instruction is obtained: "Change...'happy' to 'joyful'," which is displayed on the first voice interface, as shown in Figure 10.
[0149] Implementation Method 2:
[0150] In other examples, the first microphone control is no longer displayed on the first microphone screen while the user is performing a long press operation 1. This allows the user to be notified that the process of continuously acquiring voice information is underway by changing the displayed content on the first microphone screen. It also reduces the amount of data the phone needs to process when rendering the first microphone screen.
[0151] For example, taking a social application's chat interface as the first interactive interface, an input box is displayed in the chat interface. The user performs an input operation in the input box (e.g., clicks the input box), and the first microphone interface is displayed on the chat interface in a "embedded" form within the input method. The first microphone interface includes a first microphone control. When the phone responds to the user performing a long press operation 1 on the first microphone control, the first microphone control is no longer displayed in the first microphone interface during the long press operation 1, as shown in Figure 11.
[0152] In some examples, before the user performs a long press operation 1 on the first audio interface, the first audio interface can also display a prompt message 3 to prompt the user to enter the state of continuously acquiring voice information. That is to say, the first audio interface can also prompt the user to use the speech-to-text function by displaying prompt message 3. For example, prompt message 3 can be "Click or long press to convert speech to text", as shown in Figure 11.
[0153] In other examples, while the phone is responding to the long press operation 1 and continuously acquiring voice information, the first audio interface can also display a prompt message 4 to guide the user on how to end the continuous acquisition of voice information. For example, prompt message 4 could be "Release to end audio acquisition," as shown in Figure 11.
[0154] In one implementation, after the phone responds to the end of the long press operation 1 and stops continuously acquiring voice information, a first microphone control and a first switching control are displayed on the first microphone interface. In response to a trigger operation on the first switching control, the phone displays a text input keyboard on the first interactive interface. In some examples, the text input keyboard also includes a second switching control; in response to a trigger operation on the second switching control, the phone displays the first microphone interface on the first interactive interface. This allows users to quickly switch back to their preferred text input keyboard after completing the voice-to-text function.
[0155] Here, the display method between the first radio keypad and the text input keypad is the same as that in Implementation Method 1.
[0156] In some examples, taking a social application's chat interface as the first interactive interface, the chat interface displays a first audio interface, which includes a first audio control and a first switching control. During the process of continuously acquiring voice information while the phone responds to a long press operation 1 on the first audio control, the first audio control and the first switching control are no longer displayed. Upon the end of the long press operation 1, the phone resumes displaying the first audio control and the first switching control on the first audio interface. In this case, the phone responds to the user's click operation on the first switching control, switching the first audio interface to a text input keyboard in the chat interface.
[0157] In this case, the text input keyboard also includes a second switching control. During subsequent information interaction, when the user has a need for voice-to-text conversion again, the mobile phone can also respond to the click operation performed on the second switching control, and switch the text input keyboard to the first voice interface in the dialogue interface, as shown in Figure 12.
[0158] In one implementation, while the mobile phone is continuously acquiring voice information through a long press operation 1, it responds to an information modification instruction by modifying the first information already displayed on the first interactive interface, and then displays text information including the modified first information on the first interactive interface. The method for modifying the text information is the same as that in implementation 1, and will not be described again here.
[0159] For example, taking a social application's dialogue interface as the first interactive interface, the dialogue interface displays a first audio interface. The first audio interface includes a first audio control. During the process of the phone responding to the user's long-press operation 1 on the first audio control and continuously acquiring voice information, the phone displays the acquired voice information: "Happy New Year" in the input box. During the continuous long-press operation 1 (that is, during the continuous acquisition of voice information), the phone acquires the information modification instruction "change 'happy' to 'joyful'". In this case, the phone displays the content of the information modification instruction in text form on the first audio interface. It also highlights the first information (i.e., "happy") by changing the background color and font color of the first information in the input box. Finally, it displays the text information including the modified first information: "Happy New Year", as shown in Figure 13.
[0160] In some examples, when the information modification instruction involves a lot of content, the way to display the information modification instruction in a thumbnail or in part is the same as in Implementation 1, and will not be repeated here.
[0161] Implementation Method 3:
[0162] In some examples, while the phone is continuously acquiring voice information in response to a click operation 1 on the first microphone control, it changes the display mode of the first microphone control from a first mode to a second mode. This way, by changing the display mode of the first microphone control, the user is notified that the phone is continuously acquiring voice information, thus reducing the probability that the user unknowingly inputs invalid voice information into the phone.
[0163] The display effect when the first radio control is displayed in the first mode and the display effect when the first radio control is displayed in the second mode are the same as the display effect in Implementation Method 1, and will not be described again here.
[0164] For example, taking the memo recording interface as the first interactive interface, in response to an input operation performed on the recording interface (e.g., clicking on a blank area of the recording interface), the mobile phone displays the first radio interface in a "floating" form on the recording interface. The first radio interface includes a first radio control, which is displayed as a first icon on the first radio interface. In response to the user performing a click operation 1 on the first radio control, while continuously acquiring voice information, the first radio control is displayed as a second icon on the first radio interface, as shown in Figure 14.
[0165] In some examples, before the user performs click operation 1 on the first audio interface, the first audio interface can also display prompt message 3 to prompt the user to enter the state of continuously acquiring voice information. That is to say, the first audio interface can also prompt the user to use the speech-to-text function by displaying prompt message 3. For example, prompt message 3 can be "Click or long press to convert speech to text", as shown in Figure 14.
[0166] In other examples, while the phone is continuously acquiring voice information, the first audio interface can also display prompt message 4 to guide the user on how to stop continuously acquiring voice information. For example, prompt message 4 could be "Click to stop audio acquisition," as shown in Figure 14.
[0167] In some implementations, the first radio interface includes a first switching control. While the phone continuously acquires voice information in response to a click operation 1 on the first radio control, firstly, the phone can also respond to a trigger operation on the first switching control, displaying a text input keyboard in the first interactive interface and ending the acquisition of voice information. Secondly, the text input keyboard also includes a second switching control; the phone responds to a trigger operation on the second switching control, displaying the first radio interface in the first interactive interface. This allows the user to switch between the first radio interface and the text input keyboard while performing the speech-to-text function.
[0168] The display method between the first radio keyboard and the text input keyboard is the same as that in Implementation Method 1, and will not be described again here.
[0169] In some examples, taking the memo recording interface as an example, the recording interface displays a first audio interface, which includes a first audio control and a first switching control. While the phone is continuously acquiring voice information in response to a click operation on the first audio control, the phone also responds to a click operation on the first switching control, switching the first audio interface to a text input keyboard in the dialogue interface.
[0170] In this case, the text input keyboard also includes a second switching control. During subsequent information interaction, when the user has a need for voice-to-text conversion again, the mobile phone can also respond to the click operation performed on the second switching control and switch the text input keyboard to the first voice interface in the dialogue interface, as shown in Figure 15.
[0171] In some implementations, while the mobile phone is continuously acquiring voice information through click operation 1, it responds to a received information modification instruction by modifying the first information already displayed on the first interactive interface, and then displays text information, including the modified first information, on the first interactive interface. This allows the user to edit the converted text information without stopping voice input, reducing the complexity of text information modification operations and lowering the phone's response frequency.
[0172] The various methods of obtaining the information modification instruction, the display of the content of the information modification instruction, the display position of the first information, and the display effect of the first information are all the same as in Implementation Method 1, and will not be repeated here.
[0173] For example, taking the memo recording interface as the first interactive interface, the recording interface displays a first audio interface. The first audio interface includes a first audio control. When the phone responds to the click operation 1 on the first audio control and continuously acquires voice information, the acquired voice information "..., black pepper powder" is displayed in the recording interface. During the continuous acquisition of voice information, the phone acquires the information modification instruction "change 'black pepper powder' to 'white peppercorns'". In this case, the phone displays the content of the information modification instruction in the first audio interface in the form of text: "change 'black pepper powder' to 'white peppercorns'". The phone also highlights the first information (i.e., "black pepper powder") by changing the display background color and font color of the first information (i.e., "black pepper powder") in the recording interface. The text information including the modified first information, "white peppercorns", is then displayed in the recording interface, as shown in Figure 16.
[0174] In some examples, when the information modification instruction involves a lot of content, the way to display the information modification instruction in a thumbnail or in part is the same as in Implementation 1, and will not be repeated here.
[0175] Implementation Method 4:
[0176] In some examples, firstly, the phone displays the first radio interface. Secondly, in response to a long press operation 1 on the first radio control, the phone stops displaying the first radio control on the first radio interface while the user is performing the long press operation 1. This allows the user to be informed that the phone is continuously acquiring voice information by changing the content displayed on the first radio interface. It also reduces the amount of data the phone needs to process when rendering the first radio interface.
[0177] For example, taking a memo recording interface as the first interactive interface, the user performs an input operation in the recording interface (e.g., clicking on a blank area of the recording interface), triggering the first radio interface to be displayed in a "floating" form on the recording interface. During the user's long press operation 1, (without needing to jump to the second radio interface, or before the second radio interface appears) the first radio interface no longer displays the first radio control, as shown in Figure 17(a). The first radio interface can also display prompt information 4 to prompt the user on how to end continuous acquisition of voice information. For example, prompt information 4 can be "Release to end radio," as shown in Figure 17(a).
[0178] In other examples, firstly, in response to a long press operation 1 on the first microphone control, the phone displays a second microphone interface in the first interactive interface. Secondly, after displaying the second microphone interface, if the long press operation 1 is not yet finished, the phone continues to acquire voice information. In this way, while facilitating information interaction with other users (or sending "short voice messages"), users can still maintain their previous operating habits and quickly achieve voice-to-text conversion by performing a single long press operation to conveniently send text information.
[0179] Furthermore, during the user's long-press operation 1, the first microphone control is no longer displayed in the second microphone interface (in this case, other controls, such as voice editing controls, can be displayed). This allows the user to be alerted that the phone is currently in a continuous voice information acquisition phase by changing the displayed content of the second microphone interface. Simultaneously, it reduces the amount of data the phone needs to process when rendering the second microphone interface.
[0180] For example, taking the memo recording interface as the first interactive interface, the user performs an input operation in the recording interface (e.g., clicks on a blank area of the recording interface), triggering the first radio interface to be displayed in a "floating" form on the recording interface. The first radio interface includes a first radio control. When the user performs a long press operation 1 on the first radio control, a second radio interface is displayed. During the process of the user performing long press operation 1, the first radio control is no longer displayed in the second radio interface, as shown in Figure 17(b).
[0181] In some examples, before the user performs a long press operation 1 in the first audio interface, the first audio interface can also display a prompt message 3 to prompt the user to enter the state of continuously acquiring voice information. That is to say, the first audio interface can also prompt the user to implement the speech-to-text function by displaying prompt message 3. For example, prompt message 3 can be "Click or long press to convert speech to text", as shown in Figure 17(b).
[0182] In other examples, after the phone responds to the long press operation 1 and displays the second radio interface, if the long press operation 1 is not ended, the phone continues to acquire voice information, and the second radio interface can also display prompt message 4 to prompt the user on how to end the continuous acquisition of voice information. For example, prompt message 4 can be "Release to end radio," as shown in Figure 17(b).
[0183] In other examples, if the long press operation 1 is not finished and the phone is still acquiring voice information, the second audio interface can also display a prompt message 5 to indicate to the user that it is currently in the speech-to-text stage. For example, prompt message 5 can be "Speech to text in progress...", as shown in Figure 17(b).
[0184] In some examples, a cancel input control can also be displayed in the second radio interface, as shown in Figure 17(b). The phone responds to the triggering operation of the cancel input control (e.g., a click operation), discarding the acquired voice information.
[0185] In some implementations, after displaying the second audio interface, the phone, in response to the end of the long press operation 1, restores the first audio interface to the first interactive interface and stops acquiring voice information. Thus, after the user finishes using the speech-to-text function, the phone can automatically switch back to the first audio interface without requiring any additional user intervention. This allows the user to reselect an appropriate trigger operation (click, long press, etc.) during subsequent information interactions to re-enable the speech-to-text function.
[0186] For example, taking the memo recording interface as the first interactive interface, the recording interface displays a first audio interface, which includes a first audio control. When the phone responds to a long press operation 1 on the first audio control, a second audio interface is displayed, and the first audio control is no longer displayed on the second audio interface. If the long press operation 1 ends, the phone responds to the end of the long press operation 1 by switching the second audio interface back to the first audio interface and ending the acquisition of voice information, as shown in Figure 18.
[0187] The second radio interface is displayed in a "floating" manner within the first interactive interface. The first and second radio interfaces alternately appear in the first interactive interface through a "switching" mechanism.
[0188] In some implementations, the second audio interface includes a voice editing control. After displaying the second audio interface, the phone responds by changing the long press operation 1 to a swipe operation 1, and the swipe operation 1 ends at the voice editing control, obtaining the information modification command. In this way, the user can actively trigger the modification of text information during the speech-to-text process using the voice editing control, reducing the probability that the phone mistakenly identifies interactive information as an information modification command when modifying text information by recognizing the modification command itself.
[0189] The sliding operation 1 includes a first sliding operation. Sliding operation 1 can refer to a touch operation performed by the user using the touch device on the mobile phone. Alternatively, sliding operation 1 can also refer to a trigger operation performed by the user using an additional input device (e.g., a game controller, mouse, keyboard, etc.).
[0190] For example, taking the memo recording interface as the first interactive interface, in response to an input operation (e.g., clicking on a blank area of the recording interface), the phone displays the first audio interface in a "floating" form on the recording interface. The first audio interface includes a first audio control. In response to a long press operation 1 on the first audio control, the phone displays the second audio interface on the recording interface. If the long press operation 1 continues, and the phone is continuously acquiring voice information, the acquired voice information, "...black pepper powder," will be displayed on the recording interface.
[0191] In some examples, the second radio interface includes a voice editing control. In response to a long press operation 1 changing to a swipe operation 1 on the second radio interface, and the swipe operation 1 ending at the voice editing control, the phone receives the message modification instruction "change 'black pepper powder' to 'white peppercorns'". In this case, the phone displays the message modification instruction in text form on the second radio interface: "change 'black pepper powder' to 'white peppercorns'". Furthermore, it highlights the first message (i.e., "black pepper powder") by changing the background and font color of the first message in the recording interface, as shown in Figure 19.
[0192] In some implementations, the voice editing control is displayed in a third manner on the second audio interface. When the phone responds to a sliding operation 1 on the voice editing control and acquires voice modification information, it changes the display mode of the voice editing control from the third mode to a fourth mode. This allows the user to be notified that a voice modification command is currently being acquired by changing the display mode of the voice editing control.
[0193] The display effect of the voice editing control when displayed in the third mode differs from the display effect when displayed in the fourth mode. The display effect may include at least one of the following: display icon, display color, display position, and display size.
[0194] In some implementations, firstly, the phone responds to a long press operation 1 on the first microphone control and displays a second microphone interface. Secondly, the phone responds to the long press operation 1 on the second microphone interface changing to a swipe operation 2. After the swipe operation 2 ends, a third microphone interface is displayed, and voice information is continuously acquired. This allows users to adjust the triggering method for the voice-to-text function at any time according to their own situation during the voice-to-text process, making the triggering methods of the voice-to-text function more diverse. This reduces the probability of user fatigue and improves the user experience. Furthermore, the size of the third microphone interface is smaller than that of the second microphone interface, thus reducing the obstruction of the first interactive interface by the microphone interface during continuous acquisition of voice information.
[0195] The swipe operation 2 includes a second swipe operation. Swipe operation 2 can refer to an operation performed by the user using the touchscreen on the mobile phone. Alternatively, swipe operation 2 can also refer to an operation performed by the user using an additional input device (e.g., a game controller, mouse, keyboard, etc.).
[0196] In some examples, the swipe operation 3 includes at least two swipe methods: one is a swipe operation in a specified direction; the other is a swipe operation on a specified control. Swiping in the specified direction can include swiping upwards, swiping downwards, etc. The specified control can include the fourth radio control included in the second radio interface, etc.
[0197] In some examples, in the second radio interface, the phone can also respond to a trigger operation (e.g., a tap) on the fourth radio control to pause continuous acquisition of voice information. Alternatively, while voice acquisition is paused, the phone can also respond to a trigger operation (e.g., a tap) on the fourth radio control to resume continuous acquisition of voice information.
[0198] The third radio interface is displayed in a "floating" manner within the first interactive interface. The second and third radio interfaces alternately appear in the first interactive interface using a "switching" mechanism.
[0199] For example, taking the memo recording interface as the first interactive interface, a second audio interface is displayed in the recording interface. The phone responds to a long press operation 1 on the second audio interface by changing to a swipe operation 2 (e.g., an upward swipe). After the swipe operation 2 ends, a third audio interface is displayed, and voice information is continuously acquired. As can be seen from Figure 20, the size of the third audio interface is smaller than the size of the second audio interface.
[0200] In some examples, the third audio interface can also display a prompt message 5 to indicate to the user that the speech-to-text conversion is currently in progress. For example, prompt message 5 could be "Speech-to-text in progress...", as shown in Figure 20.
[0201] In other examples, the third radio interface may also include a third switching control. During the continuous acquisition of voice information, the mobile phone responds to the trigger operation of the third switching control (e.g., click operation), switches the third radio interface displayed in the recording interface to the first radio interface, and ends the acquisition of voice information, as shown in Figure 20.
[0202] In some implementations, the third audio interface includes a voice editing control. When the phone responds to a change from a long press (1) to a swipe (2) on the second audio interface, and continues to acquire voice information, it responds to trigger operations on the voice editing control to obtain information modification instructions. This allows users to actively modify text information during speech-to-text conversion using the voice editing control, reducing the probability that the phone might mistakenly identify interactive information as modification instructions when modifying text information through recognition of such instructions.
[0203] In some examples, the voice editing control is displayed in a third mode on the third audio interface. As the phone continuously acquires voice modification information in response to triggering operations on the voice editing control (e.g., click operations), it changes the display mode of the voice editing control from the third mode to a fourth mode. This allows the user to be notified that voice modification instructions are currently being acquired by changing the display mode of the voice editing control.
[0204] The display effect of the voice editing control when displayed in the third mode differs from the display effect when displayed in the fourth mode. The display effect may include at least one of the following: display icon, display color, display position, display size, and display background.
[0205] For example, taking the memo recording interface as the first interactive interface, a second audio interface is displayed in the recording interface. The phone responds to a long press operation 1 on the second audio interface changing to a swipe operation 2 (e.g., an upward swipe). After swipe operation 2 ends, a third audio interface is displayed, showing the continuously acquired voice information: "...black pepper powder" in the recording interface. In some examples, the third audio interface may also include a voice editing control, which is displayed with a first display background. The phone responds to a trigger operation (e.g., a click operation) on the voice editing control in the third audio interface, acquiring the information modification instruction "change 'black pepper powder' to 'white peppercorns'". In this case, the phone displays the voice editing control with a second display background in the third audio interface, as shown in Figure 21.
[0206] In some examples, when the information modification instruction involves a lot of content, the way to display the information modification instruction in a thumbnail or in part is the same as in Implementation 1, and will not be repeated here.
[0207] The following describes another speech-to-text method provided in this application embodiment, using a mobile phone as an example of the terminal device. As shown in Figure 22, the method may include the following steps S2201 to S2204.
[0208] S2201, The second interactive interface is displayed on the mobile phone.
[0209] The second interactive interface includes a text input keyboard. This interface can include the chat interface of a social application, the game interface during gameplay, or the note-taking interface of a memo or other note-taking application. The second interactive interface can be displayed on the phone's home screen or on the phone's external screen.
[0210] S2202, In response to a long press operation 2 on the first key of the text input keyboard, the mobile phone displays the fourth radio interface on the second interactive interface.
[0211] The first key can be any key on the text input keyboard, such as number keys, letter keys, symbol keys, function keys (e.g., delete key, arrow keys, space key, send key, enter key, etc.). For example, the fourth radio interface can be displayed in an "embedded" form within the second interactive interface, as shown in Figure 23.
[0212] Long press operation 2 includes a second long press operation. Long press operation 2 can refer to an operation performed by the user using the touch device on the mobile phone. Alternatively, long press operation 2 can also refer to an operation performed by the user using an additional input device (e.g., a game controller, mouse, keyboard, etc.).
[0213] In some implementations, firstly, in response to a long press operation 2 on the fourth audio interface, the phone continuously acquires voice information during the long press operation 2. Secondly, in response to the end of the long press operation 2, a text input keyboard is displayed on the second interactive interface, and the acquisition of voice information ceases. Thus, after the user completes the use of the voice-to-text function, the phone can automatically switch back to the first audio interface, allowing the user to reselect a suitable text input method in subsequent information interactions to again achieve the voice-to-text function.
[0214] In some examples, taking the second interactive interface as a memo recording interface and the first key as the space bar, the recording interface displays a text input keyboard. When the phone responds by performing a long press operation 2 on the space bar of the text input keyboard, the fourth audio interface is displayed in the recording interface, and voice information is continuously acquired. If the long press operation 2 ends, the phone responds by switching the fourth audio interface back to the text input keyboard and ending the acquisition of voice information, as shown in Figure 24.
[0215] In other examples, after the phone responds to the long press operation 2 and displays the fourth radio interface, if the long press operation 2 is not ended and the phone is still continuously acquiring voice information, the fourth radio interface can also display a prompt message 4 to guide the user on how to stop continuously acquiring voice information. For example, the prompt message 4 could be "Release to stop radio reception," as shown in Figure 24.
[0216] In other examples, after the phone responds to the long press operation 2 and displays the fourth audio interface, if the long press operation 2 is not ended and the phone is still acquiring voice information, the fourth audio interface can also display a prompt message 5 to indicate to the user that it is currently in the speech-to-text stage. For example, prompt message 5 may include "Speech to text in progress...", as shown in Figure 24. Alternatively, prompt message 5 may also include a specific icon (not shown in the figure) to describe the current speech-to-text process.
[0217] S2203: The phone responds by changing the long press operation 2 on the fourth radio interface to a swipe operation 3. After the swipe operation 3 ends, voice information is continuously acquired. This allows users to adjust the triggering method for the voice-to-text function at any time according to their own situation during the voice-to-text process, making the triggering method of the voice-to-text function more diverse. This reduces the probability of user fatigue and improves the user experience.
[0218] The sliding operation 3 includes a third sliding operation. Sliding operation 3 can refer to an operation performed by the user using the touchscreen on the mobile phone. Alternatively, sliding operation 3 can also refer to an operation performed by the user using an additional input device (e.g., a game controller, mouse, keyboard, etc.).
[0219] In some examples, the swipe operation 3 includes at least two swipe methods: one is a swipe operation in a specified direction; the other is a swipe operation on a specified control. The swipe in the specified direction can include swiping upwards, swiping downwards, etc. The specified control can include the second radio control included in the second radio interface, etc.
[0220] When swiping operation 3 is a swipe operation onto a specified control, in some embodiments, the fourth microphone interface includes a second microphone control. The phone responds to the swiping operation 3 ending at the second microphone control and continues to acquire voice information.
[0221] For example, taking the memo recording interface as the second interactive interface, a fourth audio interface is displayed in the recording interface. The fourth audio interface includes a second audio control. The phone responds to the long press operation 2 on the fourth audio interface by changing to a swipe operation 3, and the swipe operation 3 ends at the second audio control, continuously acquiring voice information, as shown in Figure 25.
[0222] When the swipe operation 3 is a swipe operation in a specified direction, in some other embodiments, the mobile phone continues to acquire voice information after swiping in the specified direction in response to the swipe operation 3.
[0223] For example, taking the memo recording interface as the second interactive interface, a fourth audio interface is displayed in the recording interface. The phone responds to the long press operation 2 on the fourth audio interface by changing to a swipe operation 3 (e.g., an upward swipe operation), and continues to acquire voice information after the swipe operation 3 ends.
[0224] In some examples, after the phone finishes swiping operation 3 and continues to acquire voice information, the fourth audio interface also includes a third audio control to prompt the user that it is currently in the process of continuously acquiring voice information, as shown in Figure 26(a).
[0225] In other examples, the fourth audio interface can also display a prompt message 5 to indicate to the user that the speech-to-text process is in progress. For example, prompt message 5 could be "Speech-to-text in progress...", as shown in Figure 26(a). Alternatively, prompt message 5 can also include a specific icon (not shown in the figure) to describe the current speech-to-text process.
[0226] In other examples, the fourth radio interface includes a voice editing control. The phone responds by changing a long press (2) to a swipe (4) on the fourth radio interface, and the swipe ends at the voice editing control. The phone then receives the message modification instruction "change 'black pepper powder' to 'white peppercorns'". In this case, the phone displays the message modification instruction in text form on the fourth radio interface: "change 'black pepper powder' to 'white peppercorns'". The phone also highlights the first message (i.e., "black pepper powder") by changing its background and font color, as shown in Figure 26(b).
[0227] In other examples, the fourth radio interface includes a cancel input control, as shown in Figure 26(b). The phone responds to a triggering action on the cancel input control (e.g., a click action), discarding the acquired voice information.
[0228] In some implementations, the voice editing control is displayed in a third manner on the fourth audio interface. When the phone responds to a sliding operation 4 on the voice editing control and acquires voice modification information, it changes the display mode of the voice editing control from the third to the fourth manner. This allows the user to be notified that a voice modification command is currently being acquired by changing the display mode of the voice editing control.
[0229] The display effect of the voice editing control when displayed in the third mode differs from the display effect when displayed in the fourth mode. The display effect may include at least one of the following: display icon, display color, display position, and display size.
[0230] In other examples, the phone can also respond to a triggering action (e.g., a tap) on a third microphone control to pause continuous acquisition of voice information. Alternatively, while voice acquisition is paused, the phone can also respond to a triggering action (e.g., a tap) on a third microphone control to resume continuous acquisition of voice information.
[0231] In some implementations, the fourth audio interface includes a third switching control. In response to the swipe operation 3 ending at the second audio control, while continuously acquiring voice information, the phone can also respond to a trigger operation on the third switching control, displaying a text input keyboard in the second interactive interface and ending the acquisition of voice information. This allows the user to switch between the fourth audio interface and the text input keyboard while performing the speech-to-text function.
[0232] Here, the fourth radio interface and the text input keyboard are alternately displayed in the second interactive interface in a "switch" manner.
[0233] In some examples, taking the second interactive interface as the memo recording interface as an example, the recording interface displays a fourth audio interface, which includes a third switching control displayed as a third icon. During the continuous acquisition of voice information, the mobile phone responds to the user's trigger operation on the third switching control displayed as a third icon, switches the fourth audio interface to a text input keyboard, and ends the acquisition of voice information, as shown in Figure 27(a).
[0234] In some examples, taking the second interactive interface as a memo recording interface as an example, the recording interface displays a fourth audio interface, which includes a third switching control displayed as a fourth icon. When the phone responds to a long press operation 2 to a swipe operation 3 (e.g., an upward swipe operation), and after the swipe operation 3 ends, while continuously acquiring voice information, the phone responds to the user's trigger operation on the third switching control displayed as a fourth icon, switching the fourth audio interface to a text input keyboard and ending the acquisition of voice information, as shown in Figure 27(b).
[0235] S2204: The mobile phone displays text information on the second interactive interface.
[0236] The text information is converted from the speech information. In the process of converting speech information into text, the mobile phone first collects speech information (or speech signals) through the microphone, and then the "speech recognition engine" built into the phone system converts the speech information into text information for output. In some examples, the speech recognition engine can use machine learning algorithms such as deep learning algorithms to train the model using a large amount of sample speech information during the process of converting speech information into text, thereby improving the model's ability to recognize and convert common speech information.
[0237] In some implementations, during the continuous acquisition of voice information, the mobile phone responds to a received information modification command by modifying the second information already displayed on the second interactive interface, and then displays text information, including the modified second information, on the second interactive interface. This allows the user to edit the converted text information without stopping voice input, reducing the complexity of the operation and decreasing the phone's response frequency.
[0238] There are several ways to obtain information modification instructions. For example, an information modification instruction can be issued by the user via voice. In another embodiment, the information modification instruction can also be issued by the user by triggering a modification control (including a delete control) in the second audio interface. Specifically, when a user issues an information modification instruction via voice, the user can input the instruction in voice form into the phone through the microphone. When a user issues an information modification instruction by triggering a modification control, the user can send the instruction to the phone by tapping the modification control.
[0239] The second information is text information related to the information modification instruction. That is, the second information is the text information that the information modification instruction requests to be modified. The second information can be displayed in multiple locations on the second interactive interface. For example, the second information can be displayed in the input box of the second interactive interface. As another example, the second information can also be displayed in a blank space on the second interactive interface.
[0240] For example, the text message already displayed in the second interactive interface is "Happy New Year," and the message modification instruction is "Change 'Happy' to 'Joyful.'" In this case, the phone responds to the message modification instruction, recognizing "Happy" as the second message and "Joyful" as the modified second message.
[0241] In some examples, the second piece of information on the second interactive interface is highlighted before it is modified. This prompts the user that the second piece of information will be modified soon, helping the user to verify the accuracy of the information to be modified and reducing the probability of incorrect modification.
[0242] In some examples, highlighting the second information can include several highlighting methods. These methods can include at least one of: changing the background color of the second information, changing the font size of the second information, and changing the font color of the second information.
[0243] If the information modification instruction is issued by the user via voice message, that is, when the voice message includes the information modification instruction, in some implementations, after receiving the information modification instruction, the mobile phone displays the content of the information modification instruction in text form on the fourth audio interface. This facilitates the user's verification of the issued information modification instruction, allowing the user to promptly adjust the information modification instruction in case of an error, reducing the probability of the user repeatedly modifying already displayed text information, and also reducing the phone's response frequency.
[0244] For example, taking the memo recording interface as the second interactive interface, a text input keyboard is displayed in the recording interface. The phone responds to a long press operation 2 on the space bar of the text input keyboard, displays the fourth audio interface, and continuously acquires voice information, displaying the acquired voice information: "..., black pepper powder" in the recording interface. During the continuous acquisition of voice information, the phone acquires the information modification instruction "change 'black pepper powder' to 'white peppercorns'". In this case, the phone displays the content of the information modification instruction in text form in the fourth audio interface. Furthermore, by changing the background color and font color of the first information (i.e., "black pepper powder") in the recording interface, the first information is highlighted, as shown in Figure 28.
[0245] In some examples, the display size of the fourth radio interface is limited. When the information modification instruction involves a lot of content, to reduce the amount of information modification instruction displayed and prevent it from obscuring other display content (such as the second radio control), only part of the information modification instruction can be displayed on the fourth radio interface. For example, only the first part (or the last part) of the information modification instruction can be displayed. In other examples, to assist the user in verifying the issued information modification instruction when the content of the instruction is large, the mobile phone responds to the received information modification instruction, analyzes the instruction, and obtains the second information and the modified second information. By combining the second information and the modified second information, the content indicated by the information modification instruction is obtained and displayed on the fourth radio interface. This achieves a thumbnail display of the information modification instruction content.
[0246] For example, during the continuous acquisition of voice information, the mobile phone receives an information modification instruction: "Change 'black pepper powder' in 'sea salt, black pepper powder' to 'white peppercorns'." After analyzing the information modification instruction, the mobile phone identifies "black pepper powder" as the second piece of information and "white peppercorns" as the modified second piece of information. By combining "black pepper powder" and "white peppercorns," the content indicated by the information modification instruction is obtained: "Change...'black pepper powder' to 'white peppercorns'," which is displayed in the fourth voice interface, as shown in Figure 29.
[0247] In some implementations, while the phone is continuously acquiring voice information, it responds to a trigger operation on the second microphone control, pauses voice information acquisition, and displays prompt message 1 on the fourth microphone interface. This allows for timely prompting to the user when the phone is in a paused voice information acquisition state, reducing the probability of the user outputting invalid voice information while in a paused state.
[0248] The prompt message 1 includes a first prompt message, which is used to inform the user that the process of the mobile phone continuously acquiring voice information has been paused.
[0249] The trigger operation can refer to a touch operation performed by the user using the touch device on the mobile phone, which can include single-point touch operations (e.g., tapping the left) and multi-point touch operations, such as tapping the second microphone control. Alternatively, the trigger operation can also refer to an operation performed by the user using an additional input device (e.g., a game controller, mouse, keyboard, etc.).
[0250] For example, taking the memo recording interface as the second interactive interface, a fourth audio interface is displayed within the recording interface, which includes the second audio control. While the phone is continuously acquiring voice information by changing from a long press operation 2 to a swipe operation 3, the phone pauses voice information acquisition in response to a click operation performed on the second audio control. The fourth audio interface then displays the prompt message "Pause audio, click to resume continuous audio."
[0251] In other examples, when the phone is paused, it resumes continuous voice reception in response to a click on the second microphone control. The fourth microphone interface then displays the message "Click to pause microphone," as shown in Figure 30.
[0252] In some implementations, if the phone fails to acquire voice information within a preset time period during the continuous acquisition of voice information, a prompt message 2 is displayed on the fourth audio interface. This allows the user to be prompted with the available voice input information even before they are familiar with the voice input rules, thus assisting them in using the voice-to-text function and improving the user experience.
[0253] The prompt message 2 includes a second prompt message, which indicates that voice input is allowed in the second interactive interface. The prompt message 3 may include various prompts, such as "Try saying 'delete…' to delete the corresponding content," "Try saying 'add…' to add the corresponding content," "Try saying 'replace… with…' to modify the corresponding content," etc. These various prompts are displayed in a loop in the second interactive interface.
[0254] The preset duration can be pre-set, such as 30 seconds. It can also be set based on the user's input habits. In some examples, when the user's speech is relatively fluent, the preset duration can be relatively short, such as 10 seconds. When the user habitually pauses during speech, the preset duration can be relatively long, such as 30 seconds.
[0255] For example, taking the second interactive interface as the memo recording interface, the fourth audio interface is displayed in the recording interface. During the continuous acquisition of voice information, if the phone does not acquire voice information within a preset time, the fourth audio interface displays prompt message 2 "Try saying 'delete…' to delete the corresponding content", as shown in Figure 31.
[0256] In other examples, prompt 2 may also be displayed as "Try saying 'Add...' to add the corresponding content" (not shown in the figure), or "Try saying 'Replace... with...' to modify the corresponding content" (not shown in the figure).
[0257] In some solutions, multiple embodiments of this application and multiple implementation methods within those embodiments can be combined and implemented as a combined solution. Optionally, some operations in the processes of each method embodiment may be arbitrarily combined, and / or the order of some operations may be arbitrarily changed. Furthermore, the execution order between the steps of each process is merely exemplary and does not constitute a limitation on the execution order between steps; other execution orders are also possible. It is not intended to indicate that the execution order is the only possible order in which these operations can be performed. Those skilled in the art will conceive of various ways to reorder the operations described in the embodiments of this application. In addition, it should be noted that the process details involved in a certain embodiment of this application are also applicable to other embodiments in a similar manner, or different embodiments may be combined.
[0258] Furthermore, some steps in the method embodiments can be equivalently replaced with other possible steps. Alternatively, some steps in the method embodiments may be optional and can be deleted in certain use cases. Or, other possible steps may be added to the method embodiments. Moreover, the various method embodiments can be implemented individually or in combination.
[0259] It is understood that, in order to achieve the above functions, the aforementioned terminal device includes hardware and / or software modules corresponding to perform each function. Based on the algorithm steps of the various examples described in conjunction with the embodiments disclosed herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application in conjunction with the embodiments, but such implementation should not be considered beyond the scope of this application.
[0260] This application embodiment can divide the terminal device into functional modules according to the above method example. For example, each function can be divided into its own functional module, or two or more functions can be integrated into one processing module. The integrated module can be implemented in hardware. It should be noted that the module division in this embodiment is illustrative and only represents one logical functional division. In actual implementation, there may be other division methods.
[0261] This application embodiment also provides a terminal device, as shown in FIG32. The terminal device may include one or more processors 3201, memory 3202, and communication interface 3203. The memory 3202 and communication interface 3203 are coupled to the processor 3201. For example, the memory 3202, communication interface 3203, and processor 3201 may be coupled together via a bus 3204. The communication interface 3203 is used for data transmission with other devices. The memory 3202 stores computer program code. The computer program code includes computer instructions, which, when executed by the processor 3201, cause the terminal device to perform the speech-to-text method described in this application embodiment.
[0262] The processor 3201 can be a processor or controller, such as a central processing unit (CPU), a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute the various exemplary logic blocks, modules, and circuits described in connection with this disclosure. The processor can also be a combination that implements computational functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0263] Bus 3204 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Bus 3204 can be categorized as an address bus, data bus, control bus, etc. For ease of representation, only one thick line is used in Figure 32, but this does not indicate that there is only one bus or one type of bus.
[0264] This application also provides a computer-readable storage medium storing computer program code. When the processor executes the computer program code, the terminal device executes the relevant method steps in the above method embodiments. This application also provides a chip including a memory and one or more processors; the memory is coupled to the processors; wherein the memory stores computer program code, which includes computer instructions. When the processor executes the computer instructions, the chip executes the relevant method steps in the above method embodiments. The terminal device, computer storage medium, and chip provided in this application are all used to execute the corresponding methods provided above. Therefore, the beneficial effects they achieve can be referred to the beneficial effects in the corresponding methods provided above, and will not be repeated here. Through the description of the above embodiments, those skilled in the art can clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0265] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0266] The units described as separate components may or may not be physically separate. A component shown as a unit can be one or more physical units, meaning it can be located in one place or distributed across multiple different locations. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs. Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated units described above can be implemented in hardware or as software functional units.
[0267] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiments of this application, in essence, or the part that contributes, or all or part of the technical solution, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0268] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for speech-to-text conversion, characterized in that, Applied to a terminal device, the method includes: Display the first interactive interface; In response to an input operation performed on the first interactive interface, a first radio interface is displayed on the first interactive interface; the first radio interface includes a first radio control. In response to a first long press operation on the first radio control, voice information is continuously acquired during the execution of the first long press operation; or, In response to a first click operation on the first radio control, voice information is continuously acquired; Text information is displayed on the first interactive interface; the text information is converted from the voice information.
2. The method according to claim 1, characterized in that, During the continuous acquisition of voice information, the method further includes: In response to receiving an information modification instruction, the first information displayed on the first interactive interface is modified; the first information is text information related to the information modification instruction. The first interactive interface displays text information, including the modified first information.
3. The method according to claim 2, characterized in that, The first information on the first interactive interface is displayed with a highlighted effect before it is modified.
4. The method according to claim 2 or 3, characterized in that, The voice information includes the information modification instruction. After receiving the information modification instruction, the method further includes: The content of the information modification instruction is displayed in text form on the first radio interface.
5. The method according to any one of claims 1-4, characterized in that, During the continuous acquisition of voice information in response to a first click operation on the first microphone control, the display mode of the first microphone control changes from a first mode to a second mode; or, During the execution of the first long press operation, the first radio control is not displayed.
6. The method according to any one of claims 1-5, characterized in that, The first radio interface includes a first switching control; in response to a first click operation on the first radio control, during the continuous acquisition of voice information, the method further includes: In response to a trigger operation on the first switching control, a text input keyboard is displayed in the first interactive interface, and the acquisition of voice information ends; the text input keyboard includes a second switching control; In response to a trigger operation on the second switching control, the first radio interface is displayed in the first interactive interface.
7. The method according to any one of claims 1-6, characterized in that, In response to a first long press operation on the first microphone control, during the first long press operation, voice information is continuously acquired, including: In response to a first long press operation on the first radio control, a second radio interface is displayed in the first interactive interface; After the second radio interface is displayed, if the first long press operation has not ended, voice information will continue to be acquired.
8. The method according to claim 7, characterized in that, The method further includes: After the second radio interface is displayed, in response to the end of the first long press operation, the first radio interface is displayed in the first interactive interface, and the acquisition of voice information ends.
9. The method according to claim 7 or 8, characterized in that, The second radio interface includes voice editing controls; the method further includes: After the second radio interface is displayed, in response to the first long press operation changing to the first swipe operation and the first swipe operation ending at the voice editing control, an information modification instruction is obtained.
10. The method according to any one of claims 1-6, characterized in that, The method further includes: In response to a first long press operation on the first radio control, a second radio interface is displayed in the first interactive interface; In response to the first long press operation on the second radio interface changing to a second swipe operation, after the second swipe operation ends, a third radio interface is displayed and voice information is continuously acquired; the size of the third radio interface is smaller than the size of the second radio interface.
11. The method according to claim 10, characterized in that, The third audio interface includes a voice editing control; during the continuous acquisition of voice information, the method further includes: In response to a trigger operation on the voice editing control, obtain information modification instructions.
12. The method according to any one of claims 1-11, characterized in that, The input operation includes: triggering a text input operation on the first interactive interface; or, triggering a text input operation in the input box displayed on the first interactive interface.
13. A method for speech-to-text conversion, characterized in that, Applied to a terminal device, the method includes: Display a second interactive interface; the second interactive interface includes a text input keyboard; In response to a second long press operation on the first key of the text input keyboard, a fourth radio interface is displayed on the second interactive interface; In response to the second long press operation on the fourth radio interface changing to a third swipe operation, voice information is continuously acquired after the third swipe operation ends. Text information is displayed on the second interactive interface; the text information is converted from the voice information.
14. The method according to claim 13, characterized in that, During the continuous acquisition of voice information, the method further includes: In response to receiving an information modification instruction, the second information displayed on the second interactive interface is modified; the second information is text information related to the information modification instruction. The second interactive interface displays text information, including the modified second information.
15. The method according to claim 14, characterized in that, The second information on the second interactive interface is displayed with a highlighted effect before it is modified.
16. The method according to claim 14 or 15, characterized in that, The voice information includes the information modification instruction. After receiving the information modification instruction, the method further includes: The content of the information modification instruction is displayed in text form on the fourth radio interface.
17. The method according to any one of claims 13-16, characterized in that, The fourth radio interface includes a second radio control; the third sliding operation ends at the second radio control.
18. The method according to claim 17, characterized in that, During the continuous acquisition of voice information, the method further includes: In response to the trigger operation of the second radio control, the acquisition of voice information is paused, and a first prompt message is displayed on the fourth radio interface; the first prompt message is used to indicate that the process of acquiring voice information has been paused.
19. The method according to any one of claims 13-18, characterized in that, The method further includes: In response to the second long press operation on the fourth radio interface, voice information is continuously acquired during the execution of the second long press operation; In response to the end of the second long press operation, the text input keyboard is displayed on the second interactive interface.
20. The method according to any one of claims 13-19, characterized in that, During the continuous acquisition of voice information, the method further includes: If no voice information is obtained within the preset time period, a second prompt message will be displayed on the second voice interface; the second prompt message is used to prompt for input of voice information.
21. A terminal device, characterized in that, The device includes a memory and one or more processors; the memory is coupled to the processors; the memory stores computer program code, the computer program code including computer instructions, which, when executed by the processor, cause the terminal device to perform the speech-to-text method as described in any one of claims 1-12, or to perform the speech-to-text method as described in any one of claims 13-20.
22. A computer-readable storage medium, characterized in that, The method includes computer instructions that, when executed on a terminal device, cause the terminal device to perform the speech-to-text method as described in any one of claims 1-12, or to perform the speech-to-text method as described in any one of claims 13-20.
23. A computer program product, characterized in that, When the computer program product is run on a terminal device, the terminal device performs the speech-to-text method as described in any one of claims 1-12, or performs the speech-to-text method as described in any one of claims 13-20.