Speech interaction method and electronic device

By providing multiple voice input methods and modes in electronic devices, and performing operations in the corresponding mode according to the user's selection, the problem of voice input errors is solved, and the accuracy and convenience of voice input are improved. It is suitable for both large-screen and small-screen devices.

WO2025261416A1PCT designated stage Publication Date: 2025-12-26HUAWEI TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/101867
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-20
Filing Date
2025-06-18
Publication Date
2025-12-26

AI Technical Summary

Technical Problem

Voice input is easily affected by homophones, user accents, and environmental noise on both large and small screen devices, leading to recognition errors and increasing the complexity of manual editing for users. Especially in scenarios where it is inconvenient to manually operate the screen, the accuracy and convenience of voice input are insufficient.

Method used

It offers a variety of voice input methods and modes. By detecting the user's selection, the electronic device performs operations in the corresponding mode, improving the accuracy and convenience of voice input. These modes include voice-based text editing, querying, creation, translation, shooting, traditional Chinese input, writing assistance, memorization, and cross-application target information retrieval.

Benefits of technology

It enables users to interact directly with electronic devices via voice input in different scenarios, freeing up their hands, improving the accuracy and convenience of voice input, and reducing manual editing operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025101867_26122025_PF_FP_ABST
    Figure CN2025101867_26122025_PF_FP_ABST
Patent Text Reader

Abstract

A speech interaction method, which is used for improving the accuracy and convenience of speech input modes. In the speech interaction method, an electronic device provides, on a display screen, selection entries for multiple speech input methods for a user to choose from, wherein the multiple speech input methods may correspond to multiple speech input modes on a one-to-one basis; and on the basis of a speech input method chosen by the user, the user can perform speech interaction with the electronic device in a corresponding speech input mode, so as to realize intelligent control over the electronic device, thereby freeing both hands of the user and improving the accuracy and convenience of speech input. Further provided is an electronic device.
Need to check novelty before this filing date? Find Prior Art

Description

A voice interaction method and electronic device

[0001] Cross-references to related applications

[0002] This application claims priority to Chinese Patent Application No. 202410805543.4, filed on June 20, 2024, entitled "A Voice Interaction Method and Electronic Device", the entire contents of which are incorporated herein by reference. Technical Field

[0003] This application relates to the field of electronic device technology, and in particular to a voice interaction method and an electronic device. Background Technology

[0004] Currently, voice input has become the primary interaction method for large-screen devices such as foldable phones and tablets. However, due to factors such as homophones, user accents, or environmental noise, electronic devices often encounter conversion errors when recognizing user voice input and converting it into corresponding text. In such cases, users need to manually edit the text to correct the errors.

[0005] Using voice input and manual correction of errors on the aforementioned large-screen devices would increase the complexity of user operation. Similar issues exist in smaller-screen devices (such as smart wearables) or in-vehicle terminals, where manual screen operation is inconvenient.

[0006] Therefore, improving the accuracy and convenience of voice input remains an important issue that urgently needs to be addressed. Summary of the Invention

[0007] This application provides a voice interaction method and an electronic device for improving the accuracy and convenience of voice input.

[0008] Firstly, this application provides a voice interaction method. The execution subject of this method is an electronic device that supports multiple voice input modes and has a display screen. In this method, the electronic device provides selection entries for multiple voice input methods on the display screen, with each voice input method corresponding to one of the multiple voice input modes; it detects a user's selection operation on the multiple voice input methods; it determines a first voice input method from the multiple voice input methods based on the selection operation, and the first voice input method corresponds to a first voice input mode among the multiple voice input modes; under the first voice input mode, it performs a first operation based on the first voice information input by the user.

[0009] Using the above method, electronic devices can perform corresponding operations based on the user's voice input information in the user's selected voice input mode, realizing voice interaction with the user. This method can effectively recognize the user's voice input requests and respond accordingly, thereby freeing the user's hands and improving the accuracy and convenience of voice input.

[0010] In one possible design, the multiple voice input modes include at least one of the following: a voice-based direct input mode; a voice-based text editing mode; a voice-based text query mode; a voice-based text creation mode; a voice-based translation input mode; a voice-based photo input mode; a voice-based traditional Chinese character input mode; a voice-based assisted writing input mode; a voice-based memory input mode; a voice-based mode for acquiring cross-application target information; and a voice-based mode for instructing cross-application target operations. It should be understood that this is merely an illustrative example of voice input modes that may be involved in the embodiments of this application and is not intended to limit the scope of the invention. In other embodiments, under different business scenarios, the electronic device may also provide users with other voice input modes, which will not be elaborated upon here.

[0011] In one possible design, if the first voice input mode is a voice-based text editing mode, the first voice information describes the first character in at least one of the following ways: common character combination description method, character decomposition description method, or semantic description method; the first operation based on the first voice information input by the user includes: replacing the second character displayed on the display screen with the first character, wherein the first character and the second character are homophones or near-homophones.

[0012] Using the above method, users can provide character instructions to electronic devices via voice input, enabling the devices to perform text editing operations based on the user's voice instructions. This frees up the user's hands and improves the accuracy and convenience of voice input.

[0013] In one possible design, if the first voice input mode is a mode for acquiring cross-application target information based on voice, the step of performing the first operation based on the first voice information input by the user in the first voice input mode includes: receiving the first voice information input by the user on a first interface provided by the electronic device in the first voice input mode; acquiring target information from a second application based on the first voice information input by the user; and displaying the target information on the first interface.

[0014] Using the above method, users can control applications across different applications via voice input without having to manually switch between them, reducing user operations and improving the accuracy and convenience of voice input.

[0015] In one possible design, providing a selection entry for multiple voice input methods on the display screen includes: providing a first button on the display screen; detecting a second operation by the user on the first button; and, based on the second operation, providing a selection entry for the multiple voice input methods on the display screen, wherein the first button is any one of a plurality of buttons, and the plurality of buttons correspond one-to-one with the multiple voice input methods. Exemplarily, the first button is a physical button; or, the first button is a virtual button.

[0016] The above method illustrates several ways to present selection entry points, such as displaying them directly or in response to user actions, and does not constitute any limitation.

[0017] In one possible design, the various voice input methods correspond one-to-one with a number of buttons, which can be physical buttons or virtual buttons.

[0018] In one possible design, the selection entry for the multiple voice input methods is included in the function menu provided by the electronic device.

[0019] In one possible design, performing the first operation based on the first voice information input by the user includes: obtaining the semantic analysis result of the first voice information; and performing the first operation based on the semantic analysis result.

[0020] In one possible design, obtaining the semantic analysis result of the first voice information includes: sending a first request message to a server, the first request message including the first voice information; and receiving first response information from the server, the first response information including the semantic analysis result of the first voice information.

[0021] In one possible design, a large-scale language model is deployed on the server, which is used to perform semantic analysis on the first speech information to obtain the semantic analysis result.

[0022] In one possible design, performing the first operation based on the first voice information input by the user further includes: acquiring first text information, the first text information including a speech recognition result of the first voice information; and displaying the first text information on the display screen.

[0023] In one possible design, obtaining the first text information includes: sending a second request message to a server, the second request message including the first voice information; and receiving second response information from the server, the second response information including the first text information obtained by recognizing the first voice information.

[0024] A second aspect provides an electronic device comprising a plurality of functional modules; the plurality of functional modules interact to implement the methods performed by the first electronic device in the first aspect and its embodiments described above. The plurality of functional modules can be implemented based on software, hardware, or a combination of software and hardware, and the plurality of functional modules can be arbitrarily combined or divided based on specific implementations.

[0025] A third aspect provides an apparatus comprising at least one processor and at least one memory, wherein the at least one memory stores computer program instructions, and when the apparatus is in operation, the at least one processor performs the method performed by the first electronic device described in the first aspect and its embodiments.

[0026] The fourth aspect also provides a program product that, when run on a device, causes the device to perform the method executed by the first electronic device in any of the above aspects and embodiments.

[0027] The fifth aspect also provides a readable storage medium storing a program that, when executed by a device, causes the device to perform the method executed by the first electronic device in any of the above aspects and embodiments.

[0028] A sixth aspect also provides a chip for reading a program stored in a memory and executing the method performed by the first electronic device in any of the above aspects and embodiments.

[0029] A seventh aspect also provides a chip system including a processor for supporting a device in performing the methods executed by the first electronic device in any of the above aspects and embodiments. In one possible design, the chip system further includes a memory for storing the necessary programs and data. The chip system may be composed of chips or may include chips and other discrete devices.

[0030] It should be noted that the beneficial effects of the various designs of the electronic devices provided in the second to seventh aspects of the embodiments of this application can be referred to the beneficial effects of any possible design in the first aspect, and will not be repeated here. Attached Figure Description

[0031] Figure 1 shows a schematic diagram of a manual editing method for a large-screen device;

[0032] Figure 2 shows a schematic diagram of the hardware structure of a possible electronic device;

[0033] Figure 3 is a software architecture block diagram of an electronic device provided in an embodiment of this application;

[0034] Figure 4 is a schematic diagram of a system architecture that may be applicable to a voice interaction method provided in an embodiment of this application;

[0035] Figure 5 is a schematic diagram of a scenario in which a voice interaction method provided in an embodiment of this application may be applied;

[0036] Figure 6a is a schematic diagram of another scenario in which a voice interaction method provided in the embodiments of this application may be applied;

[0037] Figure 6b is a schematic diagram of another scenario in which the voice interaction method provided in the embodiments of this application may be applied;

[0038] Figure 7 is a schematic diagram of another scenario in which a voice interaction method provided in the embodiments of this application may be applied;

[0039] Figure 8 is a schematic diagram of another scenario in which a voice interaction method provided in the embodiments of this application may be applied;

[0040] Figure 9 is a schematic diagram of another scenario in which a voice interaction method provided in the embodiments of this application may be applied;

[0041] Figure 10 is a schematic diagram of another scenario in which a voice interaction method provided in the embodiments of this application may be applied;

[0042] Figures 11-13 are schematic diagrams illustrating the methods for providing selection entry points for various voice input methods according to embodiments of this application;

[0043] Figures 14-18 are schematic flowcharts illustrating different examples of a voice interaction method provided in the embodiments of this application. Detailed Implementation

[0044] The embodiments of this application will now be described in detail with reference to the accompanying drawings and examples.

[0045] Currently, electronic devices are used more and more frequently in people's work and daily lives. Electronic devices with smaller screens are generally called small-screen devices, such as smart wearable watches. Electronic devices with larger screens are generally called large-screen devices, such as foldable phones and tablets.

[0046] To provide users with a better immersive experience, the screen size of some electronic devices is gradually increasing. At the same time, the manual control of electronic devices by users is decreasing.

[0047] As shown in Figure 1, taking a foldable phone as an example, when the foldable screen is unfolded along its folding axis, more text can be displayed on the foldable screen (e.g., represented by symbols like "XXX", "YYY", and "ZZZ"), making it easier for users to read and improving the reading experience. However, when the foldable screen is unfolded, the horizontal screen size is relatively wide. When editing text at a certain location on the screen, the keyboard is divided into two parts on either side of the screen. Users need to manually input or modify text by using both hands to operate the keys on the left and right sides of the keyboard, causing significant inconvenience. Illustratively, in Figure 1, the text "ZZZZZZZZZ" at the marked location is replaced with the pinyin "kao'lv". As the user clicks the corresponding pinyin k, a, o, l, and v on the keyboard, multiple selectable characters appear on the foldable screen. The user can then click "consider" to select the replacement text.

[0048] Given the problems with manual input in the aforementioned electronic devices, voice input (or voice-controlled input) has gradually become an important interaction method.

[0049] However, in scenarios using voice input, electronic devices also need to recognize the user's voice input and convert it into corresponding text for display. Due to factors such as homophones, user accents, or environmental noise, electronic devices often encounter conversion errors when recognizing and converting user-input voice into text. In such cases, the user needs to manually edit and correct the erroneous text. Using voice input and manual error correction in large-screen devices increases the complexity of user operation. Similar problems exist in smaller-screen devices (such as smart wearables) or in-vehicle terminals where manual screen operation is inconvenient.

[0050] Therefore, improving the accuracy and convenience of voice input remains an important issue that urgently needs to be addressed.

[0051] To address the aforementioned issues, this application provides a voice interaction method and an electronic device. In this method, the electronic device can offer users a variety of voice input methods to choose from. These voice input methods correspond one-to-one with various voice input modes. Based on the selected voice input method, the user can interact with the electronic device via voice in the corresponding voice input mode, thereby achieving intelligent control of the electronic device, freeing the user's hands, and improving the accuracy and convenience of voice input.

[0052] The technical solutions in this application can be applied to electronic devices, which can be any device with a display screen or associated display screen. For example, electronic devices can be mobile phones, foldable phones, tablets, wearable devices (e.g., watches, bracelets, glasses, etc.), in-vehicle devices, augmented reality (AR) / virtual reality (VR) devices, laptops, ultra-mobile personal computers (UMPCs), netbooks, personal digital assistants (PDAs), smart home devices (e.g., smart TVs, etc.). This application does not limit the specific type of electronic device.

[0053] The electronic devices to which this application can be applied can also be portable terminal devices that include other functions such as personal digital assistants and / or music players. Or electronic devices with other operating systems.

[0054] Figure 2 illustrates a possible hardware structure diagram of an electronic device. The electronic device 200 includes components such as a radio frequency (RF) circuit 210, a power supply 220, a processor 230, a memory 240, an input unit 250, a display unit 260, an audio circuit 270, a communication interface 280, and a wireless fidelity (Wi-Fi) module 290. Those skilled in the art will understand that the hardware structure of the electronic device 200 shown in Figure 2 does not constitute a limitation on the electronic device 200. The electronic device 200 provided in this application embodiment may include more or fewer components than shown, may combine two or more components, or may have different component configurations. The various components shown in Figure 2 can be implemented in hardware, software, or a combination of hardware and software, including one or more signal processing and / or application-specific integrated circuits.

[0055] The following is a detailed description of each component of the electronic device 200 with reference to Figure 2:

[0056] The RF circuit 210 can be used for receiving and transmitting data during communication or a call. Specifically, after receiving downlink data from the base station, the RF circuit 210 sends it to the processor 230 for processing; additionally, it sends uplink data to be transmitted to the base station. Typically, the RF circuit 210 includes, but is not limited to, an antenna, at least one amplifier, a transceiver, a coupler, a low-noise amplifier (LNA), a duplexer, etc.

[0057] Furthermore, the RF circuit 210 can also communicate with other devices via a wireless communication network. The wireless communication can use any communication standard or protocol, including but not limited to Global System for Mobile Communication (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, and Short Message Service (SMS).

[0058] Wi-Fi technology is a short-range wireless transmission technology. The electronic device 200 can connect to an access point (AP) via the Wi-Fi module 290, thereby enabling access to the data network. The Wi-Fi module 290 can be used for receiving and sending data during communication.

[0059] The electronic device 200 can physically connect to other devices through the communication interface 280. Optionally, the communication interface 280 can be connected to the communication interfaces of other devices via a cable to enable data transmission between the electronic device 200 and other devices.

[0060] The electronic device 200 can also perform communication services and interact with other electronic devices. Therefore, the electronic device 200 needs to have data transmission capabilities, meaning it needs to include a communication module. Although Figure 2 shows the RF circuit 210, the Wi-Fi module 290, and the communication interface 280, it is understood that the electronic device 200 contains at least one of the aforementioned components or other communication modules (such as a Bluetooth module) for data transmission.

[0061] For example, when the electronic device 200 is a mobile phone, the electronic device 200 may include the RF circuit 210, the Wi-Fi module 290, or a Bluetooth module (not shown in Figure 2); when the electronic device 200 is a tablet computer, the electronic device 200 may include the Wi-Fi module or a Bluetooth module (not shown in Figure 2); when the electronic device 200 is a smart home device, the electronic device 200 may include the Wi-Fi module 290 or a Bluetooth module (not shown in Figure 2).

[0062] The memory 240 can be used to store software programs and modules. The processor 230 executes various functional applications and data processing of the electronic device 200 by running the software programs and modules stored in the memory 240. Optionally, the memory 240 may mainly include a program storage area and a data storage area. The program storage area may store the operating system (mainly including the software programs or modules corresponding to the kernel layer, system layer, application framework layer, and application layer).

[0063] In addition, the memory 240 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device.

[0064] The input unit 250 can be used to receive editing operations on various types of data objects, such as numbers or characters, input by the user, and to generate key signal inputs related to user settings and function control of the electronic device 200. Optionally, the input unit 250 may include a touch panel 251 and other input devices 252.

[0065] The touch panel 251, also known as a touchscreen, can collect user touch operations on or near it (such as operations performed by the user using a finger, stylus, or any suitable object or accessory on or near the touch panel 251), and drive corresponding connection devices according to a pre-set program. In this embodiment, the touch panel 251 can collect user operations on it. For example, the user operation may include clicking a physical button on the touch panel 251, or the user operation may also include pressing and holding a physical button on the touch panel 251 to switch the physical button to another function.

[0066] Optionally, the other input device 252 may include, but is not limited to, one or more of the following: a physical keyboard, an infrared sensor, function keys (such as volume control buttons, power buttons, etc.), a trackball, a mouse, a joystick, etc. For example, an infrared sensor can be used to acquire the user's air gesture operations.

[0067] The display unit 260 can be used to display information input by the user or information provided to the user, as well as various menus of the electronic device 200. The display unit 260 is the display system of the electronic device 200, used to present the interface and realize human-computer interaction. The display unit 260 may include a display panel 261. Optionally, the display panel 261 can be configured using a liquid crystal display (LCD), organic light-emitting diode (OLED), or other similar forms. In this embodiment, the display unit 260 can be used to display a user interface. For example, the user interface can be the desktop (or main interface) of the electronic device, or it can be an interface displaying various possible business scenarios, such as the chat interface of an instant messaging app, a memo interface, or the interface of a multimedia video app.

[0068] The processor 230 is the control center of the electronic device 200. It connects various components via various interfaces and lines, and executes software programs and / or modules stored in the memory 240, as well as calling data stored in the memory 240, to perform various functions and process data of the electronic device 200, thereby enabling various services based on the electronic device 200. In this embodiment, the processor 230 can be used to implement a voice interaction method provided in this embodiment.

[0069] The electronic device 200 also includes a power supply 220 (such as a battery) for supplying power to various components. Optionally, the power supply 220 can be logically connected to the processor 230 through a power management system, thereby enabling the power management system to manage functions such as charging, discharging, and power consumption.

[0070] As shown in Figure 2, the electronic device 200 also includes an audio circuit 270, a microphone 271, and a speaker 272, providing an audio interface between the user and the electronic device 200. The audio circuit 270 converts audio data into signals recognizable by the speaker 272 and transmits the signals to the speaker 272, where the speaker 272 converts them into sound signals for output. The microphone 271 collects external sound signals (such as human speech or other sounds) and converts the collected external sound signals into signals recognizable by the audio circuit 270, sending them to the audio circuit 270. The audio circuit 270 can also convert the signals transmitted by the microphone 271 into audio data, and then output the audio data to the RF circuit 210 for transmission to, for example, another electronic device, or output the audio data to the memory 240 for further processing.

[0071] Although not shown in Figure 2, the electronic device 200 may also include a camera, at least one sensor, etc., which will not be described in detail here. The at least one sensor may include, but is not limited to, a pressure sensor, a barometric pressure sensor, an accelerometer, a distance sensor, a fingerprint sensor, a touch sensor, a temperature sensor, etc.

[0072] Figure 3 is a software architecture block diagram of an electronic device provided in an embodiment of this application. As shown in Figure 3, the software structure of the electronic device can be a layered architecture. For example, the software can be divided into several layers, each with a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the operating system is divided into four layers, from top to bottom: the application layer, the application framework layer (framework, FWK), the runtime and system libraries, and the kernel layer.

[0073] The application layer may include a series of application packages. As shown in Figure 3, the application layer may include a camera, settings, skin modules, user interface (UI), third-party applications, etc. Third-party applications may include gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, SMS, etc. In this embodiment, the application layer may include a target installation package of a target application that the electronic device requests to download from a server. The function files and layout files in this target installation package are adapted to the electronic device.

[0074] The application framework layer provides application programming interfaces (APIs) and a programming framework for applications in the application layer. The application framework layer can include some predefined functions. As shown in Figure 3, the application framework layer can include a window manager, content provider, view system, phone manager, resource manager, and notification manager.

[0075] The window manager is used to manage windowed applications. It can obtain the screen size, determine if a status bar is present, lock the screen, and capture screenshots. The content provider stores and retrieves data, making this data accessible to applications. This data may include videos, images, audio, made and received phone calls, browsing history and bookmarks, phone books, etc.

[0076] A view system includes visual controls, such as controls for displaying text and controls for displaying images. View systems can be used to build applications. A display interface can consist of one or more views. For example, a display interface including a text notification icon could include views for displaying text and views for displaying images.

[0077] A phone manager is used to provide communication functions for electronic devices. For example, it manages call status (including connection and disconnection).

[0078] The file explorer provides applications with various resources, such as localized strings, icons, images, layout files, video files, and more.

[0079] The notification manager allows applications to display notifications in the status bar. These notifications can be used to deliver informational messages and can disappear automatically after a short pause, requiring no user interaction. For example, the notification manager can be used to notify users of completed downloads or message alerts. The notification manager can also display notifications as icons or scrolling text in the top status bar, such as notifications from background applications, or as dialog boxes on the screen. Examples include displaying text messages in the status bar, emitting sounds, vibrating electronic devices, and flashing indicator lights.

[0080] The runtime includes the core libraries and the virtual machine. The runtime is responsible for the scheduling and management of the operating system.

[0081] The core library consists of two parts: one part contains the functionalities that the computer programming language needs to call, and the other part is the core library of the operating system. The application layer and application framework layer run in a virtual machine. Taking Java as the programming language as an example, the virtual machine executes the Java files in the application layer and application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.

[0082] System libraries can include multiple functional modules. For example: surface manager, media libraries, 3D graphics processing libraries (e.g., OpenGL ES), 2D graphics engines (e.g., SGL), image processing libraries, etc.

[0083] The Surface Manager is used to manage the display subsystem and provides the fusion of two-dimensional (2D) and three-dimensional (3D) layers for multiple applications.

[0084] The media library supports playback and recording of various common audio and video formats, as well as still image files. It supports multiple audio and video encoding formats, such as MPEG4, H.264, MP3, AAC, AMR, JPG, and PNG.

[0085] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.

[0086] A 2D graphics engine is a drawing engine for 2D graphics. It can perform drawing operations, such as drawing different voice buttons or different shapes of voice buttons on the screen of an electronic device to receive different voice input information from the user in different voice input scenarios. Alternatively, a 2D graphics engine can also draw menus on the display screen of an electronic device, providing users with selection options for various voice input methods and offering corresponding functionality.

[0087] The kernel layer is the layer between hardware and software. The kernel layer contains at least the display driver, camera driver, audio driver, and sensor driver.

[0088] The hardware layer can include various types of sensors, such as accelerometers, gyroscopes, and touch sensors.

[0089] It should be noted that the structures shown in Figures 2 and 3 are merely examples of electronic devices provided in the embodiments of this application, and cannot be used to limit the electronic devices provided in the embodiments of this application. In specific implementations, electronic devices may have more or fewer devices or modules than those shown in Figures 2 or 3.

[0090] Typically, an electronic device 200 can run multiple applications simultaneously. In a simpler scenario, one application corresponds to one process; in a more complex scenario, one application can correspond to multiple processes. Each process has a unique process ID.

[0091] It should be understood that in the embodiments of this application, "at least one of the following" or similar expressions refer to any combination of these items, including any combination of a single item or a plurality of items. For example, at least one of a, b, or c can represent: a, b, c, a and b, a and c, b and c, or a, b, and c, where a, b, and c can be single or multiple. "Multiple" refers to two or more. "And / or" is used to describe the association relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship.

[0092] Furthermore, it should be understood that in the description of this application, terms such as "first" and "second" are used only for distinguishing purposes and should not be construed as indicating or implying relative importance, nor should they be construed as indicating or implying order. For example, in the following embodiments, "first voice input method" and "second voice input method" are only used to distinguish different voice input methods and are not used to limit the specific implementation method of the voice input method, etc.

[0093] It should be understood that the hardware structure of the electronic device can be as shown in Figure 2, and the software system architecture can be as shown in Figure 3. The software programs and / or modules corresponding to the software system architecture in the electronic device can be stored in the memory 240, and the processor 230 can run the software programs and applications stored in the memory 240 to execute the flow of a voice interaction method provided in the embodiments of this application.

[0094] Figure 4 shows a schematic diagram of a system architecture that may be applicable to the voice interaction method provided in the embodiments of this application.

[0095] As shown in Figure 4A, the system architecture may include electronic device 200, user 41, and server 42.

[0096] The electronic device 200 can receive voice information input by the user 41. Based on the voice information, the electronic device 200 can send a request message to the server 42 through a communication network, requesting the server 42 to perform voice recognition, semantic analysis (or semantic understanding), and other processing on the voice information.

[0097] For example, server 42 can be equipped with an automatic speech recognition (ASR) module, which can perform speech recognition on the speech information from electronic device 200 and convert the speech information into corresponding text information to obtain the speech recognition result. Alternatively, server 42 can be equipped with a large language model (referred to as the large model), which can perform semantic analysis on the speech information from electronic device 200 and obtain the corresponding semantic analysis result.

[0098] Server 42 can send the speech recognition result or semantic analysis result as a response to the aforementioned request message to electronic device 200. Accordingly, electronic device 200 can receive the response information from server 42, display the received response information on a display screen, and / or perform corresponding operations based on the response information.

[0099] In another possible implementation, as shown in Figure 4B, the voice interaction method can also be implemented by the electronic device 200 itself, as described in this embodiment. For example, the electronic device 200 may have an ASR module deployed therein, which can be used to perform speech recognition on the voice information input by the user 41 to obtain the converted text information. Alternatively, the electronic device 200 may have a large model deployed therein, which can perform semantic analysis on the voice information input by the user 41 to obtain the corresponding semantic analysis results. The electronic device 200 can display the speech recognition results or semantic analysis results on its display screen, and can also perform corresponding operations based on the speech recognition results or semantic analysis results.

[0100] It should be understood that, in the embodiments of this application, in the scenario shown in Figure 4A, the ASR module and the large model can be deployed on the same server or on different servers. The electronic device can directly interact with the corresponding server, or indirectly interact with the corresponding server through forwarding from other devices. For example, the electronic device can send a request message containing voice information to the application server of the currently running application. The application server can forward the received request message to the server deploying the large model. The server deploying the large model performs semantic analysis on the received voice information and returns the semantic understanding result to the application server. The application server can encapsulate the semantic analysis result and feed it back to the electronic device. This embodiment of the application does not limit this.

[0101] Based on the structural diagrams in Figures 2 and 3 and the system architecture diagram in Figure 4, the following uses a mobile phone as an example to introduce several possible voice interaction scenarios involved in the embodiments of this application.

[0102] In this embodiment, the electronic device can provide users with multiple voice input method selection options on the display screen, with each voice input method corresponding to a different voice input mode. If a user selects a particular voice input method, the electronic device can perform the corresponding operation based on the user's voice input within the corresponding voice input mode.

[0103] As an example, multiple voice input modes may include at least one of the following: voice-based direct input mode; voice-based text editing mode; voice-based text query mode; voice-based text creation mode; voice-based translation input mode; voice-based photo input mode; voice-based traditional Chinese input mode; voice-based assisted writing input mode; voice-based memory input mode; voice-based mode for obtaining cross-application target information; and voice-based mode for instructing cross-application target operations.

[0104] Among them, the voice-based direct input mode is the traditional voice input mode. Users can choose the voice-based direct input mode. In the voice-based direct input mode, users can input various types of voice information into the electronic device, and the electronic device can perform voice recognition and display the user's voice information.

[0105] Voice-based text editing mode is a mode that allows editing of text content or formatting based on voice input. Users can select voice-based text editing mode. In voice-based text editing mode, the electronic device can edit the text information displayed on the screen based on the user's voice input and obtain the edited text information. The text information displayed on the screen can be the text information displayed on the interface of reading applications or other applications, or it can be the text information obtained after recognizing the user's voice input. This application embodiment does not limit the method of obtaining the modified text information. Editing operations on text content may include, but are not limited to, at least one of the operations of adding, deleting, or modifying text content. Editing text formatting may include text wrapping, first-line indentation, etc. This application embodiment does not specifically limit the editing operations involved in the voice-based text editing mode.

[0106] Voice-based text query mode is a mode that enables text information query operations based on voice input. Users can choose the voice-based text query mode. In the voice-based text query mode, the electronic device can query the number of words (or characters) in the text information based on the user's voice input, or it can query the location of a specific character or word in the text information, or it can query the frequency of a specific character or word, etc. This application embodiment does not specifically limit this query operation.

[0107] Voice-based text creation mode is a mode for creating and processing text information based on voice input. Users can choose voice-based text creation mode. In this mode, users can provide creation requirements to the electronic device via voice input. The electronic device can then perform text creation, polishing, or rewriting based on the user's requirements to obtain text information that meets the user's needs. Here, when the electronic device performs text creation, polishing, or rewriting, it can be done using the auxiliary functions of a specific application or through a cross-application approach using an application with artificial intelligence capabilities. This application embodiment does not limit this approach.

[0108] Voice-based translation input mode is a text translation processing mode that uses voice input. Users can choose this mode. In this mode, users can input voice information into the electronic device in one language, and the device can recognize the voice information, translate it into text information in another language, and display it. Voice-based translation input mode can translate between any two languages, including but not limited to Chinese, English, Japanese, Korean, and German.

[0109] Voice-based shooting input mode is a mode that enables shooting input using voice commands. Users can choose this mode. In voice-based shooting input mode, when a user points the camera of an electronic device at an image, they can use voice commands to specify the text to be captured, allowing the electronic device to extract the target text from the captured image.

[0110] The voice-based Traditional Chinese input mode is a mode that enables Traditional Chinese input using voice commands. Users can choose this mode. In this mode, the electronic device can recognize the user's voice input and convert the recognized text into Traditional Chinese characters for display.

[0111] Voice-based assisted writing input mode is a mode that implements assisted writing input function based on voice input. Users can choose the voice-based assisted writing input mode. In this mode, the electronic device can help users write text, such as greetings, birthday wishes, and holiday greetings. More intelligently, the electronic device can also help users write essays, papers, or business reports for reference. This application embodiment does not specifically limit the assisted writing input mode.

[0112] Voice-based memory input mode is a mode for acquiring memory information based on voice input. Users can choose voice-based memory input mode. In this mode, users can describe the memory information to be acquired via voice, such as user ID number, residential address, or contact number. The electronic device acquires the corresponding memory information based on the user's voice input. Here, the memory information can be information stored in the current application or information recorded in other applications, such as in the "Xiaoyi Memory" application; this embodiment does not limit this.

[0113] The voice-based method for acquiring cross-application target information is a mode for acquiring target information across applications using voice input. Users can choose this mode, in which they can acquire target information across applications within the current application. For example, they can obtain weather information for a specific region from a weather application, music, videos, and short videos from a multimedia application, addresses or phone numbers of hotels near tourist attractions from a travel application, and delivery phone numbers from a food delivery application. This application does not limit this approach.

[0114] The voice-guided cross-application target operation mode is a mode that instructs electronic devices to perform cross-application target operations based on voice input. Users can choose this mode, in which they can input voice information into the electronic device from the current application's interface. This voice information instructs the target application and the target operation, such as playing target music or video in a multimedia application, or booking a hotel near a target attraction in a travel application.

[0115] To facilitate understanding, the following examples, with reference to the accompanying diagrams, illustrate various voice interaction scenarios involving different voice input modes.

[0116] Voice interaction scenario A:

[0117] In one voice interaction scenario, users can choose a direct voice input mode to interact with electronic devices.

[0118] Referring to Figure 5, which is a schematic diagram of a possible scenario for a voice interaction method provided in this application embodiment, the interface 501 in Figure 5 shows a schematic diagram of a mobile phone's memo. In interface 501, on the newly added memo interface, the electronic device can provide the user with a selection entry for a voice input method (e.g., voice input method 1), which can be represented by button 1. The user can select the voice input method 1 by selecting button 1. For example, the user can click button 1 in interface 501 to select input method 1. Voice input method 1 can correspond to voice input mode 1, which can be, for example, a voice-based direct input mode. In voice input mode 1, the user can input voice into the electronic device, for example, voice 1. As an example, voice 1 can be the voice input of the following recipe into the electronic device's memo: "Sauté chopped onions, scallions, celery, mushroom slices, capers, black olives, and cherry tomatoes in olive oil. Add butter, sea salt, and black pepper." The electronic device can perform speech recognition on the content of speech 1 and add the corresponding speech recognition result, i.e., text 1 corresponding to speech 1, to the memo on interface 501.

[0119] The electronic device can provide the user with an entry point for another voice input method (e.g., voice input method 2). This entry point can be represented, for example, by button 2 on interface 502, 503, or 504. If the user wishes to edit the content or format of text 1, they can select voice input method 2 by selecting button 2. For example, the user can click button 2 on interface 502 to select input method 2. Voice input method 2 can correspond to voice input mode 2, which can be, for example, a voice-based text editing mode. In voice input mode 2, the user can input voice information into the electronic device, for example, as voice 2. Voice 2 can be used to edit the content or format of text 1, such as text content modification in interface 502, text line break in interface 503, and text deletion in interface 504.

[0120] As shown in interface 502, in the voice-based text editing mode, voice 2 can instruct modifications to the content of text 1. For example, voice 2 is the voice corresponding to the following text: "Change black pepper powder to white peppercorns." The electronic device can perform voice recognition on voice 2 and display the corresponding voice recognition result, i.e., the text 2 corresponding to voice 2, in the echo area next to button 2. At the same time, the electronic device can also obtain the semantic analysis result of voice 2, knowing that "black pepper powder" in text 1 recorded in the memo needs to be changed to "white peppercorns." The electronic device can also perform the corresponding modification operation based on the semantic analysis result corresponding to voice 2. Thus, the text content of the recipe recorded in the phone's memo becomes the following: "Sauté chopped onions, scallions, celery, mushroom slices, capers, black olives, and cherry tomatoes in olive oil. Add butter, sea salt, and white peppercorns," as shown in text 3-1 in interface 502.

[0121] As shown in interface 503, in the voice-based text editing mode, voice 2 can instruct on setting or modifying the format of the content in text 1. For example, voice 2 is the voice corresponding to the following text: "Start a new line from adding cream". The electronic device can perform voice recognition on voice 2 and display the corresponding voice recognition result in the echo area next to button 2, i.e., text 2 corresponding to voice 2. At the same time, the electronic device can also obtain the semantic analysis result of voice 2, knowing that a line break needs to be performed on the text in text 1 recorded in the memo that starts from "add cream". The edited text is shown as text 3-2 in interface 503.

[0122] As shown in interface 504, in voice-based text editing mode, voice 2 can instruct the deletion of content in text 1. For example, voice 2 is the voice corresponding to the following text: "Delete black pepper powder". The electronic device can perform voice recognition on voice 2 and display the corresponding voice recognition result in the echo area next to button 2, i.e., the text 2 corresponding to voice 2. At the same time, the electronic device can also obtain the semantic analysis result of voice 2, knowing that "black pepper powder" needs to be deleted from the text 1 recorded in the memo. The electronic device can also perform the corresponding deletion operation based on the semantic analysis result corresponding to voice 2. Thus, the text content of the recipe recorded in the phone's memo changes to the following content: "Sauté chopped onions, scallions, celery, mushroom slices, capers, black olives, and cherry tomatoes in olive oil. Add butter and sea salt", as shown in text 3-3 in interface 504.

[0123] It should be noted that the specific content and display format of interfaces 501 to 504 are not limited in this embodiment. For example, button 1 and / or button 2 can be located at the bottom of the interface or at the top of the interface. Button 1 and / or button 2 can be displayed separately on the interface or in a function menu on the interface. This function menu is the selection entry point for multiple buttons, which will not be elaborated here. In other embodiments, if the speech recognition of voice 1 is incorrect, resulting in erroneous text in text 1, the user can also use the method shown in interface 502 or interface 504 to input voice 2 into the electronic device in the voice-based text editing mode to edit the erroneous text. Implementation details can be found in the relevant description in conjunction with voice interaction scenario C below, which will not be elaborated here.

[0124] Voice interaction scenario B:

[0125] In a voice interaction scenario, a user can choose a voice-based text query mode to interact with an electronic device. In this mode, the electronic device can perform query operations based on the user's voice input. These operations can include, but are limited to, querying the number of words (or characters) in the text, the location of a specific word or character, or the frequency of a word or character's occurrence. Referring to Figure 6a, which illustrates another possible scenario for a voice interaction method provided in this embodiment, interface 601 in Figure 6a shows a schematic diagram of a mobile phone reading application (APP). In interface 601, the electronic device can display the content of a novel to the user, represented as text 1, illustratively indicated by "XXX", "YYY", "ZZZ", etc. Simultaneously, the electronic device can provide a selection entry for voice input method 2, such as button 2. The user can click button 2 to select voice input method 2, and correspondingly enter voice input mode 2 shown in interface 602. This voice input mode 2 can be, for example, a voice-based text query mode. In voice input mode 2, users can input voice 2 into the electronic device to perform query-like operations on text 1 on the current interface.

[0126] As an example, in the voice-based text query mode, voice 2 could be the voice corresponding to the following text: "Query the number of words (or characters) of the text on this page," used to query the number of words (or characters) of the text currently displayed. Accordingly, the electronic device can perform voice recognition on voice 2 and display the corresponding voice recognition result, i.e., text 2 corresponding to voice 2, in the echo area next to button 2. Simultaneously, the electronic device can also obtain the semantic analysis result of voice 2, knowing that it needs to query / count the number of words (or characters) of the text on the current interface. The electronic device performs the corresponding word count operation based on the semantic analysis result of voice 2 and displays the query result at the bottom of interface 602, for example, represented as text 3, indicating the number of words (or characters) of the text on the current interface: 44.

[0127] Similarly, if a user needs to search for a text (or character) or its location on interface 601, they can click on interface 601 to select voice input method 2. This will enter voice input mode 2 shown on interface 603. Voice input mode 2 could be, for example, a voice-based text query mode. In voice input mode 2, the user can input voice 2 into the electronic device to perform a query operation on text 1 on the current interface. For example, voice 2 could be the voice corresponding to the following text: "Query XX," used to query whether XX exists in text 1, or to query the location of XX in text 1. Accordingly, the electronic device can perform voice recognition on voice 2 and display the corresponding voice recognition result, i.e., the text 2 corresponding to voice 2, in the echo area next to button 2. Simultaneously, the electronic device can also obtain the semantic analysis result of voice 2, knowing that XX needs to be queried on the current interface. The electronic device performs the corresponding query operation based on the semantic analysis result of voice 2 and highlights the queried text XX in text 1 on interface 603.

[0128] Voice interaction scenario C:

[0129] In one voice interaction scenario, users can choose a voice-based text creation mode to interact with electronic devices via voice.

[0130] Figure 6b illustrates another possible scenario for the voice interaction method provided in this application embodiment. Interface 604 in Figure 6b shows a chat interface of a mobile instant messaging app. In interface 604, the user can select voice input method 1 by clicking button 1. Voice input method 1 corresponds to voice input mode 1, which can be a direct voice input mode. In the direct voice input mode, the user can input voice 1 into the electronic device, specifically the voice corresponding to the following text: "Happy Birthday, Little B". The electronic device can recognize the user's voice 1 and insert the recognized text 1 into the input box of the chat interface.

[0131] The electronic device can also input text via button 2 shown on interface 605. Button 2 is associated with voice input method 2 and corresponds to voice input mode 2, which can be a voice-based text creation mode. If the user feels the text in interface 604 is too simplistic and insufficient to express user A's blessings to user B, the user can click button 2 on interface 605 to select voice input method 2, which corresponds to voice input mode 2, also a voice-based text creation mode. In voice-based text creation mode, the user can input voice 2 into the electronic device, specifically the voice corresponding to the following text: "Please rewrite this for me." The electronic device can recognize the user's voice 2 input and display the corresponding text 2 in the echo area. The electronic device can use a large-scale language model to perform semantic analysis on the speech 2. Knowing that the user has a request to rewrite the text 1 in interface 604, the electronic device can use the large-scale language model to perform a polishing and rewriting operation on text 1 to obtain text 3, as shown below: "Dear Little B, happy birthday! May your smile be as bright as the sunshine and your life as beautiful as a poem. Every year of growth is a new beginning. May you be healthy, happy, safe, and successful in the future!" The input box of interface 605 only displays part of the content of text 3.

[0132] Similarly, users can also provide text requests directly in the voice-based text creation mode on interface 605 without providing any basic text. The electronic device can then create text 3 for the user based on the user's voice input using tools such as large-scale language models, which will not be elaborated further here.

[0133] Voice interaction scenario D:

[0134] Referring to Figure 7, which illustrates another possible scenario for the voice interaction method provided in this application embodiment, interface 701 in Figure 7 shows the interface of a mobile phone's phone app. In interface 701, the electronic device can provide the user with historical communication records, a dial pad, and button 1. Button 1 is associated with voice input method 1 and voice input mode 1. The user can click button 1 to select voice input method 1 and enter interface 702. The voice input mode 1 corresponding to voice input method 1 is, for example, a direct voice input mode. In interface 702, under voice input mode 1, the user can input the following text as voice 1: "Call XXX," so that the mobile phone's phone app can make a call to the target user "XXX." The electronic device can perform voice recognition on the user's input voice 1 and display the text 1 corresponding to voice 1 in a display area on interface 702.

[0135] Taking the Chinese language scenario as an example, taking the voice pronunciation of the target user "XXX" as the following pinyin "wang jian yu" as an example, due to the existence of homophonic words, when the electronic device performs speech recognition on Speech 1, it mistakenly recognizes the voice pronunciation "wang jian yu" of the target user as "Wang Jianyu" shown in Interface 702. Therefore, it is necessary to modify the incorrect text.

[0136] In Interface 703, the electronic device can also provide Button 2 to the user. Button 2 is associated with Speech Input Method 2 and Speech Input Mode 2. The user can click Button 2 in Interface 703 to switch to Speech Input Method 2. The speech input mode 2 corresponding to Speech Input Method 2 can be, for example, a text editing mode based on speech. In Speech Input Mode 2, the user can input Speech 2 to the electronic device to input supplementary explanations or descriptions of Speech 1 to the electronic device. Among them, for the homophonic characters to be replaced, Speech 2 can describe the homophonic characters in at least one of the following ways: common character combination description method; character decomposition description method; semantic description method. For example, in Interface 703, Speech 2 can include the speech corresponding to the following text: "Jian is the Jian with a single-person radical, and yu is the Yu of Henan." Among them, "Jian is the Jian with a single-person radical" describes the expected homophonic character "jian" in the character decomposition description method, and "yu is the Yu of Henan" describes the expected homophonic character "yu" in the semantic description method. Optionally, using the common character combination description method, Speech 2 can include the speech corresponding to the following text: "Jian is the Jian of healthy, and yu is the Yu of毫不犹豫 (without hesitation)." The electronic device can perform speech recognition on the Speech 2 input by the user and display the text 2 corresponding to Speech 2 in the echo area next to Button 2, as shown in Interface 703. At the same time, the electronic device can obtain the semantic analysis result of Speech 2 and perform corresponding operations according to the semantic analysis result. For example, modify the incorrect text "建 (Jian)" in Interface 702 to the homophonic word "健 (Jian)", and modify the incorrect text "玉 (Yu)" to the homophonic word "豫 (Yu)", as shown in Interface 704.

[0137] It should be noted that Figure 7 is only an example introduction taking the Chinese language scenario as an example, and does not constitute a limitation on the language scenario. In other language scenarios, such as English, German, French, etc., a similar method to Figure 7 can be adopted to input voice information to the electronic device in different speech input modes. The electronic device can perform corresponding operations according to the recognition result and / or semantic analysis result of the voice information, which will not be elaborated here.

[0138] The aforementioned voice interaction scenario C can also be implemented as a cross-application voice interaction scenario. For example, the interface of the phone app in Figure 7 can be replaced with the phone's desktop or the interface of other applications. Users can perform cross-application operations on the phone's desktop or other application interfaces in a manner similar to Figure 7 to make calls to target users. Similarly, during the voice interaction between the user and the electronic device, different voice information can be input to the electronic device by selecting different voice input methods, as shown in Figure 7, to achieve the purpose of making calls to target users, thus improving the accuracy and convenience of voice input.

[0139] Voice interaction scenario E:

[0140] Referring to Figure 8, which illustrates another possible scenario for the voice interaction method provided in this application embodiment, the interface 801 in Figure 8 shows a chat interface of a mobile instant messaging app. In interface 801, when user A chats with user B, the electronic device can provide the user with a selection entry for voice input method 1, for example, by associating button 1 with voice input 1 and voice input mode 1. The user can click button 1 to select voice input method 1. The voice input mode 1 corresponding to voice input method 1 can be, for example, a direct voice input mode. In voice input mode 1, the user can input voice 1 into the electronic device. Correspondingly, the electronic device can recognize voice 1 and input the text 1 obtained from recognizing voice 1 into the input box of the chat interface.

[0141] Taking an English language scenario as an example, when a user expects to express the following text "Please buy me some flours" in Voice 1, due to homophones or near-homophones, the user's intended "flour" is incorrectly identified as "flower," resulting in Text 1 displayed in the input box as "Please buy me some flowers." Therefore, the incorrect text needs to be corrected.

[0142] In interface 802, the electronic device can also provide the user with a button 2, which is associated with a voice input method 2 and a voice input mode 2. Voice input mode 2 can be, for example, a voice-based text editing mode. The user can click button 2 in interface 802 to switch to voice input method 2. In voice input mode 2, the user can input voice 2 into the electronic device to provide supplementary explanations or descriptions of voice 1. For homophones to be replaced, voice 2 can describe the homophone in at least one of the following ways: common character combination explanation; character decomposition explanation; semantic explanation. For example, in interface 803, voice 2 can include the voice corresponding to the following text: “flower change to flour which could be used to make bread,” meaning to change “flower” in text 1 to “flour,” the word that can be used to make bread. Here, “the word that can be used to make bread” is a semantic explanation of “flour.”

[0143] The electronic device can perform speech recognition on the user's input voice 2 and display the corresponding text 2 in the echo area next to button 2, as shown in interface 803. Simultaneously, the electronic device can obtain the semantic analysis results of voice 2 and perform corresponding operations based on these results. For example, it can change the erroneous text "flower" in interfaces 802 and 803 to "flour," as shown in interface 804.

[0144] It should be noted that Figure 8 is only an example using an English language scenario and chat interface, and does not constitute a limitation on the language scenario. In other language scenarios, such as Chinese, German, and French, a similar approach can be used to input voice information into the electronic device under different voice input modes. The electronic device can perform corresponding operations based on the recognition results and / or semantic analysis results of the voice information, which will not be elaborated further here.

[0145] Voice interaction scenario F:

[0146] In a voice interaction scenario, an electronic device can provide users with multiple voice input method selection options on its display screen. These multiple voice input methods correspond one-to-one with multiple voice input modes. If a user selects a particular voice input method, the electronic device can perform the corresponding operation based on the user's voice input in the corresponding voice input mode. These operations can be cross-application operations. For example, if the current interface on the display screen is the interface of application 1, with user authorization, a cross-application operation could be: the user obtains information from application 2 through voice input on the current interface. Application 2 can be from the same developer as the electronic device, or it can be from a different developer; this embodiment does not limit this.

[0147] Referring to Figure 9, which is a schematic diagram of a possible scenario for the voice interaction method provided in this application embodiment, the interface 901 in Figure 9 shows a chat diagram of an instant messaging APP on a mobile phone. In interface 901, the electronic device can provide the user with a selection entry for voice input method 2, represented by button 2. At this time, button 2 is in an inactive state, drawn with thinner lines to indicate that the voice input method 2 associated with button 2 is not active. The user can click on button 2 in interface 901 to switch button 2 to an active state, as shown in interface 902, where button 2 in the active state is drawn with thicker lines.

[0148] In interface 902, under voice input mode 2 corresponding to voice input method 2, such as the mode for obtaining cross-application target information based on voice, the user can input voice 2 into the electronic device. Voice 2 is used to express the user's expectation to obtain target information from application 2. As an example, as shown in interface 902, voice 2 can include the voice corresponding to the following text: "(In the input box) Insert my ID number". The electronic device can recognize voice 2 and display the voice recognition result corresponding to voice 2, i.e., text 2 corresponding to voice 2, in the echo area next to button 2. At the same time, the electronic device can also obtain the semantic analysis result of voice 2, knowing that user A's ID number needs to be inserted into the input box of the chat interface. The electronic device can perform cross-application target information retrieval operation. For example, if application 2 is "Xiaoyi Memory" and the target information includes the ID number of the device owner (user A), the electronic device can obtain user A's ID number from "Xiaoyi Memory" based on the semantic analysis result of voice 2, represented as 1234567890. The electronic device can insert the obtained target information into the input box of the chat interface, as shown in text 3 in interface 902.

[0149] Similarly, when user A wants to insert other target information obtained across applications into the input box of the chat interface, such as personal information like user A's contact address, weather information for a certain country or region, multimedia resources such as music, audio and video, or information related to user A's daily life such as food, clothing, housing, and transportation, such as hotel addresses and phone numbers, and food delivery phone numbers, user A can select the corresponding voice input method as shown in Figure 9 to interact with the electronic device by voice in the corresponding voice input mode, thereby freeing the user's hands and improving the accuracy and convenience of voice input.

[0150] Optionally, when user A expects (or receives information from user B) to perform a cross-application target operation, such as booking a hotel or ordering food delivery, user A can, as shown in Figure 9, select the appropriate voice input method and interact with the electronic device via voice in the corresponding voice input mode, such as the voice instruction-based cross-application target operation mode. User A can input voice instructions to the electronic device to invoke the target application (e.g., a travel app, a food delivery app, etc.) to perform the target operation, or to book a hotel or order food delivery by making a phone call. This frees up the user's hands and improves the accuracy and convenience of voice input.

[0151] Voice interaction scenario G:

[0152] Referring to Figure 10, this is a schematic diagram of a possible scenario for a voice interaction method provided in this application embodiment. Interface 1001 in Figure 10 shows a chat diagram of an instant messaging app on a mobile phone. In interface 1001, the electronic device can provide the user with a selection entry for voice input method 2, represented by button 2. At this time, button 2 is in an inactive state, drawn with thinner lines to indicate that the voice input method 2 associated with button 2 is not active. Interface 1001 can also display a keyboard. The user can click on button 2 in interface 1001 to switch button 2 to an active state, as shown in interface 1002, where button 2 in the active state is drawn with thicker lines.

[0153] In interface 1002, under voice input mode 2 corresponding to voice input method 2, such as voice-based writing assistance input mode, the user can input voice 2 into the electronic device. Voice 2 is used to express the user's desire to switch input methods. As an example, as shown in interface 1002, voice 2 can include the voice corresponding to the following text: "Switch to Xiaoyi Writing Assistance". The electronic device can recognize voice 2 and display the voice recognition result corresponding to voice 2, i.e., the text 2 corresponding to voice 2, in the echo area next to key 2. At the same time, the electronic device can also obtain the semantic analysis result of voice 2, knowing that the input method needs to be switched to "Xiaoyi Writing Assistance" in the chat interface. The electronic device can perform the operation of switching input methods, thereby switching the keyboard displayed on interface 1002 to the Xiaoyi Writing Assistance mode shown in interface 1003, so that in the Xiaoyi Writing Assistance mode, the electronic device can send greetings, blessings, etc. to user B on behalf of user A.

[0154] Similarly, if user A wishes to switch the voice input mode to at least one of the following: voice-based translation input mode, voice-based camera input mode, voice-based traditional Chinese input mode, and voice-based memory input mode, user A can interact with the electronic device via voice using a similar method as shown in Figure 10 to switch the voice input mode to at least one of these modes. Implementation details can be found in the relevant descriptions above in conjunction with Figure 10, and will not be repeated here.

[0155] It should be noted that although some voice input modes in Figures 5, 6a-6b, and 7-10 use the same button icon for illustration, in actual applications, one voice input mode can correspond to one button, or one button can be associated with several voice input modes. This application does not limit this.

[0156] It should be noted that Figures 5, 6a-6b, and 7-10 are merely illustrative examples of voice interaction scenarios in embodiments of this application and are not intended to limit the scope of the application. In other embodiments, the selection entry points for different voice input methods provided by the electronic device may have other implementations, and the operations performed by the electronic device are not limited to those described in Figures 5-10 depending on the user's selection of different voice input methods.

[0157] Taking the selection entry point for different voice input methods provided by electronic devices as an example, which can have other implementation methods, the following uses a mobile phone as an example to introduce several possible implementation methods of the voice input method selection entry point in the embodiments of this application.

[0158] In this embodiment of the application, the electronic device can provide different buttons to associate with different voice input methods, one button is associated with one voice input method, and one voice input method corresponds to one voice input mode.

[0159] In this embodiment of the application, the key associated with any voice input method provided by the electronic device can be a physical key (such as a physical button) or a virtual key.

[0160] In some optional implementations, different presentation formats of the same key can be used to represent different keys, with one voice input method associated with one presentation format and corresponding to one voice input mode. This application does not limit the implementation method of different keys associated with voice input functions. In optional implementations, different states can be set for the same key to indicate whether the function or mode associated with the key is activated. For example, the key state can include an active state and an inactive state. The active state can also be called the function-activated state, indicating that the function / mode associated with the key is activated or started. The inactive state can also be called the non-activated state, indicating that the function / mode associated with the key is not activated or started.

[0161] In one possible implementation, taking the example of implementing virtual buttons for multiple voice input methods, the user's selection operation for the multiple voice input method selection entry can include a click operation and / or a long press operation. Alternatively, taking the example of implementing physical buttons for multiple voice input methods, the user's selection operation for the multiple voice input method selection entry can include a press operation and / or a long press operation.

[0162] The following examples all use floating buttons as an example.

[0163] Example 1:

[0164] As shown in Figure 11, initially, in interface 1101, the electronic device can provide the user with a selection entry for a voice input method (e.g., voice input method 1), represented by button 1, which is currently inactive. The user can perform one or more selection operations by following the direction of the solid arrow; for example, the user can click button 1 in interface 1101 to select voice input method 1. Correspondingly, the electronic device can detect the user's operation and control button 1 to switch from an inactive state to an active state based on the user's operation, as shown in interface 1102. In interface 1102, under voice mode 1 corresponding to voice input method 1, the user can input voice 1 into the electronic device. The user can long-press button 1 in interface 1101 or long-press button 1 in interface 1102 to enter interface 1103. A function menu can be displayed in interface 1103. This function menu can include selection entries for various input methods, such as button 1, button 2, keyboard keys, and delete key. Among them, button 1 is associated with voice input method 1 as described in Figures 5, 6a-6b, and 7-10, corresponding to voice input mode 1. Button 2 is associated with voice input method 2, as described in Figures 5, 6a-6b, and 7-10, corresponding to voice input mode 2. Users can click button 2 on interface 1103 to enter interface 1104. At this time, button 2 is active, and the associated voice input method 2 is activated, allowing users to input voice information into the electronic device in voice input mode 2. Users can also long-press button 2 on interface 1104 to enter the function menu on interface 1103, allowing them to select other input methods from the menu.

[0165] Example 2:

[0166] In one alternative implementation, initially, the electronic device can provide a function menu to the user on interface 1103, including multiple input method selection options, such as key 1, key 2, keyboard keys, and a delete key. The user can perform one or more selection operations by following the direction of the dotted arrow. For example, the user can click key 1 on interface 1103 to select voice input method 1 and enter interface 1103. Alternatively, the user can click key 2 on interface 1103 to select voice input method 2 and enter interface 1104. The user can also click key 1 on interface 1102 to cancel voice input method 1 and enter interface 1101, where key 1 is in an inactive state.

[0167] It should be understood that in Figure 11, lines of different thicknesses represent different states of the keys. For example, a thin line indicates that the input method associated with the corresponding key is inactive, while a thick line indicates that the input method associated with the corresponding key is active. In other embodiments, different key states can also be distinguished by whether the key is highlighted. For example, a highlighted state indicates an active state, and a non-highlighted state indicates an inactive state. This application does not limit the presentation of key states.

[0168] Example 3:

[0169] In another possible implementation, as shown in Figure 12, initially, the electronic device can provide a function menu to the user in interface 1201, including selection entries for various voice input methods, such as button 1 and button 2. Button 1 is associated with voice input method 1 as described in Figures 5, 6a-6b, and 7-10, corresponding to voice input mode 1. Button 2 is associated with voice input method 1 as described in Figures 5, 6a-6b, and 7-10, corresponding to voice input mode 2. In interface 1201, both button 1 and button 2 are in an inactive state. For easy distinction, button 1 and button 2 in the inactive state are drawn with thinner lines.

[0170] In interface 1201, the user can select the voice input method 2 associated with button 2 by clicking button 2, entering interface 1202. At this time, button 2 is in an active state, and the user can input voice into the electronic device in voice input mode 2. In interface 1201, the user can select the voice input method 1 associated with button 1 by clicking button 1, entering interface 1203. At this time, button 1 is in an active state, and the user can input voice into the electronic device in voice input mode 1.

[0171] Optionally, users can return to the menu shown in interface 1201 by long-pressing button 2 in interface 1202 or long-pressing button 1 in interface 1203. At this time, all buttons in the menu are inactive.

[0172] It is understood that the menu in Figure 11 or 12 is merely an example and not a limitation. In other embodiments, the menu may be presented in other forms, and the buttons in the menu may include buttons other than those associated with various voice input methods.

[0173] Example 4:

[0174] Referring to Figure 13, in interface 1301, the electronic device can provide a function menu to the user in a circular manner. This menu includes selection entries for various input methods, such as key 1, key 2, keyboard, and delete key. Key 1 is associated with voice input method 1, as described in Figures 5, 6a-6b, and 7-10, corresponding to voice input mode 1. Key 2 is associated with voice input method 2, as described in Figures 5, 6a-6b, and 7-10, corresponding to voice input mode 2. In interface 1301, all items in the menu are inactive. For easy differentiation, the inactive keys are drawn with thinner lines.

[0175] In interface 1301, the user can click on button 2 to select the voice input method 2 associated with button 2, entering interface 1302. At this time, button 2 is in an active state, and in voice input mode 2, the user can input voice into the electronic device. In interface 1301, the user can click on button 1 to select the voice input method 1 associated with button 1, entering interface 1303. At this time, button 1 is in an active state, and in voice input mode 1, the user can input voice into the electronic device.

[0176] Optionally, users can return to the menu shown in interface 1301 by long-pressing button 2 in interface 1302 or long-pressing button 1 in interface 1303. At this time, all buttons in the menu are inactive.

[0177] Similarly, in Figures 12 and 13, lines of different thicknesses represent the different states of the buttons. For example, thin lines indicate that the corresponding button is inactive, while thick lines indicate that the corresponding button is active. In other embodiments, different button states can also be distinguished by whether the button is highlighted. This application does not limit the presentation of button states.

[0178] It should be understood that Figures 11-13 are merely illustrative representations of buttons associated with multiple voice input methods that can be displayed on the screen of an electronic device for user selection, and do not limit the specific type or content of the current interface on the screen. In specific implementations, the buttons shown in Figures 11-13 can float on the mobile phone desktop, or they can float on interfaces corresponding to various possible business scenarios, such as the memo interface shown in Figure 5, or the display interface of a reading app shown in Figure 6a or 6b, or the operation interface of a phone app shown in Figure 7, or the chat interface of an instant messaging app shown in Figures 8-10, etc. This application embodiment does not limit this.

[0179] Based on the different implementation methods described above, the application scenarios of the voice interaction method of this application and the implementation methods of selecting multiple voice input methods provided by electronic devices have been introduced. The implementation process of the voice interaction method of this application will be described below in conjunction with the voice interaction scenarios shown in Figures 5, 6a-6b, 7-10, and the different implementation methods of the selection entry shown in Figures 11-13, so as to understand the interface display effect that can be achieved by using the voice interaction method of this application.

[0180] Scene 1:

[0181] Referring to Figure 14, the voice interaction method may include the following steps:

[0182] S1401: Electronic devices provide multiple voice input method selection options on the display screen.

[0183] In this embodiment, the various voice input methods correspond one-to-one with the various voice input modes provided by the electronic device.

[0184] For example, as described above with reference to Figures 5, 6a-6b, and 7-13, button 1 is the selection entry for voice input method 1, which corresponds to voice input mode 1. Button 2 is the selection entry for voice input method 2, which corresponds to voice input mode 2. Multiple voice input modes can include at least one of the following: direct voice input mode; voice-based text editing mode; voice-based text query mode; voice-based text creation mode; voice-based translation input mode; voice-based photo input mode; voice-based traditional Chinese character input mode; voice-based assisted writing input mode; voice-based memory input mode; voice-based mode for obtaining cross-application target information; and voice-based mode for instructing cross-application target operations. Refer to the various voice interaction scenarios described above with reference to Figures 5, 6a-6b, and 7-10; they will not be repeated here.

[0185] S1402: User A selects from multiple voice input methods. Accordingly, the electronic device detects the user's selection of multiple voice input methods.

[0186] In this embodiment, multiple voice input methods can be associated one-to-one with multiple buttons, which can be physical buttons or virtual buttons. As an example, these multiple buttons can be different virtual buttons described above in conjunction with Figures 5, 6a-6b, and 7-13. The user's selection operation for multiple voice input methods can include click operations and / or long-press operations. Alternatively, these multiple buttons can be different physical buttons, and the user's selection operation for multiple voice input methods can include press operations and / or long-press operations.

[0187] In one example, if multiple voice input methods are directly provided on the display screen in S1401, such as multiple buttons, the user can perform a selection operation by clicking a button to select the first voice input method from multiple voice input methods, as shown in interface 1103 of Figure 11, or interface 1201 of Figure 12, or interface 1301 of Figure 13.

[0188] In another example, if in S1401, the display screen does not directly provide a selection entry for multiple voice input methods, the user can perform a second operation on a first button provided on the display screen. The first button can be any one of multiple buttons. After the electronic device detects the user's second operation on the first button, it can provide a selection entry for multiple voice input methods on the display screen according to the second operation. As an example, the second operation can include a long press operation on button 1 in interface 1101 or a long press operation on button 2 in interface 1102, as shown in Figure 11. This operation can bring up the function menu in interface 1103, i.e., the selection entry for multiple voice input methods.

[0189] It should be understood that in S1402, the selection operation performed by the user on the selection entry for multiple voice input methods can also be other types of operations. For example, if the selection entry is implemented as a physical button, the selection operation can be a pressing operation or a long press operation on the corresponding physical button, which will not be elaborated here.

[0190] S1403: The electronic device determines a first input method from multiple voice input methods based on a selection operation. The first voice input method corresponds to a first voice input mode among multiple voice input modes. Accordingly, user A interacts with the electronic device via voice in the first voice input mode.

[0191] For example, in Figure 11, in interface 1103, the electronic device detects the user's click operation on button 1, selects voice input method 1 as the first input method, enters interface 1102, and interacts with the user via voice in voice input mode 1. This voice input mode 1 is the first voice input mode. Alternatively, in interface 1103, the electronic device detects the user's click operation on button 2, selects voice input method 2 as the first input method, enters interface 1104, and interacts with the user via voice in voice input mode 2. This voice input mode 2 is the first voice input mode.

[0192] Alternatively, as shown in Figure 12, in interface 1201, the electronic device detects the user's click operation on button 2, selects voice input 2 as the first voice input method, enters interface 1202, and interacts with the user via voice in voice input mode 2, which is the first voice input mode. Alternatively, in interface 1201, the electronic device detects the user's click operation on button 1, selects voice input method 1 as the first input method, enters interface 1203, and interacts with the user via voice in voice input mode 1, which is the first voice input mode.

[0193] Alternatively, as shown in Figure 13, in interface 1301, the electronic device detects the user's click operation on button 2, selects voice input 2 as the first voice input method, enters interface 1302, and interacts with the user via voice in voice input mode 2, which is the first voice input mode. Alternatively, in interface 1301, the electronic device detects the user's click operation on button 1, selects voice input method 1 as the first input method, enters interface 1303, and interacts with the user via voice in voice input mode 1, which is the first voice input mode.

[0194] S1404: In the first voice input mode, the electronic device performs a first operation based on the first voice information input by the user.

[0195] In the embodiments of this application, the first voice information and the first operation may be different in different voice interaction scenarios.

[0196] For example, if the first voice input mode is a voice-based text editing mode and it involves replacing homophones or near-homophones, the first voice information can describe the first character in at least one of the following ways: a common character combination description method, a character decomposition description method, or a semantic description method. In S1404, the electronic device can replace the second character displayed on the display screen with the first character, wherein the first character and the second character are homophones or near-homophones. Refer to the content described above in conjunction with Figure 7 or Figure 8.

[0197] Alternatively, for example, if the first voice input mode is a voice-based text editing mode, and it involves the replacement of non-homophones or near-homophones, the first voice information can indicate the old text content to be modified, the new text content to be modified to, and a keyword associated with an editing operation. For example, "change to," "replace with," "modify to," "modify into," "replace with," etc., can all be keywords associated with the modification operation. In S1404, the electronic device can modify the old text content to the new text content. Refer to the content described above in conjunction with Figure 5.

[0198] Alternatively, for example, if the first voice input mode is a voice-based text editing mode, the first voice information can be used to describe a new first text format. In S1404, the electronic device can modify the second text format to the first text format based on the first voice information. Referring to the content described above in conjunction with Figure 5, text 1 uses the second text format, text 3 uses the first text format, and the editing operation is a text line break operation.

[0199] For example, if the first voice input mode is a voice-based text query mode, the first voice information can be used to describe the query instruction. In S1404, the electronic device can perform a query operation in the first text information based on the first voice information. Referring to the above description in conjunction with Figure 6a, one can query the number of words in the text shown in interface 602, or query XX in the text shown in interface 603.

[0200] Alternatively, for example, if the first voice input mode is a voice-based text creation mode, the first voice information can be used to instruct switching to the input method associated with the voice-based text creation mode, as shown in Figure 6b. Alternatively, the first voice information can be used to describe the requirements to be met in the creation process, such as birthday wishes, holiday greetings, etc. In S1404, the electronic device can perform a text creation operation based on the first voice information.

[0201] Alternatively, for example, if the first voice input mode is a mode for obtaining cross-application target information based on voice, the first voice information can be used to describe the target information. As shown in Figure 9, "My ID number" is a description of user A's personal information. In S1404, the electronic device can receive the first voice information input by the user on the first interface provided by the electronic device in the first voice input mode. Then, the electronic device can obtain the target information from the second application based on the first voice information input by the user and display the target information on the first interface. As shown in Figure 9, user A's ID number is entered and displayed in the input box of the chat interface.

[0202] Scene 2:

[0203] In an optional implementation, the voice interaction method of this application embodiment can be implemented collaboratively by an electronic device and a server. For example, in the architecture shown in A of Figure 4, the interface display effect achieved by the voice interaction scenario shown in Figures 5, 6a-6b, and 7-10 is obtained through the interaction between the electronic device and the server.

[0204] Referring to Figure 15, the voice interaction method may include the following steps:

[0205] S1501: In voice input mode 1 (an example of the first voice input mode), the electronic device receives voice information 1 (an example of the first voice information) input by the user. Details of this step can be found in the preceding descriptions in conjunction with Figures 5, 6a-6b, and 7-14, and will not be repeated here.

[0206] S1502: The electronic device sends a request message 1 (an example of a first request message) to the server 1, the request message 1 including voice information 1. Accordingly, the server 1 receives the request message 1 and obtains the voice information 1 from the request message 1.

[0207] Server 1 can be a server that deploys a large-scale language model (referred to as a large model), and any application on an electronic device can be authorized to communicate with server 1.

[0208] S1503: Server 1 performs semantic analysis on voice information 1 and obtains the semantic analysis results.

[0209] S1504: Server 1 sends response information 1 (an example of a first response information) to the electronic device, which includes the semantic analysis results. Accordingly, the electronic device receives response information 1.

[0210] In an optional implementation, in S1502, the electronic device can send a request message 1 to server 1 via server 2. Server 1 is an application server, which can parse and repackage the request message 1 from the electronic device before sending it to server 1. Correspondingly, in S1504, server 2 can parse and repackage the response information 1 from server 1 before sending it to the electronic device.

[0211] S1505: The electronic device performs the first operation based on the semantic analysis results. See the relevant descriptions in conjunction with Figures 5, 6a-6b, and 7-13 above; they will not be repeated here.

[0212] Scene 3:

[0213] In an optional implementation, the voice interaction method of this application embodiment can be implemented collaboratively by an electronic device and a server. For example, in the architecture shown in Figure 4A, the interface display effect achieved by the voice interaction scenario shown in Figures 5, 6a-6b, and 7-10 is obtained through the interaction between the electronic device and the server.

[0214] Referring to Figure 16, the voice interaction method may include the following steps:

[0215] S1601: In voice input mode 1 (an example of the first voice input mode), the electronic device receives voice information 1 (an example of the first voice information) input by the user. Details of this step can be found in the preceding descriptions in conjunction with Figures 5, 6a-6b, and 7-15, and will not be repeated here.

[0216] S1602: The electronic device sends a request message 1 (an example of a second request message) to the server 3, which includes voice information 1. Accordingly, the server 3 receives the request message 1 and obtains the voice information 1 from it.

[0217] Server 3 can be a server with an ASR module deployed, and any application on the electronic device can be authorized to communicate with server 3.

[0218] S1603: Server 3 performs speech recognition on speech information 1 to obtain text information 1 (an example of the first text information).

[0219] S1604: Server 3 sends response information 1 (an example of a second response information) to the electronic device, which includes text information 1. Accordingly, the electronic device receives response information 1 from server 3.

[0220] S1605: The electronic device displays text information 1 on the display screen. See the relevant descriptions in conjunction with Figures 5, 6a-6b, and 7-13 above, which will not be repeated here.

[0221] Scene 4:

[0222] In this scenario, electronic devices can interact with different servers to perform speech recognition and / or semantic analysis of voice information, thereby achieving the interface display effect shown in Figures 5, 6a-6b, and 7-10 of the voice interaction scenario.

[0223] As shown in Figure 17, the method includes the following steps:

[0224] S1701: The electronic device receives voice information 1 input by the user in voice input mode 1 (an example of the first voice input mode). The implementation details of this step can be found in the relevant descriptions in conjunction with Figures 5, 6a-6b, and 7-15 above, and will not be repeated here.

[0225] S1702: The electronic device sends a request message 1 (an example of a second request message) to the server 3, which includes voice information 1 (an example of first voice information). Accordingly, the server 3 receives the request message 1 and obtains the voice information 1 from the request message 1.

[0226] Server 3 can be a server with an ASR module deployed, and any application on the electronic device can be authorized to communicate with server 3.

[0227] S1703: Server 3 performs speech recognition on speech information 1 to obtain text information 1 (an example of the first text information).

[0228] S1704: Server 3 sends response information 1 (an example of a second response information) to the electronic device, which includes text information 1. Accordingly, the electronic device receives response information 1 from server 3.

[0229] S1705: The electronic device displays text information 1 on the display screen. See the relevant descriptions in conjunction with Figures 5, 6a-6b, and 7-13 above, which will not be repeated here.

[0230] S1706: The electronic device sends a request message 2 (an example of a first request message) to server 1, which includes voice information 1. Accordingly, server 1 receives request message 2 and obtains voice information 1 from it.

[0231] Server 1 can be a server that deploys a large-scale language model (referred to as a large model), and any application on an electronic device can be authorized to communicate with server 1.

[0232] S1707: Server 1 performs semantic analysis on voice information 1 and obtains the semantic analysis results.

[0233] S1708: Server 1 sends response information 2 (an example of a first response information) to the electronic device, which includes the semantic analysis results. Accordingly, the electronic device receives response information 2.

[0234] In an optional implementation, in S1706, the electronic device can send a request message 2 to server 1 via server 2. Server 2 is an application server that can parse and repackage the request message 2 from the electronic device before sending it to server 1. Correspondingly, in S1708, server 2 can parse and repackage the response information 2 from server 1 before sending it to the electronic device.

[0235] S1709: The electronic device performs the first operation based on the semantic analysis results. See the relevant descriptions in conjunction with Figures 5, 6a-6b, and 7-13 above; they will not be repeated here.

[0236] It should be understood that in the embodiments of this application, the execution order of S1702-S1704 and S1706-S1708 is not limited. The electronic device can execute S1702-S1704 and S1706-S1708 in parallel, or it can execute S1702-S1704 first and then execute S1706-S1708.

[0237] Scene 5:

[0238] In this scenario, the electronic device can edit incorrectly recognized text content via voice input to implement the voice interaction method of this application embodiment, as shown in Figure 5, 7, or 8. During different voice interaction stages, the user can input voice information into the electronic device using different voice input modes. These voice input modes are all variations of the first voice input mode, and the voice information input by the user in the corresponding voice input mode is an example of the first voice information. In the interaction between the electronic device and different servers, request message 1 and request message 2 are different examples of the second request message, and request message 3 is an example of the first request message. Correspondingly, response information 1 and response information 2 are different examples of the second response information, and response information 3 is an example of the first response information. Text information 1 and text information 2 are different examples of the first text information.

[0239] As shown in Figure 18, the voice interaction method may include the following steps:

[0240] S1801: In voice input mode 1, the electronic device receives voice information 1 input by the user. The implementation process of S1801 can be found in Figure 14 in conjunction with the introduction of S1401-S1404, and will not be repeated here.

[0241] S1802: The electronic device sends a request message 1 to the server 3, which includes voice information 1. Accordingly, the server 3 receives the request message 1 and obtains the voice information 1 from the request message 1.

[0242] Server 3 can be a server with an ASR module deployed, and any application on the electronic device can be authorized to communicate with server 3.

[0243] S1803: Server 3 performs speech recognition on voice information 1 to obtain text information 1.

[0244] S1804: Server 3 sends response information 1 to the electronic device, which includes text information 1. Accordingly, the electronic device receives response information 1 from server 3.

[0245] S1805: The electronic device displays text information 1 on the display screen.

[0246] S1806: In voice input mode 2, the electronic device receives voice information 2 input by the user. Details of this step can be found in the preceding descriptions in conjunction with Figures 5, 6a-6b, and 7-15, and will not be repeated here.

[0247] S1807: The electronic device sends a request message 2 to the server 3, which includes voice information 2. Accordingly, the server 3 receives the request message 2 and obtains the voice information 2 from it.

[0248] S1808: Server 3 performs speech recognition on voice information 2 to obtain text information 2.

[0249] S1809: Server 3 sends response information 2 to the electronic device, which includes text information 2. Accordingly, the electronic device receives response information 2 from server 3.

[0250] S1810: The electronic device displays text information 2 on the display screen.

[0251] S1811: The electronic device sends a request message 3 to the server 1, which includes voice information 2. Accordingly, the server 1 receives the request message 3 and obtains the voice information 2 from the request message 3.

[0252] Server 1 can be a server that deploys a large-scale language model (referred to as a large model), and any application on an electronic device can be authorized to communicate with server 1.

[0253] S1812: Server 1 performs semantic analysis on voice information 2 and obtains the semantic analysis results.

[0254] S1813: Server 1 sends response information 3 to the electronic device, which includes the semantic analysis results. Accordingly, the electronic device receives response information 3.

[0255] In an optional implementation, in S1811, the electronic device can send a request message 3 to server 1 via server 2. Server 2 is an application server that can parse and repackage the request message 3 from the electronic device before sending it to server 1. Correspondingly, in S1813, server 2 can parse and repackage the response information 3 from server 1 before sending it to the electronic device.

[0256] S1814: The electronic device performs the first operation based on the semantic analysis results. See the relevant descriptions in conjunction with Figures 5, 6a-6b, and 7-13 above, which will not be repeated here.

[0257] It should be understood that in the embodiments of this application, the execution order of S1802-S1804, S1807-S1809 and S1811-S1813 is not limited. The electronic device can execute S1802-S1804, S1807-S1809 and S1811-S1813 in parallel, or it can execute S1802-S1804 first, then S1807-S1809, and then S1811-S1813.

[0258] Based on the above embodiments, this application also provides an electronic device, which includes multiple functional modules. These multiple functional modules interact to implement the functions performed by the electronic device in the methods described in the embodiments of this application. The multiple functional modules can be implemented based on software, hardware, or a combination of both, and can be arbitrarily combined or divided based on specific implementations. For example, executing steps S1401, S1403-S1404 in the embodiment shown in FIG. 14; steps S1502-1505 in the embodiment shown in FIG. 15; and steps S1602-S1605 in the embodiment shown in FIG. 16. Similar steps are shown in FIG. 17-18.

[0259] Based on the above embodiments, this application also provides an electronic device, which includes at least one processor and at least one memory. The at least one memory stores computer program instructions. When the electronic device is running, the at least one processor executes the functions performed by the electronic device in the various methods described in the embodiments of this application. For example, executing steps S1401, S1403-S1404 in the embodiment shown in FIG. 14; executing steps S1502-S1505 in the embodiment shown in FIG. 15; and executing steps S1602-S1605 in the embodiment shown in FIG. 16. Similar steps are shown in FIG. 17 and FIG. 18.

[0260] Based on the above embodiments, this application also provides a voice interaction system, which may include the electronic device and server as described in the previous embodiments.

[0261] Based on the above embodiments, this application also provides a computer program product, which includes a computer program (also referred to as code or instructions) that, when run, causes a computer to perform the methods described in the embodiments of this application.

[0262] Based on the above embodiments, this application also provides a computer-readable storage medium storing a computer program, which, when executed by a computer, causes the computer to perform the methods described in the embodiments of this application.

[0263] Based on the above embodiments, this application also provides a chip for reading computer programs stored in a memory to implement the methods described in the embodiments of this application.

[0264] Based on the above embodiments, this application provides a chip system including a processor for supporting a computer device in implementing the methods described in the embodiments of this application. In one possible design, the chip system further includes a memory for storing necessary programs and data of the computer device. This chip system may be composed of chips or may include chips and other discrete devices. Those skilled in the art will understand that the embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0265] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in one or more blocks of the flowchart illustrations and / or one or more blocks of the block diagrams.

[0266] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means that implement the functions specified in one or more flowcharts and / or one or more block diagrams.

[0267] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, such that the instructions, which execute on the computer or other programmable apparatus, provide steps for implementing the functions specified in one or more flowcharts and / or one or more block diagrams.

[0268] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the scope of protection of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. A voice interaction method, characterized in that, Applied to an electronic device, the electronic device supports multiple voice input modes, including a first voice input mode and a second voice input mode, and has a display screen; the method includes: The display screen provides a selection entry for multiple voice input methods, including a first voice input method and a second voice input method. The first voice input mode corresponds to the first voice input method, and the second voice input mode corresponds to the second voice input method. The user's selection of a voice input method was detected. Receive the first voice information input by the user; In response to the selection operation instruction to select the first voice input method, in the first voice input mode, a first operation is performed based on the first voice information; or In response to the selection operation instruction to select the second voice input method, in the second voice input mode, the second operation is performed based on the first voice information.

2. The method according to claim 1, characterized in that, The multiple voice input modes include at least one of the following: Voice-based direct input mode; Voice-based text editing mode; Voice-based text query mode; Voice-based text creation mode; Voice-based translation input mode; Voice-based shooting input mode; Traditional Chinese input mode based on voice; Voice-based assisted writing input mode; Voice-based memory input mode; A mode for acquiring cross-application target information based on voice; A mode based on voice instructions for cross-application target operations.

3. The method according to claim 2, characterized in that, The first voice input mode includes the voice-based direct input mode, and the second voice input mode includes at least one of the following: The voice-based text editing mode; The voice-based text query mode; The voice-based text creation mode; The voice-based translation input mode; The voice-based shooting input mode; The traditional Chinese input mode based on voice; The voice-based assisted writing input mode; The voice-based memory input mode; The mode for acquiring cross-application target information based on voice; The mode of cross-application target operation based on voice instructions.

4. The method according to claim 2, characterized in that, If the first voice input mode is a voice-based text editing mode, the first voice information describes the first character in at least one of the following ways: common character combination description method, character decomposition description method, or semantic description method; The step of performing the first operation based on the first voice information includes: The second character displayed on the display screen is replaced with the first character, wherein the first character and the second character are homophones or near-homophones.

5. The method according to claim 2, characterized in that, If the first voice input mode is a mode for acquiring cross-application target information based on voice, then performing the first operation based on the first voice information in the first voice input mode includes: In the first voice input mode, the first voice information input by the user is received on the first interface provided by the electronic device; Based on the first voice information input by the user, the target information is obtained from the second application; The target information is displayed on the first interface.

6. The method according to any one of claims 1-5, characterized in that, The display screen provides multiple voice input method selection options, including: A first button is provided on the display screen; The user's third operation on the first button was detected; According to the third operation, a selection entry for the multiple voice input methods is provided on the display screen, wherein the first key is any one of multiple keys, and the multiple keys correspond one-to-one with the multiple voice input methods.

7. The method according to claim 6, characterized in that, The first button is a physical button; or, the first button is a virtual button.

8. The method according to any one of claims 1-7, characterized in that, The various voice input methods correspond one-to-one with multiple buttons, which can be physical buttons or virtual buttons.

9. The method according to any one of claims 1-8, characterized in that, The selection of various voice input methods is included in the function menu provided by the electronic device.

10. The method according to any one of claims 1-9, characterized in that, The step of performing the second operation based on the first voice information includes: Obtain the semantic analysis results of the first speech information; Based on the semantic analysis results, perform the input operation.

11. The method according to claim 10, characterized in that, The semantic analysis result of obtaining the first speech information includes: Send a first request message to the server, the first request message including the first voice information; Receive a first response from the server, the first response including the semantic analysis result of the first voice information.

12. The method according to claim 11, characterized in that, A large-scale language model is deployed on the server, which is used to perform semantic analysis on the first speech information to obtain the semantic analysis result.

13. The method according to any one of claims 10-12, characterized in that, The step of performing the first operation based on the first voice information includes: Obtain first text information, the first text information including the speech recognition result of the first speech information; Enter the first text information.

14. The method according to claim 13, characterized in that, The acquisition of the first text information includes: Send a second request message to the server, the second request message including the first voice information; Receive a second response from the server, the second response including the first text information obtained by recognizing the first voice information.

15. An electronic device, characterized in that, It includes at least one processor coupled to at least one memory, the at least one processor being configured to read a program stored in the at least one memory to perform the method as described in any one of claims 1-14.

16. A readable storage medium, characterized in that, The readable storage medium stores instructions that, when executed on the device, cause the device to perform the method as described in any one of claims 1-14.

Citation Information

Patent Citations

  • Voice input method and terminal device

    CN106933561A

  • Voice input method and device, and electronic equipment

    CN112199033A

  • Information processing method and device

    CN112306450A

  • Voice editing method and device, equipment and medium

    CN113378530A

  • Voice interaction method and electronic equipment

    CN119847473A