Voice interaction method and electronic device
By displaying voice input and editing entry points on the screen, and automatically editing text using voice input commands, the problem of manual editing caused by voice recognition errors is solved, improving the convenience and user experience of voice editing.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- HUAWEI TECH CO LTD
- Filing Date
- 2025-10-17
- Publication Date
- 2026-05-15
AI Technical Summary
Existing voice input methods are prone to errors during recognition, requiring users to manually edit, which is inconvenient and results in a poor user experience.
The voice input entry is displayed on the screen, and the editing entry is displayed in response to user operations. The input results are automatically edited through voice input editing commands, including detecting the user's first and second selection operations and voice information to achieve editing.
It improves the convenience of voice editing, reduces the need for manual modifications, and enhances the user experience.
Smart Images

Figure CN2025128442_15052026_PF_FP_ABST
Abstract
Description
A voice interaction method and electronic device
[0001] Cross-reference to related applications
[0002] This application claims priority to Chinese Patent Application No. 202411573855.3, filed on November 5, 2024, entitled "A Voice Interaction Method and Electronic Device", the entire contents of which are incorporated herein by reference. Technical Field
[0003] This application relates to the field of electronic device technology, and in particular to a voice interaction method and an electronic device. Background Technology
[0004] Currently, voice input has become the primary interaction method for large-screen devices such as smartphones and tablets. However, due to factors such as homophones, user accents, or environmental noise, electronic devices often encounter conversion errors when recognizing user-inputted speech and converting it into corresponding text. In such cases, users are required to manually edit and correct the erroneous text, which is inconvenient.
[0005] Therefore, improving the convenience of voice editing and addressing the poor user experience remain important issues that urgently need to be addressed. Summary of the Invention
[0006] This application provides a voice interaction method and an electronic device to improve the convenience of voice editing and enhance user experience.
[0007] In a first aspect, a voice interaction method is provided, applied to an electronic device with a display screen, such as a mobile phone. The method includes: displaying a voice input entry on the display screen; detecting a first selection operation by a user on the voice input entry, displaying an editing entry, and detecting first voice information input by the user to obtain a first input result; after detecting a second selection operation by the user on the editing entry, detecting second voice information input by the user to obtain a first editing instruction; editing the second input result according to the first editing instruction to obtain a third input result, wherein the second input result includes the first input result and / or historical input results; and displaying the third input result on the display screen.
[0008] Based on the method described in the first aspect, the mobile phone can detect the user's first selection operation, acquire the user's voice input information, and display the editing entry on the screen. The mobile phone can also detect the user's second selection operation on the editing entry, acquire the user's second voice input information, which is the voice of the editing command. The electronic device can also obtain the editing command based on the second voice information and edit the input result. In this editing method, the user does not need to manually modify the input result; they only need to input the editing command via voice, which improves the convenience of voice editing and enhances the user experience.
[0009] In one possible design, the second voice information includes information about the original keyword from the second input result and / or information about the target keyword from the third input result. The information about the original keyword includes at least one of the following: the pinyin of the original keyword, common character combinations of the original keyword, information about the components of the original keyword, semantic description information of the original keyword, or a pronoun, synonym, near-synonym, or similar word of the original keyword, to support the electronic device in accurately and efficiently recognizing the original keyword based on its voice information. The information about the target keyword includes at least one of the following: the pinyin of the target keyword, common character combinations of the target keyword, information about the components of the target keyword, semantic description information of the target keyword, or a pronoun, near-synonym, near-synonym, or similar word of the target keyword, to support the electronic device in accurately and efficiently recognizing the original keyword based on its voice information.
[0010] In one possible design, the information of the original keyword includes its referent. Accordingly, the electronic device can determine the original keyword from the second input result based on the referent. Based on this design, accurate and efficient keyword determination can be achieved even when the second voice information does not contain the pinyin of the original keyword.
[0011] In one possible design, the second voice information includes information about the target keyword. Accordingly, the electronic device can determine the target keyword based on the target keyword information in the second voice information, and determine the original keyword from the second input result based on the type and / or semantics of the target keyword. Therefore, even when the second voice information does not contain information about the original keyword, accurate and efficient determination of the original keyword can be achieved based on the target keyword.
[0012] In one possible design, the mobile phone can determine multiple candidate results for a keyword based on the second voice information, the keyword including the original keyword and / or the target keyword; display the multiple candidate results on the display screen; detect the user's third selection operation on the target candidate result among the multiple candidate results; and determine the keyword based on the target candidate result. Based on this design, when multiple candidate results exist for a keyword, the keyword can be determined by the user's selection of the target candidate result among the multiple candidate results.
[0013] In one possible design, the second voice information includes the pinyin of the keyword and information about the components of the keyword. The information about the components includes radical information. The mobile phone can determine the radical of the keyword based on the radical information and determine multiple candidate results for the keyword based on the radical and the pinyin. Based on this design, the mobile phone can accurately determine the keyword by combining the pinyin and component information. The components of the keyword may include radicals or non-radical components.
[0014] In one possible design, the information of the components of the keyword also includes the pinyin of the non-radical part. The mobile phone can determine multiple candidate results for the keyword based on the radical of the keyword, the pinyin of the non-radical part, and the pinyin of the keyword itself. Based on this design, the mobile phone can accurately determine the keyword by combining the pinyin of the keyword, the pinyin of the radical, and the pinyin of the non-radical part.
[0015] In one possible design, the original keywords include discontinuous text in the second input result. The mobile phone can edit the discontinuous text and the text information between the discontinuous text according to the first editing instruction to obtain the third input result. Therefore, discontinuous content in the original text can be edited based on the second voice information.
[0016] In one possible design, the text information between the two discontinuous texts is a first type of symbol, and the second voice information includes information for instructing the editing of the discontinuous text. The first type of symbol is used to segment text information within a sentence. Based on this implementation, even if the first type of symbol in the original text is omitted in the user's voice input editing instruction, the original text to be edited can be reasonably matched, improving editing accuracy and efficiency.
[0017] In one possible design, the second voice information includes information for editing the discontinuous text and the text information located between the discontinuous text. For example, if the input result corresponding to the second voice information contains the sentence structure "from A to B", the mobile phone can find "A" and "B" in the second input result and edit "A", "B" in the original text, as well as the content located between "A" and "B".
[0018] In one possible design, the information of the original keyword includes its pinyin. The mobile phone can determine the keyword from the second input result based on the pinyin and the cursor position. This design can improve the accuracy of determining the original keyword; for example, it can improve the accuracy of keyword determination when the original keyword is a polyphonic character.
[0019] In one possible design, the mobile phone can determine the keyword within the n1 characters preceding the cursor position in the second input result based on the pinyin of the keyword, where n1 is a positive integer; and / or, determine the keyword within the n2 characters following the cursor position in the second input result based on the pinyin of the keyword, where n2 is a positive integer. Alternatively, the mobile phone can determine the second input result based on the Chinese characters within the n1 characters preceding the cursor position and / or the Chinese characters within the n2 characters following the cursor position. This design can improve the accuracy of keyword recognition. For example, in scenarios where the keyword is a polyphonic character, the second input result to be edited may contain multiple keywords with different pronunciations. In this case, the keyword to be edited can be selected from these multiple keywords with different pronunciations based on the cursor position.
[0020] In one possible design, the information of the original keyword includes the pinyin of the original keyword. The electronic device can also determine the keyword based on the pinyin of the keyword and the voice input record and / or pinyin input record of the second input result. This embodiment can improve the accuracy of keyword recognition. For example, in a scenario where the keyword is a polyphonic character, the second input result to be edited may contain a polyphonic character as a keyword. In this case, the electronic device can determine the pinyin of the keyword input based on the voice input record and / or pinyin input record of the keyword in the second input result, and determine whether the keyword is the keyword to be edited based on the pinyin of the keyword input and the pinyin of the keyword in the second voice information.
[0021] In one possible design, after detecting the second voice information, the mobile phone can recognize the second voice information, obtain the input result corresponding to the second voice information, and perform semantic analysis on the input result corresponding to the second voice information to obtain the first editing instruction. Specifically, the mobile phone can perform semantic analysis based on the sentence structure of the input result corresponding to the second voice information to improve the reliability of the analysis. Alternatively, the mobile phone can perform semantic analysis using a model to improve the efficiency and reliability of the semantic analysis.
[0022] In one possible design, the mobile phone can determine one or more words in the input result corresponding to the second voice information through a word segmentation model, determine the classification information corresponding to the word through an edit parameter classification model, the classification information of any word is used to indicate the attribute of the word in the first editing instruction, and the first editing instruction is determined based on the word and the classification information.
[0023] In one possible design, the mobile phone can determine the attributes of one or more word segments in the input result corresponding to the second voice information in the first editing instruction through a large-scale language model; and determine the first editing instruction based on the word segments and the attributes.
[0024] Secondly, a user interaction method is also provided, which is applied to an electronic device with a display screen, such as a mobile phone.
[0025] The method includes detecting a second selection operation by the user on the editing entry point, detecting second voice information input by the user, and obtaining a first editing instruction. The input result corresponding to the second voice information is then displayed on the display screen.
[0026] In one possible design, the method further includes: displaying multiple candidate results of keywords in the second voice information, the keywords including original keywords in the second input result to be edited and / or target keywords in the edited third voice input result; detecting the user's third selection operation on the target candidate result among the multiple candidate results; and determining the keyword based on the target candidate result.
[0027] In one possible design, after detecting the user's second selection operation on the editing entry, the mobile phone can display an editing instruction display area on the display screen, and display the input result and / or candidate result of the keyword corresponding to the second voice information in the editing instruction display area.
[0028] In one possible design, the mobile phone can display the input result corresponding to the second voice information and / or the candidate result of the keyword in the input method display area on the display screen. The input method display area is also used to display one or more of the voice input entry, editing entry, or keyboard.
[0029] The first aspect and any possible design thereof can be implemented in combination with the second aspect and any possible design thereof.
[0030] Thirdly, an electronic device is also provided, comprising:
[0031] Processor, memory, and one or more programs;
[0032] The one or more programs are stored in the memory, and the one or more programs include instructions that, when executed by the processor, cause the electronic device to perform the method provided in the first or second aspect above.
[0033] Fourthly, a computer-readable storage medium is also provided for storing a computer program that, when run on a computer, causes the computer to perform the methods provided in the first or second aspect above.
[0034] Fifthly, a computer program product is also provided, comprising a computer program that, when run on a computer, causes the computer to perform the methods provided in the first or second aspect above.
[0035] In a sixth aspect, a chip is also provided, which is coupled to a memory in an electronic device for calling a computer program stored in the memory and executing the technical solutions provided in the first or second aspect of the embodiments of this application. In the embodiments of this application, "coupling" means that two components are directly or indirectly combined with each other.
[0036] In a seventh aspect, a chip system is also provided, the chip system including a processing circuit and a storage medium, the storage medium storing instructions; when the instructions are executed by the processing circuit, they implement the method described in the first aspect above.
[0037] For the technical effects that can be achieved in the second to seventh aspects mentioned above, please refer to the description of the technical effects that can be achieved by the corresponding design scheme in the first aspect mentioned above. This application will not repeat them here. Attached Figure Description
[0038] Figure 1 is a schematic diagram of an electronic device provided in an embodiment of this application;
[0039] Figure 2 is a schematic diagram of another structure of the electronic device provided in an embodiment of this application;
[0040] Figure 3 is a schematic diagram of an application scenario of an electronic device provided in an embodiment of this application;
[0041] Figure 4 is a schematic diagram of a voice interaction scenario provided in an embodiment of this application;
[0042] Figure 5 is a schematic diagram of another voice interaction scenario provided in an embodiment of this application;
[0043] Figure 6 is a schematic diagram of another voice interaction scenario provided in an embodiment of this application;
[0044] Figure 7 is a schematic diagram of another voice interaction scenario provided in an embodiment of this application;
[0045] Figure 8 is a schematic diagram of another voice interaction scenario provided in an embodiment of this application;
[0046] Figure 9 is a schematic diagram of another voice interaction scenario provided in an embodiment of this application;
[0047] Figure 10 is a schematic diagram of another voice interaction scenario provided in an embodiment of this application;
[0048] Figure 11 is a schematic diagram of a voice editing method provided in an embodiment of this application;
[0049] Figure 12 is a schematic diagram of another voice editing method provided in an embodiment of this application;
[0050] Figure 13 is a schematic diagram of another voice editing method provided in an embodiment of this application;
[0051] Figure 14 is a schematic diagram of another voice editing method provided in an embodiment of this application;
[0052] Figure 15 is a schematic diagram of another voice editing method provided in an embodiment of this application;
[0053] Figure 16 is a schematic diagram of another voice editing method provided in an embodiment of this application;
[0054] Figure 17 is a schematic diagram of another voice editing method provided in an embodiment of this application;
[0055] Figure 18 is a schematic diagram of another voice editing method provided in an embodiment of this application;
[0056] Figure 19 is a schematic diagram of a voice interaction method provided in an embodiment of this application;
[0057] Figure 20 is a schematic diagram of the structure of another electronic device provided in an embodiment of this application. Detailed Implementation
[0058] The embodiments of this application will now be described in detail with reference to the accompanying drawings and examples.
[0059] Currently, electronic devices are increasingly used in people's work and daily lives. With the development of voice recognition technology, voice input has become the primary interaction method for large-screen devices such as smartphones and tablets. However, due to factors such as homophones, user accents, or environmental noise, electronic devices often encounter conversion errors when recognizing user-inputted speech and converting it into corresponding text. In such cases, users need to manually edit and correct the erroneous text.
[0060] Therefore, improving the convenience of voice editing remains an important issue that urgently needs to be addressed.
[0061] To address the aforementioned problems, this application provides a voice interaction method and an electronic device. In this method, the electronic device can display a voice input field on a screen and, in response to a user's selection of the voice input field, further display an editing field. The electronic device can also detect the user's voice input after detecting the user's selection of the editing field, and obtain a first editing instruction based on the voice input. The electronic device can further edit the input result according to the first editing instruction. Users can interact with the electronic device through voice input to achieve intelligent control of the electronic device and edit input results, freeing up the user's hands and improving the convenience of voice editing.
[0062] The technical solutions in this application can be applied to electronic devices, which can be any device with a display screen or associated display screen. For example, electronic devices can be mobile phones, foldable phones, tablets, wearable devices (e.g., watches, bracelets, glasses, etc.), in-vehicle devices, augmented reality (AR) / virtual reality (VR) devices, laptops, ultra-mobile personal computers (UMPCs), netbooks, personal digital assistants (PDAs), smart home devices (e.g., smart TVs, etc.). This application does not limit the specific type of electronic device.
[0063] The electronic devices to which this application can be applied can also be portable terminal devices that include other functions such as personal digital assistants and / or music players. Or electronic devices with other operating systems.
[0064] Figure 1 illustrates a possible hardware structure diagram of an electronic device. The electronic device 200 includes components such as a radio frequency (RF) circuit 210, a power supply 220, a processor 230, a memory 240, an input unit 250, a display unit 260, an audio circuit 270, a communication interface 280, and a wireless fidelity (Wi-Fi) module 290. Those skilled in the art will understand that the hardware structure of the electronic device 200 shown in Figure 1 does not constitute a limitation on the electronic device 200. The electronic device 200 provided in this application embodiment may include more or fewer components than shown, may combine two or more components, or may have different component configurations. The various components shown in Figure 1 can be implemented in hardware, software, or a combination of hardware and software, including one or more signal processing and / or application-specific integrated circuits (ASICs).
[0065] The following is a detailed description of each component of the electronic device 200 with reference to Figure 1:
[0066] The RF circuit 210 can be used for receiving and transmitting data during communication or a call. Specifically, after receiving downlink data from the base station, the RF circuit 210 sends it to the processor 230 for processing; additionally, it sends uplink data to be transmitted to the base station. Typically, the RF circuit 210 includes, but is not limited to, an antenna, at least one amplifier, a transceiver, a coupler, a low-noise amplifier (LNA), a duplexer, etc.
[0067] Furthermore, the RF circuit 210 can also communicate with other devices via a wireless communication network. The wireless communication can use any communication standard or protocol, including but not limited to Global System for Mobile Communication (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, and Short Message Service (SMS).
[0068] Wi-Fi technology is a short-range wireless transmission technology. The electronic device 200 can connect to an access point (AP) via the Wi-Fi module 290, thereby enabling access to the data network. The Wi-Fi module 290 can be used for receiving and sending data during communication.
[0069] The electronic device 200 can physically connect to other devices through the communication interface 280. Optionally, the communication interface 280 can be connected to the communication interfaces of other devices via a cable to enable data transmission between the electronic device 200 and other devices.
[0070] The electronic device 200 can also perform communication services and interact with other electronic devices. Therefore, the electronic device 200 needs to have data transmission capabilities, meaning that it needs to include a communication module. Although Figure 1 shows the RF circuit 210, the Wi-Fi module 290, and the communication interface 280, it is understood that the electronic device 200 contains at least one of the aforementioned components or other communication modules (such as a Bluetooth module) for data transmission.
[0071] For example, when the electronic device 200 is a mobile phone, the electronic device 200 may include the RF circuit 210, the Wi-Fi module 290, or a Bluetooth module (not shown in Figure 1); when the electronic device 200 is a tablet computer, the electronic device 200 may include the Wi-Fi module or a Bluetooth module (not shown in Figure 1); when the electronic device 200 is a smart home device, the electronic device 200 may include the Wi-Fi module 290 or a Bluetooth module (not shown in Figure 1).
[0072] The memory 240 can be used to store software programs and modules. The processor 230 executes various functional applications and data processing of the electronic device 200 by running the software programs and modules stored in the memory 240. Optionally, the memory 240 may mainly include a program storage area and a data storage area. The program storage area may store the operating system (mainly including the software programs or modules corresponding to the kernel layer, system layer, application framework layer, and application layer).
[0073] In addition, the memory 240 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0074] The input unit 250 can be used to receive editing operations on various types of data objects, such as numbers or characters, input by the user, and to generate key signal inputs related to user settings and function control of the electronic device 200. Optionally, the input unit 250 may include a touch panel 251 and other input devices 252.
[0075] The touch panel 251, also known as a touchscreen, can collect user touch operations on or near it (such as operations performed by the user using a finger, stylus, or any suitable object or accessory on or near the touch panel 251), and drive corresponding connection devices according to a pre-set program. In this embodiment, the touch panel 251 can collect user operations on it. For example, the user operation may include clicking a physical button on the touch panel 251, or the user operation may also include pressing and holding a physical button on the touch panel 251 to switch the physical button to another function.
[0076] Optionally, the other input device 252 may include, but is not limited to, one or more of the following: a physical keyboard, an infrared sensor, function keys (such as volume control buttons, power buttons, etc.), a trackball, a mouse, a joystick, etc. For example, an infrared sensor can be used to acquire the user's air gesture operations.
[0077] The display unit 260 can be used to display information input by the user or information provided to the user, as well as various menus of the electronic device 200. The display unit 260 is the display system of the electronic device 200, used to present the interface and realize human-computer interaction. The display unit 260 may include a display panel 261. Optionally, the display panel 261 can be configured using a liquid crystal display (LCD), organic light-emitting diode (OLED), or other similar forms. In this embodiment, the display unit 260 can be used to display a user interface. For example, the user interface can be the desktop (or main interface) of the electronic device, or it can be an interface displaying various possible business scenarios, such as the chat interface of an instant messaging app, a memo interface, or the interface of a multimedia video app.
[0078] The processor 230 is the control center of the electronic device 200. It connects various components via various interfaces and lines, and executes software programs and / or modules stored in the memory 240, as well as calling data stored in the memory 240, to perform various functions and process data of the electronic device 200, thereby enabling various services based on the electronic device 200. In this embodiment, the processor 230 can be used to implement a voice interaction method provided in this embodiment.
[0079] The electronic device 200 also includes a power supply 220 (such as a battery) for supplying power to various components. Optionally, the power supply 220 can be logically connected to the processor 230 through a power management system, thereby enabling the power management system to manage functions such as charging, discharging, and power consumption.
[0080] As shown in Figure 1, the electronic device 200 also includes an audio circuit 270, a microphone 271, and a speaker 272, providing an audio interface between the user and the electronic device 200. The audio circuit 270 converts audio data into signals recognizable by the speaker 272 and transmits the signals to the speaker 272, where they are converted into sound signals for output. The microphone 271 collects external sound signals (such as human speech or other sounds) and converts these signals into signals recognizable by the audio circuit 270, sending them to the audio circuit 270. The audio circuit 270 can also convert the signals transmitted by the microphone 271 into audio data, outputting the audio data to the RF circuit 210 for transmission to, for example, another electronic device, or outputting the audio data to the memory 240 for further processing.
[0081] Although not shown in Figure 1, the electronic device 200 may also include a camera, at least one sensor, etc., which will not be described in detail here. The at least one sensor may include, but is not limited to, a pressure sensor, a barometric pressure sensor, an accelerometer, a distance sensor, a fingerprint sensor, a touch sensor, a temperature sensor, etc.
[0082] Figure 2 is a software architecture block diagram of an electronic device provided in an embodiment of this application. As shown in Figure 2, the software structure of the electronic device can be a layered architecture. For example, the software can be divided into several layers, each with a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the operating system is divided into four layers, from top to bottom: the application layer, the application framework layer (framework, FWK), the runtime and system libraries, and the kernel layer.
[0083] The application layer may include a series of application packages. As shown in Figure 2, the application layer may include a camera, settings, skin modules, user interface (UI), third-party applications, etc. Third-party applications may include gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, SMS, etc. In this embodiment, the application layer may include a target installation package of a target application that the electronic device requests to download from a server. The function files and layout files in this target installation package are adapted to the electronic device.
[0084] The application framework layer provides application programming interfaces (APIs) and a programming framework for applications in the application layer. The application framework layer can include some predefined functions. As shown in Figure 2, the application framework layer can include a window manager, content provider, view system, phone manager, resource manager, and notification manager.
[0085] The window manager is used to manage windowed applications. It can obtain the screen size, determine if a status bar is present, lock the screen, and capture screenshots. The content provider stores and retrieves data, making this data accessible to applications. This data may include videos, images, audio, made and received phone calls, browsing history and bookmarks, phone books, etc.
[0086] A view system includes visual controls, such as controls for displaying text and controls for displaying images. View systems can be used to build applications. A display interface can consist of one or more views. For example, a display interface including a text notification icon could include views for displaying text and views for displaying images.
[0087] A phone manager is used to provide communication functions for electronic devices. For example, it manages call status (including connection and disconnection).
[0088] The file explorer provides applications with various resources, such as localized strings, icons, images, layout files, video files, and more.
[0089] The notification manager allows applications to display notifications in the status bar. These notifications can be used to deliver informational messages and can disappear automatically after a short pause, requiring no user interaction. For example, the notification manager can be used to notify users of completed downloads or message alerts. The notification manager can also display notifications as icons or scrolling text in the top status bar, such as notifications from background applications, or as dialog boxes on the screen. Examples include displaying text messages in the status bar, emitting sounds, vibrating electronic devices, and flashing indicator lights.
[0090] The runtime includes the core libraries and the virtual machine. The runtime is responsible for the scheduling and management of the operating system.
[0091] The core library consists of two parts: one part contains the functionalities that the computer programming language needs to call, and the other part is the core library of the operating system. The application layer and application framework layer run in a virtual machine. Taking Java as the programming language as an example, the virtual machine executes the Java files in the application layer and application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.
[0092] System libraries can include multiple functional modules. For example: surface manager, media libraries, 3D graphics processing libraries (e.g., OpenGL ES), 2D graphics engines (e.g., SGL), image processing libraries, etc.
[0093] The Surface Manager is used to manage the display subsystem and provides the fusion of two-dimensional (2D) and three-dimensional (3D) layers for multiple applications.
[0094] The media library supports playback and recording of various common audio and video formats, as well as still image files. It supports multiple audio and video encoding formats, such as MPEG4, H.264, MP3, AAC, AMR, JPG, and PNG.
[0095] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.
[0096] A 2D graphics engine is a drawing engine for 2D graphics. It can perform drawing operations, such as drawing different voice buttons or different shapes of voice buttons on the screen of an electronic device to receive different voice input information from the user in different voice interaction scenarios. Alternatively, a 2D graphics engine can also draw menus on the display screen of an electronic device. These menus can contain entries for voice input modes and / or editing modes, allowing users to select the desired mode and receive corresponding functions.
[0097] The kernel layer is the layer between hardware and software. The kernel layer contains at least the display driver, camera driver, audio driver, and sensor driver.
[0098] The hardware layer can include various types of sensors, such as accelerometers, gyroscopes, and touch sensors.
[0099] It should be noted that the structures shown in Figures 1 and 2 are merely examples of electronic devices provided in the embodiments of this application, and cannot be used to limit the electronic devices provided in the embodiments of this application. In specific implementations, electronic devices may have more or fewer devices or modules than those shown in Figures 1 or 2.
[0100] Typically, an electronic device 200 can run multiple applications simultaneously. In a simpler scenario, one application corresponds to one process; in a more complex scenario, one application can correspond to multiple processes. Each process has a unique process ID.
[0101] It should be understood that in the embodiments of this application, "at least one of the following" or similar expressions refer to any combination of these items, including any combination of a single item or a plurality of items. For example, at least one of a, b, or c can represent: a, b, c, a and b, a and c, b and c, or a, b, and c, where a, b, and c can be single or multiple. "Multiple" refers to two or more. "And / or" is used to describe the association relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship.
[0102] Furthermore, it should be understood that in the description of this application, terms such as "first" and "second" are used only for the purpose of distinguishing descriptions and should not be construed as indicating or implying relative importance, nor should they be construed as indicating or implying order. For example, in the following embodiments, "first voice information" and "second voice information" are only used to distinguish different voice information and are not used to limit the specific implementation method of the voice input method, etc.
[0103] It should be understood that the hardware structure of the electronic device can be as shown in Figure 1, and the software system architecture can be as shown in Figure 2. The software programs and / or modules corresponding to the software system architecture in the electronic device can be stored in the memory 240, and the processor 230 can run the software programs and applications stored in the memory 240 to execute the flow of a voice interaction method provided in the embodiments of this application.
[0104] Figure 3 shows a schematic diagram of a system architecture that may be applicable to the voice interaction method provided in the embodiments of this application.
[0105] As shown in Figure 3A, the system architecture may include electronic device 200, user 41, and server 42.
[0106] The electronic device 200 can receive voice information input by the user 41. Based on the voice information, the electronic device 200 can send a request message to the server 42 through a communication network, requesting the server 42 to perform voice recognition, semantic analysis (or semantic understanding), and other processing on the voice information.
[0107] For example, server 42 can be equipped with an automatic speech recognition (ASR) module. This ASR module can use ASR and other technologies to perform speech recognition on the speech information from electronic device 200 and convert the speech information into corresponding text information, thus obtaining the speech recognition result. Alternatively, server 42 can be equipped with a large language model (LLM) (referred to as the large model). This large model can perform semantic analysis on the speech information from electronic device 200 and obtain the corresponding semantic analysis results.
[0108] Server 42 can send the speech recognition result or semantic analysis result as a response to the aforementioned request message to electronic device 200. Accordingly, electronic device 200 can receive the response information from server 42, display the received response information on a display screen, and / or perform corresponding operations based on the response information.
[0109] In another possible implementation, as shown in Figure 3B, the voice interaction method can also be implemented by the electronic device 200 itself, as described in this embodiment. For example, the electronic device 200 may have an ASR module deployed therein, which can be used to perform speech recognition on the voice information input by the user 41 to obtain converted text information. Alternatively, the electronic device 200 may have a large model deployed therein, which can perform semantic analysis on the voice information input by the user 41 to obtain corresponding semantic analysis results. The electronic device 200 can display the speech recognition results or semantic analysis results on its display screen, and can also perform corresponding operations based on the speech recognition results or semantic analysis results.
[0110] It should be understood that, in the embodiments of this application, in the scenario shown in Figure 3A, the ASR module and the large model can be deployed on the same server or on different servers. The electronic device can directly interact with the corresponding server, or indirectly interact with the corresponding server through forwarding from other devices. For example, the electronic device can send a request message containing voice information to the application server of the currently running application. The application server can forward the received request message to the server deploying the large model. The server deploying the large model performs semantic analysis on the received voice information and returns the semantic understanding result to the application server. The application server can encapsulate the semantic analysis result and feed it back to the electronic device. This embodiment of the application does not limit this.
[0111] Based on the structural diagrams in Figures 1-2 and the system architecture diagram in Figure 3, the following uses a mobile phone as an example to introduce several possible voice interaction scenarios involved in the embodiments of this application.
[0112] In this embodiment, the electronic device can provide a voice input entry point to the user on a display screen, and enter a voice input mode based on the user's first selection operation on this entry point, displaying an editing entry point. In voice input mode, the electronic device can acquire a first voice signal to obtain a first input result. When editing the input result is required, the user can perform a second selection operation on the editing entry point. The electronic device can detect the user's second selection operation on the editing entry point and enter editing mode. In editing mode, the user can input editing instructions through second voice information. The electronic device can detect the second voice information and obtain a first editing instruction. Furthermore, the electronic device can edit the input result according to the first editing instruction. This input result can be referred to as the second input result. The input result obtained by editing the second input result according to the first editing instruction can be referred to as the third input result.
[0113] Therefore, in this editing method, users do not need to manually modify the input results; they only need to input editing commands via voice. Furthermore, the electronic device's display screen does not need to continuously display the editing entry point; it only needs to display it after detecting the user's selection of the voice input entry point. This improves the simplicity of the display interface and reduces the power consumption of the display screen.
[0114] In voice input mode, users can input various types of voice information into the electronic device, which can then perform voice recognition and display the input. For example, the electronic device can obtain the corresponding input result based on an ASR module and / or an LLM.
[0115] In this application, voice or voice information refers to a sound signal originating from a user and / or other input source. The input result can refer to text, voice, or other information obtained by an electronic device from a voice or other form of input signal. Other forms of input signals may include manually entered input signals by the user or communication signals from other devices, and are not specifically limited thereto.
[0116] Optionally, the display screen may also show the first input result obtained based on the user's first voice information in voice input mode.
[0117] The editing mode, also known as a voice-based text editing mode, is a mode in this application that allows for editing of text content or formatting based on voice input. In editing mode, the electronic device can edit the second input result based on the user's second voice information and obtain the edited text information.
[0118] Specifically, after detecting the second voice information, the electronic device can obtain the recognition result of the second voice information through the ASR module, and then determine the semantic analysis result through semantic analysis of the recognition result. The recognition result of the second voice information can also be called the input result corresponding to the second voice information or the first editing instruction; it refers to the text information obtained by performing speech recognition on the second voice information. The semantic analysis result corresponding to the second voice information can be understood as the first editing instruction, such as an instruction to delete, add, or modify text. Subsequently, the second input result can be edited based on this voice analysis result, i.e., the first editing instruction, to obtain the third input result.
[0119] It can be understood that "the second voice information contains A" as described in this application can be understood as the second voice information containing the voice information corresponding to A, or it can be understood as the input result corresponding to the second voice information containing the text information corresponding to A.
[0120] Optionally, the display screen may also display the recognition result of the second voice information, that is, the input result corresponding to the second voice information or the first editing command.
[0121] The second input result can be understood as text information displayed in a certain area of the display screen. The text information displayed on the display screen here can be text information displayed on the interface of reading applications or other applications, or it can be text information obtained after recognizing the user's voice information. This application embodiment does not limit the method of obtaining this modified text information.
[0122] Specifically, the second input result may include the first input result entered by the user in voice mode, or it may include the user's historical input results. This historical input result may include input results from voice mode, or input results entered by the user using manual input methods such as Pinyin or handwriting. Optionally, the second input result may also include text read locally by the electronic device, and / or text from other devices.
[0123] The editing operations on text content may include, but are not limited to, at least one of the following: adding, deleting, or modifying text content. Editing text formatting may include text wrapping, first-line indentation, etc. This application does not specifically limit the editing operations involved in the voice-based text editing mode.
[0124] As an example, the second input result can be related to the cursor position in the text box. For instance, the second input result for editing could include n1 characters before the cursor position and / or n2 characters after the cursor position, where n1 and n2 are positive integers. Furthermore, the second input result to be edited could also include a complete sentence corresponding to the n1 characters before the cursor position and / or the n2 characters after the cursor position. Additionally, the second input result could be the complete sentence where the cursor is currently located, or all the text information in the current text box.
[0125] The entry point in this application can be keyed.
[0126] In one possible embodiment, the voice input entry point in this application can be a virtual button. Alternatively, in some implementations, a physical button (e.g., a physical button) can be used to implement the voice input entry point, in which case it is not necessary to display the voice input entry point on the display screen. Similarly, the editing entry point can be a virtual button. Alternatively, in some implementations, a physical button can be used to implement the editing entry point, in which case it is not necessary to display the editing entry point on the display screen.
[0127] In one possible embodiment, the voice input entry and the editing entry can be different buttons, that is, the electronic device can provide different buttons to associate different entries or modes.
[0128] In some alternative implementations, the same button can be used in different presentation formats to represent the voice input entry and the editing entry, with each corresponding to a different presentation format. For example, the voice input entry and the editing entry can be the same virtual button. The user can trigger the input mode by making a first selection operation on the button and trigger the editing mode by making a second selection operation on the button. Different presentation formats of the button can represent different entry points or modes. For example, the first presentation format of the virtual button represents the voice input entry, and the second presentation format represents the editing entry. In the first and second presentation formats, at least one of the following is different: the shape, color, size, button symbol, or button icon of the virtual button.
[0129] It is understandable that if the voice input entry and the editing entry are the same virtual button, then "displaying the editing entry" can mean that the electronic device responds to the user's first selection operation by changing the presentation of the virtual button to the presentation of the editing entry.
[0130] In optional implementations, different states can be set for the same button to indicate whether the entry or mode associated with the button is activated. For example, the button's state can include an active state and an inactive state. The active state, also known as the function-activated state, indicates that the function / mode associated with the button is activated or started. The inactive state, also known as the non-activated state, indicates that the function / mode associated with the button is not activated or started.
[0131] In one possible implementation, the "selection operation" in this application may include a user's click operation and / or long-press operation on a virtual button. The click operation may be a single click or multiple clicks. Additionally, for virtual buttons, the selection operation may also include the user ending the long-press operation at the virtual button's location. For example, the selection operation for a virtual button may include: the user starting a long-press operation at a location other than the virtual button, then dragging the long-press location to the virtual button and ending the long-press operation. Furthermore, the "selection operation" may also include a user's press operation and / or long-press operation on a physical button.
[0132] To facilitate understanding, the following is an illustrative description of the voice interaction scenarios involved in the editing mode, with reference to the accompanying diagram.
[0133] Voice interaction scenario A:
[0134] Referring to Figure 4, this is a schematic diagram of a possible scenario for which the voice interaction method provided in this application embodiment may be applied. Interface 401 in Figure 4 shows a schematic diagram of the memo interface displayed on the mobile phone screen. In interface 401, on the newly added memo interface, the mobile phone can provide the user with a voice input entry for voice input mode, represented as button 1 in interface 401. The user can select the voice input mode by selecting button 1. For example, the user can long-press or double-click button 1 in interface 401 to select the voice input mode. In voice input mode, the user can input first voice information into the mobile phone, for example, represented as voice 1. As an example, voice 1 could be the voice input of the following recipe into the mobile phone's memo: "Sauté chopped onions, scallions, celery, mushroom slices, capers, black olives, and cherry tomatoes in olive oil. Add butter, sea salt, and black pepper." The mobile phone can perform voice recognition on the content of voice 1 to obtain the first input result corresponding to voice 1, and display the first input result in the memo on interface 402, represented as text 1 in Figure 4.
[0135] Additionally, as shown in Figure 4, the mobile phone can respond to the user's selection operation on button 1 by providing an editing entry in interface 402. This editing entry is represented as button 2 in interface 402. Button 2 can represent the editing entry using text or icons. This application does not require a specific style for button 2. Optionally, the mobile phone can also respond to the user's selection operation on button 1 by providing a voice input cancellation entry in interface 402, represented as button 3 in interface 402. If the user triggers this cancellation entry, local voice input can be cancelled.
[0136] Figure 5 shows an example of an editing mode implementation in scenario A. The user can long-press button 1 on interface 501 to enter voice input mode. To edit text 1, the user can continue pressing and drag the pressed position to button 2. The phone can recognize the user's long-press position; if the long-press position is moved to the position corresponding to button 2, the user can enter editing mode. The interface corresponding to editing mode includes interface 502. Interface 502 can be found in the description of interface 402 and will not be repeated here.
[0137] In edit mode, users can input a second voice message into the phone, such as voice 2. As shown in interface 503, during the recognition of voice 2, or before the input result of voice 2 is recognized, interface 503 can display text such as "Listening to instructions" or "Please enter instructions" to improve the user experience. After the input result of voice 2 is recognized, the corresponding input result of voice 2 can be displayed in interface 504. As an example, voice 2 could be the voice input of the following recipe into the phone's memo: "Change black pepper powder to white peppercorns." The phone can perform voice recognition on the content of voice 2 to obtain the corresponding input result of voice 2. For example, the input result corresponding to the voice editing instruction can be displayed in the editing instruction display area of interface 504, as shown in text 2.
[0138] The phone can also obtain the semantic analysis result corresponding to text 2. According to the semantic analysis result, it is known that "black pepper powder" in text 1 of the memo needs to be changed to "white peppercorns," that is, the first editing instruction is to change "black pepper powder" to "white peppercorns." This voice analysis result can be understood as the first editing instruction. The phone can also perform the corresponding modification operation based on this semantic analysis result, that is, change "black pepper powder" in text 1 to "white peppercorns." Therefore, as shown in interface 505, the text content of the recipe recorded in the phone's memo is changed to the following content in text 3: "Sauté chopped onions, scallions, celery, mushroom slices, capers, black olives, and cherry tomatoes in olive oil. Add butter, sea salt, and white peppercorns."
[0139] Optionally, in this example, the phone can perform an editing operation based on the semantic analysis results of voice 2 after detecting that the user has ended the long press operation.
[0140] Figure 6 shows an example of another implementation method for the editing mode in scenario A. The user can enter voice input mode by clicking button 1 on interface 601, such as by single-clicking or double-clicking button 1. To edit text 1, the user can click or long-press button 2. Taking long-pressing button 2 as an example, the phone can recognize the user's long-press operation and enter editing mode. The interface corresponding to editing mode includes interface 602. For a description of interface 602, please refer to the description of interface 602; it will not be repeated here.
[0141] In edit mode, users can input a second voice message into the phone, such as voice 2. As shown in interface 603, during the recognition of voice 2, or before the input result of voice 2 is recognized, interface 603 can display text such as "Listening to instructions" or "Please enter instructions" to improve the user experience. After the input result of voice 2 is recognized, the corresponding input result of voice 2 can be displayed in interface 604. As an example, voice 2 can be the following recipe entered into the phone's memo: "Change black pepper powder to white peppercorns." The phone can perform voice recognition on the content of voice 2 to obtain the corresponding input result of voice 2. For example, the phone can display the input result of the voice editing instruction "Change black pepper powder to white peppercorns" in the editing instruction display area of interface 604, as shown in text 2.
[0142] The phone can also obtain the semantic analysis result corresponding to text 2. According to the semantic analysis result, it is known that "black pepper powder" in text 1 of the memo needs to be changed to "white peppercorns". This voice analysis result can be understood as the first editing instruction. The phone can also perform the corresponding modification operation according to the semantic analysis result, that is, change "black pepper powder" in text 1 to "white peppercorns". Therefore, as shown in interface 605, the text content of the recipe recorded in the phone's memo is changed to the following content in text 3: "Sauté chopped onions, scallions, celery, mushroom slices, capers, black olives, and cherry tomatoes in olive oil. Add butter, sea salt, and white peppercorns."
[0143] Optionally, in this example, the phone can perform an editing operation based on the semantic analysis results of voice 2 after detecting that the user has ended the long press operation.
[0144] It should be noted that the specific content and display format of interfaces 401 to 402, 501 to 505, and 601 to 605 are not limited in the embodiments of this application. For example, Figures 4 to 6 are merely illustrative examples of the interfaces above. Button 1 and / or button 2 in the interfaces above can be located at the bottom of the interface or at the top of the interface, etc., and this application does not make specific requirements. In addition, button 1 and / or button 2 can be displayed separately on the interface or in the function menu on the interface. This function menu is the selection entry point for multiple buttons, and this application does not make limitations. In other embodiments of this application, the user can also use any of the methods shown in Figures 4 to 6 to edit text in editing mode according to the user's voice input. The specific implementation method will not be described in detail.
[0145] In addition, Figures 4 to 6 illustrate the voice interaction method provided in this application by taking the user's input of text in the memo interface as an example. The voice interaction method in other interfaces can be implemented by reference and will not be described in detail here.
[0146] In Figures 4 to 6, the input result of the editing command can be displayed in the editing command display area of the memo interface. It is understood that this can also be replaced by displaying the input result of the editing command in the phone's input method interface. For example, the input result corresponding to voice 2 shown in text 2 in interfaces 504 and / or 604 can be replaced by displaying it in the text box of the keyboard input method shown in Figure 7. Key 1 and / or key 2 can be displayed floating on the input method interface, or they can be displayed in other locations within the input method interface; there are no specific limitations.
[0147] It is understandable that, since the accuracy of voice editing command execution depends on speech recognition and semantic analysis, the accuracy of voice editing command execution may be reduced when there are multiple possible ambiguous results in speech recognition and / or semantic analysis. To improve the accuracy of voice editing command execution, this application can eliminate ambiguous results in the recognition and / or semantic analysis of the second speech information through user feedback, as implemented in the following voice interaction scenario B. Here, speech recognition can be based on ASR technology or other technologies, and semantic analysis can be based on LLM or other methods; this application does not specifically limit the specific implementation.
[0148] Voice interaction scenario B:
[0149] After a user inputs a command via second voice information in edit mode, the phone may recognize multiple ambiguous input results corresponding to the edit command based on the voice information; these can also be called ambiguous results. In scenario B, the phone can provide the user with multiple ambiguous results based on the second voice information and determine the edit command and / or the input result corresponding to the edit command based on the user's selection. For example, the phone can determine the user's selected ambiguous result as the input result based on the user's selection of a target ambiguous result among multiple ambiguous results. For example, multiple ambiguous results may include multiple candidate results that are keywords, and the target ambiguous result or target candidate result may be the keyword selected by the user.
[0150] As shown in Figure 8, in one embodiment of scenario B, taking the chat interface 801 between user A and user B as an example, the user inputs first voice information in voice input mode. After the mobile phone recognizes the first voice information, it displays the first input result corresponding to the first voice information in the text box of the chat interface 801, denoted as text 1. Additionally, the chat interface 801 can provide one or more of buttons 1, 2, or 3. Buttons 1, 2, and 3 are described in the context of scenario A and will not be repeated here. Button 1 can serve as the voice input entry point, and button 2 can serve as the editing entry point. If the user needs to edit text 1, they can perform a second selection operation on button 2. The second selection operation is also described in the context of scenario A and will not be elaborated further.
[0151] When the mobile phone detects the user's second selection operation on Button 2 and detects the voice 2 input by the user, it can display multiple ambiguous results corresponding to the voice 2 in the interface 802.
[0152] Specifically, the voice 2 input by the user is the following voice: "Change the character '鱼' (fish) to the character '渔' (fishing) with the three-point water radical". Since the character '渔' is a homophone or near-homophone of multiple characters with the pinyin 'yu' (such as 鱼 (fish), 渝 (Chongqing), 淤 (silt), 浴 (bath)), the mobile phone may not be able to accurately recognize the character corresponding to the voice 2. For example, the text 2 in the interface 802 shows "Change the character '鱼' (fish) to the character '鱼' (fish) with the three-point water radical", that is, the character '渔' is misrecognized as '鱼', and the incorrect text needs to be modified.
[0153] In order to more accurately obtain the user's instruction, as shown in the interface 803, multiple homophones or near-homophones of the character '渔' can be displayed as ambiguous results in the chat interface, as shown in the text 3. The mobile phone can also detect the user's selection operation on the multiple ambiguous results, determine the character '渔' according to the selection operation, and modify the character '鱼' in the text 1 to the character '渔'. The mobile phone can also display the edited input result, that is, the text 4, in the interface 804.
[0154] It can be understood that the above text 2 and / or text 3 can be displayed in the editing instruction display area or in the input method. Optionally, the present application does not limit the relationship between the text 2 and the text 3. For example, the text 2 can be replaced by the text 3, or the text 2 and the text 3 can be displayed simultaneously in the interface 803. Among them, the example of replacing and displaying the text 2 with the text 3 in the interface 803 in FIG. 8 is used for illustration, but not limited thereto.
[0155] In a possible embodiment, the user's selection operation on the multiple ambiguous results may include operations such as the user clicking on the ambiguous results through the display screen, or may include operations such as the user inputting the number or index of the ambiguous results by voice. The present application does not specifically limit this.
[0156] In another embodiment of the scenario B, the mobile phone can also edit the first input result according to the first ambiguous result among the multiple ambiguous results, display the edited input result, and display one or more other ambiguous results.
[0157] As shown in Figure 9, the first input result corresponding to the first voice message input by the user is "Report to Teacher Zhang", as shown by text 1 in interface 901. Additionally, in the editing mode, the user also inputs a second voice message: "Change Zhang to Cha (查)". The input result corresponding to this second voice message is shown by text 2 in interface 902. Semantic recognition based on this text 2 may include multiple ambiguous results. The first possible ambiguous result is: replace "Zhang" in the first input result with "Cha". The second possible ambiguous result is: replace "Zhang" in the first input result with "Cha (姓查的查)". As shown in interface 903, the mobile phone can modify the first input result according to the first ambiguous result and display the modified text "Report to Teacher Cha" in interface 903, denoted as text 3. Additionally, the mobile phone can also display the second ambiguous result in interface 903, as shown by text 4, and can display "Cha (姓查的查)".
[0158] Optionally, if the mobile phone detects a selection operation by the user for text 4 or other ambiguous results, it can replace the first ambiguous result in text 3 with the ambiguous result selected by the user. That is, it can replace "Cha" in text 3 in interface 903 with "Cha (姓查的查)", so that the finally edited text is "Report to Teacher Cha (姓查的查)". If the mobile phone does not detect a selection operation by the user for text 4, it can use text 3 as the edited text, that is, in interface 904, display the edited text 3.
[0159] Among them, text 4 can be displayed in the editing instruction display area or at a position near the first input result. This application does not have specific requirements. Additionally, the modified input result corresponding to this ambiguous result can also be displayed in interface 903, which is not shown in Figure 9; for example, text 4 can be replaced with "Report to Teacher Cha (姓查的查)".
[0160] In Figures 8 and 9 above, the mobile phone can display ambiguous results in the editing instruction display area. Additionally, in another embodiment of scenario B, the mobile phone can display ambiguous results in the input method. For example, as shown in Figure 10, the first input result corresponding to the first voice message input by the user includes "... Qingqing add cream...", as shown by text 1 in interface 1001. In the editing mode, the second voice message input by the user includes: "Change Qingqing to Qingqing (轻轻)", as shown by text 2 in interface 1001. There may be multiple ambiguous results for "Qingqing (轻轻)" in this second voice message, such as Qingqing, Qingqing (青青), Qingqing (清清), and Qingqing (倾情), etc.
[0161] For example, a mobile phone can display multiple ambiguous results in the keyboard area of an input method. As shown in interface 1002 in the figure, the mobile phone can display ambiguous results such as "qing qing, qing qing, qing qing, and qing qing" in the keyboard area of the input method in interface 1002, denoted as text 3. After detecting the user's selection operation on "qing qing" in text 3, the mobile phone can display the modified input result "…… qing qing add cream ……" in interface 1004, as shown in text 5.
[0162] For another example, a mobile phone can display multiple ambiguous results in the display area of a voice input method. As shown in interface 1003 in the figure. The mobile phone can also display the modified input result "…… qing qing add cream ……" in interface 1005, as shown in text 6.
[0163] Table 1 exemplifies several situations that may lead to ambiguous results. In this application, the possible ambiguous results can be presented to the user in the manner shown in FIG. 8 or FIG. 9, and the final editing instruction can be determined based on the user's feedback.
[0164] Table 1
[0165] For example, the situation corresponding to index 1 in Table 1 is that when performing voice recognition, it is impossible to accurately recognize whether the symbol in the user's voice is displayed as a symbol or a character in the input result. Therefore, the ambiguous result 1 is the symbol form ",", and the ambiguous result 2 is the character form "comma". The situation corresponding to index 2 is that when performing voice recognition, it is impossible to accurately recognize homophones or near-homophones. Therefore, the multiple ambiguous results are multiple homophones or near-homophones. The situation corresponding to index 3 is that when performing voice recognition, it is impossible to accurately recognize the keyword "yan" and the description information of the keyword "the yan of the swallow". Therefore, the ambiguous result 1 is the keyword "Xiong Yan", and the ambiguous result 2 is the description information "the yan of Xiong Yanzi". Among them, the description information of the keyword can be referred to the description in voice interaction scenario C, which will not be elaborated here for the time. The situation corresponding to index 4 is that there are multiple mobile phone numbers of the same contact, or there are contacts with the same name. Therefore, the ambiguous results 1 and 2 are different mobile phone numbers respectively.
[0166] In various embodiments of this application, the second voice information input by the user can be voice information related to one or more keywords. Alternatively, the first editing instruction can be an editing instruction related to one or more keywords. For example, the first editing instruction can be used to edit a keyword in the second input result. The keyword can be a single character or a word. Specifically, the first editing instruction can be used to instruct the deletion of a keyword in the second input result, or to instruct the modification of the keyword to a target character. As another example, the first editing instruction can be used to edit the second input result to obtain a third input result, such that the third input result contains the keyword. Specifically, the first editing instruction can be used to instruct the addition of a keyword to the second input result, or to instruct the modification of the original text in the second input result to a keyword.
[0167] Furthermore, to facilitate the differentiation of keywords, the keywords contained in the second input result (i.e., the original text) can be called the original keywords or the text to be modified, and the keywords contained in the third input result (i.e., the edited text) can be called the target keywords or the modified text.
[0168] For example, in the example shown in Figure 8, the word "fish" is the original keyword, and the word "fishing" is the target keyword. Similarly, in the example shown in Figure 9, the word "Zhang" is the original keyword, and the word "search" is the target keyword.
[0169] It is understandable that the ambiguous results corresponding to the second speech information may include those generated during the speech recognition process of the second speech information. For example, the presence of homophones or near-homophones in the second speech information can lead to ambiguity in speech recognition analysis. For instance, the multiple ambiguous results corresponding to index 2 in Table 1 are examples of ambiguities generated during the speech recognition process.
[0170] In addition, ambiguous results may also include those generated during the semantic analysis of the input results corresponding to the second speech information. For example, the multiple ambiguous results corresponding to indices 1, 3, and 4 in Table 1 are examples of ambiguous results generated during the semantic analysis process.
[0171] Based on the embodiments in scenario B, this application can avoid the impact of ambiguous results generated in the speech recognition process and the semantic analysis process on the editing process, thereby improving the accuracy of speech editing.
[0172] Furthermore, to improve the accuracy of voice editing command execution, this application can achieve accurate keyword recognition based on the keyword description information in the second voice information through the implementation method in the following voice interaction scenario C. The keyword description information may include auxiliary information other than pinyin used to determine the keyword. Identifying keywords based on their description information can reduce or avoid the impact of ambiguity caused by homophones or near-homophones on the voice editing process, thereby improving voice editing accuracy.
[0173] Voice interaction scenario C:
[0174] In recognizing the second speech information, this application first converts the natural language in the second speech information into text form using recognition technologies such as ASR, and then performs semantic analysis on the recognition results to obtain the first editing instruction. However, if the second speech information contains easily confused speech information such as homophones or near-homophones, the recognition results of ASR and other speech recognition technologies may be inaccurate, potentially leading to errors in subsequent semantic analysis.
[0175] To improve the accuracy of speech recognition technology when homophones or near-homophones exist in the second speech information, this application can obtain keyword description information based on the second speech information, thereby improving the accuracy of keyword recognition from homophones or near-homophones and thus improving the accuracy of speech editing.
[0176] For example, the second speech information may include the speech corresponding to the descriptive information of keywords. These keywords may include original keywords and / or target keywords. This descriptive information can be used to describe homophones or near-homophones in the second speech information, which are prone to ambiguity, thereby reducing or avoiding ambiguity in speech recognition.
[0177] Taking homophones or near-homophones as an example, descriptive information can be used to describe keywords in speech. Correspondingly, when keywords contain homophones or near-homophones, the mobile phone can recognize the keywords based on the descriptive information and obtain the speech input result. Therefore, the impact of homophones or near-homophones on keyword speech recognition can be reduced or avoided, improving recognition accuracy. It is understandable that the descriptive information of keywords can also serve as a filtering condition for keywords.
[0178] As an example, the description information for keywords may include one or more of the following:
[0179] Method 1: Common character combinations containing keywords. Correspondingly, Method 1 can also be called the common character combination description method. The character combinations can include personal names, place names, nouns, idioms, etc., without specific limitations. For example, in the example shown in FIG. 9, the second voice information can include "the surname Zha's Zha", that is, the character "Zha" can be described by the common phrase (here it can be a common surname) "the surname Zha" of the keyword, which can improve the recognition accuracy of the keyword.
[0180] Method 2: Information on the components of the keyword. Method 2 can also be called the common character combination description method.
[0181] Among them, the information on the components includes pinyin, radicals, the number of strokes or the content of the strokes, etc. The number of strokes of the component can refer to the number of strokes of a part of the keyword. For example, the number of strokes of the single-person radical "亻" is 2. The content of the component strokes can refer to the specific strokes included in a part of the keyword. For example, the content of the strokes of the single-person radical "亻" is a left-falling stroke ("丿") and a vertical stroke ("丨").
[0182] For example, taking the keyword "健" as an example, the second voice information can include "jian is the character 健 with a single-person radical", which can disassemble the keyword and describe the keyword "健" through the radical "single-person radical", which can improve the recognition accuracy of the keyword. In addition, in the disassembling description method, the pinyin of the non-radical part of the keyword can be included in the second voice information to further improve the recognition accuracy. Still taking the keyword "健" as an example, the radical is "亻", and the non-radical part is "建", so the second voice information can include "jian is the single-person radical plus 建 of construction". In addition, the keyword can also be disassembled into the number of strokes and / or the content of the strokes. For example, for the keyword "人", the second voice information can include "the number of strokes is two" and / or "the strokes include a left-falling stroke and a vertical stroke" and other information. In addition, the number of strokes and / or the content of the strokes of the radical part of the keyword, and / or, the number of strokes and / or the content of the strokes of the non-radical part of the keyword can be included in the second voice information.
[0183] Method 3: Semantic description information of the keyword. This Method 3 can also be called the semantic description method. For example, "yu is the 豫 that represents Henan", and the recognition accuracy of the keyword "豫" can be improved by its meaning "the 豫 that represents Henan".
[0184] Method 4: Pronouns, synonyms or similar words of the keyword. In addition, it can also be considered that the pronouns, synonyms, synonyms or similar words of the keyword belong to the semantic description information of the keyword.
[0185] Two or more of the above-mentioned methods 1 to 4 can also be combined for implementation, that is, the description information of the keyword contains multiple types of information such as pinyin, common character combinations, component information, semantic description information, pronouns, synonyms or similar words. For example, "jian is the character '健' with a single-person radical", which explains the character '健' through the common character combination and disassembling description method.
[0186] Next, in combination with FIG. 11, an exemplary method for identifying a keyword with homophones or near-homophones based on the pinyin of the keyword, radical information, and the pinyin of non-radical components will be introduced. In the implementation shown in FIG. 11, the keyword is, for example, "芊", and the second voice information is "芊 with a grass radical and a qian", and the corresponding pinyin is "cao zi tou jia yi ge qian de qian". The ASR module of the mobile phone can obtain the input result corresponding to the second voice information, "签 with a grass radical and a qian", through voice recognition. Further, the mobile phone can determine the description information of the keyword based on the second voice information and / or the input result as the screening condition for the keyword.
[0187] As shown in FIG. 11, for example, the method for determining the screening condition is to perform segmented semantic analysis on the input result to determine the meaning of each segment of the input result.
[0188] Specifically, in the above input result, the first segment is "草字头", and the mobile phone can identify whether this segment is used to indicate a radical. After analysis, "草字头" corresponds to the radical "艹", so the first segment indicates that the radical of the keyword is "艹".
[0189] The second segment is "加一个签", which can mean that the keyword is obtained by adding the character "签" to the basis of "艹". Among them, the mobile phone can determine whether "签" is a radical. After confirmation, the mobile phone can determine that "签" does not belong to a radical. The mobile phone can further determine whether the character "签" is a polyphonic character and determine that the pinyin of the character "签" is "qian". Considering that the character "签" has homophones and near-homophones, the mobile phone can learn from the second segment that the keyword is a character composed of "艹" and the character with the pinyin "qian" or a near-homophone of "qian". For example, a character list corresponding to the character "签" can be obtained, which contains the homophones and near-homophones of the character "签".
[0190] The third segment is "的签", which can mean that the keyword is the character "签". The mobile phone can further determine whether the character "签" is a polyphonic character and determine that the pinyin of the character "签" is "qian". Considering that the character "签" has homophones and near-homophones, the mobile phone can learn from the third segment that the keyword is the character with the pinyin "qian" or a near-homophone of "qian".
[0191] Based on the analysis of the above first and second segments, the screening condition 1 for keywords can be obtained: the radical of the keyword is the grass radical, and the non-radical part is a character with the pinyin "qian" or a near-homophone of "qian". Additionally, based on the analysis of the above third segment, the screening condition 2 for keywords can be obtained: the keyword is a character with the pinyin "qian" or a near-homophone of "qian". Therefore, the mobile phone can further screen keywords according to screening condition 1 and screening condition 2.
[0192] Furthermore, Chinese characters can be screened according to screening condition 1 to obtain a set of one or more Chinese characters that meet screening condition 1, and this set may include keywords. That is, one or more characters can be obtained according to the characters in the character table corresponding to the characters "艹" and "qian", that is, the set of Chinese characters that meet screening condition 1. The radical of the one or more characters is "艹" and the non-radical part is included in the character table corresponding to the character "qian". In addition, the characters that meet screening condition 2 can be screened from the candidate characters of the keyword according to screening condition 2 to obtain one or more characters. The one or more characters meet screening condition 1 and meet screening condition 2. That is, according to the pinyin of the characters in the set of Chinese characters that meet screening condition 1 and the pinyin "qian" of the character "qian", it can be verified whether the pinyin of the characters in the set of Chinese characters that meet screening condition 1 is "qian".
[0193] The mobile phone can further determine whether there are multiple characters that meet screening condition 1 and meet screening condition 2.
[0194] If there are multiple candidate characters that meet screening condition 1 and meet screening condition 2, the mobile phone can further display these multiple candidate characters to the user and obtain the candidate character selected by the user as the keyword. It can also be understood that these multiple candidate characters can be used as multiple ambiguous results of the keyword. Therefore, in the same way as in voice interaction scenario B, the user can be requested to select one of the multiple ambiguous results as the keyword.
[0195] For example, the candidate characters that meet screening condition 1 and meet screening condition 2 include "qian, qian, qian" etc. The mobile phone can display the above candidate characters in the editing instruction display area or the input method area of the display screen and determine the keyword according to the user's selection.
[0196] As an example, referring to the method shown in Figure 8, the mobile phone can display multiple characters that satisfy both filter condition 1 and filter condition 2 on the screen. When the mobile phone detects that the user has selected a character from the candidate characters, it can use the selected character as a keyword and perform an editing action. For example, if the keyword is the original keyword, the mobile phone can modify the original keyword in the second input result to be edited to the target keyword, or delete the user-selected target keyword from the second input result to be edited. Alternatively, if the keyword is the target keyword, the mobile phone can modify the original keyword in the second input result to be edited to the user-selected target keyword, or insert the user-selected target keyword into the second input result to be edited.
[0197] As another example, referring to the method shown in Figure 9, the mobile phone can select a target character from multiple characters that satisfy both filter condition 1 and filter condition 2, edit the second input result based on the target character, and display the edited second input result on the display screen. Simultaneously, the mobile phone can also display other characters that satisfy both filter condition 1 and filter condition 2, and replace the target character in the edited second input result with the character selected by the user to obtain a third input result. If the user does not select any other characters, the edited second input result can be used as the third input result.
[0198] Additionally, if only one candidate character exists, the phone can identify that character as the keyword and edit the second input result based on that keyword to obtain the third input result. The phone can also display the edited third input result.
[0199] It is understandable that if there are no characters that satisfy both filter condition 1 and filter condition 2, that is, if there are no candidate characters with the pinyin "qian", the phone can notify the user that the editing has failed. The methods for notifying the user of editing failure may include one or more of the following: displaying text or graphics indicating editing failure on the screen, outputting an audio signal indicating editing failure through the speaker, or indicating editing failure through vibration.
[0200] Figure 11 above is a flowchart of an exemplary implementation. Adaptive modifications can be made to this flowchart, meaning that the voice interaction method shown in this application is not limited to this flowchart.
[0201] It is understandable that if there are multiple candidate characters in the second input result to be edited, and all of these candidate characters can be described by the same descriptive information in the second speech information, then the user can be prompted that there are multiple candidate characters, and the user can select one of them as the keyword. The method by which the user selects a candidate character can be the same as the method by which the user selects one ambiguous result from multiple ambiguous results, and will not be elaborated further.
[0202] In one possible embodiment, the following voice interaction scenario D provides an additional voice recognition method to improve the accuracy of voice recognition. The voice here may include first voice information and / or second voice information. For ease of description, voice interaction scenario D mainly describes the method of recognizing keywords in the second voice information.
[0203] Voice interaction scenario D:
[0204] Voice interaction scenario D provides a solution for recognizing speech based on user input information. The mobile phone can filter keywords based on the pinyin of keywords in the second voice information input by the user and the homophones or near-homophones of the pinyin in the input text information.
[0205] Specifically, the entered text information can be text information in the text box where the first input result is located or in other interfaces, or it can be the user's historical input information in other interfaces. For example, the second voice information can contain the pinyin of a keyword. If the mobile phone determines from the pinyin that the entered text information contains homophones and / or near-homophones of the keyword, then the keyword can be determined from the homophones and / or near-homophones.
[0206] In one possible implementation, the mobile phone can determine the keyword based on the cursor position. For example, it can determine whether the entered text contains homophones and / or near-homophones of the keyword within a range of n1 characters before and / or n2 characters after the cursor position in the current text box interface, thereby improving the accuracy and efficiency of keyword recognition. Here, n1 and n2 are positive integers. Furthermore, if the keyword is a polyphonic character, this implementation can search for the keyword based on its pinyin within a certain range of the original text, improving the accuracy of finding polyphonic characters.
[0207] As shown in Figure 12, taking n1 = n2 = n as an example, the second voice information input by the user in the history includes "change the tune to the 'diao' of 'hangdiao'", where the pinyin for "tiao" is "diao". Optionally, the maximum value of n can be set to N, where N is a positive integer.
[0208] The ASR module of the mobile phone can recognize the second voice message and obtain the text "Change '调' to '吊' which means hanging", that is, the input result of the second voice message is "Change '调' to '吊' which means hanging". Among them, the character '调' is the original keyword before modification, and '吊' is the target keyword after modification. Further, the mobile phone can determine whether the character '调' is a polyphonic character. If it is a polyphonic character, the mobile phone can determine that the pinyin of the character '调' is 'diao' according to the second voice message input by the user. Further, the mobile phone can obtain the position of the cursor, and based on the cursor position, determine whether there is a character with the pinyin 'diao' within the range of n steps before and n steps after the cursor position in the second voice input result. If there is a candidate character with the pinyin 'diao', it is further determined whether the candidate character is a polyphonic character. If it is not a polyphonic character, the candidate character can be determined as the keyword, and further, an editing operation can be performed according to the keyword.
[0209] If it is determined that the candidate character is a polyphonic character, it is further possible to combine the voice input record and / or pinyin input record of the candidate character to determine whether the pinyin of the candidate character is 'diao'. If the pinyin of the candidate character is 'diao', the candidate character can be determined as the keyword, and further, an editing operation can be performed according to the keyword.
[0210] In addition, if the pinyin of the candidate character is not 'diao', the candidate character is not used as the keyword, and it can be determined that there is no character with the pinyin 'diao' within the range of n steps before and n steps after the cursor position. If there is no character with the pinyin 'diao' within the range of n steps before and n steps after the cursor position, it can be determined whether n is less than N. If it is determined that n is less than N, an increment operation can be performed on n, that is, let n = n + 1, and then determine whether there is a character with the pinyin 'diao' within the range of n + 1 steps before and n + 1 steps after the cursor position in the second voice input result to determine whether there is a candidate character within the range of n + 1 steps before and n + 1 steps after the cursor position. The subsequent operations on the candidate character will not be repeated.
[0211] It can be understood that if n = N, no increment operation is performed on n, but an editing failure is prompted to the user. The way to prompt the user with an editing failure can refer to the way of prompting the user with an editing failure introduced in Voice Interaction Scenario C and will not be elaborated here.
[0212] In another possible implementation, when the keyword is a polyphonic character, the mobile phone can determine the keyword according to the pinyin of the text information input by the user historically and the pinyin of the keyword in the voice information, so as to improve the recognition accuracy of polyphonic characters.
[0213] Among them, the mobile phone can store the voice input records of the user, that is, store the pinyin of the text when the user inputs text by voice. For example, when the user inputs "tiao zheng" by voice, after ASR mode recognition, the text input by the user is "调整", then the mobile phone can store the pinyin corresponding to the two characters of "调整".
[0214] In addition, the mobile phone can also store the pinyin input records of the user through methods such as pinyin input method, for storing the pinyin of the text when the user inputs text by pinyin. For example, when the user inputs "tiao zheng" by pinyin, and the mobile phone determines through pinyin recognition that the text input by the user is "调整", then the mobile phone can store the pinyin corresponding to the two characters of "调整".
[0215] Furthermore, the mobile phone can also determine the keywords of the characters in the phrase according to the common phrases input by the user. For example, if the text information input by the user through handwriting or other means contains the word "调整", the pinyin corresponding to the word "调整" can be obtained according to the pronunciation of the common phrase, which is "tiao zheng".
[0216] Taking Figure 13 as an example below, the method of identifying keywords by pinyin according to the historical input information of the user will be introduced.
[0217] As shown in Figure 13, the second voice information in the user's historical input contains "把调改成上吊的吊", among which the pinyin of the character "调" is "diao". The ASR module of the mobile phone can recognize the second voice information and obtain the text "把调改成上吊的吊". Among them, the character "调" is the original keyword before modification, and "吊" is the target keyword after modification. Further, the mobile phone can determine whether the character "调" is a polysemous word. If it is a polysemous word, the pinyin of the character "调" can be determined as "diao" according to the second voice information input by the user. The mobile phone can also extract the context content, convert the characters in the second input result to be edited into pinyin according to the historical voice input or pinyin input records, and determine whether there is a character with the pinyin "diao" in the pinyin corresponding to the second input result to be edited according to the pinyin "diao" of the character "调" in the second voice information, and use it as the original keyword. Subsequently, the original keyword in the second input result can be modified to the character "吊" to obtain the third voice result. In addition, if there is no character with the pinyin "diao" in the pinyin corresponding to the second input result to be edited, it can prompt the user that the editing fails.
[0218] Based on the process in Figure 13, the mobile phone can determine keywords according to the pinyin of the input information, thereby improving the accuracy of keyword determination in the case of homophones or near-homophones.
[0219] It is understandable that the processes shown in Figure 12 and Figure 13 can be implemented in combination. For example, the mobile phone can first determine the characters with the same pinyin as the keyword from the range of n1 characters before and / or n2 characters after the current cursor position in the second input result to be edited. If there are no characters with the same pinyin as the keyword within the above range, then determine the characters with the same pinyin as the keyword from the range of characters outside the range of n1 characters before and / or n2 characters after the cursor position in the second input result to be edited.
[0220] Another possible implementation is as follows: If the second voice information contains the pinyin of the keyword "diao", and the second input result to be edited contains the word "adjust", the mobile phone can know from the pinyin of the historical input text information that the pinyin of "diao" in the word "adjust" is "tiao". That is, the pinyin of the word "diao" in the second voice information is different from that of the word "diao" in the second voice information. Therefore, the word "diao" in "adjust" can be excluded as a keyword, that is, the possibility of "diao" in "adjust" being a keyword is not high.
[0221] In one possible implementation, because users often omit non-textual content such as symbols during spoken expression, it becomes impossible to accurately match the original text to be edited when editing based on speech. The following section, using voice interaction scenario E as an example, introduces a method to generalize the matching of the original text based on the user's expression, reducing the impact of non-textual content such as symbols on editing.
[0222] In one possible embodiment, the mobile phone can determine whether a Chinese character is a polyphonic character through its input method. The input method can have a built-in polyphonic character table, containing polyphonic characters and their pinyin. Alternatively, the functions of the electronic device involved in determining polyphonic characters shown in Figures 11 to 13 can be implemented through the input method.
[0223] Voice interaction scenario E:
[0224] As a generalized way of matching the original text content, the mobile phone can delete the symbols contained in the input result corresponding to the user's second voice input, delete the symbols in the second input result to be edited, and match the input result after the deletion of symbols corresponding to the second voice information from the second input result after the deletion of symbols.
[0225] Specifically, deleting symbols from the second input result to be edited can mean deleting symbols used to separate text information within sentences, while retaining symbols used to separate two different sentences. Symbols used to separate text information within sentences can be called first-class symbols, and symbols used to separate two different sentences can be called second-class symbols. For example, first-class symbols may include commas, quotation marks, parentheses, and book titles. Second-class symbols may include periods, question marks, or exclamation marks.
[0226] Taking the second input result to be edited as "I recently bought some apples, bananas, and strawberries" as an example, the user's second voice input could be "Delete apples, bananas, and strawberries." Correspondingly, the editing instruction recognized by the phone's ASR module is: delete "apples, bananas, and strawberries," meaning the content to be deleted is "apples, bananas, and strawberries." It is evident that if the second input result is edited according to this instruction, it may not accurately match the original text in the second input result, leading to editing failure. Figure 14 illustrates a voice interaction method provided in this application to overcome this problem.
[0227] The determination of content to be deleted can be achieved through a first model, which can be used to output content to be deleted, added, or modified based on editing instructions. Optionally, the first model can be deployed on a server, such as a cloud server or a data network server, or it can be deployed locally on the mobile phone; there are no specific limitations.
[0228] In Figure 14, the mobile phone can obtain the input result "delete apple, banana, strawberry" corresponding to the second voice information mentioned above, and obtain the content to be deleted, "apple, banana, strawberry". The mobile phone can input the second voice information into the first model and obtain the content to be deleted from the first model. Further, the mobile phone obtains the second input result to be edited from the current text box. Optionally, in the example shown in Figure 14, the mobile phone can, based on the current cursor position, pre-select n1 characters before the cursor and / or n2 characters after the cursor in the text box, and obtain second-type symbols such as periods, question marks, or exclamation marks used to separate complete sentences based on the extracted characters, using the obtained sentence as the second input result to be edited. Alternatively, the second input result can also be considered as the complete sentence where the cursor is currently located, or all text information in the current text box.
[0229] After obtaining the second input result, the phone can also delete symbols from the second input result to be edited, resulting in the second input result after deleting the symbols: "I recently bought some apples, bananas, and strawberries". Specifically, the deleted symbols can be of the first category.
[0230] Furthermore, the phone iterates through each sentence in the second input result after the deletion symbol, determining whether the sentence contains the content to be deleted, "apple, banana, strawberry". If the sentence contains this content, then the content from "apple" to "strawberry" can be deleted from the sentence before deleting the first type of symbol (called the original sentence) to obtain the third input result. For example, if the second input result is "I recently bought some apples, bananas, and strawberries.", the second input result after deleting the symbol is "I recently bought some apples, bananas, and strawberries.", and after deleting the content from "apple" to "strawberry", the third input result is: "I recently bought some."
[0231] As another implementation of voice interaction scenario E, the second input result to be edited is "I recently bought some apples, bananas, and strawberries." The second voice message can also be "Delete the content from apples to strawberries," meaning the second voice message contains the phrase "A to B" (or "from A to B," "A to B," etc., without specific limitations), but does not need to contain the complete content to be deleted. Here, A and B can be discontinuous text in the second input result to be edited; for example, A is "apples," and B is "strawberries." In this case, the mobile phone can query the second input result to be edited, specifically the "apples" and "strawberries" to be deleted, and delete "apples" and "strawberries," as well as the text information between "apples" and "strawberries."
[0232] As shown in Figure 15, the mobile phone can determine the content to be deleted as "from apple to strawberry" based on the second voice information. The mobile phone can input the aforementioned second voice information into the first model and obtain the content to be deleted from the first model. Further, the mobile phone obtains the second input result to be edited from the current text box. Optionally, in the example shown in Figure 15, the mobile phone can, based on the current cursor position, pre-select n1 characters before the cursor and / or n2 characters after the cursor in the text box, and obtain second-type symbols such as periods, question marks, or exclamation marks used to separate complete sentences based on the extracted characters, using the obtained sentence as the second input result to be edited. Alternatively, the second input result can also be considered as the complete sentence where the cursor is currently located, or all text information within the current text box.
[0233] After receiving the second input result, the phone can further determine whether it contains content that begins with "apple" and ends with "strawberry," and delete the content that begins with "apple" and ends with "strawberry." That is, delete "apple" and "strawberry" from the second input result, as well as the text information between "apple" and "strawberry," resulting in the third input result: "I recently bought some." Alternatively, if the second input result does not contain content that begins with "apple" and ends with "strawberry," the phone can notify the user that editing failed.
[0234] In one possible embodiment, during spoken expression, the user may omit the description of the original keywords in the second input result to be edited; for example, the original keywords may be omitted from the input result corresponding to the second voice information. The following describes the editing method when keywords are omitted from the input result corresponding to the second voice information, using voice interaction scenario F as an example.
[0235] Voice interaction scenario F:
[0236] In one possible implementation, if the original keyword is omitted in the corresponding input result, and if the second voice information contains descriptive information about the original keyword, the original keyword can be determined based on the descriptive information of the original keyword in the second voice information. For example, referring to the description in voice interaction scenario C, the descriptive information of the original keyword may include common character combinations, component information, and / or semantic description information of the original keyword.
[0237] As an example of determining keywords based on semantic description information, if the original keyword in the input result corresponding to the second voice information is a pronoun containing the original keyword, the original keyword can be queried from the second input result based on the pronoun, so as to support editing methods such as modifying or deleting the original keyword.
[0238] For example, the second input result to be edited is "Navigate to City A", and the second voice input result is "Change the destination to City B". Through semantic analysis, "destination" in the second voice input result can refer to "City A" in the second input result; that is, "destination" is a pronoun for "City A". Therefore, the phone can determine the word "City A" representing the destination from the second input result, and thus change "City A" in the second input result to "City B".
[0239] As another example of determining keywords based on semantic information, if it is necessary to modify the original keyword in the second input result to a target keyword, and the second voice information contains the target keyword but not the original keyword, the target keyword can be used as a synonym or similar word of the original keyword to determine the original keyword in the second input result. That is, if the second voice information contains the target keyword but the original keyword is missing, the mobile phone can determine the original keyword based on the recognition result of the target keyword. Generally, the original keyword and the target keyword belong to the same category of characters or words. The mobile phone can determine the type and / or meaning of the original keyword based on the type of the target keyword, and further determine the original keyword from the second input result to be edited based on the type and / or meaning.
[0240] For example, the second input result to be edited is "Navigate to City A," and the second voice input result is "Change to City B." Semantic analysis reveals that "City B" in the second voice input result is the target keyword. The phone can further analyze the value of "City B," which is a place name, and can further query whether the second input result contains similar words to "City B." After the query, the phone can determine that "City A" is similar to "City B," thus confirming that the original keyword is "City A." Therefore, the phone can change "City A" to "City B."
[0241] It is understood that this application can support determining editing instructions based on the sentence structure of various second voice information to perform editing of the original keyword. These various second voice information may include voice information containing the original keyword and voice information not containing the original keyword. Taking "Navigate to City A" as an example, as shown in Figure 16, this application supports at least one or more of the following three editing instructions to perform an editing operation that changes "City A" to "City B" in the second input result:
[0242] (1) Editing command t1: Change city A to city B. In other words, the second voice information includes "change city A to city B", that is, the input result of the second voice information does not contain the original keyword "city A".
[0243] (2) Editing instruction t2: Change to City B. In other words, the second voice information includes "change to City B", that is, the input result of the second voice information contains the original keyword "City A".
[0244] (3) Editing instruction t3: Change the destination to City B. In other words, the second voice information includes "change the destination to City B", that is, the input result of the second voice information contains the original keyword "City A".
[0245] In Figure 16, the second input result can be represented as the original text s, and the first editing instruction can be represented as the editing instruction t. For example, the editing instruction t can be editing instruction t1, t2, or t3. The editing instruction t can be obtained by speech recognition of the second voice information emitted by the user. As shown in Figure 16, the mobile phone can extract the modification object a and the target text b from the editing instruction t. For example, for editing instruction t1, the modification object a is "City A", and the target text b is "City B"; for editing instruction t2, there is no modification object a, and the target text b is "City B"; for editing instruction t3, the modification object a is "destination", and the target text b is "City B".
[0246] The mobile phone further determines whether the editing instruction t contains the object to be modified, a. If the editing instruction t does not contain the object to be modified, a, then a word c that is a synonym of the target text b can be found in the original text s. For example, for editing instruction t2, since this editing instruction t2 does not contain the object to be modified, but indicates that the target text b, "City B", is included, the mobile phone can find the word "City A" that is a synonym of "City B" in the original text s, "Navigate to City A". The mobile phone can then modify "City A" in the original text s, "Navigate to City A", to "City B". The mobile phone can search for the word c using word embedding models or knowledge bases, and this application does not specifically limit the method for determining the word c.
[0247] If the editing instruction t contains the object to be modified, the phone can further determine whether the original text s contains the object to be modified.
[0248] If the original text s contains the object to be modified 'a', the mobile phone can modify the object 'a' in the original text s to the target text 'b'. For example, for the editing instruction t1, if the original text s "Navigate to City A" contains the object to be modified 'a' "City A", the mobile phone can modify the object 'a' "City A" in the original text s "Navigate to City A" to the target text 'b' "City B".
[0249] Furthermore, if the original text s does not contain the object to be modified 'a', the mobile phone can further search for the existence of a word 'd' that the object to be modified 'a' refers to in the original text s. If the word 'd' that the object to be modified 'a' refers to exists, then the word 'd' that the object to be modified 'a' refers to can be modified to the target text 'b'. Taking the editing instruction t3 as an example, "destination" can refer to "city A" in "navigate to city A", so the mobile phone can modify "city A" in the original text s "navigate to city A" to the target text 'city B'. The mobile phone can search for the word 'd' that 'a' refers to through word embedding models, knowledge bases, or specially trained referential relationship models. This application does not specifically limit the method for determining similar words.
[0250] Optionally, in the embodiment shown in Figure 16, the modified object 'a' can be used as the original keyword, and the implementation of the modified object 'a' can be in accordance with the description of the original keyword in this application. For example, the second voice information can include descriptive information of the modified object 'a'. Furthermore, the modified object 'a' can also be understood as descriptive information of the original keyword. Additionally, the target text 'b' can be used as the target keyword, and the description of the original keyword in this application can be applied. For example, the second voice information can include descriptive information of the target text 'b'.
[0251] Based on the editing method shown in Figure 16, editing instructions can be efficiently identified according to their sentence structure.
[0252] Figure 17 illustrates another way to determine the original keyword when it is missing in the input result corresponding to the second voice information. As shown in Figure 17, the mobile phone can edit the original keyword through the editing intent classification model, the word segmentation model, and the editing parameter classification model.
[0253] The editing intent classification model can be used to classify editing intents. Types of editing intents include, for example, deletion, addition, or modification. The editing intent classification model can also be used to identify whether the input result corresponding to the second voice information is an editing instruction. This model can also have other names, which are not specifically limited in this application. For example, a mobile phone can use the input result corresponding to the second voice information as input data to the editing intent classification model to obtain the output data of the model. This output data can be used to indicate the editing intent classification result corresponding to the input data. The editing intent classification result can indicate whether the input data is an editing instruction. Furthermore, if the input data is an editing instruction, the editing intent classification result can contain classification information corresponding to the input data, and this information is used to indicate types such as deletion, addition, or modification.
[0254] For example, if the output data of the editing intent classification model indicates that the input result corresponding to the second voice information is not an editing instruction, the mobile phone can prompt the user that the editing has failed, and / or display the input result corresponding to the second voice information as input text on the display screen. The method of prompting the user that the editing has failed can refer to the method of prompting the user that the editing has failed described in voice interaction scenario C, and will not be repeated here.
[0255] For example, if the output data of the editing intent classification model indicates that the input result corresponding to the second voice information is an editing instruction, the mobile phone can further use the second input result to be edited and the input result corresponding to the second voice information as input data of the word segmentation model respectively to obtain output data. The output data includes the word segmentation result of the second input result and the word segmentation result of the input result corresponding to the second voice information.
[0256] Word segmentation models can be used to segment sentences. The input data for a word segmentation model can be a sentence, and the output data can be the word segmentation result of that sentence. For example, if the second input result to be edited, i.e., the input data of the word segmentation model, is "navigate to City A," then the output result of the word segmentation model can be the word segmentation result of the sentence. The word segmentation result can be used to indicate that "navigate to City A" contains the following words: "navigate," "to," and "City A." Similarly, if the input result corresponding to the second voice information is "destination changed to City B," its corresponding word segmentation result contains the following words: "destination," "changed to," and "City B."
[0257] Optionally, the word segmentation model can be a general word segmentation model, a language model, or a specially trained edit parameter classification model.
[0258] Furthermore, the mobile phone can input the word segmentation results of the second input result and the word segmentation results of the input result corresponding to the second voice information into the editing parameter classification model respectively, and obtain the classification result.
[0259] The edit parameter classification model can be used to obtain edit parameter classification results based on word segmentation results. These results can include classification information corresponding to the word segments. This classification information can be used to determine the attributes of the word segments within editing instructions, or in other words, to define the relationship between word segmentation and editing instructions. Classification information can also be referred to as labels. For example, classification information includes at least one of the following:
[0260] (a) Background (BACKGROUND) is used to represent non-keywords in the input results corresponding to the second input results and / or the second speech information. That is, it is not necessary to edit the words with the classification information BACKGROUND, that is, the corresponding word segmentation attribute is non-keyword.
[0261] (b) REPLACE_ORIGINA can be used to represent the original keyword to be modified when the editing mode is modified, that is, the corresponding word segmentation attribute is the original keyword in the editing mode.
[0262] (c) The modified replacement (REPLACE_MODIFIED) can be used to represent the target keyword after modification when the editing mode is modified, that is, the corresponding word segmentation attribute is the target keyword in the modification editing mode.
[0263] (d) Delete (DELETE), which can be used to represent the original keyword to be deleted when the editing mode is delete.
[0264] (e) Add_input(ADD_INPUT) can be used to represent the target keyword to be added in the second input result.
[0265] (f) Add Reference Before (ADD_REFERENCE_BEFORE) can be used to indicate the reference position in the second input result. The corresponding editing operation includes adding the target keyword before the position in that window.
[0266] (g) Adding a reference after (ADD_REFERENCE_AFTER) can be used to indicate the reference position in the second input result, and the corresponding editing operation includes adding the target keyword after that position.
[0267] For example, in the word segmentation corresponding to the second input result "navigate to City A", the classification information for "navigate" and "to" can be BACKGROUND, and the classification information for "City A" can be REPLACE_ORIGINA. In the word segmentation corresponding to the second voice information "change destination to City B", the classification information for "destination" and "change to" is BACKGROUND, and the classification information for "City B" can be REPLACE_MODIFIED.
[0268] Optionally, the edit parameter classification model can be a specially trained model used to perform word segmentation classification. Furthermore, the edit parameter classification model can also include word segmentation functionality; that is, the edit parameter classification model can contain a word segmentation model, or in other words, the word segmentation model and the edit parameter classification model can be the same model.
[0269] Furthermore, the mobile phone can extract editing parameters based on the category information corresponding to the word segmentation, and then perform editing operations.
[0270] Extracting editing parameters may include determining the original keywords and / or target keywords required for the editing method based on classification information. For example, for the "modify" editing method, the original keyword could be a segment of the category information "REPLACE_ORIGINA" in the second input result, and the target keyword could be a segment of the category information "REPLACE_MODIFIED". Similarly, for the "delete" editing method, the original keyword to be deleted could be a segment of the category information "DELETE" in the second input result. And for the "add" editing method, the target keyword to be added in the second input result could be a segment of the category information "ADD_INPUT" in the input result corresponding to the second appointment information.
[0271] The phone can also perform editing operations based on editing parameters. For example, the phone can replace the segment “A city” with the category information REPLACE_ORIGINA in the second input result with the segment “B city” with the category information REPLACE_MODIFIED. Furthermore, the phone can delete the segment with the category information DELETE in the second input result, or add the segment with the category information ADD in the second input result.
[0272] Based on the editing method shown in Figure 17, editing commands can be efficiently recognized, and the recognition robustness is good.
[0273] Figure 18 illustrates another method for determining the original keyword when it is missing from the input result corresponding to the second voice information. As shown in Figure 18, the mobile phone can edit the original keyword through an editing intent classification model and an LLM model. The LLM model can be trained using training data, which at least includes the editing instruction, the original text, and editing parameters when the original keyword is missing. Therefore, the LLM model supports outputting the editing parameters corresponding to the editing instruction based on the missing original keyword.
[0274] The editing intent classification model can be referred to in the description of the model in Figure 17, and will not be repeated here. For example, if the output data of the editing intent classification model indicates that the input result corresponding to the second voice information is not an editing instruction, the mobile phone can prompt the user that editing has failed, and / or display the input result corresponding to the second voice information as input text on the display screen. The method of prompting the user that editing has failed can refer to the method of prompting the user that editing has failed described in voice interaction scenario C, and will not be elaborated here.
[0275] If the output data of the editing intent classification model indicates that the input result corresponding to the second voice information is an editing instruction, the mobile phone can further use the second input result to be edited and the input result corresponding to the second voice information as the input data of the LLM model to obtain the output data of the LLM model. The output data may include the editing parameters obtained based on the second input result and the input result corresponding to the second voice information.
[0276] For example, when modifying the editing method, the editing parameters can include the original keyword to be modified in the second input result and the target keyword in the input result corresponding to the second voice information. For instance, if the second input result is "navigate to city A" and the input result corresponding to the second voice information is "change destination to city B", the output data of the LLM model can indicate that "city A" is the original keyword and "city B" is the target keyword.
[0277] For the deletion editing method, the editing parameters can include the original keyword to be transmitted in the second input result.
[0278] For adding editing methods, the editing parameters can include the target keywords in the input results corresponding to the second voice information.
[0279] Optionally, the output data of the LLM model may include the correspondence between word segmentation and word segmentation classification information, used to indicate which editing parameters correspond to the layer through word segmentation information. The word segmentation information can be referred to in Figure 17. Alternatively, the output data of the LLM model may also indicate editing parameters in other ways, without specific limitations. In Figure 18, the way the mobile phone performs editing operations according to the editing parameters can be referred to the explanation of editing operations in Figure 17, and will not be repeated here.
[0280] Based on the editing method shown in Figure 18, editing instructions can be efficiently recognized using a large-scale language model.
[0281] It should be noted that this application uses the Chinese language scenario as an example only and does not constitute a limitation on the language scenario. In other language scenarios, such as English, German, and French, a similar approach can be adopted, which will not be elaborated here.
[0282] Figure 19 is a flowchart illustrating a voice interaction method provided in this application. This task labeling method can be applied to any of the scenarios shown in Figures 4 to 18. As shown in Figure 19, the process includes:
[0283] S101: The electronic device displays the voice input entry on the display screen.
[0284] The voice input entry can be displayed on the screen of an electronic device, either on any interface of the device's display. For example, as shown in Figure 4, the entry can be displayed in the memo interface 401 of the electronic device. As shown in Figure 8, the entry can also be displayed in the chat interface of the electronic device.
[0285] The voice input entry point can be a voice input button. For example, the voice input entry point can be button 1 as shown in interface 401 in Figure 4, button 1 as shown in interface 801 in Figure 8, or button 1 as shown in interface 1001 in Figure 10, etc.
[0286] S102: The electronic device detects the user's first selection operation for the voice input entry, displays the editing entry, and detects the first voice information input by the user to obtain the first input result.
[0287] The first input result can be the text information that the user inputs via voice this time.
[0288] The first action to choose from can be a user's click or long press on the voice input field.
[0289] As shown in Figure 4, when the electronic device detects the user's selection operation on button 1 in interface 401, it can display button 2 in interface 402. This button 2 is an example of an editing entry point. Alternatively, the editing entry point can also be button 2 in interfaces 502 to 505 as shown in Figure 5, button 2 in interfaces 602 to 605 as shown in Figure 6, button 2 in interfaces 801 to 804 as shown in Figure 8, button 2 in interfaces 901 to 904 as shown in Figure 9, or button 2 in interfaces 1001 to 1002 as shown in Figure 10.
[0290] In addition, the first input result is shown as text 1 in interface 402 in Figure 4.
[0291] S103: After detecting the user's second selection operation on the editing entry, the electronic device detects the second voice information input by the user and obtains the first editing instruction.
[0292] The second voice information can be the voice information of a voice editing command input by the user. The first editing command is the editing command obtained based on the second voice information, and can also be called a voice editing command.
[0293] The second option is for the user to click or long-press the edit entry.
[0294] S104: The electronic device edits the second input result according to the first editing instruction to obtain the third input result, wherein the second input result includes the first input result and / or the historical input result.
[0295] In other words, this application supports editing the current input result or the historical input result based on the second voice information. This application does not limit the input method of the historical input result; for example, the historical input result can be input via voice and / or keyboard.
[0296] It is understood that the electronic device edits the second input result according to the first editing instruction to obtain the third input result, as illustrated in any of the examples in Figures 11 to 19.
[0297] S105: The electronic device displays the third input result on the display screen.
[0298] For example, the third input result may include text 3 in Figure 5, text 3 in Figure 6, text 4 in Figure 8, text 3 in Figure 9, text 5 or text 6 in Figure 10.
[0299] In addition, electronic devices can send or output third-party input results. For example, in a multi-user chat scenario, an electronic device can send third-party input results to a server or other devices. Furthermore, in local applications such as memos, electronic devices can store third-party input results.
[0300] Based on the method shown in Figure 19, efficient speech editing can be achieved.
[0301] In one possible embodiment, the first editing instruction includes at least one of the editing methods such as deletion, addition, and modification. Specifically, in the deletion editing method, the electronic device can delete the original keyword from the second input result to obtain a third input result. In the addition editing method, the electronic device can add the target keyword to the second input result to obtain a third input result. In the modification editing method, the electronic device can modify the original keyword in the second input result to the target keyword to obtain a third input result.
[0302] In one possible embodiment, the second voice information includes information about the original keywords in the second input result and / or information about the target keywords in the third input result.
[0303] The information of the original keyword includes at least one of the following: the pinyin of the original keyword, common character combinations of the original keyword, information about the components of the original keyword, semantic description of the original keyword, or its referent, synonym, near-synonym, or similar word. For more information on the original keyword, please refer to the introduction of voice interaction scenario C.
[0304] That is, the second voice information input by the user may include at least one of the following: the pinyin of the original keyword, common character combinations of the original keyword, information on the components of the original keyword, semantic description information of the original keyword, or pronouns, synonyms, near-synonyms, or similar words of the original keyword, in order to support the electronic device to achieve accurate and efficient recognition of the original keyword based on the voice information of the original keyword.
[0305] Similarly, the target keyword information includes at least one of the following: the pinyin of the target keyword, common character combinations of the target keyword, information about the components of the target keyword, semantic description information of the target keyword, or a pronoun, synonym, near-synonym, or similar word of the target keyword. That is, the second voice information input by the user can include at least one of the following: the pinyin of the target keyword, common character combinations of the target keyword, information about the components of the target keyword, semantic description information of the target keyword, or a pronoun, near-synonym, near-synonym, or similar word of the target keyword, to support the electronic device in accurately and efficiently recognizing the target keyword based on the voice information of the target keyword.
[0306] In a possible embodiment, as an implementation manner of determining the original keyword according to the pronoun of the original keyword, the electronic device may determine the original keyword from the second input result according to the pronoun of the original keyword. Based on this embodiment, accurate and efficient determination of the keyword can be achieved when the original keyword's pinyin is not included in the second voice message. As shown in FIG. 16 for example, "destination" in the editing instruction t3 is the pronoun of the original keyword "Beijing". In the case where the original keyword "Beijing" is not included in the editing instruction t3, the electronic device may determine the original keyword "Beijing" according to the pronoun "destination".
[0307] In a possible embodiment, when the second voice message does not include information about the original keyword, the electronic device may determine the target keyword according to the information about the target keyword in the second voice message, and determine the original keyword from the second input result according to the type and / or semantics of the target keyword. For example, the electronic device may determine a synonym, a near synonym, or a word of the same category of the target keyword from the second input result, and the query result may be used as the original keyword. It can also be understood that the target keyword may be used as a synonym, a near synonym, or a word of the same category of the original keyword. At this time, the second voice message includes a synonym, a near synonym, or a word of the same category of the original keyword. Based on this embodiment, accurate and efficient determination of the keyword can be achieved according to the target keyword when the second voice message does not include information about the original keyword. As shown in FIG. 16 for example, "City B" in the editing instruction t2 is a word of the same category as the original keyword "City A". In the case where the original keyword "City A" is not included in the editing instruction t3, the electronic device may determine the original keyword "City A" according to "City B".
[0308] In a possible embodiment, the electronic device may determine multiple candidate results of the keyword, where the keyword here may be the original keyword and / or the target keyword. The electronic device may also display the multiple candidate results on the display screen, and detect a third selection operation of the user on the target candidate result among the multiple candidate results, and determine the keyword according to the target candidate result. Therefore, accurate determination of the keyword can be achieved when there are multiple candidate words for the keyword. As shown in FIG. 8 for example, the text 3 in the interface 803 includes homophones and near homophones of the keyword "fishing". According to the user's selection, the electronic device may determine the character "fishing" as the keyword.
[0309] In a possible embodiment, the second voice information may include the pinyin of the keyword and the information of the components of the keyword, and the information of the components of the keyword includes radical information. The electronic device may determine the radical of the keyword according to the radical information of the keyword, and determine multiple candidate results of the keyword according to the radical of the keyword and the pinyin of the keyword. Based on this implementation manner, it is possible to accurately determine the candidate characters of the keyword when there are homophonic characters in the keyword. Further, the electronic device may further improve the accuracy of the determined candidate characters according to the pinyin of the non-radical part. As shown in FIG. 13 for example, the radical information may be used to indicate "艹", and the pinyin of the non-radical part and the keyword is both "qian", and candidate characters such as "芊", "芡", and "蒨" can be determined.
[0310] In a possible embodiment, the original keyword in the second voice information includes discontinuous text in the second input result. The electronic device may edit the discontinuous text and the text information located between the discontinuous texts according to the first editing instruction to obtain a third input result. Therefore, the discontinuous content in the original text can be edited according to the second voice information.
[0311] As an implementation manner of this embodiment, the text information between the above two discontinuous texts is a first type of symbol (or referred to as a first symbol). The second voice information may include information for indicating the editing of the discontinuous text. This first type of symbol is used to divide the text information within a sentence. Based on this implementation manner, when the editing instruction in the user's voice input omits the first type of symbol in the original text, the original text to be edited can be reasonably matched, improving the editing accuracy and efficiency. As shown in FIG. 14 for example, the second input result is "I recently bought some apples, bananas, and strawberries.", and the input result corresponding to the second voice information is for example "Delete apples bananas strawberries". Correspondingly, the electronic device may delete "apples, bananas, and strawberries" in the second input result, where "," is an example of the first type of symbol.
[0312] As another implementation manner of this embodiment, the second voice information may include information for editing the discontinuous text and the text information located between the discontinuous texts. For example, the input result corresponding to the second voice information contains the sentence pattern "from A to B", that is, the first editing instruction obtained according to the second voice information is used to indicate the editing of the content from A to B in the second input result. Correspondingly, the electronic device may search for "A" and "B" in the second input result, and edit "A", "B" and the content located between "A" and "B" in the original text.
[0313] In one possible embodiment, the information of the original keyword in the second voice information includes the pinyin of the original keyword. The electronic device can determine the keyword from the second input result based on the pinyin of the keyword and the cursor position. Based on this embodiment, the accuracy of determining the original keyword can be improved. For example, if the second input result to be edited contains multiple identical keywords and only a portion of them need to be edited, the keyword to be edited can be accurately determined based on the cursor position. For instance, if the second input result contains at least two identical keywords, the electronic device can determine to edit the keyword closest to the cursor position.
[0314] In one possible embodiment, if the information of the original keyword in the second voice information includes the pinyin of the original keyword, the electronic device can determine the keyword within n1 characters before the cursor position in the second input result based on the pinyin of the keyword, where n1 is a positive integer; and / or, determine the keyword within n2 characters after the cursor position in the second input result based on the pinyin of the keyword, where n2 is a positive integer. This embodiment can improve the accuracy of keyword recognition. For example, in a scenario where the keyword is a polyphonic character, the second input result to be edited may contain multiple keywords with different pronunciations. In this case, the keyword to be edited can be selected from the multiple keywords with different pronunciations based on the cursor position.
[0315] Figure 12 illustrates an example of how keywords can be determined when n1 = n2 = n. In Figure 12, the electronic device can determine whether the keyword "diao" is present within the range of n characters before and n characters after the cursor position. Alternatively, the electronic device can determine the second input result based on the Chinese characters within the n1 characters before and / or the n2 characters after the cursor position.
[0316] In one possible embodiment, the information of the original keyword includes the pinyin of the original keyword. The electronic device can also determine the keyword based on the pinyin of the keyword and the voice input record and / or pinyin input record of the second input result. This embodiment can improve the accuracy of keyword recognition. For example, in a scenario where the keyword is a polyphonic character, the second input result to be edited may contain a polyphonic character as a keyword. In this case, the electronic device can determine the pinyin of the keyword input based on the voice input record and / or pinyin input record of the keyword in the second input result, and determine whether the keyword is the keyword to be edited based on the pinyin of the keyword input and the pinyin of the keyword in the second voice information. For example, it can be required that the keyword be edited when the pinyin of the keyword input is the same as the pinyin of the keyword in the second voice information; and if the pinyin of the keyword input is different from the pinyin of the keyword in the second voice information, the editing of the keyword in the second input result is ignored. Figure 13 shows an exemplary implementation of determining the keyword based on the pinyin of the keyword and the voice input record and / or pinyin input record of the second input result.
[0317] In one possible embodiment of S103, after detecting the second voice information, the electronic device can identify the second voice information and obtain an input result corresponding to the second voice information; perform semantic analysis on the input result corresponding to the second voice information to obtain the first editing instruction.
[0318] In this application, semantic analysis can be performed based on the sentence structure of the input result corresponding to the second voice information to obtain the first editing instruction. For example, as shown in Figure 16, the input result corresponding to the second voice information is, for example, editing instruction t1, editing instruction t2, or editing instruction t3. As described in the example of Figure 16, editing instructions t1, t2, and t3 have different sentence structures. For example, editing instruction t2 does not contain the object to be modified, editing instruction t3 contains the object to be modified but the object to be modified is not included in the original text s, and editing instruction t1 contains the object to be modified and the object to be modified is included in the original text s. The electronic device can analyze the meaning of the editing instruction based on the sentence structure of different editing instructions to perform editing.
[0319] In addition, this application can identify the semantics of editing instructions through a model.
[0320] For example, one way to perform semantic analysis on the input result corresponding to the second voice information is to: determine one or more words in the input result corresponding to the second voice information through a word segmentation model; determine the classification information corresponding to the word through an edit parameter classification model, wherein the classification information of any word is used to indicate the attribute of the word in the first editing instruction; and determine the first editing instruction based on the word and the classification information. For example, as shown in Figure 17, an electronic device can perform semantic analysis through an edit intent classification model, a word segmentation model, and an edit parameter classification model to obtain the first editing instruction.
[0321] For example, another way to perform semantic analysis on the input result corresponding to the second voice information is to determine the attributes of one or more word segments in the input result corresponding to the second voice information in the first editing instruction using a large-scale language model; and to determine the first editing instruction based on the word segments and the attributes. For example, as shown in Figure 18, an electronic device can perform semantic analysis using an LLM model to obtain the first editing instruction.
[0322] In one possible embodiment, the input result corresponding to the second voice information is displayed on the display screen. This allows for the visualization of voice input editing instructions.
[0323] Specifically, after the electronic device detects the user's second selection operation on the editing entry, it displays an editing instruction display area on the screen and shows the input result corresponding to the second voice information in the editing instruction display area. This can be understood as the editing instruction display area being used to display editing instructions in editing mode. Optionally, the editing instruction display area can also be used to display candidate results for keywords. For example, the editing instruction display area can be the display area where text 2 is located in Figures 5, 6, 8, or 9.
[0324] Additionally, the electronic device can also display the input result corresponding to the second voice information in the input method display area on the display screen. This input method display area can also display one or more of the following: a voice input entry, an editing entry, or a keyboard entry. For example, as shown in Figure 7, the input result corresponding to the second voice information can be displayed in the text display area of the keyboard input method. As shown in interface 1001 of Figure 10, text 2 can be displayed in the voice input method area; this text 2 is the input result corresponding to the second voice information.
[0325] In one possible embodiment, the electronic device can determine whether a Chinese character is a polyphonic character through an input method. The input method can have a built-in polyphonic character list, containing polyphonic characters and their pinyin. Alternatively, the functions implemented by the electronic device in this application can also be implemented through an input method.
[0326] Figure 20 is a schematic diagram of the structure of an electronic device 2000 provided in an embodiment of this application. The electronic device 2000 can be the electronic device described above (e.g., a mobile phone). As shown in Figure 20, the electronic device 2000 may include: one or more processors 2001; one or more memories 2002; a communication interface 2003; and one or more computer programs 2004. These devices can be connected via one or more communication buses 2005. The one or more computer programs 2004 are stored in the memory 2002 and configured to be executed by the one or more processors 2001. The one or more computer programs 2004 include instructions. For example, when the electronic device 2000 is the electronic device described above (e.g., a mobile phone), the instructions can be used to perform relevant steps of the electronic device as described in the corresponding embodiments above, such as performing the relevant steps of the electronic device in Figures 3 to 12. The communication interface 2003 is used to enable communication between the electronic device 2000 and other devices; for example, the communication interface can be a transceiver.
[0327] In the embodiments provided above, the methods provided by the embodiments of this application are described from the perspective of an electronic device (e.g., a mobile phone) as the executing entity. To implement the functions of the methods provided in the embodiments of this application, the electronic device may include hardware structures and / or software modules, implementing the above functions in the form of hardware structures, software modules, or a combination of hardware structures and software modules. Whether a particular function is implemented in the form of hardware structures, software modules, or a combination of hardware structures and software modules depends on the specific application and design constraints of the technical solution.
[0328] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk, SSD), etc. Where there is no conflict, the solutions in the above embodiments can be combined.
[0329] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0330] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in one or more blocks of the flowchart illustrations and / or one or more blocks of the block diagrams.
[0331] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means that implement the functions specified in one or more flowcharts and / or one or more block diagrams.
[0332] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, such that the instructions, which execute on the computer or other programmable apparatus, provide steps for implementing the functions specified in one or more flowcharts and / or one or more block diagrams.
[0333] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the scope and intent of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application is also intended to include such modifications and variations.
Claims
1. A voice interaction method, characterized in that, Applied to an electronic device, the electronic device having a display screen, including: The voice input field is displayed on the screen; Upon detecting the user's first selection operation at the voice input entry, the editing entry is displayed, and the first voice information input by the user is detected to obtain the first input result; After detecting the user's second selection operation on the editing entry, the second voice information input by the user is detected to obtain the first editing instruction; Edit the second input result according to the first editing instruction to obtain the third input result, wherein the second input result includes the first input result and / or the historical input result; The third input result is displayed on the display screen.
2. The method as described in claim 1, characterized in that, The second voice information includes information about the original keywords in the second input result and / or information about the target keywords in the third input result; The information of the original keyword includes at least one of the following: The pinyin of the original keyword, common character combinations of the original keyword, information about the components of the original keyword, or the pronouns of the original keyword; The information of the target keyword includes at least one of the following: The pinyin of the target keyword, common character combinations of the target keyword, information about the components of the target keyword, or the pronouns of the target keyword.
3. The method as described in claim 2, characterized in that, The information of the original keyword includes the pronouns of the original keyword; The method further includes: The original keyword is determined from the second input result based on the pronoun of the original keyword.
4. The method as described in claim 2, characterized in that, The second voice information includes information about the target keyword; The target keyword is determined based on the information of the target keyword; The original keyword is determined based on the type and / or semantics of the target keyword; The step of editing the second input result according to the first editing instruction to obtain the third input result includes: The original keyword in the second input result is modified to the target keyword to obtain the third input result.
5. The method as described in any one of claims 2-4, characterized in that, The method further includes: Multiple candidate results for the keyword are determined based on the second voice information, wherein the keyword includes the original keyword and / or the target keyword; The multiple candidate results are displayed on the display screen; The user's third selection operation on the target candidate result among the multiple candidate results was detected; The keyword is determined based on the target candidate results.
6. The method as described in claim 5, characterized in that, The second voice information includes the pinyin of the keyword and information about the components of the keyword, wherein the information about the components of the keyword includes radical information; The determination of multiple candidate results for keywords based on the second voice information includes: The radical of the keyword is determined based on the radical information of the keyword; Multiple candidate results for the keyword are determined based on the radical of the keyword and the pinyin of the keyword.
7. The method as described in claim 6, characterized in that, The information of the components of the keyword also includes the pinyin of the non-radical parts; The step of determining multiple candidate results for a keyword based on its radical and its pinyin includes: Multiple candidate results for the keyword are determined based on the radical of the keyword, the pinyin of the non-radical part, and the pinyin of the keyword.
8. The method according to any one of claims 2-7, characterized in that, The original keywords include discontinuous text in the second input result; The step of editing the second input result according to the first editing instruction to obtain the third input result includes: Edit the discontinuous text and the text information between the discontinuous text according to the first editing instruction to obtain the third input result.
9. The method as described in claim 8, characterized in that, The text information between the two discontinuous texts is a first type of symbol, and the second speech information includes information for instructing the editing of the discontinuous text. The first type of symbol is used to segment text information within a sentence; or, The second voice information includes information for editing the discontinuous text and the text information located between the discontinuous text.
10. The method according to any one of claims 2-9, characterized in that, The information of the original keyword includes the pinyin of the original keyword, and the method further includes: The keyword is determined from the second input result based on the pinyin of the keyword and the cursor position.
11. The method as described in claim 10, characterized in that, Determining the keyword from the second input result based on the keyword's pinyin and cursor position includes: The keyword is determined within the n1 characters preceding the cursor position in the second input result based on the pinyin of the keyword; and / or, The keyword is determined based on the pinyin of the keyword, and is located within n2 characters after the cursor position in the second input result, where n2 is a positive integer.
12. The method according to any one of claims 2-11, characterized in that, The information of the original keyword includes the pinyin of the original keyword, and the method further includes: The keyword is determined based on the pinyin of the keyword and the voice input record and / or pinyin input record of the second input result.
13. The method according to any one of claims 1-12, characterized in that, The step of detecting the second voice information input by the user to obtain the first editing instruction includes: After detecting the second voice information, the second voice information is recognized, and an input result corresponding to the second voice information is obtained; Semantic analysis is performed on the input result corresponding to the second voice information to obtain the first editing instruction.
14. The method as described in claim 13, characterized in that, The step of performing semantic analysis on the input result corresponding to the second voice information to obtain the first editing instruction includes: One or more words in the input result corresponding to the second speech information are determined by a word segmentation model; The classification information corresponding to the word segment is determined by editing the parameter classification model, and the classification information of any word segment is used to indicate the attribute of the word segment in the first editing instruction; The first editing instruction is determined based on the word segmentation and the classification information.
15. The method as described in claim 13, characterized in that, The step of performing semantic analysis on the input result corresponding to the second voice information to obtain the first editing instruction includes: The attributes of one or more word segments in the input result corresponding to the second speech information in the first editing instruction are determined by a large-scale language model. The first editing instruction is determined based on the word segmentation and the attribute.
16. An electronic device, characterized in that, include: Processor, memory, and one or more programs; The one or more programs are stored in the memory, and the one or more programs include instructions that, when executed by the processor, cause the electronic device to perform the steps of the method as described in any one of claims 1-15.
17. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store a computer program that, when run on a computer, causes the computer to perform the method as described in any one of claims 1 to 15.
18. A computer program product, characterized in that, Includes a computer program that, when run on a computer, causes the computer to perform the method as described in any one of claims 1 to 15.