Text input method and device

By using the microphone in the terminal device to recognize voice input scenarios and automatically convert voice data into text, the problem of accidental keyboard touches in foldable screen devices is solved, improving input efficiency and user experience.

WO2025200611A1PCT designated stage Publication Date: 2025-10-02HONOR DEVICE CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/139474
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-27
Filing Date
2024-12-16
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

In terminal devices with foldable screens, due to insufficient screen width or length, adjacent keys are easily touched by mistake during keyboard input, affecting input efficiency.

Method used

By using a setting of at least two microphones, the audio energy difference of the voice data collected by the microphones is detected, the voice input scenario is identified, and the voice data is automatically converted into text information for display, reducing input errors caused by accidental touches of the keyboard.

Benefits of technology

It improves text input efficiency, reduces errors caused by accidental keyboard touches, and provides a better user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000016_0000
    Figure 00000016_0000
  • Figure 00000016_0001
    Figure 00000016_0001
  • Figure 00000017_0000
    Figure 00000017_0000
Patent Text Reader

Abstract

Disclosed in the present invention are a text input method and a device. The device comprises at least two microphones, the distance between the at least two microphones being larger than a set threshold. The method comprises: in response to an input method calling operation, displaying an input interface, wherein the input interface comprises a content input control; respectively acquiring first speech data and second speech data acquired by microphones at the same time; and upon detecting that audio energy of the first speech data and audio energy of the second speech data meet a preset condition, generating first text information on the basis of the first speech data and the second speech data, and displaying the first text information in the content input control. By means of the described solution, on one hand, automatic speech-to-text conversion can be achieved without the need of user operations, and on the other hand, text input is vocally achieved, so that text input errors caused by mistaken keyboard touch can be reduced, thereby improving the input efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Text input method and device

[0001] This invention claims priority to the Chinese patent application filed with the State Intellectual Property Office of China on March 27, 2024, with application number 202410366293.9 and application name “A text input method and device”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present invention relates to the field of terminal technology, and in particular to a text input method and device. Background Art

[0003] Currently, many terminal devices are equipped with foldable screens that can be folded horizontally or vertically. When the screen is folded, due to the insufficient width or length of the screen, when using the keyboard provided by the input method to input text, adjacent keys are easily accidentally pressed, affecting input efficiency. Summary of the Invention

[0004] The purpose of the present invention is to provide a text input method and device for reducing text input errors caused by accidental keyboard touches in a small-screen display scenario and improving input efficiency.

[0005] In a first aspect, the present invention provides a text input method, which is applied to an electronic device, wherein the electronic device includes at least two microphones, and the distance between the at least two microphones is greater than a set threshold. The text input method includes: displaying an input interface in response to an input method call operation, the input interface including a first control, and the first control is used to display the input content; respectively obtaining first voice data and second voice data collected simultaneously by each microphone; detecting that the audio energy of the first voice data and the second voice data meets a preset condition; generating a first text message based on the first voice data and the second voice data, and displaying the first text message in the first control.

[0006] In this application, the user's operation of invoking the input method can be any operation that has the function of triggering the input method, such as clicking on the content search control, clicking on the message notification bar, clicking on the text input box, etc. The user's operation of invoking the input method can also have other functions while triggering the invocation of the input method, such as triggering the display of the message notified by the message notification bar. The first control can be a content input control, such as a text input box, which can be used to display any characters entered by the user, including text, numbers, symbols, etc.

[0007] In the above implementation scheme, the electronic device can recognize the user's voice input scenario in any text input scenario based on the voice data collected simultaneously by different microphones, and automatically convert the user's input voice data into text for display when it is determined to be in a voice input scenario, thereby reducing input errors caused by accidentally touching keys when using the keyboard and improving text input efficiency.

[0008] As described above, in certain implementations of the first aspect of a text input method, the electronic device further includes a first display screen and a second display screen, the second display screen and the first display screen are respectively located on different surfaces of the electronic device, the display area of ​​the second display screen is smaller than the first display screen, the first display screen includes at least two sub-display portions and is foldable; the distance between at least two microphones is greater than a set threshold, including: in the folded state, unfolded state and intermediate state of the first display screen, the distance between at least two microphones is greater than the set threshold.

[0009] In the present application, the first display screen may be the inner screen of the electronic device, and the second display screen may be the outer screen of the electronic device. The inner screen of the electronic device is a foldable screen. In the present application, the inner screen may include three different states: a folded state, an unfolded state, and an intermediate state. It can be understood that the intermediate state refers to a state between the folded state and the unfolded state, such as the state in which the user unfolds the outer screen when the outer screen is in the folded state, but has not yet fully unfolded it; or the state in which the user folds the outer screen when the outer screen is in the unfolded state, but has not yet fully folded it. When the inner screen is in the folded state, the display function can usually be achieved through the outer screen. Since the display area of ​​the outer screen is smaller than that of the inner screen, there is a problem of accidental touch when inputting text on the keyboard in the external screen display scenario.

[0010] In this implementation, by setting different microphone positions, the foldable screen device can use the voice data collected simultaneously by different microphones in the folded state, unfolded state and intermediate state to recognize the input scene, thereby improving the reliability and practicality of the solution.

[0011] In some implementations of the first aspect of the text input method described above, the text input method further includes: determining that the electronic device is in a target display scene.

[0012] In certain implementations of the first aspect of the text input method as described above, the target display scene includes a scene in which the second display screen is in a display state.

[0013] In this implementation, scenario restrictions can be added to the text input method provided by this application, that is, the electronic device only applies the text input method provided by this application to realize automatic speech-to-text input when it is determined to be in an external screen display scenario, so as to improve the input efficiency in the small-screen input scenario; in other display scenarios, the conventional keyboard-based text input method can still be applied, which is more in line with the user's text input habits.

[0014] In certain implementations of the first aspect of the text input method as described above, the target display scene includes a scene in which the first display screen is in a split-screen display state or a scene in which the first display screen is in a floating window display state.

[0015] In this implementation, scenario restrictions can be added to the text input method provided by this application, that is, the electronic device only applies the text input method provided by this application to realize automatic speech-to-text input when it is determined to be in a split-screen display scenario of the inner screen or a floating window display scenario of the inner screen, so as to improve the input efficiency in the small-screen input scenario; in other display scenarios, the conventional keyboard-based text input method can still be applied, which is more in line with the user's text input habits.

[0016] In some implementations of the first aspect of the text input method described above, the input interface further includes a second control, and when the second control is in the first display state, the input interface does not include a keyboard control.

[0017] In this application, the second control may be an input mode switching control, which may include a first display state. When in the first display state, the input interface does not display a keyboard control. In this case, the user can trigger the electronic device to perform speech-to-text input by inputting voice data. The input mode switching control can prompt the user of the current input mode through its display state, thereby improving the user experience.

[0018] As described above, in certain implementations of the first aspect, the text input method further includes: when the second control is in the first display state, in response to a triggering operation on the second control, the second control switches to the second display state; when the second control is in the second display state, the input interface includes a keyboard control.

[0019] In this application, the input mode switch control may also include a second display state. The user can switch between the first and second display states by clicking the input mode switch control. When the second control is in the second display state, the input interface includes a keyboard control, and the user can input text using the keyboard. The input mode switch control can use its display state to inform the user of the current input mode, thereby improving the user experience.

[0020] In some implementations of the first aspect of the text input method described above, the text input method further includes: when the second control is in the first display state for the first time, the input interface further includes a third control, and the third control is used to display input method prompt information.

[0021] In this application, when the input mode switch control is in the first display state for the first time, the user may not be able to directly determine the current input mode because the input interface does not include a keyboard control. In this regard, in an embodiment of this application, an input mode prompt control can also be displayed in the input interface, and input mode prompt information, such as "Please speak into the microphone", can be displayed in the control. This can help the user determine the specific operation method when the user first uses the text input method provided by this application.

[0022] In some implementations of the first aspect of the text input method described above, the text input method further includes: hiding the third control in response to a triggering operation on the third control.

[0023] In this implementation, the input method prompt control may also include a response control for receiving a user response. After the user determines the current input method based on the input method prompt information, they may click the input method prompt control or click a response control within the input method prompt control. Furthermore, the electronic device may hide the input method prompt control in response to the user's operation. In this way, the input method prompt control can be hidden after the input method prompt function is implemented, reducing obstruction of the interface content.

[0024] As described above, in certain implementations of the first aspect, the input interface further includes a fourth control, and the above-mentioned text input method further includes: in response to a long press operation on the fourth control, obtaining third voice data collected by any microphone; in response to releasing the fourth control, sending the third voice data to the target receiving end.

[0025] In this implementation, the fourth control can be another form of input method prompt control. In an online chat scenario, if the chat application supports sending both text messages and voice messages, the input method prompt control can display prompts for each supported input method. This prompt could be, for example, "Press and hold or speak into the microphone." This prompt can prompt the user to use the voice-to-text input method, while also collecting and sending voice data in response to a long press.

[0026] In some implementations of the first aspect of the text input method as described above, before respectively obtaining the first voice data and the second voice data collected simultaneously by each microphone, the text input method further includes: obtaining first posture data collected by the inertial measurement unit; and determining, based on the first posture data, that the posture characteristics of the electronic device are consistent with the preset characteristics.

[0027] In this implementation, since in the scenario of inputting voice data, the user usually lifts the electronic device from the chest position to a position close to the head, and then moves the microphone close to the lips for voice input, therefore, by increasing the recognition of the electronic device posture, the accuracy of scene recognition can be further improved, which is beneficial to improving the reliability of this solution.

[0028] In certain implementations of the first aspect of the text input method described above, detecting that the audio energy of the first voice data and the second voice data meets a preset condition includes: detecting that the audio energy difference between the first voice data and the second voice data is greater than a preset value.

[0029] As described above, in certain implementations of the first aspect of a text input method, before detecting that the audio energy of the first voice data and the second voice data meets a preset condition, the above text input method further includes: preprocessing the first voice data and the second voice data, and the preprocessing includes security detection and / or voice enhancement.

[0030] As described above, in certain implementations of the first aspect, after the first text information is displayed in the first control, the above-mentioned text input method also includes: detecting that the collection of the first voice data and the second voice data is completed, starting the first timer; in response to the timing duration of the first timer exceeding the duration threshold, sending the input content displayed in the first control to the target receiving end.

[0031] In this implementation, while collecting the first and second voice data, the collected voice data can be converted simultaneously, and the converted text information can be displayed in the content input control. When it is detected that the collection of the first and second voice data is complete, a timer can be started. Before the timer expires, the user can re-edit the content displayed in the content input control. When the timer expires, the electronic device can automatically send the content displayed in the content input control to the target receiving end.

[0032] Among them, detecting that the collection of the first voice data and the second voice data is completed can specifically be that semantic recognition can be performed on the voice data input by the user. According to the semantic recognition result, if it is determined that the semantic information is complete, it can be considered that the user has completed the input and the collection of the first voice data and the second voice data is completed; or, the pause duration of the voice input can be detected. When no voice data is detected after exceeding the set time threshold, it can be considered that the user has completed the input and the collection of the first voice data and the second voice data is completed.

[0033] As described above, in certain implementations of the first aspect, after the first text information is displayed in the first control, the above-mentioned text input method also includes: detecting that the collection of the first voice data and the second voice data is completed, starting a second timer; in response to the timing duration of the second timer exceeding a duration threshold, searching for a target object in the target database and displaying the searched target object based on the input content displayed in the first control.

[0034] In some implementations of the first aspect of the text input method described above, the text input method further includes: displaying a fifth control on the input interface, where the fifth control is used to describe the timing duration.

[0035] In this application, the fifth control may be a progress bar, which can intuitively display the time duration to the user, thereby improving the user experience.

[0036] In some implementations of the first aspect of the text input method described above, the input interface further includes at least one preset phrase; after the input interface is displayed, the text input method further includes: in response to a triggering operation on any one of the preset phrases, sending any one of the preset phrases to a target receiving end.

[0037] In a second aspect, the present technical solution provides an electronic device comprising: one or more processors; a memory; and one or more computer programs, wherein the one or more computer programs are stored in the memory, and the one or more computer programs include instructions that, when executed by the device, enable the device to execute the method in the first aspect or any possible implementation of the first aspect.

[0038] In a third aspect, the present invention further provides a computer-readable storage medium storing program code for execution by a device, wherein the program code includes instructions for executing the method in the first aspect or any possible implementation of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] FIG1 is a schematic diagram of a scenario of a text input method provided by an embodiment of the present application;

[0040] FIG2 is a schematic diagram of another scenario of the text input method provided in an embodiment of the present application;

[0041] FIG3 is a schematic diagram of another scenario of the text input method provided in an embodiment of the present application;

[0042] FIG4 is a schematic diagram of another scenario of the text input method provided in an embodiment of the present application;

[0043] FIG5 is a schematic structural diagram of an electronic device provided in an embodiment of the present application;

[0044] FIG6 is a schematic diagram of a microphone configuration method in an electronic device according to an embodiment of the present application;

[0045] FIG7 is another schematic diagram of a microphone configuration method in an electronic device according to an embodiment of the present application;

[0046] FIG8 is another schematic diagram of a microphone configuration method in an electronic device according to an embodiment of the present application;

[0047] FIG9 is another schematic diagram of a microphone configuration method in an electronic device according to an embodiment of the present application;

[0048] FIG10 is a schematic diagram of another scenario of the text input method provided in an embodiment of the present application;

[0049] FIG11 is a schematic diagram of another scenario of the text input method provided in an embodiment of the present application;

[0050] FIG12 is a schematic diagram of another scenario of the text input method provided in an embodiment of the present application;

[0051] FIG13 is a schematic diagram of another scenario of the text input method provided in an embodiment of the present application;

[0052] FIG14 is a schematic diagram of another scenario of the text input method provided in an embodiment of the present application;

[0053] FIG15 is a schematic diagram of another scenario of the text input method provided in an embodiment of the present application;

[0054] FIG16 is a schematic diagram of another scenario of the text input method provided in an embodiment of the present application;

[0055] FIG17 is a schematic diagram of another scenario of the text input method provided in an embodiment of the present application;

[0056] FIG18 is a schematic diagram of another scenario of the text input method provided in an embodiment of the present application;

[0057] FIG19 is a schematic diagram of another scenario of the text input method provided in an embodiment of the present application;

[0058] FIG20 is a schematic diagram of another scenario of the text input method provided in an embodiment of the present application;

[0059] FIG21 is a schematic diagram of another scenario of the text input method provided in an embodiment of the present application;

[0060] FIG22 is a schematic diagram of another scenario of the text input method provided in an embodiment of the present application;

[0061] FIG23 is a schematic diagram of another scenario of the text input method provided in an embodiment of the present application;

[0062] FIG24 is a schematic diagram of another scenario of the text input method provided in an embodiment of the present application;

[0063] FIG25 is a schematic diagram of another scenario of the text input method provided in an embodiment of the present application;

[0064] FIG26 is another structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0065] When using mobile devices, such as mobile phones, users often need to use input methods and the keyboards they provide to input text. However, in certain display scenarios, the width or length of the display interface is insufficient, resulting in a tightly packed keyboard layout. This makes it easy to accidentally touch adjacent keys, leading to frequent input errors and affecting input efficiency.

[0066] Figure 1 shows a display scene of a foldable screen device in a folded state. As shown in Figure 1, two screens, an inner screen and an outer screen, are respectively provided on different surfaces of the foldable screen device, wherein the inner screen can be folded in the vertical direction, and the display area of ​​the inner screen is larger than the display area of ​​the outer screen. When the inner screen is in a folded state, the display function of the device will be realized by the outer screen. At this time, when the user needs to perform text input operations on the outer screen, due to the small display area of ​​the outer screen, as shown in Figure 1, the keyboard layout displayed is too compact. When the user clicks a specific key, the adjacent keys can easily be accidentally touched, resulting in frequent erroneous input, affecting input efficiency.

[0067] Figure 2 shows the display scene of another foldable screen device in the folded state. As shown in Figure 2, the foldable screen device is provided with two screens, an inner screen and an outer screen, on different surfaces. The inner screen of the foldable screen device can be folded horizontally. Similarly, the display area of ​​the inner screen is larger than the display area of ​​the outer screen. Similar to Figure 1, when the user needs to input text on the outer screen, because the outer screen is relatively narrow, the keys of the keyboard displayed on it are also relatively compact, and there is also a high probability of accidental touches.

[0068] Figure 3 shows a schematic diagram of a split-screen display scenario, which can be a split-screen display scenario of a non-folding screen device, or a split-screen display scenario of a folding screen device in a non-folding state. As shown in Figure 3, in the split-screen display scenario, the display screen of the device is divided into two parts, the upper and lower parts, which are used to display different interface contents respectively. For example, the upper part can display video content, and the lower part can display a chat interface. When using one of the display interfaces for text input, due to the insufficient height of the display interface, on the one hand, the display layout of the keyboard will be relatively compact, affecting the accuracy when triggering the key. On the other hand, when the keyboard is displayed, the part of the text that has been entered will inevitably be obscured, resulting in the user being unable to observe whether the entered content is accurate. The user needs to frequently hide the keyboard to check the entered text, and then redisplay the keyboard to continue the input operation, which is cumbersome.

[0069] Figure 4 shows a schematic diagram of a floating window display scenario, which can be a floating window display scenario of a non-folding screen device, or a floating window display scenario of a folding screen device in a non-folding state. In the floating window display scenario, the display interface is displayed in the form of a small window floating above another display interface. In one scenario, for example, the user receives a text message while watching a video. At this time, in order not to affect the progress of the video viewing, the message reply interface can be displayed through the floating window, and the text message can be replied. However, since the display interface of the floating window is small, it is easy for the user to touch it by mistake when using the keyboard to input text.

[0070] It can be seen that in various common small-screen display scenarios of terminal devices, there is the problem of inconvenience in inputting text using the keyboard, which affects the efficiency of text input.

[0071] In order to solve the above problems, this application is proposed.

[0072] In an embodiment of the present application, an electronic device may be provided, comprising at least two microphones, each of which is positioned at a specific distance from the user. The electronic device may display an input interface upon detecting a user invoking an input method. The user invoking the input method may be any operation that triggers the input method, and the operation may also have other functions while triggering the invocation of the input method. For example, in one scenario, the user invoking the input method may be clicking on a content search control; in another scenario, the user invoking the input method may be clicking on a text input box; in yet another scenario, the user invoking the input method may be clicking on a message notification bar. In this scenario, the user invoking the input method not only triggers the input method but also triggers the display of a message in the message notification bar, allowing the user to access the message content. Furthermore, the electronic device may determine whether the current scenario is a voice input scenario by using the difference in audio energy between the voice data simultaneously collected by the at least two microphones at different locations. If the current scenario is determined to be a voice input scenario, the electronic device may activate a speech-to-text function, directly converting the voice data collected by the at least two microphones at different locations into text information, and displaying the text information in the text input box.

[0073] Through the above solution, on the one hand, text can be input without calling the keyboard, which can reduce input errors caused by accidental keyboard touches and improve input efficiency; on the other hand, the electronic device can automatically start the voice input to text function based on the voice data collected by different microphones. For users, when they need to use the voice to text function, they can directly speak the text content they want to input into the microphone, without the need for users to perform unnecessary operations, and the user experience is better.

[0074] FIG5 shows a schematic structural diagram of an electronic device 100 provided in an embodiment of the present application.

[0075] The electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display 194, and a subscriber identification module (SIM) card interface 195. The sensor module 180 may include a pressure sensor 180A, a touch sensor 180K, a gyroscope sensor 180B, an acceleration sensor 180E, and the like.

[0076] It should be understood that the structure illustrated in the embodiments of the present invention does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0077] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU). The different processing units may be independent devices or integrated into one or more processors.

[0078] The controller can generate operation control signals according to the instruction operation code and timing signal to complete the control of instruction fetching and execution.

[0079] Processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in processor 110 is a cache memory. This memory can store instructions or data that have just been used or are being recycled by processor 110. If processor 110 needs to use the same instruction or data again, it can directly access the memory. This avoids duplicate accesses, reduces processor 110 latency, and thus improves system efficiency.

[0080] In some embodiments, the processor 110 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface.

[0081] The I2C interface is a bidirectional synchronous serial bus that includes a serial data line (SDA) and a serial clock line (SCL). In some embodiments, the processor 110 may include multiple I2C bus lines. The processor 110 may be coupled to the touch sensor 180K, the charger, the flash, the camera 193, and the like via different I2C bus interfaces. For example, the processor 110 may be coupled to the touch sensor 180K via the I2C interface, enabling communication between the processor 110 and the touch sensor 180K via the I2C bus interface, thereby implementing the touch function of the electronic device 100.

[0082] The MIPI interface can be used to connect the processor 110 and the display screen 194 . The processor 110 and the display screen 194 communicate via the DSI interface to implement the display function of the electronic device 100 .

[0083] It is understood that the interface connection relationship between the modules illustrated in the embodiment of the present invention is merely an illustrative illustration and does not constitute a structural limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may also adopt different interface connection methods from the above embodiments, or a combination of multiple interface connection methods.

[0084] Electronic device 100 implements display functionality through a GPU, display screen 194, and an application processor. A GPU is a microprocessor for image processing that connects display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. Processor 110 may include one or more GPUs that execute program instructions to generate or modify display information.

[0085] Display screen 194 is used to display images, videos, etc. Display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLED, a Micro-oLed, or a quantum dot light-emitting diode (QLED).

[0086] In some embodiments, the electronic device 100 may include one or N display screens 194 , which may be located on different surfaces of the electronic device 100 , and at least one of the N display screens 194 may be a foldable display screen, where N is a positive integer greater than 1.

[0087] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 via the external memory interface 120 to implement data storage functions. For example, files such as music and videos can be stored on the external memory card.

[0088] The internal memory 121 can be used to store computer executable program codes, which include instructions. The internal memory 121 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system, an application required for at least one function (such as a text input function, etc.), etc. The data storage area may store data created during the use of the electronic device 100 (such as voice data, etc.), etc. In addition, the internal memory 121 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc. The processor 110 executes various functional applications and data processing of the electronic device 100 by running instructions stored in the internal memory 121 and / or instructions stored in a memory provided in the processor.

[0089] Pressure sensor 180A is used to sense pressure signals and convert them into electrical signals. In some embodiments, pressure sensor 180A may be located on display screen 194. When a touch operation is performed on display screen 194, electronic device 100 detects the intensity of the touch operation using pressure sensor 180A. Electronic device 100 may also calculate the location of the touch based on the detection signal from pressure sensor 180A.

[0090] The touch sensor 180K is also called a "touch-sensitive device." The touch sensor 180K can be disposed on the display screen 194. The touch sensor 180K and the display screen 194 form a touch screen, also called a "touch screen." The touch sensor 180K is used to detect touch operations applied thereto or in the vicinity thereof. The touch sensor can transmit the detected touch operations to the application processor to determine the type of touch event. Visual output related to the touch operations can be provided via the display screen 194. In other embodiments, the touch sensor 180K can also be disposed on the surface of the electronic device 100, at a location different from that of the display screen 194.

[0091] The gyro sensor 180B may be used to determine the motion posture of the electronic device 100. In some embodiments, the angular velocity of the electronic device 100 around three axes (ie, x, y, and z axes) may be determined by the gyro sensor 180B.

[0092] The accelerometer 180E can detect the magnitude of the acceleration of the electronic device 100 in various directions (generally three axes). When the electronic device 100 is stationary, it can detect the magnitude and direction of gravity. It can also be used to identify the posture of the electronic device.

[0093] The electronic device 100 can implement audio functions, such as recording, through the audio module 170 , the speaker 170A, the receiver 170B, the microphone 170C, the headphone jack 170D, and the application processor.

[0094] The audio module 170 is used to convert digital audio information into analog audio signal output, and is also used to convert analog audio input into digital audio signals. The audio module 170 can also be used to encode and decode audio signals. In some embodiments, the audio module 170 can be provided in the processor 110, or some functional modules of the audio module 170 can be provided in the processor 110.

[0095] Microphone 170C, also known as "microphone" or "microphone", is used to convert sound signals into electrical signals. When making a call or sending a voice message, the user can speak by putting their mouth close to the microphone 170C to input the sound signal into the microphone 170C. The electronic device 100 can set two microphones 170C at different positions at a specific distance apart. In other embodiments, in addition to collecting sound signals, the two microphones 170C can also achieve noise reduction. In other embodiments, the electronic device 100 can also be provided with three, four or more microphones 170C to collect sound signals, reduce noise, identify the source of sound, achieve directional recording functions, etc.

[0096] In actual scenarios, when a user needs to input voice data, they usually consciously move closer to one of the microphones of the electronic device to input. If the distance between the other microphone and the current microphone is far, then the distances between the two microphones and the sound source are significantly different, and the audio energy of the collected voice data will be significantly different. For example, the difference in audio energy will be greater than a specific threshold. In a regular conversation scenario with others, the user will not consciously move closer to any microphone of the electronic device. At this time, there is no significant difference in the distance between the sound source and different microphones. At this time, there is no significant difference in the audio energy of the voice data collected by different microphones. Therefore, the user's intention can be judged by the size of the audio energy difference of the voice data collected by different microphones.

[0097] Based on the above description, in this embodiment of the present application, upon detecting a user invoking an input method, the system can distinguish between the user's voice input scenario and a regular conversation scenario with others by detecting the audio energy of the voice data collected simultaneously by different microphones. Furthermore, when the current scenario is identified as voice input, the user's voice data can be directly converted into text information for input and display.

[0098] To implement the above technical solution, the embodiment of the present application can set the positions of different microphones in the electronic device and the distances between different microphones to ensure that the audio energy of the voice data collected by different microphones can be used to identify the voice input scene.

[0099] The following describes how to configure microphones in different types of electronic devices according to embodiments of the present application.

[0100] 1. Folding screen devices that fold vertically

[0101] For foldable screen devices that fold in the vertical direction, in the related art, referring to Figure 6, microphones are usually set near the upper and lower frames of the electronic device. With this setting, although the distance between the two microphones in the unfolded state can meet the requirements of the embodiment of the present application for recognizing voice input scenarios, when the device is in the folded state, the positions of the two microphones are close to each other. At this time, since the distance between the two microphones is too close, in the voice input scenario, the distance between the two microphones and the sound source is close, then the audio energy of the voice data collected by the two microphones will not be significantly different. At this time, it will be impossible to distinguish the voice input scenario from the normal conversation scenario through the two-way voice data.

[0102] Therefore, in the embodiment of the present application, for a foldable screen device that folds in a vertical direction, the positions of at least two microphones can be set so that the distance between the at least two microphones is greater than 10 in the folded state, the unfolded state, and the intermediate state of the electronic device. It can be understood that the intermediate state refers to a state between the folded state and the unfolded state, such as a state in which the user unfolds the outer screen when the outer screen is folded but not fully unfolded; or a state in which the user folds the outer screen when the outer screen is unfolded but not fully folded.

[0103] The value of l0 can be determined based on the difference in audio energy of the voice data collected by the two microphones in a voice input scenario when the two microphones are at different distances from each other. For example, the difference in audio energy of the voice data collected by the two microphones in a normal conversation scenario can be detected. Furthermore, in a voice input scenario, the difference in audio energy of the voice data collected by the two microphones at different distances from each other can be detected, and the distance between the two microphones when the difference in audio energy equals a specific threshold is determined as l0. The specific threshold refers to the threshold value that distinguishes between a voice input scenario and a normal conversation scenario.

[0104] For example, as shown in 7A of FIG7 , at least two microphones can be respectively arranged at two mutually perpendicular frames of the electronic device, or, as shown in 7B of FIG7 , at two side frames. In other implementations, provided that the distance between the at least two microphones is greater than 10° when the electronic device is in the folded state, the unfolded state, and the intermediate state, the at least two microphones can also be arranged in other ways, such as being respectively arranged at the upper and right frames, or the lower and left frames, or the lower and right frames of the electronic device, etc., and this embodiment of the present application does not impose any limitation on this.

[0105] It should be noted that, in the embodiment of the present application, the configuration of each microphone may be to add a new microphone and set the position of the new microphone, or to change the position of the original microphone.

[0106] 2. Folding screen devices that fold horizontally

[0107] For folding screen devices that fold in the horizontal direction, in the related art, referring to Figure 8, microphones are usually provided near the upper and lower frames of the electronic device. In the unfolded state, the two microphones are far apart, and the audio energy of the voice data collected by the two microphones in the voice input scenario has obvious differences. Moreover, in the folded state, since horizontal folding does not change the distance between the microphones, the audio energy of the two channels of voice data collected in the voice input scenario still has obvious differences. Therefore, in the embodiment of the present application, for the horizontal folding screen device, the two microphones can still be respectively provided near the upper and lower frames of the electronic device.

[0108] Alternatively, in another implementation, the two microphones may be respectively arranged at other positions of the electronic device, such as at two side frame positions, or at two frame positions perpendicular to each other. However, it should be noted that in the folded state, unfolded state, and intermediate state of the electronic device, the distance between the two microphones should be greater than the above l 10 .

[0109] 3. Non-foldable screen devices

[0110] For non-folding screen devices, in the related art, as shown in FIG9 , microphones are usually set near the upper and lower frames of the electronic device, and the two microphones are far apart. In the voice input scenario, the audio energy of the voice data collected by the two microphones has obvious differences. Therefore, in the embodiment of the present application, for non-folding screen devices, the two microphones can still be set at the positions shown in FIG9 . Alternatively, the two microphones can also be set at other positions of the electronic device, and the distance between the two microphones should be greater than the above l 10 .

[0111] In the following embodiments of the present application, an electronic device having the above-mentioned hardware structure is taken as an example to illustrate the implementation of the text input method provided by the present application.

[0112] In an embodiment of the present application, the user's input method call operation can be detected, and after the input method call operation is detected, the input interface is displayed. In an embodiment of the present application, the input interface may not display a keyboard control, and the user may input text content by direct voice input. Alternatively, in another implementation, the input interface may also display a keyboard control. At this time, the user can still implement text input by directly inputting voice. Furthermore, the electronic device can identify the current scene based on the difference in audio energy of the voice data collected simultaneously by at least two microphones. When it is determined to be a voice input scene, the voice-to-text input function can be automatically started, and the electronic device automatically converts the voice data into text information and displays it in the content input control.

[0113] In one possible implementation, the electronic device can directly display the keyboard-free input interface after detecting the user's input method call operation. Through this implementation, the keyboard-free input interface can be displayed in any text input scenario, allowing the user to input text by activating the voice-to-text input function without using the keyboard, thereby improving the user's text input efficiency.

[0114] Alternatively, in another implementation, the electronic device may first detect the current display scene after detecting the user's input method call operation. When it is determined that the current scene is a small-screen display scene, the electronic device displays the above-mentioned keyboard-free input interface so that the user can input text without operating the keyboard. Through this implementation, the above-mentioned keyboard-free input interface can be displayed in a small-screen display scene, that is, when keyboard input is inconvenient, so that the user can input text by starting the above-mentioned voice-to-text input function. In non-small-screen scenes, the input interface with a keyboard can still be displayed to provide the user with a conventional keyboard input method, which is more in line with the user's text input habits.

[0115] Among them, exemplarily, the above-mentioned small-screen display scene can be: the external screen display scene of the electronic device in the vertical folding state as shown in Figure 1, the external screen display scene of the electronic device in the horizontal folding state as shown in Figure 2, the split-screen display scene of the electronic device as shown in Figure 3, the floating window display scene of the electronic device as shown in Figure 4, etc.

[0116] The text input method provided in the embodiments of the present application can be applied to various text input scenarios of electronic devices, including but not limited to online chat scenarios, note creation scenarios, schedule creation scenarios, email sending scenarios, content search scenarios, etc. In the embodiments of the present application, the implementation of the above text input method is described in detail using the online chat scenario as an example.

[0117] In one possible implementation, in an online chat scenario, upon detecting that a user has triggered an input method call, an input interface 21 may be displayed. As shown in FIG10 , the input interface 21 may include a content input control 210, which may be used to display the text content entered by the user. Furthermore, the input interface 21 may not display a keyboard control; the user may enter text content directly through voice input, which the electronic device then automatically converts into text and displays in the content input control 210. Furthermore, in an online chat scenario, the input interface 21 may also display information such as the message sender's identification, the content of received messages, and the time the messages were received.

[0118] The detection of a user triggering an input method calling operation may be, for example, the detection of a user triggering an operation on a message notification bar.

[0119] Specifically, in an exemplary scenario, as shown in 11A of FIG11 , the electronic device receives a new message notification in the lock screen state, the lock screen interface lights up, and at the same time, the message notification bar 31 can be displayed in the lock screen interface. The message notification bar 31 can be displayed in the middle of the lock screen interface, or it can be displayed at the top or bottom of the lock screen interface. The embodiment of the present application does not limit this. The message notification bar 31 can display prompt information of the new message, such as sender information, message content, message reception time, etc. The user's input method call operation can be, for example, a trigger operation on the message notification bar 31. Specifically, the user can complete the unlock operation on the lock screen interface, and then click on the message notification bar 31. In response to the user's triggering operation on the message notification bar 31, the electronic device can display an input interface for the chat message with the sender, as shown in 11B of FIG11 . The input interface can be, for example, displayed floating above the layer where the lock screen interface is located.

[0120] Alternatively, in another exemplary scenario, the electronic device receives a new message notification while unlocked. At this point, as shown in FIG12A , the electronic device may display a message notification bar 31 in the currently displayed interface. The currently displayed interface may be, for example, the electronic device's main interface or an application interface of any application. In response to a user triggering the message notification bar 31, the electronic device may display a chat message input interface as shown in FIG12B .

[0121] Detecting that the user has triggered an input method call operation may also be detecting that the user has triggered an operation on a message input box in a chat interface. In an exemplary scenario, the user may enter the application interface through the application icon of the chat application displayed on the main interface, as shown in 13A of FIG13 , and a chat list 41 with each friend may be displayed in the application interface. In response to a trigger operation on any chat list item, a chat interface 410 with the corresponding friend may be displayed, as shown in 13B of FIG13 . A message input box 411 may be displayed in the chat interface 410, and the user's input method call operation may be, for example, a trigger operation on the message input box 411. After detecting that the user has clicked on the message input box 411, the electronic device may display the input interface, as shown in 13C of FIG13 .

[0122] Furthermore, in another implementation, as shown in FIG14 , the input interface may further include an input mode switch control 211. The input mode switch control 211 may include a first display state and a second display state. The electronic device may switch the display state of the input mode switch control 211 in response to a user triggering operation on the input mode switch control 211, such as a click operation. When the input mode switch control 211 is in the first display state, as shown in FIG15A , the input interface does not display a keyboard control. In this case, the user may directly input text content by voice input, and the electronic device automatically converts the voice data into text information and displays it in the content input control. When the input mode switch control 211 is in the second display state, as shown in FIG15B , the input interface displays a keyboard control. In this case, the user may input text content by keyboard input. The electronic device may generate corresponding text information and display it in the content input control in response to the user triggering operations on each key in the keyboard control.

[0123] This implementation provides users with an input mode switch button, allowing them to flexibly switch between different text input modes based on their needs. This includes using the keyboard control when the keyboard is displayed, or using the speech-to-text input function provided in the embodiments of the present application when the keyboard is not displayed. This allows users to meet diverse input needs.

[0124] When the keyboard control is not displayed in the input interface, in order to facilitate the user to understand the text input method of the current input interface, in an embodiment of the present application, input method prompt information can also be displayed in the input interface.

[0125] In one possible implementation, the chat application only supports sending text messages. In this case, the input method prompt information is displayed in the input interface. When it is detected that the input interface shown in Figure 10 is displayed for the first time, or the input interface shown in Figure 14 is displayed for the first time and the input method switch button is in the first display state, as shown in Figure 16, an input method prompt control 220 can be displayed in a layer above the input interface. The input method prompt control 220 can contain input method prompt information, such as "Please speak into the microphone". Furthermore, the input method prompt control 220 can also include a control 221 for receiving a user response. In response to the user triggering the control 221, the electronic device can hide the input method prompt control 220.

[0126] In another possible implementation, a chat application supports sending both text and voice messages. In this case, the input method prompt information displayed in the input interface can be, as shown in FIG17 , an input method prompt control 230 displayed in the input interface. The input method prompt control 230 can display prompt information for the various input methods supported by the current interface. This prompt can, for example, be "Press and hold or speak into the microphone." In this case, the user can enter voice data by long-pressing the input method prompt control 230. Upon detecting that the user has long-pressed the input method prompt control 230, the electronic device can collect the user's input voice data via any microphone and transmit the collected voice data to the recipient upon detecting that the user has released the input method prompt control 230. This implementation allows the sending of voice messages. Alternatively, the user can directly enter voice data into the microphone without pressing and holding the input method prompt control 230. In this case, the electronic device can automatically convert the input voice data into a text message and send it to the recipient using the voice-to-text function provided in this embodiment of the present application. This implementation allows text message entry without requiring a keyboard control.

[0127] The following describes a specific implementation method of the speech-to-text input function provided in the embodiment of the present application.

[0128] Specifically, in voice input scenarios, users typically consciously approach one of the microphones to input voice data. Therefore, the sound source is closer to one of the microphones and farther from the remaining microphones. In this case, the voice data detected by the two microphones at different locations will have a significant difference in audio energy. In contrast, in regular conversations between users and others, the distances between the sound source and the microphones are not significantly different. In this case, the voice data detected by the two microphones at different locations on the electronic device will not have a significant difference in audio energy.

[0129] Based on the above description, in an embodiment of the present application, the electronic device can receive voice data collected by at least two microphones at the same time, and then compare at least two channels of voice data. When it is determined that the audio energy of each channel of voice data meets the preset conditions, it is considered that the current scene is a voice input scene, and the voice-to-text input function is activated. On the contrary, when it is determined that the audio energy of each channel of voice data does not meet the preset conditions, it can be considered that the current scene is not a voice input scene. At this time, the currently detected voice data can be discarded, and it is determined not to activate the voice-to-text input function.

[0130] Specifically, the electronic device can input the voice data collected by each microphone simultaneously into the target model. The target model can be used to learn and compare the audio energy features of the voice data collected simultaneously by different microphones to determine whether the audio energy of the voice data collected simultaneously by different microphones meets the preset conditions, such as whether the difference in audio energy is greater than a set threshold. When it is determined that the audio energy meets the preset conditions, the current scene is determined to be a voice input scene. At this time, the target model can output a judgment result of yes, which is used to indicate that the voice input to text function is started. When it is determined that the audio energy does not meet the preset conditions, it means that the current scene does not meet the voice input scene. At this time, the target model can output a judgment result of no, which is used to indicate that the voice input to text function is not started. In the embodiment itself, the target model can be, for example, a pre-trained artificial intelligence model. The target model can be related to the distance between different microphones, or the detection algorithm of the audio energy features. When the hardware design of the microphone in the electronic device is changed or the algorithm is updated, the target model can be upgraded.

[0131] Furthermore, in the use scenario of an electronic device, the user usually holds the electronic device in front of the chest for use. When voice data needs to be input, the user needs to lift the electronic device from the chest position to a position close to the head, and then move the microphone close to the lips to input voice. Therefore, in order to further improve the accuracy of the above-mentioned voice input scenario judgment results, in another possible implementation method of the embodiment of the present application, an inertial measurement unit (IMU) can be used to detect the posture data of the electronic device. The IMU may include an acceleration sensor, a gyroscope sensor, etc. Then, based on the detected posture data, it can be determined whether the vertical displacement and posture of the electronic device meet the preset conditions. The preset conditions can be determined based on the displacement and posture of the electronic device during the process of moving from the chest position to the head position in the voice input scenario. For example, the vertical displacement of the electronic device can be within a preset range, and / or the angle between the plane of the electronic device and the ground is less than a set threshold, etc. When it is determined based on the posture data that the electronic device's vertical displacement, posture, and other conditions meet preset conditions, the electronic device may obtain voice data collected by at least two microphones, and then compare the at least two channels of voice data. If it is determined that the audio energy of each channel of voice data meets preset characteristics, it is considered that the current scenario is a voice input scenario, and the voice-to-text input function is activated. If it is determined based on the posture data that the electronic device's vertical displacement, posture, and other conditions do not meet the preset conditions, it can be determined that the current scenario is not a voice input scenario. In this case, the present method process can be directly terminated without further testing of the voice data.

[0132] Through the above implementation, scene recognition can be performed based on the posture data of the electronic device and the voice data collected simultaneously by different microphones, thereby improving the accuracy of scene recognition and thus helping to improve the reliability of this solution.

[0133] After determining to start the voice-to-text input function, as shown in FIG18 , the electronic device can convert the voice data collected by the microphone into text information in real time during the process of the microphone collecting voice data, and display the converted text information in real time in the content input control. Based on the above implementation method, the real-time conversion and display of the input voice data can be achieved during the user input of voice data. In this way, the user can check at any time whether the text content displayed in the content input control is consistent with the voice data he has entered during the process of inputting voice data, and when the user's voice data input is completed, the corresponding entire text content can be immediately displayed in the content input control, which can reduce the time the user waits for voice-to-text conversion after the voice input is completed.

[0134] Furthermore, in online chat scenarios, while converting the user's input voice data into text information in real time and displaying it, the electronic device can also detect whether the user's voice input is complete. For example, the electronic device can perform semantic recognition on the user's input voice data. If, based on the semantic recognition results, the semantic information is determined to be complete, the user can be deemed to have completed input, at which point the electronic device can determine that voice data collection is complete. Alternatively, the electronic device can detect the duration of a pause in the voice input. If no voice data is detected for a set duration exceeding a threshold, the user can be deemed to have completed input, at which point the electronic device can determine that voice data collection is complete.

[0135] When the user's voice input is completed, the electronic device can send all the text information displayed in the content input control to the information recipient. It can be understood that all the text information displayed in the content input control at this time is all the text information converted from the voice data input by the user.

[0136] In one possible implementation, after determining that the user has completed the voice input, the electronic device may immediately and automatically send all text information displayed in the content input control to the information recipient.

[0137] In another possible implementation, after determining that the user has completed voice input, the electronic device may start a timer to count. At the same time, as shown in 19A of Figure 19, the electronic device may display a countdown control 240 in the input interface. The countdown control 240 can be used to indicate the timing duration. The countdown control 240 can be in the form of a progress bar, for example. When the timing duration indicated by the countdown control 240 reaches a set duration threshold, such as 5 seconds, as shown in 19B of Figure 19, the electronic device may automatically send all text information displayed in the content input control to the information recipient. After the sending is completed, the countdown control 240 can stop being displayed.

[0138] In this implementation, after the user completes voice input, the text information displayed in the content input control can be re-edited according to actual needs within the countdown period. During the countdown period, after detecting the user's editing operation, the electronic device can stop the timer, and at the same time, the electronic device can stop displaying the above-mentioned countdown control. The user's editing operation may include, but is not limited to, deleting part or all of the text content in the content input control, pasting or copying text content in the content input control, voice data input, etc. After detecting that the user's editing operation is completed, the electronic device can restart the timer for timing, and can redisplay the above-mentioned countdown control until the timing reaches the set duration threshold, and automatically send all the text information re-edited by the user to the information recipient. Based on this implementation, the user can be provided with a second opportunity to make corrections when the voice-to-text conversion result is inaccurate or the user's voice input is incorrect.

[0139] In another possible implementation, as shown in FIG20A , the content input control may further include a send button 250. After the user completes voice input, the text message converted from the input voice data is displayed in the content input control. At this point, as shown in FIG20B , the electronic device can respond to the user triggering the send button 250 to send the text message in the content input control to the recipient. In this implementation, the user actively controls the timing of sending the text message via the send button 250, allowing the user to edit the input content at any time before sending the text message and control the sending of the text message at any time.

[0140] Furthermore, in an embodiment of the present application, in the input interface, as shown in FIG21 , an emoticon input button 212 may be displayed in the content input control. In response to a triggering operation on the emoticon input button 212, an emoticon list may be displayed. In response to a triggering operation on any emoticon in the emoticon list, the electronic device may display the emoticon in the content input control.

[0141] Based on this implementation method, emoticon options can be provided to users on the basis of text input to better meet users' diverse and interesting input needs.

[0142] Furthermore, in an embodiment of the present application, in an online chat scenario, as shown in FIG. 22A , the content input control displayed on the input interface may also display at least one preset phrase. This preset phrase can be used to reply to a received message. In response to a triggering operation on any preset phrase, as shown in FIG. 22B , the preset phrase can be directly sent to the recipient without the user having to input text. The preset phrase can, for example, be determined based on the content of the received message. Specifically, after determining that a friend message has been received, the electronic device can perform semantic recognition on the message content and, based on the recognition results, generate and display at least one possible preset phrase. Alternatively, the preset phrase can be determined based on the user's historical replies. For example, the user's historical replies can be studied to generate the most frequently used phrases from the historical replies as preset phrases for display. Alternatively, the preset phrases can be randomly assigned based on common idioms used in chat scenarios. This application does not impose any limitations on this.

[0143] Through the above implementation method, preset phrases can be provided in small-screen input scenarios where keyboard input is inconvenient, so that users can directly click on the preset phrase that matches the current conversation to reply without performing any text input operations, thereby improving text input efficiency.

[0144] In another embodiment of the present application, in a content search scenario, a user can, for example, access an application interface through an application icon displayed on the main interface. The application icon can be an icon of any application equipped with a content search function. As shown in 23A of Figure 23 , a content search control 51 can be displayed in the application interface. The content search control 51 can be, for example, a content search bar or a content search icon. In response to a triggering operation on the content search control 51, an input interface can be displayed, as shown in 23B of Figure 23 . Unlike in an online chat scenario, in a content search scenario, the input interface does not include content related to chat messages. Furthermore, similar to the aforementioned embodiment, the electronic device can receive voice data collected by at least two microphones, and then compare the at least two channels of voice data. If it is determined that the audio energy of each channel of voice data meets the preset characteristics, it is considered that the current scenario is a voice input scenario, and the voice-to-text input function is activated. The collected voice data is then automatically converted into text information for input.

[0145] Unlike online chat scenarios, in content search scenarios, once the user completes voice input, there's no need to send a message. Instead, the electronic device searches for matching target objects in the database corresponding to the current application based on the entire text displayed in the content input control. The searched objects are then displayed.

[0146] In one possible implementation, after determining that the user has completed the voice input, the electronic device may immediately and automatically search for the corresponding target object based on all text information displayed in the content input control.

[0147] In another possible implementation, as shown in 24A of FIG. 24 , after determining that the user has completed voice input, the electronic device may start a timer to count down. At the same time, a countdown control 240 may be displayed within the input interface. The countdown control 240 may be used to indicate the duration of the countdown. The countdown control 240 may be in the form of a progress bar, for example. When the duration indicated by the countdown control 240 reaches a set duration threshold, such as 5 seconds, as shown in 24B of FIG. 24 , the electronic device may automatically search for and display the corresponding target object based on all text information displayed within the content input control.

[0148] In this implementation, after the user completes the voice input, it can also be determined within the countdown time whether the text information displayed in the content input control needs to be re-edited. After detecting the user's editing operation, the electronic device can stop the timer and stop displaying the above-mentioned countdown control. The user's editing operation can be, for example, deleting part or all of the text content in the content input control, pasting or copying text content in the content input control, voice data input operation, etc. After detecting that the user's editing operation is completed, the electronic device can redisplay the above-mentioned countdown control and restart the timing until the timing reaches the set duration threshold, and automatically search for the target object that matches the current input content. Based on this implementation, the user can be provided with a second opportunity to make corrections in cases where the user's voice input is incorrect or the voice-to-text conversion result is inaccurate.

[0149] In another possible implementation, as shown in FIG25A , the content input control may further include an OK button 213. After the user completes voice input, text information converted from the input voice data is displayed in the content input control. At this point, as shown in FIG25B , the electronic device may respond to the user's triggering operation on the OK button 213 and search for a target object that matches the text information. In this implementation, the user actively controls the execution of the search through the OK button 213, thereby facilitating the user to perform secondary editing of the input content at any time before triggering the search operation.

[0150] In another embodiment of the present application, the method for triggering the voice input to text function is further described.

[0151] As shown in FIG26 , in an embodiment of the present application, the electronic device may include an Advanced Digital Signal Processing (ADSP) module and an Application Processor (AP). After the electronic device displays an input interface in response to an input method call operation, the ADSP module and the AP may be used to determine whether to activate a speech-to-text input function based on the posture data of the electronic device and the voice data collected by the microphone.

[0152] The ADSP module can be used to execute the first conditional judgment process. Specifically, the ADSP module can receive posture data collected by the IMU and voice data collected simultaneously by at least two microphones. Furthermore, the ADSP module can determine whether the posture of the electronic device meets the preset conditions based on the posture data collected by the IMU. If it is determined that the posture of the electronic device meets the preset conditions, the first conditional judgment is passed. At this point, the ADSP module can send the voice data collected by the at least two microphones to the AP.

[0153] The AP may be used to execute the second condition determination process.

[0154] Specifically, first, the AP can be used to perform security detection on received voice data.

[0155] Security testing may, for example, include voiceprint recognition of voice data, which can be used to authenticate the user. Voiceprint recognition can, for example, compare the voiceprint information of the received voice data with pre-stored voiceprint information. If the comparison is consistent, the verification is considered successful, and the subsequent judgment process can be further executed. If the comparison is inconsistent, it is considered that the current user does not have input permission, the verification fails, and the current process can be terminated. This implementation method can verify the user's identity, which is beneficial for enhancing the security of information input and protecting user data security in specific scenarios (such as transfer information input scenarios).

[0156] In addition to voiceprint recognition, security checks can also include, for example, recording playback attack detection. This feature can be used to detect whether the received voice data is recorded. If the received voice data is determined to be recorded, it indicates that it is not the user's actual voice data. In this case, the security check is considered to have failed, and the current process can be terminated. If the received voice data is determined to be not recorded, the security check is considered to have passed, and the subsequent method process can be executed.

[0157] If the security check passes, the AP can detect the audio energy of the voice data collected by at least two microphones and implement scene recognition based on the characteristics of the audio energy to determine whether to start the voice-to-text input function.

[0158] In one possible implementation, before performing audio energy detection on the voice data collected by at least two microphones, the AP can also be used to perform enhancement processing on the received voice data. Enhancement processing can be used to extract usable signals from voice data and suppress noise interference, thereby improving the quality and clarity of voice data, which is beneficial to improving the accuracy of subsequent audio energy detection results. In an embodiment of the present application, enhancement processing may include: noise recognition and noise suppression. Specifically, environmental noise, noise generated by sound signals reflected from the spatial environment, human voices that are unrelated to the input text, and other noises can be identified. Then, based on the noise recognition results, various types of noise can be suppressed, thereby improving the clarity of the voice signal. Enhancement processing can also include: sound source angle recognition and directional sound pickup. Specifically, by identifying the direction of the sound source, the voice signal consistent with the direction of the sound source can be extracted from the collected voice data, and the voice signals in other directions can be suppressed.

[0159] After the enhancement processing is completed, the AP module can input the processed voice data into the target model. The target model can be used to learn and compare the audio energy features of the voice data collected by different microphones to determine whether the difference in audio energy of the voice data collected by different microphones meets the preset conditions, for example, whether the difference in audio energy is greater than a set threshold. When it is determined that the difference in audio energy meets the preset conditions, it means that the current scene is a voice input scene. At this time, the target model can output a judgment result of yes, which is used to indicate that the voice-to-text input function is started. When it is determined that the difference in audio energy does not meet the preset conditions, it means that the current scene does not meet the voice input scene. At this time, the target model can output a judgment result of no, which is used to indicate that the voice-to-text input function is not started.

[0160] This implementation allows for voice input scenarios to be identified using voice data collected by different microphones, automatically initiating the voice-to-text input function without requiring the user to perform complex wake-up operations. This simplifies text input and facilitates efficient input on small screens. Furthermore, this method performs multiple pre-processing steps on voice data before voice detection, further improving the accuracy of voice detection and scene recognition results.

[0161] It should be understood that the electronic devices here are embodied in the form of functional units. The term "unit" here can be implemented in the form of software and / or hardware, without specific limitation. For example, a "unit" can be a software program, a hardware circuit, or a combination of the two that implements the above functions. The hardware circuit may include an application specific integrated circuit (ASIC), an electronic circuit, a processor (such as a shared processor, a proprietary processor, or a group processor, etc.) and a memory for executing one or more software or firmware programs, a merged logic circuit, and / or other suitable components that support the described functions. Whether a function is executed in hardware or in a computer software-driven hardware manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application in combination with the embodiments.

[0162] The module division in this embodiment is illustrative and represents only one logical functional division. In actual implementation, other divisions may be employed. For example, functional modules may be divided according to their respective functions, or two or more functions may be integrated into a single processing module. These integrated modules may be implemented in hardware.

[0163] An embodiment of the present application also provides an electronic device, which includes a storage medium and a central processing unit. The storage medium can be a non-volatile storage medium, and a computer executable program is stored in the storage medium. The central processing unit is connected to the non-volatile storage medium and executes the computer executable program to implement the above-mentioned text input method.

[0164] The embodiment of the present application also provides a computer-readable storage medium, which stores instructions. When the instructions are executed on a computer, the computer executes the various steps of the text input method of the embodiment of the present application.

[0165] The embodiment of the present application also provides a computer program product containing instructions. When the computer program product is run on a computer or any at least one processor, it enables the computer to execute each step of the text input method of the embodiment of the present application.

[0166] In the embodiments of the present application, "and / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent the existence of A alone, the existence of A and B at the same time, and the existence of B alone. A and B may be singular or plural. The character " / " generally indicates that the previous and subsequent associated objects are in an "or" relationship. "At least one of the following" and similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b and c can be represented by: a, b, c, ab, ac, bc, or abc, where a, b, c can be single or multiple.

[0167] Those skilled in the art will appreciate that the various units and algorithm steps described in the embodiments disclosed herein can be implemented using a combination of electronic hardware, computer software, and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0168] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0169] In the several embodiments provided in this application, if any function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of this application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0170] The above description is merely a specific embodiment of the present application. Any person skilled in the art may easily conceive of variations or substitutions within the technical scope disclosed in this application, and such variations or substitutions shall be within the scope of protection of this application. The scope of protection of this application shall be subject to the scope of protection of the claims.

Claims

1. A text input method, characterized in that: Applied to an electronic device, the electronic device includes at least two microphones, and the distance between the at least two microphones is greater than a set threshold; the method includes: In response to the input method calling operation, an input interface is displayed, wherein the input interface includes a first control, and the first control is used to display input content; respectively acquiring the first voice data and the second voice data collected simultaneously by the respective microphones; detecting that audio energy of the first voice data and the second voice data meets a preset condition; First text information is generated according to the first voice data and the second voice data, and the first text information is displayed in the first control.

2. The method according to claim 1, characterized in that The electronic device further includes a first display screen and a second display screen, wherein the second display screen and the first display screen are located on different surfaces of the electronic device, the display area of ​​the second display screen is smaller than that of the first display screen, and the first display screen includes at least two sub-display parts and is foldable; The distance between the at least two microphones is greater than a set threshold, comprising: In the folded state, the unfolded state, and the intermediate state of the first display screen, the distance between the at least two microphones is greater than a set threshold.

3. The method according to claim 2, characterized in that The method further comprises: Determine whether the electronic device is in a target display scene.

4. The method according to claim 3, characterized in that The target display scene includes a scene in which the second display screen is in a display state.

5. The method according to claim 3, characterized in that The target display scene includes a scene in which the first display screen is in a split-screen display state or a scene in which the first display screen is in a floating window display state.

6. The method according to claim 1, characterized in that The input interface further includes a second control. When the second control is in the first display state, the input interface does not include a keyboard control.

7. The method according to claim 6, characterized in that The method further comprises: When the second control is in the first display state, in response to a triggering operation on the second control, the second control switches to the second display state; when the second control is in the second display state, the input interface includes the keyboard control.

8. The method according to claim 7, characterized in that The method further comprises: When the second control is in the first display state for the first time, the input interface further includes a third control, and the third control is used to display input mode prompt information.

9. The method according to claim 8, characterized in that The method further comprises: In response to a triggering operation on the third control, the third control is hidden.

10. The method according to claim 1, characterized in that The input interface further includes a fourth control, and the method further includes: In response to a long press operation on the fourth control, acquiring third voice data collected by any one of the microphones; In response to releasing the fourth control, the third voice data is sent to a target receiving end.

11. The method according to claim 1, wherein Before respectively acquiring the first voice data and the second voice data collected simultaneously by the microphones, the method further includes: Acquiring first posture data collected by an inertial measurement unit; According to the first posture data, it is determined that the posture feature of the electronic device is consistent with the preset feature.

12. The method according to claim 1, characterized in that The detecting that the audio energy of the first voice data and the second voice data meets a preset condition includes: It is detected that the audio energy difference between the first voice data and the second voice data is greater than a preset value.

13. The method according to claim 12, characterized in that Before detecting that the audio energy of the first voice data and the second voice data meets a preset condition, the method further includes: The first voice data and the second voice data are preprocessed, where the preprocessing includes security detection and / or voice enhancement.

14. The method according to claim 1, wherein After displaying the first text information in the first control, the method further includes: Upon detecting that the collection of the first voice data and the second voice data is complete, starting a first timer; In response to the timing duration of the first timer exceeding a duration threshold, the input content displayed in the first control is sent to a target receiving end.

15. The method according to claim 1, wherein After displaying the first text information in the first control, the method further includes: Upon detecting that the collection of the first voice data and the second voice data is complete, starting a second timer; In response to the timing duration of the second timer exceeding the duration threshold, a target object is searched in the target database according to the input content displayed in the first control, and the searched target object is displayed.

16. The method according to claim 14 or 15, characterized in that The method further comprises: A fifth control is displayed on the input interface, and the fifth control is used to describe the timing duration.

17. The method according to claim 1, wherein The input interface further includes at least one preset phrase; After the input interface is displayed, the method further includes: In response to a triggering operation on any one of the preset phrases, the preset phrase is sent to a target receiving end.

18. An electronic device, characterized in that: include: one or more processors; Memory; and one or more computer programs, wherein the one or more computer programs are stored in the memory, and the one or more computer programs include instructions, which, when executed by the device, cause the device to perform the method according to any one of claims 1 to 17.

19. A storage medium, characterized in that: The storage medium stores program instructions, which, when executed on an electronic device, enable the electronic device to execute the method according to any one of claims 1 to 17.

Citation Information

Patent Citations

  • Input method and device

    CN106814879A

  • Input method, device and equipment and machine readable medium

    CN111984129A

  • Voice interaction function awakening method and electronic equipment

    CN117119102A

  • Voice input method and device, terminal equipment and storage medium

    CN117331473A

  • Electronic device and method for executing function using speech recognition thereof

    EP3160150A1