Audio input method, electronic device, storage medium, and computer program product
By combining the first and second operations of the audio input control, the input mode is determined in real time, solving the problem of single output for user voice input, realizing diversified audio content output, and improving user operation efficiency.
Patent Information
- Application Number
- PCT/CN2025/083327
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-22
- Filing Date
- 2025-03-19
- Publication Date
- 2026-01-29
AI Technical Summary
In existing technologies, the content input by users via voice can only be output in a single mode, resulting in low operational efficiency and difficulty in meeting the diverse input needs of users.
By combining the first and second operations of the audio input controls, the input mode is determined in real time, enabling diverse output of audio content. The first operation, such as a long press, and the second operation, such as a swipe gesture, allow switching input modes during audio acquisition, directly outputting or identifying the user's voice content.
It improves the diversity and ease of use of user input, reduces subsequent editing operations, and enhances user input efficiency.
Smart Images

Figure CN2025083327_29012026_PF_FP_ABST
Abstract
Description
Audio input methods, electronic devices, storage media, and computer program products
[0001] Cross-references to related applications
[0002] This application is based on and claims priority to Chinese application No. 202410986525.0, filed on July 22, 2024, the disclosure of which is incorporated herein by reference in its entirety. Technical Field
[0003] This disclosure relates to the field of computer technology, and in particular to an audio input method, electronic device, storage medium, and computer program product. Background Technology
[0004] With the development of speech recognition and internet technologies, users can more conveniently replace keyboard input with voice input, thereby improving input efficiency. For example, when using various applications, users can trigger an audio input control and speak what they want to input. After the voice input ends, the user's speech will be converted into text and displayed on the interface. Summary of the Invention
[0005] This summary section is provided to briefly introduce the concepts, which will be described in detail in the detailed description section below. This summary section is not intended to identify key or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0006] According to some embodiments of this disclosure, an audio input method is provided, comprising: acquiring audio in response to a first operation by a user on an audio input control; determining an input mode based on a second operation by the user during the first operation on the audio input control, wherein different input modes correspond to different output modes; and outputting content corresponding to the audio according to the output mode corresponding to the input mode.
[0007] According to some embodiments of the present disclosure, an electronic device is provided, including: a memory; and a processor coupled to the memory, the processor being configured to execute an audio input method of any embodiment of the present disclosure based on instructions stored in the memory.
[0008] According to some embodiments of the present disclosure, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, performs an audio input method according to any embodiment of the present disclosure.
[0009] According to some embodiments of the present disclosure, a computer program product is provided that, when the computer program product is run on a computer, causes the computer to execute the audio input method of any embodiment described in the present disclosure when it is executed.
[0010] Other features, aspects, and advantages of this disclosure will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. Attached Figure Description
[0011] Preferred embodiments of the present disclosure are described below with reference to the accompanying drawings. The accompanying drawings, which are included to provide a further understanding of the present disclosure, and which, together with the following detailed description, are incorporated in and form a part of this specification and are used to explain the present disclosure. It should be understood that the drawings described below only relate to some embodiments of the present disclosure and are not intended to limit the present disclosure. In the drawings:
[0012] Figure 1 shows a schematic flowchart of an audio input method according to some embodiments of the present disclosure.
[0013] Figures 2A and 2B show schematic diagrams of user interfaces according to some embodiments of the present disclosure.
[0014] Figure 3 shows a flowchart of an input pattern determination method according to some embodiments of the present disclosure.
[0015] Figure 4 shows a schematic diagram of the structure of an audio input device according to some embodiments of the present disclosure.
[0016] Figure 5 shows a schematic diagram of the structure of an electronic device according to some embodiments of the present disclosure.
[0017] Figure 6 shows a schematic diagram of the structure of a computer system according to some embodiments of the present disclosure.
[0018] It should be understood that, for ease of description, the dimensions of the various parts shown in the accompanying drawings are not necessarily drawn to actual scale. The same or similar reference numerals are used in the various drawings to denote the same or similar parts. Therefore, once an item is defined in one drawing, it may not be discussed further in subsequent drawings. Detailed Implementation
[0019] The technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. However, it is obvious that the described embodiments are only some embodiments of this disclosure, and not all embodiments. The following description of the embodiments is merely illustrative and is in no way intended to limit this disclosure or its application or use. It should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein.
[0020] It should be understood that the various steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect. Unless otherwise specifically stated, the relative arrangement, numerical expressions, and values of components and steps set forth in these embodiments should be interpreted as merely exemplary and do not limit the scope of this disclosure.
[0021] As used in this disclosure, the term "comprising" and its variations are open-ended terms that include at least the following elements / features but do not exclude other elements / features, i.e., "including but not limited to". Furthermore, as used in this disclosure, the term "including" and its variations are open-ended terms that include at least the following elements / features but do not exclude other elements / features, i.e., "including but not limited to". Therefore, "comprising" and "including" are synonymous. The term "based on" means "at least partially based on".
[0022] Throughout this specification, the terms "one embodiment," "some embodiments," or "embodiment" mean that a specific feature, structure, or characteristic described in connection with an embodiment is included in at least one embodiment of the invention. For example, the term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; and the term "some embodiments" means "at least some embodiments." Furthermore, the appearance of the phrases "in one embodiment," "in some embodiments," or "in an embodiment" in various places throughout the specification does not necessarily refer to the same embodiment, but may refer to the same embodiment.
[0023] It should be noted that the concepts of "first," "second," etc., used in this disclosure are used only to distinguish different devices, modules, or units, and are not intended to define the order of functions performed by these devices, modules, or units or their interdependencies. Unless otherwise specified, the concepts of "first," "second," etc., are not intended to imply that the objects described herein must be in a given temporal, spatial, rank, or any other given order.
[0024] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0025] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0026] The embodiments of this disclosure are described in detail below with reference to the accompanying drawings; however, this disclosure is not limited to these specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. Furthermore, in one or more embodiments, specific features, structures, or characteristics can be combined in any suitable manner that will be apparent to those skilled in the art from this disclosure.
[0027] Currently, voice input is commonly used in chat scenarios. The text corresponding to a user's spoken content is presented to the other party in the chat as if it were spoken. For example, when user A and user B are chatting, whether user A says "Have you eaten?" or "smile" through the audio input control, the chat interface will directly display the text of the message "Have you eaten?" or "smile". Although "smile" is an emoticon that user A wants to send to user B, or user A's current mood, the text of "smile" still appears as a user dialogue, looking no different from "Have you eaten?". That is, in related technologies, the content input by the user through voice is output in a single mode.
[0028] As users' expressive content and methods become increasingly diverse, a single output mode is no longer sufficient to meet their needs. Therefore, whether it's a social media app, a search app, or other types of applications, users can output in multiple ways. For example, they can input text, voice, emojis, images, videos, etc., or they can trigger input through different controls. These inputs obtained through different methods will be displayed on the interface accordingly.
[0029] However, for content input by users through the same control, only one output method is still supported. For example, in applications with dialogue functionality, besides users directly sending voice messages, when users input through audio input controls, only voice can be recognized as text and displayed directly on the interface. If the user's speech is merely an expression of their current mental activity or a verbal description of an action, the user needs to manually add symbols to the corresponding content using the keyboard. For example, if a user says "Hello, smile," the interface will directly display "Hello, smile." The user needs to add parentheses around the smile to indicate that "smiling" is their current mental activity. For example, the user manually edited the content as "Hello (smile)." Alternatively, the user says "Hello" to convert it to text, and then manually adds a smiling emoticon or sends an image representing a smile. Therefore, for content input by voice, users need to perform further complex operations to adjust the corresponding text to different display modes. This results in relatively low user efficiency.
[0030] Therefore, one of the technical problems that this disclosure aims to solve is: how to improve operational efficiency while increasing the diversity of user input.
[0031] To address the aforementioned problems, this disclosure provides an audio input method. During the acquisition of a user's voice or other audio data, the user can control the input mode through additional operations to ensure that the audio is output in the user's desired manner. An embodiment of the audio input method of this disclosure is described below with reference to Figure 1.
[0032] Figure 1 shows a schematic flowchart of an audio input method according to some embodiments of the present disclosure. As shown in Figure 1, the audio input method of this embodiment includes steps S102 to S106.
[0033] In step S102, audio is acquired in response to the user's first operation on the audio input control.
[0034] Audio input controls are used to trigger audio capture, such as a user's voice or singing. They can be controls on the application's interface or physical controls on the user's device. Users can trigger these controls through initial actions such as clicking or long-pressing. Users can also trigger the audio input control through specified gestures or actions, such as shaking the device or performing a specific gesture on the screen, to initiate audio capture.
[0035] In some embodiments, the first operation is a long press operation. A long press operation can prevent accidental recordings caused by user mis-touch or friction from the user's device being in their bag. While the user maintains the long press, the audio to be recorded can be continuously captured; when the user ends the long press, the audio recording is considered complete.
[0036] The embodiments of this disclosure support the scenario where the user performing the first operation and the speaking user are the same user; that is, in response to the user's first operation on the audio input control, the user's audio is captured. Furthermore, it also supports the scenario where the user performing the first operation and the speaking user are different users; that is, in response to the first user's first operation on the audio input control, the audio of a second user is captured. In practical applications, there may be situations where user A performs the first operation and, with user B's consent, user B's audio is captured. This facilitates user A assisting user B, who is currently unable to perform the operation, in capturing audio.
[0037] In step S104, during the process of the user performing the first operation on the audio input control, the input mode is determined according to the user's second operation, wherein different input modes correspond to different output modes.
[0038] In other words, during audio capture, the user can determine the input mode through a second operation. For example, in response to the second operation, the current input mode can be switched to another input mode, such as switching from the first input mode to the second input mode, or vice versa. In this way, the user can switch the input mode by performing the second operation during an audio capture process, eliminating the need for manual editing or modification after the capture is complete. This allows the user to determine the output mode of their speech while speaking, improving operational efficiency.
[0039] The second operation can be implemented in various ways. These include operations on controls within the interface, gesture operations (such as swipe gestures), operations on physical controls of the user device, and operations on the user device as a whole (such as shaking, moving, rotating, etc.). For swipe gestures, the direction can be arbitrary or along a specified direction.
[0040] In some embodiments, the first operation is a long press operation, and the second operation is a swipe gesture operation. For example, a user triggers audio capture by long-pressing an audio input control, and then can simultaneously perform a swipe gesture operation during the long press. That is, the finger that performs the long press by swiping can simultaneously maintain the long press and perform the swipe gesture, offering high operability. In this way, users can more efficiently determine the input mode while speaking.
[0041] Of course, the first and second operations can also be combined in other ways. For example, a user can tilt the device while pressing and holding the audio input control, or tap the screen while shaking the device to trigger audio capture, etc. This disclosure does not limit this.
[0042] Input modes are used to distinguish the audio input from the user, so that the audio can be converted into content corresponding to the input mode and output in the corresponding output mode. There are multiple input modes, and the number can be set as needed. Each input mode corresponds one-to-one with a different output mode.
[0043] In some embodiments, the input modes include a first input mode and a second input mode. The first input mode is used to input dialogue content, and the second input mode is used to input non-dialogue content. Non-dialogue content includes at least one of text, emoticons, and images. When the non-dialogue content is text, it describes at least one of the user's psychology, emotions, actions, behaviors, etc. As needed, there can be more input modes, or other types of input modes. For example, multiple input modes can be tailored to multiple specified objects; the first input mode is for speaking to object A, the second input mode is for speaking to object B, the third input mode is for speaking to object C, and so on.
[0044] The input mode is determined based on the user's second action, which can be based on whether the second action is performed. Furthermore, the second action may include multiple types or have multiple parameters, and the input mode is determined based on these attributes. For example, when the second action is a swipe gesture, swiping in different directions may correspond to the same input mode or different input modes.
[0045] During a single audio capture operation, the user can perform one or more secondary operations. This allows for one or more switching of input modes during a single audio capture, thereby increasing the versatility of the user's output and simplifying operation.
[0046] In step S106, the content corresponding to the audio is output according to the output mode corresponding to the input mode.
[0047] Output modes describe the type, display style, and other characteristics of the content corresponding to the captured audio, allowing for differentiation in output based on different input modes. The content corresponding to the audio can include at least one of the following: text, emoticons, images, sounds, and control information. For example, audio can be first converted to text, and then the text can be converted to content corresponding to the output mode. Control information can be information controlling application functions or objects within the application, such as controlling virtual objects in the application to perform operations corresponding to the content in the audio (e.g., actions, emoticons, appearance, etc.).
[0048] The first output mode is used to directly display the text corresponding to the audio, and the second output mode is used to display the text corresponding to the audio identified by a specified symbol, or to display non-text information associated with the text corresponding to the audio.
[0049] For example, in the audio content, content corresponding to the second input mode is identified using a specified symbol, while content corresponding to the first input mode is not identified using a symbol. Suppose the text corresponding to the captured audio is "Hello, smile," and based on the user's second operation, it is determined that "Hello" is the first output mode and "smile" is the second output mode. Then, the output content could be the text "Hello (smile)." Of course, in some embodiments, the output content could also be the text "Hello" and a smiling emoticon, or it could display the text "Hello" and control a virtual object in the current interface to perform a smiling emoticon or display a smiling image of the virtual object. Other output modes can also be used as needed, which will not be elaborated here.
[0050] During audio capture, the user's speech can be displayed on the interface in real time in the corresponding output mode. That is, the content corresponding to the user's audio can be output before the audio capture ends. Of course, the corresponding content can also be displayed on the interface after the audio capture ends.
[0051] Through the above embodiments, during the process of acquiring the user's audio using the first operation, the user can control the input mode through the second operation to output the audio in an output mode corresponding to the input mode. Thus, while supporting the diversity of user outputs, the ease of user operation is improved. Therefore, the embodiments of this disclosure can improve user input efficiency.
[0052] In some types, the location of the first operation overlaps with the audio input control. In some embodiments, the initial input mode is determined based on the triggering area of the first operation performed by the user within the audio input control. That is, the first input mode after the user triggers the audio input control. For example, in response to the triggering location of the first operation being a first area of the audio input control, the initial input mode is determined as a first input mode; in response to the triggering location of the first operation being a second area of the audio input control, the initial input mode is determined as a second input mode, and so on.
[0053] Of course, a default initial input mode can also be set. That is, regardless of where the user triggers the audio input control, a certain input mode can be set as the initial input mode. Those skilled in the art can choose according to their needs.
[0054] The following example describes how to determine the input pattern based on the user's second action.
[0055] In some embodiments, during the process of a user performing a first operation on an audio input control, in response to a second operation performed by the user being a specified operation, the current input mode is switched to another input mode.
[0056] The second operation can include one or more operations, such as multiple sub-operations, multiple types of operations, or operations with multiple attributes. It is possible to pre-set which specified operations can trigger the switching of the current input mode. For example, the switching can be triggered in response to a user performing a swipe gesture in any direction, or only when the user performs a swipe gesture in a specified direction, or only when the user performs a swipe gesture in a specified direction and the current input mode is the specified input mode.
[0057] For example, in response to the current input mode being the first input mode and the second operation performed by the user being the first specified operation, the current input mode is switched to the second input mode. If the current mode is used for inputting dialogue content, then in response to the second operation performed by the user being the first specified operation, the input mode is switched to a mode for inputting non-dialogue content.
[0058] In addition, the current second input mode can be switched back to the first input mode. In some embodiments, the current input mode is switched back to the first input mode in response to the current input mode being the second input mode and the second operation performed by the user being a second specified operation.
[0059] In one example, the first specified operation and the second specified operation are swipe gestures with different swipe directions. An embodiment of switching between different input modes via swipe gestures is described below with reference to Figures 2A and 2B.
[0060] Figures 2A and 2B illustrate schematic diagrams of user interfaces according to some embodiments of the present disclosure. As shown in Figure 2A, an audio input control 21 is provided in interface 2. The user can trigger audio acquisition by long-pressing the control 21 and end audio acquisition when the long-press operation ends.
[0061] The user presses and holds control 21 to begin speaking, as shown in Figure 2B. The initial input mode is set to the first input mode. Then, the user performs a swipe gesture (e.g., swipe right) to switch the input mode to the second input mode. Afterward, the user performs a swipe gesture in a different direction (e.g., swipe left) to switch the input mode back to the first input mode. The user can perform any number of gestures during the speaking process. Finally, the user ends the press and hold, thus ending the audio capture process.
[0062] In some embodiments, a user can also swipe right to switch the input mode from the first input mode to the second input mode, and then swipe left to switch back to the first input mode. That is, swiping in either direction will switch from the first mode to the second input mode, and then performing the reverse swipe will switch back from the second input mode to the first input mode.
[0063] When displaying an audio input control, the audio input control can have multiple display styles. For example, the style of the audio input control can be different in different input modes. Taking Figure 2 as an example, control 21 can display different colors in different input modes to indicate to the user the current input mode. If needed, text prompts can also be displayed on the interface to clearly inform the user of the current input mode, or what operation can be used to switch input modes, etc.
[0064] In some embodiments, a correspondence between a specified area of control 21 and an initial input mode can also be set. For example, when a user triggers control 21 through a first operation, if the user touches the middle part of control 21, the first input mode can be used as the initial input mode. If the user initially touches either side of control 21, the initial input mode can be determined as the second input mode. Alternatively, the user can be set to trigger a default input mode, such as the first input mode, whenever any area of control 21 is touched, and then switch the input mode in response to the user's second operation, such as switching to the second input mode.
[0065] After the audio acquisition is completed, an example of the output result is shown in the display control 22 of Figure 2B. In this example, the dialogue content that the user wants to express is displayed on the interface without any symbols, while the actions or mental activities that the user wants to express are displayed in parentheses.
[0066] The following example, using the content in control 22, illustrates the user's speaking and operation process. After long-pressing control 21, the user can say the following and perform the following actions, forming the following speaking-action sequence, where the user's actions are represented by "<>": <swipe left> - smile - <swipe right> - You're back? I'm cooking, I'll talk to you in a bit - <swipe left> - turn around and go back to cooking. The first left swipe switches the input mode to non-dialogue input mode, and the user's "smile" is then identified as non-dialogue content. The first right swipe switches the input mode to dialogue input mode, and the user's "You're back? I'm cooking, I'll talk to you in a bit" is then identified as dialogue content. The second left swipe switches the input mode back to non-dialogue input mode, and the user's "turn around and go back to cooking" is then identified as non-dialogue content. Then, the user ends the long-press operation. The above speech determines the corresponding output content based on the input pattern of each part. In Figure 2B, an exemplary output result is "(smiling) You're back? I'm cooking, I'll talk to you in a bit (turns around to cook)", that is, the dialogue content is displayed directly, and the non-dialogue content is displayed in parentheses.
[0067] Some embodiments of this disclosure support converting user audio during a single audio acquisition process into content corresponding to multiple input modes. That is, the user's audio is divided into multiple segments, each involving a different output mode. In this case, the output mode of which portion of the audio should be determined based on the time the user performs one or more second operations. Embodiments of the input mode determination method of this disclosure are described below with reference to FIG3.
[0068] Figure 3 shows a flowchart of an input pattern determination method according to some embodiments of the present disclosure. As shown in Figure 3, the input pattern determination method of this embodiment includes steps S302 to S304.
[0069] In step S302, during the process of the user performing the first operation on the audio input control, the audio or the content corresponding to the audio is divided into multiple segments according to the execution time of the user's second operation.
[0070] This execution time can be directly used as the segmentation point for dividing the audio. Alternatively, the segmentation point for dividing the audio can be determined by referring to the execution time of the second operation, but it does not have to be exactly the same.
[0071] Considering that the user's actions and the time point when the user wants to switch while speaking may not completely coincide in time, in some embodiments, the segmentation point can be determined more accurately by using the time when the user performs the second operation and the semantics corresponding to the audio.
[0072] For example, audio or its corresponding content can be divided into multiple units based on its semantics. For instance, the audio can be converted to text and then input into a text processing model to divide the text into multiple units. Each unit can be a sentence, clause, phrase, word, etc. Alternatively, the user's audio can be directly input into an audio processing model, and the segmented result output by the model can be obtained directly. Then, based on the execution time of the user's second operation, the multiple units can be divided into multiple paragraphs.
[0073] For example, when a user says, "Smile, long time no see, how have you been?", after saying "smile," there's a slight delay in the action; the second action is performed only after saying "okay." If segmented strictly by time, "smile, okay" would be identified as one output pattern, and "long time no see, how have you been?" as another. However, by using semantic segmentation, treating "smile" as one unit and "long time no see" as another, the segmentation point is determined between "smile" and "long time no see." This allows for a more accurate identification of the user's input intent.
[0074] In step S304, the input mode for each paragraph is determined based on the input mode corresponding to the user's second operation.
[0075] For example, the second operation closest to the start of each paragraph can be associated with that paragraph, and the input pattern after performing that second operation can be determined as the input pattern for that paragraph.
[0076] Then, based on the output mode corresponding to the input mode of each segment, the content corresponding to the audio can be output.
[0077] The above embodiments divide the audio or its corresponding content into segments based on the time the user performs the second operation, and then determine the input mode for each segment. This allows for the use of multiple input and output modes during a single audio acquisition process, enabling users to freely switch between input modes while speaking, thus improving the richness of audio input and the ease of operation.
[0078] The embodiments of the methods disclosed herein have been described above. The apparatus for performing the above methods is further described below.
[0079] Figure 4 shows a schematic diagram of the structure of an audio input device according to some embodiments of the present disclosure. As shown in Figure 4, the audio input device 4 of this embodiment includes: a acquisition module 401 configured to acquire audio in response to a first operation by a user on an audio input control; a determination module 402 configured to determine an input mode based on a second operation by the user during the first operation on the audio input control, wherein different input modes correspond to different output modes; and an output module 403 configured to output content corresponding to the audio according to the output mode corresponding to the input mode.
[0080] In some embodiments, the first operation is a long press operation.
[0081] In some embodiments, the second operation includes a swipe gesture.
[0082] In some embodiments, the second operation includes a swipe gesture operation in a specified direction.
[0083] In some embodiments, the determining module 402 is further configured to: determine an initial input mode based on the trigger area in the audio input control of the first operation performed by the user.
[0084] In some embodiments, the determining module 402 is further configured to: during the process of the user performing a first operation on the audio input control, in response to the second operation performed by the user being a specified operation, switch the current input mode to another input mode.
[0085] In some embodiments, the determining module 402 is further configured to: switch the current input mode to the second input mode in response to the current input mode being the first input mode and the second operation performed by the user being the first specified operation.
[0086] In some embodiments, the determining module 402 is further configured to: switch the current input mode to the first input mode in response to the current input mode being the second input mode and the second operation performed by the user being the second specified operation.
[0087] In some embodiments, the first specified operation and the second specified operation are swipe gesture operations with different swipe directions.
[0088] In some embodiments, the determining module 402 is further configured to: during the process of the user performing a first operation on the audio input control, divide the audio or the content corresponding to the audio into multiple segments according to the execution time of the user's second operation; and determine the input mode of each segment according to the input mode corresponding to the user's second operation.
[0089] In some embodiments, the determining module 402 is further configured to: divide the audio or the content corresponding to the audio into multiple units according to the semantics of the audio; and divide the multiple units into multiple paragraphs according to the execution time of the user's second operation.
[0090] In some embodiments, the output module 403 is further configured to output content corresponding to the audio according to an output mode corresponding to the input mode of each paragraph.
[0091] In some embodiments, the output modes include a first output mode and a second output mode; the first output mode is used to directly display the text corresponding to the audio; the second output mode is used to display the text corresponding to the audio identified by a specified symbol, or to display non-text information associated with the text corresponding to the audio.
[0092] In some embodiments, the input mode includes a first input mode and a second input mode, wherein the first input mode is used to input dialogue content and the second input mode is used to input non-dialogue content.
[0093] In some embodiments, non-dialogue content includes at least one of text, emoticons, and images.
[0094] In some embodiments, the output module 403 is further configured to: in the audio content, the content corresponding to the second input mode is identified using a specified symbol, while the content corresponding to the first input mode is not identified using a symbol.
[0095] In some embodiments, the audio input device 40 further includes a display module 404 configured to display an audio input control, wherein the style of the audio input control differs in different input modes.
[0096] It should be noted that the above-described units are merely logical modules divided according to their specific functions, and are not intended to limit the specific implementation method. For example, they can be implemented in software, hardware, or a combination of both. In actual implementation, the above-described units can be implemented as independent physical entities, or they can be implemented by a single entity (e.g., a processor (CPU or DSP, etc.), integrated circuit, etc.). Furthermore, the units shown in the accompanying drawings with dashed lines indicate that these units may not actually exist, and the operations / functions they perform can be implemented by the processing circuitry itself.
[0097] In addition, although not shown, the device may also include a memory that can store various information generated by the device and its constituent units during operation, programs and data used for operation, data to be transmitted by the communication unit, etc. The memory can be volatile memory and / or non-volatile memory. For example, the memory may include, but is not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), read-only memory (ROM), and flash memory. Of course, the memory may also be located outside the device. Optionally, although not shown, the device may also include a communication unit that can be used to communicate with other devices. In one example, the communication unit can be implemented in a manner known in the art, such as including communication components such as antenna arrays and / or radio frequency links, various types of interfaces, communication units, etc. These will not be described in detail here. Furthermore, the device may also include other components not shown, such as radio frequency links, baseband processing units, network interfaces, processors, controllers, etc. These will not be described in detail here.
[0098] Some embodiments of this disclosure also provide an electronic device. Figure 5 shows a schematic diagram of the structure of an electronic device according to some embodiments of this disclosure. For example, in some embodiments, the electronic device 5 can be various types of devices, such as mobile terminals including but not limited to mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. For example, the electronic device 5 may include a display panel for displaying data and / or execution results utilized in the scheme according to this disclosure. For example, the display panel can be of various shapes, such as a rectangular panel, an elliptical panel, or a polygonal panel. In addition, the display panel can be not only a planar panel, but also a curved panel, or even a spherical panel.
[0099] As shown in FIG. 5, the electronic device 5 of this embodiment includes a memory 51 and a processor 52 coupled to the memory 51. It should be noted that the components of the electronic device 5 shown in FIG. 5 are merely exemplary and not limiting; the electronic device 5 may also have other components depending on the actual application requirements. The processor 52 can control other components in the electronic device 5 to perform desired functions.
[0100] In some embodiments, memory 51 is used to store one or more computer-readable instructions. When processor 52 executes the computer-readable instructions, the computer-readable instructions are executed by processor 52 to implement the method according to any of the above embodiments. For specific implementations and related explanations of the various steps of the method, please refer to the above embodiments; repeated details will not be elaborated here.
[0101] For example, processor 52 and memory 51 can communicate with each other directly or indirectly. For example, processor 52 and memory 51 can communicate via a network. The network can include wireless networks, wired networks, and / or any combination of wireless and wired networks. Processor 52 and memory 51 can also communicate with each other via a system bus, which is not limited in this disclosure.
[0102] For example, processor 52 can be embodied in various suitable processors, processing devices, such as central processing unit (CPU), graphics processing unit (GPU), network processor (NP), etc.; it can also be digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The central processing unit (CPU) can be an x86 or ARM architecture, etc. For example, memory 51 can include any combination of various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Memory 51 can include, for example, system memory, which stores, for example, the operating system, application programs, boot loader, database, and other programs. Various application programs and various data can also be stored in the storage medium.
[0103] Furthermore, according to some embodiments of this disclosure, various operations / processes according to this disclosure, implemented via software and / or firmware, can install programs constituting the software from a storage medium or network onto a computer system with a dedicated hardware architecture, such as the computer system 60 shown in FIG. 6. When various programs are installed, the computer system is capable of performing various functions, including those described above. FIG. 6 shows a schematic diagram of the structure of a computer system according to some embodiments of this disclosure.
[0104] In Figure 6, the Central Processing Unit (CPU) 601 performs various processes according to a program stored in the Read-Only Memory (ROM) 602 or a program loaded from the storage portion 608 into the Random Access Memory (RAM) 603. The RAM 603 also stores data required as needed when the CPU 601 performs various processes, etc. The CPU is merely exemplary and can also be other types of processors, such as the various processors described above. The ROM 602, RAM 603, and storage portion 608 can be various forms of computer-readable storage media, as described below. It should be noted that although the ROM 602, RAM 603, and storage device 608 are shown separately in Figure 6, one or more of them may be combined or located in the same or different memories or storage modules.
[0105] CPU 601, ROM 602 and RAM 603 are interconnected via bus 604. Input / output interface 605 is also connected to bus 604.
[0106] The following components are connected to the input / output interface 605: input section 606, such as a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output section 607, including displays such as cathode ray tube (CRT), liquid crystal display (LCD), speakers, vibrators, etc.; storage section 608, including hard disks, magnetic tapes, etc.; and communication section 609, including network interface cards such as LAN cards, modems, etc. The communication section 609 allows communication processing to be performed via a network such as the Internet. It is readily understood that although the various devices or modules in the computer system 60 shown in Figure 6 communicate via bus 604, they can also communicate via a network or other means, wherein the network can include wireless networks, wired networks, and / or any combination of wireless and wired networks.
[0107] As needed, drive 610 is also connected to input / output interface 605. Removable media 611, such as disks, optical disks, magneto-optical disks, semiconductor memories, etc., are installed on drive 610 as needed, so that computer programs read from them can be installed into storage section 608 as needed.
[0108] When the above series of processes are implemented through software, the program constituting the software can be installed from a network such as the Internet or from a storage medium such as removable media 611.
[0109] According to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 609, or installed from a storage device 608, or installed from a ROM 602. When the computer program is executed by the CPU 601, it performs the functions defined in the methods of embodiments of this disclosure.
[0110] It should be noted that, in the context of this disclosure, a computer-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to, an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, capable of transmitting, propagating, or transmitting a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0111] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0112] In some embodiments, a computer program is also provided, comprising: instructions that, when executed by a processor, cause the processor to perform the method of any of the above embodiments. For example, the instructions may be embodied in computer program code.
[0113] In embodiments of this disclosure, computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof. These programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network (including a local area network (LAN) or a wide area network (WAN)), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0114] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0115] The modules, components, or units described in the embodiments of this disclosure can be implemented in software or hardware. The names of the modules, components, or units do not necessarily constitute a limitation on the module, component, or unit itself.
[0116] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary hardware logic components that can be used include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.
[0117] The above description is merely an embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0118] Many specific details are set forth in the description provided herein. However, it is understood that embodiments of the invention may be practiced without these specific details. In other instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of the description.
[0119] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0120] While specific embodiments of this disclosure have been described in detail by way of example, those skilled in the art should understand that the examples are for illustrative purposes only and not intended to limit the scope of this disclosure. Those skilled in the art should understand that modifications can be made to the above embodiments without departing from the scope and spirit of this disclosure. The scope of this disclosure is defined by the appended claims.
Claims
1. An audio input method, comprising: collecting audio in response to a first operation of a user on an audio input control; determining an input mode according to a second operation of the user during the first operation of the user on the audio input control, wherein different input modes correspond to different output modes; outputting content corresponding to the audio according to an output mode corresponding to the input mode.
2. The audio input method of claim 1, wherein, The first operation is a long press operation.
3. The audio input method of claim 1 or 2, wherein, The second operation includes a sliding gesture operation.
4. The audio input method of claim 3, wherein, The second operation includes a sliding gesture operation in a specified direction. 5.The audio input method of any one of claims 1 to 4, further comprising: determining an initial input mode according to a trigger area of the first operation of the user on the audio input control.
6. The audio input method of any one of claims 1 to 5, wherein, The determining an input mode according to a second operation of the user during the first operation of the user on the audio input control includes: switching a current input mode to another input mode in response to the second operation of the user being a specified operation during the first operation of the user on the audio input control.
7. The audio input method of claim 6, wherein, The switching a current input mode to another input mode in response to the second operation of the user being a specified operation includes: switching the current input mode to a second input mode in response to the current input mode being a first input mode and the second operation of the user being a first specified operation.
8. The audio input method of claim 7, wherein, The switching a current input mode to another input mode in response to the second operation of the user being a specified operation further includes: switching the current input mode to the first input mode in response to the current input mode being the second input mode and the second operation of the user being a second specified operation.
9. The audio input method of claim 8, wherein, The first specified operation and the second specified operation are sliding gesture operations with different sliding directions.
10. The audio input method of any one of claims 1 to 9, wherein, The determining an input mode according to a second operation of the user during the first operation of the user on the audio input control includes: dividing the audio or content corresponding to the audio into multiple paragraphs according to an execution time of the second operation of the user during the first operation of the user on the audio input control; determining an input mode of each paragraph according to an input mode corresponding to the second operation of the user.
11. The audio input method of claim 10, wherein, The dividing the audio or content corresponding to the audio into multiple paragraphs according to an execution time of the second operation of the user includes: dividing the audio or content corresponding to the audio into multiple units according to semantics of the audio; dividing the multiple units into multiple paragraphs according to the execution time of the second operation of the user.
12. The audio input method of claim 10 or 11, wherein, The outputting content corresponding to the audio according to an output mode corresponding to the input mode includes: outputting content corresponding to the audio according to an output mode corresponding to an input mode of each paragraph. 13.The audio input method of any one of claims 1 to 12, wherein: the output mode includes a first output mode and a second output mode; The first output mode is used to directly display text corresponding to the audio. The second output mode is used to display text corresponding to the audio using a specified symbol to identify, or to display non-text information associated with the text corresponding to the audio.
14. The audio input method of claim 1 or 13, wherein, The input mode includes a first input mode and a second input mode, the first input mode is used to input conversation content, and the second input mode is used to input non-conversation content.
15. The audio input method of claim 14, wherein, The non-conversation content includes at least one of text, an expression, and an image.
16. The audio input method of claim 14 or 15, wherein, The outputting content corresponding to the audio according to the output mode corresponding to the input mode includes: In the content of the audio, the content corresponding to the second input mode is identified using a specified symbol, and the content corresponding to the first input mode is not identified using the symbol.
17. The audio input method of any one of claims 1-16, further comprising: displaying the audio input control, wherein the style of the audio input control is different in different input modes.
18. An electronic device, comprising: a memory; and a processor coupled to the memory, the processor configured to perform the audio input method of any one of claims 1-17 based on instructions stored in the memory.
19. A computer-readable storage medium having stored thereon a computer program, which, when executed by a processor, implements the audio input method of any one of claims 1-17.
20. A computer program product, which, when executed on a computer, causes the computer to implement the audio input method of any one of claims 1-17.
Citation Information
Patent Citations
Voice input method and terminal device
CN106933561A
Voice conversation generation method and device, electronic equipment and storage medium
CN114360535A
Audio editing method and device, equipment and storage medium
CN114915836A
Audio input method, electronic device, storage medium and computer program product
CN118945562A
Voice processing method and device, equipment, storage medium and product
CN119181362A