A voice input method, apparatus, electronic device, and storage medium
By introducing a registration and event detection mechanism for the breath-activated engine application into electronic devices, the problem of difficulty in operating the voice input button with one hand has been solved, thus making voice input more convenient.
Patent Information
- Application Number
- CN202310600976.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-25
- Publication Date
- 2026-01-23
- Estimated Expiration
- 2043-05-25
AI Technical Summary
In the existing technology, the voice input method of electronic devices is not convenient enough, especially when holding the device with one hand, it is difficult to click the top and bottom buttons at the same time, which makes the voice input operation cumbersome.
By registering with the breath wake-up engine application under the preset registration conditions of the first application, and generating a task to be processed when a breath wake-up event is detected, the task is distributed to the first application for voice data processing, simplifying user operation.
Without the need to manually click the voice input button, users can trigger voice data collection and processing through breath wake-up events, improving the convenience of voice input.
Smart Images

Figure CN119028333B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of information technology, and in particular, to a voice input method and device, an electronic device, and a storage medium. BACKGROUND
[0002] With the development of voice recognition technology, voice input scenarios in electronic devices are becoming more and more common. For example, search engines, social platforms, shopping software, input methods, and other third-party applications have begun to add voice input scenarios to support users to input voice instead of text, and to perform subsequent processing such as voice recognition by the application background.
[0003] Currently, switching voice input in various application scenarios needs to be performed manually. On the one hand, as shown in Figure 1-1 and Figure 1-2 , the button position for switching voice input is difficult to click; on the other hand, users often use a single hand to hold the electronic device, but the point touch range of the single hand is difficult to simultaneously consider the banner notification button located at the top and the voice input button located at the bottom, as shown in Figure 1-3 and Figure 1-4 (OKAY indicates that the single hand can point touch, and EAST & ACCURATE and EASY indicate that the single hand can simply and / or accurately point touch; other areas indicate that it is difficult to point touch).
[0004] Therefore, a more convenient voice input method is needed to support voice input scenarios of electronic devices. SUMMARY
[0005] Embodiments of the present application aim to provide a voice input method, device, electronic device, and storage medium to improve the convenience of voice input. The specific technical solutions are as follows:
[0006] In a first aspect, the embodiments of the present application provide a voice input method, comprising:
[0007] In a case where a first application meets a preset registration condition, the first application registers with a breath wake-up engine application;
[0008] After the first application is successfully registered, when the breath wake-up engine application detects a breath wake-up event, the breath wake-up engine generates a to-be-processed task for currently collected voice data, and distributes the to-be-processed task to the first application;
[0009] The first application acquires the voice data and processes it according to the to-be-processed task.
[0010] In one embodiment of the present application, the first application is an input method application, and the preset registration condition includes at least one of a soft keyboard being pulled up and a focus being in an input box of the input method application.
[0011] In one embodiment of the present application, the to-be-processed task includes an audio session identifier.
[0012] The first application acquires and processes the voice data according to the to-be-processed task, including:
[0013] The input method application acquires the voice data according to the audio session identifier.
[0014] The input method application invokes an audio recognition algorithm to convert the voice data into text data.
[0015] The input method application inputs the text data into the input box.
[0016] In one embodiment of the present application, the method further includes:
[0017] In response to the registration of the first application, the breath wake-up engine application detects whether an application that has been registered is using a breath wake-up function.
[0018] In the absence of an application that has been registered using the breath wake-up function, the breath wake-up engine application detects whether a switch of the breath wake-up function is turned on; if not, the switch of the breath wake-up function is turned on.
[0019] In the case where the switch of the breath wake-up function is turned on, the breath wake-up engine application sets an application to which the breath wake-up function is interfaced as the first application, so as to complete the registration of the first application.
[0020] In one embodiment of the present application, the method further includes:
[0021] After the first application is successfully registered, when the breath wake-up engine application detects a breath wake-up event, other breath wake-up events are shielded.
[0022] In one embodiment of the present application, the method further includes:
[0023] In the case where the first application meets a preset release condition, the first application releases the registration to the breath wake-up engine application.
[0024] In response to the release of the registration of the first application, the breath wake-up engine application returns to a working state before the registration of the first application.
[0025] In one embodiment of the present application, the first application is an input method application, and the preset release condition includes at least one of completion of processing of the voice data, hiding of the soft keyboard, the focus not being in an input box of the input method application, and user voluntary ending.
[0026] In one embodiment of the present application, in response to the release of the registration of the first application, the breath wake-up engine application returns to a working state before the registration of the first application, including:
[0027] In response to the release of the registration of the first application, in a case where a switch of a breath wake-up function of the breath wake-up engine application is closed before the registration of the first application, the switch of the breath wake-up function is closed.
[0028] And / or,
[0029] In response to the release of the registration of the first application, in a case where a switch of a breath wake-up function of the breath wake-up engine application is opened before the registration of the first application, the switch of the breath wake-up function is opened; in a case where the breath wake-up function has been connected to a second application before the registration of the first application, the breath wake-up engine application sets an application connected by the breath wake-up function to the second application.
[0030] In a second aspect, an embodiment of the present application provides a voice input device, including:
[0031] An application registration module, configured to, in a case where a first application meets preset registration conditions, the first application registers with a breath wake-up engine application;
[0032] A task distribution module, configured to, after the registration of the first application is successful, when the breath wake-up engine application detects a breath wake-up event, the breath wake-up engine generates a to-be-processed task for currently collected voice data, and distributes the to-be-processed task to the first application;
[0033] A voice data processing module, configured to, the first application acquires and processes the voice data according to the to-be-processed task.
[0034] In one embodiment of the present application, the first application is an input method application, and the preset registration condition includes at least one of pulling up of a soft keyboard and the focus being in an input box of the input method application.
[0035] In one embodiment of the present application, the to-be-processed task includes an audio session identifier.
[0036] The voice data processing module is specifically configured to:
[0037] The input method application acquires the voice data according to the audio session identifier.
[0038] The input method application invokes an audio recognition algorithm to convert the voice data into text data;
[0039] The input method application inputs the text data into the input box.
[0040] In one embodiment of the present application, the device further comprises:
[0041] An application detection module is configured to, in response to the registration of the first application, detect whether an application registered with the breath wake-up function is being used by the breath wake-up engine application;
[0042] A switch detection module is configured to, in the absence of an application registered with the breath wake-up function being used, detect whether a switch of the breath wake-up function is on by the breath wake-up engine application; and if not, turn on the switch of the breath wake-up function.
[0043] An application setting module is configured to, in the case where the switch of the breath wake-up function is on, set an application interfaced with the breath wake-up function as the first application by the breath wake-up engine application to complete the registration of the first application.
[0044] In one embodiment of the present application, the device further comprises:
[0045] A shielding module is configured to, after the first application is successfully registered, shield other breath wake-up events when the breath wake-up event is detected by the breath wake-up engine application.
[0046] In one embodiment of the present application, the device further comprises:
[0047] A registration cancellation module is configured to, in the case where the first application meets a preset cancellation condition, cancel the registration of the first application with the breath wake-up engine application.
[0048] A state recovery module is configured to, in response to the cancellation of the registration of the first application, recover the breath wake-up engine application to a working state before the registration of the first application.
[0049] In one embodiment of the present application, the first application is an input method application, and the preset cancellation condition includes at least one of the following: completion of processing of the voice data, hiding of the soft keyboard, the focus not being in an input box of the input method application, and user-initiated ending.
[0050] In one embodiment of the present application, the state recovery module is specifically configured to:
[0051] in response to deregistration of the first application, in a case where a switch of the breath wake-up function of the breath wake-up engine application is closed before the first application is registered, closing the switch of the breath wake-up function;
[0052] and / or,
[0053] in response to deregistration of the first application, in a case where a switch of the breath wake-up function of the breath wake-up engine application is opened before the first application is registered, opening the switch of the breath wake-up function; in a case where the breath wake-up function has docked a second application before the first application is registered, the breath wake-up engine application sets the application docked by the breath wake-up function to the second application.
[0054] In a third aspect, an embodiment of the present application provides an electronic device, comprising:
[0055] a memory for storing a computer program;
[0056] a processor for executing the program stored on the memory, to implement the voice input method described above.
[0057] In a fourth aspect, an embodiment of the present application provides a computer readable storage medium, the computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the voice input method described above.
[0058] The embodiments of the present application have the following beneficial effects:
[0059] The voice input method provided by the embodiments of the present application first registers the first application to the breath wake-up engine application in a case where the first application meets a preset registration condition, so that the first application can obtain the voice data collected by the breath wake-up engine application when a breath wake-up event is detected and the to-be-processed task generated by the breath wake-up engine application; after the first application is successfully registered, when the breath wake-up event is detected by the breath wake-up engine application, the breath wake-up engine application is triggered to collect voice data by the breath wake-up event, and then generates the to-be-processed task for the collected voice data and distributes the to-be-processed task to the first application, and then the first application obtains the voice data according to the to-be-processed task and processes the voice data, that is, the voice data to be collected only needs to be triggered by the breath wake-up engine through the breath wake-up event, and then the first application can obtain the collected voice data and process the voice data in time, without the need for the user to manually click the voice input button in the first application, so that the voice data input by the user can be triggered, the operation of the user is simplified, and the convenience of voice input is improved.
[0060] Of course, implementing any product or method of the present application does not necessarily require all the advantages described above. BRIEF DESCRIPTION OF DRAWINGS
[0061] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the accompanying drawings needed to be used in the description of the embodiments or the prior art will be briefly introduced. Obviously, the accompanying drawings in the following description only only some embodiments of the present application, and other embodiments can be obtained by those skilled in the art based on these drawings.
[0062] Figure 1-1 The first example of the button position of the voice input provided by the embodiment of the present application;
[0063] Figure 1-2 The second example of the button position of the voice input provided by the embodiment of the present application;
[0064] Figure 1-3 The first example of the touch range of the single hand holding provided by the embodiment of the present application;
[0065] Figure 1-4 The second example of the touch range of the single hand holding provided by the embodiment of the present application;
[0066] Figure 1-5 The first example of the breath wake-up event provided by the embodiment of the present application;
[0067] Figure 1-6 The first example of the breath wake-up event provided by the embodiment of the present application;
[0068] Figure 2 The first example of the breath wake-up event provided by the embodiment of the present application;
[0069] Figure 3 The second example of the breath wake-up event provided by the embodiment of the present application;
[0070] Figure 4 The second example of the breath wake-up event provided by the embodiment of the present application;
[0071] Figure 5 The first example of the breath wake-up event provided by the embodiment of the present application;
[0072] Figure 6 The first example of the breath wake-up event provided by the embodiment of the present application;
[0073] Figure 7 The second example of the breath wake-up event provided by the embodiment of the present application;
[0074] Figure 8-1 The second example of the breath wake-up event provided by the embodiment of the present application;
[0075] Figure 8-2 An example diagram of a first voice input method provided by an embodiment of the present application is shown in FIG. 1.
[0076] Figure 9-1 An example diagram of a second voice input method provided by an embodiment of the present application is shown in FIG. 2.
[0077] Figure 9-2 An example diagram of a flow of a voice input method provided by an embodiment of the present application is shown in FIG. 3.
[0078] Figure 10 An example diagram of a structure of a voice input device provided by an embodiment of the present application is shown in FIG. 4. DETAILED DESCRIPTION
[0079] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art based on the present application belong to the scope of protection of the present application.
[0080] In the related art, the method for using voice input by an electronic device is not convenient enough. In order to solve this problem, the embodiments of the present application provide a voice input method, device, electronic device and storage medium.
[0081] The implementation of the embodiments of the present application will be described in detail below with reference to the drawings.
[0082] The voice input method provided by the embodiments of the present application can be applied to any electronic device with voice input capability. The electronic device can be an electronic device with display screen hardware and corresponding software support. For example, the electronic device can be a mobile phone, a tablet computer, a wearable device, a vehicle-mounted device, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), a home device, etc. The specific type of the electronic device is not limited by the present application.
[0083] In a possible embodiment, in order to more clearly describe the electronic device provided by the embodiments of the present application, an example of a possible application scenario of the electronic device provided by the embodiments of the present application will be described below. It can be understood that the following example is only a possible application scenario of the electronic device provided by the embodiments of the present application. In other possible embodiments, the electronic device provided by the embodiments of the present application can also be applied to other possible application scenarios, and the following example does not limit this.
[0084] The voice input method provided in this application is applicable to any user of an electronic device and its voice assistant, such as office workers, white-collar workers, public officials, and workers; it is applicable to any scenario where a user needs to use an electronic device and its voice assistant. Specifically, it can be a scenario where the electronic device cannot be directly operated, or a scenario where the voice assistant cannot be woken up by voice. For example, when the user is in a relatively quiet public setting such as a coffee shop, a Western restaurant, or a high-speed rail / airport lounge; when the user is in a subway station, airport, train station security checkpoint, scenic spot ticket gate, etc., and has a lot of luggage; when the user is in a busy place such as a shopping mall or supermarket and needs to select goods; when the user is walking their dog outdoors; when the user is driving into / out of a parking lot, toll station, residential area / park, etc., and has just taken their hands off the steering wheel, etc.
[0085] like Figure 1-5 As shown, Figure 1-5 The illustration shows a scenario diagram of a method for waking up an application provided in an embodiment of this application. The user can lift the electronic device, bring the bottom of the electronic device close to the mouth, and speak into the microphone.
[0086] The electronic device collects voice data from the user's speech via a microphone and gesture data from the user lifting the device via an inertial sensor. This data is then sent to a low-power ADSP (Automatic Digital Signal Processor). The ADSP uses a breath-activated wake-up model to analyze the voice and gesture data. When the ADSP detects that the similarity between the voice data and preset wake-up breath data exceeds a first threshold, and the similarity between the gesture data and preset wake-up gesture data exceeds a second threshold, the ADSP sends the voice data to the breath-activated wake-up software module. The breath-activated wake-up software module then controls the application to start.
[0087] The distance between the electronic device and the user's mouth can be maintained at 0-5cm, which facilitates the accurate acquisition of the user's voice data by the microphone of the electronic device.
[0088] In addition, the above gesture data can be wrist-raising gesture data.
[0089] in, Figure 1-5 As can be seen from the gesture states indicated by box A in the left image to the gesture states indicated by box B in the right image, users can lift electronic devices by raising their wrists, and the inertial detection sensor can collect the gesture data of raising the wrist.
[0090] Figure 1-5 As shown in the right figure, when a user raises the electronic device by raising their wrist, they can speak by bringing their mouth close to the microphone of the electronic device and emitting the breath indicated in section C. The microphone can collect this breath and the corresponding voice data.
[0091] As shown in Figure 1-6 Fig. 1 shows a scene introduction schematic diagram of a breath wake-up engine application provided by an embodiment of the present application. The breath wake-up engine application is used to prompt the user that the breath wake-up can have a natural dialogue, i.e., ask and answer. In the case of an electronic device being a mobile phone, when the user holds up the mobile phone and places the bottom of the mobile phone close to the mouth (within a distance of 5 cm), and aims at the bottom microphone, the dialogue journey of the user (the user is called "you" to enhance the interactive feeling) will be started, which also indicates that the breath wake-up engine application will detect a breath wake-up event in this case.
[0092] It should be understood that the above is an example of the scene, and does not limit the scene of the present application in any way.
[0093] First, in a first aspect of an embodiment of the present application, an electronic device is provided, as shown in Figure 2 The electronic device includes:
[0094] a memory 201 for storing a computer program;
[0095] a processor 202 for executing the program stored in the memory 201 to implement the following steps:
[0096] In a case where a first application meets a preset registration condition, the first application registers with a breath wake-up engine application;
[0097] After the first application is successfully registered, when the breath wake-up engine application detects a breath wake-up event, the breath wake-up engine generates a to-be-processed task for current collected voice data, and distributes the to-be-processed task to the first application;
[0098] The first application acquires the voice data according to the to-be-processed task and processes the voice data.
[0099] The electronic device can further include a communication bus and / or a communication interface, and the processor 202, the communication interface, and the memory 201 can communicate with each other through the communication bus.
[0100] The communication bus mentioned above can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, only one thick line is shown in the figure, but it does not mean that there is only one bus or only one type of bus.
[0101] The communication interface is used for communication between the electronic device and other devices.
[0102] The memory may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.
[0103] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0104] For ease of explanation, we will use a mobile phone as an example.
[0105] like Figure 3 As shown, in some embodiments, electronic device 300 may include processor 301 and communication module 302, etc.
[0106] Among them, processor 301 and Figure 2 The processor 202 is identical and may include one or more processing units. For example, processor 301 may include an application processor (AP), a modem processor, a graphics processor, an image signal processor (ISP), a controller, memory, a video stream codec, a digital signal processor, a baseband processor, and / or a neural network processing unit (NPU), etc. Different processing units may be independent devices or integrated into one or more processors 301.
[0107] The controller can be the nerve center and command center of the electronic device 300. The controller can generate operation control signals according to the instruction opcode and timing signals to complete the control of instruction fetching and execution.
[0108] The processor 301 may also include a memory for storing instructions and data.
[0109] In some embodiments, the memory in the processor 301 is a cache memory. This memory can store instructions or data that the processor 301 has just used or that are used repeatedly. If the processor 301 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 301, and thus improves the efficiency of the system.
[0110] In some embodiments, the processor 301 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0111] The communication module 302 may include antenna 1, antenna 2, mobile communication module, and / or wireless communication module.
[0112] like Figure 3 As shown, in some embodiments, the electronic device 300 may also include an external memory interface 305, an internal memory 304, a USB interface 306, a charging management module 307, a power management module 308, a battery 309, and a sensor module 303, etc.
[0113] An NPU (Neural Processing Unit) is a computational processor for neural networks (NNs). By borrowing the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it can rapidly process input information and continuously learn on its own. NPUs can enable intelligent cognitive applications in electronic devices, such as image recognition, facial recognition, speech recognition, and text understanding.
[0114] The charging management module 307 is used to receive charging input from the charger. The charger can be a wireless charger or a wired charger.
[0115] In some embodiments with wired charging, the charging management module 307 can receive the charging input from the wired charger through the USB interface 306.
[0116] In some embodiments with wireless charging, the charging management module 307 can receive the wireless charging input through the wireless charging coil of the electronic device 300. The charging management module 307 can also supply power to the electronic device 300 through the power management module 308 while charging the battery 309.
[0117] The power management module 308 is configured to connect the battery 309, the charging management module 307, and the processor 301. The power management module 308 receives the input from the battery 309 and / or the charging management module 307 to supply power to the processor 301, the internal memory 304, the external memory, and the communication module 302, etc. The power management module 308 can also be configured to monitor the battery capacity, the battery cycle count, the battery health status (leakage, impedance), and the like.
[0118] In some other embodiments, the power management module 308 can also be disposed in the processor 301.
[0119] In some other embodiments, the power management module 308 and the charging management module 307 can also be disposed in the same device.
[0120] The external memory interface 305 can be configured to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 300. The external memory card communicates with the processor 301 through the external memory interface 305 to perform data storage functions, such as saving music, video streams, and the like in the external memory card.
[0121] The internal memory 304 can be configured to store computer executable program codes, which include instructions. The processor 301 executes various functional applications and data processing of the electronic device 300 by running the instructions stored in the internal memory 304. The internal memory 304 can include a program storage area and a data storage area. The program storage area can store an operating system, at least one application program required by a function (such as a sound playing function, an image playing function, etc.), and the like. The data storage area can store data created during the use of the electronic device 300 (such as audio data, a phonebook, etc.), and the like. In addition, the internal memory 304 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, a universal flash storage (UFS), and the like.
[0122] The sensor module 303 in the electronic device 300 can include components such as an image sensor, a touch sensor, a pressure sensor, a gyroscope sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a distance sensor, a proximity light sensor, an ambient light sensor, a fingerprint sensor, a temperature sensor, a bone conduction sensor, and the like to implement sensing and / or acquisition functions for different signals.
[0123] Optionally, the electronic device 300 can further include peripheral devices such as a mouse, a key, an indicator light, a keyboard, a speaker, a microphone, and the like.
[0124] The keys include a power-on key, a volume key, and the like. The keys can be mechanical keys. They can also be touch keys. The electronic device 300 can receive key input and generate key signal input related to user settings and function control of the electronic device 300.
[0125] The indicator can be an indicator light, which can be used to indicate a charging state and a power change, and can also be used to indicate a message, a missed call, and a notification, and the like.
[0126] It can be understood that the structure illustrated in the embodiments does not constitute a specific limitation on the electronic device 300.
[0127] In other embodiments, the electronic device 300 can include more or fewer components than those shown, or a combination of some components, or a split of some components, or a different arrangement of components. The illustrated components can be implemented in hardware, software, or a combination of software and hardware.
[0128] In an electronic device, a software system thereof can be divided into several layers, as shown in Figure 4 Figure 4 A software structure block diagram of an electronic device provided by the embodiments of the present application. The layered architecture divides the software system of the electronic device into several layers, each of which has a clear role and division of labor. Layers communicate with each other through software interfaces. In some embodiments, the application system of the electronic device can be divided into an application layer (application), a framework layer (framework, fwk), a driver layer (hardware abstract layer, HAL), and a kernel layer (kernel) / chip layer.
[0129] The application layer can include a series of application packages, and the application layer runs the application by calling the application programming interface (API) provided by the application framework layer. The application package can include a plurality of applications, such as a breath wake-up engine application, a voice assistant application, an input method application, a call application, and a map application.
[0130] The framework layer provides APIs and a programming framework for applications in the application layer. The application framework layer includes predefined functions. For example... Figure 4 As shown, the application framework layer may include an audio trigger module (Sound Trigger). This audio trigger module can be used to control the launch of applications within the application layer.
[0131] In addition, the framework layer may also include a sound trigger module, an audio policy service module, etc. Furthermore, the application framework layer may include an audio service module, an audio trigger service module, an audio flinger module, etc. The sound trigger module is used to send wake-up events to the audio trigger module and to send notifications to the audio driver in the driver layer to instruct the breath-activated wake-up model to stop or start running. The audio policy service module is used to establish a speech recognition channel with the voice assistant application in the application layer.
[0132] Additionally, the audio service module is used to send a startup notification to the audio trigger module in response to a startup notification sent by the application in the application layer. The audio trigger module is also used to send a startup notification to the audio trigger service module in response to the startup notification sent by the audio service module. The audio trigger service module is used to send a startup notification to the sound trigger module in response to the startup notification sent by the audio trigger module. The sound trigger module is also used to send a notification to the audio policy service module to start running the breath wake-up model in response to the startup notification sent by the audio trigger module. The audio policy service module starts running the breath wake-up model in response to the notification to start running the breath wake-up model. The audio indication module is used to send a notification to the audio driver of driver layer 3 to load the breath wake-up model in response to the notification sent by the audio policy service module. The audio driver of the driver layer sends a notification to the audio digital signal processor of the breath wake-up processing device to indicate the start of running the breath wake-up model.
[0133] The driver layer is the layer between hardware and software, used to drive the hardware and make it work. Multiple drivers can be installed in the driver layer to operate the hardware. Examples include audio drivers (sound trigger HALs), sound trigger drivers, and audio stream input drivers.
[0134] The kernel / chip layer includes an audio digital signal processor (ADSP), which determines the presence of a breath wake-up event based on a breath wake-up algorithm. The ADSP can obtain the event identifier, i.e., the event ID, corresponding to the wake-up event, to distinguish between breath wake-up events and voice wake-up events. This event ID is preset in the ADSP by the developers. In one example, after detecting a breath wake-up event, the ADSP reports the breath wake-up event containing the event ID to the driver layer; the driver layer then reports the breath wake-up event to the framework layer; the framework layer reports the breath wake-up event to the breath wake-up engine application, which then processes the breath wake-up event.
[0135] It should be noted that the application layer, application framework layer, driver layer, and kernel / chip layer may also include other content, which is not specifically limited here.
[0136] The method of this application embodiment can be used to report the breath wake-up event layer by layer after the underlying breath wake-up algorithm recognizes the breath wake-up event with the event ID. This enables the breath wake-up engine application to determine the target application for voice data processing based on the business status of the registered application after receiving the breath wake-up event. The target application can then obtain voice data for data processing based on the audio session identifier or buffer size.
[0137] To address the technical problem that voice input methods for electronic devices are not convenient enough, a second aspect of the embodiments of this application is as follows: Figure 5 As shown, Figure 5 A flowchart illustrating the first voice input method provided in this application embodiment includes:
[0138] Step S501: If the first application meets the preset registration conditions, the first application registers with the breath wake-up engine application.
[0139] The first application is any application capable of processing voice data. For example, the first application can be a recording application, a video shooting application, a search engine application, a social platform application, a shopping software application, or an input method application, etc.
[0140] Preset registration conditions can be used to indicate that the first application is currently being used by a user. For example, the first application is currently inputting content or receiving relevant instructions from the user. The input content can be text data or voice data, and the relevant instructions can be implemented through finger operations or voice commands. The first application registers with the breath wake-up engine application by sending a registration message to the breath wake-up engine application, which then completes the registration based on the registration message.
[0141] In one embodiment of this application, when the first application is an input method application, the preset registration conditions include at least one of the following: the soft keyboard is raised and the focus is on the input box of the input method application. That is, the preset registration conditions include the soft keyboard in the input method application being raised by the user in preparation for input (indicating the input method application is started), and the focus (cursor) being on the input box (indicating the input focus is valid). In addition, the preset registration conditions may also include other conditions indicating that the user needs to operate the input method application, or that the user cannot currently operate the input method application directly, or that the input method application needs to collect and process voice data to meet the user's needs. In other embodiments, when the first application is a recording application, if the recording application is in the foreground and the focus is on the recording application, it is determined that the recording application meets the preset registration conditions.
[0142] When the first application is an input method application and the preset registration conditions include at least one of the following: the soft keyboard is pulled up and the focus is on the input box of the input method application, it means that the input method is about to input content. Therefore, the input method application needs to register with the breath wake-up engine application to ensure that the voice data input by the user can be processed by the input method application.
[0143] Therefore, if the first application meets the preset registration conditions, the first application will register with the Breath Awakening Engine application.
[0144] Step S502: After the first application is successfully registered, when the breath wake-up engine application detects a breath wake-up event, the breath wake-up engine generates a task to be processed for the currently collected voice data and distributes the task to be processed to the first application.
[0145] At this time, the Breath Wake-up Engine application generates a task to be processed for the collected voice data. The task to be processed is used to trigger the first application to obtain the voice data of the current user input collected by the Breath Wake-up Engine application and process the voice data. For example, the task to be processed may include parameters such as audio session ID and buffer size to distinguish the voice data. The first application can uniquely obtain the voice data based on these parameters.
[0146] Step S503: The first application obtains the voice data according to the task to be processed and processes it.
[0147] The first application retrieves voice data from the breath-activated engine application based on the task at hand and processes the voice data. For example, when the first application is an output method application, it can call a speech recognition algorithm to convert the voice data into text and input the text into an input box. In another example, when the first application is a recording application, it can call a voice processing algorithm to perform a series of operations on the voice data, such as noise reduction and waveform shaping, and finally store the processed voice data in the recording application's default directory.
[0148] As can be seen from the above, the voice input method provided in this application firstly registers with the breath wake-up engine application when the first application meets the preset registration conditions, so that the first application can obtain the voice data collected by the breath wake-up engine application when it detects a breath wake-up event and the generated pending tasks. After the first application successfully registers, when the breath wake-up engine application detects a breath wake-up event, the breath wake-up engine application is triggered by the breath wake-up event to collect voice data, and then generates pending tasks for the collected voice data and distributes them to the first application. Then, the first application obtains voice data and processes it according to the pending tasks. In other words, collecting voice data only requires triggering the breath wake-up engine through a breath wake-up event, and then the first application can obtain the collected voice data and process it in a timely manner. There is no need for the user to manually click the voice input button in the first application to trigger the collection of the voice data that the user needs to input, which simplifies the user's operation and improves the convenience of voice input.
[0149] In one embodiment of this application, the task to be processed includes an audio session identifier, which is used to distinguish each voice data collected by the breath wake-up engine application. The voice data corresponding to each task to be processed for each first application is unique, and each voice data can be distinguished by its corresponding audio session ID. In addition, in one example, the task to be processed may also include other parameters corresponding to the collected voice data, such as the data size and buffer size of the voice data.
[0150] When the first application is an input method application, such as Figure 6 As shown, step S503 above, the first application, obtains and processes the voice data according to the task to be processed, including:
[0151] Step S601: The input method application obtains the voice data based on the audio session identifier;
[0152] Step S602: The input method application calls the audio recognition algorithm to convert the voice data into text data;
[0153] Step S603: The input method application inputs the text data into the input box.
[0154] The input method application obtains the voice data currently collected by the breath wake-up engine application based on the audio session identifier included in the task to be processed generated by the breath wake-up engine application. Then, the input method application calls the built-in audio recognition algorithm to recognize the voice data, obtain the content represented by the voice data, convert the content from voice form to text form, obtain the converted text data, and then input the text data into the input box to complete the input operation.
[0155] As can be seen from the above, the voice input method provided in this application uses an input method application to obtain voice data based on the audio session identifier included in the task to be processed generated by the breath wake-up engine application, and then converts the voice data into text data and inputs it into the input box. This realizes the acquisition and input of the content required by the user, so that the user can complete the input operation without manually clicking the voice input button, which simplifies the user's operation and makes the content input simpler and faster.
[0156] In one embodiment of this application, such as Figure 7 As shown in the figure, this application embodiment provides a flowchart of a second voice input method, the method further comprising:
[0157] Step S701: In response to the registration of the first application, the breath wake-up engine application detects whether any registered applications are using the breath wake-up function.
[0158] Step S702: If no registered application is using the breath wake-up function, the breath wake-up engine application detects whether the switch of the breath wake-up function is turned on; if it is not turned on, the switch of the breath wake-up function is turned on.
[0159] Step S703: When the breath wake-up function is turned on, the breath wake-up engine application sets the application connected to the breath wake-up function as the first application to complete the registration of the first application.
[0160] When the first application registers with the Breath Wake-up Engine application, the Breath Wake-up Engine application first checks if any previously registered applications are currently using the Breath Wake-up function. If so, it means the Breath Wake-up function is already in use by a connected application and cannot be reused; the first application cannot register at this time. In one example, the Breath Wake-up Engine application can either reject the first application's registration or temporarily suspend the first application's registration request until the Breath Wake-up function returns to an idle state before responding to the first application's registration and performing subsequent checks.
[0161] If no registered application is using the breath wake-up function, it indicates that the breath wake-up function is currently idle. The breath wake-up engine application continues to check if the breath wake-up function is turned on in the electronic device. If it is turned on, the breath wake-up function can be used directly, and the application connected to the breath wake-up function can be directly registered and set as the first application. If it is not turned on, the breath wake-up engine application needs to turn on the breath wake-up function first, and then respond to the registration of the first application and set as the first application, thus completing the registration of the first application and indicating that the first application can currently use the breath wake-up function. When the breath wake-up function is turned on, the breath wake-up engine application monitors breath wake-up events; when the breath wake-up function is turned off, the breath wake-up engine application does not monitor breath wake-up events.
[0162] In one embodiment of this application, the method further includes:
[0163] After the first application is successfully registered, when the Breath Wake-up Engine application detects a Breath Wake-up event, it blocks other Breath Wake-up events.
[0164] After the first application is successfully registered, when the breath wake-up engine application detects a breath wake-up event, it will collect voice data for that event. During this process, the breath wake-up engine application blocks other breath wake-up events; that is, it will not trigger other breath wake-up events until the current event has been completed. Other breath wake-up events are only allowed to be triggered after the voice data for the current event has been collected.
[0165] As can be seen from the above, the voice input method provided in this application embodiment allows the breath wake-up engine application to detect whether the breath wake-up function is idle or not when the breath wake-up function is turned on, and to turn on the breath wake-up function when it is turned off. Furthermore, after the first application successfully registers, when the breath wake-up engine application detects a breath wake-up event, it will also block other breath wake-up events unrelated to the first application. This allows the breath wake-up function to be used by various applications, including the first application, in an orderly and effective manner, enabling them to collect and process voice data in an orderly and timely manner. This avoids invalid loop wake-up and interference during the voice data collection process, further improving the convenience of voice input and the efficiency of input and processing.
[0166] In one embodiment of this application, such as Figure 8-1 As shown in the figure, this application embodiment provides a flowchart of a third voice input method, the method further comprising:
[0167] Step S801: If the first application meets the preset deregistration conditions, the first application deregisters with the breath wake-up engine application.
[0168] In step S802, in response to the deregistration of the first application, the breath wake-up engine application is restored to its working state before the first application was registered.
[0169] The preset cancellation condition indicates that the first application is no longer accepting input or receiving commands from the user, and therefore no longer needs to prepare for collecting and processing voice data. Specifically, this could be because the first application has already finished accepting input or commands, or because the user actively closes the application. For example, the user might close the first application due to an operational error or when input or commands are not actually required. In one example, the user might close the first application or perform an operation on the application's interface indicating that input is complete. This could be because the user closes a search engine application, clicks the search button in the search interface, or the last voice command received by the search engine application was "search," or closes a social media application, clicks the send button in the message sending interface, or the last voice command received by the social media application was "send," etc. Alternatively, the first application might detect that the user has stopped accepting input or commands within a preset time period (e.g., 3 seconds, 5 seconds), such as the user not accepting input or commands within the preset time period in the search engine application's search interface, or the user not accepting input or commands within the preset time period in the social media application's message sending interface, etc.
[0170] In one embodiment of this application, the first application is an input method application. The preset cancellation conditions may include at least one of the following: voice data processing completed (i.e., the input method application has converted all input voice data into text data and entered it into the input box), the soft keyboard hidden (i.e., the soft keyboard used for inputting content in the input method application is hidden and the input method application is no longer being used), the focus is not in the input box of the input method application (i.e., the user's cursor is no longer in the input box of the input method application within a preset time period), and the user actively ending the input (i.e., the user actively ends the input). These conditions all indicate that the user no longer needs to use the input method application to input content. This allows for timely determination that the first application no longer needs to use the breath wake-up function.
[0171] In this scenario, the first application unregisters with the Breath Wake-up Engine application, indicating that it no longer needs the Breath Wake-up function. The Breath Wake-up Engine application then responds to this unregistration by reverting to its pre-registration state. This prevents the first application from continuing to affect the Breath Wake-up Engine application after it stops using the function, or from altering the engine's settings due to its use, thus improving the engine's stability.
[0172] In one embodiment of this application, step S802, in response to the deregistration of the first application, restores the breath wake-up engine application to its working state before the first application's registration, including:
[0173] In response to the deregistration of the first application, if the breath wake-up function of the breath wake-up engine application was turned off before the first application was registered, the switch of the breath wake-up function is turned off.
[0174] And / or,
[0175] In response to the deregistration of the first application, if the breath wake-up function of the breath wake-up engine application was turned on before the first application was registered, the breath wake-up function is turned on; if the breath wake-up function was already connected to the second application before the first application was registered, the breath wake-up engine application sets the application connected to the breath wake-up function as the second application.
[0176] The Breath Wake-up Engine application responds to the deregistration of the first application and restores to its previous working state. If the Breath Wake-up function was off before, it will be restored to the off state. If the Breath Wake-up function was on before, it will be restored to the on state. If the Breath Wake-up function already had a connected second application before, it will be re-set as the second application.
[0177] like Figure 8-2 As shown, taking the input method application as an example, the breath wake-up engine application can include a switch state management module and a scene awareness decision module. The scene awareness decision module detects the connection status of registered applications and distributes tasks to the target application (input method application), enabling the input method application to retrieve audio streams based on the specified audio session identifier in the task. The switch state management module manages the switch state of the breath wake-up function, determining whether to open or close the breath wake-up path based on the switch state, and also registers applications. The input method application first registers with the switch management module in the breath wake-up engine application. After registration, the switch state management module can query the switch state of the breath wake-up function. If it detects that the switch state is open, it opens the breath wake-up path, and the breath wake-up engine application can detect the breath wake-up event. If it detects that the switch state is closed, it closes the breath wake-up path, and the breath wake-up engine application cannot detect the breath wake-up event.
[0178] As can be seen from the above, the voice input method provided in this application embodiment can promptly disable the first application's use of the breath wake-up function, avoiding obstruction of other applications' use of the breath wake-up function. At the same time, it restores the on / off state of the breath wake-up function in the breath wake-up engine application and the connected applications to the state before the first application's registration, avoiding any potential continuous impact after the first application's registration is removed, and improving the stability of the breath wake-up engine application's own settings.
[0179] In one embodiment of this application, such as Figure 9-1 and Figure 9-2 As shown, an example diagram of a voice input method is provided. When the first application is an input method application, the soft keyboard in the input method application interface is brought up, and the voice input is combined with the breath wake-up technology required for breath wake-up input. Voice data is collected and a task to be processed by the input method application is generated, so that the input method application converts the voice data into text and sends the converted text.
[0180] Specifically, such as Figure 9-2As shown, the input method application first needs to meet preset registration conditions such as the soft keyboard being displayed, indicating that the input method application is started, and the input focus (cursor) being valid. If these conditions are not met, the process ends directly. If they are met, the input method application registers with the breath wake-up engine application. The breath wake-up engine application will detect whether the breath wake-up function is in use. If it is in use, the process ends directly. If it is not in use, the breath wake-up engine application will detect whether the breath wake-up function is turned on. If it is not turned on, the switch will be turned on. If it is turned on, the target application will be switched directly, that is, the connected application will be set as the input method application. Then, the breath wake-up engine application will listen for and detect the breath wake-up to obtain the current breath wake-up event, while blocking new and other breath wake-up events.
[0181] When a breath wake-up event occurs, the breath wake-up engine application collects the current voice data and generates a task to be processed for that voice data. Then, the breath wake-up engine application distributes the task to the input method application to notify the input method application to retrieve the task. The input method application then starts the recording function to acquire the voice data, performs audio recognition on the voice data, converts it into text data, and completes the input in the input box. When the input method application meets the preset release condition for the end of input, the input method application unregisters with the breath wake-up engine application. Then, the breath wake-up engine application returns to its previous working state, that is, the breath wake-up engine application checks whether the original switch state of the breath wake-up function is on. If it is, it restores the breath wake-up monitoring of the original target application (the originally connected application). Otherwise, it disables the breath wake-up function, restores the target application, and then ends.
[0182] See Figure 10 , Figure 10 A schematic diagram of a voice input device provided in this application embodiment includes:
[0183] The application registration module 1001 is used to register the first application with the breath wake-up engine application when the first application meets the preset registration conditions.
[0184] The task distribution module 1002 is used to generate a task to be processed for the currently collected voice data when the breath wake-up engine application detects a breath wake-up event after the first application has successfully registered. The task to be processed is then distributed to the first application.
[0185] The voice data processing module 1003 is used by the first application to obtain and process the voice data according to the task to be processed.
[0186] In one embodiment of this application, the first application is an input method application, and the preset registration conditions include at least one of the following: the soft keyboard is pulled up and the focus is on the input box of the input method application.
[0187] As can be seen from the above, the voice input device provided in this application first registers with the breath wake-up engine application when the first application meets the preset registration conditions, so that the first application can obtain the voice data collected by the breath wake-up engine application when it detects a breath wake-up event and the generated pending tasks. After the first application successfully registers, when the breath wake-up engine application detects a breath wake-up event, the breath wake-up engine application is triggered by the breath wake-up event to collect voice data, and then generates pending tasks for the collected voice data and distributes them to the first application. Then the first application obtains voice data and processes it according to the pending tasks. In other words, collecting voice data only requires triggering the breath wake-up engine through a breath wake-up event, and then the first application can obtain the collected voice data and process it in a timely manner. There is no need for the user to manually click the voice input button in the first application to trigger the collection of the voice data that the user needs to input, which simplifies the user's operation and improves the convenience of voice input.
[0188] In one embodiment of this application, the task to be processed includes an audio session identifier;
[0189] The voice data processing module 1003 is specifically used for:
[0190] The input method application obtains the voice data based on the audio session identifier;
[0191] The input method application calls an audio recognition algorithm to convert the voice data into text data;
[0192] The input method application inputs the text data into the input box.
[0193] As can be seen from the above, the voice input device provided in this application uses an input method to obtain voice data based on the audio session identifier included in the task to be processed generated by the breath wake-up engine application. Then, the voice data is converted into text data and input into the input box, thereby realizing the acquisition and input of the content required by the user. This allows the user to complete the input operation without manually clicking the voice input button in the first application, simplifying the user's operation and making the content input simpler and faster.
[0194] In one embodiment of this application, the apparatus further includes:
[0195] An application detection module is used to detect whether a registered application is using the breath wake-up function in response to the registration of the first application.
[0196] The switch detection module is used to detect whether the breath wake-up function is turned on when no registered application is using the breath wake-up function; if it is not turned on, the breath wake-up function is turned on.
[0197] The application settings module is used to, when the breath wake-up function is turned on, set the application that the breath wake-up function is connected to as the first application, so as to complete the registration of the first application.
[0198] As can be seen from the above, the voice input device provided in this application embodiment allows the breath wake-up engine application to detect whether the breath wake-up function is idle or not when the breath wake-up function is turned on, and to turn on the breath wake-up function when it is turned off. Furthermore, after the first application successfully registers, when the breath wake-up engine application detects a breath wake-up event, it will also block other breath wake-up events unrelated to the first application. This allows the breath wake-up function to be used by various applications, including the first application, in an orderly and effective manner, enabling these applications to collect and process voice data in an orderly and timely manner. This avoids invalid loop wake-up and interference during the voice data collection process, further improving the convenience of voice input and the efficiency of input and processing.
[0199] In one embodiment of this application, the apparatus further includes:
[0200] The blocking module is used to block other breath wake-up events when the breath wake-up engine application detects a breath wake-up event after the first application has successfully registered.
[0201] In one embodiment of this application, the apparatus further includes:
[0202] The registration deregistration module is used to deregister the first application with the breath wake-up engine application when the first application meets the preset deregistration conditions.
[0203] The state recovery module is used to restore the breath wake-up engine application to its working state before the first application was registered in response to the first application's deregistration.
[0204] In one embodiment of this application, the first application is an input method application, and the preset termination condition includes at least one of the following: the voice data processing is completed, the soft keyboard is hidden, the focus is not in the input box of the input method application, and the user actively terminates the application.
[0205] In one embodiment of this application, the state recovery module is specifically used for:
[0206] In response to the deregistration of the first application, if the breath wake-up function of the breath wake-up engine application was turned off before the first application was registered, the switch of the breath wake-up function is turned off.
[0207] And / or,
[0208] In response to the deregistration of the first application, if the breath wake-up function of the breath wake-up engine application was turned on before the first application was registered, the breath wake-up function is turned on; if the breath wake-up function was already connected to the second application before the first application was registered, the breath wake-up engine application sets the application connected to the breath wake-up function as the second application.
[0209] As can be seen from the above, the voice input device provided in this application embodiment can promptly disable the use of the breath wake-up function by the first application, avoiding obstruction of other applications using the breath wake-up function. At the same time, it restores the on / off state of the breath wake-up function in the breath wake-up engine application and the connected applications to the state before the first application registered, avoiding the possible continuous impact after the first application is removed from the registration process, and improving the stability of the breath wake-up engine application's own settings.
[0210] In another embodiment provided in this application, a computer-readable storage medium is also provided, which stores a computer program that, when executed by a processor, implements the steps of any of the above-described voice input methods.
[0211] In another embodiment provided in this application, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to execute any of the voice input methods described above.
[0212] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state disk (SSD)).
[0213] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0214] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the device and electronic device embodiments are basically similar to the method embodiments, so the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0215] The above description is merely a preferred embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application are included within the scope of protection of this application.
Claims
1. A voice input method, characterized in that, include: If the first application meets the preset registration conditions, the first application registers with the Breath Wake-up Engine application. The preset registration conditions indicate that the first application is currently being used by the user. In response to the registration of the first application, the breath wake-up engine application detects whether any registered applications are using the breath wake-up function; If no registered application is using the breath wake-up function, the breath wake-up engine application detects whether the breath wake-up function is turned on. If not turned on, turn on the switch for the breath wake-up function; When the breath wake-up function is turned on, the breath wake-up engine application sets the application connected to the breath wake-up function as the first application to complete the registration of the first application; After the first application is successfully registered, when the breath wake-up engine application detects a breath wake-up event, the breath wake-up engine generates a task to be processed for the currently collected voice data and distributes the task to be processed to the first application. The first application acquires and processes the voice data according to the task to be processed.
2. The method according to claim 1, characterized in that, The first application is an input method application, and the preset registration conditions include at least one of the following: the soft keyboard is pulled up and the focus is on the input box of the input method application.
3. The method according to claim 2, characterized in that, The tasks to be processed include audio session identifiers; The first application acquires and processes the voice data according to the task to be processed, including: The input method application obtains the voice data based on the audio session identifier; The input method application calls an audio recognition algorithm to convert the voice data into text data; The input method application inputs the text data into the input box.
4. The method according to claim 1, characterized in that, The method further includes: After the first application is successfully registered, when the breath wake-up engine application detects a breath wake-up event, it blocks other breath wake-up events.
5. The method according to claim 1, characterized in that, The method further includes: If the first application meets the preset deregistration conditions, the first application will deregister with the breath wake-up engine application. In response to the first application's deregistration, the Breath Wake-up Engine application reverts to its working state before the first application's registration.
6. The method according to claim 5, characterized in that, The first application is an input method application, and the preset termination conditions include at least one of the following: the voice data processing is completed, the soft keyboard is hidden, the focus is not in the input box of the input method application, or the user actively terminates the application.
7. The method according to claim 5, characterized in that, In response to the deregistration of the first application, the breath wake-up engine application restores to its working state before the registration of the first application, including: In response to the deregistration of the first application, if the breath wake-up function of the breath wake-up engine application was turned off before the first application was registered, the switch of the breath wake-up function is turned off. And / or, In response to the deregistration of the first application, if the breath wake-up function of the breath wake-up engine application was turned on before the first application was registered, the breath wake-up function is turned on; if the breath wake-up function was already connected to the second application before the first application was registered, the breath wake-up engine application sets the application connected to the breath wake-up function as the second application.
8. A voice input device, characterized in that, include: The application registration module is used to register the first application with the breath wake-up engine application when the first application meets the preset registration conditions. The preset registration conditions indicate that the first application is currently being used by the user. An application detection module is used to detect whether a registered application is using the breath wake-up function in response to the registration of the first application. A switch detection module is used to detect whether the breath wake-up function is turned on when no registered application is using the breath wake-up function. If not turned on, turn on the switch for the breath wake-up function; The application settings module is used to set the application that the breath wake-up function is connected to as the first application when the switch of the breath wake-up function is turned on, so as to complete the registration of the first application. The task distribution module is used to generate a task to be processed for the currently collected voice data when the breath wake-up engine application detects a breath wake-up event after the first application has successfully registered. The task to be processed is then distributed to the first application. A voice data processing module is used by the first application to obtain and process the voice data according to the task to be processed.
9. The apparatus according to claim 8, characterized in that, The first application is an input method application, and the preset registration conditions include at least one of the following: the soft keyboard is pulled up and the focus is on the input box of the input method application.
10. The apparatus according to claim 9, characterized in that, The tasks to be processed include audio session identifiers; The voice data processing module is specifically used for: The input method application obtains the voice data based on the audio session identifier; The input method application calls an audio recognition algorithm to convert the voice data into text data; The input method application inputs the text data into the input box.
11. The apparatus according to claim 8, characterized in that, The device further includes: The blocking module is used to block other breath wake-up events when the breath wake-up engine application detects a breath wake-up event after the first application has successfully registered.
12. The apparatus according to claim 8, characterized in that, The device further includes: The registration deregistration module is used to deregister the first application with the breath wake-up engine application when the first application meets the preset deregistration conditions. The state recovery module is used to restore the breath wake-up engine application to its working state before the first application was registered in response to the first application's deregistration.
13. The apparatus according to claim 12, characterized in that, The first application is an input method application, and the preset termination conditions include at least one of the following: the voice data processing is completed, the soft keyboard is hidden, the focus is not in the input box of the input method application, or the user actively terminates the application.
14. The apparatus according to claim 12, characterized in that, The state recovery module is specifically used for: In response to the deregistration of the first application, if the breath wake-up function of the breath wake-up engine application was turned off before the first application was registered, the switch of the breath wake-up function is turned off. And / or, In response to the deregistration of the first application, if the breath wake-up function of the breath wake-up engine application was turned on before the first application was registered, the breath wake-up function is turned on; if the breath wake-up function was already connected to the second application before the first application was registered, the breath wake-up engine application sets the application connected to the breath wake-up function as the second application.
15. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, when executing a program stored in memory, implements the method described in any one of claims 1-7.
16. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method described in any one of claims 1-7.
Citation Information
Patent Citations
Speech interactive awakening electronic device and method based on microphone signal and medium
CN110223711A
Voice interaction processing method and device and electronic equipment
CN111354360A
Voice data processing method and device, electronic equipment and storage medium
CN119028332A