Speech processing methods, apparatus, computer equipment and computer-readable storage media
By acquiring the priority configuration and interaction status of voice services, and using a voice engine to determine the target voice service, the problem of interactive devices being unable to recognize multi-round interactive voices is solved, achieving the effect of accurately responding to user voices in multiple business scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-31
- Publication Date
- 2026-04-03
AI Technical Summary
When multiple business scenarios are activated simultaneously, the interactive device cannot accurately recognize the user's multi-round interactive voice, resulting in the inability to provide the business scenarios required by the user.
By acquiring the priority configuration and interaction status of each voice service, the target voice service is determined using the voice engine, and semantic signals are generated based on the priority configuration and sent to the target voice service to process the user's voice.
In multi-round interaction scenarios, based on the priority configuration of voice services and the interaction status, it is ensured that the interactive device can respond to the user's voice and provide the user with the required service scenarios.
Smart Images

Figure CN116364081B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of speech recognition technology, and more specifically to a speech processing method, apparatus, computer device, and computer-readable storage medium. Background Technology
[0002] With the rapid development of the automotive industry, more and more vehicles offer voice interaction capabilities. Vehicle central control systems and other interactive devices can recognize the user's voice and then display corresponding pages on the screen based on the semantics of the voice, providing the user with the required business scenarios. When the interactive device screen is small, because only a small amount of data is displayed at a time, usually only one business scenario is active. When the interactive device screen is large, multiple business scenarios can be active simultaneously.
[0003] When multiple business scenarios are active simultaneously, for single-turn voice interactions, the interactive device can recognize the user's intent based on the voice and then provide the required business scenario based on the semantics of the voice. However, due to overly colloquial speech or unclear business scenarios mentioned in the voice, for multi-turn voice interactions, the interactive device may fail to accurately identify the requested business scenario. The interactive device's inability to reliably process multi-turn voice interactions results in its inability to provide the required business scenario based on the voice. Summary of the Invention
[0004] One objective of this invention is to provide a voice processing method to solve the problem that the prior art cannot provide the business scenarios required by the user based on voice; another objective is to provide a voice processing device; a third objective is to provide a computer device; and a fourth objective is to provide a computer-readable storage medium.
[0005] To achieve the above objectives, the technical solution adopted by the present invention is as follows:
[0006] Firstly, this application provides a speech processing method, including:
[0007] Obtain the priority configuration and interaction status of each voice service, where the interaction status includes executable and non-executable states;
[0008] If the operating status of any voice service changes, the interaction status of the voice service will be transmitted to the voice engine.
[0009] Generate semantic signals based on the received user voice;
[0010] If the voice engine determines that the interaction state of at least two voice services is executable, the target voice service is determined based on the priority configuration.
[0011] Send semantic signals to the target voice service.
[0012] In embodiments of this application, the voice processing method further includes:
[0013] If the voice engine determines that the number of voice services with an interactive state of executable is equal to one, then the voice service with an interactive state of executable is identified as the target voice service.
[0014] In the embodiments of this application, when the voice engine determines that the interaction state of at least two voice services is executable, the target voice service is determined based on priority configuration, including:
[0015] If the voice engine determines that the interaction state of at least two voice services is executable, the target voice service is determined based on the priority configuration and the order in which the voice engine receives the interaction state of each voice service.
[0016] In the embodiments of this application, when the voice engine determines that the interaction state of at least two voice services is executable, the target voice service is determined based on priority configuration, including:
[0017] If the voice engine determines that the interaction state of at least two voice services is executable, the voice service with the executable interaction state and the priority configured as high priority is identified as the target voice service.
[0018] In embodiments of this application, after sending the semantic signal to the target voice service, the method further includes:
[0019] In response to the semantic execution completion signal based on the target voice service, the interaction state of the target voice service is determined to be non-executable, and the interaction state of the target voice service is transmitted to the voice engine.
[0020] In embodiments of this application, generating semantic signals based on received user speech includes:
[0021] The received user voice is converted into a signal, and the converted user voice is then denoised to generate a semantic signal.
[0022] In the embodiments of this application, before obtaining the priority configuration and interaction state of each voice service, the method further includes:
[0023] Initialize the voice engine and each voice service.
[0024] Secondly, this application provides a voice processing apparatus, comprising:
[0025] The configuration acquisition module is used to acquire the priority configuration and interaction status of each voice service. The interaction status includes executable status and non-executable status.
[0026] The status transmission module is used to transmit the interaction status of the voice service to the voice engine when the running status of any voice service changes.
[0027] The signal generation module is used to generate semantic signals based on the received user speech;
[0028] The target determination module is used to determine the target voice service based on priority configuration when the voice engine determines that the interaction state of at least two voice services is executable.
[0029] The signal transmission module is used to send semantic signals to the target voice service.
[0030] Thirdly, this application provides a computer device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the speech processing method of the first aspect.
[0031] Fourthly, this application provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the speech processing method as described in the first aspect.
[0032] The beneficial effects of this invention are:
[0033] (1) In multi-round interaction scenarios, the target voice service is determined based on the priority configuration of the voice service and the interaction status;
[0034] (2) When the business scenario mentioned by the user's voice is unclear, the target voice service responds to the user's voice to avoid the interactive device being unable to respond to the user's voice, so that the interactive device can provide the user with the business scenario required by the user's voice. Attached Figure Description
[0035] The accompanying drawings are provided to further illustrate embodiments of the present invention and form part of the specification. They are used together with the following detailed description to explain the embodiments of the present invention, but do not constitute a limitation thereof. In the drawings:
[0036] Figure 1 A first flowchart of the speech processing method provided in this application embodiment is shown;
[0037] Figure 2 The following diagram illustrates an application example of the speech engine provided in this application embodiment;
[0038] Figure 3A second flowchart of the speech processing method provided in an embodiment of this application is shown;
[0039] Figure 4 A third flowchart of the speech processing method provided in the embodiments of this application is shown;
[0040] Figure 5 A schematic diagram of the structure of the voice processing device provided in an embodiment of this application is shown. Detailed Implementation
[0041] The specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are for illustration and explanation only and are not intended to limit the scope of the present invention.
[0042] The components of the embodiments of the invention described and illustrated herein can typically be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.
[0043] In the following, the terms “comprising,” “having,” and their cognates, which may be used in various embodiments of the invention, are intended only to indicate a particular feature, number, step, operation, element, component, or combination thereof, and should not be construed as excluding, firstly, the presence of one or more other features, numbers, steps, operations, elements, components, or combinations thereof, or adding the possibility of one or more features, numbers, steps, operations, elements, components, or combinations thereof.
[0044] Furthermore, the terms "first," "second," and "third" are used only to distinguish descriptions and should not be interpreted as indicating or implying relative importance.
[0045] Unless otherwise specified, all terms used herein (including technical and scientific terms) shall have the same meaning as commonly understood by one of ordinary skill in the art to which the various embodiments of the invention pertain. Terms (such as those defined in commonly used dictionaries) shall be interpreted as having the same meaning as in their contextual meaning in the relevant technical field and shall not be interpreted as having an idealized or overly formal meaning, unless clearly defined in the various embodiments of the invention.
[0046] Example 1
[0047] Please see Figure 1 , Figure 1A first flowchart of the speech processing method provided in an embodiment of this application is shown. Figure 1 The speech processing methods in the text include:
[0048] S110 retrieves the priority configuration and interaction status of each voice service.
[0049] Voice services refer to application services capable of recognizing and responding to user voice commands. These can include foreground controllable interfaces on interactive devices, audio playback, and navigation, etc., which will not be elaborated upon here. For ease of understanding, in the embodiments of this application, the voice service is a foreground controllable interface. Specifically, in multi-turn service scenarios, there may be situations where the interactive device displays multiple foreground controllable interfaces simultaneously. The interactive device can sequentially close the foreground controllable interfaces based on the user's multiple rounds of interactive voice commands. If the voice service mentioned by the user is unclear, it is impossible to determine which foreground controllable interface needs to be closed based on the user's voice, causing the interactive device to be unable to respond to the user's voice, and consequently, the interactive device cannot provide the user with the required service scenario based on the voice.
[0050] In the embodiments of this application, the priority configuration and interaction state of each voice service in a multi-round business scenario are obtained. The priority configuration and interaction state of each voice service are pre-set according to actual needs and are not limited here. The interaction state includes an executable state and a non-executable state. Specifically, when the interaction state of a voice service is executable, the voice service meets the requirement of responding to the user's voice and can send the semantic signals of the user's voice to the voice service. When the interaction state of a voice service is non-executable, the voice service does not meet the requirement of responding to the user's voice and does not send the semantic signals of the user's voice to the voice service. Based on the priority configuration and interaction state of each voice service, the target voice service that responds to the user's voice is determined to provide the user with the required business scenario.
[0051] S120: When the operating status of any voice service changes, the interaction status of the voice service is transmitted to the voice engine.
[0052] Typically, voice services operate in either an active or inactive state, and this state changes between active and inactive depending on user needs. In any given voice service state, the interaction status of the voice service is transmitted to the voice engine. Specifically, only the interaction status of the voice service whose state has changed can be uploaded, or the interaction status of all voice services can be uploaded simultaneously; this will not be elaborated upon here.
[0053] S130 generates semantic signals based on the received user voice.
[0054] In response to received user voice messages, audio recognition processing is performed to determine the semantics of the user voice messages, thereby generating semantic signals. This converts the user voice messages into semantic signals that can be recognized and processed by voice services, enabling voice services to respond to and process user voice messages. Ultimately, this allows interactive devices to provide the required service scenarios based on the user's voice.
[0055] In embodiments of this application, generating semantic signals based on received user speech includes:
[0056] The received user voice is converted into a signal, and the converted user voice is then denoised to generate a semantic signal.
[0057] Please see Figure 2 , Figure 2 The diagram illustrates an application example of the speech engine provided in this application embodiment.
[0058] For ease of understanding, an embodiment of this application provides a voice assistant 210. The voice assistant 210 is a computer application used to recognize and process user voice. Specifically, the user inputs analog voice signals through a microphone 220, which are then converted into digital voice signals. The voice assistant 210 performs noise reduction processing on the digital voice signals to obtain recognized audio. The voice assistant 210 transmits the recognized audio to a voice engine 230, which performs context processing and status identification on the recognized audio to generate semantic signals. The voice assistant 210 sends the semantic signals to a voice service 240, which responds to and processes the user voice, providing the user with the required service scenarios. It should be understood that the voice assistant 210 is also used to transmit the interaction status of each voice service 240 to the voice engine 230.
[0059] In the embodiments of this application, before obtaining the priority configuration and interaction state of each voice service, the method further includes:
[0060] Initialize the voice engine and each voice service.
[0061] After initializing the voice engine 230 and each voice service 240, the voice assistant 210 reads the priority configuration and interaction state of each voice service 240. By initializing the voice engine 230 and the voice services 240, default values are assigned to variables to facilitate user voice recognition. It should be understood that the voice assistant 210 can also initialize the voice engine 230 and each voice service 240 after reading their priority configuration and interaction state.
[0062] S140: If the voice engine determines that the interaction state of at least two voice services is executable, the target voice service is determined based on the priority configuration.
[0063] The voice engine determines which voice services can respond to user voice commands based on the interaction state of each voice service. If the voice engine determines that at least two voice services are in an executable state, then at least two voice services meet the requirement of responding to user voice commands. Based on the priority configuration of each voice service in an executable state, one of these voice services is selected as the target voice service.
[0064] S150 sends semantic signals to the target voice service.
[0065] Based on priority configuration, after determining the target voice service for which the user expects interaction, semantic signals are sent to the target voice service. The target voice service responds to and processes the user's voice, thereby enabling the interactive device to provide the user's desired service scenario. In multi-turn interaction scenarios, the target voice service is determined based on the priority configuration of the voice service and the interaction state. When the service scenario mentioned by the user's voice is unclear, the target voice service responds to the user's voice, avoiding the interactive device's inability to respond to the user's voice and enabling the interactive device to provide the user's required service scenario based on the user's voice.
[0066] Please see Figure 3 , Figure 3 A second flowchart of the speech processing method provided in an embodiment of this application is shown.
[0067] In the embodiments of this application, when the voice engine determines that the interaction state of at least two voice services is executable, the target voice service is determined based on priority configuration, including:
[0068] S141, if the voice engine determines that the interaction state of at least two voice services is executable, the voice service with the interaction state being executable and the priority being configured as high priority is identified as the target voice service.
[0069] For ease of understanding, the voice services in the embodiments of this application include voice service A and voice service B. When the running state of voice service A is switched to the active state, the interaction state of voice service A is determined to be executable, and the interaction state of voice service A is transmitted to the voice engine. At this time, only voice service A meets the requirement of responding to user voice. When the running state of voice service B is switched to the active state, the interaction state of voice service B is determined to be executable, and the interaction state of voice service B is transmitted to the voice engine. At this time, both voice service A and voice service B meet the requirement of responding to user voice.
[0070] After acquiring the user's voice, the voice engine identifies that the user's voice may request a response from voice service A, and may also request a response from voice service B. The voice engine then obtains the interaction state of each voice service and determines that the interaction states of voice service A and voice service B are both executable. The voice engine determines that the interaction states of at least two voice services are executable, and that the interaction states of voice service A and voice service B both meet the requirements for responding to the user's voice. Finally, it queries the priority configuration of voice services A and B.
[0071] Based on priority configuration, all voice services with an "executable" interaction state are prioritized to determine the number of voice services configured with high priority. If only one voice service is configured with high priority, it is designated as the target voice service. If voice service A is configured with high priority and voice service B is configured with low priority, voice service A, with both an "executable" interaction state and high priority, is designated as the target voice service. If voice service A is configured with low priority and voice service B is configured with high priority, voice service B, with both an "executable" interaction state and high priority, is designated as the target voice service.
[0072] In the embodiments of this application, when the voice engine determines that the interaction state of at least two voice services is executable, the target voice service is determined based on priority configuration, including:
[0073] If the voice engine determines that the interaction state of at least two voice services is executable, the target voice service is determined based on the priority configuration and the order in which the voice engine receives the interaction state of each voice service.
[0074] After acquiring the user's voice, the voice engine determines that both voice service A and voice service B are in an executable state. It then queries the priority configurations of voice services A and B. Based on these priority configurations, all voice services with executable interaction states are prioritized, and the number of highest-priority voice services is determined. If the number of highest-priority voice services is greater than one, the order in which the voice engine received the interaction states of each voice service is determined.
[0075] When voice services have the same priority, the target voice service is determined based on the order in which the voice engine receives the interaction states of each voice service. For ease of understanding, in the embodiments of this application, the voice engine receives the interaction state of voice service A first, and then receives the interaction state of voice service B. When both voice services A and B have high priority, or both have low priority, the voice engine receives the interaction state of voice service A first and determines voice service A as the target voice service.
[0076] In embodiments of this application, after sending the semantic signal to the target voice service, the method further includes:
[0077] In response to the semantic execution completion signal based on the target voice service, the interaction state of the target voice service is determined to be non-executable, and the interaction state of the target voice service is transmitted to the voice engine.
[0078] For ease of understanding, in the embodiments of this application, voice service A is the foreground controllable interface A, and voice service B is the foreground controllable interface B. The interactive device simultaneously displays foreground controllable interface A and foreground controllable interface B, meaning that both voice service A and voice service B are in an active state. In multi-turn interaction scenarios, the user can use their voice to close foreground controllable interface A and foreground controllable interface B respectively. When the user's voice indicates "close interface," since both voice service A and voice service B meet the requirements for responding to the user's voice, there is a possibility that voice service A will respond to the user's voice, and there is also a possibility that voice service B will respond to the user's voice.
[0079] In the embodiments of this application, voice service A is determined as the target voice service based on priority configuration. After the voice signal is sent to the target voice service, voice service A responds to and processes the user's voice, the foreground controllable interface A is closed, the interactive device no longer displays the foreground controllable interface A, and a semantic execution completion signal is generated. Responding to the semantic execution completion signal based on the target voice service, the interaction state of the target voice service is determined to be non-executable, and the interaction state of the target voice service is transmitted to the voice engine, avoiding repeated responses to user voice by the same voice service in multi-round interaction scenarios.
[0080] Please see Figure 4 , Figure 4 A third flowchart of the speech processing method provided in the embodiments of this application is shown.
[0081] In embodiments of this application, the voice processing method further includes:
[0082] S142, if the voice engine determines that the number of voice services with an interactive state of executable is equal to one, the voice service with an interactive state of executable is identified as the target voice service.
[0083] In multi-turn interaction scenarios, users can sequentially close foreground controllable interface A and foreground controllable interface B by repeatedly inputting their voice commands. When a user inputs their voice command for the first time, voice service A is designated as the target voice service based on priority configuration. Voice service A executes semantic signals, and its interaction state changes from executable to non-executable.
[0084] When a user inputs the same voice message a second time, the interaction state of voice service A is in an unexecutable state, while the interaction state of voice service B is in an executable state. At this point, only voice service B meets the requirement of responding to the user's voice message. If the voice engine determines that the number of voice services in an executable state is equal to one, then the voice service in an executable state is identified as the target voice service. Specifically, since only voice service B is in an executable state, voice service B is identified as the target voice service. In multi-turn interaction scenarios, even if the voice service mentioned by the user is unclear, the semantic signal of the user's voice will not be repeatedly sent to voice service A because the user inputs the same voice message. Instead, the foreground controllable interface A and the foreground controllable interface B are closed sequentially according to the user's voice. When multiple voice services are in an active state, the deactivation of one voice service will not affect the remaining voice services' ability to respond to the user's voice messages in multiple rounds of interaction, thus providing the user with the required service scenario based on the voice message.
[0085] This application provides a voice processing method, comprising: acquiring the priority configuration and interaction state of each voice service, wherein the interaction state includes an executable state and a non-executable state; transmitting the interaction state of the voice service to a voice engine when the running state of any voice service changes; generating a semantic signal based on the received user voice; determining a target voice service based on the priority configuration when the voice engine determines that the interaction states of at least two voice services are executable; and sending the semantic signal to the target voice service. In multi-turn interaction scenarios, the target voice service is determined based on the priority configuration and interaction state of the voice services. When the business scenario mentioned by the user voice is unclear, the target voice service responds to the user voice, avoiding the interaction device's inability to respond to the user voice, and enabling the interaction device to provide the user with the required business scenario based on the user voice.
[0086] Example 2
[0087] Please see Figure 5 , Figure 5 A schematic diagram of the structure of the voice processing device provided in an embodiment of this application is shown. Figure 5 The voice processing device 300 includes:
[0088] The configuration acquisition module 310 is used to acquire the priority configuration and interaction status of each voice service, wherein the interaction status includes an executable status and a non-executable status.
[0089] The status transmission module 320 is used to transmit the interaction status of the voice service to the voice engine when the operating status of any voice service changes.
[0090] The signal generation module 330 is used to generate semantic signals based on the received user speech;
[0091] The target determination module 340 is used to determine the target voice service based on priority configuration when the voice engine determines that the interaction state of at least two voice services is executable.
[0092] The signal transmission module 350 is used to send semantic signals to the target voice service.
[0093] In embodiments of this application, the voice processing device 300 further includes:
[0094] The first service determination module is used to determine the voice service with an executable interaction state as the target voice service when the number of voice services with an executable interaction state determined by the voice engine is equal to one.
[0095] In the embodiments of this application, the target determination module 340 is further configured to determine the target voice service based on priority configuration and the order in which the voice engine receives the interaction states of each voice service when the voice engine determines that the interaction states of at least two voice services are in an executable state.
[0096] In the embodiments of this application, the target determination module 340 is further configured to determine the voice service whose interaction state is executable and whose priority is configured as high priority as the target voice service when the voice engine determines that the interaction state of at least two voice services is executable.
[0097] In embodiments of this application, the voice processing device 300 further includes:
[0098] The unexecutable state upload module is used to respond to the semantic execution completion signal based on the target voice service, determine the interaction state of the target voice service as unexecutable, and transmit the interaction state of the target voice service to the voice engine.
[0099] In the embodiments of this application, the signal generation module 330 is further configured to convert the received user speech into a signal and perform noise reduction processing on the converted user speech to generate a semantic signal.
[0100] In embodiments of this application, the voice processing device 300 further includes:
[0101] The initialization module is used to initialize the voice engine and each voice service.
[0102] The voice processing device 300 is used to execute the corresponding steps in the above-described voice processing method. The specific implementation of each function will not be described in detail here. In addition, the optional examples in Embodiment 1 are also applicable to the voice processing device 300 in Embodiment 2.
[0103] This application also provides a computer device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the voice processing method as described in Embodiment 1.
[0104] In this embodiment, the configuration acquisition module 310, status transmission module 320, signal generation module 330, target determination module 340, and signal transmission module 350 are all stored as program units in the memory, and the processor executes the above-mentioned program units stored in the memory to realize the corresponding functions.
[0105] The processor contains a kernel, which retrieves the corresponding program units from memory. One or more kernels can be configured; adjusting kernel parameters can resolve issues where the system cannot provide the user's desired service scenarios based on voice input.
[0106] The memory may include non-permanent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM, and the memory includes at least one memory chip.
[0107] This application also provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the speech processing method as described in Embodiment 1.
[0108] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0109] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0110] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0111] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0112] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0113] Memory may include non-persistent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0114] Computer-readable storage media include both permanent and non-permanent, removable and non-removable media that can store information by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0115] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0116] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A speech processing method, characterized in that, include: Obtain the priority configuration and interaction status of each voice service, wherein the interaction status includes an executable state and a non-executable state; If the operating state of any of the aforementioned voice services changes, the interaction state of the voice service will be transmitted to the voice engine. Generate semantic signals based on the received user voice; If the voice engine determines that the interaction state of at least two voice services is executable, the target voice service is determined based on the priority configuration. Send the semantic signal to the target voice service; In response to the semantic execution completion signal based on the target voice service, the interaction state of the target voice service is determined to be non-executable, and the interaction state of the target voice service is transmitted to the voice engine.
2. The speech processing method according to claim 1, characterized in that, The method further includes: If the number of voice services whose interactive state is executable is determined by the voice engine to be equal to one, then the voice service whose interactive state is executable is determined as the target voice service.
3. The speech processing method according to claim 1, characterized in that, When the voice engine determines that the interaction state of at least two voice services is executable, determining the target voice service based on the priority configuration includes: When the voice engine determines that the interaction state of at least two voice services is executable, the target voice service is determined based on the priority configuration and the order in which the voice engine receives the interaction state of each voice service.
4. The speech processing method according to claim 1, characterized in that, When the voice engine determines that the interaction state of at least two voice services is executable, determining the target voice service based on the priority configuration includes: If the voice engine determines that the interaction state of at least two voice services is executable, the voice service with the executable interaction state and the priority configured as high priority is identified as the target voice service.
5. The speech processing method according to claim 1, characterized in that, The step of generating semantic signals based on the received user speech includes: The received user voice is converted into a signal, and the converted user voice is then denoised to generate a semantic signal.
6. The speech processing method according to claim 1, characterized in that, Before obtaining the priority configuration and interaction status of each voice service, the process also includes: Initialize the voice engine and each voice service.
7. A voice processing device, characterized in that, include: The configuration acquisition module is used to acquire the priority configuration and interaction status of each voice service, wherein the interaction status includes an executable status and a non-executable status; The status transmission module is used to transmit the interaction status of the voice service to the voice engine when the operating status of any of the voice services changes. The signal generation module is used to generate semantic signals based on the received user speech; The target determination module is used to determine the target voice service based on the priority configuration when the voice engine determines that the interaction state of at least two voice services is executable. A signal transmission module is used to send the semantic signal to the target voice service; The unexecutable state upload module is used to respond to the semantic execution completion signal based on the target voice service, determine the interaction state of the target voice service as unexecutable, and transmit the interaction state of the target voice service to the voice engine.
8. A computer device, characterized in that, The computer device includes a memory and a processor. The memory stores a computer program, which, when executed by the processor, implements the speech processing method as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the speech processing method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
User passive voice interaction method and device thereof, terminal, server and medium
CN113488033A