Voice processing method and device, storage medium and program product
By combining the voice content and preset voice intention judgment prompt words, and using the large language model to make multi-dimensional voice intention judgment, it solves the problem that the voice system is difficult to distinguish between normal conversations and voice commands in the wake-up state, achieving higher recognition accuracy and lower error triggering rate.
Patent Information
- Application Number
- CN202510340691.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-06-10
AI Technical Summary
The voice system is difficult to distinguish between normal conversations and voice commands of users in a wake-up state, resulting in accidentally triggering the machine function and reducing the service performance of the voice interaction system.
By obtaining the content of the input voice, combining the preset pronunciation intention to judge the prompt words, using a large language model to make multi-dimensional pronunciation intention judgments, and determining whether to respond to the input voice.
Effectively distinguish between normal conversations and voice commands of users, improve the recognition accuracy of the voice system in wake-up-free mode, reduce the chance of false triggering, and improve the accuracy of voice interaction.
Smart Images

Figure CN120126461A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of voice technology, and in particular, to a voice processing method, device, storage medium, and program product. Background Art
[0002] Voice wake-up is a key link in voice interaction technology. In traditional technologies, before a user issues a voice command, it is first necessary to use a wake-up word to wake up the in-vehicle voice recognition system from the standby state, which brings inconvenience to the user.
[0003] With the continuous development of voice interaction technology, some human-computer interaction scenarios need to support full-duplex interaction skills or wake-up-free functions to facilitate users' more convenient voice interaction operations. However, when the voice recognition system is in the wake-up-free state, it is often difficult to better distinguish between the user's normal conversation and voice interaction commands, especially resulting in accidental triggering of machine functions, seriously reducing the service performance of the voice interaction system in the wake-up-free state.
[0004] In response to the above problems, the industry has not yet proposed a better solution. Summary of the Invention
[0005] This application provides a voice processing method, device, storage medium, and program product to at least solve the problem that the voice system misidentifies the user's normal conversation as a voice command in the wake-up-free state.
[0006] In a first aspect, an embodiment of this application provides a voice processing method, including: obtaining the voice content corresponding to the input voice, and constructing a voice intention judgment prompt word according to the voice content and at least one preset voice intention judgment prompt template; determining a voice intention judgment result for at least one voice intention dimension based on a large language model and the at least one voice intention judgment prompt word; and determining whether to perform a response process on the input voice according to each of the voice intention judgment results.
[0007] In a second aspect, an embodiment of this application provides an electronic device, including: at least one processor, and a memory communicatively connected to the at least one processor, where the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the steps of the voice processing method according to any embodiment of this application.
[0008] In a third aspect, an embodiment of this application provides a storage medium, on which a computer program is stored, characterized in that when the program is executed by a processor, the steps of the voice processing method according to any embodiment of this application are implemented.
[0009] Fourthly, an embodiment of the present application provides a computer program product, including a computer program / instructions, which, when executed by a processor, implement the steps of the voice processing method according to any embodiment of the present application.
[0010] The beneficial effects of the embodiments of the present application are as follows: By combining the voice content and the preset voice intention judgment prompt words, and using the large language model to perform multi-dimensional voice intention judgment, it effectively distinguishes the user's normal conversation from the actual voice command, enables the voice system to accurately recognize the user's intention in the wake-free mode, reduces the probability of mis-triggering, and improves the accuracy of voice interaction. Description of the Drawings
[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0012] Figure 1 Shows a flowchart of an example of the voice processing method according to an embodiment of the present application; Figure 2 Shows according to Figure 1 An example of the operation flowchart of step S130 in Figure 3 Shows an effect schematic diagram of an example of the decision condition path according to an embodiment of the present application; Figure 4 Is a schematic structural diagram of an embodiment of the electronic device of the present application. Detailed Embodiments
[0013] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present application.
[0014] It should be noted that in order to implement the voice interaction function in the full-duplex interaction mode or the wake-up-free function, some current research scholars in the industry have proposed to use a rejection recognition model to recognize and process the user's voice, distinguish between invalid voice and valid voice, and filter out the invalid voice to reduce the impact of the invalid voice on the user. However, the rejection recognition model is usually trained based on a fixed corpus, and it cannot support the realization of diverse types of input voice content. In addition, due to limited model capabilities, it generally can only consider the current round of text.
[0015] Before the wide application and popularization of large model technology, since the range of content that can be interacted with is relatively narrow, in addition to vehicle control, media playback, and some common skills (such as calendars, reminders, stock queries, etc.), for the open domain, it can only cover topics such as chatting and encyclopedias, and its rejection recognition performance can still be applied.
[0016] However, with the integration of large models in the voice system, the content that people start to interact with the in-vehicle system is gradually becoming wider and deeper. For example, in addition to users starting to ask questions such as "Has the rocket of ** launched?" and "Has ** been elected?", they will also ask brain teasers such as "Is sashimi a dead fish slice?" to test artificial intelligence in reverse. In this way, the range of content requested for interaction with the voice system is more extensive, resulting in a more prominent problem of insufficient original training corpus. For example, in the past, it usually only required vehicle control or navigation to be executed. However, after the integration of large models, it is required to trigger interaction functions such as writing poems, writing compositions, and playing games through the voice system.
[0017] In addition, in the interaction scenarios of the voice system integrated with large models, it is also necessary to consider the response to some declarative sentences, such as "I went to work yesterday" and "I just had a meal". Therefore, the rejection recognition model that could operate well in the full-duplex mode originally can no longer effectively judge whether a sentence is a valid voice request.
[0018] For example, for the sentence "My friend got married", it usually undergoes rejection recognition in the traditional rejection recognition model. However, when a large model is integrated into the voice system, it may be an interaction between the user and the voice system, and it is necessary to analyze the context to identify whether rejection recognition should be performed. For example, if the previous text mentions "I'm very happy today" and the dialogue system asks "Why are you happy?", then it should not be rejected; however, if the previous text is a "vehicle control instruction", then rejection recognition is required. Another example: "Who is your uncle" usually undergoes rejection recognition in the traditional model. However, after integrating a large model and setting up user profiles in the voice system, it should not be rejected.
[0019] Therefore, in the face of a voice system integrated with a large model, the traditional rejection recognition model trained based on a fixed corpus often fails to meet the requirements, and single-round voice text analysis cannot consider the context, resulting in false rejection recognition or false triggering.
[0020] In some currently relevant technologies, some experts and scholars have also proposed to directly implement the rejection recognition function by using a large model. For example, directly specifying that the large model performs the rejection recognition detection task for the input speech. However, the large model is prone to hallucinatory and unreliable outputs. In addition, when using COT (Chain of Thought) to reduce hallucinations and improve accuracy, it will increase additional time consumption, and the prompt words are too long, resulting in non-compliance when the large model executes instructions.
[0021] It should be noted that the purpose of the above description of the currently relevant technologies is only to facilitate the public's better understanding of the inventive spirit and motivation of this application, and is not regarded as a limitation of this application. In addition, the technical solutions described in the above currently relevant technologies are not prior art, and they can also be unpublished technical solutions, such as those under research or in the laboratory stage.
[0022] In view of this, Figure 1 A flowchart showing an example of the speech processing method according to an embodiment of the present application is shown.
[0023] As Figure 1 shown, in step S110, the speech content corresponding to the input speech is obtained, and a speech intention judgment prompt word is constructed according to the speech content and at least one preset speech intention judgment prompt template.
[0024] Here, the input speech can be collected through a vehicle-mounted speech recognition system or the microphone of other speech interaction devices. After collection, the speech signal first passes through a speech signal processing module, such as denoising, feature extraction, acoustic model recognition, etc., to be converted into text content (text transcription), and noise and background interference are eliminated as much as possible.
[0025] In some embodiments, through natural language processing (NLP) technology, the input speech is initially analyzed for content text, and key information therein is identified, such as the topic, person, time, place, action, etc. Furthermore, these key information are used to fill the semantic slots of the speech intention judgment prompt template, thereby obtaining the corresponding speech intention judgment prompt word.
[0026] It should be understood that the type or business application scenario of the speech intention judgment prompt template can be diversified and can be set according to the business type of the speech system (such as in-vehicle intelligent cockpit, smart speaker, etc.), triggering the call of the large language model to judge and recognize the speech intention in a specific dimension to adapt to the analysis of different contexts.
[0027] In step S120, based on the large language model and at least one speech intent judgment prompt word, determine the speech intent judgment results for at least one speech intent dimension.
[0028] Here, the types of large language models can be diverse, and open or closed large language models can be adopted, such as the GPT series or qwen, etc., enabling the large language model to perform in-depth semantic reasoning on the prompt words based on the speech content and understand the potential complex intents in speech interactions.
[0029] In an example of the embodiment of the present application, one speech intent judgment prompt word can correspond to multiple or all speech intent dimensions, so that multi-dimensional intent judgment analysis can be achieved through a single speech intent interaction. However, in the case of a large number of speech intent dimensions, it may lead to a situation where the prompt word is too long, resulting in the output result of the large model may not comprehensively cover.
[0030] In another example of the embodiment of the present application, each speech intent judgment prompt template has a uniquely corresponding speech intent dimension. In this way, by separately invoking the large model based on different prompt words, the large model can make predictions for the corresponding speech intent dimension each time, thereby improving the accuracy of the speech intent recognition result.
[0031] It should be understood that the dimensions of speech intent judgment can be diverse, such as authentication-related, emotion intent recognition, sentence pattern recognition, etc., and should not be limited here.
[0032] Here, the large language model combines each speech intent judgment prompt word, utilizes its excellent natural language understanding ability, analyzes and judges each speech intent dimension, and in some cases, can also comprehensively analyze by combining the speech content corresponding to the input speech in the current round with the speech content or interaction results of at least one historical round, thereby improving the accuracy of the speech intent judgment result.
[0033] In some embodiments, obtain the speech input content and speech processing results corresponding to at least one historical round. Further, based on the large language model, at least one speech intent judgment prompt word, and the speech input content and speech processing results corresponding to each historical round, determine the speech intent judgment results for at least one speech intent dimension.
[0034] More specifically, in a voice interaction system, the historical turn refers to the context information related to the user's previous voice input, including the user's voice content and the system's response (processing result). By obtaining the voice input and voice processing results of the historical turn, the system can better understand the context of the current conversation. Furthermore, by combining with a large language model and comprehensively considering the current voice input, the voice input and processing results of the historical turn, the voice intention judgment is triggered through prompt words, thereby integrating the historical turn information and the voice processing results, enabling the system to continuously understand and track the user's needs, and thus more accurately predict the user's intention.
[0035] Preferably, the voice input content and voice processing results of the historical turn are only fused in the voice intention judgment analysis for one or more predefined voice intention dimensions, such as the context association dimension, etc.; in addition, for other undefined voice intention dimensions (such as the sentence pattern of the input statement), it may not need to fuse historical interaction information, thereby improving the voice processing efficiency.
[0036] In step S130, according to each voice intention judgment result, it is determined whether to perform response processing on the input voice.
[0037] In some embodiments, the voice system summarizes the voice intention judgment results from the large language model, and then comprehensively evaluates whether to perform response processing on the input voice. Exemplarily, by comprehensively evaluating whether the voice intention is a voice service instruction or a normal conversation of the user, for example, response processing is performed in the case of a voice service instruction, while no response or rejection processing is performed when it is analyzed that it belongs to the user's normal conversation.
[0038] Thus, through the intelligent decision-making mechanism, the system can flexibly handle various voice inputs in the full-duplex interaction mode, while avoiding the interference of invalid voices. By comprehensively considering the multi-dimensional analysis of voice intention and context information, the system can not only execute effective instructions in the wake-free state, but also effectively filter out unnecessary voice interference, ensuring the fluency and accuracy of the interaction service of the voice system.
[0039] Regarding the implementation details of step S120, in some examples of the embodiments of the present application, each voice intention judgment prompt word is respectively input into one or more large language models to parallelly determine the corresponding voice intention judgment results. Specifically, each voice intention judgment prompt word is respectively used to call the corresponding large language model through a preset plurality of large model API interfaces. Thus, through the parallel processing of the voice intention judgment prompt words, the voice intentions of multiple dimensions can be processed in a short time, greatly improving the analysis efficiency for diverse types of voice intentions and complex context requirements.
[0040] Figure 2 Shows according to Figure 1An operation flowchart example of step S130 in
[0041] As Figure 2 shown, in step S210, each voice intent judgment result is matched with each decision condition node in the preset decision condition path to determine the corresponding target processing flow.
[0042] Here, the decision condition path includes multiple decision condition nodes, where the first decision condition node is used to transfer to the second decision condition node or the voice feedback node according to the condition judgment result. Each decision condition node is respectively used to indicate the corresponding voice intent dimension, and the voice feedback node is used to indicate the rejection feedback action or the response feedback action.
[0043] It should be noted that the decision condition path includes multiple decision condition nodes, and each node represents a judgment logic or condition for voice intent, aiming to guide the system to the appropriate processing flow according to different voice input types and intents. In some examples, the decision path can adopt a tree structure, including multiple branches, and each branch represents a different decision path. Each path is gradually refined according to different voice intent judgment results (such as "command", "query", "rejection", etc.).
[0044] It should be understood that the first decision condition node and the second decision condition node are not specific, and they can refer to any node in the decision condition path. The voice intent dimension is judged through the first decision condition node, and the second decision condition node makes a further refined judgment according to the result of the first decision node.
[0045] In addition, the voice feedback node can be used to indicate the final feedback action, and can be divided into a rejection voice feedback node and a response voice feedback node. Through the rejection voice feedback node, the voice system can be triggered to refuse to perform interactive feedback, and through the response voice feedback node, the voice system can be triggered to execute response feedback, such as playing music, starting a small game, or calling a search engine, etc.
[0046] In some embodiments, according to the voice intent dimensions involved in each decision condition node in the decision condition path, each voice intent judgment result is matched in sequence, so as to automatically extract the target processing flow that matches the current input voice from the decision condition path. It should be understood that not all voice intent judgment results may be applied to the matching. For example, when a voice feedback node is found through the retrieval and matching of the decision condition path, even if there are still some voice intent judgment results not adopted, the corresponding target processing flow can be directly output.
[0047] In step S220, according to the voice feedback node in the target processing flow, it is determined whether to perform response processing on the input voice.
[0048] Exemplarily, if the voice feedback node in the target processing flow is a response voice feedback node, the input voice is processed for response to trigger the interaction function of the voice system. Additionally, if the voice feedback node in the target processing flow is a rejection recognition voice feedback node, the input voice is processed for rejection recognition to avoid false triggering.
[0049] Through the embodiments of the present application, by matching the voice intent judgment result with the condition nodes in the decision condition path one by one, the corresponding decision process can be obtained more accurately, avoiding the hallucination caused by directly performing voice rejection recognition detection using a large language model, and improving the accuracy of the rejection recognition detection result. In addition, based on the multi-level node design of the decision condition path, complex business requirements can be effectively supported, and misoperations or misfeedbacks can be avoided through refined path topology design. Thus, by dynamically matching the decision condition nodes and performing response or rejection recognition feedback, the voice system can operate accurately according to the voice content, context, and intent input by the user, reducing misjudgments and redundant responses, and greatly enhancing the intelligent level of the voice interaction system.
[0050] Figure 3 The effect schematic diagram of an example of the decision condition path according to the embodiments of the present application is shown.
[0051] Specifically, the voice intent dimension includes at least one of the following: whether the voice content can be understood, whether the voice content can be understood in combination with the context, whether the voice content is related to at least one preset voice interaction function, whether the voice content is an imperative sentence, whether the voice content is an interrogative sentence, or whether the voice content is a declarative sentence.
[0052] In the example as Figure 3 shown, a hierarchical design is introduced, decomposing the voice rejection recognition detection task into multiple dimensions and making a comprehensive decision. Through step-by-step decomposition and decision-making of multiple voice intent dimensions, the system can more intelligently judge the user's voice intent and achieve accurate response in the full-duplex interaction mode or the hands-free wake-up mode.
[0053] In some embodiments, the voice system decomposes the voice intent judgment process into multiple levels. Through multi-dimensional analysis of the voice input, the validity and specific voice intent of the input are checked or confirmed layer by layer. In addition, each dimension generates a specific judgment result based on an independent voice intent judgment prompt word and makes a step-by-step decision in the decision path according to the level. The dimension judgment of each layer will flow to the subsequent nodes, forming a clear logical processing chain, and finally determining the voice feedback node (response feedback or rejection recognition feedback).
[0054] The specific dimension types are as follows: Dimension 1: Whether the sentence can be understood; Dimension 2: Whether the current input content can be understood after combining the context; Dimension 3: Whether it is a traditional voice service (vehicle control, media playback, navigation, etc.); Dimension 4: Is it an imperative sentence, such as a request to perform a certain requirement (such as writing poetry, writing essays, and arithmetic); Dimension 5: Is it a question sentence, such as "Why is the sky blue?" Dimension six: Is it a declarative sentence?
[0055] In this way, we no longer rely on a single prompt word, but use different prompt words to generate different judgment dimensions, and each dimension task is not complicated, which can avoid the probability of hallucination in large models. In addition, when the semantics are complete and the sentences can be understood, the context is used to filter the conversation between people. The rejection ability in full-duplex state is improved. In addition, the use of parallel multi-dimensional output avoids the high time consumption problem originally brought by COT.
[0056] like Figure 3 As shown, after obtaining the multi-dimensional large model execution results, the large model execution results of each dimension are matched with the order of each node in the decision path.
[0057] First, identify whether the sentence itself is understandable. (Filter out disordered and messy content, such as environmental human voices) Here, the first step of the speech system is to determine whether the speech input can be correctly recognized and understood. If the user's speech is transcribed into text, but there are serious recognition errors or semantic confusion (such as meaningless text caused by noise, punctuation, dialect, etc.), the system directly transfers the speech input to the rejection feedback node.
[0058] Exemplarily, if the speech content cannot be understood, rejection feedback (such as "Sorry, I didn't hear it clearly, please say it again") is directly output; if the speech can be understood, it flows to the analysis of the next decision condition node.
[0059] Then, when the sentence itself can be understood, determine whether the content expressed by the sentence is complete. (Filter out conversations between people, such as "My uncle bought it, it's quite fun") Here, the system combines the speech input and processing results of the historical rounds in the second layer to determine whether the current speech has a clear meaning in the context. In this way, the historical round information and context are used as inputs and combined with the current speech content for judgment. If the intention of the speech cannot be determined after combining the context (such as insufficient context or ambiguous content), the system will flow to the rejection feedback node; otherwise, it will flow to the analysis of the next decision condition node.
[0060] In addition, the voice content can be matched with traditional voice services. For example, it can be determined whether the voice content is related to preset traditional voice interaction functions, such as vehicle control (e.g., "Open the window"), media playback (e.g., "Play music"), navigation (e.g., "Navigate to the gas station"), etc. If the voice content matches a traditional voice service, it will be transferred to the corresponding function processing flow; otherwise, a higher-level dimensional matching analysis will be performed.
[0061] Then, if it is a declarative sentence, consider whether it is contextually related rather than a sentence that suddenly appears. (Mainly to filter out human conversations as well). Here, first, the sentence type of the voice content will be matched and detected. For example, during analysis, determine whether it is a declarative sentence, an imperative sentence, or an interrogative sentence, etc. If the voice content is not a declarative sentence, such as "What's the weather like today?" or "Calculate 123 multiplied by 456", it will be directly transferred to the response voice feedback node to form a complete processing flow.
[0062] Through the embodiments of the present application, the voice intention judgment task is disassembled into multiple independent dimensions through hierarchical design, enabling the voice interaction system to have higher flexibility, intelligence, and accuracy. By gradually matching the condition nodes of each dimension, the voice system can accurately judge the intention of the user's voice and provide personalized and accurate responses in combination with context and historical information. At the same time, it supports multiple feedback mechanisms (rejection feedback and response feedback), improving the user experience of the full-duplex voice interaction system.
[0063] Furthermore, in the case where the recognized voice content belongs to a declarative sentence, it is identified whether it is contextually related through the context association decision node. When there is no context correlation, a post-processing is performed once to trigger a re-judgment of the intention and decision matching, preventing missed responses caused by input content deviations.
[0064] Specifically, by combining the current voice input content with the voice content of the previous rounds (including input and response), analyze whether there is a semantic association. Exemplarily, the user says in the first round: "I am very happy.", and in the second round: "I bought a new car today.", and the two are contextually related. Additionally, if the user says "Open the navigation" in the first round and "I went to the supermarket yesterday" in the second round, then the two are not contextually related.
[0065] If the context is related, the voice system will be transferred to the response voice feedback node or the rejection voice feedback node, and it is determined whether to execute a response or rejection based on the specific intention. If the context is not related, it will be transferred to the post-processing node for content supplementation or semantic correction to re-trigger the intention judgment, avoiding missed recognition or misresponse of the system caused by semantic fragmentation and improving the robustness of the system.
[0066] In some examples of the embodiments of the present application, the decision condition path further includes a post-processing node for content completion of the speech content, where the post-processing node is used to re-trigger the speech intent judgment for each speech intent dimension and the matching for the decision condition path based on the completed speech content.
[0067] Exemplarily, if the user inputs "masterpieces of Li Bai" or "scenic spots nearby", through the context, it can be completed to "what are the masterpieces of Li Bai" or "what are the scenic spots nearby", so as to complete the content ignored due to the artificial expression habit during human-computer interaction. By completing or correcting the sentence, the semantic missing of the user input content can be compensated, and the situation of missed recognition can be avoided.
[0068] Referring to Figure 3 the examples in, multiple decision condition nodes include a declarative sentence decision node and a context-related decision node. If the matching result of the declarative sentence decision node indicates that the speech content is not a declarative sentence, it flows to the response speech feedback node to perform response processing on the input speech. If the matching result of the declarative sentence decision node indicates that the speech content is a declarative sentence, it flows to the context-related decision node.
[0069] On the other hand, if the matching result of the context-related decision node indicates that the speech content is context-unrelated, it flows to the post-processing node. If the matching result of the context-related decision node indicates that the speech content is context-related, it flows to the response speech feedback node or the rejection recognition speech feedback node.
[0070] Through the embodiments of the present application, when the speech content is clear, context-related, and conforms to the system function, the speech interaction instruction of the corresponding declarative sentence is executed, such as playing music, turning on navigation, etc. In addition, when the declarative sentence is context-unrelated, the system flows to the post-processing node and uses semantic completion or correction to re-trigger the intent judgment, avoiding missed recognition due to incomplete sentences or ambiguous semantics. Thus, the high accuracy of the rejection operation of the declarative sentence by the speech system based on the large model is achieved, and the naturalness and user experience of multi-round speech interaction are improved.
[0071] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of actions combined, but those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps can be in other sequences or performed simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application. In the above embodiments, each embodiment is described with emphasis. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0072] In some embodiments, the embodiments of the present application provide a non-volatile computer-readable storage medium, in which one or more programs including execution instructions are stored, and the execution instructions can be read and executed by an electronic device (including but not limited to a computer, a server, or a network device, etc.) to execute any one of the above voice processing methods of the present application.
[0073] In some embodiments, the embodiments of the present application further provide a computer program product, which includes a computer program stored on a non-volatile computer-readable storage medium. The computer program includes program instructions, and when the program instructions are executed by a computer, the computer is enabled to execute any one of the above voice processing methods.
[0074] In some embodiments, the embodiments of the present application further provide an electronic device, which includes: at least one processor, and a memory communicatively connected to the at least one processor. Wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute a voice processing method.
[0075] Figure 4 FIG. 10 is a schematic hardware structure diagram of an electronic device for executing a voice processing method provided in another embodiment of the present application. As Figure 4 shown, the device includes: one or more processors 410 and a memory 420. Figure 4 Here, one processor 410 is taken as an example.
[0076] The device for executing the voice processing method may further include: an input device 430 and an output device 440.
[0077] The processor 410, the memory 420, the input device 430, and the output device 440 may be connected by a bus or other means. Figure 4 Here, being connected by a bus is taken as an example.
[0078] The memory 420, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the voice processing method in the embodiments of the present application. The processor 410 executes various functional applications and data processing of the server by running the non-volatile software programs, instructions, and modules stored in the memory 420, that is, implements the voice processing method in the above method embodiments.
[0079] The memory 420 may include a program storage area and a data storage area. The program storage area may store an operating system and application programs required for at least one function. The data storage area may store data created according to the use of the electronic device and the like. In addition, the memory 420 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 420 may optionally include a memory remotely disposed relative to the processor 410, and these remote memories may be connected to the electronic device through a network. Examples of the above networks include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0080] The input device 430 may receive input digital or character information, and generate signals related to user settings and function controls of the electronic device. The output device 440 may include a display device such as a display screen.
[0081] The one or more modules are stored in the memory 420, and when executed by the one or more processors 410, execute the voice processing method in any of the above method embodiments.
[0082] The above product may execute the method provided in the embodiments of the present application, and has function modules and beneficial effects corresponding to the execution of the method. For technical details not described in detail in this embodiment, reference may be made to the method provided in the embodiments of the present application.
[0083] The electronic device in the embodiments of the present application exists in various forms, including but not limited to: (1) Mobile communication devices: These devices are characterized by having mobile communication functions and mainly aim to provide voice and data communication. Such terminals include: smart phones, multimedia phones, functional phones, and low-end phones, etc.
[0084] (2) Ultra-mobile personal computer devices: These devices belong to the category of personal computers, have computing and processing functions, and generally also have the characteristic of mobile Internet access. Such terminals include: PDAs, MIDs, and UMPC devices, etc.
[0085] (3) Portable entertainment devices: These devices can display and play multimedia content. Such devices include: audio and video players, handheld game consoles, e-books, and smart toys and portable vehicle navigation devices.
[0086] (4) Other on-board electronic devices with data interaction functions, such as in-vehicle device installed on a vehicle.
[0087] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0088] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the related technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0089] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A speech processing method, comprising: Acquire the speech content corresponding to the input speech, and construct a speech intention judgment prompt word according to the speech content and at least one preset speech intention judgment prompt template; Determining a speech intention judgment result for at least one speech intention dimension based on the large language model and the at least one speech intention judgment prompt word; According to each of the speech intention judgment results, determine whether to respond to the input speech.
2. According to the method of claim 1, the step of determining the speech intent judgment result for at least one speech intent dimension based on the large language model and the at least one speech intent judgment prompt word comprises: Obtaining speech input content and speech processing results corresponding to at least one historical round; Based on the large language model, the at least one speech intent judgment prompt word, and the speech input content and speech processing results corresponding to each of the historical rounds, a speech intent judgment result for at least one speech intent dimension is determined.
3. The method according to claim 1, wherein: The voice intention dimension includes at least one of the following: whether the voice content can be understood, whether the voice content can be understood in context, whether the voice content is related to at least one preset voice interaction function, whether the voice content is an imperative sentence, whether the voice content is an interrogative sentence, or whether the voice content is a declarative sentence.
4. The method according to any one of claims 1 to 3, wherein: The determining of a speech intention judgment result for at least one speech intention dimension based on the large language model and the at least one speech intention judgment prompt word includes: Each of the speech intention judgment prompt words is input into multiple large language models respectively to determine the corresponding speech intention judgment results in parallel.
5. The method according to claim 3, wherein: Determining whether to respond to the input voice according to each of the voice intention judgment results includes: Matching each of the speech intention judgment results with each of the decision condition nodes in the preset decision condition path to determine the corresponding target processing flow; the decision condition path includes a plurality of decision condition nodes, wherein the first decision condition node is used to flow to the second decision condition node or the speech feedback node according to the condition judgment result; each of the decision condition nodes is used to indicate the corresponding speech intention dimension, and the speech feedback node is used to indicate a rejection feedback action or a response feedback action; According to the speech feedback node in the target processing flow, it is determined whether to perform response processing on the input speech.
6. The method according to claim 5, wherein: The decision condition path also includes a post-processing node for completing the speech content, wherein the post-processing node is used to re-trigger the speech intent judgment for each of the speech intent dimensions and the matching for the decision condition path based on the completed speech content.
7. The method according to claim 6, wherein: The multiple decision condition nodes include declarative sentence decision nodes and context-related decision nodes; If the matching result of the declarative sentence decision node indicates that the speech content is not a declarative sentence, the flow is transferred to the response speech feedback node to perform response processing on the input speech; and if the matching result of the declarative sentence decision node indicates that the speech content is a declarative sentence, then flowing to the context association decision node; If the matching result of the context-related decision node indicates that the speech content is context-independent, then the flow is transferred to the post-processing node; And if the matching result of the context-related decision node indicates that the speech content is context-related, the flow is transferred to the response speech feedback node or the rejection speech feedback node.
8. A storage medium having a computer program stored thereon, wherein: When the program is executed by a processor, the steps of the method described in any one of claims 1 to 7 are implemented.
9. An electronic device, comprising: At least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the steps of the method described in any one of claims 1 to 7.
10. A computer program product, comprising a computer program / instruction, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 7.
Citation Information
Cited By
Voice control method and voice control system
CN121171225A