Off-line and on-line voice module user experience optimization method, storage medium and electronic device

By receiving and parsing voice commands in the offline and online voice modules, selecting local interaction transition words, and dynamically adjusting the feedback content, the problem of awkward user interactions is solved, and a more natural and fast voice interaction experience is achieved.

CN120636384APending Publication Date: 2025-09-12QINGDAO HAIER TECH +3

Patent Information

Application Number
CN202510731932.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

When the existing offline and online voice modules are in poor network conditions or have a long execution time, the user interaction experience is stiff and the response time is long, resulting in a poor user experience.

Method used

By receiving voice commands, recognizing and parsing voice commands, selecting interactive transition words from the local corpus, executing commands and providing timely feedback on results, dynamically adjusting interactive transition words and content, and preloading or caching resources to improve response speed.

Benefits of technology

It optimizes the fluency and accuracy of voice interaction, shortens response time, enhances user experience, reduces dependence on the network, and improves system availability and reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120636384A_ABST
    Figure CN120636384A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, and provides an off-line and on-line voice module user experience optimization method, a storage medium and an electronic device, and the method comprises the steps: receiving a voice instruction inputted by a user; recognizing and analyzing the voice instruction to obtain a voice instruction classification; according to the analyzed voice instruction classification, selecting and playing corresponding interaction transition words from a local corpus; executing the analyzed voice instruction; and after the instruction is executed, feeding back an execution result to the user. According to the method and the device, the problems of rigid interaction and poor user experience of the off-line and on-line voice modules in the prior art are solved, and the interaction fluency and the optimization of the user experience are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a method for optimizing user experience of an online and offline voice module, a storage medium, and an electronic device. Background Art

[0002] With the rapid development of artificial intelligence (AI) technology, intelligent voice modules have become widely used in various smart devices, becoming a crucial interface for user-device interaction. These intelligent voice modules not only support offline voice recognition but also connect to cloud servers via the Internet, enabling a richer range of functions and services. However, in actual applications, the user experience of online voice modules still has some shortcomings.

[0003] Existing intelligent voice modules typically follow a fixed sequence of steps during user-device interaction: First, an algorithm identifies and reports the user's voice command; next, the algorithm parses the data and packages it into commands; then, the backend parses the data, executes the command, and returns the result; finally, the backend interprets the result and plays it back to the user via TTS (Text To Speech) technology. This sequence of steps provides a relatively smooth user experience when the network is good and execution efficiency is high. However, in poor network conditions or when command execution takes a long time, users have to wait for a response, resulting in a clunky interaction and a poor user experience. Summary of the Invention

[0004] The present application provides an offline voice module user experience optimization method, storage medium and electronic device, which are used to solve the problems of stiff interaction and poor user experience of offline voice modules in the prior art, and achieve optimization of interaction fluency and user experience.

[0005] This application provides a method for optimizing the user experience of an offline and online voice module, including: A method for optimizing user experience of an offline and online voice module, comprising: Receive voice commands input by the user; Recognize and analyze voice commands to obtain voice command classification; According to the classification of the parsed voice commands, the corresponding interactive transition words are selected and played from the local corpus; Execute the parsed voice command; After the instruction is executed, the execution result is fed back to the user.

[0006] According to a method for optimizing user experience of an offline and online voice module provided in the present application, the identification and parsing of voice commands specifically include: preprocessing the received voice commands; extracting feature parameters from the preprocessed voice signal; and matching the extracted feature parameters with a predefined voice command model to identify and parse the user's voice commands.

[0007] According to a method for optimizing user experience of an offline and online voice module provided by the present application, the method selects and plays corresponding interaction transition words from a local corpus based on the parsed voice command classification, specifically including: determining the corresponding interaction scenario based on the voice command classification; retrieving interaction transition words matching the interaction scenario from the local corpus; and playing the interaction transition words.

[0008] According to a method for optimizing user experience of an offline and online voice module provided by the present application, the categories of the voice command classification include control, query, entertainment and setting categories; according to the voice command classification, the corresponding interaction scenarios are determined, specifically including: determining the corresponding device operation scenario according to the control-class instructions, the interaction transition words of the device operation scenario include an operation confirmation phrase; determining the corresponding information retrieval scenario according to the query-class instructions, the interaction transition words of the information retrieval scenario include a retrieval waiting prompt; determining the corresponding media playback scenario according to the entertainment-class instructions, the interaction transition words of the media playback scenario include a playback preparation prompt; determining the corresponding parameter adjustment scenario according to the setting-class instructions, the interaction transition words of the parameter adjustment scenario include an adjustment confirmation phrase.

[0009] According to a method for optimizing user experience of an offline and online voice module provided in the present application, after playing the interactive transition words, the method further includes: monitoring the execution progress of the voice instructions; and dynamically adjusting the played interactive transition words or updating the played content according to the execution progress of the voice instructions.

[0010] According to a method for optimizing user experience of an offline and online voice module provided in the present application, the method of feeding back the execution result to the user specifically includes: generating voice feedback content corresponding to the instruction execution result; converting the voice feedback content into a voice signal through voice synthesis technology; and playing the voice signal.

[0011] According to a method for optimizing the user experience of an offline and online voice module provided in the present application, after feeding back the execution results to the user, the method further includes: predicting the user's subsequent instructions based on the user's historical interaction data and the current interaction scenario; and preloading or caching resources or data related to the subsequent instructions based on the prediction results to improve the response speed of the subsequent instructions.

[0012] This application also provides an offline and online voice module user experience optimization device, comprising: A voice receiving module, used to receive voice commands input by the user; The voice analysis module is used to recognize and analyze voice commands and obtain voice command classification; The transition word playback module is used to select and play corresponding interactive transition words from the local corpus according to the parsed voice command classification; An execution module, used to execute the parsed voice instructions; The feedback module is used to feedback the execution results to the user after the instruction is executed.

[0013] According to an offline and online voice module user experience optimization device provided by the present application, the voice analysis module includes: a voice preprocessing unit for preprocessing received voice commands; a feature extraction unit for extracting feature parameters from the preprocessed voice signal; and a pattern matching unit for matching the extracted feature parameters with a predefined voice command model to identify and parse the user's voice commands.

[0014] The present application also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute any of the above-mentioned methods for optimizing user experience of the offline and online voice modules through the computer program.

[0015] The present application also provides a computer-readable storage medium, which includes a stored program, wherein when the program is run, it is executed to implement any of the above-mentioned methods for optimizing user experience of the online and offline voice modules.

[0016] The present application also provides a computer program product, including a computer program, which, when executed by a processor, implements any of the above-mentioned methods for optimizing user experience of offline and online voice modules.

[0017] The method, storage medium, and electronic device for optimizing the user experience of an offline and online voice module, provided in this application, effectively understand user needs by receiving and accurately recognizing user voice commands. They intelligently select and play interactive transition words based on command classification, making the interaction process more natural and smooth. By executing the parsed voice commands and providing timely feedback on the results to the user, this not only improves the accuracy and efficiency of voice interaction but also significantly enhances the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the present application or the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0019] Figure 1 Schematic diagram of the hardware environment of a method for optimizing user experience of an offline and online voice module according to an embodiment of the present application; Figure 2 This is a flow chart of the method for optimizing the user experience of the offline and online voice modules provided by this application; Figure 3 This is a flow chart of the method for optimizing the user experience of an offline and online voice module using air conditioning as an example provided by this application; Figure 4 This is a structural diagram of the user experience optimization of the offline and online voice module provided by this application; Figure 5 It is a structural diagram of the electronic device provided in this application. DETAILED DESCRIPTION

[0020] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.

[0021] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0022] The main steps of the intelligent voice module during user interaction are summarized as follows: 1. Algorithm identification and reporting, 2. Algorithm data analysis and command packaging, 3. The backplane parses the data, executes and returns the execution results. 4. Analysis of the bottom board execution results, 5. Perform TTS based on the analysis results (only for online voice module). 6. Play the reply audio.

[0023] Only step 6 is user-perceptible. In a poor network environment (affecting the execution time of step 5) or a long execution time (affecting the execution time of step 3), the user needs to wait a long time to receive a response after speaking the command, which is a very awkward experience.

[0024] The existing technology only plays the interactive response after executing the entire process described above. This can result in a poor user experience in complex environments. Depending on the device, the response time is approximately 650-700ms offline and 1.8-2.3s online. The present invention adds interactive transition words. After parsing the algorithm data in step 2, the pre-defined prompt is played first. This not only makes the interaction less awkward, but also shortens the response time of the voice module, improving the user experience.

[0025] For example, in the prior art, after the user wakes up, he says the command word "cooling mode". After completing the entire operation as above, the online module broadcasts "Switched to cooling mode for you", which takes 2.0 seconds.

[0026] This application mainly solves the problem that the intelligent voice module lacks the agility and slow response of smart devices during the interaction with users, especially in an environment with poor network conditions, the voice module responds slowly and the interaction is stiff, which brings a poor experience to users.

[0027] Using this solution, after waking up, the user says the command "Cooling mode." After the algorithm analysis is completed (400ms), the online voice module plays the local offline corpus "OK." After executing subsequent operations, the online voice module then plays "Switched to Cooling mode for you." This reduces the user interaction time from 2 seconds to 400ms, greatly improving the user experience.

[0028] The following is a detailed description with reference to the embodiments.

[0029] According to one aspect of the embodiment of the present application, a method for optimizing the user experience of an online and offline voice module is provided. The method for optimizing the user experience of an online and offline voice module is widely used in smart home (Smart Home), smart home, smart home device ecology, smart residence (Intelligence House) ecology and other whole-house intelligent digital control application scenarios. Optionally, in this embodiment, the above-mentioned method for optimizing the user experience of an online and offline voice module can be applied to Figure 1 In the hardware environment shown in FIG. 1 , which is composed of a terminal device 102 and a server 104. Figure 1As shown, the server 104 is connected to the terminal device 102 via a network, and can be used to provide services (such as application services, etc.) for the terminal or a client installed on the terminal. A database can be set up on the server or independently of the server to provide data storage services for the server 104. Cloud computing and / or edge computing services can be configured on the server or independently of the server to provide data computing services for the server 104.

[0030] The aforementioned network may include, but is not limited to, at least one of the following: a wired network and a wireless network. The aforementioned wired network may include, but is not limited to, at least one of the following: a wide area network, a metropolitan area network, and a local area network. The aforementioned wireless network may include, but is not limited to, at least one of the following: Wi-Fi (Wireless Fidelity) and Bluetooth. The terminal device 102 may include, but is not limited to, a PC, a mobile phone, a tablet computer, a smart air conditioner, a smart range hood, a smart refrigerator, a smart oven, a smart stove, a smart washing machine, a smart water heater, a smart washing machine, a smart dishwasher, a smart projector, a smart TV, a smart clothes drying rack, smart curtains, a smart audio / video system, a smart socket, a smart speaker, a smart fresh air system, smart kitchen and bathroom equipment, smart bathroom equipment, a smart robot vacuum, a smart window cleaning robot, a smart robot mop, a smart air purifier, a smart steamer, a smart microwave oven, a smart kitchen appliance, a smart purifier, a smart water dispenser, a smart door lock, and the like.

[0031] Figure 2 This is one of the flowcharts of the method for upgrading the intelligent voice network device provided in an embodiment of the present application, which includes the following steps: S210: Receive a voice command input by a user.

[0032] Specifically, the system continuously monitors ambient sound through a built-in microphone or an external voice input device, preparing to receive user voice commands. When a user issues a voice command, the system captures and records the sound signal in real time. Then, using advanced speech recognition technology, the system converts the sound signal into recognizable text for subsequent processing and analysis. To ensure accurate recognition, the system also verifies and corrects the converted text, ultimately generating accurate user commands.

[0033] Receiving user voice commands greatly improves the convenience of interaction. Users can easily operate the system simply by speaking their needs, without having to use a traditional keyboard or touchscreen. This is particularly important in certain scenarios, such as driving and cooking. Receiving voice commands allows the system to more naturally understand user intent, thereby providing more personalized and intelligent services.

[0034] S220: Recognize and analyze the voice command to obtain a voice command classification.

[0035] Specifically, the system first receives voice commands from the user and then accurately identifies them using advanced speech recognition technology. After recognition, the system further analyzes the recognized text to understand the user's true intent and needs. During the analysis process, the system classifies voice commands based on pre-set rules and algorithms, for example, into different types of commands, such as query, control, and entertainment. This classification provides an important basis for subsequent interactive processing, ensuring that the system can respond appropriately and efficiently to different types of commands.

[0036] Through the above steps, the user's voice commands can be accurately identified and parsed, and the voice commands can be classified. The implementation of this link provides users with a more convenient and natural way of interaction, eliminating the need for tedious manual operations to express their needs. At the same time, the classification and processing of voice commands enables the system to more intelligently understand the user's intentions, thereby providing users with more accurate and personalized services. This not only improves the efficiency and accuracy of voice interaction, but also significantly enhances the user experience, allowing users to experience greater convenience and comfort while enjoying intelligent services.

[0037] According to a method for optimizing user experience of an offline and online voice module provided in the present application, the identification and parsing of voice commands specifically include: preprocessing the received voice commands; extracting feature parameters from the preprocessed voice signal; and matching the extracted feature parameters with a predefined voice command model to identify and parse the user's voice commands.

[0038] Specifically, the received original voice commands are preprocessed, a process that includes noise reduction and standardization to improve the clarity and quality of the voice signal. Next, the system extracts feature parameters from the preprocessed voice signal. These parameters involve the frequency, amplitude, duration, etc. of the sound, which are the basis for subsequent recognition. Finally, the system matches the extracted feature parameters with predefined voice command models. These predefined models are trained through machine learning algorithms and can accurately recognize and parse the user's voice commands, thereby classifying them into the corresponding voice command categories.

[0039] By accurately identifying and parsing users' voice commands, the system can more accurately understand their needs and provide more targeted services. This intelligent recognition not only improves the system's response speed and accuracy, but also significantly enhances the user experience. Users no longer need to worry about their voice commands being misunderstood or ignored; the system can respond quickly and accurately to meet their needs. Furthermore, by extracting feature parameters and matching them with predefined models, the system can gradually learn and adapt to users' voice habits, further improving recognition accuracy and efficiency. This personalized and intelligent service approach allows users to experience a more natural and smooth experience when interacting with the system.

[0040] S230: According to the parsed voice command classification, corresponding interactive transition words are selected from the local corpus and played.

[0041] Specifically, when a user issues a voice command and the system recognizes and analyzes it, a specific voice command classification is obtained. Based on this classification, the system further retrieves corresponding interaction transition words from the local corpus. These interaction transition words are pre-stored in the local corpus, and each instruction classification corresponds to a specific set of transition words, which are used to provide smooth and natural voice transitions during instruction execution. Once the appropriate interaction transition words are selected, the system will play them through the voice playback function, providing instant voice feedback to the user.

[0042] By selecting and playing corresponding interaction transition words from a local corpus based on the parsed voice command classification, the system can provide users with a more natural and coherent voice interaction experience. This immediate voice feedback not only allows users to feel the system's response but also guides them to the next step, thereby improving interaction efficiency and user satisfaction. Furthermore, because interaction transition words are stored in the local corpus, the system can still provide smooth voice interaction even in poor network conditions or when no network connection is available, greatly enhancing system usability and reliability.

[0043] According to a method for optimizing user experience of an offline and online voice module provided by the present application, the method selects and plays corresponding interaction transition words from a local corpus based on the parsed voice command classification, specifically including: determining the corresponding interaction scenario based on the voice command classification; retrieving interaction transition words matching the interaction scenario from the local corpus; and playing the interaction transition words.

[0044] Specifically, the system will first accurately identify and parse the received voice commands. Subsequently, based on the parsed voice command classification, it will determine the corresponding interaction scenario, such as querying the weather, playing music, or setting reminders. Next, the system will retrieve interaction transition words that match this specific interaction scenario from the local corpus. These transition words are designed to make the voice interaction process more natural and smooth. Finally, the system will play the selected interaction transition words to provide users with clear and friendly voice feedback.

[0045] By implementing the above method, the offline and online voice module can provide users with a more optimized and personalized interactive experience. Specifically, intelligently selecting and playing interactive transition words based on voice command classification not only improves the accuracy and fluency of voice interaction, but also allows users to experience a more natural and humane service when communicating with the system. In addition, the use of a local corpus reduces dependence on external networks, improves response speed, and further enhances user satisfaction. Overall, this method significantly optimizes the user experience of the offline and online voice module, making it more in line with the needs of modern intelligent interaction.

[0046] According to a method for optimizing user experience of an offline and online voice module provided by the present application, the categories of the voice command classification include control, query, entertainment and setting categories; according to the voice command classification, the corresponding interaction scenarios are determined, specifically including: determining the corresponding device operation scenario according to the control-class instructions, the interaction transition words of the device operation scenario include an operation confirmation phrase; determining the corresponding information retrieval scenario according to the query-class instructions, the interaction transition words of the information retrieval scenario include a retrieval waiting prompt; determining the corresponding media playback scenario according to the entertainment-class instructions, the interaction transition words of the media playback scenario include a playback preparation prompt; determining the corresponding parameter adjustment scenario according to the setting-class instructions, the interaction transition words of the parameter adjustment scenario include an adjustment confirmation phrase.

[0047] Specifically, the system first recognizes and parses user voice commands, classifying them into four predefined categories: control, query, entertainment, and settings. Each category corresponds to a specific interaction scenario and is associated with pre-set transition words from the local corpus to ensure natural and timely feedback.

[0048] For example, when a user issues a control command (such as "Turn on the air conditioner"), the system categorizes it as a device operation scenario and selects an action confirmation phrase from the local corpus (such as "Okay, I'll do it for you right away") as an interaction transition word. This phrase plays before the command is executed, reducing the user's perceived waiting time. For query commands (such as "What's the weather like today"), the system categorizes them as information retrieval scenarios and plays a search waiting prompt (such as "Querying, please wait") to alleviate the awkwardness caused by network latency. For entertainment commands (such as "Play Jay Chou's songs"), the system matches the media playback scenario and provides a playback preparation prompt (such as "Soon to play") to enhance the fun of the interaction. For setting commands (such as "Adjust the volume to 50%), the system matches the parameter adjustment scenario and uses an adjustment confirmation phrase (such as "Adjustment completed") as a transition word to ensure the user clearly understands the result of the action.

[0049] Through this precise matching of classification and scenarios, the system can dynamically select the most appropriate interaction transition words under different command types, which not only shortens the user's perceived response time, but also improves the naturalness and satisfaction of the interaction.

[0050] According to a method for optimizing user experience of an offline and online voice module provided in the present application, after playing the interactive transition words, the progress of voice command execution is monitored; based on the progress of voice command execution, the played interactive transition words are dynamically adjusted or the playback content is updated.

[0051] Specifically, after playing the interactive transition words, the system will enter the stage of monitoring the progress of voice command execution. By monitoring the execution of voice commands in real time, the system can accurately grasp the processing progress of the commands, such as the retrieval speed of query information, the buffer status of music playback, etc. Based on these real-time execution progress data, the system will dynamically adjust the interactive transition words played. For example, when the command execution is faster, choose a short transition word, or play transition content to soothe the user's waiting when the execution is delayed. In addition, the system can also update the playback content in a timely manner according to changes in the execution progress to ensure that the user always receives feedback that matches the current interaction status.

[0052] For example, if a user issues a simple command (such as "increase the brightness") and the system detects that the underlying device responds quickly (for example, within 200ms), the system selects a short transition word (such as "OK" or "adjusted") from the local corpus to quickly confirm the action, avoiding redundant feedback and giving the user an immediate response, making the interaction process concise and efficient.

[0053] For example, if a user requests a complex action (such as "Find nearby restaurants") and the system detects network latency or prolonged data processing (e.g., more than 1.5 seconds), it will first play a basic prompt (e.g., "Looking for restaurants for you..."). If the wait time increases, the system will add reassuring content (e.g., "There are many restaurants in the list, we'll be there soon...") to alleviate the user's anxiety. After the action is complete, the system will announce the complete result (e.g., "We found five high-rated restaurants, recommending the first one..."). This phased feedback ensures that the system is processing the request, reducing the awkwardness of waiting.

[0054] For example, if a user command is executed in multiple steps (such as "Download and install the update"), the system monitors the progress of each stage (e.g., 30% downloaded, installing in progress). Dynamically adjust the interaction transition words at different stages: In stage 1 (download started): "Downloading the update package, please wait..." In stage 2 (50% downloaded): "Half downloaded, will be completed soon..." In stage 3 (installing): "Installing the update, will be completed soon..." By dynamically adjusting the interaction transition words, users can clearly understand the progress of the task and avoid misjudging system failures due to prolonged silence.

[0055] By monitoring the progress of voice command execution in real time and dynamically adjusting the interactive transition words played or updating the playback content based on the progress, the method provided by this application significantly improves the user experience of the offline and online voice modules. This dynamic adjustment mechanism makes the voice interaction process more flexible and intelligent, and can provide users with timely and accurate feedback based on actual conditions. It reduces the user's anxiety while waiting for the command to be executed and increases the transparency and controllability of the interaction. At the same time, the real-time update of the playback content based on the execution progress also ensures the freshness and accuracy of user information, further improving user trust and satisfaction with the voice interaction system.

[0056] S240: Execute the parsed voice command.

[0057] Specifically, based on the specific instruction content identified and parsed in the previous steps, the system will call the corresponding function or service to complete the user's request. For example, if the user's instruction is "turn on the lights," the system will control the lighting system in the smart home device and turn it on. If the user's instruction is "check today's weather," the system will go online to obtain the latest weather forecast information and display it to the user. Whether controlling hardware devices or providing information services, the system will ensure the accurate execution of instructions to meet the user's needs.

[0058] By accurately executing user voice commands, the system provides convenient and efficient services, significantly enhancing the user experience. Users no longer need to navigate complex operations or interfaces to complete tasks; simply state their needs, and the system automatically completes them. This intelligent service approach not only saves users time and energy, but also brings technology closer to our daily lives, improving convenience and comfort.

[0059] S250: After the instruction is executed, the execution result is fed back to the user.

[0060] Specifically, when the system completes the execution of the user's voice command, it will immediately enter the result feedback stage. First, the system will organize the results of the command execution to ensure the accuracy and completeness of the information. Next, these execution results will be converted into a format that is easy for the user to understand, such as text descriptions, data lists, or graphical displays. Subsequently, the system will clearly convey the execution results to the user through speech synthesis or text display. In the case of voice feedback, the system will ensure the clarity and speed of the voice so that the user can easily understand it. If it is a text display, the typesetting and layout will be optimized to improve the readability of the results.

[0061] Providing timely feedback to users after a command is executed enhances transparency, ensuring users clearly understand whether the command was executed correctly and the results, thereby increasing their trust in the system. Timely feedback reduces user wait time and anxiety, improving interaction efficiency. Furthermore, clear and accurate feedback allows users to make further decisions or actions more quickly, facilitating a smoother interaction process.

[0062] According to a method for optimizing user experience of an offline and online voice module provided in the present application, the method of feeding back the execution result to the user specifically includes: generating voice feedback content corresponding to the instruction execution result; converting the voice feedback content into a voice signal through voice synthesis technology; and playing the voice signal.

[0063] Specifically, the system generates voice feedback content corresponding to the results of the command execution. This means that based on the specific execution status and results of the user's voice command, the system will organize a descriptive text to accurately reflect the status of the command execution or the information obtained. Subsequently, through advanced speech synthesis technology, the system converts this text into a natural and fluent voice signal. This step ensures that the voice expression of the feedback content is clear, easy to understand, and highly humane. Finally, the system plays this voice signal through a speaker or other audio output device, thereby intuitively conveying the execution result of the command to the user.

[0064] By providing execution results to users through the aforementioned method, the introduction of voice feedback allows users to accurately and conveniently obtain command execution results without viewing the device screen, which is particularly important in certain usage scenarios (such as driving). Secondly, the application of speech synthesis technology ensures the clarity and comprehensibility of feedback content, improving the communication efficiency between users and the voice interaction system. Furthermore, this voice feedback method enhances the user experience and makes users feel a more natural and friendly interaction atmosphere. Overall, this method not only optimizes the process and efficiency of voice interaction, but also significantly improves user satisfaction.

[0065] According to a method for optimizing the user experience of an offline and online voice module provided in this application, after the execution results are fed back to the user, the user's subsequent instructions are predicted based on the user's historical interaction data and the current interaction scenario; based on the prediction results, resources or data related to the subsequent instructions are preloaded or cached to improve the response speed of the subsequent instructions.

[0066] Specifically, after feeding back the execution results to the user, the system will further use the user's historical interaction data and the current interaction scenario to predict the user's subsequent instructions. By deeply analyzing and learning the user's interaction habits and patterns, the system can intelligently infer the subsequent operation instructions that the user may issue in a specific scenario. Based on these prediction results, the system will proactively preload or cache resources or data related to the predicted instructions, such as music files, weather information, or navigation maps. In this way, when the user actually issues a subsequent instruction, the system can quickly obtain the required resources from the preloaded or cached data, thereby greatly improving the response speed and processing efficiency of the instructions.

[0067] By predicting the user's subsequent instructions and preloading or caching related resources in advance, the method provided by this application significantly optimizes the user experience of the offline and online voice module. First, the prediction mechanism enables the system to adapt to user needs more intelligently and provide more personalized services. Secondly, the preloading and caching strategy effectively reduces user waiting time and improves the response speed of subsequent instructions, allowing users to experience a smoother and more efficient interactive experience. In addition, this method also reduces the risk of interaction interruption caused by network delays or data transmission, further enhancing the stability and reliability of the system. Overall, this method not only improves the intelligence and efficiency of voice interaction, but also brings users a more convenient and comfortable user experience.

[0068] like Figure 3 The figure shows the implementation flow chart using online voice interaction of air conditioner as an example.

[0069] 1. The user issues the command "turn on the air conditioner" through the voice module.

[0070] 2. The voice module recognizes the voice commands issued by the user.

[0071] 3. The system analyzes the recognized voice commands and clarifies the user's intent.

[0072] 4. The system plays the preset audio "OK, please wait" to inform the user that the command is being processed.

[0073] 5. Send a control command to the air conditioner baseboard to turn on the air conditioner.

[0074] 6. Receive and confirm the result of the bottom board executing the command.

[0075] 7. The system parses the results returned by the backplane and generates a TTS (text-to-speech) request based on the results and sends it to the AI ​​cloud.

[0076] 8. After receiving the TTS result, confirm that the TTS information has been successfully generated.

[0077] 9. The player plays the TTS information and provides feedback to the user on the operation results, such as "The air conditioner is turned on."

[0078] The entire control process is now completed.

[0079] The flowchart clearly shows the complete process from the user issuing a voice command to the system feedback results, reflecting the intelligence and efficiency of the voice-controlled air-conditioning system.

[0080] Compared to existing solutions, after the user speaks the command in step 1, they must wait until step 9 is completed before receiving a response. This can prolong the entire process and lead to a poor user experience in poor network conditions or when the command execution takes a long time. This solution, after parsing the command (step 3), plays the pre-built corpus in the voice module based on the command's classification. After the entire process is complete, the execution result is announced, significantly improving both the interactive response time and the user experience.

[0081] Through the above solution, the intelligence and agility of the intelligent voice module can be increased at a lower cost, the interactive fun is increased, the interaction time is shortened, and the user experience is improved; as the user's interaction time and flexibility with voice intelligent products increase, adding transitional voice can not only shorten the user's interaction waiting time, but also make the interaction more natural and appropriate, improve the user experience, and will not cause too much burden on the resource usage of the voice module itself.

[0082] The following describes the user experience optimization of the offline and online voice modules provided in this application. The user experience optimization of the offline and online voice modules described below and the user experience optimization method of the offline and online voice modules described above can be referenced to each other.

[0083] Figure 4This is a schematic diagram of the user experience optimization structure of the online and offline voice module provided by an embodiment of the present invention, which includes: The voice receiving module 410 is used to receive voice commands input by the user; The voice analysis module 420 is used to recognize and analyze voice commands to obtain voice command classification; The transition word playing module 430 is used to select and play corresponding interactive transition words from the local corpus according to the parsed voice instruction classification; An execution module 440 is used to execute the parsed voice command; The feedback module 450 is used to feed back the execution result to the user after the instruction is executed.

[0084] According to an offline and online voice module user experience optimization device provided by the present application, the voice analysis module 420 includes: a voice preprocessing unit for preprocessing received voice commands; a feature extraction unit for extracting feature parameters from the preprocessed voice signal; and a pattern matching unit for matching the extracted feature parameters with a predefined voice command model to identify and parse the user's voice commands.

[0085] Figure 5 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 5 As shown, the electronic device may include: a processor 510, a communications interface 520, a memory 530, and a communications bus 540, wherein the processor 510, the communications interface 520, and the memory 530 communicate with each other via the communications bus 540. The processor 510 may call the logic instructions in the memory 530 to execute the offline voice module user experience optimization method, which includes: receiving a voice command input by a user; recognizing and parsing the voice command to obtain a voice command classification; selecting and playing corresponding interactive transition words from a local corpus based on the parsed voice command classification; executing the parsed voice command; and after the command execution is completed, providing feedback to the user on the execution result.

[0086] In addition, the logical instructions in the above-mentioned memory 530 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present application, or the part that contributes to the existing technology, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program code.

[0087] On the other hand, the present application also provides a computer program product, which includes a computer program, which can be stored on a computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the offline and online voice module user experience optimization method provided by the above methods, which includes: receiving voice instructions input by the user; recognizing and parsing the voice instructions to obtain a voice instruction classification; selecting and playing corresponding interactive transition words from a local corpus based on the parsed voice instruction classification; executing the parsed voice instructions; and after the instruction is executed, feeding back the execution result to the user.

[0088] On the other hand, the present application also provides a computer-readable storage medium, which includes a stored program, wherein when the program is run, the offline and online voice module user experience optimization method provided by the above methods is executed, and the method includes: receiving voice instructions input by the user; recognizing and parsing the voice instructions to obtain voice instruction classification; according to the parsed voice instruction classification, selecting and playing corresponding interactive transition words from the local corpus; executing the parsed voice instructions; after the instruction is executed, feeding back the execution result to the user.

[0089] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0090] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.

[0091] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for optimizing user experience of an offline and online voice module, characterized in that: include: Receive voice commands input by the user; Recognize and analyze voice commands to obtain voice command classification; According to the classification of the parsed voice commands, the corresponding interactive transition words are selected and played from the local corpus; Execute the parsed voice command; After the instruction is executed, the execution result is fed back to the user.

2. The method for optimizing user experience of an online and offline voice module according to claim 1, wherein: The recognition and analysis of voice commands specifically includes: Preprocessing the received voice commands; Extracting feature parameters from the preprocessed speech signal; The extracted feature parameters are matched with the predefined voice command model to recognize and parse the user's voice commands.

3. The method for optimizing user experience of an online and offline voice module according to claim 1, wherein: The process of selecting and playing corresponding interactive transition words from the local corpus according to the parsed voice command classification specifically includes: Determine the corresponding interaction scenario based on the voice command classification; Retrieving interaction transition words matching the interaction scenario from a local corpus; Play the interaction transition words.

4. The method for optimizing user experience of an online and offline voice module according to claim 3, wherein: The voice command classification categories include control, query, entertainment and setting; According to the voice command classification, the corresponding interaction scenario is determined, including: Determine a corresponding device operation scenario according to the control instruction, wherein the interactive transition words of the device operation scenario include an operation confirmation phrase; Determining a corresponding information retrieval scenario according to a query instruction, wherein the interactive transition words of the information retrieval scenario include a retrieval waiting prompt; Determining a corresponding media playback scenario based on the entertainment instruction, wherein the interactive transition words of the media playback scenario include a playback preparation prompt; A corresponding parameter adjustment scenario is determined according to the setting instruction, and the interactive transition words of the parameter adjustment scenario include an adjustment confirmation phrase.

5. The method for optimizing user experience of an online and offline voice module according to claim 1, wherein: After playing the interactive transition words, the method further includes: Monitor the progress of voice command execution; Dynamically adjust the interactive transition words played or update the playback content according to the execution progress of the voice command.

6. The method for optimizing user experience of an online and offline voice module according to claim 1, wherein: Feedback of the execution result to the user specifically includes: Generate voice feedback content corresponding to the command execution result; Converting the speech feedback content into a speech signal through speech synthesis technology; Play the voice signal.

7. The method for optimizing user experience of an offline and online voice module according to any one of claims 1 to 6, wherein: After feeding back the execution result to the user, the method further includes: Predict the user's subsequent instructions based on historical user interaction data and current interaction scenarios; Based on the prediction results, resources or data related to subsequent instructions are preloaded or cached to improve the response speed of subsequent instructions.

8. An optimization of user experience of online and offline voice modules, characterized in that: include: A voice receiving module, used to receive voice commands input by the user; The voice analysis module is used to recognize and analyze voice commands and obtain voice command classification; The transition word playback module is used to select and play corresponding interactive transition words from the local corpus according to the parsed voice command classification; An execution module, used to execute the parsed voice instructions; The feedback module is used to feedback the execution results to the user after the instruction is executed.

9. The device for optimizing user experience of an offline and online voice module according to claim 8, wherein: The speech analysis module includes: A voice preprocessing unit, used to preprocess received voice commands; A feature extraction unit, configured to extract feature parameters from the preprocessed speech signal; The pattern matching unit is used to match the extracted feature parameters with a predefined voice command model to recognize and parse the user's voice command.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein the program executes the method according to any one of claims 1 to 7 when executed.

11. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to execute the method according to any one of claims 1 to 7 through the computer program.

Citation Information

Patent Citations

  • Voice interaction method and related device

    CN117334182A

  • Multi-mode interactive intelligent control system

    CN118226967A

  • Voice instruction processing method and device and electronic equipment

    CN118782032A

  • Intelligent prompt word generation method and device for vehicle-mounted voice assistant

    CN119107942A

  • Voice instruction processing method, apparatus and system, and storage medium

    WO2024002298A1

Cited By

  • Man-machine conversation implementation method and device, electronic equipment and computer storage medium

    CN121256005A

  • Man-machine conversation implementation method and device, electronic equipment and computer storage medium

    CN121256005B