Real-time voice conversation method and system based on LLM technology and applied to operation and maintenance platform
By applying real-time voice dialogue method based on LLM technology on the operation and maintenance platform, the problems of low accuracy of speech recognition and large response delay in the existing technology are solved, efficient and intelligent voice interaction and task execution are achieved, and user experience and system performance are improved.
Patent Information
- Application Number
- CN202510235575.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-05-27
AI Technical Summary
The voice control system of the existing operation and maintenance platform has problems such as low speech recognition accuracy, large response delay, lack of intelligent understanding capabilities and poor system scalability, and it is difficult to meet the needs of complex instructions and real-time interaction.
The real-time voice dialogue method based on LLM technology is adopted, and semantic analysis is performed through the LLMAgent module to identify user intentions and generate task execution plans. Combined with the interaction control of the Unity platform, real-time voice interaction and task execution are realized, and the system performance is continuously improved through feedback optimization mechanisms.
It realizes efficient real-time voice interaction, with a response delay of less than 100ms, improves speech recognition accuracy and intent understanding accuracy, supports intelligent analysis and execution of complex instructions, and improves the sustainability of human-computer interaction experience and system performance.
Smart Images

Figure CN120048276A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a real-time voice dialogue method and system based on LLM technology applied to an operation and maintenance platform. Background Art
[0002] With the rapid development of artificial intelligence technology, large language models have made breakthrough progress in the fields of natural language processing and human-computer interaction. At present, the implementation of intelligent voice interaction control in the operation and maintenance platform faces the following technical challenges:
[0003] 1. Limitations of voice control on existing operation and maintenance platforms:
[0004] The voice control system of the traditional operation and maintenance platform mainly relies on simple command word matching and rule parsing, which has the following problems:
[0005] The accuracy of speech recognition is low, especially in complex environments, it is easily disturbed, the interaction mode is single, only supports preset fixed commands, lacks flexibility, cannot understand the context, and has difficulty handling continuous conversations and complex commands.
[0006] 2. Technical defects of traditional voice control systems. The existing technical solutions mainly have the following deficiencies:
[0007] It lacks intelligent understanding capabilities and is unable to accurately analyze the user's true intentions. The task execution process lacks dynamic adjustment and optimization mechanisms. The feedback mechanism is simple and cannot provide a personalized interactive experience. The system has poor scalability and is difficult to adapt to new functions and scenario requirements.
[0008] 3. Application status of large language models in interactive control. Currently, the application of LLM technology in the field of interactive control is still in its infancy:
[0009] Most existing solutions are general-purpose dialogue systems that lack professional optimization for operation and maintenance scenarios, are difficult to meet real-time requirements, have high response delays, consume large computing resources, have high deployment costs, and their security and reliability need to be improved. Summary of the invention
[0010] The technical problem to be solved by the present invention is to provide a real-time voice dialogue method and system based on LLM technology applied to an operation and maintenance platform, which is used to solve the technical problems of poor voice interaction effect, large response delay, low accuracy and so on in the existing Unity platform.
[0011] In order to solve the above technical problems, the technical solution of the present invention is as follows:
[0012] In a first aspect, a real-time voice dialogue method based on LLM technology applied to an operation and maintenance platform, the method comprising:
[0013] Step 1: Receive the user's voice input and perform noise reduction preprocessing to convert the voice signal into text data;
[0014] Step 2: Perform semantic analysis on the text data through the LLMAgent module, identify the intent of the user's instruction, and generate a task execution plan based on the intent;
[0015] Step 3: Decompose the task execution plan into a sequence of basic operations executable on the Unity platform and determine the priority of task execution;
[0016] Step 4: Execute the sequence of basic operations in the Unity environment, synchronize the operation status in real time, and feedback the execution result;
[0017] Step 5: Collect user interaction data and system performance data to optimize the LLM model parameters and decision-making strategies.
[0018] Furthermore, the specific steps for identifying the intent of the user's instruction in Step 2 include:
[0019] Perform a preset intent classification on the input text. If a preset intent is matched, directly trigger the corresponding response action;
[0020] If no preset intent is matched, then perform multi-step semantic reasoning based on the LLM through a deep thinking system to generate an execution plan for complex tasks.
[0021] Furthermore, converting the voice signal into text data in Step 1 includes:
[0022] Real-time monitor the voice input signal through voice activity detection, and dynamically adjust the energy threshold to distinguish valid voice from background noise;
[0023] When valid voice is detected, adopt a parallel processing architecture to synchronously execute the speech recognition and speech synthesis tasks, and achieve low-latency parallel processing of voice input and output through multi-threaded scheduling;
[0024] During the speech recognition process, real-time detect preset keywords to trigger specific operations, including wake-up words, pause words, and exit words, and dynamically adjust the current task priority or switch the operation process according to the keyword detection results.
[0025] Furthermore, the response time of voice activity detection is less than 100 milliseconds. The real-time voice interruption mechanism includes:
[0026] Based on the real-time environmental noise level, dynamically optimize the energy threshold of voice activity detection. The adjustment of the energy threshold is achieved through a sliding window algorithm, and the threshold parameter is updated every 50 milliseconds to match the current acoustic environment characteristics;
[0027] When a valid voice input is detected and the response time is less than 100 milliseconds, the currently executing speech synthesis task is interrupted;
[0028] The speech synthesis thread resources are released by the priority scheduler and allocated to the speech recognition thread for real-time processing of the new voice input;
[0029] If the speech synthesis task is a critical operation, the interruption is delayed until the critical operation is completed, and the newly input voice is temporarily stored in the cache queue;
[0030] After the task switch is triggered, the semantic state, task execution progress, and user instruction history of the current conversation are stored in the temporary memory stack;
[0031] After the new speech recognition task is completed, the saved context state is loaded. If the context saving fails, the user instruction is re-parsed by the intent understanding unit of the LLMAgent module to resume the conversation flow.
[0032] Further, determining the priority of task execution in step 3 includes:
[0033] Dynamically adjust the execution order according to the task urgency;
[0034] The task status is tracked in real time by the execution monitoring subunit;
[0035] Optimize the task allocation strategy based on the resource occupancy rate.
[0036] Further, reducing the interaction latency through a multi-threaded parallel processing architecture includes:
[0037] Implement streaming processing of voice data using a circular buffer;
[0038] Dynamically allocate computing resources to balance the load of speech recognition and synthesis;
[0039] Optimize the task switching efficiency through the priority scheduler.
[0040] In a second aspect, a real-time voice dialogue system based on LLM technology applied to an operation and maintenance platform, which is used to execute the method described above, includes:
[0041] A speech recognition module for converting user voice input into text data, and the speech recognition module includes a speech preprocessing unit, a feature extraction unit, and a speech-to-text unit;
[0042] The LLMAgent module includes an intent understanding unit, a task planning unit, and a decision execution unit, and the intent understanding unit adopts a dual-system architecture of fast response and in-depth thinking;
[0043] The Unity interaction control module is used to execute control instructions and perform status synchronization;
[0044] A feedback optimization module for collecting user feedback and performing performance optimization;
[0045] A real-time duplex communication module for supporting parallel processing of voice input and output, including a voice interruption mechanism, a parallel speech recognition and synthesis mechanism, and a keyword trigger mechanism;
[0046] Among them, the LLMAgent module receives the text data, decomposes tasks through intention understanding, and generates corresponding Unity control instructions.
[0047] In a third aspect, a computing device includes:
[0048] One or more processors;
[0049] A storage device for storing one or more programs, which when executed by the one or more processors cause the one or more processors to implement the above method.
[0050] In a fourth aspect, a computer-readable storage medium stores a program that implements the above method when executed by a processor.
[0051] The above solution of the present invention has at least the following beneficial effects:
[0052] Efficient real-time voice interaction is achieved, with a response delay of less than 100 ms. The speech recognition accuracy and intention understanding accuracy are improved through LLM technology. A modular design is adopted, and each functional unit is decoupled, facilitating maintenance and expansion. A feedback optimization mechanism is introduced, and the system performance can be continuously improved. It supports intelligent parsing and execution of complex instructions, enhancing the human-computer interaction experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 It is a schematic diagram of the overall system architecture of the present invention;
[0054] Figure 2 It is a schematic diagram of the structure of the real-time duplex communication module of the present invention;
[0055] Figure 3 It is a flowchart of the dual-system architecture operation of the LLMAgent module of the present invention;
[0056] Figure 4 It is a timing diagram of the voice interruption mechanism of the present invention;
[0057] Figure 5 It is a data flow diagram of the parallel voice processing of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0058] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.
[0059] As Figure 1 shown, an embodiment of the present invention provides a real-time voice dialogue method based on LLM technology applied to an operation and maintenance platform. The method includes:
[0060] Step 1, receiving user voice input and performing noise reduction preprocessing, and converting the voice signal into text data;
[0061] Step 2, performing semantic analysis on the text data through the LLMAgent module, identifying the intention of the user instruction, and generating a task execution plan based on the intention;
[0062] Step 3, decomposing the task execution plan into a basic operation sequence executable by the Unity platform, and determining the priority of task execution;
[0063] Step 4, executing the basic operation sequence in the Unity environment, synchronizing the operation status in real time and feeding back the execution result;
[0064] Step 5, collecting user interaction data and system performance data, and optimizing the LLM model parameters and decision-making strategies.
[0065] In the embodiments of the present invention, through real-time voice conversations, users can interact with the system more intuitively without the need for cumbersome keyboard input or click operations. This greatly improves the interaction efficiency between the user and the system, making the operation more convenient. The voice interaction method is more in line with the habits of natural human communication and can provide a more friendly and user-friendly operation experience for users. Users can communicate with the system just like having a conversation with a person, reducing the usage difficulty and learning cost. Through the semantic analysis of the LLMAgent module, the system can accurately identify the intention of the user's instructions and generate corresponding task execution plans, which enables the system to flexibly handle various complex instructions and improves the intelligence and accuracy of operation and maintenance operations. By decomposing the task execution plan into a basic operation sequence executable on the Unity platform and determining the priority of task execution, the system can efficiently and orderly complete tasks, which avoids conflicts and chaos between tasks and improves the overall task execution efficiency. By collecting user interaction data and system performance data, the system can continuously optimize the LLM model parameters and decision-making strategies. This continuous optimization ability enables the system to gradually adapt to the needs and operation habits of different users and improve the personalization and accuracy of services. Automated voice interaction and task execution reduce the need for manual intervention, thereby reducing the operation and maintenance costs. At the same time, by providing real-time feedback on the execution results, the system can promptly detect and handle problems, reducing the time cost of troubleshooting and repair.
[0066] In a preferred embodiment of the present invention, the specific steps for identifying the intention of the user's instructions in step 2 include:
[0067] Perform a preset intention classification on the input text. If a preset intention is matched, directly trigger the corresponding response action;
[0068] If no preset intention is matched, then through a deep thinking system, perform multi-step semantic reasoning based on the LLM to generate an execution plan for complex tasks.
[0069] In the embodiments of the present invention, by performing preset intent classification on the input text, the system can quickly identify and respond to those common or standard user instructions. This fast matching mechanism significantly reduces the processing time and improves the system's instant response ability, thereby providing a smoother experience for users. Preset intent classification is usually constructed based on a large amount of prior knowledge and historical data, so it has high accuracy and reliability. When the user instruction matches the preset intent, the system can confidently execute the corresponding response actions, reducing misoperations and unnecessary confirmation steps. For the input text that does not match the preset intent, the system performs multi-step semantic reasoning through a deep thinking system and an LLM. This reasoning ability enables the system to understand and parse more complex and non-standard user instructions and generate corresponding execution plans, which provides greater flexibility for users and supports more personalized and innovative operation requirements. The multi-step semantic reasoning based on the LLM is not limited to the understanding of the current instruction, but can also make comprehensive judgments by combining context information. This ability enhances the intelligence level of the system, enabling it to continuously learn and evolve during the conversation process and better adapt to the changes and needs of users. By comprehensively applying preset intent classification and semantic reasoning, the system can more comprehensively meet the different needs of users, whether the instructions are simple or complex. This comprehensive and efficient processing method significantly improves the user experience and satisfaction and enhances the user's trust and dependence on the system.
[0070] In a preferred embodiment of the present invention, in step 1, converting the voice signal into text data includes:
[0071] Real-time monitoring of the voice input signal through voice activity detection, and dynamically adjusting the energy threshold to distinguish effective voice from background noise;
[0072] When effective voice is detected, a parallel processing architecture is used to synchronously execute the speech recognition and speech synthesis tasks, and low-latency parallel processing of voice input and output is achieved through multi-threaded scheduling;
[0073] During the speech recognition process, preset keywords are detected in real time to trigger specific operations, including wake-up words, pause words, and exit words, and the current task priority or operation process is dynamically adjusted according to the keyword detection results.
[0074] In the embodiments of the present invention, the voice activity detection is used to monitor the voice input signal in real time and dynamically adjust the energy threshold. The system can more accurately distinguish the valid voice from the background noise. This dynamic adjustment mechanism helps to reduce the noise interference, thereby improving the accuracy of speech recognition. The parallel processing architecture is adopted to synchronously execute the speech recognition and speech synthesis tasks, and the multi-threaded scheduling is used to achieve the low-latency parallel processing of voice input and output. This parallel processing method significantly improves the real-time performance of the system, enabling the user to obtain feedback faster and enhancing the fluency and immediacy of the interaction. During the speech recognition process, the preset keywords are detected in real time to trigger specific operations, such as wake-up words, pause words, and exit words. This mechanism allows the user to control the operation process of the system through simple voice commands without the need for complex menu navigation or manual operations, thus improving the convenience and flexibility of the operation. According to the keyword detection results, the current task priority is dynamically adjusted or the operation process is switched. The system can manage the tasks and execution processes more intelligently. This dynamic adjustment helps to ensure that important tasks are given priority and avoid unnecessary resource waste, thereby improving the overall task execution efficiency. Through the comprehensive application of the above technical means, step 1 provides the user with a more efficient, accurate, and convenient voice interaction experience. The user can communicate with the system more naturally, reducing the operation difficulty and learning cost, thus enhancing the user's satisfaction and acceptance of the system.
[0075] In a preferred embodiment of the present invention, the response time of the voice activity detection is less than 100 milliseconds. The real-time voice interruption mechanism includes:
[0076] Based on the real-time environmental noise level, the energy threshold of the voice activity detection is dynamically optimized. The adjustment of the energy threshold is implemented through a sliding window algorithm, and the threshold parameter is updated every 50 milliseconds to match the current acoustic environment characteristics;
[0077] When a valid voice input is detected and the response time is less than 100 milliseconds, the currently executing speech synthesis task is interrupted;
[0078] The speech synthesis thread resources are released through the priority scheduler and allocated to the speech recognition thread to enable the real-time processing of the new voice input;
[0079] If the speech synthesis task is a critical operation, the interruption is delayed until the critical operation is completed, and the newly input voice is temporarily stored in the cache queue;
[0080] When the task switch is triggered, the semantic state, task execution progress, and user instruction history of the current conversation are stored in the temporary memory stack;
[0081] After the new speech recognition task is completed, load the saved context state. If the context saving fails, re-parse the user instruction through the intent understanding unit of the LLMAgent module to resume the conversation flow.
[0082] In the embodiment of the present invention, the response time of the voice activity detection is less than 100 milliseconds, ensuring that the system can quickly capture the user's voice input and make a response. This low-latency response ability greatly improves the real-time performance and fluency of user interaction. By dynamically optimizing the energy threshold of the voice activity detection based on the real-time environmental noise level, the system can better adapt to different acoustic environments, updating the threshold parameters every 50 milliseconds, enabling the system to match the characteristics of the current environment in real time and improving the accuracy and stability of speech recognition. When a valid voice input is detected, the system can interrupt the ongoing speech synthesis task and quickly release resources through the priority scheduler and allocate them to the new speech recognition task. This flexible task scheduling mechanism ensures that the new voice input can be processed in real time and improves the concurrent processing ability of the system. If the ongoing speech synthesis task is a critical operation, the system will delay the interruption until the critical operation is completed and temporarily store the newly input voice in the cache queue. This design avoids the loss of critical information or operation errors caused by interruption and ensures the stability and reliability of the system. After triggering the task switch, the system stores the semantic state of the current conversation, the task execution progress, and the user instruction history in the temporary memory stack. After the new speech recognition task is completed, by loading the saved context state, the system can quickly resume the previous conversation flow, improving the coherence and efficiency of user interaction. Even if the context saving fails, the system can re-parse the user instruction through the intent understanding unit of the LLM Agent module to ensure the smooth recovery of the conversation flow.
[0083] In a preferred embodiment of the present invention, determining the priority of task execution in step 3 includes:
[0084] Dynamically adjust the execution order according to the task urgency;
[0085] Real-time track the task status through the execution monitoring subunit;
[0086] Optimize the task allocation strategy based on the resource occupancy rate.
[0087] In the embodiments of the present invention, by dynamically adjusting the execution order according to the urgency of tasks, it is possible to ensure that urgent and important tasks are given priority. This flexible priority adjustment mechanism helps reduce task waiting time, speed up task completion, and thus improve the overall task execution efficiency. The execution monitoring subunit can track the task status in real time and provide real-time feedback on the task execution progress to system administrators or users. This transparent monitoring mechanism helps detect and solve problems in a timely manner to ensure that tasks can proceed smoothly as expected. Optimizing the task allocation strategy based on resource occupancy can ensure more reasonable utilization of system resources. By balancing the resource requirements of different tasks, it is possible to avoid some tasks over-occupying resources and causing other tasks to be blocked, thereby improving the overall utilization efficiency of system resources. Reasonable task priority setting and task allocation strategies help reduce the risk of system overload and avoid system crashes or performance degradation caused by task conflicts or resource contention, which helps maintain the stable operation of the system and enhance the user experience. By optimizing the task execution priority, the system can respond to user needs more quickly, reduce user waiting time, and provide higher-quality services, which will directly improve user satisfaction and loyalty.
[0088] In a preferred embodiment of the present invention, the multi-thread parallel processing architecture reduces the interaction delay, including:
[0089] Implementing the streaming processing of voice data using a circular buffer;
[0090] Dynamically allocating computing resources to balance the load of speech recognition and synthesis;
[0091] Optimizing the task switching efficiency through a priority scheduler.
[0092] In the embodiments of the present invention, a circular buffer is adopted to implement the streaming processing of voice data, which can effectively manage the input and output of voice data, avoid data congestion and loss. This streaming processing method can ensure the continuous and stable transmission of data, thereby improving the data processing efficiency and reducing the interaction delay. Dynamically allocating computing resources to balance the loads of speech recognition and synthesis can ensure the efficient utilization of system resources. When the load of the speech recognition or speech synthesis task is heavy, the system can dynamically adjust the resource allocation to meet the processing requirements of different tasks, avoid resource waste or bottlenecks, and thus improve the overall performance. Optimizing the task switching efficiency through a priority scheduler can ensure that important or urgent tasks are processed first. This priority scheduling mechanism can reduce the delay during task switching, improve the response speed and flexibility of the system for multi-task processing, and further reduce the interaction delay perceived by users. Through the comprehensive application of the above technical means, the multi-threaded parallel processing architecture significantly reduces the interaction delay, enabling users to obtain smoother and more immediate feedback when interacting with the system by voice. This fast response ability improves the user experience and makes the user feel that the system is more intelligent and efficient. The multi-threaded parallel processing architecture also has good scalability. With the development of technology and the change of user requirements, the system can easily add more threads or adjust the resource allocation strategy to adapt to higher processing requirements and more complex application scenarios.
[0093] An embodiment of the present invention further provides a real-time voice dialogue system based on LLM technology applied to an operation and maintenance platform, including:
[0094] A speech recognition module for converting user voice input into text data, where the speech recognition module includes a speech preprocessing unit, a feature extraction unit, and a speech-to-text unit;
[0095] The LLMAgent module includes an intention understanding unit, a task planning unit, and a decision execution unit, where the intention understanding unit adopts a dual-system architecture of rapid response and in-depth thinking;
[0096] A Unity interaction control module for executing control instructions and synchronizing states;
[0097] A feedback optimization module for collecting user feedback and performing performance optimization;
[0098] A real-time duplex communication module for supporting the parallel processing of voice input and output, including a voice interruption mechanism, a parallel speech recognition and synthesis mechanism, and a keyword trigger mechanism;
[0099] Among them, the LLMAgent module receives the text data, decomposes tasks through intention understanding, and generates corresponding Unity control instructions.
[0100] Figure 1Schematic diagram of the overall system architecture of the present invention;
[0101] It shows five core modules of the system: Unity client, intelligent building brain, core services, knowledge services, and external services. The Unity client includes a voice processing unit and an interaction layer; the intelligent building brain includes a voice service unit and an ArchAgent central control unit; the core services include functional units such as device control, meeting room management, catering service, and attendance management; the knowledge services include a Qdrant vector database and a document retrieval unit; the external services include third-party service interfaces such as weather APIs and Wikipedia.
[0102] Figure 2 Schematic diagram of the structure of the real-time duplex communication module of the present invention;
[0103] It shows three core mechanisms of the real-time duplex communication module: voice interruption mechanism, parallel voice processing mechanism, and keyword trigger mechanism. Among them, the voice interruption mechanism includes a VAD detector, energy threshold detection, and a status manager; parallel voice processing includes a circular buffer, a priority scheduler, and a switching controller; keyword trigger includes local keyword recognition and scenario configuration management.
[0104] Figure 3 Flowchart of the dual-system architecture operation of the LLMAgent module of the present invention;
[0105] It shows the operation processes of the quick response system and the in-depth thinking system. The quick response system quickly matches preset intents through an intent classifier and triggers a response; the in-depth thinking system conducts in-depth understanding processing, multi-step reasoning analysis through an LLM, generates an execution plan, and finally integrates the execution results to generate a response.
[0106] Figure 4 Timing diagram of the voice interruption mechanism of the present invention;
[0107] It shows the complete interaction timing from user voice input to system response, including the whole process of VAD detection, STT speech recognition, status management, voice processing, LLM service processing, and TTS speech synthesis output, reflecting the real-time interaction ability of the system.
[0108] Figure 5 Data flow diagram of the parallel voice processing of the present invention;
[0109] It shows the parallel processing process of audio data, including an audio input stream, a circular buffer, speech recognition processing, speech synthesis processing, and an audio output stream, reflecting the parallel processing architecture of the system.
[0110] Solid lines in the figure represent the control flow direction, and dashed lines represent the data flow direction. Each module communicates through standard interfaces to ensure the scalability and maintainability of the system.
[0111] It should be noted that this system corresponds to the above method, and all implementation manners in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.
[0112] An embodiment of the present invention further provides a computing device, including: a processor and a memory storing a computer program. When the computer program is run by the processor, it executes the method as described above. All implementation manners in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.
[0113] An embodiment of the present invention further provides a computer-readable storage medium storing instructions. When the instructions are run on a computer, the computer is caused to execute the method as described above. All implementation manners in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.
[0114] The above is the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A real-time voice dialogue method based on LLM technology applied to an operation and maintenance platform, characterized in that: The method comprises: Step 1, receiving user voice input and performing noise reduction preprocessing to convert the voice signal into text data; Step 2, performing semantic analysis on the text data through the LLMAgent module, identifying the intention of the user instruction, and generating a task execution plan based on the intention; Step 3, decomposing the task execution plan into basic operation sequences executable by the Unity platform, and determining the priority of task execution; Step 4, executing the basic operation sequence in the Unity environment, synchronizing the operation status in real time and feeding back the execution result; Step 5: Collect user interaction data and system performance data to optimize LLM model parameters and decision strategies.
2. The real-time voice dialogue method based on LLM technology applied to the operation and maintenance platform according to claim 1 is characterized in that: The intention of identifying the user's instruction in step 2 specifically includes: Classify the input text into preset intents, and directly trigger the corresponding response action if it matches the preset intent; If the preset intention is not matched, the deep thinking system will perform multi-step semantic reasoning based on LLM to generate an execution plan for complex tasks.
3. The real-time voice dialogue method based on LLM technology applied to the operation and maintenance platform according to claim 2 is characterized in that: In step 1, the speech signal is converted into text data, including: Through voice activity detection, the voice input signal is monitored in real time, and the energy threshold is dynamically adjusted to distinguish valid speech from background noise; When valid speech is detected, a parallel processing architecture is used to synchronously execute speech recognition and speech synthesis tasks, and low-latency parallel processing of speech input and output is achieved through multi-threaded scheduling; During the speech recognition process, preset keywords are detected in real time to trigger specific operations, including wake-up words, pause words, and exit words, and the current task priority is dynamically adjusted or the operation process is switched based on the keyword detection results.
4. The real-time voice dialogue method based on LLM technology applied to the operation and maintenance platform according to claim 3 is characterized in that: Voice activity detection with a response time of less than 100 milliseconds and real-time voice interruption mechanism, including: Dynamically optimize the energy threshold of voice activity detection based on the real-time environmental noise level. The energy threshold is adjusted by a sliding window algorithm, and the threshold parameter is updated every 50 milliseconds to match the current acoustic environment characteristics. When valid voice input is detected and the response time is less than 100 milliseconds, the speech synthesis task being executed is interrupted; Release speech synthesis thread resources through the priority scheduler and allocate them to the speech recognition thread to enable real-time processing of new speech input; If the speech synthesis task is a critical operation, the interruption is delayed until the critical operation is completed, and the new input speech is temporarily stored in the cache queue; When the task switch is triggered, the semantic state of the current dialogue, the task execution progress and the user command history are stored in a temporary memory stack; After the new speech recognition task is completed, the saved context state is loaded. If the context saving fails, the user instructions are re-parsed through the intention understanding unit of the LLMAgent module to restore the dialogue process.
5. The real-time voice dialogue method based on LLM technology applied to the operation and maintenance platform according to claim 4 is characterized in that: In step 3, the priority of task execution is determined, including: Dynamically adjust the execution order according to the urgency of the task; Track task status in real time through execution monitoring subunit; Optimize task allocation strategy based on resource occupancy.
6. The real-time voice dialogue method based on LLM technology applied to the operation and maintenance platform according to claim 5 is characterized in that: Reduce interactive latency through multi-threaded parallel processing architecture, including: Use a ring buffer to implement streaming processing of voice data; Dynamically allocate computing resources to balance the load of speech recognition and synthesis; Optimize task switching efficiency through priority scheduler.
7. A real-time voice dialogue system based on LLM technology applied to the operation and maintenance platform, characterized in that: The system is used to perform the method according to any one of claims 1 to 6, comprising: A speech recognition module, used to convert user speech input into text data, the speech recognition module includes a speech preprocessing unit, a feature extraction unit and a speech-to-text unit; The LLMAgent module includes an intention understanding unit, a task planning unit, and a decision execution unit. The intention understanding unit adopts a dual-system architecture of rapid response and deep thinking. Unity interactive control module, used to execute control instructions and synchronize states; Feedback optimization module, used to collect user feedback and perform performance optimization; Real-time duplex communication module to support parallel processing of speech input and output, including speech interruption mechanism, parallel speech recognition and synthesis mechanism, and keyword triggering mechanism; The LLMAgent module receives the text data, performs task decomposition through intent understanding, and generates corresponding Unity control instructions.
8. A computing device, characterized in that include: one or more processors; A storage device for storing one or more programs, when the one or more programs are executed by the one or more processors, the one or more processors implement the method as claimed in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a program, which, when executed by a processor, implements the method according to any one of claims 1 to 6.
Citation Information
Cited By
Information interaction method, system and device, electronic equipment and storage medium
CN120315798A
Voice data interaction feedback control processing method based on large language model
CN120510846A
A speech data interactive feedback control processing method based on large language model
CN120510846B
Pet equipment voice control method and system based on intelligent switching
CN120808775A
Pet device voice control method and system based on intelligent switching
CN120808775B